AI Spend Operations Runbook
Use this runbook when AI Spend shows elevated attention items or when usage governance alerts fire. The goal is to keep spend controls accurate, keep imported usage visible, and make cost differences easy to resolve without exposing prompts, secrets, tokens, or raw provider responses.
Primary Signals
- Usage ledger ingest lag: P95 time from usage event time to ledger write.
- Usage ledger write failures: failed ledger writes by source and governance mode.
- External usage import failures: failed import batches by source system.
- Budget decision latency: P95 budget preflight response time by action.
- Budget settlement failures: failed reservation settlement attempts by invocation status.
- Cost matching variance: estimated and billed cost differences by provider and model.
- Compatible-completion persistence failures:
tutela_external_usage_persistence_failed_total{reason}. Any increase means a provider response was returned after its usage/outbox/cost transaction failed or could not be attempted.
First Response
- Open the AI Cost Telemetry & Budget Health dashboard and review allocation coverage, budget enforcement, and the Usage Governance row.
- Open AI Spend and check Usage, Cost Matching, Imports, Budgets, and AI Runs for the same time window.
- Identify whether the issue is isolated to one source system, provider, model, budget scope, or tenant.
- Preserve the current budget settings until the cause is understood. If spend risk is immediate, pause the affected integration or tighten only the affected budget scope.
- Record the alert name, time window, tenant, source, provider, and request ID when one is available.
Cost Allocation Boundary
Grafana shows aggregate, low-cardinality operational telemetry. It is useful for provider spend trends, allocation coverage, budget actions, ledger health, and reconciliation. It is not the authoritative chargeback ledger and must not expose user names, email addresses, directory-group IDs, or other identity values as Prometheus labels.
Use AI Spend for user-, group-, department-, project-, and cost-center-aware allocation. Treat unassigned as an attribution defect to investigate, not as a department. Allocation should be resolved from canonical authenticated identity and an event-time organizational snapshot; do not infer ownership from browser-supplied metadata, a current group membership that may have changed, or a Grafana series label.
Until the authoritative allocation workflow supports every organizational dimension, report unmatched usage in an explicit unallocated bucket. Do not silently redistribute it across known users or groups.
Usage Ledger Ingest Lag
High ingest lag means usage records are arriving late or writes are slowed.
- Check analytics service health, ClickHouse availability, and Kafka consumer lag.
- Compare lag by
sourceandgovernance_modeto identify whether routed usage or imported usage is affected. - If lag is isolated to imported usage, move to the External Usage Imports section.
- If lag is broad, reduce optional reporting traffic and check ClickHouse write latency before changing budget settings.
Usage Ledger Write Failures
Write failures can affect AI Spend accuracy and budget settlement history.
- Check analytics logs for wrapped ledger write errors.
- Confirm ClickHouse is reachable and usage ledger tables are present.
- Confirm every configured analytics source topic and its derived
tutela.dlq.<source-topic>recovery topic exist. The default analytics sources aretutela.inspection.events,tutela.policy.violations,tutela.discovery.events,tutela.token.usage,tutela.mcp.audit, andtutela.runtime.invocations. Analytics start-up creates any recovery topic that is missing and exits — loggingrequired Kafka DLQ topics are not ready— if one still cannot be reached, so an analytics container that is running has already confirmed its recovery topics. Start-up does not check the source topics; confirm those yourself by listing them. - Compare metadata-only
token_usageandusage_ledger_eventsevent IDs. The analytics startup reconciliation restores ledger records that were durably written totoken_usagebut are missing fromusage_ledger_events; a recurring reconciliation repeats this check every five minutes. - Treat a reconciliation error or a reconciliation that does not converge within its bounded startup window as a failed analytics deployment. Do not certify AI Spend while the ledger remains incomplete.
- Check whether failures align with a recent migration, deploy, or storage pressure event.
- Keep enforcement paths active. Budget preflight should continue to protect configured scopes even when a later write fails.
For OpenAI-, Anthropic-, and Gemini-compatible completions, one server-owned
usage_event_id follows the provider execution through budget settlement, the
PostgreSQL usage record, the delivery outbox, and cost attribution. A repeated
request_id is correlation, not idempotency, and must not collapse two billed
executions. Stateless compatible calls deliberately have no message_id.
If tutela_external_usage_persistence_failed_total increases, search AI Proxy
logs for failed to persist external completion usage and the matching bounded
reason. The response may already have reached the caller, so do not retry the
provider call blindly. Restore PostgreSQL/outbox health and reconcile using the
logged usage_event_id; never substitute the caller-controlled request ID.
Reconciliation uses metadata-only usage records and does not recover prompts, arguments, outputs, resource identifiers, credentials, or tokens. An invalid source timestamp or attribution envelope is rejected rather than replaced with the current time, preserving truthful historical evidence.
External Usage Imports
Import failures affect visibility for external agent or model activity.
- Open AI Spend Imports and filter by the source system shown in the alert.
- Confirm connector credentials, source API availability, and source-side rate limits.
- Re-run the affected import only after the source problem is resolved.
- If records are rejected, compare the import payload shape to the public usage import contract and correct missing actor, provider, model, status, time, or usage fields.
- Imported usage is visibility-only unless the record contains accepted enforcement evidence.
Budgets
Budget decision latency affects request startup. Settlement failures affect post-call accounting.
- For decision latency, check analytics service latency and ClickHouse read latency for budget assignments, usage, and active reservations.
- For settlement failures, check reservation IDs, tenant scope, and ledger write health.
- If a budget blocks unexpectedly, review the matched budget, period, active reservations, and current usage in AI Spend.
- Delete or disable a budget only when it is clearly mis-scoped. Prefer narrowing scope over removing protection.
Cost Matching
Cost matching variance means provider-reported or billed cost differs from Tutela's estimate.
- Filter Cost Matching by provider and model.
- Check whether variance is concentrated in imported records, routed records, or a specific pricing window.
- Confirm provider pricing, currency, billing period, and model name mapping.
- If provider billing is delayed, leave the item unresolved until the next billing export arrives.
- If Tutela pricing is stale, update the pricing snapshot and re-run matching for the affected window.
Escalation
Escalate when:
- Usage ledger writes fail for more than 15 minutes.
- Budget decision P95 latency remains above 250ms for more than 15 minutes.
- Settlement failures continue after ClickHouse and analytics health recover.
- A cost variance materially changes customer billing or chargeback reporting.
Include alert name, time window, tenant, source system, provider/model, request ID when present, and the latest AI Spend filters used.