Skip to main content

High Latency

Target of six alerts in /opt/tutela/deploy/monitoring/alerts/service-alerts.yml, all severity warning, all for: 5m:

AlertMeasuresThreshold
HighLatencyInspectortutela_api_request_duration_ms_bucket{job="inspector"}P95 > 200 ms
HighLatencyClassifierT1tutela_classification_latency_seconds_bucket{job="classifier", tier="T1"}P95 > 10 ms
HighLatencyPolicytutela_policy_evaluation_ms_bucket{job="policy"}P95 > 5 ms
AIGatewayGovernancePhaseLatencyHightutela_ai_gateway_phase_duration_ms_bucket, governance phases onlyP95 > 1000 ms
AIGatewayProviderPhaseLatencyHightutela_ai_gateway_phase_duration_ms_bucket{phase=~"tutela_provider_request|tutela_provider_stream_start"}P95 > 30000 ms
MCPGatewayPhaseLatencyHightutela_mcp_phase_duration_ms_bucketP95 > 5000 ms

Each is histogram_quantile(0.95, sum(rate(<bucket>[5m])) by (le, ...)).

What fired, in plain terms​

The 95th-percentile time for that operation crossed its budget for five consecutive minutes. One in twenty operations is now slower than the number in the table.

The three service alerts track the product's published performance budgets. The three phase alerts are much looser ceilings on the AI and MCP paths, sized to catch a stall rather than a slowdown.

What this does NOT mean​

  • It is not an average. P95 is the slow tail. A P95 of 250 ms is perfectly compatible with a P50 of 8 ms and a healthy service for 19 of every 20 requests.
  • AIGatewayProviderPhaseLatencyHigh is not Tutela's latency. The tutela_provider_request and tutela_provider_stream_start phases measure the upstream model provider. Above 30 s means the provider is slow or the request is enormous. There is no local fix. Those two phases are deliberately excluded from the governance alert so the two causes never blur.
  • It does not mean requests failed. Slow is not failed. If they were also failing you would additionally have high-error-rate.md.
  • A quantile over sparse data is not meaningful. histogram_quantile interpolates inside whichever bucket contains the 95th percentile. With only a few observations in a five-minute window it effectively reports that bucket's upper bound — which is why a nearly idle service can report an alarming P95.

First three checks​

1. Confirm there is enough traffic for the quantile to mean anything.

docker exec tutela-prometheus \
/bin/promtool query instant http://localhost:9090 'sum(rate(tutela_api_request_duration_ms_count{job="inspector"}[5m]))'

Below roughly 1 request/second, treat the percentile as advisory and read the raw count and P50 before acting.

2. Compare the percentile band — whole distribution, or just the tail?

for q in 0.5 0.95 0.99; do
printf 'P%s: ' "$q"
docker exec tutela-prometheus \
/bin/promtool query instant http://localhost:9090 \
"histogram_quantile($q, sum(rate(tutela_api_request_duration_ms_bucket{job=\"inspector\"}[5m])) by (le))"
done

P50 flat with P95 climbing is contention or a slow subset of requests. The whole band shifting up is saturation.

3. Split by phase to find where the time goes.

For the AI and MCP gateways the phase label localises it directly:

docker exec tutela-prometheus \
/bin/promtool query instant http://localhost:9090 'topk(5, histogram_quantile(0.95, sum(rate(tutela_ai_gateway_phase_duration_ms_bucket[5m])) by (le, phase)))'

For the inspector the split is classifier vs policy — use the Tutela - Inspector Path Triage Grafana dashboard and the queries in inspector-failclosed-triage.md. The Tutela - E2E Latency and Tutela - AI Gateway Latency dashboards (/opt/tutela/deploy/monitoring/grafana/dashboards/) carry the same breakdowns.

Real problem or artefact?​

SignalReading
Low request rate, spiky P95Artefact. Sparse-histogram interpolation.
P95 climbs, P50 flat, errors flatTail latency — contention or GC. Real but rarely urgent.
P50 and P95 both climbSaturation. Check high-cpu.md and the dependency chain.
Inspector latency up and classifier latency upThe classifier is the cause; the inspector is blocked on it. Fix one, not both.
Provider phase alert aloneUpstream provider. Confirm on the provider's status page; not locally actionable.
Latency up and tutela_inspector_semaphore_high_water_mark near 100Serious. The publish path is saturating and data loss is next — go to event-dropped.md.

Escalation​

Inspector latency is on the customer's inline path: above budget, users feel it on every prompt. Classifier T1 latency is the strongest predictor of inspector latency, because the inspector blocks on it.

Sustained latency has a second consequence beyond slowness. When the classifier stops answering inside its timeout the inspector applies its fail mode, and the deployment default is fail-open — so the visible symptom flips from "slow" to "fast and ungoverned". Watch InspectorFailOpenElevated alongside this alert; that transition matters more than the latency number.

Escalate with the percentile band from check 2, the phase breakdown from check 3, the request rate, and whether a deploy or traffic change preceded it.