High Latency
Target of six alerts in /opt/tutela/deploy/monitoring/alerts/service-alerts.yml, all
severity warning, all for: 5m:
| Alert | Measures | Threshold |
|---|---|---|
HighLatencyInspector | tutela_api_request_duration_ms_bucket{job="inspector"} | P95 > 200 ms |
HighLatencyClassifierT1 | tutela_classification_latency_seconds_bucket{job="classifier", tier="T1"} | P95 > 10 ms |
HighLatencyPolicy | tutela_policy_evaluation_ms_bucket{job="policy"} | P95 > 5 ms |
AIGatewayGovernancePhaseLatencyHigh | tutela_ai_gateway_phase_duration_ms_bucket, governance phases only | P95 > 1000 ms |
AIGatewayProviderPhaseLatencyHigh | tutela_ai_gateway_phase_duration_ms_bucket{phase=~"tutela_provider_request|tutela_provider_stream_start"} | P95 > 30000 ms |
MCPGatewayPhaseLatencyHigh | tutela_mcp_phase_duration_ms_bucket | P95 > 5000 ms |
Each is histogram_quantile(0.95, sum(rate(<bucket>[5m])) by (le, ...)).
What fired, in plain terms
The 95th-percentile time for that operation crossed its budget for five consecutive minutes. One in twenty operations is now slower than the number in the table.
The three service alerts track the product's published performance budgets. The three phase alerts are much looser ceilings on the AI and MCP paths, sized to catch a stall rather than a slowdown.
What this does NOT mean
- It is not an average. P95 is the slow tail. A P95 of 250 ms is perfectly compatible with a P50 of 8 ms and a healthy service for 19 of every 20 requests.
AIGatewayProviderPhaseLatencyHighis not Tutela's latency. Thetutela_provider_requestandtutela_provider_stream_startphases measure the upstream model provider. Above 30 s means the provider is slow or the request is enormous. There is no local fix. Those two phases are deliberately excluded from the governance alert so the two causes never blur.- It does not mean requests failed. Slow is not failed. If they were also failing you would additionally have high-error-rate.md.
- A quantile over sparse data is not meaningful.
histogram_quantileinterpolates inside whichever bucket contains the 95th percentile. With only a few observations in a five-minute window it effectively reports that bucket's upper bound — which is why a nearly idle service can report an alarming P95.
First three checks
1. Confirm there is enough traffic for the quantile to mean anything.
docker exec tutela-prometheus \
/bin/promtool query instant http://localhost:9090 'sum(rate(tutela_api_request_duration_ms_count{job="inspector"}[5m]))'
Below roughly 1 request/second, treat the percentile as advisory and read the raw count and P50 before acting.
2. Compare the percentile band — whole distribution, or just the tail?
for q in 0.5 0.95 0.99; do
printf 'P%s: ' "$q"
docker exec tutela-prometheus \
/bin/promtool query instant http://localhost:9090 \
"histogram_quantile($q, sum(rate(tutela_api_request_duration_ms_bucket{job=\"inspector\"}[5m])) by (le))"
done
P50 flat with P95 climbing is contention or a slow subset of requests. The whole band shifting up is saturation.
3. Split by phase to find where the time goes.
For the AI and MCP gateways the phase label localises it directly:
docker exec tutela-prometheus \
/bin/promtool query instant http://localhost:9090 'topk(5, histogram_quantile(0.95, sum(rate(tutela_ai_gateway_phase_duration_ms_bucket[5m])) by (le, phase)))'
For the inspector the split is classifier vs policy — use the
Tutela - Inspector Path Triage Grafana dashboard and the queries in
inspector-failclosed-triage.md. The
Tutela - E2E Latency and Tutela - AI Gateway Latency dashboards
(/opt/tutela/deploy/monitoring/grafana/dashboards/) carry the same breakdowns.
Real problem or artefact?
| Signal | Reading |
|---|---|
| Low request rate, spiky P95 | Artefact. Sparse-histogram interpolation. |
| P95 climbs, P50 flat, errors flat | Tail latency — contention or GC. Real but rarely urgent. |
| P50 and P95 both climb | Saturation. Check high-cpu.md and the dependency chain. |
| Inspector latency up and classifier latency up | The classifier is the cause; the inspector is blocked on it. Fix one, not both. |
| Provider phase alert alone | Upstream provider. Confirm on the provider's status page; not locally actionable. |
Latency up and tutela_inspector_semaphore_high_water_mark near 100 | Serious. The publish path is saturating and data loss is next — go to event-dropped.md. |
Escalation
Inspector latency is on the customer's inline path: above budget, users feel it on every prompt. Classifier T1 latency is the strongest predictor of inspector latency, because the inspector blocks on it.
Sustained latency has a second consequence beyond slowness. When the classifier
stops answering inside its timeout the inspector applies its fail mode, and the
deployment default is fail-open — so the visible symptom flips from "slow"
to "fast and ungoverned". Watch InspectorFailOpenElevated alongside this
alert; that transition matters more than the latency number.
Escalate with the percentile band from check 2, the phase breakdown from check 3, the request rate, and whether a deploy or traffic change preceded it.
Related runbooks
- inspector-failclosed-triage.md — what happens when slowness becomes a timeout.
- high-cpu.md — saturation as a cause.
- redis-latency.md — a slow cache surfaces as service latency.
- cache-hit-rate.md — a falling hit rate moves work onto the slow path.
- high-error-rate.md — when timeouts become 5xx.