High Error Rate
Target of the HighErrorRateWarning and HighErrorRateCritical alerts
(/opt/tutela/deploy/monitoring/alerts/service-alerts.yml).
What fired
(
sum(rate(tutela_api_request_total{status=~"5.."}[5m])) by (job)
/
sum(rate(tutela_api_request_total[5m])) by (job)
) > 0.01 # warning; > 0.05 critical
Both for: 5m. More than 1% (warning) or 5% (critical) of {{ $labels.job }}'s
requests returned a 5xx over the last five minutes.
tutela_api_request_total is the shared HTTP counter, labelled method, path
and status. The ratio is per job, so a service handling four requests a
minute crosses 1% on a single error.
What this does NOT mean
- It does not mean customers saw errors. 5xx on
analyticsordspm-bridgeis internal; onlygatewayfailures are directly customer-facing. - It does not include 4xx. A flood of 401s, 403s or 429s — an expired credential, a misconfigured client, a rate-limited integration — will NOT fire this alert. Absence of this alert is not evidence that clients are succeeding.
- It does not mean policy stopped enforcing. A blocked prompt is a
successful 200 response carrying a
blockdecision, not a 5xx. Governance failure has its own alerts — see inspector-failclosed-triage.md. - It is a rate ratio, not an error count. At low traffic the denominator is tiny and the ratio is noise. Always read the absolute rate before reacting.
First three checks
1. Get the absolute numbers, not just the ratio.
docker exec tutela-prometheus \
/bin/promtool query instant http://localhost:9090 'sum(rate(tutela_api_request_total{status=~"5.."}[5m])) by (job)'
Under roughly 0.05 errors/s this is a handful of requests, not an outage.
2. Find which endpoint and which status.
docker exec tutela-prometheus \
/bin/promtool query instant http://localhost:9090 'topk(10, sum(rate(tutela_api_request_total{status=~"5..",job="<job>"}[5m])) by (path, status))'
A single path means a broken handler or a bad query. Every path at once means
a shared dependency — database, Redis, or an upstream service.
3. Read the errors themselves.
docker logs --tail=300 tutela-<service> \
| grep -E '"level":"(error|warn)"'
Every service logs structured JSON carrying a request_id. Take one
request_id from a failing request and follow it across services — the gateway
propagates X-Request-ID on every internal call, so one id reconstructs the
whole path. In Loki:
{service="<service>", level="error"}
Real problem or artefact?
| Signal | Reading |
|---|---|
| Ratio high, absolute rate < 0.05/s | Artefact of a small denominator. Watch, do not page. |
| Ratio steps up exactly at a deploy | Regression in the new build. Check /v1/version; roll back before debugging. |
Errors on one path only | Handler or query bug. A 500 on a route that previously worked is usually a schema/migration mismatch. |
| Errors across every path on one job | That service's dependency. Check /readyz and its database/Redis/Kafka connections. |
| Errors across every job at once | Shared infrastructure — the database, or the network. Do not chase individual services. |
| Warning fires but critical never does | Chronic low-level failure. Real, but not an outage; often one broken integration retrying. |
Expect to receive both alerts during a bad outage. The inhibition rule in
alertmanager.yml suppresses a warning only when a critical with the same
alertname and service is firing, and these are two different alertnames.
That is the configured behaviour, not a routing bug.
Escalation
5xx on gateway above 5% is a customer-visible outage — escalate immediately
with the failing path, status, a sample request_id, and the running build
from /v1/version.
Never "fix" this alert by widening the threshold. The thresholds encode the 99.9% SLO that error-budget.md measures against; moving them changes what the product promises.
Related runbooks
- service-down.md — when the service stops answering entirely.
- high-latency.md — timeouts often surface as 5xx upstream.
- error-budget.md — the accumulated cost of these errors.
- redis-errors.md — a common shared-dependency cause.