Skip to main content

High Error Rate

Target of the HighErrorRateWarning and HighErrorRateCritical alerts (/opt/tutela/deploy/monitoring/alerts/service-alerts.yml).

What fired​

(
sum(rate(tutela_api_request_total{status=~"5.."}[5m])) by (job)
/
sum(rate(tutela_api_request_total[5m])) by (job)
) > 0.01 # warning; > 0.05 critical

Both for: 5m. More than 1% (warning) or 5% (critical) of {{ $labels.job }}'s requests returned a 5xx over the last five minutes.

tutela_api_request_total is the shared HTTP counter, labelled method, path and status. The ratio is per job, so a service handling four requests a minute crosses 1% on a single error.

What this does NOT mean​

  • It does not mean customers saw errors. 5xx on analytics or dspm-bridge is internal; only gateway failures are directly customer-facing.
  • It does not include 4xx. A flood of 401s, 403s or 429s — an expired credential, a misconfigured client, a rate-limited integration — will NOT fire this alert. Absence of this alert is not evidence that clients are succeeding.
  • It does not mean policy stopped enforcing. A blocked prompt is a successful 200 response carrying a block decision, not a 5xx. Governance failure has its own alerts — see inspector-failclosed-triage.md.
  • It is a rate ratio, not an error count. At low traffic the denominator is tiny and the ratio is noise. Always read the absolute rate before reacting.

First three checks​

1. Get the absolute numbers, not just the ratio.

docker exec tutela-prometheus \
/bin/promtool query instant http://localhost:9090 'sum(rate(tutela_api_request_total{status=~"5.."}[5m])) by (job)'

Under roughly 0.05 errors/s this is a handful of requests, not an outage.

2. Find which endpoint and which status.

docker exec tutela-prometheus \
/bin/promtool query instant http://localhost:9090 'topk(10, sum(rate(tutela_api_request_total{status=~"5..",job="<job>"}[5m])) by (path, status))'

A single path means a broken handler or a bad query. Every path at once means a shared dependency — database, Redis, or an upstream service.

3. Read the errors themselves.

docker logs --tail=300 tutela-<service> \
| grep -E '"level":"(error|warn)"'

Every service logs structured JSON carrying a request_id. Take one request_id from a failing request and follow it across services — the gateway propagates X-Request-ID on every internal call, so one id reconstructs the whole path. In Loki:

{service="<service>", level="error"}

Real problem or artefact?​

SignalReading
Ratio high, absolute rate < 0.05/sArtefact of a small denominator. Watch, do not page.
Ratio steps up exactly at a deployRegression in the new build. Check /v1/version; roll back before debugging.
Errors on one path onlyHandler or query bug. A 500 on a route that previously worked is usually a schema/migration mismatch.
Errors across every path on one jobThat service's dependency. Check /readyz and its database/Redis/Kafka connections.
Errors across every job at onceShared infrastructure — the database, or the network. Do not chase individual services.
Warning fires but critical never doesChronic low-level failure. Real, but not an outage; often one broken integration retrying.

Expect to receive both alerts during a bad outage. The inhibition rule in alertmanager.yml suppresses a warning only when a critical with the same alertname and service is firing, and these are two different alertnames. That is the configured behaviour, not a routing bug.

Escalation​

5xx on gateway above 5% is a customer-visible outage — escalate immediately with the failing path, status, a sample request_id, and the running build from /v1/version.

Never "fix" this alert by widening the threshold. The thresholds encode the 99.9% SLO that error-budget.md measures against; moving them changes what the product promises.