Skip to main content

Error Budget Consumed

Target of the ErrorBudget80PercentConsumed (warning) and ErrorBudget95PercentConsumed (critical) alerts (/opt/tutela/deploy/monitoring/alerts/service-alerts.yml).

What fired​

(
1 - (
sum(rate(tutela_api_request_total{status!~"5.."}[30d])) by (job)
/
sum(rate(tutela_api_request_total[30d])) by (job)
)
) / (1 - 0.999) > 0.80 # warning; > 0.95 critical

Both for: 5m.

The SLO is 99.9% of requests non-5xx, hardcoded as the 0.999 in the expression. That leaves 0.1% as the monthly error budget. The expression divides the observed failure ratio by that 0.1% allowance, so the result is the fraction of the budget already spent: 0.80 means 80% spent, with 20% of the month's allowable failure left.

The alerts' own guidance: at 80%, "consider reducing deployment velocity"; at 95%, "freeze deployments and investigate".

What this does NOT mean​

  • It does not mean anything is broken right now. This is a 30-day retrospective. A single bad hour three weeks ago can hold this alert on while the service is perfectly healthy today. It measures accumulated cost, not current state.
  • It is not a count of failed requests. It is a ratio of two rate() results. Requests are weighted by the rate at which they arrived, so an outage during peak traffic costs far more budget than the same outage overnight.
  • It does not cover latency, correctness, or availability of governance. The SLO here is 5xx-only. A service that answers every request with a fast, wrong, ungoverned 200 spends zero error budget. This alert would stay silent through a complete enforcement outage.
  • It does not cover 4xx. See high-error-rate.md.

Known measurement gap — read before quoting the number​

The window is 30 days; the default retention is 15. Prometheus is started with --storage.tsdb.retention.time=${PROMETHEUS_RETENTION_TIME:-15d} in the deployment's compose file (/opt/tutela/docker-compose.yaml). A rate(...[30d]) range vector can only contain the samples that still exist, so on a default deployment this alert computes a 15-day budget and calls it monthly.

The practical consequences:

  • The number is not wrong so much as shorter-window than labelled. A half-month of good traffic cannot dilute a bad day the way a full month would, so the alert is biased toward firing sooner than a true 30-day SLO.
  • Raise PROMETHEUS_RETENTION_TIME to at least 31d if the 30-day figure is going to be reported to anyone. Retention costs disk — see disk-space.md before raising it.
  • Do not "fix" this by shortening the alert's window to [15d]. That would silently redefine the SLO. Either lengthen retention or state the real window when reporting.

There are no recording rules and no multi-window burn-rate alerts in this deployment: these two thresholds are the whole error-budget implementation. A 30-day range vector evaluated every 15 seconds is also expensive — if Prometheus itself is struggling, this rule group is a place to look.

First three checks​

1. Read the current budget consumption per job.

docker exec tutela-prometheus \
/bin/promtool query instant http://localhost:9090 '(1 - (sum(rate(tutela_api_request_total{status!~"5.."}[30d])) by (job) / sum(rate(tutela_api_request_total[30d])) by (job))) / (1 - 0.999)'

2. Establish how much data actually backs it.

docker exec tutela-prometheus \
wget -qO- 'http://localhost:9090/api/v1/status/runtimeinfo' \
| jq -r '"retention: \(.data.storageRetention) oldest sample: \(.data.startTime)"'

If the retention is shorter than 30 days, the value from check 1 is a shorter-window number. Say so when you report it.

3. Find when the budget was spent — this is almost always one outage.

Grafana is the only monitoring surface this deployment publishes (Prometheus publishes no port), and a trend over time is what it is for. Open it from the console under Settings → Diagnostics, or at https://<your-console-host>/grafana, choose the Prometheus datasource, and run this over the window named below:

sum(rate(tutela_api_request_total{status=~"5..",job="<job>"}[1h]))

Window: last 30 days, 1h step.

Every hour with a non-zero rate is a candidate. Correlate the spikes with deploy times and with any HighErrorRate* alerts that fired then.

Real problem or artefact?​

SignalReading
Budget high, no current 5xxRetrospective. Find the past outage (check 3); nothing to fix now.
Budget climbing and HighErrorRateCritical firingLive failure in progress. Work high-error-rate.md first — this alert is bookkeeping.
Budget jumps immediately after Prometheus restarts or data is lostArtefact. A shortened effective window changes the denominator. Confirm with check 2.
Budget high on a low-traffic jobWeak signal. A job serving a handful of requests spends its whole "budget" on one error. Read the absolute count.
Budget high on gatewayTake seriously. That is the customer-facing SLO.
Alert fires on a freshly deployed stackArtefact. Days of history do not exist yet; the ratio is computed from a tiny sample.

Because the window is 30 days, this alert clears slowly — often weeks after the outage that caused it. That is by design. Do not silence it to make it quiet; note the cause and let it age out.

Escalation​

This alert is a delivery-process signal, not a paging signal. The action it asks for is organisational: slow down at 80%, stop shipping at 95% and fix reliability. Route it to whoever owns release cadence, not to whoever is on call tonight — unless a live failure is also firing, in which case that failure is the priority.

When reporting the number, state the real window (check 2). Reporting a 15-day figure as a monthly SLO is the kind of small inaccuracy that erodes trust in every other number on the dashboard.