Error Budget Consumed
Target of the ErrorBudget80PercentConsumed (warning) and
ErrorBudget95PercentConsumed (critical) alerts
(/opt/tutela/deploy/monitoring/alerts/service-alerts.yml).
What fired
(
1 - (
sum(rate(tutela_api_request_total{status!~"5.."}[30d])) by (job)
/
sum(rate(tutela_api_request_total[30d])) by (job)
)
) / (1 - 0.999) > 0.80 # warning; > 0.95 critical
Both for: 5m.
The SLO is 99.9% of requests non-5xx, hardcoded as the 0.999 in the
expression. That leaves 0.1% as the monthly error budget. The expression divides
the observed failure ratio by that 0.1% allowance, so the result is the
fraction of the budget already spent: 0.80 means 80% spent, with 20% of the
month's allowable failure left.
The alerts' own guidance: at 80%, "consider reducing deployment velocity"; at 95%, "freeze deployments and investigate".
What this does NOT mean
- It does not mean anything is broken right now. This is a 30-day retrospective. A single bad hour three weeks ago can hold this alert on while the service is perfectly healthy today. It measures accumulated cost, not current state.
- It is not a count of failed requests. It is a ratio of two
rate()results. Requests are weighted by the rate at which they arrived, so an outage during peak traffic costs far more budget than the same outage overnight. - It does not cover latency, correctness, or availability of governance. The SLO here is 5xx-only. A service that answers every request with a fast, wrong, ungoverned 200 spends zero error budget. This alert would stay silent through a complete enforcement outage.
- It does not cover 4xx. See high-error-rate.md.
Known measurement gap — read before quoting the number
The window is 30 days; the default retention is 15. Prometheus is started
with --storage.tsdb.retention.time=${PROMETHEUS_RETENTION_TIME:-15d} in the
deployment's compose file (/opt/tutela/docker-compose.yaml). A
rate(...[30d]) range vector can only contain the samples that still exist, so
on a default deployment this alert computes a 15-day budget and calls it
monthly.
The practical consequences:
- The number is not wrong so much as shorter-window than labelled. A half-month of good traffic cannot dilute a bad day the way a full month would, so the alert is biased toward firing sooner than a true 30-day SLO.
- Raise
PROMETHEUS_RETENTION_TIMEto at least31dif the 30-day figure is going to be reported to anyone. Retention costs disk — see disk-space.md before raising it. - Do not "fix" this by shortening the alert's window to
[15d]. That would silently redefine the SLO. Either lengthen retention or state the real window when reporting.
There are no recording rules and no multi-window burn-rate alerts in this deployment: these two thresholds are the whole error-budget implementation. A 30-day range vector evaluated every 15 seconds is also expensive — if Prometheus itself is struggling, this rule group is a place to look.
First three checks
1. Read the current budget consumption per job.
docker exec tutela-prometheus \
/bin/promtool query instant http://localhost:9090 '(1 - (sum(rate(tutela_api_request_total{status!~"5.."}[30d])) by (job) / sum(rate(tutela_api_request_total[30d])) by (job))) / (1 - 0.999)'
2. Establish how much data actually backs it.
docker exec tutela-prometheus \
wget -qO- 'http://localhost:9090/api/v1/status/runtimeinfo' \
| jq -r '"retention: \(.data.storageRetention) oldest sample: \(.data.startTime)"'
If the retention is shorter than 30 days, the value from check 1 is a shorter-window number. Say so when you report it.
3. Find when the budget was spent — this is almost always one outage.
Grafana is the only monitoring surface this deployment publishes (Prometheus
publishes no port), and a trend over time is what it is for. Open it from the
console under Settings → Diagnostics, or at
https://<your-console-host>/grafana, choose the Prometheus datasource, and run
this over the window named below:
sum(rate(tutela_api_request_total{status=~"5..",job="<job>"}[1h]))
Window: last 30 days, 1h step.
Every hour with a non-zero rate is a candidate. Correlate the spikes with deploy
times and with any HighErrorRate* alerts that fired then.
Real problem or artefact?
| Signal | Reading |
|---|---|
| Budget high, no current 5xx | Retrospective. Find the past outage (check 3); nothing to fix now. |
Budget climbing and HighErrorRateCritical firing | Live failure in progress. Work high-error-rate.md first — this alert is bookkeeping. |
| Budget jumps immediately after Prometheus restarts or data is lost | Artefact. A shortened effective window changes the denominator. Confirm with check 2. |
| Budget high on a low-traffic job | Weak signal. A job serving a handful of requests spends its whole "budget" on one error. Read the absolute count. |
Budget high on gateway | Take seriously. That is the customer-facing SLO. |
| Alert fires on a freshly deployed stack | Artefact. Days of history do not exist yet; the ratio is computed from a tiny sample. |
Because the window is 30 days, this alert clears slowly — often weeks after the outage that caused it. That is by design. Do not silence it to make it quiet; note the cause and let it age out.
Escalation
This alert is a delivery-process signal, not a paging signal. The action it asks for is organisational: slow down at 80%, stop shipping at 95% and fix reliability. Route it to whoever owns release cadence, not to whoever is on call tonight — unless a live failure is also firing, in which case that failure is the priority.
When reporting the number, state the real window (check 2). Reporting a 15-day figure as a monthly SLO is the kind of small inaccuracy that erodes trust in every other number on the dashboard.
Related runbooks
- high-error-rate.md — the live version of this signal.
- service-down.md — outages are what spend the budget.
- disk-space.md — before raising Prometheus retention.