Redis High Error Rate
Target of the RedisHighErrorRate alert
(/opt/tutela/deploy/monitoring/alerts/service-alerts.yml).
What fired
(
sum(rate(tutela_redis_operations_total{status="error"}[5m])) by (service)
/
sum(rate(tutela_redis_operations_total[5m])) by (service)
) > 0.05
for: 5m, severity warning.
More than 5% of {{ $labels.service }}'s Redis operations returned an error
over five minutes. tutela_redis_operations_total is labelled
service, operation, status, and every operation is recorded as either
success or error.
Which services and operations are instrumented
Only these call sites record Redis operations, so only these can raise the
alert. Read the operation label — it tells you which subsystem is failing:
service | operation values | What it is |
|---|---|---|
gateway | sso_state_set, sso_state_get, sso_state_del | Single-use SSO login state. |
gateway | sso_exchange_grant_set, sso_exchange_grant_get, sso_exchange_grant_del | The single-use code an SSO login exchanges for its session. |
gateway | get, set, del | Session and cache storage. |
gateway | incr, expire | The failed-login / account-lockout counter. |
gateway | evalsha, script_load | The rate limiter's server-side script. |
ai-proxy | rate_limit | Per-tenant request throttling. |
mcp-gateway | session_get, session_set | Shared MCP session store. |
Redis is used in a few places beyond these — the ai-proxy quota check and the runtime governance session projection among them — that record nothing. A quiet alert does not show that every Redis interaction is healthy.
What this does NOT mean
- It does not mean Redis is down. A 5% error ratio means 95% of operations succeeded. A total outage looks different: the ratio pins at 1.0, or the series disappears entirely because no operations are attempted.
- It does not mean requests failed. Most of these call sites degrade rather than fail. Errors here become fail-opens (redis-fail-open.md) or cache misses (cache-hit-rate.md), not customer-visible 5xx.
- It does not include "key not found". A missing key is a normal miss, not an error. If it were counted, every first lookup would inflate this alert.
- It is a ratio at low volume.
gatewayperformssso_stateoperations only during SSO logins. On a quiet night a couple of failed callbacks can push the gateway's ratio past 5% while nothing is wrong with Redis. - There is no error-type breakdown to reach for. A
tutela_redis_errors_totalmetric with anerror_typelabel is declared, but nothing in the running services ever writes to it — querying it returns nothing, which is a dead end rather than a clean result. Classify the failure from the service's own logs instead (check 3).
First three checks
1. Get the absolute error rate and the failing operation.
docker exec tutela-prometheus \
/bin/promtool query instant http://localhost:9090 'sum(rate(tutela_redis_operations_total{status="error"}[5m])) by (service, operation)'
The operation label is the fastest localiser: incr/expire is the lockout
counter, evalsha/script_load the rate limiter, sso_state_* the login
flow, session_* the MCP session store. What KIND of failure it is comes from
the logs in check 3, not from a metric.
2. Is this one service's client, or Redis itself?
docker exec tutela-prometheus \
/bin/promtool query instant http://localhost:9090 'sum(rate(tutela_redis_operations_total{status="error"}[5m])) by (service)'
All three services erroring together points at Redis or the network. One service alone points at that service's client — its connection pool, its credentials, or its TLS material.
Then confirm reachability from the host, through the gateway's readiness
endpoint. The gateway publishes port 8080 and serves it as plain HTTP whether
or not the deployment runs with internal mTLS — it is the customer-facing edge,
and internal mTLS does not apply to it — so this one form is right in every
posture:
curl -s -w '\nHTTP %{http_code}\n' http://127.0.0.1:8080/readyz
You will see one of three bodies. {"status":"not_ready","reason":"redis unavailable"}
with HTTP 503 means Redis is configured but the gateway cannot open a
connection to it. {"status":"ready",...} with HTTP 200 means it can — the
errors are then in individual operations, not in reachability. And
{"status":"not_ready","reason":"database unavailable"} with HTTP 503 means
the database check, which runs first, failed — readiness stops there, so it
tells you nothing about Redis until the database is back.
3. Look at the connection pool and the server.
docker exec tutela-prometheus \
/bin/promtool query instant http://localhost:9090 '{__name__=~"tutela_redis_connections_(active|idle|stale)"}'
A rising stale count with errors means the pool is holding dead connections —
typical after a failover or an idle-timeout mismatch.
The connection-pool gauges are reported by the gateway only — it is the one
service that publishes its pool statistics. An empty result for ai-proxy or
mcp-gateway means the metric is not collected for them, not that their pools
are healthy.
On the Compose deployment on AWS, Redis itself is managed ElastiCache, not a container, so there is nothing on the host to inspect for the server side: check the ElastiCache console for failover events, CPU, evictions and connection counts over the alert window.
On Kubernetes this differs. The Helm chart runs Redis as a pod in your
cluster when redis.enabled is true in your chart values (it is true by
default), so the server side is inspectable there:
kubectl -n <namespace> get pods -l app.kubernetes.io/name=redis
kubectl -n <namespace> logs -l app.kubernetes.io/name=redis --tail=200
kubectl -n <namespace> exec <redis-pod> -- redis-cli -a '<redis.auth.password>' info stats
If you set redis.enabled: false and pointed REDIS_URL at a store you run
yourself, that store's own console is the place to look.
Also read the dependent service's own logs — the error text is usually explicit:
docker logs --tail=200 tutela-<service> \
| grep -i redis
Real problem or artefact?
| Signal | Reading |
|---|---|
gateway only, sso_state_*, very low absolute rate | Artefact of a small denominator. A few failed SSO callbacks, not a Redis problem. |
| All three services, errors rising together | Real. Redis or the network between it and the stack. |
Errors spike briefly then stop, stale connections rose | Failover or restart. Pool recovered. Record the window. |
| Logs show timeouts, and latency is also elevated | Redis is saturated, not broken. Go to redis-latency.md. |
| Logs show auth or TLS failures | Credential or certificate problem, not capacity. Check REDIS_URL and the REDIS_*_CERT_PATH settings. |
Errors alongside RedisFailOpen* firing | Expected pairing. The errors are the cause; the fail-opens are the consequence and the more serious of the two. |
Escalation
By itself this is a warning-grade degradation signal. It becomes urgent
when it is paired with RedisFailOpenCritical — at that point a control is
being bypassed, and redis-fail-open.md is the runbook that
matters. Escalate on the bypass, not on the error ratio.
Escalate with the failing operation, the error text from the logs, whether one
or all services are affected, the connection-pool figures, and whether a
failover or configuration change preceded it.
Restarting the dependent service clears a poisoned connection pool but does nothing if Redis itself is the problem, and it drops any in-process cache alongside it — expect a cache-hit-rate.md alert to follow.
Related runbooks
- redis-fail-open.md — what these errors cause.
- redis-latency.md — the usual precursor.
- cache-hit-rate.md — the same store seen from the caller.
- service-down.md — when a Redis-dependent service fails readiness entirely.