Skip to main content

Redis High Error Rate

Target of the RedisHighErrorRate alert (/opt/tutela/deploy/monitoring/alerts/service-alerts.yml).

What fired​

(
sum(rate(tutela_redis_operations_total{status="error"}[5m])) by (service)
/
sum(rate(tutela_redis_operations_total[5m])) by (service)
) > 0.05

for: 5m, severity warning.

More than 5% of {{ $labels.service }}'s Redis operations returned an error over five minutes. tutela_redis_operations_total is labelled service, operation, status, and every operation is recorded as either success or error.

Which services and operations are instrumented​

Only these call sites record Redis operations, so only these can raise the alert. Read the operation label — it tells you which subsystem is failing:

serviceoperation valuesWhat it is
gatewaysso_state_set, sso_state_get, sso_state_delSingle-use SSO login state.
gatewaysso_exchange_grant_set, sso_exchange_grant_get, sso_exchange_grant_delThe single-use code an SSO login exchanges for its session.
gatewayget, set, delSession and cache storage.
gatewayincr, expireThe failed-login / account-lockout counter.
gatewayevalsha, script_loadThe rate limiter's server-side script.
ai-proxyrate_limitPer-tenant request throttling.
mcp-gatewaysession_get, session_setShared MCP session store.

Redis is used in a few places beyond these — the ai-proxy quota check and the runtime governance session projection among them — that record nothing. A quiet alert does not show that every Redis interaction is healthy.

What this does NOT mean​

  • It does not mean Redis is down. A 5% error ratio means 95% of operations succeeded. A total outage looks different: the ratio pins at 1.0, or the series disappears entirely because no operations are attempted.
  • It does not mean requests failed. Most of these call sites degrade rather than fail. Errors here become fail-opens (redis-fail-open.md) or cache misses (cache-hit-rate.md), not customer-visible 5xx.
  • It does not include "key not found". A missing key is a normal miss, not an error. If it were counted, every first lookup would inflate this alert.
  • It is a ratio at low volume. gateway performs sso_state operations only during SSO logins. On a quiet night a couple of failed callbacks can push the gateway's ratio past 5% while nothing is wrong with Redis.
  • There is no error-type breakdown to reach for. A tutela_redis_errors_total metric with an error_type label is declared, but nothing in the running services ever writes to it — querying it returns nothing, which is a dead end rather than a clean result. Classify the failure from the service's own logs instead (check 3).

First three checks​

1. Get the absolute error rate and the failing operation.

docker exec tutela-prometheus \
/bin/promtool query instant http://localhost:9090 'sum(rate(tutela_redis_operations_total{status="error"}[5m])) by (service, operation)'

The operation label is the fastest localiser: incr/expire is the lockout counter, evalsha/script_load the rate limiter, sso_state_* the login flow, session_* the MCP session store. What KIND of failure it is comes from the logs in check 3, not from a metric.

2. Is this one service's client, or Redis itself?

docker exec tutela-prometheus \
/bin/promtool query instant http://localhost:9090 'sum(rate(tutela_redis_operations_total{status="error"}[5m])) by (service)'

All three services erroring together points at Redis or the network. One service alone points at that service's client — its connection pool, its credentials, or its TLS material.

Then confirm reachability from the host, through the gateway's readiness endpoint. The gateway publishes port 8080 and serves it as plain HTTP whether or not the deployment runs with internal mTLS — it is the customer-facing edge, and internal mTLS does not apply to it — so this one form is right in every posture:

curl -s -w '\nHTTP %{http_code}\n' http://127.0.0.1:8080/readyz

You will see one of three bodies. {"status":"not_ready","reason":"redis unavailable"} with HTTP 503 means Redis is configured but the gateway cannot open a connection to it. {"status":"ready",...} with HTTP 200 means it can — the errors are then in individual operations, not in reachability. And {"status":"not_ready","reason":"database unavailable"} with HTTP 503 means the database check, which runs first, failed — readiness stops there, so it tells you nothing about Redis until the database is back.

3. Look at the connection pool and the server.

docker exec tutela-prometheus \
/bin/promtool query instant http://localhost:9090 '{__name__=~"tutela_redis_connections_(active|idle|stale)"}'

A rising stale count with errors means the pool is holding dead connections — typical after a failover or an idle-timeout mismatch.

The connection-pool gauges are reported by the gateway only — it is the one service that publishes its pool statistics. An empty result for ai-proxy or mcp-gateway means the metric is not collected for them, not that their pools are healthy.

On the Compose deployment on AWS, Redis itself is managed ElastiCache, not a container, so there is nothing on the host to inspect for the server side: check the ElastiCache console for failover events, CPU, evictions and connection counts over the alert window.

On Kubernetes this differs. The Helm chart runs Redis as a pod in your cluster when redis.enabled is true in your chart values (it is true by default), so the server side is inspectable there:

kubectl -n <namespace> get pods -l app.kubernetes.io/name=redis
kubectl -n <namespace> logs -l app.kubernetes.io/name=redis --tail=200
kubectl -n <namespace> exec <redis-pod> -- redis-cli -a '<redis.auth.password>' info stats

If you set redis.enabled: false and pointed REDIS_URL at a store you run yourself, that store's own console is the place to look.

Also read the dependent service's own logs — the error text is usually explicit:

docker logs --tail=200 tutela-<service> \
| grep -i redis

Real problem or artefact?​

SignalReading
gateway only, sso_state_*, very low absolute rateArtefact of a small denominator. A few failed SSO callbacks, not a Redis problem.
All three services, errors rising togetherReal. Redis or the network between it and the stack.
Errors spike briefly then stop, stale connections roseFailover or restart. Pool recovered. Record the window.
Logs show timeouts, and latency is also elevatedRedis is saturated, not broken. Go to redis-latency.md.
Logs show auth or TLS failuresCredential or certificate problem, not capacity. Check REDIS_URL and the REDIS_*_CERT_PATH settings.
Errors alongside RedisFailOpen* firingExpected pairing. The errors are the cause; the fail-opens are the consequence and the more serious of the two.

Escalation​

By itself this is a warning-grade degradation signal. It becomes urgent when it is paired with RedisFailOpenCritical — at that point a control is being bypassed, and redis-fail-open.md is the runbook that matters. Escalate on the bypass, not on the error ratio.

Escalate with the failing operation, the error text from the logs, whether one or all services are affected, the connection-pool figures, and whether a failover or configuration change preceded it.

Restarting the dependent service clears a poisoned connection pool but does nothing if Redis itself is the problem, and it drops any in-process cache alongside it — expect a cache-hit-rate.md alert to follow.