Redis High Latency
Target of the RedisHighLatency alert
(/opt/tutela/deploy/monitoring/alerts/service-alerts.yml).
What fired
histogram_quantile(0.95,
sum(rate(tutela_redis_operation_duration_ms_bucket[5m])) by (le, service)
) > 50
for: 5m, severity warning.
One in twenty Redis operations from {{ $labels.service }} took longer than
50 ms. For an in-memory store on the same network, healthy is single-digit
milliseconds — 50 ms means something is queueing.
The instrumented callers are the same set as
redis-errors.md: the gateway's SSO-state and SSO
exchange-grant, session/cache, account-lockout and rate-limiter operations, the
ai-proxy's rate_limit, and the mcp-gateway's session store.
Histogram limits — read before quoting a number
The histogram's buckets are 0.5, 1, 2, 5, 10, 25, 50, 100, 250, 500 ms.
Two consequences:
- Resolution above 50 ms is coarse. The next boundaries are 100, 250, 500. A reported P95 of "170 ms" is interpolated inside the 100–250 bucket, not measured. Treat values above 50 ms as "over budget by roughly this much", never as a precise figure.
- Anything slower than 500 ms lands in
+Infand the quantile becomes an extrapolation. A P95 reported at or near 500 ms usually means "much worse than 500 ms".
Durations at or above 1 ms are recorded as whole milliseconds; only sub-millisecond operations keep finer precision. That quantization is invisible at this alert's threshold but matters if you are comparing healthy baselines.
The alert groups by service, not by operation, so a single slow operation
drags the whole service's P95. Always split by operation before concluding
anything.
What this does NOT mean
- It does not mean Redis is failing. Slow operations still succeed. If they were failing you would also have redis-errors.md.
- It does not mean the Redis server is slow. This measures the round trip from the caller: client-side pool contention, DNS, TLS handshakes, or network latency all land in this number. On the Compose deployment on AWS, Redis is managed ElastiCache across the network (on Kubernetes it is a pod in your cluster) — a slow path is at least as likely as a slow server.
- It does not mean users are waiting 50 ms extra. These are not all on the request path, and none of them are on the classification or policy path.
- It is P95, not typical. See the tail-vs-average discussion in high-latency.md.
First three checks
1. Split by operation and compare the percentile band.
docker exec tutela-prometheus \
/bin/promtool query instant http://localhost:9090 'histogram_quantile(0.95, sum(rate(tutela_redis_operation_duration_ms_bucket[5m])) by (le, service, operation))'
docker exec tutela-prometheus \
/bin/promtool query instant http://localhost:9090 'histogram_quantile(0.5, sum(rate(tutela_redis_operation_duration_ms_bucket[5m])) by (le, service))'
A healthy P50 with a bad P95 is contention or pool starvation. Both elevated is a genuinely slow store or network.
2. Confirm there is enough traffic for the quantile to mean anything.
docker exec tutela-prometheus \
/bin/promtool query instant http://localhost:9090 'sum(rate(tutela_redis_operation_duration_ms_count[5m])) by (service)'
gateway only touches Redis during SSO logins. At a fraction of an operation
per second its P95 is a description of two or three logins, not a measurement.
3. Check the connection pool, then the server.
docker exec tutela-prometheus \
/bin/promtool query instant http://localhost:9090 '{__name__=~"tutela_redis_connections_(active|idle|stale)"}'
active climbing with idle at zero is pool exhaustion: callers are queueing
for a connection, and the wait is being counted as Redis latency. That is a
client-side fix, not a server one.
The connection-pool gauges are reported by the gateway only — it is the one
service that publishes its pool statistics. An empty result for ai-proxy or
mcp-gateway means the metric is not collected for them, not that their pools
are healthy.
For the server, on the Compose deployment on AWS: Redis is managed ElastiCache,
not a container, so read its CloudWatch metrics for the alert window —
EngineCPUUtilization, CurrConnections, evictions, swap and failover events —
and compare the cluster's availability zone with the host's; a cross-zone hop is
a common reason a healthy store reads as slow from here.
On Kubernetes this differs. The Helm chart runs Redis as a pod in your
cluster when redis.enabled is true in your chart values (it is true by
default), so the server side is inspectable there, and the slow log is the
direct answer to "is the store slow, or the path to it":
kubectl -n <namespace> get pods -l app.kubernetes.io/name=redis -o wide
kubectl -n <namespace> exec <redis-pod> -- redis-cli -a '<redis.auth.password>' slowlog get 20
An empty slow log with a rising P95 here means the time is spent between the
caller and the pod — compare the pod's node with the callers' nodes. If you set
redis.enabled: false and pointed REDIS_URL at a store you run yourself, that
store's own metrics are the place to look.
Real problem or artefact?
| Signal | Reading |
|---|---|
gateway alone, well under 1 op/s | Artefact. Sparse histogram over a handful of SSO logins. |
| P95 just above 50 ms, P50 healthy, no errors | Marginal. Real but low-impact; watch rather than act. |
| P95 and P50 both climbing | Real. Store, network, or pool. Continue to check 3. |
active high and idle zero | Pool exhaustion in the caller. Not Redis's fault. |
| Latency rising then errors following | The normal progression. Act now, before it becomes redis-errors.md and then redis-fail-open.md. |
| Step change after a deploy or failover | Configuration or topology change — new endpoint, TLS newly enabled, different availability zone. |
| P95 pinned near 500 ms | Almost certainly worse than it reads. Treat as severe. |
Escalation
On its own, warning-grade. The reason to act early is the progression: slow Redis becomes failing Redis, and failing Redis becomes a bypassed control via redis-fail-open.md. Catching it at the latency stage avoids the enforcement gap entirely.
Escalate with: the per-operation P95 from check 1, the P50 for comparison, the operation rate, the connection-pool figures, and whether TLS or the endpoint changed recently.
Related runbooks
- redis-errors.md — what this becomes if it worsens.
- redis-fail-open.md — the control bypass at the end of that chain.
- cache-hit-rate.md — the same store from the caller's side.
- high-latency.md — service latency, including the tail-vs-average reasoning that applies here too.