Redis Fail-Open and Bounded Control Degradation
Target of the RedisFailOpenWarning, RedisFailOpenCritical,
RedisControlDegraded, and RedisControlDegradedCritical alerts
(/opt/tutela/deploy/monitoring/alerts/service-alerts.yml).
What fired
rate(tutela_redis_fail_open_total[5m]) > 0.1 # warning, for: 2m
rate(tutela_redis_fail_open_total[5m]) > 1 # critical, for: 1m
A control that depends on Redis could not reach it and allowed the operation anyway. This is the deliberate fail-open path: the code chose availability over enforcement, and incremented this counter so the choice is visible rather than silent.
Read the operation label — it decides what was actually bypassed.
The separate degraded-control signal means Redis was unavailable but the control continued enforcing through a bounded per-instance fallback:
sum(rate(tutela_redis_degraded_control_total[5m])) by (service, operation, mode)
For AI Proxy rate and quota controls, degradation is no longer allow-all. Each replica keeps tenant-scoped request buckets and daily token usage locally. That is weaker than a shared global ceiling (N replicas can each use their local allowance), but it remains bounded and keeps outage-period usage in the local ceiling after Redis recovers. It does not replay failed or timed-out writes into the shared counter because Redis may already have applied an ambiguous write; the shared fleet-wide view can therefore remain incomplete until the daily reset even after connectivity returns.
The remaining fail-open call sites
tutela_redis_fail_open_total remains for operations that genuinely allow work
without their normal shared Redis state:
service / operation | What is bypassed |
|---|---|
mcp-gateway / session_get | A session lookup could not be served from the shared store. |
mcp-gateway / session_create | A session could not be persisted to the shared store. |
For the two mcp-gateway session operations the cost is session continuity
across replicas, not rate limiting. Do not report an MCP session fail-open as
"rate limiting disabled".
What this does NOT mean
- It does not mean the gateway's rate limiting failed open. The gateway
deliberately does not fail open. On a Redis outage it degrades to a
per-instance in-memory bucket and keeps limiting, and it records that on a
different metric,
tutela_gateway_ratelimit_redis_fallback_total{reason}. That metric has no alert rule — so gateway rate-limit degradation is invisible here and must be queried directly. - It does not mean content inspection or policy stopped. Classification, policy evaluation and blocking do not depend on Redis. A prompt that would be blocked is still blocked. The inspector's own fail mode is a separate mechanism — see inspector-failclosed-triage.md.
- It does not mean AI Proxy limits disappeared. AI Proxy rate and quota
controls emit
tutela_redis_degraded_control_totaland keep a bounded local ceiling when Redis is unavailable. - It does not mean requests failed. Fail-open means they succeeded — that is the problem. Nothing looks wrong from the outside.
First three checks
1. Establish exactly what is being bypassed, and how hard.
docker exec tutela-prometheus \
/bin/promtool query instant http://localhost:9090 'rate(tutela_redis_fail_open_total[5m])'
docker exec tutela-prometheus \
/bin/promtool query instant http://localhost:9090 'increase(tutela_redis_fail_open_total[1h])'
docker exec tutela-prometheus \
/bin/promtool query instant http://localhost:9090 'sum(rate(tutela_redis_degraded_control_total[5m])) by (service, operation, mode)'
A degraded rate_limit or quota_* row means enforcement is per-replica rather
than shared. Check traffic and spend alongside it and restore Redis promptly.
2. Is Redis reachable, or is it simply not configured?
The AI Proxy degraded-control signal fires both when REDIS_URL is unset and
when Redis is down. Distinguish them from the host — the service image has no
shell, and the value is a credential, so check for its presence without
printing it:
docker inspect tutela-ai-proxy --format '{{range .Config.Env}}{{println .}}{{end}}' \
| grep -q '^REDIS_URL=.' && echo 'REDIS_URL is set' || echo 'REDIS_URL is UNSET'
An empty or unset REDIS_URL means the shared fleet-wide rate and quota
ceilings are unavailable. AI Proxy still enforces bounded per-instance limits,
but that is a production configuration defect rather than a healthy state. The
deployment's compose file makes REDIS_URL mandatory
(${REDIS_URL:?REDIS_URL is required}), so an unset value means someone is
running a modified configuration.
If it is set, test the connection from the stack's side. On the Compose
deployment on AWS, Redis is managed ElastiCache, not a container, so there is no
redis container to inspect (on Kubernetes there is a Redis pod — see below);
the quickest confirmation that the problem is Redis rather than one service's
client is the gateway's readiness body:
curl -s http://127.0.0.1:8080/readyz
It reports {"status":"not_ready","reason":"redis unavailable"} when Redis is
configured but unreachable. (The gateway serves port 8080 as plain HTTP in
every posture — internal mTLS does not apply to it — so this form needs no
certificate even when the rest of the stack runs mTLS.) For the server side on
the Compose deployment on AWS, read the ElastiCache console: failover events,
CurrConnections, EngineCPUUtilization and evictions over the alert window.
On Kubernetes this differs. The Helm chart runs Redis as a pod in your
cluster when redis.enabled is true in your chart values (it is true by
default), so the server side is inspectable there:
kubectl -n <namespace> get pods -l app.kubernetes.io/name=redis
kubectl -n <namespace> logs -l app.kubernetes.io/name=redis --tail=200
kubectl -n <namespace> exec <redis-pod> -- redis-cli -a '<redis.auth.password>' info clients
A pod that is Running with a healthy connected_clients count while the
gateway reports redis unavailable points at the network path or REDIS_URL,
not at the store. If you set redis.enabled: false and pointed REDIS_URL at a
store you run yourself, that store's own console is the place to look.
3. Check the correlated Redis health signals.
docker exec tutela-prometheus \
/bin/promtool query instant http://localhost:9090 'sum(rate(tutela_redis_operations_total{status="error"}[5m])) by (service, operation)'
docker exec tutela-prometheus \
/bin/promtool query instant http://localhost:9090 'tutela_redis_connections_active'
Query both bounded degradation signals:
docker exec tutela-prometheus \
/bin/promtool query instant http://localhost:9090 'sum(rate(tutela_gateway_ratelimit_redis_fallback_total[5m])) by (reason)'
docker exec tutela-prometheus \
/bin/promtool query instant http://localhost:9090 'sum(rate(tutela_redis_degraded_control_total[5m])) by (service, operation, mode)'
Real problem or artefact?
| Signal | Reading |
|---|---|
Steady degraded rate on ai-proxy, REDIS_URL unset | Configuration defect. Limits are only per-instance. Restore the required shared Redis configuration. |
| Burst during a Redis failover or restart | Expected transient. Confirm it drains; record the per-instance enforcement window. |
| Sustained critical rate with Redis errors rising | Real Redis outage. Work the outage; this alert is the consequence. |
| Fail-opens with no Redis errors at all | The client is not reaching Redis — DNS, TLS, credentials, or an unset URL. Check 2. |
mcp-gateway/session_* only, gateway healthy | Session store degraded. Sessions may not survive a replica hop; rate limiting is unaffected. |
| Alert clears but you never found a cause | Do not close it. A silent bypass window is exactly what this counter exists to make non-silent. |
Escalation
Treat sustained AI Proxy degraded-control activity as a shared-control outage with cost exposure, not merely an infrastructure ticket. Limits remain bounded per replica, but the fleet-wide ceiling is relaxed. Record the window; if spend or abuse is a concern, reconcile usage for that period against ai-spend-operations.md.
Escalate with the operation breakdown, the duration, whether REDIS_URL is
configured, and the Redis error rate.
There is no environment variable that disables the bounded AI Proxy
fallback. It is the compiled safety behavior whenever shared Redis is absent
or unavailable.
The two fail-mode settings your deployment does expose in /opt/tutela/.env
(FAIL_MODE and POLICY_FAIL_MODE, both read by the inspector) govern
inspection — what happens to a request when the classifier or the policy
service cannot answer — not Redis, and changing them will not affect this
alert.
Related runbooks
- redis-errors.md — the failures that cause the fail-open.
- redis-latency.md — slow Redis becomes failed Redis.
- cache-hit-rate.md — the same store, seen as misses.
- inspector-failclosed-triage.md — the product's other, unrelated fail-open decision.
- ai-spend-operations.md — reconciling spend after a shared-control degradation window.