Skip to main content

Redis Fail-Open and Bounded Control Degradation

Target of the RedisFailOpenWarning, RedisFailOpenCritical, RedisControlDegraded, and RedisControlDegradedCritical alerts (/opt/tutela/deploy/monitoring/alerts/service-alerts.yml).

What fired​

rate(tutela_redis_fail_open_total[5m]) > 0.1 # warning, for: 2m
rate(tutela_redis_fail_open_total[5m]) > 1 # critical, for: 1m

A control that depends on Redis could not reach it and allowed the operation anyway. This is the deliberate fail-open path: the code chose availability over enforcement, and incremented this counter so the choice is visible rather than silent.

Read the operation label — it decides what was actually bypassed.

The separate degraded-control signal means Redis was unavailable but the control continued enforcing through a bounded per-instance fallback:

sum(rate(tutela_redis_degraded_control_total[5m])) by (service, operation, mode)

For AI Proxy rate and quota controls, degradation is no longer allow-all. Each replica keeps tenant-scoped request buckets and daily token usage locally. That is weaker than a shared global ceiling (N replicas can each use their local allowance), but it remains bounded and keeps outage-period usage in the local ceiling after Redis recovers. It does not replay failed or timed-out writes into the shared counter because Redis may already have applied an ambiguous write; the shared fleet-wide view can therefore remain incomplete until the daily reset even after connectivity returns.

The remaining fail-open call sites​

tutela_redis_fail_open_total remains for operations that genuinely allow work without their normal shared Redis state:

service / operationWhat is bypassed
mcp-gateway / session_getA session lookup could not be served from the shared store.
mcp-gateway / session_createA session could not be persisted to the shared store.

For the two mcp-gateway session operations the cost is session continuity across replicas, not rate limiting. Do not report an MCP session fail-open as "rate limiting disabled".

What this does NOT mean​

  • It does not mean the gateway's rate limiting failed open. The gateway deliberately does not fail open. On a Redis outage it degrades to a per-instance in-memory bucket and keeps limiting, and it records that on a different metric, tutela_gateway_ratelimit_redis_fallback_total{reason}. That metric has no alert rule — so gateway rate-limit degradation is invisible here and must be queried directly.
  • It does not mean content inspection or policy stopped. Classification, policy evaluation and blocking do not depend on Redis. A prompt that would be blocked is still blocked. The inspector's own fail mode is a separate mechanism — see inspector-failclosed-triage.md.
  • It does not mean AI Proxy limits disappeared. AI Proxy rate and quota controls emit tutela_redis_degraded_control_total and keep a bounded local ceiling when Redis is unavailable.
  • It does not mean requests failed. Fail-open means they succeeded — that is the problem. Nothing looks wrong from the outside.

First three checks​

1. Establish exactly what is being bypassed, and how hard.

docker exec tutela-prometheus \
/bin/promtool query instant http://localhost:9090 'rate(tutela_redis_fail_open_total[5m])'

docker exec tutela-prometheus \
/bin/promtool query instant http://localhost:9090 'increase(tutela_redis_fail_open_total[1h])'

docker exec tutela-prometheus \
/bin/promtool query instant http://localhost:9090 'sum(rate(tutela_redis_degraded_control_total[5m])) by (service, operation, mode)'

A degraded rate_limit or quota_* row means enforcement is per-replica rather than shared. Check traffic and spend alongside it and restore Redis promptly.

2. Is Redis reachable, or is it simply not configured?

The AI Proxy degraded-control signal fires both when REDIS_URL is unset and when Redis is down. Distinguish them from the host — the service image has no shell, and the value is a credential, so check for its presence without printing it:

docker inspect tutela-ai-proxy --format '{{range .Config.Env}}{{println .}}{{end}}' \
| grep -q '^REDIS_URL=.' && echo 'REDIS_URL is set' || echo 'REDIS_URL is UNSET'

An empty or unset REDIS_URL means the shared fleet-wide rate and quota ceilings are unavailable. AI Proxy still enforces bounded per-instance limits, but that is a production configuration defect rather than a healthy state. The deployment's compose file makes REDIS_URL mandatory (${REDIS_URL:?REDIS_URL is required}), so an unset value means someone is running a modified configuration.

If it is set, test the connection from the stack's side. On the Compose deployment on AWS, Redis is managed ElastiCache, not a container, so there is no redis container to inspect (on Kubernetes there is a Redis pod — see below); the quickest confirmation that the problem is Redis rather than one service's client is the gateway's readiness body:

curl -s http://127.0.0.1:8080/readyz

It reports {"status":"not_ready","reason":"redis unavailable"} when Redis is configured but unreachable. (The gateway serves port 8080 as plain HTTP in every posture — internal mTLS does not apply to it — so this form needs no certificate even when the rest of the stack runs mTLS.) For the server side on the Compose deployment on AWS, read the ElastiCache console: failover events, CurrConnections, EngineCPUUtilization and evictions over the alert window.

On Kubernetes this differs. The Helm chart runs Redis as a pod in your cluster when redis.enabled is true in your chart values (it is true by default), so the server side is inspectable there:

kubectl -n <namespace> get pods -l app.kubernetes.io/name=redis
kubectl -n <namespace> logs -l app.kubernetes.io/name=redis --tail=200
kubectl -n <namespace> exec <redis-pod> -- redis-cli -a '<redis.auth.password>' info clients

A pod that is Running with a healthy connected_clients count while the gateway reports redis unavailable points at the network path or REDIS_URL, not at the store. If you set redis.enabled: false and pointed REDIS_URL at a store you run yourself, that store's own console is the place to look.

3. Check the correlated Redis health signals.

docker exec tutela-prometheus \
/bin/promtool query instant http://localhost:9090 'sum(rate(tutela_redis_operations_total{status="error"}[5m])) by (service, operation)'

docker exec tutela-prometheus \
/bin/promtool query instant http://localhost:9090 'tutela_redis_connections_active'

Query both bounded degradation signals:

docker exec tutela-prometheus \
/bin/promtool query instant http://localhost:9090 'sum(rate(tutela_gateway_ratelimit_redis_fallback_total[5m])) by (reason)'

docker exec tutela-prometheus \
/bin/promtool query instant http://localhost:9090 'sum(rate(tutela_redis_degraded_control_total[5m])) by (service, operation, mode)'

Real problem or artefact?​

SignalReading
Steady degraded rate on ai-proxy, REDIS_URL unsetConfiguration defect. Limits are only per-instance. Restore the required shared Redis configuration.
Burst during a Redis failover or restartExpected transient. Confirm it drains; record the per-instance enforcement window.
Sustained critical rate with Redis errors risingReal Redis outage. Work the outage; this alert is the consequence.
Fail-opens with no Redis errors at allThe client is not reaching Redis — DNS, TLS, credentials, or an unset URL. Check 2.
mcp-gateway/session_* only, gateway healthySession store degraded. Sessions may not survive a replica hop; rate limiting is unaffected.
Alert clears but you never found a causeDo not close it. A silent bypass window is exactly what this counter exists to make non-silent.

Escalation​

Treat sustained AI Proxy degraded-control activity as a shared-control outage with cost exposure, not merely an infrastructure ticket. Limits remain bounded per replica, but the fleet-wide ceiling is relaxed. Record the window; if spend or abuse is a concern, reconcile usage for that period against ai-spend-operations.md.

Escalate with the operation breakdown, the duration, whether REDIS_URL is configured, and the Redis error rate.

There is no environment variable that disables the bounded AI Proxy fallback. It is the compiled safety behavior whenever shared Redis is absent or unavailable. The two fail-mode settings your deployment does expose in /opt/tutela/.env (FAIL_MODE and POLICY_FAIL_MODE, both read by the inspector) govern inspection — what happens to a request when the classifier or the policy service cannot answer — not Redis, and changing them will not affect this alert.