Low Cache Hit Rate
Target of the LowCacheHitRate alert
(/opt/tutela/deploy/monitoring/alerts/service-alerts.yml).
What fired
(
sum(rate(tutela_cache_hits_total[5m])) by (service, key_prefix)
/
(sum(rate(tutela_cache_hits_total[5m])) by (service, key_prefix)
+ sum(rate(tutela_cache_misses_total[5m])) by (service, key_prefix))
) < 0.8
and
(sum(rate(tutela_cache_hits_total[5m])) by (service, key_prefix)
+ sum(rate(tutela_cache_misses_total[5m])) by (service, key_prefix)) > 0.1
for: 10m, severity warning.
Fewer than 80% of lookups against {{ $labels.service }} /
{{ $labels.key_prefix }} were hits. The second clause is a built-in
noise guard: the rule only fires when the cache is doing more than 0.1
lookups/second, so a nearly idle cache cannot alert on a single miss.
The six prefixes, and what each one means
key_prefix is not decoration — the right response is completely different per
prefix. These are the only six instrumented call sites:
service / key_prefix | Backing store | What a miss costs |
|---|---|---|
policy / active_policies | In-process map, TTL-bounded | A database read for the tenant's active policy set, on the policy-evaluation path. |
ai-proxy / model_pricing | In-process, 5-minute TTL | A pricing refresh. Affects cost attribution, not request success. |
ai-proxy / topic_restrictions | In-process | A policy-service call before completion. |
mcp-gateway / session | Redis | A session lookup falls through; see redis-fail-open.md. |
gateway / sso_state | Redis | Not a cache miss in the usual sense — see below. |
gateway / sso_exchange_grant | Redis | Not a cache miss in the usual sense — see below. |
What this does NOT mean
gateway/sso_stateandgateway/sso_exchange_grantare not caches. Both are single-use, replay-prevention tokens: the SSO callback looks the state up and immediately deletes it, and the client then redeems a one-time code for its session, which is likewise consumed on first use. A "hit" is a valid callback or redemption; a "miss" is one whose token is absent — expired, already used, or never issued. A falling hit rate on either prefix is a security and login-failure signal, not a performance one, and the fix is never "make the cache bigger". Investigate it as failed, replayed or forged SSO logins, and correlate with login failures.- Three of the six are in-process, not Redis.
policy/active_policies,ai-proxy/model_pricingandai-proxy/topic_restrictionslive in the service's own memory. Redis being healthy tells you nothing about them, and every restart empties them. - It does not mean anything failed. A miss is served from the source of truth. Users see latency, not errors.
- It does not mean the cache is broken. For a TTL cache, a hit rate near
1 - (TTL_expiries / lookups)is the design, not a defect. A low-traffic tenant that queries once every ten minutes against a five-minute TTL will miss every single time and is behaving exactly as intended.
First three checks
1. Identify which prefix, and get the absolute rates.
docker exec tutela-prometheus \
/bin/promtool query instant http://localhost:9090 'sum(rate(tutela_cache_hits_total[5m])) by (service, key_prefix)'
docker exec tutela-prometheus \
/bin/promtool query instant http://localhost:9090 'sum(rate(tutela_cache_misses_total[5m])) by (service, key_prefix)'
Branch here. If the prefix is sso_state or sso_exchange_grant, stop
treating this as a cache problem and go to check 3. Otherwise continue to
check 2.
2. Did the process restart, or did traffic change shape?
docker exec tutela-prometheus \
/bin/promtool query instant http://localhost:9090 'time() - process_start_time_seconds'
An in-process cache is empty after a restart and takes a TTL period to warm. A uptime shorter than the alert's 10-minute window explains the whole thing.
For the Redis-backed prefixes, check whether the store itself is failing — misses and errors look similar from here:
docker exec tutela-prometheus \
/bin/promtool query instant http://localhost:9090 'sum(rate(tutela_cache_errors_total[5m])) by (service, operation, error_type)'
Where that store runs depends on the deployment: on the Compose deployment on
AWS it is managed ElastiCache, with nothing on the host to inspect; on
Kubernetes the Helm chart runs Redis as a pod in your cluster when
redis.enabled is true in your chart values (the default), and the pod's own
logs and redis-cli info stats are available to you.
redis-errors.md gives the commands for each.
3. For the two SSO prefixes only — investigate them as failed logins.
docker logs --since=1h tutela-gateway \
| grep -iE 'sso|saml|oidc' | grep -iE 'error|invalid|expired|state'
A burst of misses concentrated in time suggests one broken identity provider configuration or a replay attempt. A steady low rate is usually users abandoning the login flow and returning after the state expired. Correlate with user reports of failed SSO login before escalating as an attack.
Real problem or artefact?
| Signal | Reading |
|---|---|
| Fires within ~10 min of a deploy or restart, in-process prefix | Artefact. Cold cache warming. Wait one TTL. |
| Hit rate steady but low on a low-traffic prefix | Artefact of TTL vs. traffic shape. Expected behaviour. |
Hit rate collapses on policy/active_policies while traffic is flat | Real. Policies are being invalidated or the TTL is being reset — expect policy-evaluation latency to follow. |
Miss rate rises with tutela_cache_errors_total rising | Not a hit-rate problem. The store is failing. Go to redis-errors.md. |
sso_state or sso_exchange_grant hit rate drops | Not a cache problem at all. Failed or replayed SSO logins — check 3. |
| Many distinct tenants appearing at once | Real but benign: more tenants means more first-lookup misses. |
Escalation
For the four genuine caches this is a performance-efficiency signal, and
rarely urgent on its own. It becomes urgent when paired with latency: a
collapsing policy/active_policies hit rate plus HighLatencyPolicy is the
policy path degrading, and that is on the enforcement path.
For the two SSO prefixes escalate to whoever owns authentication, not to whoever owns performance. Include the miss rate, the window, and the gateway log excerpt.
Do not respond to this alert by raising TTLs without understanding the miss
pattern. A longer TTL on active_policies means enforcement runs against staler
policy — trading correctness of governance for a hit-rate number.
Related runbooks
- redis-errors.md — when misses are actually failures.
- redis-latency.md — a slow cache defeats the point of one.
- redis-fail-open.md — what the
mcp-gateway/sessionprefix does when Redis is unreachable. - high-latency.md — the downstream symptom.