Skip to main content

Low Cache Hit Rate

Target of the LowCacheHitRate alert (/opt/tutela/deploy/monitoring/alerts/service-alerts.yml).

What fired​

(
sum(rate(tutela_cache_hits_total[5m])) by (service, key_prefix)
/
(sum(rate(tutela_cache_hits_total[5m])) by (service, key_prefix)
+ sum(rate(tutela_cache_misses_total[5m])) by (service, key_prefix))
) < 0.8
and
(sum(rate(tutela_cache_hits_total[5m])) by (service, key_prefix)
+ sum(rate(tutela_cache_misses_total[5m])) by (service, key_prefix)) > 0.1

for: 10m, severity warning.

Fewer than 80% of lookups against {{ $labels.service }} / {{ $labels.key_prefix }} were hits. The second clause is a built-in noise guard: the rule only fires when the cache is doing more than 0.1 lookups/second, so a nearly idle cache cannot alert on a single miss.

The six prefixes, and what each one means​

key_prefix is not decoration — the right response is completely different per prefix. These are the only six instrumented call sites:

service / key_prefixBacking storeWhat a miss costs
policy / active_policiesIn-process map, TTL-boundedA database read for the tenant's active policy set, on the policy-evaluation path.
ai-proxy / model_pricingIn-process, 5-minute TTLA pricing refresh. Affects cost attribution, not request success.
ai-proxy / topic_restrictionsIn-processA policy-service call before completion.
mcp-gateway / sessionRedisA session lookup falls through; see redis-fail-open.md.
gateway / sso_stateRedisNot a cache miss in the usual sense — see below.
gateway / sso_exchange_grantRedisNot a cache miss in the usual sense — see below.

What this does NOT mean​

  • gateway / sso_state and gateway / sso_exchange_grant are not caches. Both are single-use, replay-prevention tokens: the SSO callback looks the state up and immediately deletes it, and the client then redeems a one-time code for its session, which is likewise consumed on first use. A "hit" is a valid callback or redemption; a "miss" is one whose token is absent — expired, already used, or never issued. A falling hit rate on either prefix is a security and login-failure signal, not a performance one, and the fix is never "make the cache bigger". Investigate it as failed, replayed or forged SSO logins, and correlate with login failures.
  • Three of the six are in-process, not Redis. policy/active_policies, ai-proxy/model_pricing and ai-proxy/topic_restrictions live in the service's own memory. Redis being healthy tells you nothing about them, and every restart empties them.
  • It does not mean anything failed. A miss is served from the source of truth. Users see latency, not errors.
  • It does not mean the cache is broken. For a TTL cache, a hit rate near 1 - (TTL_expiries / lookups) is the design, not a defect. A low-traffic tenant that queries once every ten minutes against a five-minute TTL will miss every single time and is behaving exactly as intended.

First three checks​

1. Identify which prefix, and get the absolute rates.

docker exec tutela-prometheus \
/bin/promtool query instant http://localhost:9090 'sum(rate(tutela_cache_hits_total[5m])) by (service, key_prefix)'

docker exec tutela-prometheus \
/bin/promtool query instant http://localhost:9090 'sum(rate(tutela_cache_misses_total[5m])) by (service, key_prefix)'

Branch here. If the prefix is sso_state or sso_exchange_grant, stop treating this as a cache problem and go to check 3. Otherwise continue to check 2.

2. Did the process restart, or did traffic change shape?

docker exec tutela-prometheus \
/bin/promtool query instant http://localhost:9090 'time() - process_start_time_seconds'

An in-process cache is empty after a restart and takes a TTL period to warm. A uptime shorter than the alert's 10-minute window explains the whole thing.

For the Redis-backed prefixes, check whether the store itself is failing — misses and errors look similar from here:

docker exec tutela-prometheus \
/bin/promtool query instant http://localhost:9090 'sum(rate(tutela_cache_errors_total[5m])) by (service, operation, error_type)'

Where that store runs depends on the deployment: on the Compose deployment on AWS it is managed ElastiCache, with nothing on the host to inspect; on Kubernetes the Helm chart runs Redis as a pod in your cluster when redis.enabled is true in your chart values (the default), and the pod's own logs and redis-cli info stats are available to you. redis-errors.md gives the commands for each.

3. For the two SSO prefixes only — investigate them as failed logins.

docker logs --since=1h tutela-gateway \
| grep -iE 'sso|saml|oidc' | grep -iE 'error|invalid|expired|state'

A burst of misses concentrated in time suggests one broken identity provider configuration or a replay attempt. A steady low rate is usually users abandoning the login flow and returning after the state expired. Correlate with user reports of failed SSO login before escalating as an attack.

Real problem or artefact?​

SignalReading
Fires within ~10 min of a deploy or restart, in-process prefixArtefact. Cold cache warming. Wait one TTL.
Hit rate steady but low on a low-traffic prefixArtefact of TTL vs. traffic shape. Expected behaviour.
Hit rate collapses on policy/active_policies while traffic is flatReal. Policies are being invalidated or the TTL is being reset — expect policy-evaluation latency to follow.
Miss rate rises with tutela_cache_errors_total risingNot a hit-rate problem. The store is failing. Go to redis-errors.md.
sso_state or sso_exchange_grant hit rate dropsNot a cache problem at all. Failed or replayed SSO logins — check 3.
Many distinct tenants appearing at onceReal but benign: more tenants means more first-lookup misses.

Escalation​

For the four genuine caches this is a performance-efficiency signal, and rarely urgent on its own. It becomes urgent when paired with latency: a collapsing policy/active_policies hit rate plus HighLatencyPolicy is the policy path degrading, and that is on the enforcement path.

For the two SSO prefixes escalate to whoever owns authentication, not to whoever owns performance. Include the miss rate, the window, and the gateway log excerpt.

Do not respond to this alert by raising TTLs without understanding the miss pattern. A longer TTL on active_policies means enforcement runs against staler policy — trading correctness of governance for a hit-rate number.