Kafka Consumer Lag
Target of the KafkaConsumerLagWarning and KafkaConsumerLagCritical alerts
(/opt/tutela/deploy/monitoring/alerts/service-alerts.yml).
What fired
tutela_kafka_consumer_lag > 1000 # warning
tutela_kafka_consumer_lag > 10000 # critical
Both for: 5m. Consumer group {{ $labels.consumer_group }} on topic
{{ $labels.topic }} is that many messages behind the head of the partition.
Lag is the distance between the partition's high-water mark and the consumer's committed position. Growing lag means the consumer is reading slower than producers are writing; downstream data — analytics, cases, notifications — is stale by that many messages.
Coverage gap — which consumers this alert can actually see
Read this before triaging. tutela_kafka_consumer_lag is registered in
three places with three different label sets, and only one of them ever reports
a real value:
| Registration | Labels | Behaviour |
|---|---|---|
| classifier feedback consumer | topic, partition, consumer_group | Real lag. Computed as high-water mark minus committed position per assigned partition, refreshed on an interval. |
| analytics consumer | topic, consumer_group | Always 0. The series is pre-created at start-up with a zero value so the label set exists, and nothing ever computes real lag into it. |
| shared instrumentation library | service, topic, partition | Never populated. Nothing in the running services writes to it. It also carries no consumer_group label, so even if something did, these alerts would not group it as expected. |
There is no Kafka exporter in any compose stack, and Prometheus scrapes only
the Tutela services plus itself — the broker is not a scrape target
(/opt/tutela/deploy/monitoring/prometheus.yml).
The consequence, stated plainly:
- The only consumer that can ever raise this alert is the classifier's
classifier-feedbackgroup ontutela.chat.feedback. - The main evidence pipeline —
analytics-consumerreadingtutela.inspection.events,tutela.policy.violations,tutela.discovery.events,tutela.token.usage,tutela.mcp.audit,tutela.runtime.invocationsinto ClickHouse — reports a permanent zero. If analytics falls a million messages behind, this alert stays silent. notification-consumerandintegrations-violation-forwarderare not instrumented for lag at all.
So a quiet KafkaConsumerLag alert is not evidence that consumers are
keeping up. For everything except the classifier feedback consumer you must
measure lag directly (check 2 below). Closing this gap needs either a broker-side
exporter or a real lag computation in the analytics consumer; until then, treat
the alert as covering one topic only.
What this does NOT mean
- It does not mean messages were lost. Lag is backlog, not loss. The
messages are still on the broker and will be processed. Actual loss has
different signals — see event-dropped.md and
tutela_analytics_dlq_write_errors_total. - It does not mean the consumer is down. A stopped consumer shows frozen lag that then grows steadily; a slow consumer shows lag that grows and recedes. A crashed consumer may report no series at all rather than high lag.
- It does not mean the broker is unhealthy. Lag is a consumer-side property.
- Silence does not mean zero lag. See the coverage gap above.
First three checks
1. Read the series that exist, and note which consumer they belong to.
docker exec tutela-prometheus \
/bin/promtool query instant http://localhost:9090 'tutela_kafka_consumer_lag'
If the only non-zero row is a classifier row, that is the one consumer this
alert covers. Rows reading 0 from analytics are placeholders, not evidence.
2. Measure real lag at the broker — the authoritative check.
This is the step that covers the consumers the metric does not. Which broker
you have depends on the deployment profile — KAFKA_BROKERS in
/opt/tutela/.env names it.
In the profile that runs the bundled Redpanda broker (tutela-redpanda is
present in docker ps), ask it directly:
docker exec tutela-redpanda \
rpk group describe analytics-consumer
In the production profile the broker is managed Amazon MSK, there is no broker
container, and rpk is not installed on the host. Run the same tool from the
Redpanda image against the MSK bootstrap brokers (they use TLS):
docker run --rm --network host docker.redpanda.com/redpandadata/redpanda:v24.2.9 \
rpk group describe analytics-consumer \
-X brokers="$(grep '^KAFKA_BROKERS=' /opt/tutela/.env | cut -d= -f2- | tr -d '"')" \
-X tls.enabled=true
Either form prints CURRENT-OFFSET, LOG-END-OFFSET and LAG per partition.
The consumer groups to check are analytics-consumer, notification-consumer,
integrations-violation-forwarder and classifier-feedback.
3. Decide whether the consumer is stuck or merely slow.
docker logs --tail=200 tutela-analytics
Then look at whether it is still making progress:
docker exec tutela-prometheus \
/bin/promtool query instant http://localhost:9090 'sum(rate(tutela_kafka_messages_processed_total[5m])) by (topic, status)'
A processing rate of zero with growing lag is a stuck consumer — restart it after capturing logs. A healthy processing rate with growing lag is a throughput shortfall: producers are simply outrunning consumers.
Also check whether failures are being diverted rather than processed:
docker exec tutela-prometheus \
/bin/promtool query instant http://localhost:9090 'increase(tutela_analytics_dlq_write_errors_total[1h])'
Real problem or artefact?
| Signal | Reading |
|---|---|
Lag reported only on classifier / tutela.chat.feedback | The one covered consumer. Real. |
Every series reads exactly 0 | Placeholder values, not evidence of health. Use check 2. |
| Lag growing steadily, processing rate zero | Stuck consumer. Real failure. |
| Lag growing, processing rate healthy | Throughput shortfall. Real, capacity-bound. |
| Lag spikes then drains after a deploy | Expected. A restarted consumer rejoins and catches up; look for it to reach zero. |
| Lag high right after a partition rebalance | Transient. Give it a rebalance interval before acting. |
Escalation
Lag on analytics-consumer is the one that matters most, and it is precisely
the one this alert cannot see — a sustained backlog there means the dashboard,
cases and compliance exports are stale, and customers are looking at data
that does not reflect reality. Verify it by hand with check 2 during any
evidence-related investigation rather than relying on the alert.
Escalate with: the per-partition LAG output from check 2, the processing rate
from check 3, the consumer's recent logs, and whether the lag is growing,
flat or draining.
Restarting a consumer clears a stuck one but does nothing for a throughput shortfall — and it re-triggers a rebalance, briefly making lag worse. Establish which case you are in before restarting.
Related runbooks
- event-dropped.md — the producer-side loss this is often confused with.
- service-down.md — a consumer that is not running at all.
- high-latency.md — slow consumers are often slow for the same reason services are.
- disk-space.md — an unbounded backlog is retained on the broker's volume.