Skip to main content

Kafka Consumer Lag

Target of the KafkaConsumerLagWarning and KafkaConsumerLagCritical alerts (/opt/tutela/deploy/monitoring/alerts/service-alerts.yml).

What fired​

tutela_kafka_consumer_lag > 1000 # warning
tutela_kafka_consumer_lag > 10000 # critical

Both for: 5m. Consumer group {{ $labels.consumer_group }} on topic {{ $labels.topic }} is that many messages behind the head of the partition.

Lag is the distance between the partition's high-water mark and the consumer's committed position. Growing lag means the consumer is reading slower than producers are writing; downstream data — analytics, cases, notifications — is stale by that many messages.

Coverage gap — which consumers this alert can actually see​

Read this before triaging. tutela_kafka_consumer_lag is registered in three places with three different label sets, and only one of them ever reports a real value:

RegistrationLabelsBehaviour
classifier feedback consumertopic, partition, consumer_groupReal lag. Computed as high-water mark minus committed position per assigned partition, refreshed on an interval.
analytics consumertopic, consumer_groupAlways 0. The series is pre-created at start-up with a zero value so the label set exists, and nothing ever computes real lag into it.
shared instrumentation libraryservice, topic, partitionNever populated. Nothing in the running services writes to it. It also carries no consumer_group label, so even if something did, these alerts would not group it as expected.

There is no Kafka exporter in any compose stack, and Prometheus scrapes only the Tutela services plus itself — the broker is not a scrape target (/opt/tutela/deploy/monitoring/prometheus.yml).

The consequence, stated plainly:

  • The only consumer that can ever raise this alert is the classifier's classifier-feedback group on tutela.chat.feedback.
  • The main evidence pipeline — analytics-consumer reading tutela.inspection.events, tutela.policy.violations, tutela.discovery.events, tutela.token.usage, tutela.mcp.audit, tutela.runtime.invocations into ClickHouse — reports a permanent zero. If analytics falls a million messages behind, this alert stays silent.
  • notification-consumer and integrations-violation-forwarder are not instrumented for lag at all.

So a quiet KafkaConsumerLag alert is not evidence that consumers are keeping up. For everything except the classifier feedback consumer you must measure lag directly (check 2 below). Closing this gap needs either a broker-side exporter or a real lag computation in the analytics consumer; until then, treat the alert as covering one topic only.

What this does NOT mean​

  • It does not mean messages were lost. Lag is backlog, not loss. The messages are still on the broker and will be processed. Actual loss has different signals — see event-dropped.md and tutela_analytics_dlq_write_errors_total.
  • It does not mean the consumer is down. A stopped consumer shows frozen lag that then grows steadily; a slow consumer shows lag that grows and recedes. A crashed consumer may report no series at all rather than high lag.
  • It does not mean the broker is unhealthy. Lag is a consumer-side property.
  • Silence does not mean zero lag. See the coverage gap above.

First three checks​

1. Read the series that exist, and note which consumer they belong to.

docker exec tutela-prometheus \
/bin/promtool query instant http://localhost:9090 'tutela_kafka_consumer_lag'

If the only non-zero row is a classifier row, that is the one consumer this alert covers. Rows reading 0 from analytics are placeholders, not evidence.

2. Measure real lag at the broker — the authoritative check.

This is the step that covers the consumers the metric does not. Which broker you have depends on the deployment profile — KAFKA_BROKERS in /opt/tutela/.env names it.

In the profile that runs the bundled Redpanda broker (tutela-redpanda is present in docker ps), ask it directly:

docker exec tutela-redpanda \
rpk group describe analytics-consumer

In the production profile the broker is managed Amazon MSK, there is no broker container, and rpk is not installed on the host. Run the same tool from the Redpanda image against the MSK bootstrap brokers (they use TLS):

docker run --rm --network host docker.redpanda.com/redpandadata/redpanda:v24.2.9 \
rpk group describe analytics-consumer \
-X brokers="$(grep '^KAFKA_BROKERS=' /opt/tutela/.env | cut -d= -f2- | tr -d '"')" \
-X tls.enabled=true

Either form prints CURRENT-OFFSET, LOG-END-OFFSET and LAG per partition. The consumer groups to check are analytics-consumer, notification-consumer, integrations-violation-forwarder and classifier-feedback.

3. Decide whether the consumer is stuck or merely slow.

docker logs --tail=200 tutela-analytics

Then look at whether it is still making progress:

docker exec tutela-prometheus \
/bin/promtool query instant http://localhost:9090 'sum(rate(tutela_kafka_messages_processed_total[5m])) by (topic, status)'

A processing rate of zero with growing lag is a stuck consumer — restart it after capturing logs. A healthy processing rate with growing lag is a throughput shortfall: producers are simply outrunning consumers.

Also check whether failures are being diverted rather than processed:

docker exec tutela-prometheus \
/bin/promtool query instant http://localhost:9090 'increase(tutela_analytics_dlq_write_errors_total[1h])'

Real problem or artefact?​

SignalReading
Lag reported only on classifier / tutela.chat.feedbackThe one covered consumer. Real.
Every series reads exactly 0Placeholder values, not evidence of health. Use check 2.
Lag growing steadily, processing rate zeroStuck consumer. Real failure.
Lag growing, processing rate healthyThroughput shortfall. Real, capacity-bound.
Lag spikes then drains after a deployExpected. A restarted consumer rejoins and catches up; look for it to reach zero.
Lag high right after a partition rebalanceTransient. Give it a rebalance interval before acting.

Escalation​

Lag on analytics-consumer is the one that matters most, and it is precisely the one this alert cannot see — a sustained backlog there means the dashboard, cases and compliance exports are stale, and customers are looking at data that does not reflect reality. Verify it by hand with check 2 during any evidence-related investigation rather than relying on the alert.

Escalate with: the per-partition LAG output from check 2, the processing rate from check 3, the consumer's recent logs, and whether the lag is growing, flat or draining.

Restarting a consumer clears a stuck one but does nothing for a throughput shortfall — and it re-triggers a rebalance, briefly making lag worse. Establish which case you are in before restarting.