Inspector Event Dropped
Target of the InspectorEventDropped alert
(/opt/tutela/deploy/monitoring/alerts/service-alerts.yml).
What fired
increase(tutela_inspector_event_dropped_total[5m]) > 0
for: 1m, severity critical. Any drop at all fires this alert — the
threshold is zero, deliberately.
The counter is incremented at exactly one point in the inspector: the branch taken when a non-blocking hand-off onto the publish queue finds no free slot. The inspector does not wait — it discards the event immediately, carries on serving, and logs an error naming the inspection, the tenant and the action that was decided.
The publish queue holds 5000 in-flight publishes, and reports a high-water warning once it passes 4000 (80%). Drops begin only when all 5000 slots are occupied.
What this does NOT mean
This distinction is the whole point of the runbook:
- It does NOT mean a prompt went ungoverned. The inspection already completed and the enforcement decision — allow, block, redact, coach — was already returned to the caller and applied. What is lost is the event describing that decision.
- It does NOT mean traffic was blocked or degraded. Callers see nothing. Responses are normal, and marginally faster.
- It is NOT a Kafka failure. Kafka may be perfectly healthy. This is
admission control before the publish attempt. Kafka produce failures are
counted separately by
tutela_kafka_publish_errors_total, and events that exhausted their retries land intutela_inspector_event_dlq_total— the DLQ path is recoverable by replay; this one is not. - It is NOT a "timeout", despite the alert description's wording. The 30s timeout in that code path bounds each publish goroutine; the drop itself is instantaneous saturation. Sizing timeouts will not fix it.
What it does mean: audit and analytics evidence has been permanently destroyed. Those inspections will never appear in ClickHouse, in the activity feed, in a case, or in a compliance export. The record of what the product decided is gone, with no replay path.
That is why this is critical at a threshold of one, and why it must never be presented as "governance held". Any compliance or audit report covering this window is incomplete, and should say so.
First three checks
1. How many events, and is it still happening?
docker exec tutela-prometheus \
/bin/promtool query instant http://localhost:9090 'increase(tutela_inspector_event_dropped_total[1h])'
docker exec tutela-prometheus \
/bin/promtool query instant http://localhost:9090 'tutela_inspector_semaphore_high_water_mark'
Read the high-water mark carefully: it is written only while utilisation is at or above the 80% warning threshold, and it never decays. After a burst it stays frozen at the last high value it saw, so a reading of 95 does not by itself mean saturation is happening now. Sample it a few times a minute apart: a value that is still moving means the queue is still under pressure, a value that has not changed is a leftover from the last burst. A gauge that has never been written at all is absent rather than zero.
2. Identify the blast radius — which tenants, which decisions.
Every drop logs inspection_id, tenant_id and action:
docker logs --since=1h tutela-inspector \
| grep 'EVENT DROPPED'
In Loki:
{service="inspector"} |= "EVENT DROPPED"
Record the tenant_id set and the action distribution. A dropped block is
materially worse than a dropped allow: it is an enforcement action with no
evidence.
3. Find what is holding the 5000 slots.
Publishes occupy a slot until they finish or their 30s budget expires, so saturation means the downstream publish path is slow or stuck:
docker exec tutela-prometheus \
/bin/promtool query instant http://localhost:9090 'sum(rate(tutela_inspector_event_publish_retry_total[5m])) by (attempt)'
docker exec tutela-prometheus \
/bin/promtool query instant http://localhost:9090 'increase(tutela_inspector_event_dlq_total[1h])'
Rising retries mean the broker is rejecting or timing out. A rising DLQ counter means retries are being exhausted — those events are recoverable by replay, the dropped ones are not. Then check the broker itself:
docker ps --filter name=tutela-redpanda --format 'table {{.Names}}\t{{.Status}}'
docker logs --tail=200 tutela-redpanda
(In the production profile the broker is managed Amazon MSK and there is no
tutela-redpanda container: read the cluster's CloudWatch metrics and broker
logs instead, and use the broker-side lag check in
kafka-consumer-lag.md.)
Real problem or artefact?
There is no artefact case. The counter is incremented only on a real discard, and it has no labels to misread.
| Signal | Reading |
|---|---|
| Single small burst, high-water back under 80 | A traffic spike outran the publisher. Real evidence loss, bounded. Record the window. |
| Sustained drops, high-water pinned near 100 | The publish path is not keeping up at all. Broker or network problem — check 3. |
| Drops with rising publish retries | Broker degraded. Fix the broker; the semaphore is the symptom. |
| Drops with flat retries and a healthy broker | The inspector is receiving more traffic than 5000 concurrent publishes can absorb. Capacity problem — scale the inspector. |
Drops alongside InspectorFailOpenElevated | Two different failures at once. Fail-open means prompts were not inspected; this means inspections were not recorded. Report both separately — they are not the same gap. |
Escalation
Escalate immediately, and treat it as a data-integrity failure, not a
performance problem. Include the count, the time window, the affected tenant_id
list, and the action distribution from check 2.
The remediation for the loss itself is: none. Dropped events cannot be reconstructed. The only honest action is to record the window so that any audit export, compliance report, or certification evidence covering it is annotated as incomplete.
Preventing recurrence is a capacity question — inspector replicas, publisher throughput, broker health — not a tuning question. Raising the semaphore capacity trades memory for a longer queue; it does not create publish throughput, and if the broker is the bottleneck it only delays the same drop.
Related runbooks
- kafka-consumer-lag.md — the downstream half of this pipeline, and its own measurement gap.
- high-latency.md — a saturating publish path usually shows as latency first.
- inspector-failclosed-triage.md — the other inspector failure mode, which affects enforcement rather than evidence.