Skip to main content

Inspector Event Dropped

Target of the InspectorEventDropped alert (/opt/tutela/deploy/monitoring/alerts/service-alerts.yml).

What fired​

increase(tutela_inspector_event_dropped_total[5m]) > 0

for: 1m, severity critical. Any drop at all fires this alert — the threshold is zero, deliberately.

The counter is incremented at exactly one point in the inspector: the branch taken when a non-blocking hand-off onto the publish queue finds no free slot. The inspector does not wait — it discards the event immediately, carries on serving, and logs an error naming the inspection, the tenant and the action that was decided.

The publish queue holds 5000 in-flight publishes, and reports a high-water warning once it passes 4000 (80%). Drops begin only when all 5000 slots are occupied.

What this does NOT mean​

This distinction is the whole point of the runbook:

  • It does NOT mean a prompt went ungoverned. The inspection already completed and the enforcement decision — allow, block, redact, coach — was already returned to the caller and applied. What is lost is the event describing that decision.
  • It does NOT mean traffic was blocked or degraded. Callers see nothing. Responses are normal, and marginally faster.
  • It is NOT a Kafka failure. Kafka may be perfectly healthy. This is admission control before the publish attempt. Kafka produce failures are counted separately by tutela_kafka_publish_errors_total, and events that exhausted their retries land in tutela_inspector_event_dlq_total — the DLQ path is recoverable by replay; this one is not.
  • It is NOT a "timeout", despite the alert description's wording. The 30s timeout in that code path bounds each publish goroutine; the drop itself is instantaneous saturation. Sizing timeouts will not fix it.

What it does mean: audit and analytics evidence has been permanently destroyed. Those inspections will never appear in ClickHouse, in the activity feed, in a case, or in a compliance export. The record of what the product decided is gone, with no replay path.

That is why this is critical at a threshold of one, and why it must never be presented as "governance held". Any compliance or audit report covering this window is incomplete, and should say so.

First three checks​

1. How many events, and is it still happening?

docker exec tutela-prometheus \
/bin/promtool query instant http://localhost:9090 'increase(tutela_inspector_event_dropped_total[1h])'

docker exec tutela-prometheus \
/bin/promtool query instant http://localhost:9090 'tutela_inspector_semaphore_high_water_mark'

Read the high-water mark carefully: it is written only while utilisation is at or above the 80% warning threshold, and it never decays. After a burst it stays frozen at the last high value it saw, so a reading of 95 does not by itself mean saturation is happening now. Sample it a few times a minute apart: a value that is still moving means the queue is still under pressure, a value that has not changed is a leftover from the last burst. A gauge that has never been written at all is absent rather than zero.

2. Identify the blast radius — which tenants, which decisions.

Every drop logs inspection_id, tenant_id and action:

docker logs --since=1h tutela-inspector \
| grep 'EVENT DROPPED'

In Loki:

{service="inspector"} |= "EVENT DROPPED"

Record the tenant_id set and the action distribution. A dropped block is materially worse than a dropped allow: it is an enforcement action with no evidence.

3. Find what is holding the 5000 slots.

Publishes occupy a slot until they finish or their 30s budget expires, so saturation means the downstream publish path is slow or stuck:

docker exec tutela-prometheus \
/bin/promtool query instant http://localhost:9090 'sum(rate(tutela_inspector_event_publish_retry_total[5m])) by (attempt)'

docker exec tutela-prometheus \
/bin/promtool query instant http://localhost:9090 'increase(tutela_inspector_event_dlq_total[1h])'

Rising retries mean the broker is rejecting or timing out. A rising DLQ counter means retries are being exhausted — those events are recoverable by replay, the dropped ones are not. Then check the broker itself:

docker ps --filter name=tutela-redpanda --format 'table {{.Names}}\t{{.Status}}'
docker logs --tail=200 tutela-redpanda

(In the production profile the broker is managed Amazon MSK and there is no tutela-redpanda container: read the cluster's CloudWatch metrics and broker logs instead, and use the broker-side lag check in kafka-consumer-lag.md.)

Real problem or artefact?​

There is no artefact case. The counter is incremented only on a real discard, and it has no labels to misread.

SignalReading
Single small burst, high-water back under 80A traffic spike outran the publisher. Real evidence loss, bounded. Record the window.
Sustained drops, high-water pinned near 100The publish path is not keeping up at all. Broker or network problem — check 3.
Drops with rising publish retriesBroker degraded. Fix the broker; the semaphore is the symptom.
Drops with flat retries and a healthy brokerThe inspector is receiving more traffic than 5000 concurrent publishes can absorb. Capacity problem — scale the inspector.
Drops alongside InspectorFailOpenElevatedTwo different failures at once. Fail-open means prompts were not inspected; this means inspections were not recorded. Report both separately — they are not the same gap.

Escalation​

Escalate immediately, and treat it as a data-integrity failure, not a performance problem. Include the count, the time window, the affected tenant_id list, and the action distribution from check 2.

The remediation for the loss itself is: none. Dropped events cannot be reconstructed. The only honest action is to record the window so that any audit export, compliance report, or certification evidence covering it is annotated as incomplete.

Preventing recurrence is a capacity question — inspector replicas, publisher throughput, broker health — not a tuning question. Raising the semaphore capacity trades memory for a longer queue; it does not create publish throughput, and if the broker is the bottleneck it only delays the same drop.