Disk Space Running Low
Target of the DiskSpaceRunningLowWarning and DiskSpaceRunningLowCritical
alerts (/opt/tutela/deploy/monitoring/alerts/service-alerts.yml).
What fired
(node_filesystem_avail_bytes{mountpoint="/"} / node_filesystem_size_bytes{mountpoint="/"}) * 100 < 20 # warning
(node_filesystem_avail_bytes{mountpoint="/"} / node_filesystem_size_bytes{mountpoint="/"}) * 100 < 10 # critical
Both for: 5m. Intended meaning: the root filesystem on {{ $labels.instance }}
is below 20% (warning) or 10% (critical) free.
Coverage gap — these alerts cannot fire in the Compose deployment
node_filesystem_avail_bytes and node_filesystem_size_bytes are node
exporter metrics, and there is no node exporter in the Compose deployment. Prometheus
scrapes only the twelve Tutela service targets plus itself
(/opt/tutela/deploy/monitoring/prometheus.yml); no infrastructure exporter of any kind is
present in either compose stack.
With no series, the expression yields an empty vector on every evaluation and the alert never fires. Confirm:
docker exec tutela-prometheus \
/bin/promtool query instant http://localhost:9090 'count(node_filesystem_avail_bytes) or vector(0)'
{} => 0 means no series exist: nothing in this deployment produces the
metric (it comes from the node exporter, which is not part of the stack), so the
alert is inert.
The count(...) or vector(0) wrapper is deliberate. A bare metric name
prints NOTHING when the series is absent -- and so does a command that never
reached Prometheus, which would let a broken command read as proof of the
gap. This form prints three distinguishable things: {} => 0 (no producer,
the documented gap), {} => N (a producer exists, so investigate for real),
or a loud query error: ... (the command itself did not run).
There is no disk alerting in the Compose deployment.
On Kubernetes this may not apply. The Helm chart ships these rules as a
PrometheusRule only when monitoring.enabled is true in your chart values
(it is false by default), and ships no exporters: your cluster's Prometheus supplies
them, and node-exporter is commonly present in a standard cluster monitoring
install. Run the confirmation query above against your own Prometheus before
concluding anything — {} => N means your cluster does produce the metric and
the alert is live for you.
In the Compose deployment, disk exhaustion will present as services failing — Postgres refusing writes, ClickHouse rejecting inserts, the broker refusing produces — rather than as this alert. Check disk by hand during any unexplained write-failure investigation.
The {{ $labels.instance }} in the alert text is also misleading: even with a
node exporter, instance would identify the exporter target, and these are the
only alerts in the ruleset that assume a host-level target exists at all.
What actually fills up
The volumes that grow, from /opt/tutela/deploy/docker-compose.aws.yaml:
| Volume | Mount | What grows it |
|---|---|---|
clickhouse-data | /var/lib/clickhouse | Inspection events and analytics. Usually the largest by far. TTL is configurable per customer. |
redpanda-data | /var/lib/redpanda/data | The event log. Grows without bound if a consumer stalls — see kafka-consumer-lag.md. |
prometheus-data | /prometheus | Metrics. Bounded by --storage.tsdb.retention.time, default 15d. |
loki-data | /loki | Logs. retention_period: 168h with retention_enabled: true in /opt/tutela/deploy/monitoring/loki/config.yaml. |
clamav-data | /var/lib/clamav | Malware signature database. Grows slowly, bounded. |
grafana-data, alertmanager-data, promtail-positions | — | Small; rarely the cause. |
Postgres and Redis are not containers — they are managed RDS and
ElastiCache, with their own storage and their own CloudWatch alarms. Their disk
is not visible from this host at all. In the production profile the broker is
managed MSK as well, so redpanda-data exists only in the profile that runs
the bundled broker.
What this does NOT mean
- Silence does not mean there is space. Per the gap above, silence is the only possible state.
- It would not mean the right filesystem. The rule pins
mountpoint="/". Docker volumes usually live under/var/lib/docker/volumes, which may be a separate filesystem; a full data volume with a healthy root would not have fired this alert even with an exporter present. - A full disk is not a slow disk. Latency and capacity are unrelated problems.
First three checks
1. Measure the host and the volumes directly.
df -h /
df -h /var/lib/docker
docker system df -v
docker system df -v lists per-volume sizes and is the fastest way to find
which of the volumes above is responsible.
2. Check the two that grow fastest.
ClickHouse — find the largest tables and confirm TTL is doing its job:
docker exec tutela-clickhouse \
clickhouse-client --query "
SELECT database, table, formatReadableSize(sum(bytes_on_disk)) AS size
FROM system.parts WHERE active
GROUP BY database, table ORDER BY sum(bytes_on_disk) DESC LIMIT 10"
The broker — a stalled consumer means retained data grows without bound:
docker exec tutela-redpanda \
rpk group describe analytics-consumer
Large, growing lag here is a disk problem waiting to happen. Work kafka-consumer-lag.md first — draining the backlog is usually what reclaims the space.
3. Confirm retention is actually configured as intended.
docker exec tutela-prometheus \
wget -qO- 'http://localhost:9090/api/v1/status/runtimeinfo' \
| jq -r '.data.storageRetention'
grep -n 'retention' /opt/tutela/deploy/monitoring/loki/config.yaml
If someone raised PROMETHEUS_RETENTION_TIME — for example to make the 30-day
window in error-budget.md honest — Prometheus storage grows
proportionally. That is a common cause of a slow squeeze.
Real problem or artefact?
| Signal | Reading |
|---|---|
| This alert firing | Unexpected. With no node exporter it should be impossible — find out what is exporting node_filesystem_*. |
df -h / above 90% | Real, and urgent. Databases fail hard, not gracefully, when writes cannot land. |
One volume dominating in docker system df -v | Real. Address that service's retention rather than adding disk. |
| Broker volume growing with consumer lag | Symptom, not cause. Fix the consumer. |
| Steady growth tracking customer traffic | Expected. Plan capacity or shorten TTL. |
| Sudden jump with no traffic change | Look for a log loop — a service erroring at high rate fills Loki quickly. |
Escalation
The primary escalation is the gap: this deployment has no disk alerting. Closing it requires a node exporter (or host-level monitoring outside this stack), which is a monitoring-configuration change, not an on-call fix.
For an actual squeeze, escalate with df -h output, the per-volume breakdown,
the ClickHouse table sizes, and current retention settings.
Deleting data to recover space destroys audit evidence. ClickHouse holds the inspection record that compliance exports read from — shortening its TTL is a retention-policy decision with customer-facing consequences, not a cleanup action. Reclaim space from Prometheus, Loki or the broker backlog first.
Related runbooks
- kafka-consumer-lag.md — the most common cause of unbounded broker growth.
- error-budget.md — raising Prometheus retention costs disk.
- high-memory.md and high-cpu.md — the other two resource alerts, with the same class of gap.
- service-down.md — how a full disk actually announces itself.