Skip to main content

Disk Space Running Low

Target of the DiskSpaceRunningLowWarning and DiskSpaceRunningLowCritical alerts (/opt/tutela/deploy/monitoring/alerts/service-alerts.yml).

What fired​

(node_filesystem_avail_bytes{mountpoint="/"} / node_filesystem_size_bytes{mountpoint="/"}) * 100 < 20 # warning
(node_filesystem_avail_bytes{mountpoint="/"} / node_filesystem_size_bytes{mountpoint="/"}) * 100 < 10 # critical

Both for: 5m. Intended meaning: the root filesystem on {{ $labels.instance }} is below 20% (warning) or 10% (critical) free.

Coverage gap — these alerts cannot fire in the Compose deployment​

node_filesystem_avail_bytes and node_filesystem_size_bytes are node exporter metrics, and there is no node exporter in the Compose deployment. Prometheus scrapes only the twelve Tutela service targets plus itself (/opt/tutela/deploy/monitoring/prometheus.yml); no infrastructure exporter of any kind is present in either compose stack.

With no series, the expression yields an empty vector on every evaluation and the alert never fires. Confirm:

docker exec tutela-prometheus \
/bin/promtool query instant http://localhost:9090 'count(node_filesystem_avail_bytes) or vector(0)'

{} => 0 means no series exist: nothing in this deployment produces the metric (it comes from the node exporter, which is not part of the stack), so the alert is inert.

The count(...) or vector(0) wrapper is deliberate. A bare metric name prints NOTHING when the series is absent -- and so does a command that never reached Prometheus, which would let a broken command read as proof of the gap. This form prints three distinguishable things: {} => 0 (no producer, the documented gap), {} => N (a producer exists, so investigate for real), or a loud query error: ... (the command itself did not run).

There is no disk alerting in the Compose deployment. On Kubernetes this may not apply. The Helm chart ships these rules as a PrometheusRule only when monitoring.enabled is true in your chart values (it is false by default), and ships no exporters: your cluster's Prometheus supplies them, and node-exporter is commonly present in a standard cluster monitoring install. Run the confirmation query above against your own Prometheus before concluding anything — {} => N means your cluster does produce the metric and the alert is live for you.

In the Compose deployment, disk exhaustion will present as services failing — Postgres refusing writes, ClickHouse rejecting inserts, the broker refusing produces — rather than as this alert. Check disk by hand during any unexplained write-failure investigation.

The {{ $labels.instance }} in the alert text is also misleading: even with a node exporter, instance would identify the exporter target, and these are the only alerts in the ruleset that assume a host-level target exists at all.

What actually fills up​

The volumes that grow, from /opt/tutela/deploy/docker-compose.aws.yaml:

VolumeMountWhat grows it
clickhouse-data/var/lib/clickhouseInspection events and analytics. Usually the largest by far. TTL is configurable per customer.
redpanda-data/var/lib/redpanda/dataThe event log. Grows without bound if a consumer stalls — see kafka-consumer-lag.md.
prometheus-data/prometheusMetrics. Bounded by --storage.tsdb.retention.time, default 15d.
loki-data/lokiLogs. retention_period: 168h with retention_enabled: true in /opt/tutela/deploy/monitoring/loki/config.yaml.
clamav-data/var/lib/clamavMalware signature database. Grows slowly, bounded.
grafana-data, alertmanager-data, promtail-positions—Small; rarely the cause.

Postgres and Redis are not containers — they are managed RDS and ElastiCache, with their own storage and their own CloudWatch alarms. Their disk is not visible from this host at all. In the production profile the broker is managed MSK as well, so redpanda-data exists only in the profile that runs the bundled broker.

What this does NOT mean​

  • Silence does not mean there is space. Per the gap above, silence is the only possible state.
  • It would not mean the right filesystem. The rule pins mountpoint="/". Docker volumes usually live under /var/lib/docker/volumes, which may be a separate filesystem; a full data volume with a healthy root would not have fired this alert even with an exporter present.
  • A full disk is not a slow disk. Latency and capacity are unrelated problems.

First three checks​

1. Measure the host and the volumes directly.

df -h /
df -h /var/lib/docker
docker system df -v

docker system df -v lists per-volume sizes and is the fastest way to find which of the volumes above is responsible.

2. Check the two that grow fastest.

ClickHouse — find the largest tables and confirm TTL is doing its job:

docker exec tutela-clickhouse \
clickhouse-client --query "
SELECT database, table, formatReadableSize(sum(bytes_on_disk)) AS size
FROM system.parts WHERE active
GROUP BY database, table ORDER BY sum(bytes_on_disk) DESC LIMIT 10"

The broker — a stalled consumer means retained data grows without bound:

docker exec tutela-redpanda \
rpk group describe analytics-consumer

Large, growing lag here is a disk problem waiting to happen. Work kafka-consumer-lag.md first — draining the backlog is usually what reclaims the space.

3. Confirm retention is actually configured as intended.

docker exec tutela-prometheus \
wget -qO- 'http://localhost:9090/api/v1/status/runtimeinfo' \
| jq -r '.data.storageRetention'

grep -n 'retention' /opt/tutela/deploy/monitoring/loki/config.yaml

If someone raised PROMETHEUS_RETENTION_TIME — for example to make the 30-day window in error-budget.md honest — Prometheus storage grows proportionally. That is a common cause of a slow squeeze.

Real problem or artefact?​

SignalReading
This alert firingUnexpected. With no node exporter it should be impossible — find out what is exporting node_filesystem_*.
df -h / above 90%Real, and urgent. Databases fail hard, not gracefully, when writes cannot land.
One volume dominating in docker system df -vReal. Address that service's retention rather than adding disk.
Broker volume growing with consumer lagSymptom, not cause. Fix the consumer.
Steady growth tracking customer trafficExpected. Plan capacity or shorten TTL.
Sudden jump with no traffic changeLook for a log loop — a service erroring at high rate fills Loki quickly.

Escalation​

The primary escalation is the gap: this deployment has no disk alerting. Closing it requires a node exporter (or host-level monitoring outside this stack), which is a monitoring-configuration change, not an on-call fix.

For an actual squeeze, escalate with df -h output, the per-volume breakdown, the ClickHouse table sizes, and current retention settings.

Deleting data to recover space destroys audit evidence. ClickHouse holds the inspection record that compliance exports read from — shortening its TTL is a retention-policy decision with customer-facing consequences, not a cleanup action. Reclaim space from Prometheus, Loki or the broker backlog first.