High Memory Usage
Target of the HighMemoryUsage alert
(/opt/tutela/deploy/monitoring/alerts/service-alerts.yml).
What fired
(
process_resident_memory_bytes
/
machine_memory_bytes
) * 100 > 80
for: 5m, severity warning.
Intended meaning: the service's resident set exceeds 80% of the machine's total memory.
Coverage gap — this alert cannot fire in the Compose deployment
machine_memory_bytes has no producer in the Compose deployment. It is a cAdvisor
metric, and there is no cAdvisor, node exporter, or any other
infrastructure exporter in either compose stack. Prometheus scrapes only the
twelve Tutela service targets plus itself
(/opt/tutela/deploy/monitoring/prometheus.yml).
In PromQL, dividing by a metric that has no series produces an empty vector. An empty vector never exceeds a threshold, so the rule evaluates to nothing on every cycle and the alert can never fire — regardless of how much memory any service uses.
Confirm it directly:
docker exec tutela-prometheus \
/bin/promtool query instant http://localhost:9090 'count(machine_memory_bytes) or vector(0)'
{} => 0 means no series exist: nothing in this deployment produces the
metric (it comes from cAdvisor, which is not part of the stack), so the
alert is inert.
The count(...) or vector(0) wrapper is deliberate. A bare metric name
prints NOTHING when the series is absent -- and so does a command that never
reached Prometheus, which would let a broken command read as proof of the
gap. This form prints three distinguishable things: {} => 0 (no producer,
the documented gap), {} => N (a producer exists, so investigate for real),
or a loud query error: ... (the command itself did not run).
The numerator is fine — process_resident_memory_bytes comes free from the Go
client's default collector, so twelve of the thirteen scrape jobs report it
(the eleven Go services and Prometheus itself). The classifier does not: it runs in Prometheus multiprocess mode
(PROMETHEUS_MULTIPROC_DIR in /opt/tutela/deploy/docker-compose.aws.yaml), where the
default process collector is not registered.
There is no memory alerting in the Compose deployment. Treat that as the
finding: anyone relying on this alert there to catch a leak is relying on
nothing.
On Kubernetes this may not apply. The Helm chart ships these rules as a
PrometheusRule only when monitoring.enabled is true in your chart values
(it is false by default), and ships no exporters: your cluster's Prometheus supplies
them, and cAdvisor (via the kubelet) is commonly present in a standard cluster monitoring
install. Run the confirmation query above against your own Prometheus before
concluding anything — {} => N means your cluster does produce the metric and
the alert is live for you.
The comparison is also the wrong one
Even with cAdvisor present, comparing a process's RSS to the machine's total
memory would be the wrong denominator here. Every Tutela container carries an
explicit memory limit in /opt/tutela/deploy/docker-compose.aws.yaml:
| Container | Limit | Container | Limit |
|---|---|---|---|
| the eleven Go services | 512M each | web | 256M |
classifier | ${CLASSIFIER_MEMORY_LIMIT:-4G} | clickhouse | 2G |
redpanda | 1536M | clamav | 2G |
prometheus, loki, grafana | 512M each | alertmanager, promtail | 128M each |
A Go service is OOM-killed at 512 MB, which on a 16 GB host is 3% of machine memory — nowhere near the 80% threshold. The limit is what kills the process, so the limit is what should be measured against.
What to use instead
docker stats --no-stream \
--format 'table {{.Name}}\t{{.MemUsage}}\t{{.MemPerc}}\t{{.CPUPerc}}'
MemPerc is memory as a percentage of the container's limit — the number
that predicts an OOM kill. Above ~85% sustained, the container is at risk.
For trend rather than a snapshot, the RSS series is still useful even without a denominator:
docker exec tutela-prometheus \
/bin/promtool query instant http://localhost:9090 'process_resident_memory_bytes / 1024 / 1024'
Compare that against the 512 MB limit yourself.
What this does NOT mean
- Silence does not mean healthy. Per the gap above, silence is the only possible state. Absence of this alert is not evidence about memory.
- It would not mean a leak even if it fired. High steady memory is a cache or a large working set. A leak is memory that climbs and never returns — identifiable only from a trend, not a threshold.
- It does not detect the failure that actually happens. The real failure mode is the OOM killer terminating the container. That surfaces as service-down.md and a restart, not as this alert.
First three checks
1. Find containers near their limit.
docker stats --no-stream \
--format 'table {{.Name}}\t{{.MemUsage}}\t{{.MemPerc}}'
2. Check whether anything has already been OOM-killed.
docker ps --filter name=tutela- --format 'table {{.Names}}\t{{.Status}}'
docker inspect tutela-<service> \
--format '{{.State.OOMKilled}} {{.State.ExitCode}} {{.RestartCount}}'
OOMKilled=true, or ExitCode 137, is the memory limit being enforced. A
climbing RestartCount with no other explanation is usually this.
3. Distinguish a leak from a working set — look at the trend, not the value.
Grafana is the only monitoring surface this deployment publishes (Prometheus
publishes no port), and a trend over time is what it is for. Open it from the
console under Settings → Diagnostics, or at
https://<your-console-host>/grafana, choose the Prometheus datasource, and run
this over the window named below:
process_resident_memory_bytes{job="<job>"} / 1024 / 1024
Window: last 24 hours.
Sawtooth that returns to a baseline after GC is a working set. A monotonic climb across restarts-free hours is a leak. For Go services, confirm with the runtime series:
docker exec tutela-prometheus \
/bin/promtool query instant http://localhost:9090 'go_memstats_heap_inuse_bytes / 1024 / 1024'
Heap growing while RSS grows is an application leak. RSS growing with a flat
heap is off-heap — cgo, mmap, or goroutine stacks; check
go_goroutines for an unbounded goroutine count.
Real problem or artefact?
| Signal | Reading |
|---|---|
| This alert firing | Investigate the monitoring stack — with no machine_memory_bytes producer it should be impossible. Something non-standard is exporting that series. |
MemPerc above 85% sustained | Real risk. An OOM kill is coming. |
OOMKilled=true in docker inspect | Already happened. Restart hid it — check RestartCount. |
| Memory high and flat since startup | Working set. Not a leak. |
| Memory climbing monotonically over hours | Leak. Capture a heap profile before restarting. |
Memory high on classifier | Expected — it has a 4G limit and loads models. Compare to its own limit, not the others'. |
Escalation
The primary escalation is the gap itself: this deployment has no working
memory alert. Closing it means either adding a container-metrics exporter or
rewriting the rule against process_resident_memory_bytes and the known
per-container limit. Both change what the alert means, so both belong to the
owner of the monitoring configuration rather than to an on-call fix.
For a real memory problem found through docker stats, escalate with MemPerc,
the 24-hour trend from check 3, the OOMKilled/RestartCount values, and the
container's configured limit. Restarting clears the symptom and destroys the
evidence — capture the trend first.
Related runbooks
- high-cpu.md — the sibling resource alert, with its own ceiling problem.
- disk-space.md — the third resource alert with no producer.
- service-down.md — how an OOM kill actually surfaces.
- high-latency.md — GC pressure from memory shows up as latency first.