Skip to main content

High Memory Usage

Target of the HighMemoryUsage alert (/opt/tutela/deploy/monitoring/alerts/service-alerts.yml).

What fired​

(
process_resident_memory_bytes
/
machine_memory_bytes
) * 100 > 80

for: 5m, severity warning.

Intended meaning: the service's resident set exceeds 80% of the machine's total memory.

Coverage gap — this alert cannot fire in the Compose deployment​

machine_memory_bytes has no producer in the Compose deployment. It is a cAdvisor metric, and there is no cAdvisor, node exporter, or any other infrastructure exporter in either compose stack. Prometheus scrapes only the twelve Tutela service targets plus itself (/opt/tutela/deploy/monitoring/prometheus.yml).

In PromQL, dividing by a metric that has no series produces an empty vector. An empty vector never exceeds a threshold, so the rule evaluates to nothing on every cycle and the alert can never fire — regardless of how much memory any service uses.

Confirm it directly:

docker exec tutela-prometheus \
/bin/promtool query instant http://localhost:9090 'count(machine_memory_bytes) or vector(0)'

{} => 0 means no series exist: nothing in this deployment produces the metric (it comes from cAdvisor, which is not part of the stack), so the alert is inert.

The count(...) or vector(0) wrapper is deliberate. A bare metric name prints NOTHING when the series is absent -- and so does a command that never reached Prometheus, which would let a broken command read as proof of the gap. This form prints three distinguishable things: {} => 0 (no producer, the documented gap), {} => N (a producer exists, so investigate for real), or a loud query error: ... (the command itself did not run).

The numerator is fine — process_resident_memory_bytes comes free from the Go client's default collector, so twelve of the thirteen scrape jobs report it (the eleven Go services and Prometheus itself). The classifier does not: it runs in Prometheus multiprocess mode (PROMETHEUS_MULTIPROC_DIR in /opt/tutela/deploy/docker-compose.aws.yaml), where the default process collector is not registered.

There is no memory alerting in the Compose deployment. Treat that as the finding: anyone relying on this alert there to catch a leak is relying on nothing. On Kubernetes this may not apply. The Helm chart ships these rules as a PrometheusRule only when monitoring.enabled is true in your chart values (it is false by default), and ships no exporters: your cluster's Prometheus supplies them, and cAdvisor (via the kubelet) is commonly present in a standard cluster monitoring install. Run the confirmation query above against your own Prometheus before concluding anything — {} => N means your cluster does produce the metric and the alert is live for you.

The comparison is also the wrong one​

Even with cAdvisor present, comparing a process's RSS to the machine's total memory would be the wrong denominator here. Every Tutela container carries an explicit memory limit in /opt/tutela/deploy/docker-compose.aws.yaml:

ContainerLimitContainerLimit
the eleven Go services512M eachweb256M
classifier${CLASSIFIER_MEMORY_LIMIT:-4G}clickhouse2G
redpanda1536Mclamav2G
prometheus, loki, grafana512M eachalertmanager, promtail128M each

A Go service is OOM-killed at 512 MB, which on a 16 GB host is 3% of machine memory — nowhere near the 80% threshold. The limit is what kills the process, so the limit is what should be measured against.

What to use instead​

docker stats --no-stream \
--format 'table {{.Name}}\t{{.MemUsage}}\t{{.MemPerc}}\t{{.CPUPerc}}'

MemPerc is memory as a percentage of the container's limit — the number that predicts an OOM kill. Above ~85% sustained, the container is at risk.

For trend rather than a snapshot, the RSS series is still useful even without a denominator:

docker exec tutela-prometheus \
/bin/promtool query instant http://localhost:9090 'process_resident_memory_bytes / 1024 / 1024'

Compare that against the 512 MB limit yourself.

What this does NOT mean​

  • Silence does not mean healthy. Per the gap above, silence is the only possible state. Absence of this alert is not evidence about memory.
  • It would not mean a leak even if it fired. High steady memory is a cache or a large working set. A leak is memory that climbs and never returns — identifiable only from a trend, not a threshold.
  • It does not detect the failure that actually happens. The real failure mode is the OOM killer terminating the container. That surfaces as service-down.md and a restart, not as this alert.

First three checks​

1. Find containers near their limit.

docker stats --no-stream \
--format 'table {{.Name}}\t{{.MemUsage}}\t{{.MemPerc}}'

2. Check whether anything has already been OOM-killed.

docker ps --filter name=tutela- --format 'table {{.Names}}\t{{.Status}}'
docker inspect tutela-<service> \
--format '{{.State.OOMKilled}} {{.State.ExitCode}} {{.RestartCount}}'

OOMKilled=true, or ExitCode 137, is the memory limit being enforced. A climbing RestartCount with no other explanation is usually this.

3. Distinguish a leak from a working set — look at the trend, not the value.

Grafana is the only monitoring surface this deployment publishes (Prometheus publishes no port), and a trend over time is what it is for. Open it from the console under Settings → Diagnostics, or at https://<your-console-host>/grafana, choose the Prometheus datasource, and run this over the window named below:

process_resident_memory_bytes{job="<job>"} / 1024 / 1024

Window: last 24 hours.

Sawtooth that returns to a baseline after GC is a working set. A monotonic climb across restarts-free hours is a leak. For Go services, confirm with the runtime series:

docker exec tutela-prometheus \
/bin/promtool query instant http://localhost:9090 'go_memstats_heap_inuse_bytes / 1024 / 1024'

Heap growing while RSS grows is an application leak. RSS growing with a flat heap is off-heap — cgo, mmap, or goroutine stacks; check go_goroutines for an unbounded goroutine count.

Real problem or artefact?​

SignalReading
This alert firingInvestigate the monitoring stack — with no machine_memory_bytes producer it should be impossible. Something non-standard is exporting that series.
MemPerc above 85% sustainedReal risk. An OOM kill is coming.
OOMKilled=true in docker inspectAlready happened. Restart hid it — check RestartCount.
Memory high and flat since startupWorking set. Not a leak.
Memory climbing monotonically over hoursLeak. Capture a heap profile before restarting.
Memory high on classifierExpected — it has a 4G limit and loads models. Compare to its own limit, not the others'.

Escalation​

The primary escalation is the gap itself: this deployment has no working memory alert. Closing it means either adding a container-metrics exporter or rewriting the rule against process_resident_memory_bytes and the known per-container limit. Both change what the alert means, so both belong to the owner of the monitoring configuration rather than to an on-call fix.

For a real memory problem found through docker stats, escalate with MemPerc, the 24-hour trend from check 3, the OOMKilled/RestartCount values, and the container's configured limit. Restarting clears the symptom and destroys the evidence — capture the trend first.