High CPU Usage
Target of the HighCPUUsage alert
(/opt/tutela/deploy/monitoring/alerts/service-alerts.yml).
What fired
rate(process_cpu_seconds_total[5m]) * 100 > 80
for: 5m, severity warning.
The unit is cores, not percent of the machine
This is the single most misread expression in the ruleset, so read it carefully.
process_cpu_seconds_total counts CPU-seconds the process has consumed.
rate(...[5m]) is therefore CPU-seconds consumed per wall-clock second —
which is a count of cores in use, not a percentage of anything. 1.0 means one
full core. Multiplying by 100 does not convert it to a percentage of the host; it
just scales cores by 100.
So the threshold > 80 means "more than 0.8 of one CPU core". On a
four-core host a process using 0.8 cores is at 20% of the machine, and this
alert fires. On a 32-core host it fires at 2.5% of the machine.
The alert's own description says "CPU usage is above 80%". That wording is misleading; the number is 0.8 cores.
Coverage gap — this alert cannot fire on the Compose deployment
Two independent reasons, both verifiable. Note both are properties of the
COMPOSE deployment's limits: on Kubernetes the chart's defaults are the same
0.5 CPU (500m), but if you raise a pod's CPU limit above 0.8 this alert
becomes reachable again — and correct. (On Kubernetes the chart's alert rules,
and so this alert, exist only when monitoring.enabled is true in your
chart values; it is false by default.)
1. Every Go service is capped below the threshold. In
/opt/tutela/deploy/docker-compose.aws.yaml, eleven of the thirteen Tutela
containers carry deploy.resources.limits.cpus: '0.5', and web is lower still
at 0.25. A container limited to 0.5 CPU cannot consume more than 0.5
CPU-seconds per second, so the expression's ceiling for those is
0.5 * 100 = 50 — permanently below the threshold of 80. A service pinned at
100% of its own limit, throttling hard, reports 50 and stays silent.
2. The classifier does not export the metric. The classifier is the one
container with headroom (CLASSIFIER_CPU_LIMIT defaults to 6), but it runs
prometheus_fastapi_instrumentator with PROMETHEUS_MULTIPROC_DIR set
(/opt/tutela/deploy/docker-compose.aws.yaml). In multiprocess mode the registry serves
only the memory-mapped application metrics; prometheus_client's default
process collector — the source of process_cpu_seconds_total — is not part of
it.
Confirm both for yourself before relying on either:
docker exec tutela-prometheus \
/bin/promtool query instant http://localhost:9090 'rate(process_cpu_seconds_total[5m]) * 100'
If classifier is absent from that output, reason 2 holds. If every value is at
or below 50, reason 1 holds.
What to use instead: read the live number and compare it with the container's own limit, which is what actually matters:
docker stats --no-stream \
--format 'table {{.Name}}\t{{.CPUPerc}}\t{{.MemUsage}}\t{{.MemPerc}}'
Read CPUPerc carefully: it is in the same unit as the alert — cores × 100.
Docker computes it from the container's share of the host's online CPUs, not
from the container's limit. A service pinned at cpus: '0.5' therefore tops out
at about 50% there — the very value this alert reports and cannot act on — a
0.25 container tops out near 25%, and the classifier (limit 6 by default)
can read up to 600%. The reading that means "at its cap" is limit × 100,
never a flat 100. Take each container's limit from the running container, not
from memory:
docker inspect --format '{{.Name}} {{.HostConfig.NanoCpus}}' $(docker ps -q --filter name=tutela-)
NanoCpus is the limit in billionths of a core: 500000000 is 0.5 CPU, so
that container's ceiling in docker stats is 50%. A 0 means no limit was
applied and the ceiling is the host's core count × 100.
Whether a container at its ceiling is actually being held back is recorded by
the kernel, not by docker stats. On the shipped Ubuntu host (cgroup v2):
cat /sys/fs/cgroup$(sed -n 's/^0:://p' /proc/$(docker inspect --format '{{.State.Pid}}' tutela-<service>)/cgroup)/cpu.stat
nr_throttled and throttled_usec rising between two readings mean the
container wanted more CPU than its limit allowed. That is the number to act on.
What this does NOT mean
- It does not mean the host is busy. It is per-process and unit-cores. Host CPU is not measured anywhere in this stack — there is no node exporter.
- It does not mean the service is slow. High CPU with healthy latency is a service doing its job. Pair it with high-latency.md before concluding anything.
- It does not mean a leak. CPU is instantaneous work, not accumulated state.
- Silence does not mean idle. Per the gap above, silence is the expected state on this stack regardless of load.
First three checks
1. Get the real utilisation, against the real limit.
docker stats --no-stream \
--format 'table {{.Name}}\t{{.CPUPerc}}\t{{.MemPerc}}'
A container at its own limit × 100 (50% for a 0.5 container) is at its
cap; confirm it is being held back with the cpu.stat reading above, since
docker stats alone cannot tell "busy" from "throttled". Note the container
name (tutela-<service>).
2. Decide whether the work is legitimate.
docker exec tutela-prometheus \
/bin/promtool query instant http://localhost:9090 'sum(rate(tutela_api_request_total[5m])) by (job)'
CPU tracking request rate is capacity. CPU high with flat or zero request rate is a hot loop, a retry storm, or a background job — go to check 3.
3. Find what it is doing.
docker logs --tail=300 tutela-<service> \
| grep -E '"level":"(warn|error)"'
A retry storm against a failing dependency is the most common cause of CPU with no traffic. Correlate with redis-errors.md, kafka-consumer-lag.md and service-down.md for the dependency that is being retried.
Go services also expose runtime detail that separates GC pressure from real work:
docker exec tutela-prometheus \
/bin/promtool query instant http://localhost:9090 'rate(go_gc_duration_seconds_sum[5m])'
Heavy GC time usually means a memory problem presenting as a CPU problem — see high-memory.md.
Real problem or artefact?
| Signal | Reading |
|---|---|
| Alert firing at all on the AWS stack | Unexpected — the ceiling analysis says it should not. Verify the container limits were actually applied. |
docker stats at the container's limit × 100, nr_throttled rising, latency rising | Real. The service is throttling; add capacity or raise the limit. |
docker stats at limit × 100, latency flat | Working as intended at its cap. Headroom is gone but nothing is wrong yet. |
docker stats reads 100% on a 0.5 container | Impossible as a throttling signal — that container cannot exceed 50%. Either the limit was not applied (NanoCpus is 0) or you are reading a different container. |
| CPU high, request rate zero | Retry storm or background job. Check 3. |
| CPU spike at container start | Artefact. Startup, migrations, cache warming. Ignore under a minute. |
| High GC time alongside high CPU | Memory pressure, not CPU demand. Go to high-memory.md. |
Escalation
The real escalation for this alert is not an outage — it is that the alert does not work as configured. Fixing it properly means either alerting on CPU relative to the container's limit (which requires a container-metrics exporter, absent from this stack) or lowering the threshold to match the 0.5-CPU cap. Both are changes to what the alert means and belong to whoever owns the monitoring configuration.
For a live capacity problem found via docker stats, escalate with the
CPUPerc figure next to the container's limit (NanoCpus), two readings of the
cpu.stat throttling counters, the request rate, and the latency percentiles.
Related runbooks
- high-memory.md — the sibling resource alert, which has a larger gap.
- high-latency.md — the symptom that makes CPU matter.
- disk-space.md — the third resource alert with no producer.
- service-down.md — a saturated service fails its scrape before it fails anything else.