Service Down
Target of the ServiceDown alert
(/opt/tutela/deploy/monitoring/alerts/service-alerts.yml).
What fired
up == 0
for: 1m, severity critical.
up is not produced by the service. Prometheus synthesises it for every scrape
target: 1 when the last scrape returned HTTP 200 with parseable metrics, 0
when the scrape failed for any reason. So this alert says "Prometheus could
not scrape {{ $labels.job }} for a minute" — which is a superset of "the
service is down".
The thirteen scrape jobs — twelve services plus Prometheus itself — are pinned
in /opt/tutela/deploy/monitoring/prometheus.yml:
| Job | Target | Job | Target |
|---|---|---|---|
prometheus | localhost:9090 | notification | notification:8086 |
gateway | gateway:8080 | integrations | integrations:8087 |
discovery | discovery:8083 | dspm-bridge | dspm-bridge:8088 |
classifier | classifier:8002 | ai-proxy | ai-proxy:8089 |
inspector | inspector:8082 | mcp-gateway | mcp-gateway:8090 |
policy | policy:8084 | aws-runtime-gateway | aws-runtime-gateway:8095 |
analytics | analytics:8085 |
What this does NOT mean
- It does not mean the process exited. Every job sends an
X-Internal-Authheader read from a per-job file under/tmp/prometheus-secrets/inside thetutela-prometheuscontainer, which writes those files at start-up from theINTERNAL_API_KEY_*values in/opt/tutela/.env. The file names are pinned in/opt/tutela/deploy/monitoring/prometheus.ymland do not all follow the job name:gatewayreadsprometheus_internal_key,integrationsreadsintegration_internal_key(singular), and the rest arediscovery_internal_key,classifier_internal_key,inspector_internal_key,policy_internal_key,analytics_internal_key,notification_internal_key,dspm_bridge_internal_key,ai_proxy_internal_key,mcp_gateway_internal_keyandaws_runtime_gateway_internal_key. A rotated, missing or mismatched internal key makes/metricsreturn 401 andupgo to0while the service happily serves customer traffic. Check this before restarting anything. - It does not mean customers are affected. Only
gatewayis on the customer path by itself; the rest are reached through it. A scrape failure ondspm-bridgeis invisible to users. - It does not mean the service is unhealthy.
upis about the/metricsendpoint, not/healthzor/readyz. A service can beup == 1and failing every request — that is whatInspectorHiddenFailureWhileUpand the error-rate alerts are for. - It does not distinguish "down" from "never started". A target that has never been scraped successfully since Prometheus booted looks identical to one that just died.
First three checks
1. Is the container running, or is this a scrape problem?
docker ps --filter name=tutela- --format 'table {{.Names}}\t{{.Status}}'
docker ps --filter name=tutela-<service> --format 'table {{.Names}}\t{{.Status}}'
A container in running (healthy) with up == 0 means the scrape is failing,
not the service. Go to check 2. A container that is exited, restarting or
absent is a real outage — go to check 3.
2. Ask Prometheus why the scrape failed.
Prometheus records the reason for every failed scrape, which is more reliable than reproducing the request by hand — it already holds the credentials and it saw the actual failure:
docker exec tutela-prometheus \
wget -qO- 'http://localhost:9090/api/v1/targets?state=active' \
| jq -r '.data.activeTargets[] | select(.health!="up")
| "\(.labels.job)\t\(.health)\t\(.lastError)"'
lastError distinguishes the cases directly: connection refused (process
down), 401 Unauthorized (internal key mismatch), context deadline exceeded
(service alive but saturated — treat as saturation, not an outage),
no such host (the container is gone or the network was recreated).
3. Read the service's own health, then its last logs.
Every one of the twelve service containers carries a health probe that the
deployment runs on a schedule, using whichever scheme the services talk to each
other with — http, or https when the deployment was installed with internal
mTLS. Read that result first; it is right in either posture and needs nothing
from you:
docker inspect --format '{{.State.Health.Status}}' tutela-<service>
docker inspect --format '{{json .State.Health.Log}}' tutela-<service> | jq '.[-1]'
The first prints healthy, unhealthy or starting; the second prints the
most recent probe — its exit code and whatever it wrote.
To probe by hand: the service images carry no shell and no curl — only the
service and the /healthcheck binary the probe uses (the classifier's is
python -m app.healthcheck, same behaviour). It prints nothing and exits 0
when the endpoint answers 2xx; otherwise it prints healthcheck failed: ...
(the HTTP status, or the connection error) and exits 1. Use the scheme the
deployment itself uses. Read it once:
grep -E '^INTERNAL_SERVICE_SCHEME=' /opt/tutela/.env
https means internal mTLS is on; http, or no line at all, means it is off.
Then, with <scheme> set to that value, address the service by its own name:
docker exec tutela-<service> /healthcheck <scheme>://<service>:<port>/healthz; echo "exit=$?"
docker exec tutela-<service> /healthcheck <scheme>://<service>:<port>/readyz; echo "exit=$?"
docker exec tutela-classifier python -m app.healthcheck <scheme>://classifier:8002/readyz; echo "exit=$?"
docker logs --tail=200 tutela-<service>
This is exactly the form the deployment's own probe runs, so it works in both
postures: under mTLS the probe already holds the client certificate and CA it
needs, and the service certificate is issued for the service's name. Do not
substitute http://localhost:<port> — under mTLS the service is listening for
TLS, so that form returns healthcheck failed for a perfectly healthy service
and reads as a false "down". The one exception is the gateway, which serves
plain HTTP on 8080 in every posture: http://localhost:8080/... is always
right for it.
/healthz is liveness, /readyz is readiness — a service that is live but not
ready is waiting on a dependency (database, Kafka, Redis). /healthcheck only
tells you that readiness failed; to read which dependency, use the port the
service publishes on the host's loopback interface (the gateway publishes
8080 on every interface):
| Service | Host port | Service | Host port |
|---|---|---|---|
gateway | 8080 | notification | 127.0.0.1:8086 |
classifier | 127.0.0.1:8002 | integrations | 127.0.0.1:8087 |
inspector | 127.0.0.1:8082 | ai-proxy | 127.0.0.1:8089 |
discovery | 127.0.0.1:8083 | mcp-gateway | 127.0.0.1:8090 |
analytics | 127.0.0.1:8085 | policy, dspm-bridge, aws-runtime-gateway | not published — use /healthcheck and the logs |
curl -s http://127.0.0.1:<host-port>/readyz
The body names the dependency, for example
{"status":"not_ready","reason":"redis unavailable"}. When the deployment runs
with internal mTLS enabled (INTERNAL_SERVICE_SCHEME=https in
/opt/tutela/.env) every service other than the gateway listens with TLS and
requires a client certificate, so for those use https:// and present the
material under /opt/tutela/mtls/ (the gateway's 8080 stays plain HTTP):
curl -s --cacert /opt/tutela/mtls/ca/ca.crt \
--cert /opt/tutela/mtls/gateway/tls.crt --key /opt/tutela/mtls/gateway/tls.key \
https://127.0.0.1:<host-port>/readyz
Real problem or artefact?
| Signal | Reading |
|---|---|
One job down, container exited non-zero | Real outage. Crash loop — read the logs before restarting, or you destroy the evidence. |
One job down, container running, lastError is 401 | Artefact. Internal key rotation; the service is fine. |
One job down, lastError is context deadline exceeded | Not an outage. The service is too slow to answer a 10s scrape (scrape_timeout: 10s). Treat as saturation. |
All jobs down at once, including prometheus | Artefact. Prometheus itself, its network, or the host is the problem — not every service failing at once. |
job="prometheus" alone is down | Prometheus cannot scrape itself; the alerting pipeline is degraded. Nothing else can be trusted until it recovers. |
| Alert fires during a deploy, clears within 1-2m | Expected. Rolling restart. Confirm against the deploy timestamp before escalating. |
Escalation
ServiceDown on gateway is customer-visible and pages immediately: every
customer request enters through it. ServiceDown on inspector, policy or
classifier means governance is degraded even if the gateway is up — pair this
runbook with inspector-failclosed-triage.md,
because the inspector's fail mode decides whether that degradation blocks
traffic or silently stops inspecting it.
Escalate with: the lastError string from check 2, the last 200 log lines, the
container state, and whether a deploy was in flight. /v1/version on the
gateway identifies the running build.
Related runbooks
- high-error-rate.md — the service answers, but with 5xx.
- high-latency.md — scrape timeouts are usually latency.
- inspector-failclosed-triage.md — what an inspector or policy outage does to enforcement.
- error-budget.md — the outage's cost against the SLO.