Skip to main content

Service Down

Target of the ServiceDown alert (/opt/tutela/deploy/monitoring/alerts/service-alerts.yml).

What fired​

up == 0

for: 1m, severity critical.

up is not produced by the service. Prometheus synthesises it for every scrape target: 1 when the last scrape returned HTTP 200 with parseable metrics, 0 when the scrape failed for any reason. So this alert says "Prometheus could not scrape {{ $labels.job }} for a minute" — which is a superset of "the service is down".

The thirteen scrape jobs — twelve services plus Prometheus itself — are pinned in /opt/tutela/deploy/monitoring/prometheus.yml:

JobTargetJobTarget
prometheuslocalhost:9090notificationnotification:8086
gatewaygateway:8080integrationsintegrations:8087
discoverydiscovery:8083dspm-bridgedspm-bridge:8088
classifierclassifier:8002ai-proxyai-proxy:8089
inspectorinspector:8082mcp-gatewaymcp-gateway:8090
policypolicy:8084aws-runtime-gatewayaws-runtime-gateway:8095
analyticsanalytics:8085

What this does NOT mean​

  • It does not mean the process exited. Every job sends an X-Internal-Auth header read from a per-job file under /tmp/prometheus-secrets/ inside the tutela-prometheus container, which writes those files at start-up from the INTERNAL_API_KEY_* values in /opt/tutela/.env. The file names are pinned in /opt/tutela/deploy/monitoring/prometheus.yml and do not all follow the job name: gateway reads prometheus_internal_key, integrations reads integration_internal_key (singular), and the rest are discovery_internal_key, classifier_internal_key, inspector_internal_key, policy_internal_key, analytics_internal_key, notification_internal_key, dspm_bridge_internal_key, ai_proxy_internal_key, mcp_gateway_internal_key and aws_runtime_gateway_internal_key. A rotated, missing or mismatched internal key makes /metrics return 401 and up go to 0 while the service happily serves customer traffic. Check this before restarting anything.
  • It does not mean customers are affected. Only gateway is on the customer path by itself; the rest are reached through it. A scrape failure on dspm-bridge is invisible to users.
  • It does not mean the service is unhealthy. up is about the /metrics endpoint, not /healthz or /readyz. A service can be up == 1 and failing every request — that is what InspectorHiddenFailureWhileUp and the error-rate alerts are for.
  • It does not distinguish "down" from "never started". A target that has never been scraped successfully since Prometheus booted looks identical to one that just died.

First three checks​

1. Is the container running, or is this a scrape problem?

docker ps --filter name=tutela- --format 'table {{.Names}}\t{{.Status}}'
docker ps --filter name=tutela-<service> --format 'table {{.Names}}\t{{.Status}}'

A container in running (healthy) with up == 0 means the scrape is failing, not the service. Go to check 2. A container that is exited, restarting or absent is a real outage — go to check 3.

2. Ask Prometheus why the scrape failed.

Prometheus records the reason for every failed scrape, which is more reliable than reproducing the request by hand — it already holds the credentials and it saw the actual failure:

docker exec tutela-prometheus \
wget -qO- 'http://localhost:9090/api/v1/targets?state=active' \
| jq -r '.data.activeTargets[] | select(.health!="up")
| "\(.labels.job)\t\(.health)\t\(.lastError)"'

lastError distinguishes the cases directly: connection refused (process down), 401 Unauthorized (internal key mismatch), context deadline exceeded (service alive but saturated — treat as saturation, not an outage), no such host (the container is gone or the network was recreated).

3. Read the service's own health, then its last logs.

Every one of the twelve service containers carries a health probe that the deployment runs on a schedule, using whichever scheme the services talk to each other with — http, or https when the deployment was installed with internal mTLS. Read that result first; it is right in either posture and needs nothing from you:

docker inspect --format '{{.State.Health.Status}}' tutela-<service>
docker inspect --format '{{json .State.Health.Log}}' tutela-<service> | jq '.[-1]'

The first prints healthy, unhealthy or starting; the second prints the most recent probe — its exit code and whatever it wrote.

To probe by hand: the service images carry no shell and no curl — only the service and the /healthcheck binary the probe uses (the classifier's is python -m app.healthcheck, same behaviour). It prints nothing and exits 0 when the endpoint answers 2xx; otherwise it prints healthcheck failed: ... (the HTTP status, or the connection error) and exits 1. Use the scheme the deployment itself uses. Read it once:

grep -E '^INTERNAL_SERVICE_SCHEME=' /opt/tutela/.env

https means internal mTLS is on; http, or no line at all, means it is off. Then, with <scheme> set to that value, address the service by its own name:

docker exec tutela-<service> /healthcheck <scheme>://<service>:<port>/healthz; echo "exit=$?"
docker exec tutela-<service> /healthcheck <scheme>://<service>:<port>/readyz; echo "exit=$?"
docker exec tutela-classifier python -m app.healthcheck <scheme>://classifier:8002/readyz; echo "exit=$?"
docker logs --tail=200 tutela-<service>

This is exactly the form the deployment's own probe runs, so it works in both postures: under mTLS the probe already holds the client certificate and CA it needs, and the service certificate is issued for the service's name. Do not substitute http://localhost:<port> — under mTLS the service is listening for TLS, so that form returns healthcheck failed for a perfectly healthy service and reads as a false "down". The one exception is the gateway, which serves plain HTTP on 8080 in every posture: http://localhost:8080/... is always right for it.

/healthz is liveness, /readyz is readiness — a service that is live but not ready is waiting on a dependency (database, Kafka, Redis). /healthcheck only tells you that readiness failed; to read which dependency, use the port the service publishes on the host's loopback interface (the gateway publishes 8080 on every interface):

ServiceHost portServiceHost port
gateway8080notification127.0.0.1:8086
classifier127.0.0.1:8002integrations127.0.0.1:8087
inspector127.0.0.1:8082ai-proxy127.0.0.1:8089
discovery127.0.0.1:8083mcp-gateway127.0.0.1:8090
analytics127.0.0.1:8085policy, dspm-bridge, aws-runtime-gatewaynot published — use /healthcheck and the logs
curl -s http://127.0.0.1:<host-port>/readyz

The body names the dependency, for example {"status":"not_ready","reason":"redis unavailable"}. When the deployment runs with internal mTLS enabled (INTERNAL_SERVICE_SCHEME=https in /opt/tutela/.env) every service other than the gateway listens with TLS and requires a client certificate, so for those use https:// and present the material under /opt/tutela/mtls/ (the gateway's 8080 stays plain HTTP):

curl -s --cacert /opt/tutela/mtls/ca/ca.crt \
--cert /opt/tutela/mtls/gateway/tls.crt --key /opt/tutela/mtls/gateway/tls.key \
https://127.0.0.1:<host-port>/readyz

Real problem or artefact?​

SignalReading
One job down, container exited non-zeroReal outage. Crash loop — read the logs before restarting, or you destroy the evidence.
One job down, container running, lastError is 401Artefact. Internal key rotation; the service is fine.
One job down, lastError is context deadline exceededNot an outage. The service is too slow to answer a 10s scrape (scrape_timeout: 10s). Treat as saturation.
All jobs down at once, including prometheusArtefact. Prometheus itself, its network, or the host is the problem — not every service failing at once.
job="prometheus" alone is downPrometheus cannot scrape itself; the alerting pipeline is degraded. Nothing else can be trusted until it recovers.
Alert fires during a deploy, clears within 1-2mExpected. Rolling restart. Confirm against the deploy timestamp before escalating.

Escalation​

ServiceDown on gateway is customer-visible and pages immediately: every customer request enters through it. ServiceDown on inspector, policy or classifier means governance is degraded even if the gateway is up — pair this runbook with inspector-failclosed-triage.md, because the inspector's fail mode decides whether that degradation blocks traffic or silently stops inspecting it.

Escalate with: the lastError string from check 2, the last 200 log lines, the container state, and whether a deploy was in flight. /v1/version on the gateway identifies the running build.