Inspector Fail-Closed Triage
Use this runbook when the browser extension or AI Chat blocks work with Inspection unavailable (fail-closed) or other inspection-unavailable behavior. Browser-extension availability defaults to fail-open; a fail-closed block should indicate an explicit managed/admin setting or route override.
Route-Scoped Availability Policy
Inspection availability is controlled per route. The deployment default is open and is used unless an administrator sets a tenant override for the specific route:
- Browser extension and file uploads:
browser_extension_file - Portal AI Chat:
portal_chat - API and SDK requests:
api_sdk - MCP and enforcement routes:
mcp_enforcement
The Settings page shows the effective mode, its source, and the precedence (tenant_override before deployment_default). The older Portal Chat toggle is retained for compatibility, but it updates only portal_chat; it does not change file uploads, API/SDK, or MCP behavior. A file-upload fail-closed selection must appear on the Browser extension and file uploads row before /v1/inspect/file should block solely because inspection is unavailable.
Fail-open is the default availability decision. It applies only when inspection is unavailable; it does not bypass an active policy match or turn a policy block into an allow. Fail-closed is an explicit administrator-controlled decision for environments that prefer stopping work over allowing uninspected traffic during service degradation.
What To Capture
- Timestamp of the user-visible failure
- Target site or app
- Inspection flow if known:
prompt,response,paste,file - Request ID from the extension diagnostics flow if available
- Effective route and fail mode shown in Settings
Single-Pane Workflow
- Open the
Tutela - Inspector Path TriageGrafana dashboard. - Paste the request ID from the extension diagnostic details view into Grafana Logs or Loki.
- Branch from the dashboard signal:
- classifier spike: inspect classifier metrics and
classifier errorlogs - policy fallback: inspect policy metrics and fallback state
- auth, transport, or validation failure: inspect inspector logs for request-level errors
- classifier spike: inspect classifier metrics and
Key Queries
sum(rate(tutela_inspection_result_total[5m])) by (action, action_reason)
sum(rate(tutela_inspector_failure_total[5m])) by (stage, flow, reason)
increase(tutela_classification_error_total[5m])
max(tutela_policy_fallback_active)
{service="inspector", level=~"warn|error"}
Operator Checklist
- Confirm whether the dashboard shows
inspection_unavailableblocks. - Use request ID correlation before widening the search window.
- If
stage="classifier", inspect classifier latency and error growth. - If
stage="policy", inspect policy availability and fallback counters. - If classifier and policy are quiet, inspect auth, validation, or WebSocket failures in inspector logs.
User-Safe Failure Reasons
The inspector keeps the response free of prompts and sensitive values while exposing a normalized reason for triage. Common values are:
classifier_unavailableorclassifier_timeoutpolicy_unavailableorpolicy_timeoutservice_auth_failureinspection_dependency_unavailable
These reasons identify the failed stage; they are not policy-match results. A fail-open response is allowed with an auditable degraded result by default. A fail-closed response blocks because inspection could not complete and the route or deployment was explicitly configured to fail closed.