Skip to main content

Inspector Fail-Closed Triage

Use this runbook when the browser extension or AI Chat blocks work with Inspection unavailable (fail-closed) or other inspection-unavailable behavior. Browser-extension availability defaults to fail-open; a fail-closed block should indicate an explicit managed/admin setting or route override.

Route-Scoped Availability Policy​

Inspection availability is controlled per route. The deployment default is open and is used unless an administrator sets a tenant override for the specific route:

  • Browser extension and file uploads: browser_extension_file
  • Portal AI Chat: portal_chat
  • API and SDK requests: api_sdk
  • MCP and enforcement routes: mcp_enforcement

The Settings page shows the effective mode, its source, and the precedence (tenant_override before deployment_default). The older Portal Chat toggle is retained for compatibility, but it updates only portal_chat; it does not change file uploads, API/SDK, or MCP behavior. A file-upload fail-closed selection must appear on the Browser extension and file uploads row before /v1/inspect/file should block solely because inspection is unavailable.

Fail-open is the default availability decision. It applies only when inspection is unavailable; it does not bypass an active policy match or turn a policy block into an allow. Fail-closed is an explicit administrator-controlled decision for environments that prefer stopping work over allowing uninspected traffic during service degradation.

What To Capture​

  • Timestamp of the user-visible failure
  • Target site or app
  • Inspection flow if known: prompt, response, paste, file
  • Request ID from the extension diagnostics flow if available
  • Effective route and fail mode shown in Settings

Single-Pane Workflow​

  1. Open the Tutela - Inspector Path Triage Grafana dashboard.
  2. Paste the request ID from the extension diagnostic details view into Grafana Logs or Loki.
  3. Branch from the dashboard signal:
    • classifier spike: inspect classifier metrics and classifier error logs
    • policy fallback: inspect policy metrics and fallback state
    • auth, transport, or validation failure: inspect inspector logs for request-level errors

Key Queries​

sum(rate(tutela_inspection_result_total[5m])) by (action, action_reason)
sum(rate(tutela_inspector_failure_total[5m])) by (stage, flow, reason)
increase(tutela_classification_error_total[5m])
max(tutela_policy_fallback_active)
{service="inspector", level=~"warn|error"}

Operator Checklist​

  • Confirm whether the dashboard shows inspection_unavailable blocks.
  • Use request ID correlation before widening the search window.
  • If stage="classifier", inspect classifier latency and error growth.
  • If stage="policy", inspect policy availability and fallback counters.
  • If classifier and policy are quiet, inspect auth, validation, or WebSocket failures in inspector logs.

User-Safe Failure Reasons​

The inspector keeps the response free of prompts and sensitive values while exposing a normalized reason for triage. Common values are:

  • classifier_unavailable or classifier_timeout
  • policy_unavailable or policy_timeout
  • service_auth_failure
  • inspection_dependency_unavailable

These reasons identify the failed stage; they are not policy-match results. A fail-open response is allowed with an auditable degraded result by default. A fail-closed response blocks because inspection could not complete and the route or deployment was explicitly configured to fail closed.