Skip to documentation content
Documentation/Administer
Administer

Operator runbook index

Symptom-first routing to the owning operational contract, evidence and stop condition.

AudienceSRE · on-call · platform operationsReading timeReferenceCommands in contextNo terminal requiredReviewed source9552586 · 1 August 2026

First five minutes

  • Confirm environment, release and whether a deployment/change is in progress
  • Check /health, /ready and /health/clock without restarting first
  • Open active/failed OperationRuns and identify the owning step
  • Capture error_code, correlation and dependency/provider identifiers
  • Freeze duplicate submissions and unrelated mutations

Symptom routing

SymptomGo toStop condition
API not readyObservability & supportDo not enable lifecycle flags
Agent missing/incompatibleInstall the agentNo work assignment
Provider validation/throttlingCloud providersNo client retry loop
Cluster run stuckOperationRunsDo not create a duplicate intent
Add-on health failedBuilt-in add-onsDo not open dashboard or install dependants
Evidence checksum failedTroubleshootingTreat artifact as unusable
Backup verify failedBackup & restoreBlock upgrades
Unknown stable codeError catalogEscalate if absent

During a failed change

Use the OperationRun as the incident timeline. A service restart is not a universal recovery action: it can move leases, obscure in-memory clues or trigger startup recovery. Decide cancel, retry, cleanup, rollback or restore from the durable state contract.

Handoff

Build a complete secret-safe support package