Documentation/Administer
Administer · v0.0.1

Operator runbook index

Symptom-first routing to the owning operational contract, evidence and stop condition.

AudienceSRE · on-call · platform operationsReading timeReferenceCommands in contextNo terminal requiredReviewed sourcefcc5871 · 30 July 2026
RC

Exact release scope. This page documents the reviewed v0.0.1 source at fcc5871. Check release status before enabling a gated capability.

First five minutes

  • Confirm environment, release and whether a deployment/change is in progress
  • Check /health, /ready and /health/clock without restarting first
  • Open active/failed OperationRuns and identify the owning step
  • Capture error_code, correlation and dependency/provider identifiers
  • Freeze duplicate submissions and unrelated mutations

Symptom routing

SymptomGo toStop condition
API not readyObservability & supportDo not enable lifecycle flags
Agent missing/incompatibleInstall the agentNo work assignment
Provider validation/throttlingCloud providersNo client retry loop
Cluster run stuckOperationRunsDo not create a duplicate intent
Add-on health failedBuilt-in add-onsDo not open dashboard or install dependants
Evidence checksum failedTroubleshootingTreat artifact as unusable
Backup verify failedBackup & restoreBlock upgrades
Unknown stable codeError catalogEscalate if absent

During a failed change

Use the OperationRun as the incident timeline. A service restart is not a universal recovery action: it can move leases, obscure in-memory clues or trigger startup recovery. Decide cancel, retry, cleanup, rollback or restore from the durable state contract.

Handoff

Build a complete secret-safe support package