Add recovery_playbook.py with the narrow-to-broad recovery ladder, symptom routing, attempt-log helpers, and escalation metrics. Wire the attempt-log gate into restart_coordinator so rolling/full/host restarts require prior insufficient narrower attempts (or break-glass). Document the ladder and update gitea_request_mcp_restart for prior_recovery_attempts_json. Co-Authored-By: Grok 4.5 <[email protected]>
4.8 KiB
MCP scoped recovery playbook (#669)
Parent: #655
Vision / roadmap: #652 · #653
Class matrix: #663 · docs/mcp-restart-classes.md
Coordinator: #658 · restart_coordinator.py
Audit lineage: #665
Decision
Full-server MCP reset is a last resort. Prefer the narrowest recovery that
can clear the symptom. The coordinator refuses rolling_mcp_restart,
full_mcp_restart, and host_restart unless:
- The inventory carries a prior attempt log of at least one insufficient narrower recovery, or
- Break-glass is authorized
(
request_break_glass+GITEA_BREAKGLASS_RESTART_AUTHORIZATION).
Break-glass still never bypasses the #663 class matrix (role/permission).
Ladder (narrow → broad)
| Rank | Action | Self-service | Implementation / delegation |
|---|---|---|---|
| 0 | client_reconnect |
yes | Host auto-reconnect / client reconnect · #584 · docs/mcp-namespace-eof-recovery.md |
| 1 | capability_refresh |
yes | gitea_resolve_task_capability + gitea_whoami · #610 · #685 |
| 2 | session_reconnect |
yes | Runtime rebind + explicit worktree_path · #543 · #618 |
| 3 | configuration_reload |
no | Class configuration_reload · console reload · #642 |
| 4 | lease_recovery |
no | Lock/lease recovery paths · #702 · #753 · #790 |
| 5 | worker_restart |
no | Class worker_restart · #663 |
| 6 | role_runtime_restart |
no | Class role_runtime_restart · console restart · #642/#663 |
| 7 | connector_restart |
no | Class connector_restart · #663 |
| 8 | rolling_mcp_restart |
no | Class rolling_mcp_restart · design #668 · attempt log required |
| 9 | full_mcp_restart |
no | Class full_mcp_restart · attempt log required |
| 10 | host_restart |
no | Class host_restart · attempt log required |
Machine-readable source of truth: recovery_playbook.RECOVERY_LADDER and
recovery_playbook.ladder_document().
Attempt log shape
Each prior attempt is a mapping:
{
"action": "client_reconnect",
"outcome": "insufficient",
"reason": "transport still closed after IDE reconnect",
"actor": "prgs-controller-12345",
"recorded_at": "2026-07-25T21:00:00+00:00"
}
Outcomes that count toward escalation: failed, insufficient, denied,
unresolved, timeout, error.
Pass attempts into the coordinator via inventory
prior_recovery_attempts or the MCP tool argument
prior_recovery_attempts_json on gitea_request_mcp_restart.
Helper: recovery_playbook.build_attempt_record(...).
Symptom → first rung
recovery_playbook.recommend_actions(symptoms=[...]) maps symptoms such as
transport_eof, stale_capability, stale_lease, daemon_corrupt to the
narrowest recommended action, then walks the ladder. Soft recommendations
never replace the hard gate on broad restarts.
Enforcement points
recovery_playbook.assess_escalation— pure gate.restart_coordinator.evaluate_restart_impact— whenrestart_classis set (policy-enforced path), broad classes require the gate; report fieldsattempt_log_satisfied,playbook_escalation,break_glass.gitea_request_mcp_restart— accepts attempt JSON and env-authorized break-glass; never restarts a process.
Metrics
recovery_playbook.recovery_metrics(attempts) reports the fraction of
successful recoveries that avoided full/host restart
(fraction_avoided_full_restart).
Non-goals
- HA multi-instance execution (#668 design only here).
- Normalizing
pkill(#630 contamination stays forbidden). - Silent mutation of leases or processes from the playbook itself.
Manual process kills
Remain forbidden and contaminating (#630). The playbook never recommends them.