# MCP scoped recovery playbook (#669) **Parent:** [#655](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/655) **Vision / roadmap:** [#652](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/652) · [#653](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/653) **Class matrix:** [#663](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/663) · `docs/mcp-restart-classes.md` **Coordinator:** [#658](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/658) · `restart_coordinator.py` **Audit lineage:** [#665](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/665) ## Decision Full-server MCP reset is a **last resort**. Prefer the narrowest recovery that can clear the symptom. The coordinator **refuses** `rolling_mcp_restart`, `full_mcp_restart`, and `host_restart` unless: 1. The inventory carries a prior **attempt log** of at least one *insufficient* narrower recovery, **or** 2. **Break-glass** is authorized (`request_break_glass` + `GITEA_BREAKGLASS_RESTART_AUTHORIZATION`). Break-glass still never bypasses the #663 class matrix (role/permission). ## Ladder (narrow → broad) | Rank | Action | Self-service | Implementation / delegation | |---:|---|---|---| | 0 | `client_reconnect` | yes | Host auto-reconnect / client reconnect · [#584](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/584) · `docs/mcp-namespace-eof-recovery.md` | | 1 | `capability_refresh` | yes | `gitea_resolve_task_capability` + `gitea_whoami` · [#610](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/610) · [#685](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/685) | | 2 | `session_reconnect` | yes | Runtime rebind + explicit `worktree_path` · [#543](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/543) · [#618](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/618) | | 3 | `configuration_reload` | no | Class `configuration_reload` · console reload · [#642](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/642) | | 4 | `lease_recovery` | no | Lock/lease recovery paths · [#702](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/702) · [#753](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/753) · [#790](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/790) | | 5 | `worker_restart` | no | Class `worker_restart` · #663 | | 6 | `role_runtime_restart` | no | Class `role_runtime_restart` · console restart · #642/#663 | | 7 | `connector_restart` | no | Class `connector_restart` · #663 | | 8 | `rolling_mcp_restart` | no | Class `rolling_mcp_restart` · design [#668](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/668) · **attempt log required** | | 9 | `full_mcp_restart` | no | Class `full_mcp_restart` · **attempt log required** | | 10 | `host_restart` | no | Class `host_restart` · **attempt log required** | Machine-readable source of truth: `recovery_playbook.RECOVERY_LADDER` and `recovery_playbook.ladder_document()`. ## Attempt log shape Each prior attempt is a mapping: ```json { "action": "client_reconnect", "outcome": "insufficient", "reason": "transport still closed after IDE reconnect", "actor": "prgs-controller-12345", "recorded_at": "2026-07-25T21:00:00+00:00" } ``` Outcomes that count toward escalation: `failed`, `insufficient`, `denied`, `unresolved`, `timeout`, `error`. Pass attempts into the coordinator via inventory `prior_recovery_attempts` or the MCP tool argument `prior_recovery_attempts_json` on `gitea_request_mcp_restart`. Helper: `recovery_playbook.build_attempt_record(...)`. ## Symptom → first rung `recovery_playbook.recommend_actions(symptoms=[...])` maps symptoms such as `transport_eof`, `stale_capability`, `stale_lease`, `daemon_corrupt` to the narrowest recommended action, then walks the ladder. Soft recommendations never replace the hard gate on broad restarts. ## Enforcement points 1. **`recovery_playbook.assess_escalation`** — pure gate. 2. **`restart_coordinator.evaluate_restart_impact`** — when `restart_class` is set (policy-enforced path), broad classes require the gate; report fields `attempt_log_satisfied`, `playbook_escalation`, `break_glass`. 3. **`gitea_request_mcp_restart`** — accepts attempt JSON and env-authorized break-glass; never restarts a process. ## Metrics `recovery_playbook.recovery_metrics(attempts)` reports the fraction of successful recoveries that avoided full/host restart (`fraction_avoided_full_restart`). ## Non-goals * HA multi-instance execution (#668 design only here). * Normalizing `pkill` (#630 contamination stays forbidden). * Silent mutation of leases or processes from the playbook itself. ## Manual process kills Remain forbidden and contaminating (#630). The playbook never recommends them.