Files
Gitea-Tools/docs/mcp-recovery-playbook.md
T
sysadminandGrok 4.5 461e1dac78 feat(mcp): enforce scoped recovery playbook before full restart (Closes #669)
Add recovery_playbook.py with the narrow-to-broad recovery ladder, symptom
routing, attempt-log helpers, and escalation metrics. Wire the attempt-log
gate into restart_coordinator so rolling/full/host restarts require prior
insufficient narrower attempts (or break-glass). Document the ladder and
update gitea_request_mcp_restart for prior_recovery_attempts_json.

Co-Authored-By: Grok 4.5 <[email protected]>
2026-07-25 17:35:11 -04:00

95 lines
4.8 KiB
Markdown

# MCP scoped recovery playbook (#669)
**Parent:** [#655](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/655)
**Vision / roadmap:** [#652](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/652) · [#653](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/653)
**Class matrix:** [#663](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/663) · `docs/mcp-restart-classes.md`
**Coordinator:** [#658](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/658) · `restart_coordinator.py`
**Audit lineage:** [#665](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/665)
## Decision
Full-server MCP reset is a **last resort**. Prefer the narrowest recovery that
can clear the symptom. The coordinator **refuses** `rolling_mcp_restart`,
`full_mcp_restart`, and `host_restart` unless:
1. The inventory carries a prior **attempt log** of at least one *insufficient*
narrower recovery, **or**
2. **Break-glass** is authorized
(`request_break_glass` + `GITEA_BREAKGLASS_RESTART_AUTHORIZATION`).
Break-glass still never bypasses the #663 class matrix (role/permission).
## Ladder (narrow → broad)
| Rank | Action | Self-service | Implementation / delegation |
|---:|---|---|---|
| 0 | `client_reconnect` | yes | Host auto-reconnect / client reconnect · [#584](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/584) · `docs/mcp-namespace-eof-recovery.md` |
| 1 | `capability_refresh` | yes | `gitea_resolve_task_capability` + `gitea_whoami` · [#610](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/610) · [#685](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/685) |
| 2 | `session_reconnect` | yes | Runtime rebind + explicit `worktree_path` · [#543](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/543) · [#618](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/618) |
| 3 | `configuration_reload` | no | Class `configuration_reload` · console reload · [#642](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/642) |
| 4 | `lease_recovery` | no | Lock/lease recovery paths · [#702](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/702) · [#753](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/753) · [#790](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/790) |
| 5 | `worker_restart` | no | Class `worker_restart` · #663 |
| 6 | `role_runtime_restart` | no | Class `role_runtime_restart` · console restart · #642/#663 |
| 7 | `connector_restart` | no | Class `connector_restart` · #663 |
| 8 | `rolling_mcp_restart` | no | Class `rolling_mcp_restart` · design [#668](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/668) · **attempt log required** |
| 9 | `full_mcp_restart` | no | Class `full_mcp_restart` · **attempt log required** |
| 10 | `host_restart` | no | Class `host_restart` · **attempt log required** |
Machine-readable source of truth: `recovery_playbook.RECOVERY_LADDER` and
`recovery_playbook.ladder_document()`.
## Attempt log shape
Each prior attempt is a mapping:
```json
{
"action": "client_reconnect",
"outcome": "insufficient",
"reason": "transport still closed after IDE reconnect",
"actor": "prgs-controller-12345",
"recorded_at": "2026-07-25T21:00:00+00:00"
}
```
Outcomes that count toward escalation: `failed`, `insufficient`, `denied`,
`unresolved`, `timeout`, `error`.
Pass attempts into the coordinator via inventory
`prior_recovery_attempts` or the MCP tool argument
`prior_recovery_attempts_json` on `gitea_request_mcp_restart`.
Helper: `recovery_playbook.build_attempt_record(...)`.
## Symptom → first rung
`recovery_playbook.recommend_actions(symptoms=[...])` maps symptoms such as
`transport_eof`, `stale_capability`, `stale_lease`, `daemon_corrupt` to the
narrowest recommended action, then walks the ladder. Soft recommendations
never replace the hard gate on broad restarts.
## Enforcement points
1. **`recovery_playbook.assess_escalation`** — pure gate.
2. **`restart_coordinator.evaluate_restart_impact`** — when `restart_class` is
set (policy-enforced path), broad classes require the gate; report fields
`attempt_log_satisfied`, `playbook_escalation`, `break_glass`.
3. **`gitea_request_mcp_restart`** — accepts attempt JSON and env-authorized
break-glass; never restarts a process.
## Metrics
`recovery_playbook.recovery_metrics(attempts)` reports the fraction of
successful recoveries that avoided full/host restart
(`fraction_avoided_full_restart`).
## Non-goals
* HA multi-instance execution (#668 design only here).
* Normalizing `pkill` (#630 contamination stays forbidden).
* Silent mutation of leases or processes from the playbook itself.
## Manual process kills
Remain forbidden and contaminating (#630). The playbook never recommends them.