feat(mcp): enforce scoped recovery playbook before full restart (Closes #669)
Add recovery_playbook.py with the narrow-to-broad recovery ladder, symptom routing, attempt-log helpers, and escalation metrics. Wire the attempt-log gate into restart_coordinator so rolling/full/host restarts require prior insufficient narrower attempts (or break-glass). Document the ladder and update gitea_request_mcp_restart for prior_recovery_attempts_json. Co-Authored-By: Grok 4.5 <[email protected]>
This commit is contained in:
@@ -0,0 +1,94 @@
|
||||
# MCP scoped recovery playbook (#669)
|
||||
|
||||
**Parent:** [#655](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/655)
|
||||
**Vision / roadmap:** [#652](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/652) · [#653](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/653)
|
||||
**Class matrix:** [#663](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/663) · `docs/mcp-restart-classes.md`
|
||||
**Coordinator:** [#658](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/658) · `restart_coordinator.py`
|
||||
**Audit lineage:** [#665](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/665)
|
||||
|
||||
## Decision
|
||||
|
||||
Full-server MCP reset is a **last resort**. Prefer the narrowest recovery that
|
||||
can clear the symptom. The coordinator **refuses** `rolling_mcp_restart`,
|
||||
`full_mcp_restart`, and `host_restart` unless:
|
||||
|
||||
1. The inventory carries a prior **attempt log** of at least one *insufficient*
|
||||
narrower recovery, **or**
|
||||
2. **Break-glass** is authorized
|
||||
(`request_break_glass` + `GITEA_BREAKGLASS_RESTART_AUTHORIZATION`).
|
||||
|
||||
Break-glass still never bypasses the #663 class matrix (role/permission).
|
||||
|
||||
## Ladder (narrow → broad)
|
||||
|
||||
| Rank | Action | Self-service | Implementation / delegation |
|
||||
|---:|---|---|---|
|
||||
| 0 | `client_reconnect` | yes | Host auto-reconnect / client reconnect · [#584](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/584) · `docs/mcp-namespace-eof-recovery.md` |
|
||||
| 1 | `capability_refresh` | yes | `gitea_resolve_task_capability` + `gitea_whoami` · [#610](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/610) · [#685](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/685) |
|
||||
| 2 | `session_reconnect` | yes | Runtime rebind + explicit `worktree_path` · [#543](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/543) · [#618](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/618) |
|
||||
| 3 | `configuration_reload` | no | Class `configuration_reload` · console reload · [#642](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/642) |
|
||||
| 4 | `lease_recovery` | no | Lock/lease recovery paths · [#702](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/702) · [#753](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/753) · [#790](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/790) |
|
||||
| 5 | `worker_restart` | no | Class `worker_restart` · #663 |
|
||||
| 6 | `role_runtime_restart` | no | Class `role_runtime_restart` · console restart · #642/#663 |
|
||||
| 7 | `connector_restart` | no | Class `connector_restart` · #663 |
|
||||
| 8 | `rolling_mcp_restart` | no | Class `rolling_mcp_restart` · design [#668](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/668) · **attempt log required** |
|
||||
| 9 | `full_mcp_restart` | no | Class `full_mcp_restart` · **attempt log required** |
|
||||
| 10 | `host_restart` | no | Class `host_restart` · **attempt log required** |
|
||||
|
||||
Machine-readable source of truth: `recovery_playbook.RECOVERY_LADDER` and
|
||||
`recovery_playbook.ladder_document()`.
|
||||
|
||||
## Attempt log shape
|
||||
|
||||
Each prior attempt is a mapping:
|
||||
|
||||
```json
|
||||
{
|
||||
"action": "client_reconnect",
|
||||
"outcome": "insufficient",
|
||||
"reason": "transport still closed after IDE reconnect",
|
||||
"actor": "prgs-controller-12345",
|
||||
"recorded_at": "2026-07-25T21:00:00+00:00"
|
||||
}
|
||||
```
|
||||
|
||||
Outcomes that count toward escalation: `failed`, `insufficient`, `denied`,
|
||||
`unresolved`, `timeout`, `error`.
|
||||
|
||||
Pass attempts into the coordinator via inventory
|
||||
`prior_recovery_attempts` or the MCP tool argument
|
||||
`prior_recovery_attempts_json` on `gitea_request_mcp_restart`.
|
||||
|
||||
Helper: `recovery_playbook.build_attempt_record(...)`.
|
||||
|
||||
## Symptom → first rung
|
||||
|
||||
`recovery_playbook.recommend_actions(symptoms=[...])` maps symptoms such as
|
||||
`transport_eof`, `stale_capability`, `stale_lease`, `daemon_corrupt` to the
|
||||
narrowest recommended action, then walks the ladder. Soft recommendations
|
||||
never replace the hard gate on broad restarts.
|
||||
|
||||
## Enforcement points
|
||||
|
||||
1. **`recovery_playbook.assess_escalation`** — pure gate.
|
||||
2. **`restart_coordinator.evaluate_restart_impact`** — when `restart_class` is
|
||||
set (policy-enforced path), broad classes require the gate; report fields
|
||||
`attempt_log_satisfied`, `playbook_escalation`, `break_glass`.
|
||||
3. **`gitea_request_mcp_restart`** — accepts attempt JSON and env-authorized
|
||||
break-glass; never restarts a process.
|
||||
|
||||
## Metrics
|
||||
|
||||
`recovery_playbook.recovery_metrics(attempts)` reports the fraction of
|
||||
successful recoveries that avoided full/host restart
|
||||
(`fraction_avoided_full_restart`).
|
||||
|
||||
## Non-goals
|
||||
|
||||
* HA multi-instance execution (#668 design only here).
|
||||
* Normalizing `pkill` (#630 contamination stays forbidden).
|
||||
* Silent mutation of leases or processes from the playbook itself.
|
||||
|
||||
## Manual process kills
|
||||
|
||||
Remain forbidden and contaminating (#630). The playbook never recommends them.
|
||||
Reference in New Issue
Block a user