Files
Gitea-Tools/docs/mcp-recovery-playbook.md
T
sysadminandGrok 4.5 461e1dac78 feat(mcp): enforce scoped recovery playbook before full restart (Closes #669)
Add recovery_playbook.py with the narrow-to-broad recovery ladder, symptom
routing, attempt-log helpers, and escalation metrics. Wire the attempt-log
gate into restart_coordinator so rolling/full/host restarts require prior
insufficient narrower attempts (or break-glass). Document the ladder and
update gitea_request_mcp_restart for prior_recovery_attempts_json.

Co-Authored-By: Grok 4.5 <[email protected]>
2026-07-25 17:35:11 -04:00

4.8 KiB

MCP scoped recovery playbook (#669)

Parent: #655
Vision / roadmap: #652 · #653
Class matrix: #663 · docs/mcp-restart-classes.md
Coordinator: #658 · restart_coordinator.py
Audit lineage: #665

Decision

Full-server MCP reset is a last resort. Prefer the narrowest recovery that can clear the symptom. The coordinator refuses rolling_mcp_restart, full_mcp_restart, and host_restart unless:

  1. The inventory carries a prior attempt log of at least one insufficient narrower recovery, or
  2. Break-glass is authorized (request_break_glass + GITEA_BREAKGLASS_RESTART_AUTHORIZATION).

Break-glass still never bypasses the #663 class matrix (role/permission).

Ladder (narrow → broad)

Rank Action Self-service Implementation / delegation
0 client_reconnect yes Host auto-reconnect / client reconnect · #584 · docs/mcp-namespace-eof-recovery.md
1 capability_refresh yes gitea_resolve_task_capability + gitea_whoami · #610 · #685
2 session_reconnect yes Runtime rebind + explicit worktree_path · #543 · #618
3 configuration_reload no Class configuration_reload · console reload · #642
4 lease_recovery no Lock/lease recovery paths · #702 · #753 · #790
5 worker_restart no Class worker_restart · #663
6 role_runtime_restart no Class role_runtime_restart · console restart · #642/#663
7 connector_restart no Class connector_restart · #663
8 rolling_mcp_restart no Class rolling_mcp_restart · design #668 · attempt log required
9 full_mcp_restart no Class full_mcp_restart · attempt log required
10 host_restart no Class host_restart · attempt log required

Machine-readable source of truth: recovery_playbook.RECOVERY_LADDER and recovery_playbook.ladder_document().

Attempt log shape

Each prior attempt is a mapping:

{
  "action": "client_reconnect",
  "outcome": "insufficient",
  "reason": "transport still closed after IDE reconnect",
  "actor": "prgs-controller-12345",
  "recorded_at": "2026-07-25T21:00:00+00:00"
}

Outcomes that count toward escalation: failed, insufficient, denied, unresolved, timeout, error.

Pass attempts into the coordinator via inventory prior_recovery_attempts or the MCP tool argument prior_recovery_attempts_json on gitea_request_mcp_restart.

Helper: recovery_playbook.build_attempt_record(...).

Symptom → first rung

recovery_playbook.recommend_actions(symptoms=[...]) maps symptoms such as transport_eof, stale_capability, stale_lease, daemon_corrupt to the narrowest recommended action, then walks the ladder. Soft recommendations never replace the hard gate on broad restarts.

Enforcement points

  1. recovery_playbook.assess_escalation — pure gate.
  2. restart_coordinator.evaluate_restart_impact — when restart_class is set (policy-enforced path), broad classes require the gate; report fields attempt_log_satisfied, playbook_escalation, break_glass.
  3. gitea_request_mcp_restart — accepts attempt JSON and env-authorized break-glass; never restarts a process.

Metrics

recovery_playbook.recovery_metrics(attempts) reports the fraction of successful recoveries that avoided full/host restart (fraction_avoided_full_restart).

Non-goals

  • HA multi-instance execution (#668 design only here).
  • Normalizing pkill (#630 contamination stays forbidden).
  • Silent mutation of leases or processes from the playbook itself.

Manual process kills

Remain forbidden and contaminating (#630). The playbook never recommends them.