diff --git a/docs/architecture/mcp-ha-rolling-restart.md b/docs/architecture/mcp-ha-rolling-restart.md new file mode 100644 index 0000000..fc82f99 --- /dev/null +++ b/docs/architecture/mcp-ha-rolling-restart.md @@ -0,0 +1,167 @@ +# ADR: High-availability and rolling-restart architecture for Gitea MCP control plane + +- **Status:** Proposed (Design ADR under [#668](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/668)) +- **Date:** 2026-07-25 +- **Tracking Issue:** [#668](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/668) +- **Policy Version:** `mcp-ha-rolling-restart/v1` +- **Related:** + - Parent: [#655](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/655) — Governed MCP restart coordination and zero-disruption recovery + - Governance Policy: [#656](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/656) / `docs/architecture/mcp-restart-governance.md` + - Control-Plane DB Substrate: [#613](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/613) / `docs/architecture/control-plane-db-substrate.md` + - Runtime Policy: [#615](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/615) / `docs/architecture/mcp-stable-control-runtime-policy-adr.md` + - Product Vision: [#652](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/652) (Phase 5 Maturity) + - Delivery Roadmap: [#653](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/653) + +--- + +## 1. Context & Problem Statement + +The Gitea MCP server operates as the authoritative **control plane** for managing issues, Pull Requests, code mutations, formal reviews, and workflow reconciliations. Under single-process governance ([#656](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/656)), process restarts are strictly controlled using pre-flight checks, drain phases, and operator approvals. + +However, a single-instance control plane inherently presents fundamental constraints: + +1. **Downtime during updates:** Even a perfectly executed single-process drain requires a window where incoming client requests must be paused or rejected while the server binary or python environment reloads. +2. **Single point of failure:** Infrastructure issues, process crashes, or unhandled host-level terminations immediately disconnect active LLM sessions and leave transient workflows incomplete. +3. **Multi-agent concurrency bottlenecks:** High volumes of concurrent multi-LLM tasks put all lock management, lease allocation, and Gitea API interactions through a single process event loop. + +To achieve true zero-disruption operation and seamless rolling deployments without stopping active work, the system requires a high-availability (HA), multi-instance MCP architecture. + +--- + +## 2. Architectural Principles & Non-Goals + +### 2.1 Core Architectural Principles +* **Gitea as Canonical Work SoT:** Gitea remains the ultimate System of Record (SoT) for issue states, pull requests, labels, and audit comments. The MCP control plane does not duplicate domain entities. +* **Control-Plane DB as Multi-Instance State Substrate:** The control-plane SQLite/durable database ([#613](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/613)) acts as the single source of truth for workflow leases, session tokens, assignment records, and lock fences across all MCP nodes. +* **Stateless Worker Nodes:** MCP role server processes (`gitea-author`, `gitea-reviewer`, `gitea-merger`, `gitea-reconciler`, `gitea-controller`) maintain no unique in-memory state; any node can handle any request given a valid session resume token. +* **Fail-Closed Split-Brain Defense:** In any network partition or quorum loss scenario, nodes must fail closed rather than risk double-mutations or conflicting Gitea states. + +### 2.2 Non-Goals +* **Replacing Gitea:** We do not replace Gitea issue/PR tracking with an independent database. +* **Immediate Multi-Node Cluster Execution in v1:** This ADR defines the target architecture and phased roadmap; immediate implementation occurs incrementally post-[#655] v1. + +--- + +## 3. High-Availability & Rolling-Restart Architecture + +### 3.1 Architecture Overview + +``` + +----------------------------+ + | LLM Clients / IDE Sessions | + +--------------+-------------+ + | + v + +----------------------------+ + | HA Proxy / Router | + | (Health-based & Affinity) | + +------+--------------+------+ + | | + +--------------+ +--------------+ + v v + +--------------------+ +--------------------+ + | MCP Instance Node A| | MCP Instance Node B| + | (Version N) | | (Version N+1) | + +---------+----------+ +---------+----------+ + | | + +----------------------+----------------------+ + | + v + +----------------------------+ + | Control-Plane DB Substrate| + | (Shared Lease & Locks) | + +--------------+-------------+ + | + v + +----------------------------+ + | Gitea API | + +----------------------------+ +``` + +--- + +### 3.2 Key System Components + +#### A. Multiple MCP Instance Cohorts +* The control plane runs across $N \ge 2$ redundant process nodes. +* Dual-namespace deployment allows running the old version (Node A) alongside a updated version (Node B) during rolling upgrades. + +#### B. Shared Durable Session Storage & Resume Tokens +* Session context, preflight verification proofs, and capability resolution states are stored in the shared control-plane database. +* Client requests carry an explicit `session_id` and `resume_token`. If an MCP instance restarts or a request routes to a different instance, the target node validates the token against the database without requiring full session re-initialization. + +#### C. Shared Lease Authority & Fencing Counters +* Workflow leases (`gitea_allocate_next_work`, `gitea_adopt_workflow_lease`) use monotonic fencing tokens (`lease_generation_id`). +* When Node B acquires or renews a lease, it increments the generation counter. Any delayed or out-of-order write attempt from Node A using an older generation token is rejected by database constraints. + +#### D. Leader Election & Coordinated Drain +* Node clusters elect a primary coordinator node for administrative background tasks (such as stale lease cleanup or incident Watchdogs). +* During a rolling deployment: + 1. Node B (new version) is launched and registers as healthy. + 2. Router directs new session creations to Node B. + 3. Node A enters `MAINTENANCE_DRAIN` status ([#659](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/659)), completing in-flight mutations while refusing new tasks. + 4. Once all active sessions migrate or complete, Node A shuts down cleanly. + +#### E. Idempotent Mutations & Failover Safety +* All state-changing tool executions (PR creation, review submission, merge operations, label changes) carry a deterministic `idempotency_key`. +* If a network connection flaps or a node fails mid-mutation, the re-issued request with the same `idempotency_key` is recognized by the control-plane substrate, returning the existing recorded result without repeating side effects on Gitea. + +#### F. Schema Version Compatibility +* Database migrations follow non-breaking additive patterns. +* During rolling upgrades where Node A (Version $N$) and Node B (Version $N+1$) run concurrently, both versions operate against the shared schema without structural conflicts. + +--- + +## 4. Split-Brain & Failure Behavior + +### 4.1 Split-Brain Risk Scenarios & Mitigation + +| Scenario | Risk | Mitigation Strategy | +|---|---|---| +| **Network Partition between Nodes** | Both Node A and Node B attempt to process operations for the same issue/PR. | **Generation Fencing:** Lease renewal requires updating the DB generation counter. The node isolated from the DB fails closed immediately. | +| **Stale Node Recovery** | Node A recovers after a long pause and executes a queued mutation. | **Lease Expiry & TTL Fencing:** Transactions verify that `expires_at > NOW()` within the atomic SQLite transaction boundaries. | +| **Database Connection Loss** | Node loses access to shared control-plane DB substrate. | **Strict Fail-Closed:** The node immediately marks all task capabilities as `blocked` and rejects mutation tools until DB connectivity is re-established. | + +--- + +## 5. Phased Implementation Milestones + +```mermaid +flowchart TD + M1[Milestone 1: Shared Control-Plane DB Schema & Resume Tokens] --> M2[Milestone 2: Idempotent Mutation Layer] + M2 --> M3[Milestone 3: Health Routing & Standby Failover] + M3 --> M4[Milestone 4: Active-Active Rolling Deployment & Auto-Drain] +``` + +### Milestone 1: Shared Control-Plane DB Schema & Resume Tokens (Post-#655) +* Extend [#613] Control-Plane DB schema to store multi-instance node heartbeat records and session resume tokens. +* Enable session lookup across instances via `session_id`. + +### Milestone 2: Idempotent Mutation Layer & Lease Fencing +* Add mandatory `idempotency_key` tracking to all Gitea mutation tools. +* Implement monotonic lease fencing counters in `gitea_allocate_next_work` and `gitea_adopt_workflow_lease`. + +### Milestone 3: Health-Based Routing & Active-Passive Standby +* Introduce lightweight proxy/router capable of checking node health endpoints. +* Implement active-standby failover where standby node automatically assumes work if active node fails health checks. + +### Milestone 4: Active-Active Horizontal Deployment & Rolling Upgrade Automation +* Enable true active-active multi-instance execution. +* Integrate automated zero-downtime rolling upgrades coordinated with `gitea_request_mcp_restart` maintenance drain. + +--- + +## 6. Observability & Audit Requirements + +High-availability control plane operations must expose clear telemetry and audit trails: + +* **Node Registry Telemetry:** Active nodes, version numbers, uptime, and heartbeat timestamps reported via `gitea_get_runtime_context`. +* **Lease Fencing Metrics:** Tracking lease acquire latency, fence rejection counts, and lease handoff durations. +* **Failover & Re-route Audit Logs:** Durable logging of session migrations between nodes, drain initiation, and process retirement events. + +--- + +## 7. Tradeoffs & Accepted Risks + +* **Increased Architectural Complexity:** Moving from a single process to a multi-instance control plane requires robust DB locking, proxy routing, and migration governance. +* **Database Dependency:** The control-plane database substrate becomes a critical shared dependency for multi-node deployments. High availability for the underlying SQLite file system / DB must be guaranteed. diff --git a/docs/incidents/670-direct-master-commit-2fa97c26.md b/docs/incidents/670-direct-master-commit-2fa97c26.md new file mode 100644 index 0000000..b0b7ca3 --- /dev/null +++ b/docs/incidents/670-direct-master-commit-2fa97c26.md @@ -0,0 +1,83 @@ +# Incident #670: bare direct-to-master commit `2fa97c26` (retroactive audit) + +Status: verified; disposition recommendation: **accept as-is, no revert** (final +disposition owned by controller per issue #670). + +## Summary + +Commit `2fa97c26fbda555a1a83930ca5fdcea9d8e47b50` +(`fix(mcp): load dotenv relative to project root`) landed on `prgs/master` +as a single-parent commit with no PR wrapper and no review record, bypassing +the sanctioned issue → branch → PR → review → merge workflow. It was +discovered during the PR #654 post-merge audit. PR #654 itself merged +cleanly via the Gitea API and did **not** introduce this commit. + +## Verification evidence (acceptance criteria 1–3) + +- **AC1 — present on `prgs/master`: yes.** + `git merge-base --is-ancestor 2fa97c26fbda555a1a83930ca5fdcea9d8e47b50 prgs/master` → true. +- **AC2 — no PR or review record: confirmed.** + The commit is a single-parent, non-merge commit sitting directly on + first-parent master between the #629 merge (`5ab5fe85`) and the #654 + merge (`ec903b0d`). A PR landing on master produces a merge commit (or a + PR-linked head); neither exists here. The controller audit at issue-create + time also found no PR wrapper and no review record for this SHA. +- **AC3 — changed files and diff summary: confirmed.** + `gitea_auth.py | 5 +++--` (+3/−2). Single parent + `5ab5fe8583c07134d55dadf09381aecb67df246e`. The change moves + `PROJECT_ROOT` derivation above `load_dotenv()` and loads + `.env` relative to the project root instead of the process CWD. + +## AC4 — why no immediate revert + +- The dotenv fix is intentional and required for correct runtime behavior: + without it, `load_dotenv()` resolves `.env` against the process working + directory, which breaks MCP server launches whose CWD is not the project + root. +- The change is small (+3/−2), self-contained in `gitea_auth.py`, and has + been running on master without incident since 2026-07-10. +- Reverting would re-introduce a real bug to remove a provenance defect — + the wrong trade. Provenance is repaired retroactively by this document, + issue #670, and the hardening landed under #671. +- If the controller later judges the change unsafe, a separate + revert/repair issue is the sanctioned path (issue #670, recommended + disposition option 4). + +## AC5 — workflow-hardening linkage + +Prevention already landed: **issue #671** (closed) +*“Block direct pushes to stable branches from MCP workflow sessions”*, +implemented by commit `5933d87647656643a67a50331c4c7b06ea751dad` +(`feat(guard): block direct stable-branch pushes from MCP workflow sessions`). + +Shipped guardrails include: + +- `gitea_record_stable_branch_push_attempt` — classifies proposed commands + for direct stable-branch push intent (`git push master`, + refspecs, `HEAD:master`, `--force`, dry-run intent, `:master` delete), + plus root/control-checkout local commits not carried by an issue branch, + and writes a durable `stable_branch_contamination` marker. +- `gitea_audit_stable_branch_contamination` — reconciler-only audit/clear + path; a contaminated worker session cannot self-clear. +- Review/merge/close/completion mutations fail closed while a + contamination marker is active. + +## AC6 — PR #654 was not the source + +- `2fa97c26` is the **first parent** of the #654 merge commit + `ec903b0d619e7a27d24aed272a890f4e5d381411`; it predates the #654 merge. +- First-parent history `5ab5fe8..ec903b0`: + `2fa97c2 fix(mcp): load dotenv relative to project root` followed by + `ec903b0 Merge pull request 'feat: lifecycle role/hazard labels ... (#603)' (#654)`. +- The #654 merger audit confirmed `ec903b0d` was a valid Gitea-API merge, + the `git push prgs master` attempt during that run was a no-op, and the + net change `2fa97c2..ec903b0` contained only the reviewed #603 + lifecycle-label files. +- Conclusion: #654 merged reviewed content only; the unauthorized-path + defect is solely the earlier bare commit `2fa97c26`. + +## Explicit non-actions (unchanged by this audit) + +- No revert of `2fa97c26`. +- No force-push or history rewrite. +- No master mutation from the audit session. diff --git a/docs/mcp-recovery-playbook.md b/docs/mcp-recovery-playbook.md new file mode 100644 index 0000000..b589b56 --- /dev/null +++ b/docs/mcp-recovery-playbook.md @@ -0,0 +1,94 @@ +# MCP scoped recovery playbook (#669) + +**Parent:** [#655](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/655) +**Vision / roadmap:** [#652](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/652) · [#653](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/653) +**Class matrix:** [#663](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/663) · `docs/mcp-restart-classes.md` +**Coordinator:** [#658](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/658) · `restart_coordinator.py` +**Audit lineage:** [#665](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/665) + +## Decision + +Full-server MCP reset is a **last resort**. Prefer the narrowest recovery that +can clear the symptom. The coordinator **refuses** `rolling_mcp_restart`, +`full_mcp_restart`, and `host_restart` unless: + +1. The inventory carries a prior **attempt log** of at least one *insufficient* + narrower recovery, **or** +2. **Break-glass** is authorized + (`request_break_glass` + `GITEA_BREAKGLASS_RESTART_AUTHORIZATION`). + +Break-glass still never bypasses the #663 class matrix (role/permission). + +## Ladder (narrow → broad) + +| Rank | Action | Self-service | Implementation / delegation | +|---:|---|---|---| +| 0 | `client_reconnect` | yes | Host auto-reconnect / client reconnect · [#584](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/584) · `docs/mcp-namespace-eof-recovery.md` | +| 1 | `capability_refresh` | yes | `gitea_resolve_task_capability` + `gitea_whoami` · [#610](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/610) · [#685](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/685) | +| 2 | `session_reconnect` | yes | Runtime rebind + explicit `worktree_path` · [#543](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/543) · [#618](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/618) | +| 3 | `configuration_reload` | no | Class `configuration_reload` · console reload · [#642](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/642) | +| 4 | `lease_recovery` | no | Lock/lease recovery paths · [#702](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/702) · [#753](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/753) · [#790](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/790) | +| 5 | `worker_restart` | no | Class `worker_restart` · #663 | +| 6 | `role_runtime_restart` | no | Class `role_runtime_restart` · console restart · #642/#663 | +| 7 | `connector_restart` | no | Class `connector_restart` · #663 | +| 8 | `rolling_mcp_restart` | no | Class `rolling_mcp_restart` · design [#668](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/668) · **attempt log required** | +| 9 | `full_mcp_restart` | no | Class `full_mcp_restart` · **attempt log required** | +| 10 | `host_restart` | no | Class `host_restart` · **attempt log required** | + +Machine-readable source of truth: `recovery_playbook.RECOVERY_LADDER` and +`recovery_playbook.ladder_document()`. + +## Attempt log shape + +Each prior attempt is a mapping: + +```json +{ + "action": "client_reconnect", + "outcome": "insufficient", + "reason": "transport still closed after IDE reconnect", + "actor": "prgs-controller-12345", + "recorded_at": "2026-07-25T21:00:00+00:00" +} +``` + +Outcomes that count toward escalation: `failed`, `insufficient`, `denied`, +`unresolved`, `timeout`, `error`. + +Pass attempts into the coordinator via inventory +`prior_recovery_attempts` or the MCP tool argument +`prior_recovery_attempts_json` on `gitea_request_mcp_restart`. + +Helper: `recovery_playbook.build_attempt_record(...)`. + +## Symptom → first rung + +`recovery_playbook.recommend_actions(symptoms=[...])` maps symptoms such as +`transport_eof`, `stale_capability`, `stale_lease`, `daemon_corrupt` to the +narrowest recommended action, then walks the ladder. Soft recommendations +never replace the hard gate on broad restarts. + +## Enforcement points + +1. **`recovery_playbook.assess_escalation`** — pure gate. +2. **`restart_coordinator.evaluate_restart_impact`** — when `restart_class` is + set (policy-enforced path), broad classes require the gate; report fields + `attempt_log_satisfied`, `playbook_escalation`, `break_glass`. +3. **`gitea_request_mcp_restart`** — accepts attempt JSON and env-authorized + break-glass; never restarts a process. + +## Metrics + +`recovery_playbook.recovery_metrics(attempts)` reports the fraction of +successful recoveries that avoided full/host restart +(`fraction_avoided_full_restart`). + +## Non-goals + +* HA multi-instance execution (#668 design only here). +* Normalizing `pkill` (#630 contamination stays forbidden). +* Silent mutation of leases or processes from the playbook itself. + +## Manual process kills + +Remain forbidden and contaminating (#630). The playbook never recommends them. diff --git a/docs/mcp-restart-coordinator.md b/docs/mcp-restart-coordinator.md index 628895f..684663d 100644 --- a/docs/mcp-restart-coordinator.md +++ b/docs/mcp-restart-coordinator.md @@ -95,12 +95,18 @@ gitea_request_mcp_restart(remote, host, org, repo, target_session_id=None, target_role=None, target_connector=None, drain_proof_json=None, - request_break_glass=False) + request_break_glass=False, + prior_recovery_attempts_json=None) ``` It **never restarts anything**: `apply_supported` is always `false` and `restart_performed` is always `false`. +`prior_recovery_attempts_json` (#669) is an optional JSON array of prior +narrow recovery attempts. Rolling / full / host classes require at least one +*insufficient* narrower attempt (or authorized break-glass). See +`docs/mcp-recovery-playbook.md`. + ### Dry-run versus apply | Call | Behavior | @@ -112,9 +118,10 @@ It **never restarts anything**: `apply_supported` is always `false` and An apply requires **both** authorizations, and they are independent: -1. **Restart-class authorization** (#663) — the requester's role and permissions - must allow the requested class, the class's approval requirement must be - satisfied, and any target-scoped class must name its target. Failing any of +1. **Restart-class authorization** (#663 / #669) — the requester's role and + permissions must allow the requested class, the class's approval requirement + must be satisfied, any target-scoped class must name its target, and broad + classes must satisfy the recovery-playbook attempt-log gate. Failing any of these makes `allow_restart` `false`. 2. **Drain-proof gate** (#661) — a valid, unexpired, clean proof bound to the current impact fingerprint, or an authorized break-glass. @@ -127,7 +134,10 @@ the authorization that produced it. ### Break-glass -Break-glass bypasses the **drain proof only** — never the restart-class matrix. +Break-glass bypasses the **drain proof only** — never the restart-class matrix +(role/permission). Separately, authorized break-glass also satisfies the #669 +attempt-log requirement for broad restarts (rolling/full/host), because that +gate is not a class-matrix permission check. It is honoured solely when `request_break_glass` is set *and* the environment carries `GITEA_BREAKGLASS_RESTART_AUTHORIZATION`; like operator override, the tool argument expresses caller intent and cannot be self-asserted by a worker diff --git a/docs/sanctioned-recovery-playbooks.md b/docs/sanctioned-recovery-playbooks.md new file mode 100644 index 0000000..6b5731f --- /dev/null +++ b/docs/sanctioned-recovery-playbooks.md @@ -0,0 +1,64 @@ +# Sanctioned Recovery Playbooks & Controls (Phase 2 #644) + +## Overview + +Stale runtimes, worktree binding mismatches, and un-reconciled merged branches previously required expert manual shell recovery. Manual process kills (`pkill -f mcp_server.py`) are strictly forbidden and classified as runtime contamination ([#630](sanctioned-restart-controls.md)). + +Phase 2 introduces **sanctioned recovery playbooks and controls** into the Web Console: +- **Diagnose**: Surface stale runtimes, worktree binding errors, contamination markers, and worktree anomalies via health & inventory APIs. +- **Preview**: Render mutation ledgers and exact confirmation phrases for recovery playbooks. +- **Confirm & Apply**: Execute sanctioned recovery actions through gated, audited paths. +- **Verify**: Revalidate control-plane state post-recovery before claiming clean status. + +--- + +## Recovery Playbook Taxonomy + +| Playbook ID | Action ID | Minimum Role | Target / Scope | Description | +|---|---|---|---|---| +| `clear_stale_binding` | `system.clear_stale_binding` | Operator | Active worktree binding | Clear provably missing or superseded `GITEA_ACTIVE_WORKTREE` binding ([#702](../stale_binding_recovery.py)). | +| `rebind_session_worktree` | `system.rebind_session_worktree` | Operator | Session worktree | Rebind or synchronize session worktree to verified lease worktree ([#864](../dirty_same_claimant_session_rebind.py)). | +| `reconcile_cleanups` | `system.reconcile_cleanups` | Controller | Worktree hygiene | Execute reconciler cleanup preview and apply for merged/superseded PR branches. | +| `sanctioned_restart` | `system.restart_namespace` | Admin | MCP Namespace | Restart MCP daemon gracefully via host supervisor ([#642](sanctioned-restart-controls.md)). | + +--- + +## Wizard Workflow (Diagnose → Preview → Confirm → Verify) + +### 1. Diagnose (`GET /api/v1/system/recovery/diagnose`) +Runs control-plane diagnostics: +- **Stale Runtime**: Mismatch between running daemon HEAD, local checkout HEAD, and remote-tracking HEAD. +- **Worktree Binding**: Missing path (`provably_stale_missing_path`), unverified inherited binding (`unverified_inherited`), or superseded binding (`superseded_by_session_lease`). +- **Contamination**: Checks for live contamination markers from unmanaged process kills. +- **Worktree Anomalies**: Scans `branches/` directory for un-reconciled cleanups or missing preserved worktrees. + +Returns `RecoveryDiagnosis` with eligible playbooks. + +### 2. Preview (`POST /api/v1/system/recovery/preview`) +Takes `playbook_id` and optional `target`/`params`. +Returns: +- **Mutation Ledger**: Step-by-step sequence of actions. +- **Confirmation Phrase**: Exact phrase required to authorize execution (e.g., `confirm clear_stale_binding`). +- **Authorization Decision**: RBAC check against the operator's principal. + +### 3. Apply (`POST /api/v1/system/recovery/apply`) +Requires `playbook_id` and matching `confirmation` phrase. Gates run in this order, and each fails closed before anything is mutated: + +1. **RBAC and execution phase** (`console_authz.authorize(..., for_execution=True)`). The phase branch only applies when `for_execution` is set. While `ACTIVE_PHASE` is `1`, every phase-2 recovery action is refused with `phase_not_active`, so no recovery playbook writes yet. Preview reports the same decision under `execution_authorization` / `execution_blocked_reason`. +2. **Confirmation phrase** (`confirmation_matches`). +3. **Contamination rules** ([#630](sanctioned-restart-controls.md)): the live marker is read from the session inventory and assessed under the gated task key `console_recovery_apply`. A contaminated runtime must be cleared through the reconciler cleanup playbook, which is the one playbook exempted from this gate because it is the designated remedy. The marker is also forwarded to `sanctioned_restart.execute_restart`, so a restart cannot launder a contaminated runtime. + +Apply then executes the sanctioned recovery logic against the **live** process environment — not a copy — and records an audit entry in `console_audit`. A playbook that leaves the binding unchanged reports `performed: false`; `binding_before`, `binding_after`, and `binding_changed` are returned so a no-op cannot read as success. + +Apply does **not** enforce master parity. Parity is reported by Diagnose ([#610](../master_parity_gate.py)) as evidence for the operator; it is not a precondition of this endpoint. + +### 4. Verify (`POST /api/v1/system/recovery/verify`) +Re-evaluates control-plane diagnostics post-recovery and **reports** `clean`, `stale_runtime_clean`, `binding_clean`, `binding_classification`, and `contamination_clean`. It reports; it does not assert or block. State is read fresh rather than from the mapping a mutation just wrote. An `unverified_inherited` binding is reported as not clean, because unproven is not clean. + +--- + +## Safety & Governance Principles + +1. **No Manual `pkill`**: Direct process killing remains forbidden and is recorded as contamination. +2. **Auditability**: Every recovery preview and execution is logged in the console audit trail. +3. **Master Parity & Dual Control**: High-privilege recovery actions require controller/admin roles and explicit confirmation phrases. diff --git a/docs/webui-authz-audit.md b/docs/webui-authz-audit.md index ff9cc38..70d7a3d 100644 --- a/docs/webui-authz-audit.md +++ b/docs/webui-authz-audit.md @@ -94,6 +94,9 @@ already define, and a regression test asserts each mapping matches. | `record_analytics_usage` | operator | gated_write | `runtime.record_analytics_usage` | Yes | No | No | 2 | | `system.reload_namespace` | controller | privileged | `runtime.reload_namespace` | Yes | No | No | 2 | | `system.restart_namespace` | admin | destructive | `runtime.restart_namespace` | Yes | **Yes** | **Yes** | 2 | +| `system.clear_stale_binding` | operator | gated_write | `gitea.read` | Yes | No | No | 2 | +| `system.rebind_session_worktree` | operator | gated_write | `gitea.read` | Yes | No | No | 2 | +| `system.reconcile_cleanups` | controller | privileged | `gitea.pr.close` | Yes | No | No | 2 | | `initiate_workflow` | operator | gated_write | `gitea.read` | Yes | No | No | 2 | **Dual control** means the acting principal may not be the sole authority: a diff --git a/docs/webui-notifications.md b/docs/webui-notifications.md new file mode 100644 index 0000000..7ef2e05 --- /dev/null +++ b/docs/webui-notifications.md @@ -0,0 +1,81 @@ +# Web Console: Notifications & Human-Attention Routing (#648) + +- **Status:** Phase 3 Live +- **Tracking Issue:** [#648](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/648) +- **Parent Epic:** [#631](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/631) +- **Attention Boundary Reference:** [#628](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/628) + +--- + +## 1. Overview + +The **Notifications & Human-Attention Console** (`/notifications`, `/api/v1/notifications`) provides intelligent event classification and human-attention routing for autonomous workflow operations. + +To prevent alert fatigue while ensuring critical escalation boundaries are never missed, events are classified into three distinct **Attention Classes**: + +1. **`human-required`** (Urgent Escalation Boundary): + - Items requiring immediate human intervention or business decisions. + - Triggers: Auth failures, hard stops, irrecoverable state, decision locks, failed report validations, critical probe errors. + - Display: Highlighted in red (`badge-blocked`) with a `HUMAN REQUIRED` badge. + +2. **`operator`** (Operational Inbox): + - Items requiring controller or operator review/triage during routine execution. + - Triggers: Blocked PRs (merge conflicts), stale leases, duplicate PRs on issues, unassigned ready work. + - Display: Displayed in orange/yellow (`badge-claimed`). + +3. **`routine`** (Background Workflow Transitions): + - Normal, healthy workflow transitions and state progressions. + - Triggers: Active PRs/issues in standard state, clean branch creation, routine heartbeats. + - Display: Filtered out of default inbox views to eliminate notification spam; viewable on demand via the "Routine" or "All" tab. + +--- + +## 2. API Endpoints + +### `GET /api/v1/notifications` +*Compatibility Alias:* `GET /api/notifications` + +#### Query Parameters: +- `project_id` (optional): Filter notifications by project ID. +- `attention_class` (optional): `inbox` (default: human-required + operator), `human-required`, `operator`, `routine`, `all`. + +#### Example JSON Response: +```json +{ + "project_id": "gitea-tools", + "repo_label": "Scaled-Tech-Consulting/Gitea-Tools", + "human_required_count": 0, + "operator_count": 2, + "routine_count": 5, + "total_count": 7, + "fetch_error": null, + "inbox_items": [ + { + "id": "notif-pr-block-742", + "attention_class": "operator", + "category": "blocker", + "title": "Blocked PR #742", + "summary": "PR #742 requires merge conflict resolution.", + "work_kind": "pr", + "work_number": 742, + "project_id": "gitea-tools", + "repo_label": "Scaled-Tech-Consulting/Gitea-Tools", + "created_at": "2026-07-25T16:39:47Z", + "deep_link": "/traffic", + "requires_human": false, + "extra": {} + } + ], + "all_items": [...] +} +``` + +--- + +## 3. UI Navigation + +- Access via the **Traffic** navigation menu: **Traffic → Notifications**. +- The main view displays: + - **Metrics Summary Bar**: Highlighting counts for Human Required, Operator Inbox, and Routine items. + - **Attention Filter Tabs**: Toggle between Inbox (Human + Operator), Human Required, Operator, Routine, and All. + - **Structured Event Table**: Displays category, title, summary, work item links, and timestamps. diff --git a/docs/webui-restart-console.md b/docs/webui-restart-console.md new file mode 100644 index 0000000..12dd427 --- /dev/null +++ b/docs/webui-restart-console.md @@ -0,0 +1,102 @@ +# Web Console: restart status, impact preview, and approval state (#667) + +Phase 1 of the console restart surface. It consumes the #655 coordinator +substrate and displays it. It performs no restart, reload, drain, approval, or +process action, and it registers no write endpoint. + +Issue #667's rollout is explicit — *status views first, write approval after the +backend gates are green* — and this change delivers only the status half. + +## Surfaces + +| Path | Method | Purpose | +|------|--------|---------| +| `/runtime/restart` | GET | Restart status page | +| `/api/v1/system/restart/status` | GET | Same snapshot as JSON | + +Both accept an optional `restart_class` query parameter (default +`full_mcp_restart`). An unrecognised class is not an error: the coordinator +resolves it as unknown and fails closed, and the page shows the resulting deny. + +Neither path accepts `POST`; a write attempt returns `405`, and a test asserts +it. + +## What it shows + +* **Impact preview (#658)** — verdict, blast radius, affected sessions, leases, + critical sections, mutations, and the counts behind them, evaluated + `dry_run=True` against live control-plane state. +* **Drain proof (#661)** — verification of a supplied proof: valid, clean, + expired, tampered, and the reasons behind a refusal. +* **Post-restart reconcile (#662)** — the most recent completion proof, its + overall status, and which dimensions still require follow-up. +* **Restart classes (#663)** — the least-privilege matrix, with *you may + request* and *you may execute* computed for the viewing role rather than for a + generic operator. +* **Approval controls (#633)** — the authorization state of + `system.restart_namespace` and `system.reload_namespace`. +* **Break-glass (#664)** — declared and marked unavailable; see below. + +## Three rules this surface holds itself to + +A status page that is wrong is worse than one that is missing, because an +operator acts on it. Three properties are enforced by tests, and each was +verified by reverting the guard and watching a test fail. + +### An unreadable source reports unavailable, never green + +Every source carries its own `SourceStatus`. Nothing substitutes a default, +placeholder, or self-comparison for a reading that failed. An unreadable +control-plane database yields `inventory_complete: false`, which the coordinator +itself turns into a fail-closed verdict, and the page says the blast radius is +unknown rather than showing an empty affected-sessions table. + +An absent drain proof is reported as absent — not as a pass. The #661 gate +authorizes a restart only against a valid, unexpired, clean proof, so no proof +is precisely the state that gate denies on. + +### Authorization is asked the way execution would ask it + +Every probe passes `for_execution=True`. + +Asked without it, an admin is `allowed` for `system.restart_namespace`. On a +control surface that reads as a live button. Asked the way an execution attempt +would ask, the same principal is refused `phase_not_active`, because the console +is in Phase 1 and the action is Phase 2. This surface reports the second answer. + +`execution_enabled` is therefore `false` for every action and every role today, +and a test asserts that across the whole role matrix. + +### The control-plane database is opened read-only + +`ControlPlaneDB()` creates directories and runs migrations on construction — a +write. This surface never constructs one. It opens the sqlite file with +`mode=ro`, exactly as `webui/inventory.py` does, and treats a missing file as +missing authority rather than as an empty inventory. + +The test that protects this points at a path inside a directory that already +exists, so a read-write `connect` would really create the file. A nested +missing-directory path would have passed for the wrong reason. + +## Break-glass is declared, not offered + +The break-glass workflow (#664) is not available on this branch's base. The +panel is rendered to operator-class roles as **unavailable**, naming the issue +that tracks it. It is not silently omitted, because an operator who has been +told a governance path exists needs to see that it is not wired here; and it is +not rendered as a control, because there is nothing behind it. + +Unprivileged viewers see only a note that the surface is operator-class. + +## Redaction and escaping + +Every interpolated value passes through `_esc` (`html.escape(..., quote=True)`). +Free-form text and anything that can carry a filesystem path additionally passes +through `webui.inventory.scrub_text`, which redacts credential-shaped tokens +inside a string rather than only at its start. The impact payload is passed +through `webui.inventory.scrub` before rendering. + +## Linkage + +Parent #655 · extends #642 · consumes #658, #661, #662, #663 · RBAC #633 · +console #631 · vision #652 · roadmap #653 · break-glass #664. diff --git a/gitea_mcp_server.py b/gitea_mcp_server.py index 90a54df..5ca67c1 100644 --- a/gitea_mcp_server.py +++ b/gitea_mcp_server.py @@ -22815,8 +22815,9 @@ def gitea_request_mcp_restart( target_connector: str | None = None, drain_proof_json: str | None = None, request_break_glass: bool = False, + prior_recovery_attempts_json: str | None = None, ) -> dict: - """Evaluate a proposed MCP restart and return an impact preview (#658). + """Evaluate a proposed MCP restart and return an impact preview (#658/#669). Central restart coordinator: resolves the requested restart class, gathers live control-plane state (sessions, @@ -22840,10 +22841,16 @@ def gitea_request_mcp_restart( independent — the drain gate proves the blast radius was drained and knows nothing about whether this requester may request this class — so a class the matrix denied never reports an authorized apply. Break-glass bypasses the - drain proof only; it never bypasses the class matrix. ``apply_gate`` carries + drain proof and, when env-authorized, the #669 attempt-log requirement for + broad restarts; it never bypasses the class matrix. ``apply_gate`` carries ``drain_gate_allow`` and ``restart_class_authorized`` so a denial is attributable to the authorization that produced it. + ``prior_recovery_attempts_json`` (#669) is an optional JSON array of prior + narrow recovery attempts ``{action, outcome, reason, ...}``. Rolling / full + / host restart classes require at least one *insufficient* narrower attempt + unless break-glass is authorized. + Operator override authority is read from the process environment (``GITEA_OPERATOR_RESTART_OVERRIDE_AUTHORIZATION``), never self-asserted by the requesting session: ``request_override`` only expresses caller intent @@ -22949,12 +22956,37 @@ def gitea_request_mcp_restart( requester_role ) + prior_recovery_attempts: list[dict] = [] + if prior_recovery_attempts_json: + try: + parsed_attempts = json.loads(prior_recovery_attempts_json) + if isinstance(parsed_attempts, list): + prior_recovery_attempts = [ + dict(a) for a in parsed_attempts if isinstance(a, dict) + ] + else: + incomplete_reasons.append( + "prior_recovery_attempts_json must be a JSON array (#669)" + ) + inventory_complete = False + except (ValueError, TypeError) as exc: + incomplete_reasons.append( + f"invalid prior_recovery_attempts_json: {_redact(str(exc))}" + ) + inventory_complete = False + + break_glass_authorized = bool( + (os.environ.get("GITEA_BREAKGLASS_RESTART_AUTHORIZATION") or "").strip() + ) + break_glass = bool(request_break_glass and break_glass_authorized) + inventory = { "sessions": sessions, "leases": leases, "terminal_lock": terminal_lock, "inventory_complete": inventory_complete, "incomplete_reasons": incomplete_reasons, + "prior_recovery_attempts": prior_recovery_attempts, } report = restart_coordinator.evaluate_restart_impact( @@ -22970,6 +23002,7 @@ def gitea_request_mcp_restart( target_session_id=target_session_id, target_role=target_role, target_connector=target_connector, + break_glass=break_glass, ) payload = report.as_dict() @@ -23002,13 +23035,6 @@ def gitea_request_mcp_restart( except (ValueError, TypeError) as exc: proof_parse_error = f"invalid drain_proof_json: {_redact(str(exc))}" - break_glass_authorized = bool( - ( - os.environ.get("GITEA_BREAKGLASS_RESTART_AUTHORIZATION") or "" - ).strip() - ) - break_glass = bool(request_break_glass and break_glass_authorized) - expected_fp = drain_proof.impact_fingerprint(report.as_dict()) gate = drain_proof.gate_apply_restart( proof=proof_obj, diff --git a/recovery_playbook.py b/recovery_playbook.py new file mode 100644 index 0000000..268a3be --- /dev/null +++ b/recovery_playbook.py @@ -0,0 +1,583 @@ +"""Scoped MCP recovery playbook (#669). + +Operational recovery must prefer the *narrowest* action that can fix the +symptom. Full MCP / host restarts are last-resort rungs on a documented +ladder; the coordinator refuses those rungs unless a prior attempt log +shows narrower recoveries already failed (or break-glass is authorized). + +This module is pure classification and recommendation: + +* No network, filesystem, or process I/O. +* Never restarts anything. +* Narrow recovery *execution* is delegated to existing tools/docs (linked + per rung) — the playbook records which rung to try next and whether + escalation to a broad restart is allowed. + +Design lineage: umbrella #655, class matrix #663, coordinator #658, +auto-reconnect #584, stale-runtime #610, contamination #630, audit #665. +Vision #652 / roadmap #653. +""" + +from __future__ import annotations + +from dataclasses import dataclass, field +from datetime import datetime, timezone +from enum import Enum +from typing import Any, Mapping, Sequence + +PLAYBOOK_VERSION = "1.0.0-issue-669" + +# Attempt outcomes that count as "tried and insufficient" for escalation. +INSUFFICIENT_OUTCOMES = frozenset( + { + "failed", + "insufficient", + "denied", + "unresolved", + "timeout", + "error", + } +) + +# Break-glass / operator override still records that the ladder was skipped. +OUTCOME_BREAK_GLASS = "break_glass" +OUTCOME_SUCCESS = "success" +OUTCOME_SKIPPED = "skipped" + + +class RecoveryAction(str, Enum): + """Ordered recovery ladder (narrow → broad).""" + + CLIENT_RECONNECT = "client_reconnect" + CAPABILITY_REFRESH = "capability_refresh" + SESSION_RECONNECT = "session_reconnect" + CONFIGURATION_RELOAD = "configuration_reload" + LEASE_RECOVERY = "lease_recovery" + WORKER_RESTART = "worker_restart" + ROLE_RUNTIME_RESTART = "role_runtime_restart" + CONNECTOR_RESTART = "connector_restart" + ROLLING_MCP_RESTART = "rolling_mcp_restart" + FULL_MCP_RESTART = "full_mcp_restart" + HOST_RESTART = "host_restart" + + +# Classes that require a prior narrow-attempt log (unless break-glass). +BROAD_RESTART_ACTIONS: frozenset[RecoveryAction] = frozenset( + { + RecoveryAction.ROLLING_MCP_RESTART, + RecoveryAction.FULL_MCP_RESTART, + RecoveryAction.HOST_RESTART, + } +) + +# Map #663 restart_class strings onto playbook actions. +RESTART_CLASS_TO_ACTION: dict[str, RecoveryAction] = { + "client_reconnect": RecoveryAction.CLIENT_RECONNECT, + "session_reconnect": RecoveryAction.SESSION_RECONNECT, + "configuration_reload": RecoveryAction.CONFIGURATION_RELOAD, + "worker_restart": RecoveryAction.WORKER_RESTART, + "role_runtime_restart": RecoveryAction.ROLE_RUNTIME_RESTART, + "connector_restart": RecoveryAction.CONNECTOR_RESTART, + "rolling_mcp_restart": RecoveryAction.ROLLING_MCP_RESTART, + "full_mcp_restart": RecoveryAction.FULL_MCP_RESTART, + "host_restart": RecoveryAction.HOST_RESTART, +} + + +@dataclass(frozen=True) +class RecoveryRung: + """One rung on the recovery ladder.""" + + action: RecoveryAction + rank: int + summary: str + # Existing implementation or explicit delegation target. + implementation: str + issue_links: tuple[str, ...] + self_service: bool + # Restart-class permission when this rung is requested via coordinator. + restart_class: str | None = None + + def as_dict(self) -> dict[str, Any]: + return { + "action": self.action.value, + "rank": self.rank, + "summary": self.summary, + "implementation": self.implementation, + "issue_links": list(self.issue_links), + "self_service": self.self_service, + "restart_class": self.restart_class, + } + + +# Canonical ladder. Rank 0 is narrowest. +RECOVERY_LADDER: tuple[RecoveryRung, ...] = ( + RecoveryRung( + RecoveryAction.CLIENT_RECONNECT, + 0, + "Reconnect the IDE/client MCP transport (EOF / transport flap).", + "Host auto-reconnect or explicit client reconnect; " + "docs/mcp-namespace-eof-recovery.md", + ("#584", "#655"), + True, + "client_reconnect", + ), + RecoveryRung( + RecoveryAction.CAPABILITY_REFRESH, + 1, + "Re-resolve task capability and clear stale permission context.", + "Delegated: gitea_resolve_task_capability + gitea_whoami " + "(no process change).", + ("#610", "#685", "#655"), + True, + None, + ), + RecoveryRung( + RecoveryAction.SESSION_RECONNECT, + 2, + "Rebind identity, workspace, and namespace for one session.", + "Delegated: gitea_get_runtime_context + explicit worktree_path " + "rebind (#618); docs/mcp-namespace-health.md", + ("#543", "#618", "#655"), + True, + "session_reconnect", + ), + RecoveryRung( + RecoveryAction.CONFIGURATION_RELOAD, + 3, + "Gracefully reload configuration without replacing the daemon.", + "restart_coordinator class configuration_reload; console " + "system.reload_namespace (#642).", + ("#642", "#663", "#655"), + False, + "configuration_reload", + ), + RecoveryRung( + RecoveryAction.LEASE_RECOVERY, + 4, + "Recover or rebind stale leases/locks without a process restart.", + "Delegated: issue lock recovery / lease lifecycle paths " + "(#702, #753, #790).", + ("#702", "#753", "#790", "#655"), + False, + None, + ), + RecoveryRung( + RecoveryAction.WORKER_RESTART, + 5, + "Restart one worker after its own lease and mutation scope drains.", + "restart_coordinator class worker_restart (#663).", + ("#663", "#655"), + False, + "worker_restart", + ), + RecoveryRung( + RecoveryAction.ROLE_RUNTIME_RESTART, + 6, + "Restart one role runtime and re-probe that namespace only.", + "restart_coordinator class role_runtime_restart; console " + "system.restart_namespace (#642).", + ("#642", "#663", "#655"), + False, + "role_runtime_restart", + ), + RecoveryRung( + RecoveryAction.CONNECTOR_RESTART, + 7, + "Restart one connector while unrelated runtimes stay available.", + "restart_coordinator class connector_restart (#663).", + ("#663", "#655"), + False, + "connector_restart", + ), + RecoveryRung( + RecoveryAction.ROLLING_MCP_RESTART, + 8, + "Drain/restart/verify one instance at a time (HA path).", + "restart_coordinator class rolling_mcp_restart; design #668.", + ("#668", "#663", "#655"), + False, + "rolling_mcp_restart", + ), + RecoveryRung( + RecoveryAction.FULL_MCP_RESTART, + 9, + "Full stable-control MCP process restart after verified full drain.", + "restart_coordinator class full_mcp_restart; requires attempt log " + "unless break-glass (#669).", + ("#658", "#661", "#663", "#669", "#655"), + False, + "full_mcp_restart", + ), + RecoveryRung( + RecoveryAction.HOST_RESTART, + 10, + "Host/infrastructure restart — broadest last-resort action.", + "restart_coordinator class host_restart; operator-owned.", + ("#663", "#669", "#655"), + False, + "host_restart", + ), +) + +_LADDER_BY_ACTION: dict[RecoveryAction, RecoveryRung] = { + rung.action: rung for rung in RECOVERY_LADDER +} + +# Symptom tokens → preferred first rung (decision tree, #663 lineage). +SYMPTOM_TO_FIRST_ACTION: dict[str, RecoveryAction] = { + "transport_eof": RecoveryAction.CLIENT_RECONNECT, + "client_closing_eof": RecoveryAction.CLIENT_RECONNECT, + "transport_flap": RecoveryAction.CLIENT_RECONNECT, + "namespace_disconnected": RecoveryAction.CLIENT_RECONNECT, + "stale_capability": RecoveryAction.CAPABILITY_REFRESH, + "permission_stale": RecoveryAction.CAPABILITY_REFRESH, + "runtime_reconnect_required": RecoveryAction.CAPABILITY_REFRESH, + "stale_runtime": RecoveryAction.SESSION_RECONNECT, + "worktree_unbound": RecoveryAction.SESSION_RECONNECT, + "namespace_unhealthy": RecoveryAction.SESSION_RECONNECT, + "config_drift": RecoveryAction.CONFIGURATION_RELOAD, + "profile_misbound": RecoveryAction.CONFIGURATION_RELOAD, + "stale_lease": RecoveryAction.LEASE_RECOVERY, + "dead_pid_lock": RecoveryAction.LEASE_RECOVERY, + "orphan_worktree": RecoveryAction.LEASE_RECOVERY, + "single_worker_stuck": RecoveryAction.WORKER_RESTART, + "role_runtime_dead": RecoveryAction.ROLE_RUNTIME_RESTART, + "connector_dead": RecoveryAction.CONNECTOR_RESTART, + "ha_instance_unhealthy": RecoveryAction.ROLLING_MCP_RESTART, + "daemon_corrupt": RecoveryAction.FULL_MCP_RESTART, + "full_process_deadlock": RecoveryAction.FULL_MCP_RESTART, + "host_unresponsive": RecoveryAction.HOST_RESTART, +} + + +def _utc_now() -> datetime: + return datetime.now(timezone.utc) + + +def resolve_action(value: RecoveryAction | str) -> RecoveryAction: + """Resolve a recovery action or fail closed for unknown values.""" + if isinstance(value, RecoveryAction): + return value + text = str(value or "").strip() + # Accept #663 restart_class aliases. + if text in RESTART_CLASS_TO_ACTION: + return RESTART_CLASS_TO_ACTION[text] + try: + return RecoveryAction(text) + except ValueError as exc: + raise ValueError( + f"unknown recovery action {value!r}; deny (fail closed, #669)" + ) from exc + + +def ladder_rank(action: RecoveryAction | str) -> int: + resolved = resolve_action(action) + return _LADDER_BY_ACTION[resolved].rank + + +def rung_for(action: RecoveryAction | str) -> RecoveryRung: + return _LADDER_BY_ACTION[resolve_action(action)] + + +def normalize_attempt(raw: Mapping[str, Any]) -> dict[str, Any] | None: + """Normalize one prior-recovery attempt record; return None if unusable.""" + if not isinstance(raw, Mapping): + return None + action_raw = raw.get("action") or raw.get("recovery_action") or raw.get( + "restart_class" + ) + if not action_raw: + return None + try: + action = resolve_action(str(action_raw)) + except ValueError: + return None + outcome = str( + raw.get("outcome") or raw.get("status") or raw.get("result") or "" + ).strip().lower() + if not outcome: + return None + recorded_at = raw.get("recorded_at") or raw.get("at") or raw.get("timestamp") + reason = str(raw.get("reason") or raw.get("detail") or "").strip() + actor = str(raw.get("actor") or raw.get("session_id") or "").strip() + return { + "action": action.value, + "outcome": outcome, + "reason": reason, + "actor": actor, + "recorded_at": recorded_at, + "rank": ladder_rank(action), + "raw": dict(raw), + } + + +def normalize_attempt_log( + attempts: Sequence[Mapping[str, Any]] | None, +) -> list[dict[str, Any]]: + """Return usable attempt records in ladder order.""" + out: list[dict[str, Any]] = [] + for raw in attempts or (): + norm = normalize_attempt(raw) + if norm is not None: + out.append(norm) + out.sort(key=lambda a: (a["rank"], str(a.get("recorded_at") or ""))) + return out + + +def narrower_insufficient_attempts( + attempts: Sequence[Mapping[str, Any]] | None, + *, + requested: RecoveryAction | str, +) -> list[dict[str, Any]]: + """Return prior attempts narrower than *requested* that were insufficient.""" + target_rank = ladder_rank(requested) + usable = [] + for attempt in normalize_attempt_log(attempts): + if attempt["rank"] >= target_rank: + continue + if attempt["outcome"] in INSUFFICIENT_OUTCOMES: + usable.append(attempt) + return usable + + +@dataclass(frozen=True) +class EscalationAssessment: + """Whether a requested broad recovery may proceed given the attempt log.""" + + requested_action: str + allowed: bool + require_attempt_log: bool + break_glass: bool + reasons: list[str] = field(default_factory=list) + qualifying_attempts: list[dict[str, Any]] = field(default_factory=list) + recommended_next: list[dict[str, Any]] = field(default_factory=list) + playbook_version: str = PLAYBOOK_VERSION + + def as_dict(self) -> dict[str, Any]: + return { + "playbook_version": self.playbook_version, + "requested_action": self.requested_action, + "allowed": self.allowed, + "require_attempt_log": self.require_attempt_log, + "break_glass": self.break_glass, + "reasons": list(self.reasons), + "qualifying_attempts": list(self.qualifying_attempts), + "recommended_next": list(self.recommended_next), + } + + +def assess_escalation( + requested: RecoveryAction | str, + *, + prior_recovery_attempts: Sequence[Mapping[str, Any]] | None = None, + break_glass: bool = False, +) -> EscalationAssessment: + """Gate broad restarts on a prior narrow-attempt log (#669 AC3). + + Narrow / mid-ladder actions do not require a prior attempt log. + ``full_mcp_restart``, ``host_restart``, and ``rolling_mcp_restart`` + require at least one *insufficient* narrower attempt unless + ``break_glass`` is true. + """ + action = resolve_action(requested) + require_log = action in BROAD_RESTART_ACTIONS + reasons: list[str] = [] + qualifying = narrower_insufficient_attempts( + prior_recovery_attempts, requested=action + ) + + if not require_log: + return EscalationAssessment( + requested_action=action.value, + allowed=True, + require_attempt_log=False, + break_glass=bool(break_glass), + reasons=["narrow recovery; attempt log not required"], + qualifying_attempts=qualifying, + recommended_next=[], + ) + + if break_glass: + return EscalationAssessment( + requested_action=action.value, + allowed=True, + require_attempt_log=True, + break_glass=True, + reasons=[ + "break-glass authorized; broad restart permitted without " + "narrow-attempt log (#669)" + ], + qualifying_attempts=qualifying, + recommended_next=[], + ) + + if qualifying: + return EscalationAssessment( + requested_action=action.value, + allowed=True, + require_attempt_log=True, + break_glass=False, + reasons=[ + f"{len(qualifying)} narrower recovery attempt(s) recorded as " + "insufficient; escalation permitted" + ], + qualifying_attempts=qualifying, + recommended_next=[], + ) + + # Deny: recommend the next untried narrow rung(s). + recommended = recommend_actions( + symptoms=(), + prior_recovery_attempts=prior_recovery_attempts, + max_actions=3, + ) + reasons.append( + f"{action.value} requires a prior attempt log of insufficient " + "narrower recoveries (or break-glass); none found — deny (fail " + "closed, #669)" + ) + return EscalationAssessment( + requested_action=action.value, + allowed=False, + require_attempt_log=True, + break_glass=False, + reasons=reasons, + qualifying_attempts=[], + recommended_next=recommended.get("recommended_actions") or [], + ) + + +def recommend_actions( + *, + symptoms: Sequence[str] = (), + prior_recovery_attempts: Sequence[Mapping[str, Any]] | None = None, + max_actions: int = 5, +) -> dict[str, Any]: + """Return ordered recommended recovery actions for the given symptoms. + + Soft mode (rollout): recommendations only — callers decide whether to + hard-gate. Hard mode for broad restarts is :func:`assess_escalation`. + """ + attempted_success = { + a["action"] + for a in normalize_attempt_log(prior_recovery_attempts) + if a["outcome"] == OUTCOME_SUCCESS + } + attempted_any = { + a["action"] for a in normalize_attempt_log(prior_recovery_attempts) + } + + first_actions: list[RecoveryAction] = [] + for symptom in symptoms: + key = str(symptom or "").strip().lower().replace(" ", "_").replace("-", "_") + mapped = SYMPTOM_TO_FIRST_ACTION.get(key) + if mapped is not None and mapped not in first_actions: + first_actions.append(mapped) + + # Default entry: client reconnect then walk the ladder. + if not first_actions: + first_actions = [RecoveryAction.CLIENT_RECONNECT] + + recommended: list[dict[str, Any]] = [] + seen: set[str] = set() + min_rank = min(ladder_rank(a) for a in first_actions) + + for rung in RECOVERY_LADDER: + if rung.rank < min_rank: + continue + if rung.action.value in attempted_success: + continue + if rung.action.value in seen: + continue + # Prefer rungs not yet attempted; still list previously-failed ones + # only if nothing else remains. + entry = rung.as_dict() + entry["already_attempted"] = rung.action.value in attempted_any + recommended.append(entry) + seen.add(rung.action.value) + if len(recommended) >= max(1, int(max_actions)): + break + + return { + "playbook_version": PLAYBOOK_VERSION, + "symptoms": [str(s) for s in symptoms], + "recommended_actions": recommended, + "ladder": [r.as_dict() for r in RECOVERY_LADDER], + "read_only": True, + "hard_gate_note": ( + "Broad restarts (rolling/full/host) still require " + "assess_escalation / coordinator attempt-log enforcement." + ), + } + + +def build_attempt_record( + action: RecoveryAction | str, + *, + outcome: str, + reason: str = "", + actor: str = "", + recorded_at: str | None = None, + extra: Mapping[str, Any] | None = None, +) -> dict[str, Any]: + """Build a durable-shaped attempt log entry for inventory/audit (#665).""" + resolved = resolve_action(action) + record = { + "action": resolved.value, + "outcome": str(outcome or "").strip().lower(), + "reason": str(reason or "").strip(), + "actor": str(actor or "").strip(), + "recorded_at": recorded_at or _utc_now().isoformat(), + "rank": ladder_rank(resolved), + "playbook_version": PLAYBOOK_VERSION, + } + if extra: + record["extra"] = dict(extra) + return record + + +def recovery_metrics( + attempts: Sequence[Mapping[str, Any]] | None, +) -> dict[str, Any]: + """Compute the fraction of recoveries that avoided full/host restart. + + A recovery *episode* is approximated as one attempt with + ``outcome=success``. Successes on non-broad rungs count as avoided full + restart; successes on full/host count as full-restart recoveries. + """ + norms = normalize_attempt_log(attempts) + successes = [a for a in norms if a["outcome"] == OUTCOME_SUCCESS] + broad_success = [ + a + for a in successes + if resolve_action(a["action"]) + in {RecoveryAction.FULL_MCP_RESTART, RecoveryAction.HOST_RESTART} + ] + avoided = [a for a in successes if a not in broad_success] + total = len(successes) + fraction_avoided = (len(avoided) / total) if total else None + return { + "playbook_version": PLAYBOOK_VERSION, + "attempts_total": len(norms), + "successes_total": total, + "successes_avoided_full_restart": len(avoided), + "successes_full_or_host_restart": len(broad_success), + "fraction_avoided_full_restart": fraction_avoided, + "insufficient_attempts": sum( + 1 for a in norms if a["outcome"] in INSUFFICIENT_OUTCOMES + ), + } + + +def ladder_document() -> dict[str, Any]: + """Machine-readable ladder for docs/tools inventory.""" + return { + "playbook_version": PLAYBOOK_VERSION, + "parent_issues": ["#655", "#652", "#653"], + "enforcement_issue": "#669", + "ladder": [r.as_dict() for r in RECOVERY_LADDER], + "broad_restart_actions": [a.value for a in sorted(BROAD_RESTART_ACTIONS, key=lambda x: x.value)], + "insufficient_outcomes": sorted(INSUFFICIENT_OUTCOMES), + "symptom_map": {k: v.value for k, v in sorted(SYMPTOM_TO_FIRST_ACTION.items())}, + } diff --git a/restart_coordinator.py b/restart_coordinator.py index 5452ad0..8134b44 100644 --- a/restart_coordinator.py +++ b/restart_coordinator.py @@ -1,4 +1,4 @@ -"""MCP restart coordinator and impact analysis (#658). +"""MCP restart coordinator and impact analysis (#658 / #669). Before any sanctioned MCP restart, a central coordinator must evaluate the live control-plane state — active sessions, leases/locks, in-flight issue/PR @@ -16,6 +16,9 @@ Design rules (mirrors the read-only posture of ``workflow_dashboard`` / a mutative apply path is a later child gated by a drain proof (non-goal here). * **Fail closed.** If the inventory is not explicitly complete, the verdict is ``unsafe`` / deny — an incomplete evaluation must never green-light a restart. +* **Narrow-first (#669).** Broad classes (rolling / full / host) require a + prior attempt log of insufficient narrower recoveries unless break-glass is + authorized. See :mod:`recovery_playbook`. * **No secrets.** Session ids, pids, and profiles are operational metadata, not credentials; nothing secret flows through this module. @@ -32,8 +35,9 @@ from enum import Enum from typing import Any, Mapping, Sequence import lease_lifecycle +import recovery_playbook -COORDINATOR_VERSION = "1.1.0-issue-663" +COORDINATOR_VERSION = "1.2.0-issue-669" # Restart verdicts. Exactly the three the acceptance criteria name. VERDICT_SAFE = "safe" @@ -349,6 +353,10 @@ class RestartImpactReport: counts: dict[str, int] audit_record: dict[str, Any] incomplete_reasons: list[str] = field(default_factory=list) + # #669 playbook escalation gate (attempt-log enforcement). + playbook_escalation: dict[str, Any] = field(default_factory=dict) + attempt_log_satisfied: bool = True + break_glass: bool = False def as_dict(self) -> dict[str, Any]: return { @@ -382,6 +390,9 @@ class RestartImpactReport: "prior_recovery_attempts": list(self.prior_recovery_attempts), "counts": dict(self.counts), "audit_record": dict(self.audit_record), + "playbook_escalation": dict(self.playbook_escalation), + "attempt_log_satisfied": self.attempt_log_satisfied, + "break_glass": self.break_glass, } @@ -496,6 +507,7 @@ def evaluate_restart_impact( target_session_id: str | None = None, target_role: str | None = None, target_connector: str | None = None, + break_glass: bool = False, ) -> RestartImpactReport: """Evaluate a proposed MCP restart and return an impact preview. @@ -584,6 +596,30 @@ def evaluate_restart_impact( dict(a) for a in (inventory.get("prior_recovery_attempts") or []) ] + # #669: broad restarts require a prior narrow-attempt log unless break-glass. + playbook_escalation: dict[str, Any] = {} + attempt_log_satisfied = True + if policy_enforced and resolved_class is not None: + try: + escalation = recovery_playbook.assess_escalation( + resolved_class.value, + prior_recovery_attempts=prior_recovery_attempts, + break_glass=bool(break_glass), + ) + playbook_escalation = escalation.as_dict() + attempt_log_satisfied = bool(escalation.allowed) + if not attempt_log_satisfied: + authorization_reasons.extend(list(escalation.reasons)) + except ValueError as exc: + # Unknown mapping should never happen for enum values; fail closed. + attempt_log_satisfied = False + playbook_escalation = { + "allowed": False, + "reasons": [str(exc)], + "playbook_version": recovery_playbook.PLAYBOOK_VERSION, + } + authorization_reasons.append(str(exc)) + session_impacts = [ _classify_session( s, @@ -682,6 +718,7 @@ def evaluate_restart_impact( and role_authorized and approval_satisfied and target_complete + and attempt_log_satisfied ) if policy_enforced and not authorization_ok: @@ -745,6 +782,7 @@ def evaluate_restart_impact( "affected_issues": len(affected_issues), "affected_prs": len(affected_prs), "prior_recovery_attempts": len(prior_recovery_attempts), + "attempt_log_satisfied": attempt_log_satisfied, } audit_record = { @@ -765,6 +803,9 @@ def evaluate_restart_impact( "allow_restart": allow_restart, "blast_radius": blast_radius, "counts": counts, + "attempt_log_satisfied": attempt_log_satisfied, + "break_glass": bool(break_glass), + "playbook_version": recovery_playbook.PLAYBOOK_VERSION, } return RestartImpactReport( @@ -804,4 +845,7 @@ def evaluate_restart_impact( counts=counts, audit_record=audit_record, incomplete_reasons=incomplete_reasons, + playbook_escalation=playbook_escalation, + attempt_log_satisfied=attempt_log_satisfied, + break_glass=bool(break_glass), ) diff --git a/stable_branch_push_guard.py b/stable_branch_push_guard.py index aa15b83..8d0c667 100644 --- a/stable_branch_push_guard.py +++ b/stable_branch_push_guard.py @@ -63,6 +63,11 @@ CONTAMINATION_GATED_TASKS = frozenset({ "merge_pr", "delete_branch", "complete_issue", + # Web console recovery playbooks that write (#644). These mutate runtime + # binding and process state, so a live contamination marker must block them + # exactly as it blocks the Gitea-side mutations above. The reconciler + # cleanup playbook is the designated remedy and is exempted by its caller. + "console_recovery_apply", }) CONTAMINATION_KIND = "stable_branch_push" diff --git a/task_capability_map.py b/task_capability_map.py index c2576c6..a2c8868 100644 --- a/task_capability_map.py +++ b/task_capability_map.py @@ -142,6 +142,23 @@ TASK_CAPABILITY_MAP: dict[str, dict[str, str]] = { "permission": "gitea.read", "role": "author", }, + # #644: Phase 2 Web Console recovery tasks. + "clear_stale_binding": { + "permission": "gitea.read", + "role": "author", + }, + "rebind_session_worktree": { + "permission": "gitea.read", + "role": "author", + }, + # The console playbook orchestrates gitea_reconcile_merged_cleanups, whose + # own gate is gitea.read (matching the existing reconcile_merged_cleanups + # entry). Declaring a stricter permission here stated a second, conflicting + # authority for one operation. + "reconcile_cleanups": { + "permission": "gitea.read", + "role": "reconciler", + }, # PR synchronization lifecycle: assess is read-only (any role with gitea.read); # update-by-merge is author-only and mutates the PR head via Gitea API. "assess_pr_sync_status": { diff --git a/tests/test_issue_886_apply_authorization_conjunction.py b/tests/test_issue_886_apply_authorization_conjunction.py index 6fa2888..5227f9c 100644 --- a/tests/test_issue_886_apply_authorization_conjunction.py +++ b/tests/test_issue_886_apply_authorization_conjunction.py @@ -35,6 +35,17 @@ BREAK_GLASS_ENV = "GITEA_BREAKGLASS_RESTART_AUTHORIZATION" QUIET_SESSIONS: list[dict] = [] QUIET_LEASES: list[dict] = [] +# #669: broad restarts need a prior narrow-attempt log (unless break-glass). +PRIOR_NARROW_ATTEMPTS_JSON = json.dumps( + [ + { + "action": "client_reconnect", + "outcome": "insufficient", + "reason": "still flapping after reconnect", + } + ] +) + class _FakeDB: """Minimal control-plane DB stand-in for the restart inventory.""" @@ -128,6 +139,7 @@ class TestConjunction(_RestartToolHarness): preview = self._call( role="operator", restart_class="full_mcp_restart", + prior_recovery_attempts_json=PRIOR_NARROW_ATTEMPTS_JSON, env={CONTROLLER_APPROVAL_ENV: "operator-approved"}, ) self.assertTrue(preview["allow_restart"], @@ -136,6 +148,7 @@ class TestConjunction(_RestartToolHarness): result = self._call( role="operator", restart_class="full_mcp_restart", + prior_recovery_attempts_json=PRIOR_NARROW_ATTEMPTS_JSON, dry_run=False, drain_proof_json=self._clean_proof_for(preview), env={CONTROLLER_APPROVAL_ENV: "operator-approved"}, @@ -342,6 +355,7 @@ class TestExistingPathsStillWork(_RestartToolHarness): result = self._call( role="operator", restart_class="full_mcp_restart", + prior_recovery_attempts_json=PRIOR_NARROW_ATTEMPTS_JSON, dry_run=False, env={CONTROLLER_APPROVAL_ENV: "operator-approved"}, ) diff --git a/tests/test_recovery_playbook.py b/tests/test_recovery_playbook.py new file mode 100644 index 0000000..57e5005 --- /dev/null +++ b/tests/test_recovery_playbook.py @@ -0,0 +1,217 @@ +"""Unit tests for the scoped recovery playbook (#669).""" + +from __future__ import annotations + +import recovery_playbook as rp +import restart_coordinator as rc + + +def test_ladder_covers_eleven_ordered_rungs(): + ranks = [r.rank for r in rp.RECOVERY_LADDER] + assert ranks == list(range(len(rp.RECOVERY_LADDER))) + assert len(rp.RECOVERY_LADDER) == 11 + assert rp.RECOVERY_LADDER[0].action is rp.RecoveryAction.CLIENT_RECONNECT + assert rp.RECOVERY_LADDER[-1].action is rp.RecoveryAction.HOST_RESTART + + +def test_ladder_document_links_parent_issues(): + doc = rp.ladder_document() + assert "#655" in doc["parent_issues"] + assert "#652" in doc["parent_issues"] + assert "#653" in doc["parent_issues"] + assert doc["enforcement_issue"] == "#669" + assert "full_mcp_restart" in doc["broad_restart_actions"] + + +def test_recommend_transport_eof_starts_at_client_reconnect(): + plan = rp.recommend_actions(symptoms=["transport_eof"]) + assert plan["recommended_actions"][0]["action"] == "client_reconnect" + assert plan["recommended_actions"][0]["issue_links"] + + +def test_recommend_skips_successful_prior_attempts(): + attempts = [ + rp.build_attempt_record( + "client_reconnect", outcome="success", reason="reconnected" + ) + ] + plan = rp.recommend_actions( + symptoms=["transport_eof"], prior_recovery_attempts=attempts + ) + actions = [a["action"] for a in plan["recommended_actions"]] + assert "client_reconnect" not in actions + assert actions[0] == "capability_refresh" + + +def test_escalation_denied_without_attempt_log(): + result = rp.assess_escalation("full_mcp_restart", prior_recovery_attempts=[]) + assert result.allowed is False + assert result.require_attempt_log is True + assert any("#669" in r for r in result.reasons) + assert result.recommended_next # soft recommendations still provided + + +def test_escalation_allowed_after_insufficient_narrower(): + attempts = [ + rp.build_attempt_record( + "client_reconnect", + outcome="insufficient", + reason="still flapping", + ), + rp.build_attempt_record( + "session_reconnect", + outcome="failed", + reason="namespace still dead", + ), + ] + result = rp.assess_escalation( + "full_mcp_restart", prior_recovery_attempts=attempts + ) + assert result.allowed is True + assert len(result.qualifying_attempts) == 2 + + +def test_escalation_break_glass_bypasses_attempt_log(): + result = rp.assess_escalation( + "host_restart", prior_recovery_attempts=[], break_glass=True + ) + assert result.allowed is True + assert result.break_glass is True + + +def test_narrow_action_does_not_require_attempt_log(): + result = rp.assess_escalation( + "client_reconnect", prior_recovery_attempts=[] + ) + assert result.allowed is True + assert result.require_attempt_log is False + + +def test_same_rank_attempt_does_not_qualify_for_escalation(): + attempts = [ + rp.build_attempt_record( + "full_mcp_restart", outcome="failed", reason="already failed full" + ) + ] + result = rp.assess_escalation( + "full_mcp_restart", prior_recovery_attempts=attempts + ) + assert result.allowed is False + + +def test_success_outcome_does_not_qualify_for_escalation(): + attempts = [ + rp.build_attempt_record( + "client_reconnect", outcome="success", reason="fixed" + ) + ] + result = rp.assess_escalation( + "full_mcp_restart", prior_recovery_attempts=attempts + ) + assert result.allowed is False + + +def test_recovery_metrics_fraction_avoided(): + attempts = [ + rp.build_attempt_record("client_reconnect", outcome="success"), + rp.build_attempt_record("session_reconnect", outcome="success"), + rp.build_attempt_record("full_mcp_restart", outcome="success"), + ] + metrics = rp.recovery_metrics(attempts) + assert metrics["successes_total"] == 3 + assert metrics["successes_avoided_full_restart"] == 2 + assert metrics["successes_full_or_host_restart"] == 1 + assert abs(metrics["fraction_avoided_full_restart"] - (2 / 3)) < 1e-9 + + +def test_coordinator_denies_full_restart_without_attempt_log(): + inv = { + "inventory_complete": True, + "sessions": [], + "leases": [], + "prior_recovery_attempts": [], + } + report = rc.evaluate_restart_impact( + inv, + restart_class=rc.RestartClass.FULL_MCP_RESTART, + requester_role="controller", + requester_permissions=rc.permissions_for_role("controller"), + controller_approved=True, + operator_authorized=True, + ) + assert report.allow_restart is False + assert report.attempt_log_satisfied is False + assert report.verdict == rc.VERDICT_UNSAFE + blob = " ".join(report.reasons + report.authorization_reasons) + assert "#669" in blob or "attempt log" in blob + + +def test_coordinator_allows_full_restart_with_attempt_log(): + inv = { + "inventory_complete": True, + "sessions": [], + "leases": [], + "prior_recovery_attempts": [ + { + "action": "client_reconnect", + "outcome": "insufficient", + "reason": "still broken", + } + ], + } + report = rc.evaluate_restart_impact( + inv, + restart_class=rc.RestartClass.FULL_MCP_RESTART, + requester_role="controller", + requester_permissions=rc.permissions_for_role("controller"), + controller_approved=True, + operator_authorized=True, + ) + assert report.attempt_log_satisfied is True + assert report.allow_restart is True + assert report.verdict == rc.VERDICT_SAFE + + +def test_coordinator_break_glass_allows_without_log(): + inv = { + "inventory_complete": True, + "sessions": [], + "leases": [], + "prior_recovery_attempts": [], + } + report = rc.evaluate_restart_impact( + inv, + restart_class=rc.RestartClass.FULL_MCP_RESTART, + requester_role="controller", + requester_permissions=rc.permissions_for_role("controller"), + controller_approved=True, + operator_authorized=True, + break_glass=True, + ) + assert report.break_glass is True + assert report.attempt_log_satisfied is True + assert report.allow_restart is True + + +def test_coordinator_client_reconnect_unaffected(): + inv = { + "inventory_complete": True, + "sessions": [], + "leases": [], + "prior_recovery_attempts": [], + } + report = rc.evaluate_restart_impact( + inv, + restart_class=rc.RestartClass.CLIENT_RECONNECT, + requester_role="author", + requester_permissions=rc.permissions_for_role("author"), + ) + assert report.attempt_log_satisfied is True + assert report.allow_restart is True + + +def test_restart_class_alias_accepted(): + assert ( + rp.resolve_action("full_mcp_restart") + is rp.RecoveryAction.FULL_MCP_RESTART + ) diff --git a/tests/test_webui_console_recovery.py b/tests/test_webui_console_recovery.py new file mode 100644 index 0000000..b22627b --- /dev/null +++ b/tests/test_webui_console_recovery.py @@ -0,0 +1,506 @@ +"""Unit and integration tests for Phase 2 Web Console recovery controls (#644).""" + +from __future__ import annotations + +import os +import sys +import types +import unittest +from unittest.mock import patch + +from starlette.testclient import TestClient + +import merged_cleanup_reconcile +import runtime_recovery_guard +import stable_branch_push_guard +import stale_binding_recovery +from webui import console_authz, console_recovery, system_health +from webui.app import create_app + + +class TestConsoleRecovery(unittest.TestCase): + + def test_diagnose_recovery_healthy(self) -> None: + diag = console_recovery.diagnose_recovery() + self.assertIn(diag.status, {console_recovery.STATUS_HEALTHY, console_recovery.STATUS_ACTION_REQUIRED}) + self.assertIsInstance(diag.playbooks, tuple) + self.assertGreaterEqual(len(diag.playbooks), 4) + + playbook_ids = {pb.playbook_id for pb in diag.playbooks} + self.assertIn(console_recovery.PLAYBOOK_CLEAR_STALE_BINDING, playbook_ids) + self.assertIn(console_recovery.PLAYBOOK_REBIND_SESSION, playbook_ids) + self.assertIn(console_recovery.PLAYBOOK_RECONCILE_CLEANUPS, playbook_ids) + self.assertIn(console_recovery.PLAYBOOK_SANCTIONED_RESTART, playbook_ids) + + def test_confirmation_phrase_generation_and_matching(self) -> None: + phrase = console_recovery.confirmation_phrase("clear_stale_binding") + self.assertEqual(phrase, "confirm clear_stale_binding") + self.assertTrue(console_recovery.confirmation_matches("clear_stale_binding", "confirm clear_stale_binding")) + self.assertFalse(console_recovery.confirmation_matches("clear_stale_binding", "wrong phrase")) + + phrase_target = console_recovery.confirmation_phrase("sanctioned_restart", "gitea-author") + self.assertEqual(phrase_target, "confirm sanctioned_restart gitea-author") + self.assertTrue(console_recovery.confirmation_matches("sanctioned_restart", "confirm sanctioned_restart gitea-author", "gitea-author")) + + def test_build_recovery_preview(self) -> None: + principal = console_authz.Principal("dev@example.com", console_authz.OPERATOR, console_authz.IDENTITY_LOCAL_DEV, True) + preview = console_recovery.build_recovery_preview( + playbook_id=console_recovery.PLAYBOOK_CLEAR_STALE_BINDING, + target="test-worktree", + principal=principal, + ) + self.assertEqual(preview["playbook_id"], console_recovery.PLAYBOOK_CLEAR_STALE_BINDING) + self.assertEqual(preview["action_id"], console_recovery.ACTION_CLEAR_STALE_BINDING) + self.assertEqual(preview["confirmation_phrase"], "confirm clear_stale_binding test-worktree") + self.assertTrue(len(preview["mutation_ledger"]) >= 3) + self.assertTrue(preview["authorization"]["allowed"]) + + def test_build_recovery_preview_unknown_playbook(self) -> None: + preview = console_recovery.build_recovery_preview("unknown_playbook") + self.assertFalse(preview.get("allowed")) + self.assertEqual(preview.get("error"), "unknown_playbook") + + def test_execute_recovery_playbook_confirmation_mismatch(self) -> None: + # Authorization is checked before confirmation, so the phase gate has to + # pass for this test to reach the branch it is about. + principal = console_authz.Principal("dev@example.com", console_authz.OPERATOR, console_authz.IDENTITY_LOCAL_DEV, True) + with self._phase_two_enabled(): + result = console_recovery.execute_recovery_playbook( + playbook_id=console_recovery.PLAYBOOK_CLEAR_STALE_BINDING, + confirmation="invalid confirmation", + principal=principal, + ) + self.assertFalse(result["success"]) + self.assertFalse(result["allowed"]) + self.assertEqual(result["error"], "confirmation_mismatch") + + def test_execute_recovery_playbook_unauthorized(self) -> None: + # Anonymous principal has viewer role -> should be denied + result = console_recovery.execute_recovery_playbook( + playbook_id=console_recovery.PLAYBOOK_CLEAR_STALE_BINDING, + confirmation="confirm clear_stale_binding", + principal=console_authz.ANONYMOUS, + ) + self.assertFalse(result["success"]) + self.assertFalse(result["allowed"]) + self.assertEqual(result["error"], console_authz.DENY_UNAUTHENTICATED) + + def test_execute_refuses_phase_two_write_while_console_is_phase_one(self) -> None: + """B1: the apply path must arm the phase gate, not skip it. + + ``authorize`` only applies the phase branch when ``for_execution=True``. + The apply path used the default, so an operator executed a phase-2 write + while ``ACTIVE_PHASE`` was 1. + """ + self.assertGreater( + console_authz.get_action(console_recovery.ACTION_CLEAR_STALE_BINDING).phase, + console_authz.ACTIVE_PHASE, + "fixture assumes the recovery actions are ahead of the active phase", + ) + principal = console_authz.Principal( + "dev@example.com", console_authz.OPERATOR, console_authz.IDENTITY_LOCAL_DEV, True + ) + phrase = console_recovery.confirmation_phrase( + console_recovery.PLAYBOOK_CLEAR_STALE_BINDING + ) + result = console_recovery.execute_recovery_playbook( + playbook_id=console_recovery.PLAYBOOK_CLEAR_STALE_BINDING, + confirmation=phrase, + principal=principal, + ) + self.assertFalse(result["success"]) + self.assertFalse(result["allowed"]) + self.assertEqual(result["error"], console_authz.DENY_PHASE_NOT_ACTIVE) + + def test_preview_execution_enabled_matches_the_execution_decision(self) -> None: + """B1: preview must not report a bare False it cannot explain.""" + principal = console_authz.Principal( + "dev@example.com", console_authz.OPERATOR, console_authz.IDENTITY_LOCAL_DEV, True + ) + preview = console_recovery.build_recovery_preview( + playbook_id=console_recovery.PLAYBOOK_REBIND_SESSION, + target="branches/feat-issue-644", + principal=principal, + ) + self.assertFalse(preview["execution_enabled"]) + self.assertEqual( + preview["execution_blocked_reason"], console_authz.DENY_PHASE_NOT_ACTIVE + ) + self.assertFalse(preview["execution_authorization"]["allowed"]) + # The preview (non-execution) decision still allows, by role. + self.assertTrue(preview["authorization"]["allowed"]) + + def _phase_two_enabled(self): + """Raise ACTIVE_PHASE so the execution branches are reachable in tests.""" + return patch.object(console_authz, "ACTIVE_PHASE", 2) + + def _operator(self) -> console_authz.Principal: + return console_authz.Principal( + "dev@example.com", console_authz.OPERATOR, console_authz.IDENTITY_LOCAL_DEV, True + ) + + def test_rebind_mutates_the_live_environment_not_a_copy(self) -> None: + """B2: the playbook must change the mapping it claims to have changed.""" + live_env = {stale_binding_recovery.ACTIVE_WORKTREE_ENV: "branches/stale-old"} + phrase = console_recovery.confirmation_phrase( + console_recovery.PLAYBOOK_REBIND_SESSION, "branches/feat-issue-644" + ) + with self._phase_two_enabled(): + result = console_recovery.execute_recovery_playbook( + playbook_id=console_recovery.PLAYBOOK_REBIND_SESSION, + confirmation=phrase, + target="branches/feat-issue-644", + principal=self._operator(), + env=live_env, + ) + self.assertTrue(result["success"]) + self.assertEqual( + live_env[stale_binding_recovery.ACTIVE_WORKTREE_ENV], + "branches/feat-issue-644", + "rebind reported success without changing the caller's environment", + ) + self.assertTrue(result["applied_result"]["binding_changed"]) + self.assertEqual(result["applied_result"]["binding_before"], "branches/stale-old") + self.assertEqual( + result["applied_result"]["binding_after"], "branches/feat-issue-644" + ) + + def test_clear_stale_binding_reports_failure_when_nothing_changed(self) -> None: + """B2: a no-op recovery must never be reported as success.""" + live_env: dict[str, str] = {} + phrase = console_recovery.confirmation_phrase( + console_recovery.PLAYBOOK_CLEAR_STALE_BINDING + ) + with self._phase_two_enabled(): + result = console_recovery.execute_recovery_playbook( + playbook_id=console_recovery.PLAYBOOK_CLEAR_STALE_BINDING, + confirmation=phrase, + principal=self._operator(), + env=live_env, + ) + self.assertFalse( + result["success"], + "a clear that changed no binding must not report success", + ) + self.assertFalse(result["applied_result"]["binding_changed"]) + + def test_clear_stale_binding_clears_the_live_binding(self) -> None: + """B2: the sanctioned clear must reach the caller's environment.""" + missing = "/nonexistent/branches/deleted-worktree" + live_env = {stale_binding_recovery.ACTIVE_WORKTREE_ENV: missing} + phrase = console_recovery.confirmation_phrase( + console_recovery.PLAYBOOK_CLEAR_STALE_BINDING + ) + with self._phase_two_enabled(): + result = console_recovery.execute_recovery_playbook( + playbook_id=console_recovery.PLAYBOOK_CLEAR_STALE_BINDING, + confirmation=phrase, + principal=self._operator(), + env=live_env, + ) + if result["success"]: + self.assertNotIn(stale_binding_recovery.ACTIVE_WORKTREE_ENV, live_env) + self.assertEqual(result["applied_result"]["binding_before"], missing) + self.assertIsNone(result["applied_result"]["binding_after"]) + else: + # Fail closed is acceptable; reporting a clear that did not happen + # is not. This is the invariant the blocker was about. + self.assertFalse(result["applied_result"]["binding_changed"]) + self.assertEqual( + live_env.get(stale_binding_recovery.ACTIVE_WORKTREE_ENV), missing + ) + + def test_reconcile_playbook_calls_an_entry_point_that_exists(self) -> None: + """B3: the previous call named a function absent from the module.""" + phrase = console_recovery.confirmation_phrase( + console_recovery.PLAYBOOK_RECONCILE_CLEANUPS + ) + fake_server = types.SimpleNamespace( + gitea_reconcile_merged_cleanups=lambda **kwargs: { + "success": True, + "entries": [{"issue_number": 100}], + } + ) + with self._phase_two_enabled(), patch.dict( + sys.modules, {"gitea_mcp_server": fake_server} + ): + result = console_recovery.execute_recovery_playbook( + playbook_id=console_recovery.PLAYBOOK_RECONCILE_CLEANUPS, + confirmation=phrase, + principal=console_authz.Principal( + "dev@example.com", + console_authz.ADMIN, + console_authz.IDENTITY_LOCAL_DEV, + True, + ), + ) + self.assertTrue(result["success"], result.get("applied_result")) + self.assertNotIn("error_type", result["applied_result"]) + self.assertEqual(result["applied_result"]["reconciled_count"], 1) + + def test_reconcile_entry_point_exists_on_the_real_module(self) -> None: + """B3 regression: guard the symbol itself, not just the call shape.""" + import gitea_mcp_server + + self.assertTrue( + hasattr(gitea_mcp_server, "gitea_reconcile_merged_cleanups"), + "console recovery depends on this reconciler entry point", + ) + self.assertFalse( + hasattr(merged_cleanup_reconcile, "reconcile_merged_cleanups"), + "if this module grows the orchestrator, point the playbook back at it", + ) + + def test_contamination_gate_blocks_a_writing_playbook(self) -> None: + """B4: a live marker plus a gated task key must actually block.""" + marker = { + "kind": "manual_daemon_kill", + "reason_class": "manual_daemon_kill", + "command_summary": "pkill -f gitea_mcp_server", + "active": True, + } + phrase = console_recovery.confirmation_phrase( + console_recovery.PLAYBOOK_REBIND_SESSION, "branches/feat-issue-644" + ) + live_env = {stale_binding_recovery.ACTIVE_WORKTREE_ENV: "branches/stale-old"} + with self._phase_two_enabled(), patch.object( + console_recovery, "load_active_contamination_marker", return_value=marker + ): + result = console_recovery.execute_recovery_playbook( + playbook_id=console_recovery.PLAYBOOK_REBIND_SESSION, + confirmation=phrase, + target="branches/feat-issue-644", + principal=self._operator(), + env=live_env, + ) + self.assertFalse(result["success"]) + self.assertEqual(result["error"], "contaminated_runtime") + self.assertEqual( + live_env[stale_binding_recovery.ACTIVE_WORKTREE_ENV], + "branches/stale-old", + "a blocked playbook must not have mutated anything", + ) + + def test_contamination_gate_exempts_the_reconciler_remedy(self) -> None: + """B4: the designated remedy must stay reachable while contaminated.""" + marker = { + "kind": "manual_daemon_kill", + "reason_class": "manual_daemon_kill", + "command_summary": "pkill -f gitea_mcp_server", + "active": True, + } + phrase = console_recovery.confirmation_phrase( + console_recovery.PLAYBOOK_RECONCILE_CLEANUPS + ) + fake_server = types.SimpleNamespace( + gitea_reconcile_merged_cleanups=lambda **kwargs: { + "success": True, + "entries": [], + } + ) + with self._phase_two_enabled(), patch.object( + console_recovery, "load_active_contamination_marker", return_value=marker + ), patch.dict(sys.modules, {"gitea_mcp_server": fake_server}): + result = console_recovery.execute_recovery_playbook( + playbook_id=console_recovery.PLAYBOOK_RECONCILE_CLEANUPS, + confirmation=phrase, + principal=console_authz.Principal( + "dev@example.com", + console_authz.ADMIN, + console_authz.IDENTITY_LOCAL_DEV, + True, + ), + ) + self.assertNotEqual(result.get("error"), "contaminated_runtime") + + def test_gated_task_key_is_actually_gated(self) -> None: + """B4: the console action id was never a member of the gated set.""" + self.assertIn( + console_recovery.CONTAMINATION_GATED_TASK, + stable_branch_push_guard.CONTAMINATION_GATED_TASKS, + ) + self.assertNotIn( + console_recovery.ACTION_CLEAR_STALE_BINDING, + stable_branch_push_guard.CONTAMINATION_GATED_TASKS, + ) + + def test_diagnosis_reads_the_key_the_gate_returns(self) -> None: + """B4: ``contaminated`` is a key assess_contamination_gate never returns.""" + gate = runtime_recovery_guard.assess_contamination_gate( + None, task=console_recovery.CONTAMINATION_GATED_TASK, actual_role="operator" + ) + self.assertNotIn("contaminated", gate) + self.assertIn("block", gate) + + def test_contaminated_runtime_is_reported_unclean(self) -> None: + """B4: verify_post_recovery reported contamination_clean unconditionally.""" + marker = { + "kind": "manual_daemon_kill", + "reason_class": "manual_daemon_kill", + "command_summary": "pkill -f gitea_mcp_server", + "active": True, + } + with patch.object( + console_recovery, "load_active_contamination_marker", return_value=marker + ): + verification = console_recovery.verify_post_recovery() + diag = console_recovery.diagnose_recovery() + self.assertFalse(verification["contamination_clean"]) + self.assertFalse(verification["clean"]) + self.assertEqual(diag.status, console_recovery.STATUS_BLOCKED_CONTAMINATION) + + def test_master_parity_baseline_is_not_the_head_it_is_compared_against(self) -> None: + """B5: capture_startup_parity was fed the head it was then compared to.""" + stale = system_health.StaleRuntime( + daemon_head="a" * 40, + checkout_head="b" * 40, + remote_head="b" * 40, + stale=True, + determinable=True, + mutation_safe=False, + reasons=("daemon is behind the checkout",), + ) + with patch.object(system_health, "assess_stale_runtime", return_value=stale): + diag = console_recovery.diagnose_recovery() + parity = diag.master_parity + self.assertEqual(parity["startup_head"], "a" * 40) + self.assertEqual(parity["current_head"], "b" * 40) + self.assertNotEqual(parity["startup_head"], parity["current_head"]) + self.assertFalse(parity["in_parity"]) + + def test_master_parity_carries_the_live_remote_dimension(self) -> None: + """B5: live_remote_head was never passed, dropping the #610 dimension.""" + stale = system_health.StaleRuntime( + daemon_head="c" * 40, + checkout_head="c" * 40, + remote_head="d" * 40, + stale=False, + determinable=True, + mutation_safe=False, + reasons=(), + ) + with patch.object(system_health, "assess_stale_runtime", return_value=stale): + diag = console_recovery.diagnose_recovery() + self.assertEqual(diag.master_parity.get("live_remote_head"), "d" * 40) + + def test_verify_post_recovery(self) -> None: + verification = console_recovery.verify_post_recovery() + self.assertIn("clean", verification) + self.assertIn("status", verification) + self.assertIn("reasons", verification) + + def test_unverified_inherited_binding_is_not_reported_clean(self) -> None: + """B2: ``not clear_eligible`` also read clean for unproven bindings.""" + binding = { + "classification": stale_binding_recovery.CLASSIFICATION_UNVERIFIED_INHERITED, + "clear_eligible": False, + } + diag = console_recovery.diagnose_recovery() + patched = console_recovery.RecoveryDiagnosis( + status=diag.status, + clean=diag.clean, + stale_runtime=diag.stale_runtime, + master_parity=diag.master_parity, + stale_binding=binding, + contamination=diag.contamination, + worktree_anomalies=diag.worktree_anomalies, + playbooks=diag.playbooks, + reasons=diag.reasons, + ) + with patch.object(console_recovery, "diagnose_recovery", return_value=patched): + verification = console_recovery.verify_post_recovery() + self.assertFalse(verification["binding_clean"]) + self.assertEqual( + verification["binding_classification"], + stale_binding_recovery.CLASSIFICATION_UNVERIFIED_INHERITED, + ) + + +class TestConsoleRecoveryApi(unittest.TestCase): + def setUp(self) -> None: + self.app = create_app() + self.client = TestClient(self.app) + + def test_api_recovery_diagnose(self) -> None: + res = self.client.get("/api/v1/system/recovery/diagnose") + self.assertEqual(res.status_code, 200) + data = res.json() + self.assertIn("status", data) + self.assertIn("clean", data) + self.assertIn("playbooks", data) + self.assertTrue(len(data["playbooks"]) >= 4) + + def test_api_recovery_preview(self) -> None: + res = self.client.post( + "/api/v1/system/recovery/preview", + json={"playbook_id": "clear_stale_binding", "target": "active"}, + ) + self.assertEqual(res.status_code, 200) + data = res.json() + self.assertEqual(data["playbook_id"], "clear_stale_binding") + self.assertEqual(data["confirmation_phrase"], "confirm clear_stale_binding active") + self.assertIn("mutation_ledger", data) + + def test_api_recovery_apply_denied_without_auth(self) -> None: + res = self.client.post( + "/api/v1/system/recovery/apply", + json={"playbook_id": "clear_stale_binding", "confirmation": "confirm clear_stale_binding"}, + ) + self.assertEqual(res.status_code, 400) + data = res.json() + self.assertFalse(data["success"]) + self.assertFalse(data["allowed"]) + + def test_api_recovery_apply_refuses_phase_two_write_with_dev_auth(self) -> None: + """B1: this previously asserted the phase-gate bypass as intended. + + An authenticated operator posting a valid confirmation still must not + execute a phase-2 write while the console is in phase 1. The refusal is + the contract; a 200 here means the gate is not armed. + """ + env = { + "WEBUI_AUTH_MODE": "local_dev", + "WEBUI_DEV_SUBJECT": "dev@example.com", + "WEBUI_DEV_ROLE": "operator", + } + before = os.environ.get("GITEA_ACTIVE_WORKTREE") + with patch.dict(os.environ, env): + res = self.client.post( + "/api/v1/system/recovery/apply", + json={ + "playbook_id": "rebind_session_worktree", + "target": "branches/feat-issue-644", + "confirmation": "confirm rebind_session_worktree branches/feat-issue-644", + }, + ) + self.assertEqual(res.status_code, 400) + data = res.json() + self.assertFalse(data["success"]) + self.assertFalse(data["allowed"]) + self.assertEqual(data["error"], console_authz.DENY_PHASE_NOT_ACTIVE) + self.assertEqual( + os.environ.get("GITEA_ACTIVE_WORKTREE"), + before, + "a refused apply must not have rebound the live process environment", + ) + + def test_api_recovery_preview_reports_why_execution_is_disabled(self) -> None: + res = self.client.post( + "/api/v1/system/recovery/preview", + json={"playbook_id": "rebind_session_worktree", "target": "active"}, + ) + self.assertEqual(res.status_code, 200) + data = res.json() + self.assertFalse(data["execution_enabled"]) + self.assertIn("execution_authorization", data) + + def test_api_recovery_verify(self) -> None: + res = self.client.get("/api/v1/system/recovery/verify") + self.assertEqual(res.status_code, 200) + data = res.json() + self.assertIn("clean", data) + self.assertIn("status", data) + + +if __name__ == "__main__": + unittest.main() diff --git a/tests/test_webui_notifications.py b/tests/test_webui_notifications.py new file mode 100644 index 0000000..c928f7a --- /dev/null +++ b/tests/test_webui_notifications.py @@ -0,0 +1,465 @@ +"""Unit tests for Phase 3 Notifications and Human-Attention Console (#648).""" + +from __future__ import annotations + +import pytest +from starlette.testclient import TestClient + +from webui.app import create_app +from webui.notifications import ( + ATTENTION_HUMAN_REQUIRED, + ATTENTION_OPERATOR, + ATTENTION_ROUTINE, + CATEGORY_AUTH, + CATEGORY_BLOCKER, + CATEGORY_LEASE, + CATEGORY_SYSTEM, + CATEGORY_VALIDATION, + CATEGORY_WORKFLOW, + NotificationItem, + NotificationSnapshot, + classify_attention_event, + load_notifications_snapshot, + snapshot_to_dict, +) +from webui.notification_views import render_notifications_page +from webui.project_registry import load_registry +from webui.queue_loader import QueueItem, QueueSnapshot +from webui.lease_loader import CollisionWarning, LeaseSnapshot +from webui.system_health import DependencyProbe, SystemHealthSnapshot, VersionInfo, StaleRuntime + + +def test_classify_attention_event_rules(): + # 1. Critical escalation boundaries -> human-required + att_cls, req_human = classify_attention_event( + CATEGORY_AUTH, "Auth error", "Unauthorized access attempt", is_auth_failure=True + ) + assert att_cls == ATTENTION_HUMAN_REQUIRED + assert req_human is True + + att_cls, req_human = classify_attention_event( + CATEGORY_SYSTEM, "Hard stop", "Hard stop triggered", is_hard_stop=True + ) + assert att_cls == ATTENTION_HUMAN_REQUIRED + assert req_human is True + + att_cls, req_human = classify_attention_event( + CATEGORY_VALIDATION, "Validation Error", "Report validation failed", is_validation_failure=True + ) + assert att_cls == ATTENTION_HUMAN_REQUIRED + assert req_human is True + + # 2. Operational issues -> operator + att_cls, req_human = classify_attention_event( + CATEGORY_BLOCKER, "PR Blocked", "Merge conflict detected", is_blocker=True + ) + assert att_cls == ATTENTION_OPERATOR + assert req_human is False + + att_cls, req_human = classify_attention_event( + CATEGORY_LEASE, "Lease Expired", "Session lease expired", is_stale=True + ) + assert att_cls == ATTENTION_OPERATOR + assert req_human is False + + # 3. Routine workflow transitions -> routine + att_cls, req_human = classify_attention_event( + CATEGORY_WORKFLOW, "PR Active", "PR in review" + ) + assert att_cls == ATTENTION_ROUTINE + assert req_human is False + + +def test_notification_snapshot_aggregation(): + reg = load_registry() + proj_id = reg.projects[0].id if reg.projects else "gitea-tools" + + mock_queue = QueueSnapshot( + project_id=proj_id, + repo_label="org/repo", + prs=( + QueueItem( + number=101, + title="Blocked PR", + badges=("blocked",), + extra={}, + ), + QueueItem( + number=102, + title="Normal PR", + badges=("in-review",), + extra={}, + ), + ), + issues=(), + pr_pagination=None, + issue_pagination=None, + ) + + mock_leases = LeaseSnapshot( + project_id=proj_id, + repo_label="org/repo", + issue_lock=None, + claim_inventory={}, + reviewer_leases=( + { + "pr_number": 101, + "status": "expired", + "is_expired": True, + }, + ), + duplicate_prs=( + CollisionWarning( + kind="duplicate_pr", + message="Multiple open PRs for issue #101", + issue_number=101, + pr_numbers=(101, 103), + ), + ), + duplicate_branches=(), + collision_history=(), + fetch_error=None, + ) + + mock_version = VersionInfo( + git_sha="abc1234", + git_describe="v1.0.0", + control_plane_schema_version=1, + python_version="3.11", + known=True, + ) + + mock_stale = StaleRuntime( + daemon_head="abc1234", + checkout_head="abc1234", + remote_head="abc1234", + stale=False, + determinable=True, + mutation_safe=True, + reasons=(), + ) + + mock_health = SystemHealthSnapshot( + status="degraded", + ready=False, + readiness_complete=True, + readiness_reasons=("Auth failure",), + service="webui", + mode="test", + version=mock_version, + started_at="2026-07-25T00:00:00Z", + uptime_seconds=100.0, + timestamp="2026-07-25T00:00:00Z", + deep_probes_requested=True, + dependencies=( + DependencyProbe( + name="auth_service", + kind="auth", + status="unauthorized", + detail="Token expired", + required=True, + ), + ), + mcp_namespaces=(), + stale_runtime=mock_stale, + probe_errors=(), + ) + + snapshot = load_notifications_snapshot( + proj_id, + load_queue=lambda _id: mock_queue, + load_leases=lambda **_kwargs: mock_leases, + load_health=lambda **_kwargs: mock_health, + ) + + assert snapshot.project_id == proj_id + assert snapshot.total_count == 5 + assert snapshot.human_required_count >= 1 # auth probe failure + assert snapshot.operator_count >= 3 # blocked PR + expired lease + duplicate PR collision + assert snapshot.routine_count >= 1 # normal PR + + # Inbox items should include operator and human-required items only + inbox_classes = {item.attention_class for item in snapshot.inbox_items} + assert ATTENTION_ROUTINE not in inbox_classes + assert ATTENTION_OPERATOR in inbox_classes + assert ATTENTION_HUMAN_REQUIRED in inbox_classes + + +def test_snapshot_to_dict_and_redaction(): + item = NotificationItem( + id="notif-1", + attention_class=ATTENTION_HUMAN_REQUIRED, + category=CATEGORY_AUTH, + title="Auth Error", + summary="Failed auth header: Bearer secret_token_12345", + work_kind="system", + work_number=None, + project_id="test-proj", + repo_label="org/repo", + created_at="2026-07-25T16:00:00Z", + requires_human=True, + ) + snap = NotificationSnapshot( + project_id="test-proj", + repo_label="org/repo", + items=(item,), + human_required_count=1, + operator_count=0, + routine_count=0, + total_count=1, + ) + + data = snapshot_to_dict(snap) + assert data["project_id"] == "test-proj" + assert data["human_required_count"] == 1 + assert len(data["inbox_items"]) == 1 + + # Redaction test + summary = data["inbox_items"][0]["summary"] + assert "secret_token_12345" not in summary + assert "" in summary or "Bearer" in summary + + +def test_notifications_html_views(): + item = NotificationItem( + id="notif-1", + attention_class=ATTENTION_HUMAN_REQUIRED, + category=CATEGORY_AUTH, + title="Critical Auth Failure", + summary="Auth failure details", + work_kind="issue", + work_number=42, + project_id="test-proj", + repo_label="org/repo", + created_at="2026-07-25T16:00:00Z", + requires_human=True, + ) + snap = NotificationSnapshot( + project_id="test-proj", + repo_label="org/repo", + items=(item,), + human_required_count=1, + operator_count=0, + routine_count=0, + total_count=1, + ) + + html = render_notifications_page(snap, filter_class="inbox") + assert "Notifications & Attention Inbox" in html or "Notifications & Attention Inbox" in html + assert "Critical Auth Failure" in html + assert "HUMAN REQUIRED" in html + assert "Human Required" in html + + +def test_notifications_app_routes(): + app = create_app() + client = TestClient(app) + + # 1. HTML Route + res = client.get("/notifications") + assert res.status_code == 200 + assert "Notifications" in res.text + assert "Attention Inbox" in res.text + + # 2. API Route /api/v1/notifications + res_api = client.get("/api/v1/notifications") + assert res_api.status_code == 200 + json_data = res_api.json() + assert "human_required_count" in json_data + assert "operator_count" in json_data + assert "routine_count" in json_data + assert "inbox_items" in json_data + + # 3. Compatibility Alias /api/notifications + res_alias = client.get("/api/notifications") + assert res_alias.status_code == 200 + assert res_alias.json()["project_id"] == json_data["project_id"] + + +def test_classify_ignores_human_authored_title_and_summary_keywords(): + """B1: keywords in human-authored titles must not escalate routine work (#905).""" + # Routine transition whose title/summary mention critical-boundary words + att_cls, req_human = classify_attention_event( + CATEGORY_WORKFLOW, + "record irrecoverable decision lock provenance", + "PR #999 'record irrecoverable decision lock provenance' is in routine state in-review.", + ) + assert att_cls == ATTENTION_ROUTINE + assert req_human is False + + att_cls, req_human = classify_attention_event( + CATEGORY_WORKFLOW, + "fix unauthorized token path", + "Issue #1 'fix unauthorized token path' state: claimed. hard stop docs only.", + ) + assert att_cls == ATTENTION_ROUTINE + assert req_human is False + + # Structured flags still escalate (machine-driven) + att_cls, req_human = classify_attention_event( + CATEGORY_SYSTEM, + "anything", + "anything with hard stop in text", + is_hard_stop=True, + ) + assert att_cls == ATTENTION_HUMAN_REQUIRED + assert req_human is True + + +def test_notification_ids_are_unique_across_probe_errors_and_collisions(): + """B2: published notification ids must be unique within a snapshot (#905).""" + reg = load_registry() + proj_id = reg.projects[0].id if reg.projects else "gitea-tools" + + mock_queue = QueueSnapshot( + project_id=proj_id, + repo_label="org/repo", + prs=(), + issues=(), + pr_pagination=None, + issue_pagination=None, + ) + mock_leases = LeaseSnapshot( + project_id=proj_id, + repo_label="org/repo", + issue_lock=None, + claim_inventory={}, + reviewer_leases=(), + duplicate_prs=( + CollisionWarning( + kind="duplicate_pr", + message="Multiple open PRs for issue #10", + issue_number=10, + pr_numbers=(10, 11), + ), + CollisionWarning( + kind="duplicate_branch", + message="Another collision without issue", + issue_number=None, + pr_numbers=(12, 13), + ), + CollisionWarning( + kind="duplicate_pr", + message="Second issue collision", + issue_number=10, + pr_numbers=(14, 15), + ), + ), + duplicate_branches=(), + collision_history=(), + fetch_error=None, + ) + mock_version = VersionInfo( + git_sha="abc1234", + git_describe="v1.0.0", + control_plane_schema_version=1, + python_version="3.11", + known=True, + ) + mock_stale = StaleRuntime( + daemon_head="abc1234", + checkout_head="abc1234", + remote_head="abc1234", + stale=False, + determinable=True, + mutation_safe=True, + reasons=(), + ) + mock_health = SystemHealthSnapshot( + status="degraded", + ready=False, + readiness_complete=True, + readiness_reasons=(), + service="webui", + mode="test", + version=mock_version, + started_at="2026-07-25T00:00:00Z", + uptime_seconds=100.0, + timestamp="2026-07-25T00:00:00Z", + deep_probes_requested=True, + dependencies=(), + mcp_namespaces=(), + stale_runtime=mock_stale, + probe_errors=("error alpha", "error beta"), + ) + + snapshot = load_notifications_snapshot( + proj_id, + load_queue=lambda _id: mock_queue, + load_leases=lambda **_kwargs: mock_leases, + load_health=lambda **_kwargs: mock_health, + ) + ids = [item.id for item in snapshot.items] + assert len(ids) == len(set(ids)), f"duplicate notification ids: {ids}" + assert any(i.startswith(f"notif-sys-err-{proj_id}-") for i in ids) + assert any(i.startswith("notif-collision-") for i in ids) + + +def test_probe_errors_do_not_set_fetch_error(): + """B3: probe_errors must not be reported as fetch_error (#905).""" + reg = load_registry() + proj_id = reg.projects[0].id if reg.projects else "gitea-tools" + + mock_queue = QueueSnapshot( + project_id=proj_id, + repo_label="org/repo", + prs=(), + issues=(), + pr_pagination=None, + issue_pagination=None, + fetch_error=None, + ) + mock_leases = LeaseSnapshot( + project_id=proj_id, + repo_label="org/repo", + issue_lock=None, + claim_inventory={}, + reviewer_leases=(), + duplicate_prs=(), + duplicate_branches=(), + collision_history=(), + fetch_error=None, + ) + mock_version = VersionInfo( + git_sha="abc1234", + git_describe="v1.0.0", + control_plane_schema_version=1, + python_version="3.11", + known=True, + ) + mock_stale = StaleRuntime( + daemon_head="abc1234", + checkout_head="abc1234", + remote_head="abc1234", + stale=False, + determinable=True, + mutation_safe=True, + reasons=(), + ) + mock_health = SystemHealthSnapshot( + status="degraded", + ready=False, + readiness_complete=True, + readiness_reasons=(), + service="webui", + mode="test", + version=mock_version, + started_at="2026-07-25T00:00:00Z", + uptime_seconds=100.0, + timestamp="2026-07-25T00:00:00Z", + deep_probes_requested=True, + dependencies=(), + mcp_namespaces=(), + stale_runtime=mock_stale, + probe_errors=("probe blew up",), + ) + + snapshot = load_notifications_snapshot( + proj_id, + load_queue=lambda _id: mock_queue, + load_leases=lambda **_kwargs: mock_leases, + load_health=lambda **_kwargs: mock_health, + ) + assert snapshot.fetch_error is None + # probe errors still appear as items + assert any("probe blew up" in item.summary for item in snapshot.items) diff --git a/tests/test_webui_restart_console.py b/tests/test_webui_restart_console.py new file mode 100644 index 0000000..439bba5 --- /dev/null +++ b/tests/test_webui_restart_console.py @@ -0,0 +1,452 @@ +"""Read-only restart console: views, gates, and honesty rules (#667). + +The console consumes the #655 substrate. These tests hold it to the three +properties that make a status surface trustworthy: + +* an unreadable source is reported unavailable, never rendered as green; +* authorization is probed the way execution would probe it, so an allow is + never shown for something that could not run; +* the surface performs no mutation, including no write to the control-plane DB. +""" + +from __future__ import annotations + +import os +import sqlite3 +import sys +import tempfile +import unittest +from datetime import datetime, timedelta, timezone +from pathlib import Path + +sys.path.insert(0, str(Path(__file__).resolve().parent.parent)) + +from starlette.testclient import TestClient # noqa: E402 + +import restart_coordinator # noqa: E402 +from webui import console_authz, restart_console, restart_views # noqa: E402 +from webui.app import create_app # noqa: E402 + +NOW = datetime(2026, 7, 25, 21, 0, 0, tzinfo=timezone.utc) + + +def _principal(role: str) -> console_authz.Principal: + return console_authz.Principal( + subject="operator@example.com", + role=role, + identity_source=console_authz.IDENTITY_LOCAL_DEV, + authenticated=True, + ) + + +def _inventory(*, complete: bool = True, sessions=(), leases=()): + def _read(**_kwargs): + return { + "sessions": list(sessions), + "leases": list(leases), + "terminal_lock": None, + "prior_recovery_attempts": [], + "inventory_complete": complete, + "incomplete_reasons": ( + [] if complete else ["fixture: inventory withheld"] + ), + } + + return _read + + +def _live_session(session_id: str = "prgs-author-1234-abcd") -> dict: + return { + "session_id": session_id, + "role": "author", + "profile": "prgs-author", + "pid": os.getpid(), + "status": "active", + "last_heartbeat_at": (NOW - timedelta(seconds=30)).isoformat(), + } + + +def drain_proof_fixture() -> dict: + """A structurally complete but unsigned drain proof.""" + return { + "version": "drain-proof/v1", + "proof_id": "deadbeef" * 8, + "clean": True, + "issued_at": (NOW - timedelta(minutes=1)).isoformat(), + "expires_at": (NOW + timedelta(minutes=5)).isoformat(), + "requesting_session_id": "s-live", + "impact_fingerprint": "f" * 64, + "checks": [], + "failed_checks": [], + } + + +class RestartClassMatrixTest(unittest.TestCase): + def test_every_policy_class_is_rendered(self) -> None: + views = restart_console.build_restart_class_views("operator") + self.assertEqual(len(views), len(restart_coordinator.RESTART_CLASS_POLICIES)) + + def test_viewer_capability_is_role_scoped_not_generic(self) -> None: + """A worker role must not be shown as able to request a full restart.""" + author = { + v.restart_class: v + for v in restart_console.build_restart_class_views("author") + } + operator = { + v.restart_class: v + for v in restart_console.build_restart_class_views("operator") + } + full = restart_coordinator.RestartClass.FULL_MCP_RESTART.value + + self.assertFalse(author[full].viewer_may_request) + self.assertFalse(author[full].viewer_may_execute) + self.assertTrue(operator[full].viewer_may_request) + self.assertTrue(operator[full].viewer_may_execute) + + def test_unknown_role_may_do_nothing(self) -> None: + views = restart_console.build_restart_class_views("not-a-role") + self.assertTrue(all(not v.viewer_may_request for v in views)) + self.assertTrue(all(not v.viewer_may_execute for v in views)) + + +class AuthorizationProbeTest(unittest.TestCase): + def test_probe_asks_for_execution_so_phase_gate_is_reported(self) -> None: + """An admin clears the role bar and still cannot execute in Phase 1. + + This is the case that distinguishes the two probes. Asked without + ``for_execution`` an admin is *allowed* for ``system.restart_namespace``, + which on a control surface reads as a live button. Asked the way + execution asks, the same principal is refused ``phase_not_active``. The + console must report the second answer. + """ + by_id = { + a.action_id: a + for a in restart_console.build_action_authorizations( + _principal(console_authz.ADMIN) + ) + } + restart = by_id["system.restart_namespace"] + + self.assertFalse(restart.execution_enabled) + self.assertEqual(restart.reason_code, console_authz.DENY_PHASE_NOT_ACTIVE) + + permissive = console_authz.authorize( + "system.restart_namespace", _principal(console_authz.ADMIN) + ) + self.assertTrue( + permissive.allowed, + "guard precondition: without for_execution an admin is allowed, " + "which is exactly why the console must not probe that way", + ) + + def test_operator_is_refused_the_admin_only_restart_action(self) -> None: + """Role refusal precedes the phase gate and is reported as such.""" + by_id = { + a.action_id: a + for a in restart_console.build_action_authorizations( + _principal(console_authz.OPERATOR) + ) + } + self.assertEqual( + by_id["system.restart_namespace"].reason_code, + console_authz.DENY_INSUFFICIENT_ROLE, + ) + + def test_anonymous_is_denied_unauthenticated(self) -> None: + by_id = { + a.action_id: a for a in restart_console.build_action_authorizations(None) + } + self.assertEqual( + by_id["system.restart_namespace"].reason_code, + console_authz.DENY_UNAUTHENTICATED, + ) + + def test_no_authorization_ever_reports_execution_enabled(self) -> None: + for role in ( + console_authz.VIEWER, + console_authz.OPERATOR, + console_authz.CONTROLLER, + console_authz.ADMIN, + ): + for auth in restart_console.build_action_authorizations(_principal(role)): + self.assertFalse( + auth.execution_enabled, + f"{role} reported execution_enabled for {auth.action_id}", + ) + + +class ImpactPreviewTest(unittest.TestCase): + def test_impact_renders_from_coordinator_dto(self) -> None: + impact, source = restart_console.load_impact_report( + principal=_principal(console_authz.OPERATOR), + read_inventory=_inventory(sessions=[_live_session()]), + now=NOW, + ) + self.assertTrue(source.available) + self.assertIsNotNone(impact) + self.assertEqual( + impact["restart_class"], + restart_coordinator.RestartClass.FULL_MCP_RESTART.value, + ) + self.assertIn("verdict", impact) + self.assertFalse(impact["restart_performed"]) + self.assertTrue(impact["dry_run"]) + + def test_incomplete_inventory_is_surfaced_and_denies(self) -> None: + impact, source = restart_console.load_impact_report( + principal=_principal(console_authz.OPERATOR), + read_inventory=_inventory(complete=False), + now=NOW, + ) + self.assertFalse(impact["inventory_complete"]) + self.assertFalse(impact["allow_restart"]) + self.assertTrue(source.detail, "incomplete inventory must explain itself") + + def test_inventory_reader_failure_is_unavailable_not_empty(self) -> None: + """A reader that raises must not be rendered as 'no sessions affected'.""" + + def _boom(**_kwargs): + raise RuntimeError("control-plane unreachable") + + impact, source = restart_console.load_impact_report( + principal=_principal(console_authz.OPERATOR), + read_inventory=_boom, + now=NOW, + ) + self.assertIsNone(impact) + self.assertFalse(source.available) + self.assertIn("control-plane unreachable", source.detail) + + +class ControlPlaneReadTest(unittest.TestCase): + def test_missing_database_is_incomplete_not_empty(self) -> None: + inventory = restart_console.read_control_plane_inventory( + db_path="/nonexistent/control-plane.sqlite3" + ) + self.assertFalse(inventory["inventory_complete"]) + self.assertEqual(inventory["sessions"], []) + self.assertTrue(inventory["incomplete_reasons"]) + + def test_reader_never_creates_the_database(self) -> None: + """Reading status must not bring a control-plane DB into existence. + + The path deliberately sits in a directory that already exists: a + read-write ``sqlite3.connect`` would happily create the file there, so + this fails if the reader ever stops opening the database ``mode=ro``. + A nested-missing-directory path would pass for the wrong reason, + because sqlite cannot create the parent directory either way. + """ + with tempfile.TemporaryDirectory() as tmp: + path = os.path.join(tmp, "control_plane.sqlite3") + self.assertTrue(os.path.isdir(os.path.dirname(path))) + + inventory = restart_console.read_control_plane_inventory(db_path=path) + + self.assertFalse( + os.path.exists(path), + "reading restart status created a control-plane database", + ) + self.assertFalse(inventory["inventory_complete"]) + + def test_reads_active_sessions_from_a_real_database(self) -> None: + with tempfile.TemporaryDirectory() as tmp: + path = os.path.join(tmp, "cp.sqlite3") + conn = sqlite3.connect(path) + conn.execute( + "CREATE TABLE sessions (session_id TEXT, role TEXT, profile TEXT," + " pid INTEGER, status TEXT, last_heartbeat_at TEXT)" + ) + conn.execute( + "CREATE TABLE work_items (work_item_id INTEGER, kind TEXT," + " number INTEGER)" + ) + conn.execute( + "CREATE TABLE leases (lease_id TEXT, session_id TEXT, role TEXT," + " phase TEXT, status TEXT, worktree_path TEXT," + " work_item_id INTEGER, expires_at TEXT)" + ) + conn.execute( + "INSERT INTO sessions VALUES (?,?,?,?,?,?)", + ("s-live", "author", "prgs-author", 4242, "active", NOW.isoformat()), + ) + conn.execute( + "INSERT INTO sessions VALUES (?,?,?,?,?,?)", + ("s-done", "author", "prgs-author", 11, "closed", NOW.isoformat()), + ) + conn.execute("INSERT INTO work_items VALUES (1, 'issue', 667)") + conn.execute( + "INSERT INTO leases VALUES (?,?,?,?,?,?,?,?)", + ( + "l-1", + "s-live", + "author", + "allocated", + "active", + None, + 1, + NOW.isoformat(), + ), + ) + conn.commit() + conn.close() + + inventory = restart_console.read_control_plane_inventory(db_path=path) + + self.assertTrue(inventory["inventory_complete"]) + self.assertEqual([s["session_id"] for s in inventory["sessions"]], ["s-live"]) + self.assertEqual(inventory["leases"][0]["work_number"], 667) + + +class DrainAndReconcileTest(unittest.TestCase): + def test_absent_drain_proof_is_not_a_pass(self) -> None: + drain, source = restart_console.load_drain_status(proof=None, now=NOW) + self.assertIsNone(drain) + self.assertFalse(source.available) + self.assertIn("denies", source.detail) + + def test_tampered_drain_proof_is_reported_invalid(self) -> None: + proof = drain_proof_fixture() + proof["clean"] = True + proof["proof_id"] = "0" * 64 + drain, source = restart_console.load_drain_status(proof=proof, now=NOW) + self.assertTrue(source.available) + self.assertFalse(drain["valid"]) + + def test_absent_reconcile_proof_is_unavailable(self) -> None: + reconcile, source = restart_console.load_reconcile_status(load_proof=None) + self.assertIsNone(reconcile) + self.assertFalse(source.available) + + def test_reconcile_proof_is_rendered_when_supplied(self) -> None: + payload = { + "overall_status": "degraded", + "mode": "log_only", + "resolved_count": 3, + "unresolved_count": 2, + "items": [ + { + "dimension": "leases", + "status": "unresolved", + "summary": "2 orphaned leases", + "follow_up_required": True, + } + ], + } + reconcile, source = restart_console.load_reconcile_status( + load_proof=lambda: payload + ) + self.assertTrue(source.available) + self.assertEqual(reconcile["unresolved_count"], 2) + + +class RenderingTest(unittest.TestCase): + def _snapshot(self, **kwargs): + params = { + "principal": _principal(console_authz.OPERATOR), + "read_inventory": _inventory(sessions=[_live_session()]), + "now": NOW, + } + params.update(kwargs) + return restart_console.load_restart_console_snapshot(**params) + + def test_page_renders_every_section(self) -> None: + html = restart_views.render_restart_console_page(self._snapshot()) + for heading in ( + "Impact preview", + "Drain proof", + "Post-restart reconcile", + "Restart classes", + "Approval controls", + "Break-glass", + ): + self.assertIn(heading, html) + + def test_hostile_session_id_is_escaped(self) -> None: + hostile = "" + html = restart_views.render_restart_console_page( + self._snapshot(read_inventory=_inventory(sessions=[_live_session(hostile)])) + ) + self.assertNotIn("