# ADR: MCP restart governance and authorization policy - **Status:** Accepted (policy effective immediately for LLM and operator sessions; enforcement tooling may lag) - **Date:** 2026-07-23 - **Tracking issue:** [#656](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/656) - **Policy version:** `restart-governance/v1` - **Related:** - Umbrella: [#655](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/655) — governed MCP restart coordination and zero-disruption recovery - Vision: [#652](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/652) — MCP Control Plane Web Console product vision (§A system health and process control) - Roadmap: [#653](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/653) — Control Plane Web Console phased delivery (Phase 2 restart controls) - Contamination guard: [#630](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/630) — blocks manual process-kill recovery - Console restart UX: [#642](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/642) — sanctioned restart and graceful reload - Existing restart / reconnect paths to inventory: [#591](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/591) — auto-restart on master advance (closed); [#584](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/584) — host auto-reconnect on transport flap - Stable-control runtime split: `docs/architecture/mcp-stable-control-runtime-policy-adr.md` (#615) - Client-namespace health: `docs/mcp-namespace-health.md` (#543) - Reconnect-only EOF recovery: `docs/mcp-namespace-eof-recovery.md` ## 1. Context The Gitea MCP server is the **control plane** for real issue and PR mutations (create, comment, lock, review, merge, reconcile). The same process serves every role namespace (`gitea-author`, `gitea-reviewer`, `gitea-merger`, `gitea-reconciler`, `gitea-controller`) and holds the in-memory capability-gate code loaded at startup. Restarting that process is destructive to concurrent work: - It resets every session's identity, preflight, and capability-lease binding. - It can interrupt a mutation mid-critical-section (a lock acquire, a review submit, a merge), leaving durable state half-written. - Relaunching from the wrong checkout or worktree silently changes which code the control plane runs, defeating master-parity gates (#420 / #615). Today there is **no durable written policy** stating who may restart MCP, under what conditions, that restart is a last resort, and how controller approval, automated safety gates, and break-glass interact. Operators and LLM sessions therefore invent restart behavior ad hoc, which makes concurrent multi-role work unsafe. #630 and #642 need this policy as their backbone. This ADR defines that policy. It does **not** implement coordinator code or HA multi-instance restart (those are later children of #655). ## 2. Decision ### 2.1 v1 decision (recorded) **Restart authority in v1 is `controller approval + automated safety gates`.** A restart of the stable control runtime is authorized only when **both** hold: 1. A **controller** role explicitly approves the restart, recording an audit entry (who, why, scope, affected sessions), **and** 2. The **automated safety gates** pass: a completed drain acknowledgement (no affected session is mid-critical-section) or a declared break-glass incident (§2.5). Quorum among multiple controllers is **not** required day-one. It is deferred unless a later investigation (tracked under #653) proves single-controller approval is insufficient. This ADR records the v1 decision so enforcement code (#630) has a fixed target; changing it requires a superseding ADR. ### 2.2 Restart is a last resort — the recovery ladder Restart is the **last** rung. Before any restart, exhaust the narrower recoveries, in order: 1. **Reconnect** the IDE/client MCP namespace (transport EOF, `client is closing: EOF`, transient `#584` flap). No process change. See `docs/mcp-namespace-eof-recovery.md`. 2. **Refresh / rebind** the session workspace: re-run `gitea_whoami`, `gitea_resolve_task_capability`, and pass an explicit validated `worktree_path`. Fixes stale session context without touching the process. 3. **Scoped restart** of a single misbehaving namespace/service (where the deployment supports per-service restart) rather than the whole control plane. 4. **Full restart** of the stable control runtime process — operator-owned, controller-approved, drained. 5. **Host / infrastructure restart** — the broadest action; same authorization as a full restart plus infrastructure ownership. A session **must** try rungs 1–2 and record why they were insufficient before requesting a restart at rung 3 or above. Skipping straight to restart is a policy violation. ### 2.3 Authorization matrix | Role | Reconnect (1) | Refresh/rebind (2) | Scoped restart (3) | Full restart (4) | Host restart (5) | |---|---|---|---|---|---| | **author** | self | self | request only | **forbidden** | forbidden | | **reviewer** | self | self | request only | **forbidden** | forbidden | | **merger** | self | self | request only | **forbidden** | forbidden | | **reconciler** | self | self | request only | **forbidden** | forbidden | | **controller** | self | self | **approve** (+gates) | **approve** (+gates) | request to operator | | **operator** | self | self | execute (controller-approved) | execute (controller-approved) | execute (controller-approved) | | **admin** | self | self | execute | execute | execute (break-glass) | Legend: *self* = may perform for its own client session; *request only* = may raise a restart request but not authorize or execute it; *approve* = may authorize under §2.1 gates; *execute* = may perform the process action after the authorization is recorded. Key invariants: - **No LLM worker role (author/reviewer/merger/reconciler) may perform or authorize a full or host restart.** They may only reconnect/rebind their own client and file a restart request. - **Controller approval authorizes; operator/admin executes.** The approving controller and the executing operator may be the same human, but both the approval and the execution are audited. - Privileged process actions (full restart, host restart) are reserved to **operator/admin**, never to an automated worker. ### 2.4 Approved conditions A restart at rung 3+ is approved only under one of these recorded conditions: - **No affected sessions:** the control plane has no live session that would be interrupted (verified, not assumed). - **Full drain acknowledged:** every affected session has drained (no open critical section — no held mutation lease mid-write) and the drain is acknowledged in the audit record. - **Controller + gates:** controller approval plus passing automated safety gates (§2.1), the standard v1 path. - **Quorum:** not required in v1; reserved for a future superseding ADR. - **Break-glass:** an incident-backed emergency exception (§2.5). Restart **never** bypasses mutation gates mid-critical-section. Drain before restart is mandatory except under break-glass with a declared incident. ### 2.5 Break-glass Break-glass is a **separate, narrower** authorization path for emergencies where the normal drain-and-approve path cannot complete (e.g. the control plane is wedged and cannot drain). Break-glass conditions: - A declared incident record exists (id, timestamp, declarer) **before** the action. - The action is taken by **operator or admin** authority only — never by an LLM worker role, and never unilaterally by an operator with active peers when a controller is reachable. - The scope is the minimum necessary rung of the ladder. - A **mandatory post-hoc audit** entry is filed: what was restarted, why the normal path was impossible, which sessions were affected, and the incident id. Break-glass suspends the drain requirement, not the audit requirement. ### 2.6 Explicit prohibitions - **A unilateral LLM or operator full restart while active peer sessions exist is forbidden.** An LLM worker role must not kill, restart, or relaunch the MCP process; a lone operator must not full-restart over live peer work without controller approval or a break-glass incident. - Process-kill recovery is forbidden as a routine tool (#630). This ADR does not introduce a kill path. - Ambiguous policy state **denies** restart (§4). ## 3. Security requirements - Full restart and host restart are **privileged**; only operator/admin execute them, only after a controller approval or break-glass incident is recorded. - Break-glass is a distinct authorization path with its own audit mandate; it is never the default and never silent. - **Every approval and every restart action is audited** (who approved, who executed, scope, affected sessions, condition, policy version). No restart is authorized without a durable audit entry. ## 4. Failure behavior **Ambiguous policy → deny restart.** If it cannot be established that a restart is authorized under §2 — unknown affected-session state, missing controller approval, absent break-glass incident, or an unclassifiable request — the safe action is to **refuse** the restart and stop with a recovery report, never to restart on assumption. ## 5. Policy IDs (for enforcement code) Enforcement code — the restart coordinator (a later child of #655), the #630 contamination guard, and the #642 console restart UX — binds to these stable policy identifiers rather than to prose: | Policy ID | Statement | |---|---| | `RG-01` | Restart is last resort; rungs 1–2 must be tried and recorded first (§2.2). | | `RG-02` | v1 authority = controller approval + automated safety gates (§2.1). | | `RG-03` | No LLM worker role performs or authorizes full/host restart (§2.3). | | `RG-04` | Full/host restart executed by operator/admin only, post approval (§2.3). | | `RG-05` | Drain before restart is mandatory except break-glass with incident (§2.4). | | `RG-06` | Break-glass requires a pre-declared incident and post-hoc audit (§2.5). | | `RG-07` | Unilateral LLM/operator full restart with active peers is forbidden (§2.6). | | `RG-08` | Ambiguous policy state denies restart (§4). | The `restart-governance/v1` **policy version** field is emitted on future restart audit events so approvals can be reconciled against the policy revision in force. ## 6. Dogfooding Gitea-Tools governs its own MCP control plane by this policy. Author, reviewer, merger, and reconciler sessions operating on this repository use the recovery ladder (§2.2) — reconnect and rebind, never self-restart — and any real restart of the Gitea-Tools stable control runtime follows the controller-approval + drain path defined here. ## 7. Acceptance and cross-links This ADR is the authoritative restart-governance policy. It **must** stay cross-linked from the safety model and the web-console deployment boundary: - `docs/safety-model.md` § Process restart governance references this ADR. - `docs/webui-deployment.md` references this ADR for restart/reload disposition. It is linked to its issue lineage — umbrella **#655**, vision **#652**, roadmap **#653**, contamination guard **#630**, and console restart UX **#642** — in § Related above. ## 8. Non-goals - Implementing the restart coordinator or approval state machine (#630, later children of #655). - Implementing HA multi-instance restart or quorum machinery. - Introducing any process-kill or auto-restart tool; existing auto-restart behavior must be inventoried before any new restart tool is enabled.