The Related section and cross-reference lines described #652, #653, #630, #642, and #591 by roles they do not hold. Align each description with the linked issue's actual title and scope: - #652 is the Control Plane Web Console product vision (restart controls live in its capability area A), not a restart-specific vision. - #653 is the console phased-delivery roadmap; restart controls are Phase 2. - #630 is the manual process-kill contamination guard, not the coordinator; the coordinator remains an unimplemented later child of #655. - #642 is the sanctioned restart / graceful reload console UX. - #591 is auto-restart on master advance (closed); only #584 is transport-flap reconnect. They were previously collapsed into one transport-recovery label. Documentation-only wording change. Policy IDs RG-01..RG-08, the policy version restart-governance/v1, the authorization matrix, and every normative statement are unchanged. Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
224 lines
11 KiB
Markdown
224 lines
11 KiB
Markdown
# ADR: MCP restart governance and authorization policy
|
||
|
||
- **Status:** Accepted (policy effective immediately for LLM and operator sessions; enforcement tooling may lag)
|
||
- **Date:** 2026-07-23
|
||
- **Tracking issue:** [#656](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/656)
|
||
- **Policy version:** `restart-governance/v1`
|
||
- **Related:**
|
||
- Umbrella: [#655](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/655) — governed MCP restart coordination and zero-disruption recovery
|
||
- Vision: [#652](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/652) — MCP Control Plane Web Console product vision (§A system health and process control)
|
||
- Roadmap: [#653](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/653) — Control Plane Web Console phased delivery (Phase 2 restart controls)
|
||
- Contamination guard: [#630](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/630) — blocks manual process-kill recovery
|
||
- Console restart UX: [#642](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/642) — sanctioned restart and graceful reload
|
||
- Existing restart / reconnect paths to inventory: [#591](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/591) — auto-restart on master advance (closed); [#584](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/584) — host auto-reconnect on transport flap
|
||
- Stable-control runtime split: `docs/architecture/mcp-stable-control-runtime-policy-adr.md` (#615)
|
||
- Client-namespace health: `docs/mcp-namespace-health.md` (#543)
|
||
- Reconnect-only EOF recovery: `docs/mcp-namespace-eof-recovery.md`
|
||
|
||
## 1. Context
|
||
|
||
The Gitea MCP server is the **control plane** for real issue and PR mutations
|
||
(create, comment, lock, review, merge, reconcile). The same process serves every
|
||
role namespace (`gitea-author`, `gitea-reviewer`, `gitea-merger`,
|
||
`gitea-reconciler`, `gitea-controller`) and holds the in-memory capability-gate
|
||
code loaded at startup.
|
||
|
||
Restarting that process is destructive to concurrent work:
|
||
|
||
- It resets every session's identity, preflight, and capability-lease binding.
|
||
- It can interrupt a mutation mid-critical-section (a lock acquire, a review
|
||
submit, a merge), leaving durable state half-written.
|
||
- Relaunching from the wrong checkout or worktree silently changes which code
|
||
the control plane runs, defeating master-parity gates (#420 / #615).
|
||
|
||
Today there is **no durable written policy** stating who may restart MCP, under
|
||
what conditions, that restart is a last resort, and how controller approval,
|
||
automated safety gates, and break-glass interact. Operators and LLM sessions
|
||
therefore invent restart behavior ad hoc, which makes concurrent multi-role work
|
||
unsafe. #630 and #642 need this policy as their backbone.
|
||
|
||
This ADR defines that policy. It does **not** implement coordinator code or HA
|
||
multi-instance restart (those are later children of #655).
|
||
|
||
## 2. Decision
|
||
|
||
### 2.1 v1 decision (recorded)
|
||
|
||
**Restart authority in v1 is `controller approval + automated safety gates`.**
|
||
|
||
A restart of the stable control runtime is authorized only when **both** hold:
|
||
|
||
1. A **controller** role explicitly approves the restart, recording an audit
|
||
entry (who, why, scope, affected sessions), **and**
|
||
2. The **automated safety gates** pass: a completed drain acknowledgement (no
|
||
affected session is mid-critical-section) or a declared break-glass incident
|
||
(§2.5).
|
||
|
||
Quorum among multiple controllers is **not** required day-one. It is deferred
|
||
unless a later investigation (tracked under #653) proves single-controller
|
||
approval is insufficient. This ADR records the v1 decision so enforcement code
|
||
(#630) has a fixed target; changing it requires a superseding ADR.
|
||
|
||
### 2.2 Restart is a last resort — the recovery ladder
|
||
|
||
Restart is the **last** rung. Before any restart, exhaust the narrower
|
||
recoveries, in order:
|
||
|
||
1. **Reconnect** the IDE/client MCP namespace (transport EOF, `client is
|
||
closing: EOF`, transient `#584` flap). No process change. See
|
||
`docs/mcp-namespace-eof-recovery.md`.
|
||
2. **Refresh / rebind** the session workspace: re-run `gitea_whoami`,
|
||
`gitea_resolve_task_capability`, and pass an explicit validated
|
||
`worktree_path`. Fixes stale session context without touching the process.
|
||
3. **Scoped restart** of a single misbehaving namespace/service (where the
|
||
deployment supports per-service restart) rather than the whole control plane.
|
||
4. **Full restart** of the stable control runtime process — operator-owned,
|
||
controller-approved, drained.
|
||
5. **Host / infrastructure restart** — the broadest action; same authorization
|
||
as a full restart plus infrastructure ownership.
|
||
|
||
A session **must** try rungs 1–2 and record why they were insufficient before
|
||
requesting a restart at rung 3 or above. Skipping straight to restart is a
|
||
policy violation.
|
||
|
||
### 2.3 Authorization matrix
|
||
|
||
| Role | Reconnect (1) | Refresh/rebind (2) | Scoped restart (3) | Full restart (4) | Host restart (5) |
|
||
|---|---|---|---|---|---|
|
||
| **author** | self | self | request only | **forbidden** | forbidden |
|
||
| **reviewer** | self | self | request only | **forbidden** | forbidden |
|
||
| **merger** | self | self | request only | **forbidden** | forbidden |
|
||
| **reconciler** | self | self | request only | **forbidden** | forbidden |
|
||
| **controller** | self | self | **approve** (+gates) | **approve** (+gates) | request to operator |
|
||
| **operator** | self | self | execute (controller-approved) | execute (controller-approved) | execute (controller-approved) |
|
||
| **admin** | self | self | execute | execute | execute (break-glass) |
|
||
|
||
Legend: *self* = may perform for its own client session; *request only* = may
|
||
raise a restart request but not authorize or execute it; *approve* = may
|
||
authorize under §2.1 gates; *execute* = may perform the process action after the
|
||
authorization is recorded.
|
||
|
||
Key invariants:
|
||
|
||
- **No LLM worker role (author/reviewer/merger/reconciler) may perform or
|
||
authorize a full or host restart.** They may only reconnect/rebind their own
|
||
client and file a restart request.
|
||
- **Controller approval authorizes; operator/admin executes.** The approving
|
||
controller and the executing operator may be the same human, but both the
|
||
approval and the execution are audited.
|
||
- Privileged process actions (full restart, host restart) are reserved to
|
||
**operator/admin**, never to an automated worker.
|
||
|
||
### 2.4 Approved conditions
|
||
|
||
A restart at rung 3+ is approved only under one of these recorded conditions:
|
||
|
||
- **No affected sessions:** the control plane has no live session that would be
|
||
interrupted (verified, not assumed).
|
||
- **Full drain acknowledged:** every affected session has drained
|
||
(no open critical section — no held mutation lease mid-write) and the drain is
|
||
acknowledged in the audit record.
|
||
- **Controller + gates:** controller approval plus passing automated safety
|
||
gates (§2.1), the standard v1 path.
|
||
- **Quorum:** not required in v1; reserved for a future superseding ADR.
|
||
- **Break-glass:** an incident-backed emergency exception (§2.5).
|
||
|
||
Restart **never** bypasses mutation gates mid-critical-section. Drain before
|
||
restart is mandatory except under break-glass with a declared incident.
|
||
|
||
### 2.5 Break-glass
|
||
|
||
Break-glass is a **separate, narrower** authorization path for emergencies where
|
||
the normal drain-and-approve path cannot complete (e.g. the control plane is
|
||
wedged and cannot drain).
|
||
|
||
Break-glass conditions:
|
||
|
||
- A declared incident record exists (id, timestamp, declarer) **before** the
|
||
action.
|
||
- The action is taken by **operator or admin** authority only — never by an LLM
|
||
worker role, and never unilaterally by an operator with active peers when a
|
||
controller is reachable.
|
||
- The scope is the minimum necessary rung of the ladder.
|
||
- A **mandatory post-hoc audit** entry is filed: what was restarted, why the
|
||
normal path was impossible, which sessions were affected, and the incident id.
|
||
|
||
Break-glass suspends the drain requirement, not the audit requirement.
|
||
|
||
### 2.6 Explicit prohibitions
|
||
|
||
- **A unilateral LLM or operator full restart while active peer sessions
|
||
exist is forbidden.** An LLM worker role must not kill, restart, or relaunch
|
||
the MCP process; a lone operator must not full-restart over live peer work
|
||
without controller approval or a break-glass incident.
|
||
- Process-kill recovery is forbidden as a routine tool (#630). This ADR does not
|
||
introduce a kill path.
|
||
- Ambiguous policy state **denies** restart (§4).
|
||
|
||
## 3. Security requirements
|
||
|
||
- Full restart and host restart are **privileged**; only operator/admin execute
|
||
them, only after a controller approval or break-glass incident is recorded.
|
||
- Break-glass is a distinct authorization path with its own audit mandate; it is
|
||
never the default and never silent.
|
||
- **Every approval and every restart action is audited** (who approved, who
|
||
executed, scope, affected sessions, condition, policy version). No restart is
|
||
authorized without a durable audit entry.
|
||
|
||
## 4. Failure behavior
|
||
|
||
**Ambiguous policy → deny restart.** If it cannot be established that a
|
||
restart is authorized under §2 — unknown affected-session state, missing
|
||
controller approval, absent break-glass incident, or an unclassifiable request —
|
||
the safe action is to **refuse** the restart and stop with a recovery report,
|
||
never to restart on assumption.
|
||
|
||
## 5. Policy IDs (for enforcement code)
|
||
|
||
Enforcement code — the restart coordinator (a later child of #655), the #630
|
||
contamination guard, and the #642 console restart UX — binds to these stable
|
||
policy identifiers rather than to prose:
|
||
|
||
| Policy ID | Statement |
|
||
|---|---|
|
||
| `RG-01` | Restart is last resort; rungs 1–2 must be tried and recorded first (§2.2). |
|
||
| `RG-02` | v1 authority = controller approval + automated safety gates (§2.1). |
|
||
| `RG-03` | No LLM worker role performs or authorizes full/host restart (§2.3). |
|
||
| `RG-04` | Full/host restart executed by operator/admin only, post approval (§2.3). |
|
||
| `RG-05` | Drain before restart is mandatory except break-glass with incident (§2.4). |
|
||
| `RG-06` | Break-glass requires a pre-declared incident and post-hoc audit (§2.5). |
|
||
| `RG-07` | Unilateral LLM/operator full restart with active peers is forbidden (§2.6). |
|
||
| `RG-08` | Ambiguous policy state denies restart (§4). |
|
||
|
||
The `restart-governance/v1` **policy version** field is emitted on future
|
||
restart audit events so approvals can be reconciled against the policy revision
|
||
in force.
|
||
|
||
## 6. Dogfooding
|
||
|
||
Gitea-Tools governs its own MCP control plane by this policy. Author, reviewer,
|
||
merger, and reconciler sessions operating on this repository use the recovery
|
||
ladder (§2.2) — reconnect and rebind, never self-restart — and any real restart
|
||
of the Gitea-Tools stable control runtime follows the controller-approval +
|
||
drain path defined here.
|
||
|
||
## 7. Acceptance and cross-links
|
||
|
||
This ADR is the authoritative restart-governance policy. It **must** stay
|
||
cross-linked from the safety model and the web-console deployment boundary:
|
||
|
||
- `docs/safety-model.md` § Process restart governance references this ADR.
|
||
- `docs/webui-deployment.md` references this ADR for restart/reload disposition.
|
||
|
||
It is linked to its issue lineage — umbrella **#655**, vision **#652**, roadmap
|
||
**#653**, contamination guard **#630**, and console restart UX **#642** — in
|
||
§ Related above.
|
||
|
||
## 8. Non-goals
|
||
|
||
- Implementing the restart coordinator or approval state machine (#630, later
|
||
children of #655).
|
||
- Implementing HA multi-instance restart or quorum machinery.
|
||
- Introducing any process-kill or auto-restart tool; existing auto-restart
|
||
behavior must be inventoried before any new restart tool is enabled.
|