feat: enforce MCP restart class permissions (#663)

This commit is contained in:
2026-07-24 18:10:22 -04:00
parent 870843f999
commit 714190e02a
7 changed files with 721 additions and 18 deletions
+53
View File
@@ -0,0 +1,53 @@
# MCP restart classes and blast-radius permissions (#663)
This is the machine-enforced class matrix used by
`restart_coordinator.RESTART_CLASS_POLICIES`. It implements the narrower-first
recovery ladder from #655 and the authorization policy from #656, using the
path inventory from #657 and the impact coordinator from #658. Product and
delivery lineage: vision #652 and roadmap #653.
Unknown class names are denied. The coordinator requires both the class
permission and an eligible request role. Approval gates are additional: a
caller cannot turn a request permission into execution authority.
| Restart class | Required permission | Expected blast radius | Drain requirement | Approval requirement | Audit requirement | Recovery behavior |
|---|---|---|---|---|---|---|
| `client_reconnect` | `mcp.reconnect.client` | none | none | self service | class, actor, client namespace, reason, outcome | Reconnect only the caller's client transport. No daemon or peer work changes. |
| `session_reconnect` | `mcp.reconnect.session` | low | requesting-session safe point | self service | class, actor, session, reason, outcome | Rebind identity, capability, and workspace state for one session. |
| `worker_restart` | `mcp.restart.worker.request` | low | target worker | controller approval + automated gates | class, actor, worker, approval, scoped drain, outcome | Restart one worker after its own leases and mutations drain. |
| `role_runtime_restart` | `mcp.restart.role_runtime.request` | medium | target role runtime | controller approval + automated gates | class, actor, role namespace, approval, scoped drain, outcome | Restart and re-probe one role runtime; unrelated roles remain available. |
| `connector_restart` | `mcp.restart.connector.request` | medium | target connector | controller approval + automated gates | class, actor, connector, approval, scoped drain, outcome | Restart one connector while unrelated runtimes remain available. |
| `configuration_reload` | `mcp.reload.configuration.request` | low | mutation quiesce | controller approval + automated gates | class, actor, configuration revision, approval, outcome | Gracefully reload configuration without replacing the daemon. |
| `rolling_mcp_restart` | `mcp.restart.rolling.request` | medium | one instance at a time | controller approval + automated gates | class, actor, instance order, approval, per-instance drains, outcome | Drain, restart, verify, and restore each instance before advancing. |
| `full_mcp_restart` | `mcp.restart.full.request` | high | all sessions and mutations | controller approval + automated gates | class, actor, full impact, approval, full drain proof, outcome | Replace the complete MCP runtime only after a verified full drain. |
| `host_restart` | `mcp.restart.host.request` | high | all host work | controller approval + infrastructure operator | class, actor, host/change or incident id, approval, full drain proof, outcome | Hand off to infrastructure ownership and reconcile every runtime afterward. |
## Drain boundary
Only `full_mcp_restart` and `host_restart` set `full_drain_required=true`.
Reconnects and configuration reloads do not disrupt peer sessions. Worker,
role-runtime, and connector restarts evaluate only their explicitly named
target. Rolling restart drains one instance at a time. Missing required target
scope denies the request rather than silently widening it to a full restart.
## Permission and approval boundary
Author, reviewer, merger, and reconciler roles may self-request reconnects and
request scoped worker/role/connector/reload recovery. They cannot request
rolling, full, or host restart classes. Controller/operator/admin roles may
request the broader classes, while execution remains operator/admin-owned.
Controller approval is independently required for every class above a session
reconnect. Host restart additionally requires infrastructure-operator proof.
The MCP request tool derives class permissions from its authenticated runtime
role. It does not accept caller-supplied permissions. Controller and operator
authorization are read from the already-running daemon environment, never
from a request argument.
## Audit and failure behavior
Every impact audit and every console restart/reload audit includes a
`restart_class` field. The impact audit also includes the exact
`required_permission`. Unknown classes, missing permissions, ineligible roles,
missing approval, missing scoped targets, and incomplete inventory all deny
fail closed. Manual process kills remain forbidden and contaminating (#630).