# Sanctioned restart and graceful reload controls (#642) Sessions used to recover MCP connectivity by killing the host daemon (`pkill -f mcp_server.py`, #630). That is forbidden and stays forbidden: it kills every namespace on the host, contaminates whichever session survives, and leaves no audit trail. This document describes the sanctioned replacement, implemented in `webui/sanctioned_restart.py`. ## What the console will and will not do The console **never** restarts anything. It authorizes an intent, records it, and hands off to a host supervisor. There is no code path in which the console sends a signal, spawns a process, or renders a kill command — a regression test asserts the module contains no `subprocess`, `signal`, `os.kill`, `os.system`, or `popen` reference, and that no returned payload contains a kill command. ## Operations | Mode | Action | Minimum role | Behaviour | |------|--------|--------------|-----------| | `reload` | `system.reload_namespace` | controller | Host supervisor reloads the namespace in place, draining in-flight requests. | | `restart` | `system.restart_namespace` | admin | Host supervisor restarts the namespace. In-flight requests are lost. | Scope is always exactly one namespace. A fleet-wide restart is an explicit non-goal: `all`, `*`, `fleet`, and an empty scope are refused with `fleet_scope_not_permitted`, because that is precisely the blast radius the forbidden kill already had. An unrecognised namespace is refused rather than passed through to the host. ## The gate sequence `assess_restart_request()` applies every gate in order and reports the first failure with a stable reason code: | Order | Gate | Reason code on failure | |-------|------|------------------------| | 1 | Mode is `restart` or `reload` | `unknown_mode` | | 2 | Scope is a single known namespace | `fleet_scope_not_permitted`, `unknown_namespace` | | 3 | Principal holds the required console role | `unauthorized` | | 4 | Confirmation phrase supplied | `confirmation_required` | | 5 | Confirmation names this namespace and mode | `confirmation_mismatch` | | 6 | Out-of-band operator authorization present | `operator_authorization_missing` | | 7 | Runtime is not contaminated | `contaminated_runtime` | | 8 | Host restart hook configured | `restart_hook_not_configured` | Passing every gate yields `host_action_required`, never "restarted". ### Confirmation binds the namespace The required phrase is `" "` — for example `restart gitea-author`. Binding the namespace into the phrase is the point: a confirmation typed for one namespace cannot be replayed against another. ### Operator authorization is not self-assertable Host daemon maintenance is authorized out of band through `GITEA_OPERATOR_DAEMON_MAINTENANCE_AUTHORIZATION`, read from the process environment and nowhere else (#630; #710 finding F1). A worker session cannot set an environment variable for an already-running daemon, so this cannot be faked the way a tool argument could. ### The host hook `GITEA_SANCTIONED_RESTART_HOOK` holds an opaque reference the *host* resolves — a supervisor label such as a launchd job name, never a command line. With no hook configured the request is refused; the console does not fall back to a process kill. The value is read server-side and never rendered to a client. ## Manual kill remains contamination `classify_restart_command()` classifies an operator-proposed recovery command. A manual `pkill`/`kill`/`killall` of the MCP daemon is contamination, not a restart: it returns `clean_claim_allowed: false` and builds a durable contamination marker (redacted command only, never secrets) naming `system.restart_namespace` as the sanctioned alternative. A live, uncleared contamination marker also blocks a restart. This is stricter than #630's task-scoped gate, which deliberately lets a contaminated worker keep commenting and handing off: restarting a contaminated runtime would launder the contamination rather than resolve it. Clear the marker through the reconciler path first. ## Post-restart health verification After the host supervisor acts, `verify_post_restart_health()` decides whether the session may claim to be clean: | Status | Meaning | Clean claim | |--------|---------|-------------| | `clean` | Required tool callable, proven through the live client namespace | Allowed | | `unproven` | Reported healthy without live client-namespace evidence | Refused | | `unhealthy` | Probe failed | Refused | Only `probe_source=client_namespace` evidence clears a session. Static tool registration is not proof, and neither is an offline subprocess probe — an IDE client can hold a registered tool list while live calls fail with `client is closing: EOF` (see [`mcp-namespace-health.md`](mcp-namespace-health.md)). ## Audit Every attempt — allowed or denied — is recorded through `webui.console_audit` with actor, target namespace, mode, result, and reason code, and is redacted before it is persisted. `system.restart_namespace` is break-glass, so its records are retained for 730 days. Records carry `process_kill_executed: false`, which is a fact about the code path rather than a claim: no such path exists. ## Environment variables | Variable | Purpose | |----------|---------| | `GITEA_SANCTIONED_RESTART_HOOK` | Host supervisor reference; absent means restart is refused. | | `GITEA_OPERATOR_DAEMON_MAINTENANCE_AUTHORIZATION` | Out-of-band operator authorization reference. | | `WEBUI_AUDIT_LOG` | Console audit sink; absent means records are built but not persisted. | ## Non-goals * No unrestricted `kill` from the UI, in any role, in any phase. * No fleet-wide restart. * No silent auto-restart loop: every attempt is confirmed and audited. * This does not implement the Phase 1 health API (#634).