Compare commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
903536d3c0 | ||
|
|
87f7b5385d | ||
|
|
261f0b82de | ||
|
|
103d0df289 | ||
|
|
ef1ad41678 | ||
|
|
205207abb0 | ||
|
|
a87a7d1da2 | ||
|
|
00c67067f1 | ||
|
|
18bca47977 | ||
|
|
2a5d6571ec | ||
|
|
37c3e5dc39 | ||
|
|
d2eaca4949 | ||
|
|
4ff4d2acc9 | ||
|
|
fc8fe329d2 | ||
|
|
e593444eea | ||
|
|
996e7094fe | ||
|
|
ab33337a94 |
@@ -0,0 +1,122 @@
|
||||
# Sanctioned restart and graceful reload controls (#642)
|
||||
|
||||
Sessions used to recover MCP connectivity by killing the host daemon
|
||||
(`pkill -f mcp_server.py`, #630). That is forbidden and stays forbidden: it
|
||||
kills every namespace on the host, contaminates whichever session survives, and
|
||||
leaves no audit trail. This document describes the sanctioned replacement,
|
||||
implemented in `webui/sanctioned_restart.py`.
|
||||
|
||||
## What the console will and will not do
|
||||
|
||||
The console **never** restarts anything. It authorizes an intent, records it,
|
||||
and hands off to a host supervisor. There is no code path in which the console
|
||||
sends a signal, spawns a process, or renders a kill command — a regression test
|
||||
asserts the module contains no `subprocess`, `signal`, `os.kill`, `os.system`,
|
||||
or `popen` reference, and that no returned payload contains a kill command.
|
||||
|
||||
## Operations
|
||||
|
||||
| Mode | Action | Minimum role | Behaviour |
|
||||
|------|--------|--------------|-----------|
|
||||
| `reload` | `system.reload_namespace` | controller | Host supervisor reloads the namespace in place, draining in-flight requests. |
|
||||
| `restart` | `system.restart_namespace` | admin | Host supervisor restarts the namespace. In-flight requests are lost. |
|
||||
|
||||
Scope is always exactly one namespace. A fleet-wide restart is an explicit
|
||||
non-goal: `all`, `*`, `fleet`, and an empty scope are refused with
|
||||
`fleet_scope_not_permitted`, because that is precisely the blast radius the
|
||||
forbidden kill already had. An unrecognised namespace is refused rather than
|
||||
passed through to the host.
|
||||
|
||||
## The gate sequence
|
||||
|
||||
`assess_restart_request()` applies every gate in order and reports the first
|
||||
failure with a stable reason code:
|
||||
|
||||
| Order | Gate | Reason code on failure |
|
||||
|-------|------|------------------------|
|
||||
| 1 | Mode is `restart` or `reload` | `unknown_mode` |
|
||||
| 2 | Scope is a single known namespace | `fleet_scope_not_permitted`, `unknown_namespace` |
|
||||
| 3 | Principal holds the required console role | `unauthorized` |
|
||||
| 4 | Confirmation phrase supplied | `confirmation_required` |
|
||||
| 5 | Confirmation names this namespace and mode | `confirmation_mismatch` |
|
||||
| 6 | Out-of-band operator authorization present | `operator_authorization_missing` |
|
||||
| 7 | Runtime is not contaminated | `contaminated_runtime` |
|
||||
| 8 | Host restart hook configured | `restart_hook_not_configured` |
|
||||
|
||||
Passing every gate yields `host_action_required`, never "restarted".
|
||||
|
||||
### Confirmation binds the namespace
|
||||
|
||||
The required phrase is `"<mode> <namespace>"` — for example
|
||||
`restart gitea-author`. Binding the namespace into the phrase is the point: a
|
||||
confirmation typed for one namespace cannot be replayed against another.
|
||||
|
||||
### Operator authorization is not self-assertable
|
||||
|
||||
Host daemon maintenance is authorized out of band through
|
||||
`GITEA_OPERATOR_DAEMON_MAINTENANCE_AUTHORIZATION`, read from the process
|
||||
environment and nowhere else (#630; #710 finding F1). A worker session cannot
|
||||
set an environment variable for an already-running daemon, so this cannot be
|
||||
faked the way a tool argument could.
|
||||
|
||||
### The host hook
|
||||
|
||||
`GITEA_SANCTIONED_RESTART_HOOK` holds an opaque reference the *host* resolves —
|
||||
a supervisor label such as a launchd job name, never a command line. With no
|
||||
hook configured the request is refused; the console does not fall back to a
|
||||
process kill. The value is read server-side and never rendered to a client.
|
||||
|
||||
## Manual kill remains contamination
|
||||
|
||||
`classify_restart_command()` classifies an operator-proposed recovery command.
|
||||
A manual `pkill`/`kill`/`killall` of the MCP daemon is contamination, not a
|
||||
restart: it returns `clean_claim_allowed: false` and builds a durable
|
||||
contamination marker (redacted command only, never secrets) naming
|
||||
`system.restart_namespace` as the sanctioned alternative.
|
||||
|
||||
A live, uncleared contamination marker also blocks a restart. This is stricter
|
||||
than #630's task-scoped gate, which deliberately lets a contaminated worker keep
|
||||
commenting and handing off: restarting a contaminated runtime would launder the
|
||||
contamination rather than resolve it. Clear the marker through the reconciler
|
||||
path first.
|
||||
|
||||
## Post-restart health verification
|
||||
|
||||
After the host supervisor acts, `verify_post_restart_health()` decides whether
|
||||
the session may claim to be clean:
|
||||
|
||||
| Status | Meaning | Clean claim |
|
||||
|--------|---------|-------------|
|
||||
| `clean` | Required tool callable, proven through the live client namespace | Allowed |
|
||||
| `unproven` | Reported healthy without live client-namespace evidence | Refused |
|
||||
| `unhealthy` | Probe failed | Refused |
|
||||
|
||||
Only `probe_source=client_namespace` evidence clears a session. Static tool
|
||||
registration is not proof, and neither is an offline subprocess probe — an IDE
|
||||
client can hold a registered tool list while live calls fail with
|
||||
`client is closing: EOF` (see
|
||||
[`mcp-namespace-health.md`](mcp-namespace-health.md)).
|
||||
|
||||
## Audit
|
||||
|
||||
Every attempt — allowed or denied — is recorded through
|
||||
`webui.console_audit` with actor, target namespace, mode, result, and reason
|
||||
code, and is redacted before it is persisted. `system.restart_namespace` is
|
||||
break-glass, so its records are retained for 730 days. Records carry
|
||||
`process_kill_executed: false`, which is a fact about the code path rather than
|
||||
a claim: no such path exists.
|
||||
|
||||
## Environment variables
|
||||
|
||||
| Variable | Purpose |
|
||||
|----------|---------|
|
||||
| `GITEA_SANCTIONED_RESTART_HOOK` | Host supervisor reference; absent means restart is refused. |
|
||||
| `GITEA_OPERATOR_DAEMON_MAINTENANCE_AUTHORIZATION` | Out-of-band operator authorization reference. |
|
||||
| `WEBUI_AUDIT_LOG` | Console audit sink; absent means records are built but not persisted. |
|
||||
|
||||
## Non-goals
|
||||
|
||||
* No unrestricted `kill` from the UI, in any role, in any phase.
|
||||
* No fleet-wide restart.
|
||||
* No silent auto-restart loop: every attempt is confirmed and audited.
|
||||
* This does not implement the Phase 1 health API (#634).
|
||||
@@ -91,6 +91,8 @@ already define, and a regression test asserts each mapping matches.
|
||||
| `close_pr` | controller | privileged | `gitea.pr.close` | Yes | No | No | 3 |
|
||||
| `merge_pr` | controller | privileged | `gitea.pr.merge` | Yes | **Yes** | **Yes** | 3 |
|
||||
| `delete_branch` | admin | destructive | `gitea.branch.delete` | Yes | **Yes** | **Yes** | 3 |
|
||||
| `system.reload_namespace` | controller | privileged | `runtime.reload_namespace` | Yes | No | No | 2 |
|
||||
| `system.restart_namespace` | admin | destructive | `runtime.restart_namespace` | Yes | **Yes** | **Yes** | 2 |
|
||||
|
||||
**Dual control** means the acting principal may not be the sole authority: a
|
||||
second distinct principal must confirm. **Break-glass** means the action is
|
||||
@@ -102,6 +104,13 @@ honouring it.
|
||||
`delete_branch` is admin-only rather than controller because it is the one
|
||||
irreversible action in the set.
|
||||
|
||||
`system.restart_namespace` is admin-only for the same reason: restarting a
|
||||
namespace drops every in-flight request on it. `system.reload_namespace` drains
|
||||
first, so it is privileged but not destructive. Neither action is ever executed
|
||||
by the console — both hand off to a host supervisor, and neither exposes a raw
|
||||
process kill. See
|
||||
[`sanctioned-restart-controls.md`](sanctioned-restart-controls.md) (#642).
|
||||
|
||||
### Authorization decision
|
||||
|
||||
`authorize(action_id, principal, for_execution=False)` returns a decision
|
||||
@@ -199,8 +208,8 @@ breaking the request it describes.
|
||||
| Class | Applies to | Default |
|
||||
|-------|-----------|---------|
|
||||
| `standard` | Routine gated writes | 90 days |
|
||||
| `privileged` | `review_pr`, `close_pr`, and any unclassifiable action | 365 days |
|
||||
| `break_glass` | `merge_pr`, `delete_branch` | 730 days |
|
||||
| `privileged` | `review_pr`, `close_pr`, `system.reload_namespace`, and any unclassifiable action | 365 days |
|
||||
| `break_glass` | `merge_pr`, `delete_branch`, `system.restart_namespace` | 730 days |
|
||||
|
||||
Each record carries its own class, day count, and computed `expires_at`, so
|
||||
retention is auditable per record rather than inferred from file age. An
|
||||
|
||||
+20
-6
@@ -66,6 +66,8 @@ status, onboarding checklist state, and the fail-closed error payloads (#635).
|
||||
| `/api/prompts` | JSON prompt export with workflow hashes |
|
||||
| `/runtime` | MCP runtime health and stale detection (#430) |
|
||||
| `/api/runtime` | JSON runtime health export |
|
||||
| `/policy` | Workflow policy and guardrail configuration visibility (#646) |
|
||||
| `/api/v1/policy` | Versioned JSON guardrail inventory (redacted, read-only) |
|
||||
| `/audit` | Report audit paste + validator preview (#431) |
|
||||
| `/api/audit` | JSON validator preview (POST `report_text`, optional `task_kind`) |
|
||||
| `/worktrees` | Worktree hygiene dashboard (#432) |
|
||||
@@ -78,7 +80,6 @@ status, onboarding checklist state, and the fail-closed error payloads (#635).
|
||||
| `/sessions` | Phase 1 shell stub — session inventory (backed by #636) |
|
||||
| `/inventory` | Phase 1 shell stub — unified inventory (backed by #636) |
|
||||
| `/timeline` | Phase 1 shell stub — workflow event timeline |
|
||||
| `/policy` | Phase 1 shell stub — capability/role policy placeholder |
|
||||
| `/insights` | Phase 1 shell stub — operational insights placeholder |
|
||||
|
||||
Most routes are GET-only. POST/PUT/PATCH/DELETE return `405` with
|
||||
@@ -239,12 +240,25 @@ health, workflow/schema SHA-256 hashes, and stale-runtime warnings when the
|
||||
checkout is behind merged safety-gate changes. Restart guidance links to #420;
|
||||
no tokens or MCP restart actions are exposed.
|
||||
|
||||
## Policy & guardrail visibility (#646)
|
||||
|
||||
`/policy` (HTML) and `/api/v1/policy` (JSON) surface a **read-only** projection
|
||||
of the major workflow guardrails — role separation/RBAC, lease lifecycle,
|
||||
author worktree binding, merge confirmation, secret redaction, contamination
|
||||
containment, allocator policy, audit logging, and mutation gating. Each entry
|
||||
carries source pointers to the file/module/doc that owns it, a compact active
|
||||
value derived from the existing safe policy accessors, and — where a documented
|
||||
default is declared — a diff of active vs documented. The whole payload is run
|
||||
through the console redaction pass before it is emitted, so a planted or
|
||||
accidental secret degrades to the placeholder rather than reaching a client.
|
||||
The view never edits policy and exposes no gate-weakening toggle.
|
||||
|
||||
## Application shell — Phase 1 (#638)
|
||||
|
||||
The console shell (`webui/layout.py`) renders a grouped navigation driven by a
|
||||
single nav-config module, `webui/nav.py`. Nav groups follow the epic #631
|
||||
Phase 1 information architecture: **Health, Traffic, Runtime/Sessions,
|
||||
Projects, Inventory, Timeline, Policy** (placeholder), and **Insights**
|
||||
Projects, Inventory, Timeline, Policy** (live via #646), and **Insights**
|
||||
(placeholder). Live views and Phase 1 placeholders (`stub`) are declared in one
|
||||
place so the layout and the route table cannot drift.
|
||||
|
||||
@@ -254,10 +268,10 @@ a **mode: read-only** badge — plus a **Docs** link to this document. No
|
||||
privileged action controls are present in the Phase 1 shell.
|
||||
|
||||
Not-yet-implemented surfaces (`/sessions`, `/inventory`, `/timeline`,
|
||||
`/policy`, `/insights`) resolve to graceful read-only stub pages instead of
|
||||
404s; their backing views land in later child issues of #631 (the inventory
|
||||
surfaces are backed by #636). Mutating methods on stub routes still fail closed
|
||||
with `read-only-mvp`.
|
||||
`/insights`) resolve to graceful read-only stub pages instead of 404s; their
|
||||
backing views land in later child issues of #631 (the inventory surfaces are
|
||||
backed by #636). `/policy` is a live read-only surface (#646), not a stub.
|
||||
Mutating methods on stub routes still fail closed with `read-only-mvp`.
|
||||
|
||||
## System-health dashboard (#639)
|
||||
|
||||
|
||||
@@ -363,6 +363,21 @@ TASK_CAPABILITY_MAP: dict[str, dict[str, str]] = {
|
||||
"role": "controller",
|
||||
},
|
||||
|
||||
# #642: sanctioned host-daemon lifecycle controls. Deliberately *not* a
|
||||
# ``gitea.*`` operation — restarting an MCP namespace is a host action, not
|
||||
# a Gitea API call, and no configured Gitea profile should be able to
|
||||
# satisfy it by accident. Authority comes from the console RBAC model plus
|
||||
# out-of-band operator authorization (#630); these entries exist so the
|
||||
# console cannot invent an authority the capability layer never declared.
|
||||
"restart_namespace": {
|
||||
"permission": "runtime.restart_namespace",
|
||||
"role": "controller",
|
||||
},
|
||||
"reload_namespace": {
|
||||
"permission": "runtime.reload_namespace",
|
||||
"role": "controller",
|
||||
},
|
||||
|
||||
# #601 first-class lease lifecycle — inspect/list need read; mutations gate on
|
||||
# ownership in the control-plane DB (not a separate Gitea write permission).
|
||||
"list_workflow_leases": {
|
||||
|
||||
@@ -0,0 +1,243 @@
|
||||
"""Tests for the read-only workflow policy/guardrail visibility view (#646).
|
||||
|
||||
Covers issue #646 acceptance criteria:
|
||||
|
||||
1. Console lists major guardrails with source pointers.
|
||||
2. Secrets redacted.
|
||||
3. Tests ensure sample secrets never appear.
|
||||
4. Docs explain read-only nature (asserted here for the page copy; the doc
|
||||
itself is covered by inspection).
|
||||
"""
|
||||
|
||||
import json
|
||||
import sys
|
||||
import unittest
|
||||
from pathlib import Path
|
||||
|
||||
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
|
||||
|
||||
from starlette.testclient import TestClient
|
||||
|
||||
from webui import console_redaction
|
||||
from webui import policy_inventory
|
||||
from webui.app import create_app
|
||||
from webui.policy_inventory import (
|
||||
PolicyEntry,
|
||||
PolicyInventorySnapshot,
|
||||
SourcePointer,
|
||||
load_policy_inventory,
|
||||
snapshot_to_dict,
|
||||
)
|
||||
from webui.policy_views import render_policy_page
|
||||
|
||||
|
||||
def _entry(key, category, *, active=None, error=None):
|
||||
return PolicyEntry(
|
||||
key=key,
|
||||
title=key.replace("_", " ").title(),
|
||||
category=category,
|
||||
summary=f"summary for {key}",
|
||||
sources=(SourcePointer("src", f"{key}.py", "module"),),
|
||||
active=active,
|
||||
documented_default=None,
|
||||
diff=None,
|
||||
error=error,
|
||||
)
|
||||
|
||||
|
||||
def _snapshot(entries):
|
||||
return PolicyInventorySnapshot(
|
||||
schema_version=1,
|
||||
read_only=True,
|
||||
note="read-only projection",
|
||||
entries=tuple(entries),
|
||||
categories=tuple(dict.fromkeys(e.category for e in entries)),
|
||||
build_errors=(),
|
||||
)
|
||||
|
||||
# The guardrail categories issue #646 names as in-scope.
|
||||
_EXPECTED_CATEGORIES = {
|
||||
"role_separation",
|
||||
"lease_rules",
|
||||
"worktree_rules",
|
||||
"merge_confirmation",
|
||||
"redaction",
|
||||
"contamination",
|
||||
"allocator_policy",
|
||||
"audit_logging",
|
||||
"mutation_gating",
|
||||
}
|
||||
|
||||
|
||||
class TestPolicyInventoryModel(unittest.TestCase):
|
||||
def test_major_guardrails_present(self):
|
||||
snapshot = load_policy_inventory()
|
||||
categories = {e.category for e in snapshot.entries}
|
||||
self.assertEqual(_EXPECTED_CATEGORIES, categories)
|
||||
self.assertGreaterEqual(len(snapshot.entries), len(_EXPECTED_CATEGORIES))
|
||||
|
||||
def test_every_guardrail_has_source_pointers(self):
|
||||
# AC1: source attribution (file/module/doc) for every guardrail.
|
||||
snapshot = load_policy_inventory()
|
||||
for entry in snapshot.entries:
|
||||
with self.subTest(entry=entry.key):
|
||||
self.assertTrue(entry.sources, "guardrail must carry source pointers")
|
||||
for source in entry.sources:
|
||||
self.assertTrue(source.path)
|
||||
self.assertIn(source.kind, {"module", "doc", "script", "config"})
|
||||
|
||||
def test_diff_reported_where_documented_default_declared(self):
|
||||
snapshot = load_policy_inventory()
|
||||
checked_any = False
|
||||
for entry in snapshot.entries:
|
||||
if entry.documented_default is None:
|
||||
self.assertIsNone(entry.diff)
|
||||
continue
|
||||
checked_any = True
|
||||
self.assertIsNotNone(entry.diff)
|
||||
self.assertEqual(
|
||||
entry.diff["status"],
|
||||
"matches_documented_default",
|
||||
f"{entry.key} drifted from its documented default: {entry.diff}",
|
||||
)
|
||||
self.assertTrue(checked_any, "at least one guardrail should declare a default")
|
||||
|
||||
def test_live_projections_populate_active(self):
|
||||
snapshot = load_policy_inventory()
|
||||
by_key = {e.key: e for e in snapshot.entries}
|
||||
for key in ("role_separation", "redaction", "audit_logging"):
|
||||
self.assertIsNone(by_key[key].error, f"{key} projection failed")
|
||||
self.assertIsInstance(by_key[key].active, dict)
|
||||
|
||||
def test_build_entry_is_fail_soft_on_projection_error(self):
|
||||
def _boom():
|
||||
raise RuntimeError("projection exploded")
|
||||
|
||||
row = (
|
||||
"redaction",
|
||||
"Secret redaction",
|
||||
"redaction",
|
||||
"summary",
|
||||
(SourcePointer("x", "webui/console_redaction.py", "module"),),
|
||||
_boom,
|
||||
{"redact_before_persist": True},
|
||||
)
|
||||
entry = policy_inventory._build_entry(row)
|
||||
self.assertIsNone(entry.active)
|
||||
self.assertIsNotNone(entry.error)
|
||||
self.assertEqual(entry.diff["status"], "active_unavailable")
|
||||
|
||||
|
||||
class TestPolicyRedaction(unittest.TestCase):
|
||||
def test_real_snapshot_has_no_secret_shapes(self):
|
||||
# AC3: the real emitted payload never carries a known secret shape.
|
||||
payload = snapshot_to_dict(load_policy_inventory())
|
||||
self.assertEqual(console_redaction.scan_for_secrets(payload), [])
|
||||
|
||||
def test_planted_keychain_secret_is_redacted(self):
|
||||
# AC2/AC3: a secret planted in an active projection is masked before emit.
|
||||
snapshot = _snapshot([
|
||||
_entry(
|
||||
"redaction",
|
||||
"redaction",
|
||||
active={"leaked": "keychain:prgs-author-super-secret", "roles": ["author"]},
|
||||
)
|
||||
])
|
||||
payload = snapshot_to_dict(snapshot)
|
||||
blob = json.dumps(payload)
|
||||
self.assertNotIn("keychain:prgs-author-super-secret", blob)
|
||||
self.assertEqual(console_redaction.scan_for_secrets(payload), [])
|
||||
|
||||
def test_planted_credential_assignment_is_redacted(self):
|
||||
snapshot = _snapshot([
|
||||
_entry(
|
||||
"audit_logging",
|
||||
"audit_logging",
|
||||
active={"leaked": "token=abcd1234efgh5678", "append_only": True},
|
||||
)
|
||||
])
|
||||
payload = snapshot_to_dict(snapshot)
|
||||
blob = json.dumps(payload)
|
||||
self.assertNotIn("abcd1234efgh5678", blob)
|
||||
self.assertEqual(console_redaction.scan_for_secrets(payload), [])
|
||||
|
||||
def test_planted_secret_is_redacted_in_html_emit(self):
|
||||
# AC2/AC3: HTML emit path runs redaction before rendering HTML cards.
|
||||
snapshot = _snapshot([
|
||||
_entry(
|
||||
"redaction",
|
||||
"redaction",
|
||||
active={"leaked": "keychain:prgs-author-super-secret", "roles": ["author"]},
|
||||
)
|
||||
])
|
||||
html_output = render_policy_page(snapshot)
|
||||
self.assertNotIn("keychain:prgs-author-super-secret", html_output)
|
||||
self.assertEqual(console_redaction.scan_for_secrets(html_output), [])
|
||||
|
||||
|
||||
class TestPolicyRoutes(unittest.TestCase):
|
||||
def setUp(self):
|
||||
self.client = TestClient(create_app())
|
||||
|
||||
def test_policy_html_lists_guardrails_with_sources(self):
|
||||
response = self.client.get("/policy")
|
||||
self.assertEqual(response.status_code, 200)
|
||||
text = response.text
|
||||
self.assertIn("Workflow policy", text)
|
||||
self.assertIn("Role separation and RBAC", text)
|
||||
self.assertIn("Source pointers", text)
|
||||
self.assertIn("task_capability_map.py", text)
|
||||
self.assertIn("docs/safety-model.md", text)
|
||||
|
||||
def test_policy_html_states_read_only(self):
|
||||
# AC4: the page explains its read-only nature.
|
||||
text = self.client.get("/policy").text
|
||||
self.assertIn("read-only", text.lower())
|
||||
self.assertNotIn("<form", text.lower())
|
||||
|
||||
def test_policy_html_has_no_secret_shapes(self):
|
||||
text = self.client.get("/policy").text
|
||||
self.assertEqual(console_redaction.scan_for_secrets(text), [])
|
||||
|
||||
def test_api_v1_policy_returns_inventory(self):
|
||||
response = self.client.get("/api/v1/policy")
|
||||
self.assertEqual(response.status_code, 200)
|
||||
data = response.json()
|
||||
self.assertEqual(data["schema_version"], policy_inventory.SCHEMA_VERSION)
|
||||
self.assertTrue(data["read_only"])
|
||||
self.assertEqual(data["entry_count"], len(data["entries"]))
|
||||
self.assertEqual(set(data["categories"]), _EXPECTED_CATEGORIES)
|
||||
|
||||
def test_policy_is_read_only_no_post(self):
|
||||
# AC4 / non-goal: no mutation endpoint.
|
||||
response = self.client.post("/policy")
|
||||
self.assertIn(response.status_code, (404, 405))
|
||||
|
||||
def test_nav_links_policy(self):
|
||||
# Policy is graduated live: not in STUB_PAGES and not labeled ·stub in nav.
|
||||
from webui.nav import STUB_PAGES, iter_nav_items
|
||||
|
||||
self.assertNotIn("/policy", STUB_PAGES)
|
||||
policy_items = [i for i in iter_nav_items() if i.href == "/policy"]
|
||||
self.assertEqual(len(policy_items), 1)
|
||||
self.assertEqual(policy_items[0].status, "live")
|
||||
text = self.client.get("/").text
|
||||
self.assertIn('href="/policy">Policy</a>', text)
|
||||
self.assertNotIn('href="/policy" class="nav-stub"', text)
|
||||
|
||||
|
||||
class TestPolicyViewFailSoft(unittest.TestCase):
|
||||
def test_page_renders_when_a_projection_errors(self):
|
||||
snapshot = _snapshot([
|
||||
_entry("role_separation", "role_separation", error="active projection unavailable: boom"),
|
||||
_entry("redaction", "redaction", active={"redact_before_persist": True}),
|
||||
])
|
||||
page = render_policy_page(snapshot)
|
||||
# The errored guardrail surfaces its error; other guardrails still render.
|
||||
self.assertIn("Active value unavailable", page)
|
||||
self.assertIn("Redaction", page)
|
||||
self.assertIn("Workflow policy", page)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
unittest.main()
|
||||
@@ -0,0 +1,502 @@
|
||||
"""Sanctioned restart / graceful reload control tests (#642).
|
||||
|
||||
Acceptance criteria under test:
|
||||
|
||||
1. The sanctioned restart path is implemented behind gates (capability,
|
||||
confirmation, operator authorization, host hook).
|
||||
2. Manual ``pkill`` stays forbidden and is classified as contamination.
|
||||
3. Post-restart mutations require clean health/session proof.
|
||||
4. Authorized restart preview, unauthorized deny, contamination classification.
|
||||
5. No entry point exposes a raw kill.
|
||||
"""
|
||||
|
||||
import json
|
||||
import os
|
||||
import tempfile
|
||||
import unittest
|
||||
|
||||
import mcp_namespace_health
|
||||
import runtime_recovery_guard
|
||||
from task_capability_map import TASK_CAPABILITY_MAP
|
||||
from webui import console_audit, console_authz, gated_actions, sanctioned_restart
|
||||
|
||||
NAMESPACE = "gitea-author"
|
||||
|
||||
# An operator-authorized, hook-configured host. Passed explicitly so no test
|
||||
# depends on (or mutates) the real process environment.
|
||||
READY_ENV = {
|
||||
sanctioned_restart.RESTART_HOOK_ENV: "launchd:cc.prgs.gitea-author",
|
||||
runtime_recovery_guard.OPERATOR_AUTHORIZATION_ENV: "ops-ticket-4821",
|
||||
}
|
||||
|
||||
|
||||
def admin(subject: str = "[email protected]") -> console_authz.Principal:
|
||||
return console_authz.Principal(
|
||||
subject=subject,
|
||||
role=console_authz.ADMIN,
|
||||
identity_source=console_authz.IDENTITY_ACCESS_PROXY,
|
||||
authenticated=True,
|
||||
)
|
||||
|
||||
|
||||
def viewer() -> console_authz.Principal:
|
||||
return console_authz.Principal(
|
||||
subject="[email protected]",
|
||||
role=console_authz.VIEWER,
|
||||
identity_source=console_authz.IDENTITY_ACCESS_PROXY,
|
||||
authenticated=True,
|
||||
)
|
||||
|
||||
|
||||
class TestCapabilityWiring(unittest.TestCase):
|
||||
"""AC1: authority is declared, not invented by the console."""
|
||||
|
||||
def test_actions_resolve_through_the_capability_map(self):
|
||||
for action_id in (
|
||||
sanctioned_restart.ACTION_RESTART_NAMESPACE,
|
||||
sanctioned_restart.ACTION_RELOAD_NAMESPACE,
|
||||
):
|
||||
with self.subTest(action=action_id):
|
||||
action = console_authz.get_action(action_id)
|
||||
self.assertIsNotNone(action)
|
||||
self.assertIn(action.task_key, TASK_CAPABILITY_MAP)
|
||||
self.assertEqual(
|
||||
action.mcp_permission,
|
||||
TASK_CAPABILITY_MAP[action.task_key]["permission"],
|
||||
)
|
||||
|
||||
def test_restart_permission_is_not_a_gitea_operation(self):
|
||||
"""No configured Gitea profile should satisfy a host restart."""
|
||||
permission = TASK_CAPABILITY_MAP["restart_namespace"]["permission"]
|
||||
self.assertFalse(permission.startswith("gitea."))
|
||||
|
||||
def test_restart_is_destructive_dual_control_break_glass(self):
|
||||
action = console_authz.get_action(
|
||||
sanctioned_restart.ACTION_RESTART_NAMESPACE
|
||||
)
|
||||
self.assertEqual(action.action_class, console_authz.CLASS_DESTRUCTIVE)
|
||||
self.assertEqual(action.minimum_role, console_authz.ADMIN)
|
||||
self.assertTrue(action.dual_control)
|
||||
self.assertTrue(action.break_glass)
|
||||
self.assertTrue(action.requires_confirmation)
|
||||
|
||||
def test_reload_is_privileged_but_not_destructive(self):
|
||||
action = console_authz.get_action(
|
||||
sanctioned_restart.ACTION_RELOAD_NAMESPACE
|
||||
)
|
||||
self.assertEqual(action.action_class, console_authz.CLASS_PRIVILEGED)
|
||||
self.assertTrue(action.requires_confirmation)
|
||||
|
||||
|
||||
class TestPreview(unittest.TestCase):
|
||||
"""AC4: an authorized preview renders the plan without executing it."""
|
||||
|
||||
def test_preview_lists_the_mutation_ledger(self):
|
||||
preview = sanctioned_restart.build_restart_preview(
|
||||
NAMESPACE, principal=admin(), env=READY_ENV
|
||||
)
|
||||
steps = [entry["step"] for entry in preview["mutation_ledger"]]
|
||||
self.assertEqual(
|
||||
steps, ["quiesce", "host_restart_hook", "health_recheck", "audit"]
|
||||
)
|
||||
self.assertTrue(preview["scope_valid"])
|
||||
self.assertTrue(preview["post_restart_verification_required"])
|
||||
|
||||
def test_reload_preview_drains_instead_of_restarting(self):
|
||||
preview = sanctioned_restart.build_restart_preview(
|
||||
NAMESPACE, sanctioned_restart.MODE_RELOAD,
|
||||
principal=admin(), env=READY_ENV,
|
||||
)
|
||||
steps = [entry["step"] for entry in preview["mutation_ledger"]]
|
||||
self.assertIn("host_graceful_reload", steps)
|
||||
self.assertNotIn("host_restart_hook", steps)
|
||||
|
||||
def test_preview_never_enables_execution(self):
|
||||
preview = sanctioned_restart.build_restart_preview(
|
||||
NAMESPACE, principal=admin(), env=READY_ENV
|
||||
)
|
||||
self.assertFalse(preview["execution_enabled"])
|
||||
self.assertFalse(preview["authorization"]["execution_enabled"])
|
||||
|
||||
def test_confirmation_phrase_binds_the_namespace(self):
|
||||
self.assertTrue(
|
||||
sanctioned_restart.confirmation_matches(
|
||||
NAMESPACE, sanctioned_restart.MODE_RESTART,
|
||||
"restart gitea-author",
|
||||
)
|
||||
)
|
||||
# A phrase typed for one namespace must not authorize another.
|
||||
self.assertFalse(
|
||||
sanctioned_restart.confirmation_matches(
|
||||
"gitea-merger", sanctioned_restart.MODE_RESTART,
|
||||
"restart gitea-author",
|
||||
)
|
||||
)
|
||||
|
||||
|
||||
class TestGates(unittest.TestCase):
|
||||
"""AC1/AC4: every gate denies with a stable reason code."""
|
||||
|
||||
def _assess(self, **kwargs):
|
||||
params = {
|
||||
"principal": admin(),
|
||||
"confirmation": f"restart {NAMESPACE}",
|
||||
"env": READY_ENV,
|
||||
}
|
||||
params.update(kwargs)
|
||||
namespace = params.pop("namespace", NAMESPACE)
|
||||
mode = params.pop("mode", sanctioned_restart.MODE_RESTART)
|
||||
return sanctioned_restart.assess_restart_request(
|
||||
namespace, mode, **params
|
||||
)
|
||||
|
||||
def test_authorized_confirmed_request_passes_every_gate(self):
|
||||
result = self._assess()
|
||||
self.assertTrue(result["allowed"])
|
||||
self.assertEqual(
|
||||
result["reason_code"], sanctioned_restart.ALLOW_HOST_ACTION_REQUIRED
|
||||
)
|
||||
|
||||
def test_passing_every_gate_is_not_an_execution_grant(self):
|
||||
"""An allowed request still never lets the console touch the process."""
|
||||
result = self._assess()
|
||||
self.assertTrue(result["allowed"])
|
||||
self.assertFalse(result["execution_enabled"])
|
||||
self.assertFalse(result["console_executes"])
|
||||
|
||||
def test_unauthorized_principal_is_denied(self):
|
||||
result = self._assess(principal=viewer())
|
||||
self.assertFalse(result["allowed"])
|
||||
self.assertEqual(
|
||||
result["reason_code"], sanctioned_restart.DENY_UNAUTHORIZED
|
||||
)
|
||||
|
||||
def test_anonymous_principal_is_denied(self):
|
||||
result = self._assess(principal=None)
|
||||
self.assertFalse(result["allowed"])
|
||||
self.assertEqual(
|
||||
result["reason_code"], sanctioned_restart.DENY_UNAUTHORIZED
|
||||
)
|
||||
|
||||
def test_missing_confirmation_is_denied(self):
|
||||
result = self._assess(confirmation=None)
|
||||
self.assertFalse(result["allowed"])
|
||||
self.assertEqual(
|
||||
result["reason_code"], sanctioned_restart.DENY_CONFIRMATION_MISSING
|
||||
)
|
||||
|
||||
def test_confirmation_for_another_namespace_is_denied(self):
|
||||
result = self._assess(confirmation="restart gitea-merger")
|
||||
self.assertFalse(result["allowed"])
|
||||
self.assertEqual(
|
||||
result["reason_code"], sanctioned_restart.DENY_CONFIRMATION_MISMATCH
|
||||
)
|
||||
|
||||
def test_missing_operator_authorization_is_denied(self):
|
||||
env = {sanctioned_restart.RESTART_HOOK_ENV: "launchd:cc.prgs.author"}
|
||||
result = self._assess(env=env)
|
||||
self.assertFalse(result["allowed"])
|
||||
self.assertEqual(
|
||||
result["reason_code"],
|
||||
sanctioned_restart.DENY_OPERATOR_AUTHORIZATION,
|
||||
)
|
||||
|
||||
def test_missing_host_hook_is_denied_without_kill_fallback(self):
|
||||
env = {
|
||||
runtime_recovery_guard.OPERATOR_AUTHORIZATION_ENV: "ops-ticket-1",
|
||||
}
|
||||
result = self._assess(env=env)
|
||||
self.assertFalse(result["allowed"])
|
||||
self.assertEqual(
|
||||
result["reason_code"], sanctioned_restart.DENY_HOOK_NOT_CONFIGURED
|
||||
)
|
||||
|
||||
def test_fleet_scope_is_refused(self):
|
||||
for scope in ("all", "*", "fleet"):
|
||||
with self.subTest(scope=scope):
|
||||
result = self._assess(
|
||||
namespace=scope, confirmation=f"restart {scope}"
|
||||
)
|
||||
self.assertFalse(result["allowed"])
|
||||
self.assertEqual(
|
||||
result["reason_code"], sanctioned_restart.DENY_FLEET_SCOPE
|
||||
)
|
||||
|
||||
def test_unknown_namespace_is_refused(self):
|
||||
result = self._assess(
|
||||
namespace="gitea-nope", confirmation="restart gitea-nope"
|
||||
)
|
||||
self.assertFalse(result["allowed"])
|
||||
self.assertEqual(
|
||||
result["reason_code"], sanctioned_restart.DENY_UNKNOWN_NAMESPACE
|
||||
)
|
||||
|
||||
def test_unknown_mode_is_refused(self):
|
||||
result = self._assess(mode="obliterate")
|
||||
self.assertFalse(result["allowed"])
|
||||
self.assertEqual(
|
||||
result["reason_code"], sanctioned_restart.DENY_UNKNOWN_MODE
|
||||
)
|
||||
|
||||
def test_live_contamination_marker_blocks_restart(self):
|
||||
marker = runtime_recovery_guard.build_contamination_record(
|
||||
reason_class=runtime_recovery_guard.REASON_MANUAL_DAEMON_KILL,
|
||||
command_redacted="pkill -f mcp_server.py",
|
||||
)
|
||||
result = self._assess(contamination_marker=marker)
|
||||
self.assertFalse(result["allowed"])
|
||||
self.assertEqual(
|
||||
result["reason_code"], sanctioned_restart.DENY_CONTAMINATED_RUNTIME
|
||||
)
|
||||
|
||||
def test_reconciler_cleared_marker_no_longer_blocks(self):
|
||||
marker = runtime_recovery_guard.build_contamination_record(
|
||||
reason_class=runtime_recovery_guard.REASON_MANUAL_DAEMON_KILL,
|
||||
command_redacted="pkill -f mcp_server.py",
|
||||
)
|
||||
marker = dict(marker, cleared_by_reconciler=True)
|
||||
result = self._assess(contamination_marker=marker)
|
||||
self.assertTrue(result["allowed"])
|
||||
|
||||
|
||||
class TestExecutionNeverKills(unittest.TestCase):
|
||||
"""AC5: no path exposes or runs a raw process kill."""
|
||||
|
||||
def test_authorized_execution_defers_to_the_host_supervisor(self):
|
||||
result = sanctioned_restart.execute_restart(
|
||||
NAMESPACE,
|
||||
principal=admin(),
|
||||
confirmation=f"restart {NAMESPACE}",
|
||||
env=READY_ENV,
|
||||
)
|
||||
self.assertTrue(result["allowed"])
|
||||
self.assertFalse(result["success"])
|
||||
self.assertFalse(result["process_kill_executed"])
|
||||
self.assertEqual(
|
||||
result["outcome"], sanctioned_restart.ALLOW_HOST_ACTION_REQUIRED
|
||||
)
|
||||
|
||||
def test_denied_execution_reports_the_refusing_gate(self):
|
||||
result = sanctioned_restart.execute_restart(
|
||||
NAMESPACE, principal=viewer(), confirmation=f"restart {NAMESPACE}",
|
||||
env=READY_ENV,
|
||||
)
|
||||
self.assertFalse(result["allowed"])
|
||||
self.assertEqual(
|
||||
result["outcome"], sanctioned_restart.DENY_UNAUTHORIZED
|
||||
)
|
||||
self.assertFalse(result["process_kill_executed"])
|
||||
|
||||
def test_module_never_spawns_a_process(self):
|
||||
path = os.path.join(
|
||||
os.path.dirname(os.path.dirname(os.path.abspath(__file__))),
|
||||
"webui", "sanctioned_restart.py",
|
||||
)
|
||||
with open(path, encoding="utf-8") as handle:
|
||||
source = handle.read()
|
||||
for forbidden in (
|
||||
"import subprocess", "import signal", "os.kill", "os.system",
|
||||
"popen",
|
||||
):
|
||||
with self.subTest(forbidden=forbidden):
|
||||
self.assertNotIn(forbidden, source.lower())
|
||||
|
||||
def test_no_surface_returns_a_kill_command(self):
|
||||
payloads = [
|
||||
sanctioned_restart.build_restart_preview(
|
||||
NAMESPACE, principal=admin(), env=READY_ENV
|
||||
),
|
||||
sanctioned_restart.restart_policy(),
|
||||
sanctioned_restart.execute_restart(
|
||||
NAMESPACE, principal=admin(),
|
||||
confirmation=f"restart {NAMESPACE}", env=READY_ENV,
|
||||
),
|
||||
]
|
||||
for payload in payloads:
|
||||
rendered = json.dumps(payload, default=str).lower()
|
||||
self.assertNotIn("kill -9", rendered)
|
||||
self.assertNotIn("pkill -f", rendered)
|
||||
|
||||
def test_policy_declares_no_raw_kill_and_no_silent_restart(self):
|
||||
policy = sanctioned_restart.restart_policy()
|
||||
self.assertFalse(policy["raw_kill_exposed"])
|
||||
self.assertFalse(policy["console_executes_process_kill"])
|
||||
self.assertFalse(policy["fleet_scope_permitted"])
|
||||
self.assertFalse(policy["silent_auto_restart_permitted"])
|
||||
self.assertTrue(policy["audit_required"])
|
||||
|
||||
|
||||
class TestContaminationClassification(unittest.TestCase):
|
||||
"""AC2: manual pkill is contamination, and it blocks clean claims."""
|
||||
|
||||
def test_manual_daemon_pkill_is_contamination(self):
|
||||
result = sanctioned_restart.classify_restart_command(
|
||||
"pkill -f mcp_server.py"
|
||||
)
|
||||
self.assertTrue(result["contamination"])
|
||||
self.assertFalse(result["clean_claim_allowed"])
|
||||
self.assertIsNotNone(result["contamination_marker"])
|
||||
self.assertEqual(
|
||||
result["sanctioned_alternative"],
|
||||
sanctioned_restart.ACTION_RESTART_NAMESPACE,
|
||||
)
|
||||
|
||||
def test_broad_process_kill_is_contamination(self):
|
||||
result = sanctioned_restart.classify_restart_command("killall -9 Python")
|
||||
self.assertTrue(result["contamination"])
|
||||
self.assertFalse(result["clean_claim_allowed"])
|
||||
|
||||
def test_marker_names_the_sanctioned_alternative(self):
|
||||
result = sanctioned_restart.classify_restart_command(
|
||||
"pkill -f mcp_server.py"
|
||||
)
|
||||
marker = result["contamination_marker"]
|
||||
self.assertIn(
|
||||
sanctioned_restart.ACTION_RESTART_NAMESPACE, marker["detail"]
|
||||
)
|
||||
|
||||
def test_benign_command_is_not_contamination(self):
|
||||
result = sanctioned_restart.classify_restart_command("git status")
|
||||
self.assertFalse(result["contamination"])
|
||||
self.assertTrue(result["clean_claim_allowed"])
|
||||
|
||||
def test_no_command_is_not_contamination(self):
|
||||
result = sanctioned_restart.classify_restart_command(None)
|
||||
self.assertFalse(result["contamination"])
|
||||
self.assertTrue(result["clean_claim_allowed"])
|
||||
|
||||
|
||||
class TestPostRestartHealth(unittest.TestCase):
|
||||
"""AC3: a clean post-restart claim needs live client-namespace proof."""
|
||||
|
||||
def test_live_client_probe_clears_the_session(self):
|
||||
result = sanctioned_restart.verify_post_restart_health(
|
||||
NAMESPACE,
|
||||
probe_result={"success": True},
|
||||
probe_source=mcp_namespace_health.PROBE_SOURCE_CLIENT,
|
||||
registered_tools=["gitea_whoami"],
|
||||
required_tool="gitea_whoami",
|
||||
)
|
||||
self.assertEqual(result["status"], sanctioned_restart.HEALTH_CLEAN)
|
||||
self.assertTrue(result["clean_claim_allowed"])
|
||||
self.assertTrue(result["mutations_allowed"])
|
||||
|
||||
def test_offline_probe_does_not_clear_the_session(self):
|
||||
result = sanctioned_restart.verify_post_restart_health(
|
||||
NAMESPACE,
|
||||
probe_result={"success": True},
|
||||
probe_source=mcp_namespace_health.PROBE_SOURCE_OFFLINE,
|
||||
registered_tools=["gitea_whoami"],
|
||||
required_tool="gitea_whoami",
|
||||
)
|
||||
self.assertFalse(result["clean_claim_allowed"])
|
||||
self.assertFalse(result["mutations_allowed"])
|
||||
|
||||
def test_failed_probe_is_unhealthy(self):
|
||||
result = sanctioned_restart.verify_post_restart_health(
|
||||
NAMESPACE,
|
||||
probe_result={"success": False, "error": "client is closing: EOF"},
|
||||
probe_source=mcp_namespace_health.PROBE_SOURCE_CLIENT,
|
||||
registered_tools=["gitea_whoami"],
|
||||
required_tool="gitea_whoami",
|
||||
)
|
||||
self.assertEqual(result["status"], sanctioned_restart.HEALTH_UNHEALTHY)
|
||||
self.assertFalse(result["clean_claim_allowed"])
|
||||
|
||||
def test_static_registration_alone_never_clears_the_session(self):
|
||||
result = sanctioned_restart.verify_post_restart_health(
|
||||
NAMESPACE,
|
||||
registered_tools=["gitea_whoami"],
|
||||
required_tool="gitea_whoami",
|
||||
)
|
||||
self.assertFalse(result["clean_claim_allowed"])
|
||||
|
||||
|
||||
class TestAuditEmission(unittest.TestCase):
|
||||
"""Every restart attempt is audited with actor, target, and result."""
|
||||
|
||||
def _run(self, principal, sink):
|
||||
prior = os.environ.get(console_audit.AUDIT_LOG_ENV)
|
||||
os.environ[console_audit.AUDIT_LOG_ENV] = sink
|
||||
try:
|
||||
return sanctioned_restart.execute_restart(
|
||||
NAMESPACE,
|
||||
principal=principal,
|
||||
confirmation=f"restart {NAMESPACE}",
|
||||
env=READY_ENV,
|
||||
request_id="req-642",
|
||||
)
|
||||
finally:
|
||||
if prior is None:
|
||||
os.environ.pop(console_audit.AUDIT_LOG_ENV, None)
|
||||
else:
|
||||
os.environ[console_audit.AUDIT_LOG_ENV] = prior
|
||||
|
||||
def test_allowed_attempt_is_written_with_actor_and_target(self):
|
||||
with tempfile.TemporaryDirectory() as tmp:
|
||||
sink = os.path.join(tmp, "audit.jsonl")
|
||||
result = self._run(admin(), sink)
|
||||
self.assertTrue(result["audit"]["written"])
|
||||
with open(sink, encoding="utf-8") as handle:
|
||||
record = json.loads(handle.read().strip())
|
||||
self.assertEqual(
|
||||
record["action"], sanctioned_restart.ACTION_RESTART_NAMESPACE
|
||||
)
|
||||
self.assertEqual(record["target"]["namespace"], NAMESPACE)
|
||||
self.assertEqual(record["target"]["mode"], "restart")
|
||||
self.assertEqual(record["result"], console_audit.RESULT_ALLOWED)
|
||||
self.assertEqual(record["actor"]["subject"], "[email protected]")
|
||||
self.assertFalse(record["metadata"]["process_kill_executed"])
|
||||
|
||||
def test_denied_attempt_is_audited_too(self):
|
||||
with tempfile.TemporaryDirectory() as tmp:
|
||||
sink = os.path.join(tmp, "audit.jsonl")
|
||||
self._run(viewer(), sink)
|
||||
with open(sink, encoding="utf-8") as handle:
|
||||
record = json.loads(handle.read().strip())
|
||||
self.assertEqual(record["result"], console_audit.RESULT_DENIED)
|
||||
self.assertEqual(
|
||||
record["reason_code"], sanctioned_restart.DENY_UNAUTHORIZED
|
||||
)
|
||||
|
||||
def test_restart_audit_uses_break_glass_retention(self):
|
||||
action = console_authz.get_action(
|
||||
sanctioned_restart.ACTION_RESTART_NAMESPACE
|
||||
)
|
||||
self.assertEqual(
|
||||
console_audit.retention_class_for(action),
|
||||
console_audit.RETENTION_BREAK_GLASS,
|
||||
)
|
||||
|
||||
|
||||
class TestRegistrySurface(unittest.TestCase):
|
||||
"""AC5: the console surfaces the control, still disabled, with no kill."""
|
||||
|
||||
def test_registry_exposes_both_actions_disabled(self):
|
||||
registry = gated_actions.load_action_registry()
|
||||
for action_id in (
|
||||
sanctioned_restart.ACTION_RESTART_NAMESPACE,
|
||||
sanctioned_restart.ACTION_RELOAD_NAMESPACE,
|
||||
):
|
||||
with self.subTest(action=action_id):
|
||||
action = registry.get(action_id)
|
||||
self.assertIsNotNone(action)
|
||||
self.assertFalse(action.enabled)
|
||||
|
||||
def test_registry_preview_names_the_namespace_target(self):
|
||||
preview = gated_actions.preview_action(
|
||||
sanctioned_restart.ACTION_RESTART_NAMESPACE, namespace=NAMESPACE
|
||||
)
|
||||
target = preview["mutation_ledger"][0]["target"]
|
||||
self.assertIn(NAMESPACE, target)
|
||||
self.assertFalse(preview["enabled"])
|
||||
|
||||
def test_registry_attempt_fails_closed(self):
|
||||
result = gated_actions.attempt_action(
|
||||
sanctioned_restart.ACTION_RESTART_NAMESPACE, namespace=NAMESPACE
|
||||
)
|
||||
self.assertFalse(result["success"])
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
unittest.main()
|
||||
@@ -46,6 +46,8 @@ from webui.worktree_scanner import load_hygiene_snapshot, snapshot_to_dict as wo
|
||||
from webui.worktree_views import render_worktrees_page
|
||||
from webui.runtime_health import load_runtime_snapshot, snapshot_to_dict as runtime_snapshot_to_dict
|
||||
from webui.runtime_views import render_runtime_page
|
||||
from webui.policy_inventory import load_policy_inventory, snapshot_to_dict as policy_snapshot_to_dict
|
||||
from webui.policy_views import render_policy_page
|
||||
from webui.timeline import load_timeline, snapshot_to_dict as timeline_snapshot_to_dict
|
||||
from webui.system_health import (
|
||||
API_PATH as SYSTEM_HEALTH_API_PATH,
|
||||
@@ -302,6 +304,17 @@ async def api_runtime(_request: Request) -> JSONResponse:
|
||||
return JSONResponse(runtime_snapshot_to_dict(load_runtime_snapshot()))
|
||||
|
||||
|
||||
async def policy(_request: Request) -> HTMLResponse:
|
||||
snapshot = load_policy_inventory()
|
||||
return HTMLResponse(
|
||||
render_page(title="Policy", body_html=render_policy_page(snapshot))
|
||||
)
|
||||
|
||||
|
||||
async def api_v1_policy(_request: Request) -> JSONResponse:
|
||||
return JSONResponse(policy_snapshot_to_dict(load_policy_inventory()))
|
||||
|
||||
|
||||
async def _parse_audit_form(request: Request) -> tuple[str, str | None]:
|
||||
if request.method == "GET":
|
||||
return "", None
|
||||
@@ -607,6 +620,8 @@ def create_app(*, bind_host: str | None = None) -> Starlette:
|
||||
Route("/api/prompts", api_prompts, methods=["GET"]),
|
||||
Route("/runtime", runtime, methods=["GET"]),
|
||||
Route("/api/runtime", api_runtime, methods=["GET"]),
|
||||
Route("/policy", policy, methods=["GET"]),
|
||||
Route("/api/v1/policy", api_v1_policy, methods=["GET"]),
|
||||
Route("/api/v1/timeline", api_v1_timeline, methods=["GET"]),
|
||||
Route("/audit", audit, methods=["GET", "POST"]),
|
||||
Route("/api/audit", api_audit, methods=["GET", "POST"]),
|
||||
|
||||
@@ -236,6 +236,34 @@ _ACTION_SPECS: tuple[ConsoleAction, ...] = (
|
||||
phase=3,
|
||||
summary="Remove a remote feature branch.",
|
||||
),
|
||||
# #642: sanctioned daemon lifecycle. These exist so operators have an
|
||||
# audited path off `pkill -f mcp_server.py` (#630). Restart drops every
|
||||
# in-flight request on a namespace, so it carries the same dual-control and
|
||||
# break-glass weight as a merge; reload drains first and is privileged but
|
||||
# not destructive. Neither ever exposes a raw kill: execution is handed to
|
||||
# a host supervisor by ``webui.sanctioned_restart``.
|
||||
ConsoleAction(
|
||||
action_id="system.reload_namespace",
|
||||
task_key="reload_namespace",
|
||||
action_class=CLASS_PRIVILEGED,
|
||||
minimum_role=CONTROLLER,
|
||||
requires_confirmation=True,
|
||||
dual_control=False,
|
||||
break_glass=False,
|
||||
phase=2,
|
||||
summary="Gracefully reload one MCP namespace via the host supervisor.",
|
||||
),
|
||||
ConsoleAction(
|
||||
action_id="system.restart_namespace",
|
||||
task_key="restart_namespace",
|
||||
action_class=CLASS_DESTRUCTIVE,
|
||||
minimum_role=ADMIN,
|
||||
requires_confirmation=True,
|
||||
dual_control=True,
|
||||
break_glass=True,
|
||||
phase=2,
|
||||
summary="Restart one MCP namespace via the host supervisor.",
|
||||
),
|
||||
)
|
||||
|
||||
ACTIONS: dict[str, ConsoleAction] = {a.action_id: a for a in _ACTION_SPECS}
|
||||
|
||||
@@ -110,6 +110,8 @@ def _format_target(action_id: str, params: dict[str, Any]) -> str:
|
||||
)
|
||||
if action_id == "create_issue":
|
||||
return f"issue {params.get('title', '?')!r}"
|
||||
if action_id in {"system.restart_namespace", "system.reload_namespace"}:
|
||||
return f"MCP namespace {params.get('namespace', '?')!r}"
|
||||
return "unspecified"
|
||||
|
||||
|
||||
@@ -165,6 +167,17 @@ def build_action_registry() -> ActionRegistry:
|
||||
"gitea_create_issue_comment", "Post a PR review thread comment."),
|
||||
("close_pr", "Close PR", "close_pr", "gitea_edit_pr",
|
||||
"Close a pull request without merge."),
|
||||
# #642: the sanctioned replacement for the forbidden manual daemon-kill
|
||||
# recovery path (#630). The "tool" is a host supervisor hook, not an MCP
|
||||
# call — the console never signals a process. Preview and gating live in
|
||||
# ``webui.sanctioned_restart``; these stay disabled like every other
|
||||
# registry entry.
|
||||
("system.reload_namespace", "Reload MCP namespace", "reload_namespace",
|
||||
"host.supervisor_reload",
|
||||
"Gracefully reload one MCP namespace via the host supervisor."),
|
||||
("system.restart_namespace", "Restart MCP namespace",
|
||||
"restart_namespace", "host.supervisor_restart",
|
||||
"Restart one MCP namespace via the host supervisor."),
|
||||
)
|
||||
actions = tuple(
|
||||
GatedAction(
|
||||
|
||||
+2
-6
@@ -5,7 +5,7 @@ the ``webui/app.py`` route table stay aligned with epic #631. Read-only: every
|
||||
destination is a GET view or a Phase 1 placeholder. No mutation links.
|
||||
|
||||
Nav groups follow the #631 Phase 1 information architecture: Health, Traffic,
|
||||
Runtime/Sessions, Projects, Inventory, Timeline, Policy (placeholder), and
|
||||
Runtime/Sessions, Projects, Inventory, Timeline, Policy (live via #646), and
|
||||
Insights (placeholder). Later-phase surfaces are declared as ``stub`` items and
|
||||
backed by ``STUB_PAGES`` so their nav links resolve to a graceful placeholder
|
||||
instead of a 404.
|
||||
@@ -60,7 +60,7 @@ NAV_GROUPS: tuple[NavGroup, ...] = (
|
||||
NavItem("/timeline", "Timeline", "stub"),
|
||||
)),
|
||||
NavGroup("Policy", (
|
||||
NavItem("/policy", "Policy", "stub"),
|
||||
NavItem("/policy", "Policy"),
|
||||
NavItem("/prompts", "Prompts"),
|
||||
)),
|
||||
NavGroup("Insights", (
|
||||
@@ -88,10 +88,6 @@ STUB_PAGES: dict[str, tuple[str, str]] = {
|
||||
"Timeline",
|
||||
"Workflow event timeline across issues and PRs. A later Phase 1 surface.",
|
||||
),
|
||||
"/policy": (
|
||||
"Policy",
|
||||
"Capability and role policy surface. Placeholder until a later phase.",
|
||||
),
|
||||
"/insights": (
|
||||
"Insights",
|
||||
"Aggregate operational insights and trends. Placeholder until a later "
|
||||
|
||||
@@ -0,0 +1,387 @@
|
||||
"""Read-only workflow policy and guardrail inventory for the web UI (#646).
|
||||
|
||||
Policy and guardrails live in code, profiles, docs, and skills. An operator
|
||||
cannot *see* the active workflow policy configuration from the console without
|
||||
reading the repository tree. This module projects the major guardrails into a
|
||||
redacted, machine-readable inventory with source attribution (file / module /
|
||||
doc), so the console can render them as HTML tables with source pointers.
|
||||
|
||||
Design constraints (Phase 3, #646):
|
||||
|
||||
- **Read-only projection.** Nothing here edits policy or exposes a toggle that
|
||||
could weaken a gate. It reports what is already enforced elsewhere.
|
||||
- **Source attribution without secrets.** Every guardrail carries pointers to
|
||||
the file/module/doc that owns it. Live values are compact summaries derived
|
||||
from the safe policy accessors that already exist (``rbac_matrix``,
|
||||
``redaction_policy``, ``audit_policy``); raw regex, tokens, and endpoints are
|
||||
never embedded.
|
||||
- **Redact before emit.** ``snapshot_to_dict`` runs the whole payload through
|
||||
``console_redaction.redact_payload`` so a planted or accidental secret in any
|
||||
projected value degrades to the placeholder rather than reaching a client.
|
||||
- **Fail soft.** A projection that raises is recorded as a per-entry error and
|
||||
never takes the page down; a guardrail is still listed with its sources.
|
||||
- **Diff vs documented defaults where feasible.** When a guardrail declares a
|
||||
documented invariant, the active projection is compared against it and the
|
||||
result is reported; otherwise the diff is explicitly ``None`` with a reason.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from dataclasses import dataclass
|
||||
from typing import Any, Callable
|
||||
|
||||
from webui import console_audit
|
||||
from webui import console_authz
|
||||
from webui import console_redaction
|
||||
|
||||
SCHEMA_VERSION = 1
|
||||
|
||||
READ_ONLY_NOTE = (
|
||||
"Read-only projection of guardrails enforced in code, profiles, docs, and "
|
||||
"skills. This view never edits policy and exposes no gate-weakening toggle."
|
||||
)
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class SourcePointer:
|
||||
"""Where a guardrail is defined. Attribution only — never a secret."""
|
||||
|
||||
label: str
|
||||
path: str
|
||||
kind: str # "module" | "doc" | "script" | "config"
|
||||
anchor: str | None = None
|
||||
|
||||
def to_dict(self) -> dict[str, Any]:
|
||||
return {
|
||||
"label": self.label,
|
||||
"path": self.path,
|
||||
"kind": self.kind,
|
||||
"anchor": self.anchor,
|
||||
}
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class PolicyEntry:
|
||||
key: str
|
||||
title: str
|
||||
category: str
|
||||
summary: str
|
||||
sources: tuple[SourcePointer, ...]
|
||||
active: dict[str, Any] | None
|
||||
documented_default: dict[str, Any] | None
|
||||
diff: dict[str, Any] | None
|
||||
error: str | None = None
|
||||
|
||||
def to_dict(self) -> dict[str, Any]:
|
||||
return {
|
||||
"key": self.key,
|
||||
"title": self.title,
|
||||
"category": self.category,
|
||||
"summary": self.summary,
|
||||
"sources": [s.to_dict() for s in self.sources],
|
||||
"active": self.active,
|
||||
"documented_default": self.documented_default,
|
||||
"diff": self.diff,
|
||||
"error": self.error,
|
||||
}
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class PolicyInventorySnapshot:
|
||||
schema_version: int
|
||||
read_only: bool
|
||||
note: str
|
||||
entries: tuple[PolicyEntry, ...]
|
||||
categories: tuple[str, ...]
|
||||
build_errors: tuple[str, ...]
|
||||
|
||||
|
||||
def _diff_active_vs_default(
|
||||
active: dict[str, Any] | None,
|
||||
documented_default: dict[str, Any] | None,
|
||||
) -> dict[str, Any] | None:
|
||||
"""Compare only the keys the documented default declares.
|
||||
|
||||
Returns ``None`` when no documented default is declared (diff not feasible)
|
||||
or when the active projection is unavailable. Otherwise reports, per
|
||||
declared key, whether the active value matches the documented invariant.
|
||||
"""
|
||||
if not documented_default:
|
||||
return None
|
||||
if not active:
|
||||
return {"status": "active_unavailable", "checked": {}}
|
||||
checked: dict[str, Any] = {}
|
||||
matches = True
|
||||
for key, expected in documented_default.items():
|
||||
observed = active.get(key)
|
||||
ok = observed == expected
|
||||
matches = matches and ok
|
||||
checked[key] = {"expected": expected, "observed": observed, "matches": ok}
|
||||
return {
|
||||
"status": "matches_documented_default" if matches else "drift_detected",
|
||||
"checked": checked,
|
||||
}
|
||||
|
||||
|
||||
# ── Live projections (compact, safe, fail-soft) ──────────────────────────────
|
||||
# Each returns a small dict of already-safe machine values. They are module
|
||||
# level so tests can substitute one to prove the redaction pass runs.
|
||||
|
||||
|
||||
def _project_role_separation() -> dict[str, Any]:
|
||||
matrix = console_authz.rbac_matrix()
|
||||
return {
|
||||
"model_version": matrix.get("model_version"),
|
||||
"active_phase": matrix.get("active_phase"),
|
||||
"roles": [r.get("role") for r in matrix.get("roles", [])],
|
||||
"privileged_action_count": len(matrix.get("privileged_actions", [])),
|
||||
"default_decision": matrix.get("default_decision"),
|
||||
"execution_enabled": matrix.get("execution_enabled"),
|
||||
}
|
||||
|
||||
|
||||
def _project_redaction() -> dict[str, Any]:
|
||||
policy = console_redaction.redaction_policy()
|
||||
return {
|
||||
"policy_version": policy.get("policy_version"),
|
||||
"placeholder": policy.get("placeholder"),
|
||||
"applies_to": policy.get("applies_to"),
|
||||
"console_detector_count": len(policy.get("console_rules", [])),
|
||||
"redact_before_persist": policy.get("redact_before_persist"),
|
||||
"failure_mode": policy.get("failure_mode"),
|
||||
}
|
||||
|
||||
|
||||
def _project_audit() -> dict[str, Any]:
|
||||
policy = console_audit.audit_policy()
|
||||
return {
|
||||
"schema_version": policy.get("schema_version"),
|
||||
"required_field_count": len(policy.get("required_fields", [])),
|
||||
"results": policy.get("results"),
|
||||
"retention_defaults_days": policy.get("retention_defaults_days"),
|
||||
"append_only": policy.get("append_only"),
|
||||
"redact_before_persist": policy.get("redact_before_persist"),
|
||||
"enabled": policy.get("enabled"),
|
||||
}
|
||||
|
||||
|
||||
def _static(value: dict[str, Any]) -> Callable[[], dict[str, Any]]:
|
||||
return lambda: dict(value)
|
||||
|
||||
|
||||
# ── Guardrail catalog ────────────────────────────────────────────────────────
|
||||
# One row per major guardrail. ``project`` yields the active value (may raise;
|
||||
# caught per entry). ``documented_default`` drives the feasible diff.
|
||||
|
||||
_CatalogRow = tuple[
|
||||
str,
|
||||
str,
|
||||
str,
|
||||
str,
|
||||
tuple[SourcePointer, ...],
|
||||
Callable[[], dict[str, Any]] | None,
|
||||
dict[str, Any] | None,
|
||||
]
|
||||
|
||||
_CATALOG: tuple[_CatalogRow, ...] = (
|
||||
(
|
||||
"role_separation",
|
||||
"Role separation and RBAC",
|
||||
"role_separation",
|
||||
"Author, reviewer, merger, and reconciler capabilities are disjoint and "
|
||||
"role-exclusive; self-review and self-merge are always blocked. The "
|
||||
"console RBAC model defaults to deny.",
|
||||
(
|
||||
SourcePointer("task capability map", "task_capability_map.py", "module"),
|
||||
SourcePointer("role/namespace gate", "role_namespace_gate.py", "module"),
|
||||
SourcePointer("console RBAC", "webui/console_authz.py", "module"),
|
||||
),
|
||||
_project_role_separation,
|
||||
{"default_decision": "deny", "execution_enabled": False},
|
||||
),
|
||||
(
|
||||
"lease_rules",
|
||||
"Issue and PR lease lifecycle",
|
||||
"lease_rules",
|
||||
"Durable work is claimed through issue locks and control-plane leases "
|
||||
"with freshness, expiry, and dead-session recovery; abandoned or stale "
|
||||
"claims are reclaimed only through the sanctioned recovery path.",
|
||||
(
|
||||
SourcePointer("issue lock store", "issue_lock_store.py", "module"),
|
||||
SourcePointer("branch cleanup guard", "branch_cleanup_guard.py", "module"),
|
||||
SourcePointer("safety model §5", "docs/safety-model.md", "doc", "5-mutation-gating"),
|
||||
),
|
||||
None,
|
||||
None,
|
||||
),
|
||||
(
|
||||
"worktree_rules",
|
||||
"Author worktree binding",
|
||||
"worktree_rules",
|
||||
"Author mutations require a validated worktree under branches/ derived "
|
||||
"from the active issue lock; silent fallback to the stable control "
|
||||
"checkout or master is forbidden (#618).",
|
||||
(
|
||||
SourcePointer("author worktree gate", "author_mutation_worktree.py", "module"),
|
||||
SourcePointer("worktree bootstrap", "scripts/worktree-start", "script"),
|
||||
SourcePointer("workflow scope guard", "workflow_scope_guard.py", "module"),
|
||||
),
|
||||
None,
|
||||
None,
|
||||
),
|
||||
(
|
||||
"merge_confirmation",
|
||||
"Explicit merge confirmation",
|
||||
"merge_confirmation",
|
||||
"A merge fails closed unless the caller passes the exact confirmation "
|
||||
"phrase for that PR; reviewing never implies merging.",
|
||||
(
|
||||
SourcePointer("merge path", "merge_pr.py", "module"),
|
||||
SourcePointer("merge tool gate", "gitea_mcp_server.py", "module"),
|
||||
),
|
||||
_static({"required_confirmation_format": "MERGE PR <n>", "auto_merge": False}),
|
||||
{"auto_merge": False},
|
||||
),
|
||||
(
|
||||
"redaction",
|
||||
"Secret redaction",
|
||||
"redaction",
|
||||
"Every console surface runs the shared gitea_audit pass then console "
|
||||
"patterns before any payload, HTML, log line, or audit record leaves "
|
||||
"the server; unredactable values fail closed to the placeholder.",
|
||||
(
|
||||
SourcePointer("console redaction", "webui/console_redaction.py", "module"),
|
||||
SourcePointer("shared redaction", "gitea_audit.py", "module"),
|
||||
SourcePointer("safety model §3", "docs/safety-model.md", "doc", "3-secret-redaction"),
|
||||
),
|
||||
_project_redaction,
|
||||
{"redact_before_persist": True},
|
||||
),
|
||||
(
|
||||
"contamination",
|
||||
"Contamination containment",
|
||||
"contamination",
|
||||
"A session contaminated by a direct stable-branch push or a manual MCP "
|
||||
"daemon kill is blocked from review, merge, close, and completion "
|
||||
"mutations until cleared (reconciler-exempt).",
|
||||
(
|
||||
SourcePointer("contamination gates", "gitea_mcp_server.py", "module"),
|
||||
SourcePointer("stable-branch audit", "workflow_scope_guard.py", "module"),
|
||||
),
|
||||
None,
|
||||
None,
|
||||
),
|
||||
(
|
||||
"allocator_policy",
|
||||
"Work allocation policy",
|
||||
"allocator_policy",
|
||||
"Workers do not self-select exclusive work; the controller-owned "
|
||||
"allocator ranks the complete queue by priority then PRs-before-issues "
|
||||
"then ascending number, honoring dependency edges and foreign claims.",
|
||||
(
|
||||
SourcePointer("allocator", "gitea_mcp_server.py", "module"),
|
||||
SourcePointer("safety model §5", "docs/safety-model.md", "doc", "5-mutation-gating"),
|
||||
),
|
||||
_static(
|
||||
{
|
||||
"self_select_exclusive_work": False,
|
||||
"ranking": "priority desc, PRs before issues, number asc",
|
||||
"respects_dependency_edges": True,
|
||||
"respects_foreign_claims": True,
|
||||
}
|
||||
),
|
||||
{"self_select_exclusive_work": False},
|
||||
),
|
||||
(
|
||||
"audit_logging",
|
||||
"Audit logging",
|
||||
"audit_logging",
|
||||
"Console intent and authorization outcomes are recorded to an "
|
||||
"append-only, redact-before-persist audit log; MCP mutations are "
|
||||
"recorded by gitea_audit and correlated by request id.",
|
||||
(
|
||||
SourcePointer("console audit", "webui/console_audit.py", "module"),
|
||||
SourcePointer("MCP audit", "gitea_audit.py", "module"),
|
||||
SourcePointer("safety model §1", "docs/safety-model.md", "doc", "1-audit-logging-and-confirmation"),
|
||||
),
|
||||
_project_audit,
|
||||
{"append_only": True, "redact_before_persist": True},
|
||||
),
|
||||
(
|
||||
"mutation_gating",
|
||||
"Mutation gating and master parity",
|
||||
"mutation_gating",
|
||||
"Mutations fail closed while the running server is stale relative to "
|
||||
"master, and every mutation is preceded by identity and capability "
|
||||
"resolution in a fixed pre-flight order.",
|
||||
(
|
||||
SourcePointer("mutation gate", "gitea_mcp_server.py", "module"),
|
||||
SourcePointer("safety model §5", "docs/safety-model.md", "doc", "5-mutation-gating"),
|
||||
),
|
||||
_static(
|
||||
{
|
||||
"stale_runtime_blocks_mutations": True,
|
||||
"preflight_order": "whoami -> resolve_task_capability -> mutation",
|
||||
}
|
||||
),
|
||||
{"stale_runtime_blocks_mutations": True},
|
||||
),
|
||||
)
|
||||
|
||||
|
||||
def _build_entry(row: _CatalogRow) -> PolicyEntry:
|
||||
key, title, category, summary, sources, project, documented_default = row
|
||||
active: dict[str, Any] | None = None
|
||||
error: str | None = None
|
||||
if project is not None:
|
||||
try:
|
||||
active = project()
|
||||
except Exception as exc: # noqa: BLE001 — fail soft; never take the page down
|
||||
active = None
|
||||
error = f"active projection unavailable: {exc}"
|
||||
diff = _diff_active_vs_default(active, documented_default)
|
||||
return PolicyEntry(
|
||||
key=key,
|
||||
title=title,
|
||||
category=category,
|
||||
summary=summary,
|
||||
sources=sources,
|
||||
active=active,
|
||||
documented_default=documented_default,
|
||||
diff=diff,
|
||||
error=error,
|
||||
)
|
||||
|
||||
|
||||
def load_policy_inventory() -> PolicyInventorySnapshot:
|
||||
"""Build the read-only guardrail inventory. Never raises for one bad entry."""
|
||||
entries: list[PolicyEntry] = []
|
||||
build_errors: list[str] = []
|
||||
for row in _CATALOG:
|
||||
try:
|
||||
entries.append(_build_entry(row))
|
||||
except Exception as exc: # noqa: BLE001 — one row must not break the rest
|
||||
build_errors.append(f"{row[0]}: {exc}")
|
||||
categories = tuple(dict.fromkeys(e.category for e in entries))
|
||||
return PolicyInventorySnapshot(
|
||||
schema_version=SCHEMA_VERSION,
|
||||
read_only=True,
|
||||
note=READ_ONLY_NOTE,
|
||||
entries=tuple(entries),
|
||||
categories=categories,
|
||||
build_errors=tuple(build_errors),
|
||||
)
|
||||
|
||||
|
||||
def snapshot_to_dict(snapshot: PolicyInventorySnapshot) -> dict[str, Any]:
|
||||
"""Serialize the snapshot, redacting the entire payload before it is emitted."""
|
||||
payload = {
|
||||
"schema_version": snapshot.schema_version,
|
||||
"read_only": snapshot.read_only,
|
||||
"note": snapshot.note,
|
||||
"categories": list(snapshot.categories),
|
||||
"entry_count": len(snapshot.entries),
|
||||
"entries": [entry.to_dict() for entry in snapshot.entries],
|
||||
"build_errors": list(snapshot.build_errors),
|
||||
}
|
||||
return console_redaction.redact_payload(payload)
|
||||
@@ -0,0 +1,145 @@
|
||||
"""HTML views for the workflow policy and guardrail inventory (#646)."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import html
|
||||
import json
|
||||
from typing import Any
|
||||
|
||||
from webui import console_redaction
|
||||
from webui.policy_inventory import (
|
||||
PolicyEntry,
|
||||
PolicyInventorySnapshot,
|
||||
SourcePointer,
|
||||
snapshot_to_dict,
|
||||
)
|
||||
|
||||
|
||||
def _source_pointer(source: dict[str, Any] | SourcePointer) -> str:
|
||||
if isinstance(source, dict):
|
||||
label = str(source.get("label") or "")
|
||||
path = str(source.get("path") or "")
|
||||
anchor = source.get("anchor")
|
||||
kind = str(source.get("kind") or "")
|
||||
else:
|
||||
label = source.label
|
||||
path = source.path
|
||||
anchor = source.anchor
|
||||
kind = source.kind
|
||||
|
||||
if anchor:
|
||||
path = f"{path}#{anchor}"
|
||||
return (
|
||||
f"<li>{html.escape(label)} — "
|
||||
f"<code>{html.escape(path)}</code> "
|
||||
f"<span class='muted'>({html.escape(kind)})</span></li>"
|
||||
)
|
||||
|
||||
|
||||
def _active_block(entry: dict[str, Any] | PolicyEntry) -> str:
|
||||
error = entry.get("error") if isinstance(entry, dict) else entry.error
|
||||
active = entry.get("active") if isinstance(entry, dict) else entry.active
|
||||
if error:
|
||||
return (
|
||||
"<p class='muted'><strong>Active value unavailable:</strong> "
|
||||
f"{html.escape(error)}</p>"
|
||||
)
|
||||
if not active:
|
||||
return "<p class='muted'>No live projection for this guardrail.</p>"
|
||||
pretty = json.dumps(active, indent=2, sort_keys=True, default=str)
|
||||
return f"<pre class='prompt-text'>{html.escape(pretty)}</pre>"
|
||||
|
||||
|
||||
def _diff_block(entry: dict[str, Any] | PolicyEntry) -> str:
|
||||
diff = entry.get("diff") if isinstance(entry, dict) else entry.diff
|
||||
documented_default = entry.get("documented_default") if isinstance(entry, dict) else entry.documented_default
|
||||
if diff is None:
|
||||
if documented_default is None:
|
||||
return "<p class='muted'>Diff vs documented default: not feasible (no declared default).</p>"
|
||||
return "<p class='muted'>Diff vs documented default: unavailable.</p>"
|
||||
status = str(diff.get("status") if isinstance(diff, dict) else "unknown")
|
||||
badge = "badge-claimed" if status == "matches_documented_default" else "badge-blocked"
|
||||
rows = []
|
||||
checked = diff.get("checked") if isinstance(diff, dict) else {}
|
||||
if isinstance(checked, dict):
|
||||
for key, cell in checked.items():
|
||||
cell_dict = cell if isinstance(cell, dict) else {}
|
||||
marker = "✓" if cell_dict.get("matches") else "✗"
|
||||
rows.append(
|
||||
"<tr>"
|
||||
f"<td><code>{html.escape(str(key))}</code></td>"
|
||||
f"<td><code>{html.escape(str(cell_dict.get('expected')))}</code></td>"
|
||||
f"<td><code>{html.escape(str(cell_dict.get('observed')))}</code></td>"
|
||||
f"<td>{marker}</td>"
|
||||
"</tr>"
|
||||
)
|
||||
table = ""
|
||||
if rows:
|
||||
table = (
|
||||
"<table class='detail'><thead><tr>"
|
||||
"<th>Key</th><th>Documented</th><th>Active</th><th>Match</th>"
|
||||
"</tr></thead><tbody>"
|
||||
f"{''.join(rows)}</tbody></table>"
|
||||
)
|
||||
return (
|
||||
f"<p class='meta'>Diff vs documented default: "
|
||||
f"<span class='badge {badge}'>{html.escape(status)}</span></p>"
|
||||
f"{table}"
|
||||
)
|
||||
|
||||
|
||||
def _entry_card(entry: dict[str, Any] | PolicyEntry) -> str:
|
||||
title = str(entry.get("title") if isinstance(entry, dict) else entry.title)
|
||||
category = str(entry.get("category") if isinstance(entry, dict) else entry.category)
|
||||
summary = str(entry.get("summary") if isinstance(entry, dict) else entry.summary)
|
||||
sources_data = entry.get("sources", []) if isinstance(entry, dict) else entry.sources
|
||||
|
||||
sources = "".join(_source_pointer(s) for s in sources_data)
|
||||
return (
|
||||
"<div class='prompt-card'>"
|
||||
f"<h3>{html.escape(title)} "
|
||||
f"<span class='badge'>{html.escape(category)}</span></h3>"
|
||||
f"<p>{html.escape(summary)}</p>"
|
||||
"<p class='meta'><strong>Source pointers</strong></p>"
|
||||
f"<ul>{sources}</ul>"
|
||||
"<p class='meta'><strong>Active configuration</strong></p>"
|
||||
f"{_active_block(entry)}"
|
||||
f"{_diff_block(entry)}"
|
||||
"</div>"
|
||||
)
|
||||
|
||||
|
||||
def render_policy_page(snapshot: PolicyInventorySnapshot | dict[str, Any]) -> str:
|
||||
if isinstance(snapshot, PolicyInventorySnapshot):
|
||||
payload = snapshot_to_dict(snapshot)
|
||||
elif isinstance(snapshot, dict):
|
||||
payload = console_redaction.redact_payload(snapshot)
|
||||
else:
|
||||
payload = {}
|
||||
|
||||
categories_list = payload.get("categories") or []
|
||||
categories = ", ".join(html.escape(str(c)) for c in categories_list) or "none"
|
||||
entries_list = payload.get("entries") or []
|
||||
cards = "".join(_entry_card(e) for e in entries_list)
|
||||
build_errors = ""
|
||||
errors_list = payload.get("build_errors") or []
|
||||
if errors_list:
|
||||
items = "".join(
|
||||
f"<li>{html.escape(str(err))}</li>" for err in errors_list
|
||||
)
|
||||
build_errors = (
|
||||
"<div class='stub'><p><strong>Some guardrails could not be built:"
|
||||
f"</strong></p><ul>{items}</ul></div>"
|
||||
)
|
||||
note = str(payload.get("note") or "")
|
||||
schema_version = payload.get("schema_version") or 1
|
||||
return (
|
||||
"<h2>Workflow policy & guardrails</h2>"
|
||||
f"<p class='muted'>{html.escape(note)}</p>"
|
||||
f"<p class='meta'>Schema v{schema_version} · "
|
||||
f"{len(entries_list)} guardrails · categories: {categories}</p>"
|
||||
f"{build_errors}"
|
||||
f"{cards}"
|
||||
"<p class='muted'>This page is read-only. It reports enforced policy "
|
||||
"and never edits or weakens a gate. Secret values are redacted.</p>"
|
||||
)
|
||||
@@ -0,0 +1,613 @@
|
||||
"""Sanctioned MCP restart and graceful reload controls (#642, Phase 2).
|
||||
|
||||
Sessions have historically recovered MCP connectivity by killing the host
|
||||
daemon (``pkill -f mcp_server.py``, #630). That path stays forbidden: it kills
|
||||
every namespace on the host, contaminates the surviving session, and leaves no
|
||||
audit trail. This module is the sanctioned replacement.
|
||||
|
||||
A restart is modelled as a *gated action*, never as a command:
|
||||
|
||||
1. **Capability** — the console action resolves through ``console_authz``
|
||||
against ``task_capability_map``, so the console cannot invent an authority
|
||||
the MCP layer does not already define.
|
||||
2. **Preview** — :func:`build_restart_preview` renders a mutation ledger and
|
||||
the exact confirmation phrase. It never returns a shell command.
|
||||
3. **Confirmation** — the operator echoes a phrase naming the exact namespace
|
||||
and mode. A phrase for one namespace never authorizes another.
|
||||
4. **Operator authorization** — host daemon maintenance is authorized out of
|
||||
band through the environment (#630, and #710 finding F1: a worker session
|
||||
cannot set an env var for an already-running daemon, so this cannot be
|
||||
self-asserted the way a tool argument could).
|
||||
5. **Execution** — :func:`execute_restart` never spawns a process. Once every
|
||||
gate passes it hands the request to the configured host-managed restart
|
||||
hook; with no hook configured it fails closed.
|
||||
6. **Health recheck** — :func:`verify_post_restart_health` requires live
|
||||
client-namespace probe evidence before any post-restart clean claim.
|
||||
|
||||
Manual ``pkill`` remains forbidden and is classified as contamination by
|
||||
:func:`classify_restart_command`, which blocks clean claims (#630 AC3).
|
||||
|
||||
This module performs no I/O beyond reading its own environment configuration,
|
||||
imports no MCP client, and holds no credential.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import os
|
||||
from dataclasses import asdict, dataclass
|
||||
from typing import Any
|
||||
|
||||
import mcp_namespace_health
|
||||
import runtime_recovery_guard
|
||||
from webui import console_audit, console_authz
|
||||
|
||||
# --- Operations -------------------------------------------------------------
|
||||
|
||||
MODE_RESTART = "restart"
|
||||
MODE_RELOAD = "reload"
|
||||
MODES: tuple[str, ...] = (MODE_RESTART, MODE_RELOAD)
|
||||
|
||||
ACTION_RESTART_NAMESPACE = "system.restart_namespace"
|
||||
ACTION_RELOAD_NAMESPACE = "system.reload_namespace"
|
||||
|
||||
ACTION_FOR_MODE: dict[str, str] = {
|
||||
MODE_RESTART: ACTION_RESTART_NAMESPACE,
|
||||
MODE_RELOAD: ACTION_RELOAD_NAMESPACE,
|
||||
}
|
||||
|
||||
# Namespaces the console may target. An unlisted name fails closed rather than
|
||||
# being passed through to a host hook.
|
||||
KNOWN_NAMESPACES: tuple[str, ...] = tuple(
|
||||
sorted(
|
||||
set(mcp_namespace_health.DEFAULT_NAMESPACES)
|
||||
| {"gitea-author", "gitea-reviewer", "gitea-merger",
|
||||
"gitea-reconciler", "gitea-controller"}
|
||||
)
|
||||
)
|
||||
|
||||
# Scope tokens that would mean "everything at once". Explicit non-goal: the
|
||||
# console never offers a fleet-wide restart, because that is the blast radius
|
||||
# `pkill -f mcp_server.py` already had.
|
||||
_FLEET_TOKENS = frozenset({"*", "all", "fleet", "any", ""})
|
||||
|
||||
# --- Environment configuration ----------------------------------------------
|
||||
# Read server-side only; the value is an opaque host hook reference (e.g. a
|
||||
# launchd label), never a command line, and is never rendered to a client.
|
||||
RESTART_HOOK_ENV = "GITEA_SANCTIONED_RESTART_HOOK"
|
||||
|
||||
# --- Reason codes -----------------------------------------------------------
|
||||
|
||||
DENY_UNKNOWN_MODE = "unknown_mode"
|
||||
DENY_UNKNOWN_NAMESPACE = "unknown_namespace"
|
||||
DENY_FLEET_SCOPE = "fleet_scope_not_permitted"
|
||||
DENY_UNAUTHORIZED = "unauthorized"
|
||||
DENY_CONFIRMATION_MISSING = "confirmation_required"
|
||||
DENY_CONFIRMATION_MISMATCH = "confirmation_mismatch"
|
||||
DENY_OPERATOR_AUTHORIZATION = "operator_authorization_missing"
|
||||
DENY_HOOK_NOT_CONFIGURED = "restart_hook_not_configured"
|
||||
DENY_CONTAMINATED_RUNTIME = "contaminated_runtime"
|
||||
|
||||
ALLOW_HOST_ACTION_REQUIRED = "host_action_required"
|
||||
|
||||
# Post-restart verification outcomes.
|
||||
HEALTH_CLEAN = "clean"
|
||||
HEALTH_UNPROVEN = "unproven"
|
||||
HEALTH_UNHEALTHY = "unhealthy"
|
||||
|
||||
|
||||
def _clean(value: Any) -> str:
|
||||
return str(value or "").strip()
|
||||
|
||||
|
||||
# --- Mutation ledger --------------------------------------------------------
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class RestartLedgerEntry:
|
||||
"""One planned step, shown before anything is asked of the host."""
|
||||
|
||||
sequence: int
|
||||
step: str
|
||||
summary: str
|
||||
executes_process_kill: bool = False
|
||||
|
||||
|
||||
def _mutation_ledger(namespace: str, mode: str) -> tuple[RestartLedgerEntry, ...]:
|
||||
if mode == MODE_RELOAD:
|
||||
middle = RestartLedgerEntry(
|
||||
sequence=2,
|
||||
step="host_graceful_reload",
|
||||
summary=(
|
||||
f"Ask the configured host supervisor to reload {namespace} "
|
||||
"in place, draining in-flight requests. The console does not "
|
||||
"signal the process itself."
|
||||
),
|
||||
)
|
||||
else:
|
||||
middle = RestartLedgerEntry(
|
||||
sequence=2,
|
||||
step="host_restart_hook",
|
||||
summary=(
|
||||
f"Ask the configured host supervisor to restart {namespace}. "
|
||||
"The console never sends a signal and never runs a kill."
|
||||
),
|
||||
)
|
||||
return (
|
||||
RestartLedgerEntry(
|
||||
sequence=1,
|
||||
step="quiesce",
|
||||
summary=(
|
||||
f"Stop admitting new gated mutations for {namespace} and "
|
||||
"record the intent before anything restarts."
|
||||
),
|
||||
),
|
||||
middle,
|
||||
RestartLedgerEntry(
|
||||
sequence=3,
|
||||
step="health_recheck",
|
||||
summary=(
|
||||
f"Re-probe {namespace} through the live client namespace and "
|
||||
"prove the required tool is callable again."
|
||||
),
|
||||
),
|
||||
RestartLedgerEntry(
|
||||
sequence=4,
|
||||
step="audit",
|
||||
summary=(
|
||||
"Append actor, target namespace, mode, and result to the "
|
||||
"console audit log."
|
||||
),
|
||||
),
|
||||
)
|
||||
|
||||
|
||||
# --- Confirmation -----------------------------------------------------------
|
||||
|
||||
|
||||
def confirmation_phrase(namespace: str, mode: str) -> str:
|
||||
"""Exact phrase an operator must echo, naming the namespace and mode.
|
||||
|
||||
Binding the namespace into the phrase is the point: a confirmation typed
|
||||
for ``gitea-author`` cannot be replayed against ``gitea-merger``.
|
||||
"""
|
||||
return f"{_clean(mode)} {_clean(namespace)}"
|
||||
|
||||
|
||||
def confirmation_matches(
|
||||
namespace: str, mode: str, confirmation: str | None
|
||||
) -> bool:
|
||||
"""Compare *confirmation* to the required phrase (exact, whitespace-trimmed)."""
|
||||
return _clean(confirmation) == confirmation_phrase(namespace, mode)
|
||||
|
||||
|
||||
# --- Scope validation -------------------------------------------------------
|
||||
|
||||
|
||||
def _validate_scope(namespace: str, mode: str) -> tuple[str, str] | None:
|
||||
"""Return ``(reason_code, detail)`` when the scope is refused."""
|
||||
ns = _clean(namespace)
|
||||
md = _clean(mode)
|
||||
|
||||
if md not in MODES:
|
||||
return (
|
||||
DENY_UNKNOWN_MODE,
|
||||
f"Mode {md!r} is not one of {', '.join(MODES)}.",
|
||||
)
|
||||
if ns.lower() in _FLEET_TOKENS:
|
||||
return (
|
||||
DENY_FLEET_SCOPE,
|
||||
(
|
||||
"Fleet-wide restart is an explicit non-goal: it reproduces the "
|
||||
"blast radius of `pkill -f mcp_server.py` (#630). Restart one "
|
||||
"namespace at a time."
|
||||
),
|
||||
)
|
||||
if ns not in KNOWN_NAMESPACES:
|
||||
return (
|
||||
DENY_UNKNOWN_NAMESPACE,
|
||||
f"Namespace {ns!r} is not a known MCP namespace.",
|
||||
)
|
||||
return None
|
||||
|
||||
|
||||
# --- Host hook --------------------------------------------------------------
|
||||
|
||||
|
||||
def restart_hook(env: dict[str, str] | None = None) -> dict[str, Any]:
|
||||
"""Report the configured host-managed restart hook.
|
||||
|
||||
The hook is a reference the *host* resolves (a supervisor label), not a
|
||||
command this process runs. ``configured=False`` fails restart closed.
|
||||
"""
|
||||
source = env if env is not None else os.environ
|
||||
reference = _clean(source.get(RESTART_HOOK_ENV))
|
||||
return {
|
||||
"configured": bool(reference),
|
||||
"reference": reference or None,
|
||||
"source": RESTART_HOOK_ENV if reference else None,
|
||||
"self_assertable": False,
|
||||
"console_executes_process": False,
|
||||
}
|
||||
|
||||
|
||||
# --- Preview ----------------------------------------------------------------
|
||||
|
||||
|
||||
def build_restart_preview(
|
||||
namespace: str,
|
||||
mode: str = MODE_RESTART,
|
||||
*,
|
||||
principal: console_authz.Principal | None = None,
|
||||
env: dict[str, str] | None = None,
|
||||
) -> dict[str, Any]:
|
||||
"""Render the dry-run preview for a restart/reload request.
|
||||
|
||||
Read-only: no authorization is granted, no host is contacted, and the
|
||||
result never contains a shell command.
|
||||
"""
|
||||
ns = _clean(namespace)
|
||||
md = _clean(mode)
|
||||
action_id = ACTION_FOR_MODE.get(md, ACTION_RESTART_NAMESPACE)
|
||||
action = console_authz.get_action(action_id)
|
||||
decision = console_authz.authorize(action_id, principal)
|
||||
scope_error = _validate_scope(ns, md)
|
||||
hook = restart_hook(env)
|
||||
operator = runtime_recovery_guard.operator_authorization(env)
|
||||
|
||||
return {
|
||||
"action_id": action_id,
|
||||
"namespace": ns,
|
||||
"mode": md,
|
||||
"scope_valid": scope_error is None,
|
||||
"scope_reason_code": scope_error[0] if scope_error else None,
|
||||
"scope_detail": scope_error[1] if scope_error else None,
|
||||
"required_role": action.minimum_role if action else None,
|
||||
"required_permission": action.mcp_permission if action else None,
|
||||
"action_class": action.action_class if action else None,
|
||||
"dual_control": action.dual_control if action else True,
|
||||
"break_glass": action.break_glass if action else True,
|
||||
"requires_confirmation": True,
|
||||
"confirmation_phrase": confirmation_phrase(ns, md),
|
||||
"mutation_ledger": [asdict(entry) for entry in _mutation_ledger(ns, md)],
|
||||
"authorization": decision.to_dict(),
|
||||
"operator_authorization": operator,
|
||||
"restart_hook": hook,
|
||||
"execution_enabled": False,
|
||||
"raw_process_kill_exposed": False,
|
||||
"known_namespaces": list(KNOWN_NAMESPACES),
|
||||
"post_restart_verification_required": True,
|
||||
}
|
||||
|
||||
|
||||
# --- Gate -------------------------------------------------------------------
|
||||
|
||||
|
||||
def assess_restart_request(
|
||||
namespace: str,
|
||||
mode: str = MODE_RESTART,
|
||||
*,
|
||||
principal: console_authz.Principal | None = None,
|
||||
confirmation: str | None = None,
|
||||
contamination_marker: dict[str, Any] | None = None,
|
||||
env: dict[str, str] | None = None,
|
||||
) -> dict[str, Any]:
|
||||
"""Decide whether a restart request may proceed to the host hook.
|
||||
|
||||
Every gate must pass. The first failure wins and is reported with a stable
|
||||
reason code; a pass never means "restarted", only "may be handed to the
|
||||
configured host hook".
|
||||
"""
|
||||
ns = _clean(namespace)
|
||||
md = _clean(mode)
|
||||
action_id = ACTION_FOR_MODE.get(md, ACTION_RESTART_NAMESPACE)
|
||||
preview = build_restart_preview(ns, md, principal=principal, env=env)
|
||||
|
||||
def refuse(reason_code: str, detail: str) -> dict[str, Any]:
|
||||
return {
|
||||
"allowed": False,
|
||||
"gates_passed": False,
|
||||
"reason_code": reason_code,
|
||||
"detail": detail,
|
||||
"action_id": action_id,
|
||||
"namespace": ns,
|
||||
"mode": md,
|
||||
"preview": preview,
|
||||
"execution_enabled": False,
|
||||
}
|
||||
|
||||
scope_error = _validate_scope(ns, md)
|
||||
if scope_error is not None:
|
||||
return refuse(*scope_error)
|
||||
|
||||
# Authority is checked as an authorization decision, not an execution
|
||||
# grant. ``for_execution=True`` asks "may the console perform this write?",
|
||||
# and the answer here is permanently no: step 2 of the ledger is a request
|
||||
# to the host supervisor, so the console's Phase 2 execution gate is not
|
||||
# the relevant gate. Every branch below keeps ``execution_enabled`` False
|
||||
# and :func:`execute_restart` never touches a process.
|
||||
decision = console_authz.authorize(action_id, principal)
|
||||
if not decision.allowed:
|
||||
return refuse(DENY_UNAUTHORIZED, decision.detail)
|
||||
|
||||
if not _clean(confirmation):
|
||||
return refuse(
|
||||
DENY_CONFIRMATION_MISSING,
|
||||
(
|
||||
"Type the confirmation phrase "
|
||||
f"{preview['confirmation_phrase']!r} to proceed."
|
||||
),
|
||||
)
|
||||
if not confirmation_matches(ns, md, confirmation):
|
||||
return refuse(
|
||||
DENY_CONFIRMATION_MISMATCH,
|
||||
(
|
||||
"Confirmation does not name this namespace and mode; expected "
|
||||
f"{preview['confirmation_phrase']!r}."
|
||||
),
|
||||
)
|
||||
|
||||
operator = preview["operator_authorization"]
|
||||
if not operator["authorized"]:
|
||||
return refuse(
|
||||
DENY_OPERATOR_AUTHORIZATION,
|
||||
(
|
||||
"Host daemon maintenance requires out-of-band operator "
|
||||
"authorization via "
|
||||
f"{runtime_recovery_guard.OPERATOR_AUTHORIZATION_ENV}."
|
||||
),
|
||||
)
|
||||
|
||||
# #630's task-scoped gate deliberately lets a contaminated worker keep
|
||||
# commenting and handing off. Restart is stricter and unconditional: a
|
||||
# runtime already contaminated by a manual kill must be reconciled before
|
||||
# it is restarted again, or the restart just launders the contamination.
|
||||
if contamination_marker and not contamination_marker.get(
|
||||
"cleared_by_reconciler"
|
||||
):
|
||||
return refuse(
|
||||
DENY_CONTAMINATED_RUNTIME,
|
||||
(
|
||||
"A live contamination marker is present; clear it through the "
|
||||
"reconciler path before restarting."
|
||||
),
|
||||
)
|
||||
|
||||
hook = preview["restart_hook"]
|
||||
if not hook["configured"]:
|
||||
return refuse(
|
||||
DENY_HOOK_NOT_CONFIGURED,
|
||||
(
|
||||
"No host-managed restart hook is configured "
|
||||
f"({RESTART_HOOK_ENV}). The console will not fall back to a "
|
||||
"process kill."
|
||||
),
|
||||
)
|
||||
|
||||
return {
|
||||
"allowed": True,
|
||||
"gates_passed": True,
|
||||
"reason_code": ALLOW_HOST_ACTION_REQUIRED,
|
||||
"detail": (
|
||||
"Every gate passed. The restart must be performed by the "
|
||||
"configured host supervisor; the console does not signal the "
|
||||
"process."
|
||||
),
|
||||
"action_id": action_id,
|
||||
"namespace": ns,
|
||||
"mode": md,
|
||||
"preview": preview,
|
||||
"execution_enabled": False,
|
||||
"console_executes": False,
|
||||
"console_active_phase": console_authz.ACTIVE_PHASE,
|
||||
}
|
||||
|
||||
|
||||
# --- Execution --------------------------------------------------------------
|
||||
|
||||
|
||||
def execute_restart(
|
||||
namespace: str,
|
||||
mode: str = MODE_RESTART,
|
||||
*,
|
||||
principal: console_authz.Principal | None = None,
|
||||
confirmation: str | None = None,
|
||||
contamination_marker: dict[str, Any] | None = None,
|
||||
env: dict[str, str] | None = None,
|
||||
request_id: str | None = None,
|
||||
session_id: str | None = None,
|
||||
) -> dict[str, Any]:
|
||||
"""Run every gate, audit the outcome, and hand off to the host.
|
||||
|
||||
This function never spawns a process, never sends a signal, and never
|
||||
builds a command line. ``success`` is False in both directions: a refused
|
||||
request is refused, and an authorized request still requires the host
|
||||
supervisor to act.
|
||||
"""
|
||||
assessment = assess_restart_request(
|
||||
namespace,
|
||||
mode,
|
||||
principal=principal,
|
||||
confirmation=confirmation,
|
||||
contamination_marker=contamination_marker,
|
||||
env=env,
|
||||
)
|
||||
action_id = assessment["action_id"]
|
||||
allowed = assessment["allowed"]
|
||||
|
||||
audit = console_audit.record_event(
|
||||
action_id=action_id,
|
||||
result=(
|
||||
console_audit.RESULT_ALLOWED if allowed
|
||||
else console_audit.RESULT_DENIED
|
||||
),
|
||||
principal=principal,
|
||||
target={"namespace": assessment["namespace"], "mode": assessment["mode"]},
|
||||
reason_code=assessment["reason_code"],
|
||||
detail=assessment["detail"],
|
||||
request_id=request_id,
|
||||
session_id=session_id,
|
||||
metadata={
|
||||
"gates_passed": assessment["gates_passed"],
|
||||
"process_kill_executed": False,
|
||||
"post_restart_verification_required": True,
|
||||
},
|
||||
)
|
||||
|
||||
return {
|
||||
"success": False,
|
||||
"outcome": (
|
||||
ALLOW_HOST_ACTION_REQUIRED if allowed else assessment["reason_code"]
|
||||
),
|
||||
"allowed": allowed,
|
||||
"detail": assessment["detail"],
|
||||
"namespace": assessment["namespace"],
|
||||
"mode": assessment["mode"],
|
||||
"action_id": action_id,
|
||||
"process_kill_executed": False,
|
||||
"host_hook": assessment["preview"]["restart_hook"],
|
||||
"next_action": (
|
||||
"Have the host supervisor perform the restart, then call "
|
||||
"verify_post_restart_health with live client-namespace evidence "
|
||||
"before claiming a clean session."
|
||||
if allowed
|
||||
else assessment["detail"]
|
||||
),
|
||||
"assessment": assessment,
|
||||
"audit": audit,
|
||||
}
|
||||
|
||||
|
||||
# --- Contamination classification -------------------------------------------
|
||||
|
||||
|
||||
def classify_restart_command(
|
||||
command: str | None,
|
||||
*,
|
||||
mcp_pids: list[Any] | tuple[Any, ...] | None = None,
|
||||
session_id: str | None = None,
|
||||
remote: str | None = None,
|
||||
role: str | None = None,
|
||||
) -> dict[str, Any]:
|
||||
"""Classify an operator-proposed recovery command (#630 AC2).
|
||||
|
||||
A manual ``pkill``/``kill`` of the MCP daemon is contamination, not a
|
||||
restart. When contaminating, a durable marker is returned so downstream
|
||||
gated mutations and clean claims fail closed.
|
||||
"""
|
||||
classification = runtime_recovery_guard.classify_recovery_command(
|
||||
command, mcp_pids=mcp_pids
|
||||
)
|
||||
contaminating = bool(classification.get("contamination"))
|
||||
|
||||
marker = None
|
||||
if contaminating:
|
||||
marker = runtime_recovery_guard.build_contamination_record(
|
||||
reason_class=(
|
||||
classification.get("reason_class")
|
||||
or runtime_recovery_guard.REASON_MANUAL_DAEMON_KILL
|
||||
),
|
||||
command_redacted=classification.get("redacted_command"),
|
||||
session_id=session_id,
|
||||
remote=remote,
|
||||
role=role,
|
||||
detail=(
|
||||
"Manual daemon kill is forbidden; use the sanctioned "
|
||||
f"{ACTION_RESTART_NAMESPACE} gated action instead."
|
||||
),
|
||||
)
|
||||
|
||||
return {
|
||||
"contamination": contaminating,
|
||||
"sanctioned": not contaminating and not classification.get("process_kill"),
|
||||
"clean_claim_allowed": not contaminating,
|
||||
"reason_class": classification.get("reason_class"),
|
||||
"redacted_command": classification.get("redacted_command"),
|
||||
"classification": classification,
|
||||
"contamination_marker": marker,
|
||||
"sanctioned_alternative": ACTION_RESTART_NAMESPACE,
|
||||
}
|
||||
|
||||
|
||||
# --- Post-restart health verification ---------------------------------------
|
||||
|
||||
|
||||
def verify_post_restart_health(
|
||||
namespace: str,
|
||||
*,
|
||||
probe_result: dict[str, Any] | None = None,
|
||||
probe_source: str | None = None,
|
||||
registered_tools: list[str] | tuple[str, ...] | None = None,
|
||||
required_tool: str | None = None,
|
||||
profile: str | None = None,
|
||||
) -> dict[str, Any]:
|
||||
"""Require live proof a namespace is callable before any clean claim (AC3).
|
||||
|
||||
Static registration is not proof and neither is an offline subprocess
|
||||
probe: only ``probe_source=client_namespace`` evidence can clear a
|
||||
post-restart session for mutations.
|
||||
"""
|
||||
ns = _clean(namespace)
|
||||
health = mcp_namespace_health.classify_namespace_probe(
|
||||
ns,
|
||||
required_tool=required_tool,
|
||||
registered_tools=registered_tools,
|
||||
probe_result=probe_result,
|
||||
profile=profile,
|
||||
probe_source=probe_source,
|
||||
)
|
||||
healthy = bool(health.get("healthy"))
|
||||
proven = bool(health.get("ide_namespace_proven"))
|
||||
|
||||
if healthy and proven:
|
||||
status = HEALTH_CLEAN
|
||||
elif healthy:
|
||||
status = HEALTH_UNPROVEN
|
||||
else:
|
||||
status = HEALTH_UNHEALTHY
|
||||
|
||||
reasons = list(health.get("reasons") or [])
|
||||
if status == HEALTH_UNPROVEN:
|
||||
reasons.append(
|
||||
"Namespace reported healthy without live client-namespace "
|
||||
"evidence; a post-restart clean claim requires "
|
||||
f"probe_source={mcp_namespace_health.PROBE_SOURCE_CLIENT!r}."
|
||||
)
|
||||
|
||||
return {
|
||||
"namespace": ns,
|
||||
"status": status,
|
||||
"healthy": healthy,
|
||||
"ide_namespace_proven": proven,
|
||||
"clean_claim_allowed": status == HEALTH_CLEAN,
|
||||
"mutations_allowed": status == HEALTH_CLEAN,
|
||||
"reasons": reasons,
|
||||
"health": health,
|
||||
}
|
||||
|
||||
|
||||
# --- Policy surface ---------------------------------------------------------
|
||||
|
||||
|
||||
def restart_policy() -> dict[str, Any]:
|
||||
"""Machine-readable description of the sanctioned restart contract."""
|
||||
return {
|
||||
"policy_version": 1,
|
||||
"modes": list(MODES),
|
||||
"actions": [ACTION_RESTART_NAMESPACE, ACTION_RELOAD_NAMESPACE],
|
||||
"known_namespaces": list(KNOWN_NAMESPACES),
|
||||
"fleet_scope_permitted": False,
|
||||
"console_executes_process_kill": False,
|
||||
"raw_kill_exposed": False,
|
||||
"requires_confirmation": True,
|
||||
"confirmation_binds_namespace": True,
|
||||
"operator_authorization_env": (
|
||||
runtime_recovery_guard.OPERATOR_AUTHORIZATION_ENV
|
||||
),
|
||||
"restart_hook_env": RESTART_HOOK_ENV,
|
||||
"manual_kill_classified_as": runtime_recovery_guard.CONTAMINATION_KIND,
|
||||
"post_restart_clean_claim_requires": (
|
||||
mcp_namespace_health.PROBE_SOURCE_CLIENT
|
||||
),
|
||||
"audit_required": True,
|
||||
"silent_auto_restart_permitted": False,
|
||||
}
|
||||
Reference in New Issue
Block a user