Compare commits

..
Author SHA1 Message Date
sysadmin 22e0a41bd5 feat(restart): audit lifecycle events and durable incidents (#665)
Add restart_audit with mcp.restart.* event schema, redacted emission via
gitea_audit, correlation ids, and incident materialization for failed drain
and break-glass. Wire gitea_request_mcp_restart to always audit impact
previews and fail closed on privileged apply when the audit sink is enabled
but write fails.

Closes #665
2026-07-25 17:10:16 -04:00
sysadmin 76f293eb28 Merge pull request 'fix(gate): stop classifying stale-runtime blocks as permission denials (Closes #897)' (#901) from fix/issue-897-permission-stale-runtime-classification into master 2026-07-25 06:56:16 -05:00
jcwalker3 54559aebc3 Merge branch 'master' into fix/issue-897-permission-stale-runtime-classification 2026-07-25 06:42:27 -05:00
sysadmin 715863799f Merge pull request 'feat(webui): Runtime and session view (Phase 1) (Closes #641)' (#898) from feat/issue-641-runtime-session-view into master 2026-07-25 05:50:48 -05:00
jcwalker3 daf7ed4c2b Merge branch 'master' into feat/issue-641-runtime-session-view 2026-07-25 02:41:46 -05:00
sysadmin 8598537a35 Merge pull request 'feat: enforce MCP restart class permissions' (#886) from feat/issue-663-restart-classes into master 2026-07-25 02:34:24 -05:00
sysadmin 6e6ca94338 fix(gate): stop classifying stale-runtime blocks as permission denials (Closes #897)
Stale-runtime and runtime-mode mutation refusals previously shared the
permission-denial channel, so permission_report claimed a missing op the
active profile already held and recommended gitea_activate_profile.
Typed blocker_kind payloads report reconnect-only recovery for staleness,
omit permission_report for non-permission gates, and fail closed when a
permission_report would invent a missing permission the profile holds.
2026-07-25 01:47:35 -04:00
jcwalker3 d5d121a21b Merge branch 'master' into feat/issue-641-runtime-session-view 2026-07-25 00:27:10 -05:00
sysadminandClaude Opus 4.8 1ca2b50406 fix(webui): authority-aware ownership + redacted contamination text (#641)
Addresses the two blockers from the PR #898 review at a81db754.

B1 - degraded ownership inventory was rendered as affirmative absence.
_build_session_rows read the sessions/leases/locks sections without
consulting their status, so a session row emitted lease_ids=() and
worktree_paths=() whether the session genuinely held nothing or the
lease store simply could not be read. The renderer printed both as
"none" and "unbound", contradicting the ownership_authority_complete
invariant documented on InventorySnapshot.

SessionRow now carries lease_authority and worktree_authority. A
worktree binding is correlated through lease work numbers, so it is
unproven when either the leases or the locks section fails to read --
this covers the narrow variant where locks hold real worktree paths but
a degraded leases section leaves work_numbers empty. The renderer emits
"unknown (inventory <status>)" with an authority-unproven badge instead
of none/unbound, the card names the unreadable sections, and an empty
session list from an unreadable sessions section no longer reads as
"no sessions recorded". snapshot_to_dict exports
ownership_authority_complete, ownership_section_status, and per-row
lease_authority / worktree_authority so /api/sessions consumers can
distinguish the two cases.

B2 - contamination payload strings bypassed redaction.
_inspect_contamination copied command_summary, session_id, role and
reason_class out of the marker payload with only str(), while every
inventory-sourced field on the same page arrives through
webui.inventory.scrub(). The write-time redactor
stable_branch_push_guard.redact_command is a narrow denylist that leaves
absolute $HOME paths, -H 'X-Api-Key: <value>', --password <value>, and
PRIVATE_KEY=<value> intact, and this is the first web surface to render
command_summary at all.

Adds webui.inventory.scrub_text(), which collapses $HOME and redacts
credential-shaped tokens and URL userinfo anywhere inside a string rather
than only at its start, and routes the marker payload through it. scrub()
and every existing caller are untouched. The command_summary field is
kept: it is the #630 evidence naming which daemon was killed. The module
docstring claiming absolute paths were already collapsed is corrected.

Also: removes the locks_by_session_hint dead loop and its discard (N1),
adds the missing trailing newline to webui/runtime_views.py (N4), drops
an unused dataclasses.field import, and documents both honesty rules in
docs/webui-local-dev.md.

Tests: tests/test_webui_sessions_view.py grows from 12 to 26 cases,
covering degraded and unavailable ownership sections in both the HTML and
JSON paths, the locks-readable/leases-degraded variant, a guard against
over-correcting clean inventory into "unknown", the previously untested
expired-lease flag, HTML escaping of hostile values in clean and degraded
renders, and each secret class the write-time denylist misses. The
STATUS_UNAVAILABLE import that was present but unused is now exercised.

Validation, from the issue worktree with venv/bin/python (Python 3.14.5,
pytest 9.1.1):

  pytest tests/test_webui_sessions_view.py tests/test_webui_*.py -q
    -> 512 passed, 376 subtests (was 498 / 372; +14 new tests)
  pytest tests/test_issue_854_semantic_container_exclusion.py -q
    -> 13 passed, 8 subtests
  13-file runtime/health/inventory/restart set
    -> 1 failed, 223 passed; the single failure is
       test_runtime_clarity.py::TestRuntimeClarity::
       test_activate_profile_succeeds_when_enabled, the identical test and
       assertion the reviewer recorded on master at 7af40fb5, so it is
       baseline-equivalent and not introduced here.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-07-25 01:25:28 -04:00
jcwalker3 a81db75402 Merge branch 'master' into feat/issue-641-runtime-session-view 2026-07-24 22:34:40 -05:00
sysadminandClaude Opus 4.8 619f679077 feat(webui): Runtime and session view (Phase 1) (Closes #641)
Compose runtime health with inventory sessions/namespaces/worktrees into a
live /sessions page and JSON API. Surface stale PID/lease flags and durable
contamination markers when detectable. Recovery links name sanctioned
reconnect/restart paths only — no kill controls.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-07-24 22:45:27 -04:00
15 changed files with 3520 additions and 86 deletions
+93
View File
@@ -0,0 +1,93 @@
# MCP restart audit events and incidents (#665)
Restarts and recovery attempts leave a forensic trail. Failed drains and
break-glass paths also raise durable Gitea incident issues so unsafe restarts
cannot be silently repeated.
Parent umbrella: **#655**. Related: impact coordinator **#658**, drain proof
**#661**, restart classes **#663**, break-glass **#664**, post-restart reconcile
**#662**, vision **#652**, roadmap **#653**, console recovery **#642**.
## Components
| Piece | Where | Responsibility |
|-------|-------|----------------|
| Event schema + emission | `restart_audit.py` | `mcp.restart.*` vocabulary, redacted payload builder, append-only sink via `gitea_audit` |
| Fail-closed privileged gate | `restart_audit.require_audit_or_deny` | When `GITEA_AUDIT_LOG` is set and the write fails, privileged apply is denied |
| Incident descriptors | `restart_audit.build_incident_descriptor` | Durable follow-up issues (failed drain, break-glass, reconcile unresolved, unguarded) |
| Materializer | `restart_audit.materialize_incident` | Injected `create_issue_fn` (network kept out of pure tests) |
| Wiring | `gitea_request_mcp_restart` | Correlation id, impact-preview audit, apply-gate / break-glass audit + incident creation |
## Event vocabulary
| Event type | When |
|------------|------|
| `mcp.restart.impact_preview` | Every `gitea_request_mcp_restart` evaluation |
| `mcp.restart.drain_enter` | Drain window starts (schema reserved; emit from drain path) |
| `mcp.restart.drain_exit` | Drain window ends |
| `mcp.restart.drain_proof` | Drain-proof verification result |
| `mcp.restart.apply_gate` | Apply hard gate (`dry_run=False`) |
| `mcp.restart.break_glass` | Authorized break-glass bypass |
| `mcp.restart.post_restart_reconcile` | Post-restart reconcile outcome |
| `mcp.restart.narrower_recovery` | Narrower recovery attempt recorded |
| `mcp.restart.unguarded_detected` | Unguarded restart path detected |
All free text is redacted before sink write or issue body assembly. Emission
never raises; callers decide fail-closed policy.
## Correlation
Each restart lifecycle mints a short `correlation_id` (`rst-` + 16 hex) shared
across impact preview → apply gate → incident descriptors so operators can join
the trail.
## Privileged deny-on-audit-fail
Rollout policy (issue #665):
1. Configure `GITEA_AUDIT_LOG` so writes land.
2. Only then enforce deny when a privileged restart path cannot audit.
When audit is **not** configured, privileged apply still proceeds (no false
denials during rollout). When audit **is** configured and the write fails,
`apply_authorized` is cleared.
## Incidents
| Kind | Trigger |
|------|---------|
| `restart_failed_drain` | Apply denied by drain hard gate / failed proof |
| `restart_break_glass` | Any authorized break-glass apply |
| `restart_reconcile_unresolved` | Post-restart reconcile left work unresolved |
| `restart_unguarded_detected` | Unguarded restart attempt detected |
Break-glass **always** creates an incident descriptor (and a Gitea issue when
the create path is available). Failed drain does the same. Incident bodies
include correlation id, session, class, scope, proof id, and redacted reasons.
Default labels: `mcp-health`, `safety`, `observability`, `status:ready`,
`type:bug`, `workflow-hardening`.
## Tool payload surface
`gitea_request_mcp_restart` returns:
* `correlation_id` — lifecycle join key
* `restart_audit.impact_preview_written` — sink success for the preview event
* `restart_audit.apply_gate_written` — sink success for apply/break-glass (apply only)
* `restart_audit.incident_result` — materialization outcome when an incident was required
* `incident` — durable descriptor (when gate requires follow-up)
## Security
* No secrets in audit payloads or issue bodies.
* This module never restarts a process.
* Drain proof verification remains #661; audit only records the decision.
* Incident creation failures are recorded in `incident_result.reasons` and never
crash the restart evaluation path (audit write failure still fails closed for
privileged apply when the sink is enabled).
## Tests
See `tests/test_restart_audit.py`: schema, redaction, emission, deny policy,
incident materialization mocks, break-glass / failed-drain selection.
+1
View File
@@ -33,6 +33,7 @@ recovery behavior for all nine classes.
| `ControlPlaneDB.list_sessions` | `control_plane_db.py` | Read-only session inventory (the process-level unit a restart kills). |
| `gitea_request_mcp_restart` | `gitea_mcp_server.py` | MCP tool: gathers inventory from the #613 DB, calls the coordinator, returns the report, and on `dry_run=False` runs the #661 drain-proof hard gate. Never restarts a process. |
| `drain_proof.gate_apply_restart` | `drain_proof.py` | The #661 hard gate: verifies a drain proof against the current impact fingerprint, or records an authorized break-glass bypass. |
| `restart_audit` | `restart_audit.py` | #665 forensic trail: `mcp.restart.*` events via `gitea_audit`, correlation ids, durable incidents for failed drain / break-glass. See [`mcp-restart-audit.md`](./mcp-restart-audit.md). |
## Dimensions evaluated
+43 -6
View File
@@ -77,7 +77,9 @@ status, onboarding checklist state, and the fail-closed error payloads (#635).
| `/api/actions/{id}/preview` | Mutation ledger preview (GET, read-only) |
| `/leases` | Lease and collision visibility (#433) |
| `/api/leases` | JSON lease/collision export |
| `/sessions` | Phase 1 shell stub — session inventory (backed by #636) |
| `/sessions` | Runtime and session view (#641) — health + inventory sessions/namespaces/worktrees |
| `/api/sessions` | JSON export for the runtime/session view |
| `/api/v1/sessions` | Versioned alias of `/api/sessions` |
| `/inventory` | Phase 1 shell stub — unified inventory (backed by #636) |
| `/timeline` | Phase 1 shell stub — workflow event timeline |
| `/policy` | Phase 1 shell stub — capability/role policy placeholder |
@@ -284,11 +286,46 @@ The header carries two read-only status badges — an **environment** badge
a **mode: read-only** badge — plus a **Docs** link to this document. No
privileged action controls are present in the Phase 1 shell.
Not-yet-implemented surfaces (`/sessions`, `/inventory`, `/timeline`,
`/policy`, `/insights`) resolve to graceful read-only stub pages instead of
404s; their backing views land in later child issues of #631 (the inventory
surfaces are backed by #636). Mutating methods on stub routes still fail closed
with `read-only-mvp`.
Not-yet-implemented surfaces (`/inventory`, `/timeline`, `/policy`,
`/insights`) resolve to graceful read-only stub pages instead of 404s; their
backing views land in later child issues of #631 (the inventory surfaces are
backed by #636). Mutating methods on stub routes still fail closed with
`read-only-mvp`.
### Runtime and sessions (#641)
`/sessions` is a live Phase 1 read-only view that composes:
* runtime health from `#430` (profile, role, identity, master parity, stale warning)
* control-plane sessions / leases and filesystem locks / worktrees / namespaces from `#636`
* durable contamination markers when detectable (`#630` runtime recovery, `#671` stable-branch push)
It surfaces stale indicators (dead PID, expired lease) and never silences an
active contamination marker. Recovery links point only at sanctioned
reconnect/operator restart docs (`docs/mcp-namespace-eof-recovery.md`,
`docs/mcp-namespace-health.md`, `docs/mcp-restart-path-inventory.md`, this
document). The page does **not** restart, kill, or take over sessions; manual
`pkill` of MCP daemons is contamination, not recovery.
Honesty rules specific to this view:
* **Ownership columns never assert absence they cannot prove.** When the
`leases` or `locks` section is degraded or unavailable, the Leases and
Worktree-binding cells render `unknown (inventory <status>)` with an
*authority unproven* badge instead of `none` / `unbound`, and a caveat names
the unreadable sections. A worktree binding is correlated through lease work
numbers, so it is unproven when *either* section fails to read.
`/api/sessions` carries the same facts as `ownership_authority_complete`,
`ownership_section_status`, and per-row `lease_authority` /
`worktree_authority`, so a JSON consumer can tell "holds none" from "could
not be read".
* **Contamination text is redacted at the display boundary.** Marker payloads
(`command_summary`, `reason_class`, `session_id`, `role`) are
operator-supplied free text that does not arrive through inventory scrubbing,
so they pass through `webui.inventory.scrub_text`, which collapses `$HOME` and
redacts credential-shaped tokens and URL userinfo *anywhere* in the string.
The write-time redactor is a narrow denylist and is not relied on. The field
itself is kept — it is the `#630` evidence naming which daemon was killed.
## System-health dashboard (#639)
+409 -70
View File
@@ -2069,6 +2069,7 @@ import lease_policy # noqa: E402
import workflow_dashboard # noqa: E402 # #605 live queue/lease dashboard
import restart_coordinator # noqa: E402 # #658 MCP restart coordinator/impact
import drain_proof # noqa: E402 # #661 pre-restart drain proof and hard gate
import restart_audit # noqa: E402 # #665 restart audit events + incidents
import incident_bridge # noqa: E402
import sentry_observability # noqa: E402 (#606 optional Sentry observability)
import sentry_incident_bridge # noqa: E402 (#607 Sentry→Gitea incident bridge)
@@ -9552,15 +9553,13 @@ def gitea_edit_pr(
if closing:
gate_reasons = _profile_operation_gate("gitea.pr.close")
if gate_reasons:
return {
"success": False,
"performed": False,
"pr_number": pr_number,
"requested_state": "closed",
"required_permission": "gitea.pr.close",
"reasons": gate_reasons,
"permission_report": _permission_block_report("gitea.pr.close"),
}
return _build_operation_gate_refusal(
"gitea.pr.close",
gate_reasons,
pr_number=pr_number,
requested_state="closed",
required_permission="gitea.pr.close",
)
h, o, r = _resolve(remote, host, org, repo)
auth = _auth(h)
@@ -13823,13 +13822,17 @@ def gitea_view_issue(
def _permission_block_report(required_operation: str,
identity: str | None = None) -> dict:
"""Structured, LLM-safe explanation of a permission denial (#142).
"""Structured, LLM-safe explanation of a permission denial (#142, #897).
Built only after a gate has already refused; it adds guidance to the
refusal and never widens any permission, performs network I/O, or
raises (fail-soft: degrades to a minimal fail-closed report). Names
configured profiles only never auth references, tokens, endpoint
URLs, or keychain IDs.
#897: never fabricate a missing permission when the active profile
already allows the operation. That path is a diagnostic defect (the
refusal was not a permission denial), not a cue to switch profiles.
"""
report = {
"requested_operation": required_operation,
@@ -13841,6 +13844,7 @@ def _permission_block_report(required_operation: str,
"matching_configured_profiles": [],
"runtime_switching_supported": False,
"different_mcp_namespace_required": True,
"diagnostic_defect": False,
"exact_safe_next_action": (
"Ask the operator to fix GITEA_MCP_CONFIG/GITEA_MCP_PROFILE; "
"the active profile could not be resolved (fail closed)."),
@@ -13855,6 +13859,32 @@ def _permission_block_report(required_operation: str,
report["active_allowed_operations"] = (
profile.get("allowed_operations") or [])
# #897: fail closed as a diagnostic defect when the active profile
# already holds the operation — callers must not invent "missing".
try:
holds, _hold_reason = gitea_config.check_operation(
required_operation,
profile.get("allowed_operations") or [],
profile.get("forbidden_operations") or [],
)
except Exception:
holds = False
if holds:
report["missing_permission"] = None
report["required_permission"] = required_operation
report["diagnostic_defect"] = True
report["different_mcp_namespace_required"] = False
report["exact_safe_next_action"] = (
"Diagnostic defect: the active profile already allows "
f"{required_operation}. This is not a permission denial — "
"inspect blocker_kind / reasons (stale-runtime or runtime-mode). "
"Do not call gitea_activate_profile or switch MCP sessions."
)
report["matching_configured_profiles"] = [
p for p in [profile.get("profile_name")] if p
]
return report
matching = []
try:
config = gitea_config.load_config() or {}
@@ -13903,6 +13933,205 @@ def _permission_block_report(required_operation: str,
return report
def _reason_is_stale_runtime(reason: str) -> bool:
"""True when *reason* is a master-parity / stale-daemon refusal (#897)."""
r = (reason or "").lower()
if not r:
return False
if "stale relative to live master" in r:
return True
if "server code is stale" in r:
return True
if "daemon is stale" in r:
return True
if "started at commit" in r and "workspace master is now" in r:
return True
if "mcp server started at" in r and "stale" in r:
return True
if "restart the server to load the current capability gates" in r:
return True
if "restart/reconnect before mutating" in r:
return True
return False
def _reason_is_runtime_mode(reason: str) -> bool:
"""True when *reason* is a stable-control / runtime-mode refusal (#897)."""
r = (reason or "").lower()
if not r:
return False
if _reason_is_stale_runtime(reason):
return False
if "runtime mode could not be assessed" in r:
return True
if "runtime mode is" in r:
return True
if "stable control runtime" in r:
return True
if "dev-test" in r and ("runtime" in r or "production" in r):
return True
if "development worktree" in r or "dev worktree" in r:
return True
if "launched from a 'branches/" in r or "launched from a \"branches/" in r:
return True
if "process-root / active-workspace alignment" in r:
return True
if "namespace" in r and "reproof" in r:
return True
return False
def _reason_is_permission(reason: str) -> bool:
"""True when *reason* is a genuine profile-permission denial (#897)."""
r = (reason or "").lower()
if not r:
return False
if _reason_is_stale_runtime(reason) or _reason_is_runtime_mode(reason):
return False
if "profile could not be resolved" in r:
return True
if "profile has no configured allowed operations" in r:
return True
if "profile forbids" in r:
return True
if "profile is not allowed to" in r:
return True
if "unrecognized forbidden operation" in r:
return True
return False
def _classify_operation_gate_reasons(reasons: list[str]) -> dict:
"""Partition gate reasons into stale / runtime-mode / permission (#897)."""
stale: list[str] = []
runtime_mode: list[str] = []
permission: list[str] = []
other: list[str] = []
for reason in reasons or []:
if _reason_is_stale_runtime(reason):
stale.append(reason)
elif _reason_is_runtime_mode(reason):
runtime_mode.append(reason)
elif _reason_is_permission(reason):
permission.append(reason)
else:
other.append(reason)
return {
"stale_runtime": stale,
"runtime_mode": runtime_mode,
"permission": permission,
"other": other,
}
def _stale_runtime_reconnect_action() -> str:
"""Sanctioned recovery for a stale daemon — reconnect only (#685/#897)."""
return (
"Reconnect the IDE/client MCP session so the server reloads at the "
"current master head. Do not call gitea_activate_profile or switch "
"MCP role sessions — profile switching does not clear a stale daemon."
)
def _build_operation_gate_refusal(
required_operation: str,
reasons: list[str],
**extra_fields,
) -> dict:
"""Structured gate refusal with typed blockers (#897).
Stale-runtime and runtime-mode refusals never attach a
``permission_report`` and never recommend profile switching. True
permission denials still get ``permission_report``. When both apply,
causes are reported separately under distinct fields.
"""
classified = _classify_operation_gate_reasons(reasons)
stale = classified["stale_runtime"]
runtime_mode = classified["runtime_mode"]
permission = classified["permission"]
other = classified["other"]
blocked: dict = {
"success": False,
"performed": False,
"reasons": list(reasons),
"mutation_performed": False,
"session_context_audit": session_ctx.mutation_context_audit_fields(),
"gate_reason_classes": {
"stale_runtime": list(stale),
"runtime_mode": list(runtime_mode),
"permission": list(permission),
"other": list(other),
},
}
if stale:
parity = _current_master_parity()
blocked["blocker_kind"] = "runtime_reconnect_required"
blocked["restart_required"] = True
blocked["stop_required"] = True
blocked["startup_head"] = parity.get("startup_head")
blocked["current_head"] = parity.get("current_head")
blocked["daemon_start_head"] = (
parity.get("daemon_start_head") or parity.get("startup_head")
)
blocked["local_head"] = (
parity.get("local_head") or parity.get("current_head")
)
blocked["live_remote_head"] = parity.get("live_remote_head")
blocked["live_stale"] = bool(parity.get("live_stale"))
blocked["live_known"] = bool(parity.get("live_known"))
blocked["exact_safe_next_action"] = _stale_runtime_reconnect_action()
if permission or other:
blocked["permission_block_reasons"] = list(permission) + list(other)
blocked["stale_runtime_reasons"] = list(stale)
# Never attach permission_report for a staleness refusal.
blocked.update(extra_fields)
return blocked
if runtime_mode:
blocked["blocker_kind"] = "runtime_mode_blocked"
blocked["restart_required"] = False
blocked["stop_required"] = True
blocked["exact_safe_next_action"] = (
"Real workflow mutations run only on the promoted stable control "
"runtime. Promote/reload the stable runtime; do not call "
"gitea_activate_profile or switch MCP role sessions to clear a "
"runtime-mode block."
)
if permission or other:
blocked["permission_block_reasons"] = list(permission) + list(other)
blocked["runtime_mode_reasons"] = list(runtime_mode)
blocked.update(extra_fields)
return blocked
# Pure permission (or unclassified-as-permission) denial.
blocked["blocker_kind"] = "permission_denied"
blocked["permission_report"] = _permission_block_report(required_operation)
blocked.update(extra_fields)
return blocked
def _permission_report_for_gate_reasons(
required_operation: str,
reasons: list[str] | None,
) -> dict | None:
"""Attach ``permission_report`` only for true permission denials (#897).
Call sites that historically always attached a permission report after
``_profile_operation_gate`` should use this so stale/runtime refusals
do not emit a fabricated missing-permission payload.
"""
if not reasons:
return None
classified = _classify_operation_gate_reasons(reasons)
if classified["stale_runtime"] or classified["runtime_mode"]:
return None
if not (classified["permission"] or classified["other"]):
return None
return _permission_block_report(required_operation)
def _role_for_operation(op: str) -> str | None:
# Normalize op first
try:
@@ -14079,7 +14308,7 @@ def _master_parity_block(op: str) -> list[str]:
def _profile_operation_gate(op: str) -> list[str]:
"""Profile permission check for a single gated operation (#126, #216, #420).
"""Profile permission check for a single gated operation (#126, #216, #420, #897).
Issue discussion comments are gated separately from the gitea.pr.*
review/merge family: listing requires ``gitea.read``, creating requires
@@ -14092,21 +14321,26 @@ def _profile_operation_gate(op: str) -> list[str]:
capability gate that has since been merged, and when the runtime itself is
not the promoted stable control runtime (#615) -- a dev/test or unknown
runtime holds production credentials but has not been promoted.
#897: collect *all* independent refusal classes (stale, runtime-mode,
permission) rather than short-circuiting after the first. Callers that
only need a boolean still treat any non-empty list as blocked; typed
consumers (``_build_operation_gate_refusal``) can separate causes.
"""
stale_reasons = _master_parity_block(op)
if stale_reasons:
return stale_reasons
runtime_reasons = _runtime_mode_block(op)
if runtime_reasons:
return runtime_reasons
reasons: list[str] = []
reasons.extend(_master_parity_block(op))
reasons.extend(_runtime_mode_block(op))
try:
profile = get_profile()
except Exception as exc:
return [f"profile could not be resolved (fail closed): {_redact(str(exc))}"]
reasons.append(
f"profile could not be resolved (fail closed): {_redact(str(exc))}"
)
return reasons
op_ok, op_reason = gitea_config.check_operation(
op, profile["allowed_operations"], profile["forbidden_operations"])
if op_ok:
return []
return reasons
if _try_auto_switch_for_operation(op):
try:
@@ -14114,17 +14348,26 @@ def _profile_operation_gate(op: str) -> list[str]:
op_ok, op_reason = gitea_config.check_operation(
op, profile["allowed_operations"], profile["forbidden_operations"])
if op_ok:
return []
return reasons
except Exception as exc:
return [f"profile could not be resolved (fail closed): {_redact(str(exc))}"]
reasons.append(
f"profile could not be resolved (fail closed): {_redact(str(exc))}"
)
return reasons
if op_reason == "no-allowed-operations":
return ["profile has no configured allowed operations (fail closed)"]
if op_reason == "forbidden":
return [f"profile forbids '{op}'"]
if op_reason == "invalid-forbidden-entry":
return ["profile has an unrecognized forbidden operation entry (fail closed)"]
return [f"profile is not allowed to {op}"]
reasons.append(
"profile has no configured allowed operations (fail closed)"
)
elif op_reason == "forbidden":
reasons.append(f"profile forbids '{op}'")
elif op_reason == "invalid-forbidden-entry":
reasons.append(
"profile has an unrecognized forbidden operation entry (fail closed)"
)
else:
reasons.append(f"profile is not allowed to {op}")
return reasons
def _mutation_config_authority_block(required_operation: str) -> dict | None:
@@ -14344,10 +14587,14 @@ def _session_context_mutation_block(
def _profile_permission_block(required_operation: str, **extra_fields) -> dict | None:
"""Structured permission denial for gated tools (#69, #142).
"""Structured operation-gate denial for gated tools (#69, #142, #897).
Returns a block dict when the active profile forbids *required_operation*,
or ``None`` when the gate passes. Never performs network I/O.
the daemon is stale, or the runtime mode is not mutation-safe or
``None`` when the gate passes. Never performs network I/O.
#897: stale-runtime and runtime-mode refusals are typed
(``blocker_kind``) and never carry a ``permission_report``.
"""
req_role = "reviewer" if any(required_operation.startswith(p) for p in (
"gitea.pr.approve", "gitea.pr.merge", "gitea.pr.request_changes", "gitea.pr.review"
@@ -14357,15 +14604,9 @@ def _profile_permission_block(required_operation: str, **extra_fields) -> dict |
reasons = _profile_operation_gate(required_operation)
if reasons:
blocked = {
"success": False,
"performed": False,
"reasons": reasons,
"permission_report": _permission_block_report(required_operation),
"session_context_audit": session_ctx.mutation_context_audit_fields(),
}
blocked.update(extra_fields)
return blocked
return _build_operation_gate_refusal(
required_operation, reasons, **extra_fields
)
auth_block = _mutation_config_authority_block(required_operation)
if auth_block is not None:
@@ -14494,20 +14735,14 @@ def gitea_acquire_reviewer_pr_lease(
"""Acquire a per-PR reviewer lease before review/merge mutations (#407)."""
read_block = _profile_operation_gate("gitea.read")
if read_block:
return {
"success": False,
"acquired": False,
"reasons": read_block,
"permission_report": _permission_block_report("gitea.read"),
}
return _build_operation_gate_refusal(
"gitea.read", read_block, acquired=False
)
comment_block = _profile_operation_gate("gitea.pr.comment")
if comment_block:
return {
"success": False,
"acquired": False,
"reasons": comment_block,
"permission_report": _permission_block_report("gitea.pr.comment"),
}
return _build_operation_gate_refusal(
"gitea.pr.comment", comment_block, acquired=False
)
# task=acquire_reviewer_pr_lease so verify_preflight_purity runs shared #604
# anti-stomp for the declared lease-acquire mutation inventory entry.
@@ -14623,20 +14858,14 @@ def gitea_acquire_merger_pr_lease(
"""
read_block = _profile_operation_gate("gitea.read")
if read_block:
return {
"success": False,
"acquired": False,
"reasons": read_block,
"permission_report": _permission_block_report("gitea.read"),
}
return _build_operation_gate_refusal(
"gitea.read", read_block, acquired=False
)
comment_block = _profile_operation_gate("gitea.pr.comment")
if comment_block:
return {
"success": False,
"acquired": False,
"reasons": comment_block,
"permission_report": _permission_block_report("gitea.pr.comment"),
}
return _build_operation_gate_refusal(
"gitea.pr.comment", comment_block, acquired=False
)
merge_block = _profile_operation_gate("gitea.pr.merge")
if merge_block:
return {
@@ -19375,14 +19604,12 @@ def gitea_update_pr_branch_by_merge(
# Permission: author branch push / PR mutation surface.
push_block = _profile_operation_gate("gitea.branch.push")
if push_block:
return {
"success": False,
"performed": False,
"mutation_allowed": False,
"reasons": push_block,
"permission_report": _permission_block_report("gitea.branch.push"),
"role_kind": role,
}
return _build_operation_gate_refusal(
"gitea.branch.push",
push_block,
mutation_allowed=False,
role_kind=role,
)
if role != "author":
pre = pr_sync_status.assess_update_pr_branch_preflight(
@@ -22524,6 +22751,38 @@ def gitea_request_mcp_restart(
# and a durable incident is raised. Break-glass is the only bypass and its
# authorization is read from the environment, never self-asserted.
payload["apply_supported"] = False
# #665: correlation id threads impact preview → apply gate → incidents.
correlation_id = restart_audit.new_correlation_id()
payload["correlation_id"] = correlation_id
auth_user = None
try:
# remote-only: host override is for operator diagnostics, not required here
auth_user = (gitea_whoami(remote=remote) or {}).get("username")
except Exception: # noqa: BLE001 — identity is best-effort for audit
auth_user = None
preview_audit = restart_audit.record_restart_lifecycle(
event_type=restart_audit.EVENT_IMPACT_PREVIEW,
outcome=str(report.verdict or "unknown"),
correlation_id=correlation_id,
remote=remote,
org=o,
repo=r,
requesting_session_id=sid,
restart_class=restart_class,
profile_name=profile_name,
authenticated_username=auth_user,
reasons=list(report.reasons or []),
details={
"dry_run": True,
"allow_restart": bool(report.allow_restart),
"inventory_complete": inventory_complete,
},
privileged=False,
)
payload["restart_audit"] = {
"correlation_id": correlation_id,
"impact_preview_written": preview_audit["audit_written"],
}
if not dry_run:
proof_obj: dict | None = None
proof_parse_error: str | None = None
@@ -22583,6 +22842,86 @@ def gitea_request_mcp_restart(
)
if not gate.allow and gate.incident is not None:
payload["incident"] = gate.incident
# #665: audit apply-gate + materialize durable incidents for failed
# drain and break-glass. Privileged apply denies if audit is enabled
# and the sink write fails.
incident_desc = restart_audit.incident_from_apply_gate(
gate_payload={
**gate_payload,
"incident": gate.incident,
"allow": gate.allow,
},
break_glass=break_glass,
correlation_id=correlation_id,
requesting_session_id=sid,
restart_class=restart_class,
remote=remote,
org=o,
repo=r,
)
if incident_desc is not None:
payload["incident"] = incident_desc
def _create_restart_incident_issue(
*, title, body, labels, org=None, repo=None, **_kw
):
return gitea_create_issue(
title=title,
body=body,
labels=labels,
remote=remote,
host=h,
org=org or o,
repo=repo or r,
)
apply_audit = restart_audit.record_restart_lifecycle(
event_type=(
restart_audit.EVENT_BREAK_GLASS
if break_glass
else restart_audit.EVENT_APPLY_GATE
),
outcome=(
"break_glass"
if break_glass
else ("allow" if payload["apply_authorized"] else "deny")
),
correlation_id=correlation_id,
remote=remote,
org=o,
repo=r,
requesting_session_id=sid,
restart_class=restart_class,
profile_name=profile_name,
authenticated_username=auth_user,
reasons=list(gate_payload.get("reasons") or []),
details={
"apply_authorized": payload["apply_authorized"],
"drain_gate_allow": gate_payload.get("drain_gate_allow"),
"restart_class_authorized": restart_class_authorized,
"break_glass": break_glass,
"proof_id": gate_payload.get("proof_id"),
},
privileged=True,
create_incident=incident_desc,
create_issue_fn=_create_restart_incident_issue
if incident_desc is not None
else None,
dry_run_incident=False,
)
payload["restart_audit"] = {
"correlation_id": correlation_id,
"impact_preview_written": preview_audit["audit_written"],
"apply_gate_written": apply_audit["audit_written"],
"incident_result": apply_audit.get("incident_result"),
}
if apply_audit["deny_reasons"]:
payload["apply_authorized"] = False
payload["reasons"] = list(payload.get("reasons") or []) + list(
apply_audit["deny_reasons"]
)
payload["success"] = True
return payload
+426
View File
@@ -0,0 +1,426 @@
"""MCP restart lifecycle audit events and incident materialization (#665).
Restarts and recovery attempts must leave a forensic trail: impact previews,
drain enter/exit, drain-proof results, apply gate verdicts, break-glass, and
post-restart reconcile outcomes. Failed drains and break-glass must also raise
durable Gitea incident issues so they cannot be silently repeated.
This module is the pure + sink layer for that trail:
* **Schema** — ``mcp.restart.*`` event names and a redacted payload builder.
* **Emission** — append-only via :mod:`gitea_audit` (off when ``GITEA_AUDIT_LOG``
is unset; privileged apply can still *require* a successful write).
* **Incidents** — descriptors for failed drain / break-glass / unguarded restart,
plus an optional materializer that creates a Gitea issue through an injected
``create_issue_fn`` (keeps this module free of network I/O in tests).
Design rules:
* **No secrets.** All free text is redacted before write or issue body assembly.
* **Never raises from emission.** ``emit_restart_event`` returns False on sink
failure so callers can decide fail-closed policy for privileged restarts.
* **Does not restart.** Audit never executes a process restart.
* **Drain proof stays #661.** This module records what the gate decided; it
does not re-verify proofs.
"""
from __future__ import annotations
import uuid
from datetime import datetime, timezone
from typing import Any, Callable, Mapping, Sequence
import gitea_audit
# ── Event vocabulary (stable identifiers for operators + tests) ───────────────
EVENT_IMPACT_PREVIEW = "mcp.restart.impact_preview"
EVENT_DRAIN_ENTER = "mcp.restart.drain_enter"
EVENT_DRAIN_EXIT = "mcp.restart.drain_exit"
EVENT_DRAIN_PROOF = "mcp.restart.drain_proof"
EVENT_APPLY_GATE = "mcp.restart.apply_gate"
EVENT_BREAK_GLASS = "mcp.restart.break_glass"
EVENT_POST_RESTART_RECONCILE = "mcp.restart.post_restart_reconcile"
EVENT_NARROWER_RECOVERY = "mcp.restart.narrower_recovery"
EVENT_UNGUARDED_DETECTED = "mcp.restart.unguarded_detected"
RESTART_EVENT_TYPES: frozenset[str] = frozenset(
{
EVENT_IMPACT_PREVIEW,
EVENT_DRAIN_ENTER,
EVENT_DRAIN_EXIT,
EVENT_DRAIN_PROOF,
EVENT_APPLY_GATE,
EVENT_BREAK_GLASS,
EVENT_POST_RESTART_RECONCILE,
EVENT_NARROWER_RECOVERY,
EVENT_UNGUARDED_DETECTED,
}
)
# Incident kinds (durable Gitea issues).
INCIDENT_FAILED_DRAIN = "restart_failed_drain"
INCIDENT_BREAK_GLASS = "restart_break_glass"
INCIDENT_RECONCILE_UNRESOLVED = "restart_reconcile_unresolved"
INCIDENT_UNGUARDED = "restart_unguarded_detected"
DEFAULT_INCIDENT_LABELS: tuple[str, ...] = (
"mcp-health",
"safety",
"observability",
"status:ready",
"type:bug",
"workflow-hardening",
)
CreateIssueFn = Callable[..., dict[str, Any]]
def _utc_now_iso() -> str:
return datetime.now(timezone.utc).isoformat()
def new_correlation_id() -> str:
"""Mint a short correlation id shared across a restart lifecycle."""
return f"rst-{uuid.uuid4().hex[:16]}"
def build_restart_event(
*,
event_type: str,
outcome: str,
correlation_id: str | None = None,
remote: str | None = None,
org: str | None = None,
repo: str | None = None,
requesting_session_id: str | None = None,
restart_class: str | None = None,
profile_name: str | None = None,
authenticated_username: str | None = None,
reasons: Sequence[str] | None = None,
details: Mapping[str, Any] | None = None,
now: str | None = None,
) -> dict[str, Any]:
"""Build a redacted ``mcp.restart.*`` audit event.
Raises ``ValueError`` on unknown event types so a typo cannot silently land
under a free-form action name.
"""
name = str(event_type or "").strip()
if name not in RESTART_EVENT_TYPES:
raise ValueError(
f"unknown restart audit event_type {name!r}; expected one of "
f"{sorted(RESTART_EVENT_TYPES)}"
)
redacted_reasons = [
gitea_audit.redact(str(r)) for r in (reasons or []) if str(r).strip()
]
redacted_details = gitea_audit.redact(dict(details or {}))
if not isinstance(redacted_details, dict):
redacted_details = {"value": redacted_details}
event = gitea_audit.build_event(
action=name,
result=str(outcome or "unknown"),
remote=remote,
repository=f"{org}/{repo}" if org and repo else None,
profile_name=profile_name,
authenticated_username=authenticated_username,
reason="; ".join(redacted_reasons) if redacted_reasons else None,
request_metadata={
"event_family": "mcp.restart",
"correlation_id": correlation_id or new_correlation_id(),
"restart_class": restart_class,
"requesting_session_id": requesting_session_id,
"org": org,
"repo": repo,
"details": redacted_details,
"reasons": redacted_reasons,
},
now=now or _utc_now_iso(),
operation=name,
)
event["action_type"] = "restart_lifecycle"
event["event_type"] = name
event["correlation_id"] = (event.get("request_metadata") or {}).get(
"correlation_id"
)
return event
def emit_restart_event(event: Mapping[str, Any], *, path: str | None = None) -> bool:
"""Append *event* to the audit sink. Never raises. Returns write success."""
try:
return bool(gitea_audit.write_event(dict(event), path=path))
except Exception:
return False
def require_audit_or_deny(
*,
privileged: bool,
written: bool,
audit_enabled: bool | None = None,
) -> list[str]:
"""Return deny reasons when a privileged restart path fails to audit.
When audit is not configured (``GITEA_AUDIT_LOG`` unset), privileged apply
still proceeds under the rollout policy "enable audit before enforcing
deny-on-audit-fail" — but *if* audit is enabled and the write fails,
privileged apply is denied (fail closed).
"""
enabled = (
gitea_audit.audit_enabled() if audit_enabled is None else bool(audit_enabled)
)
if not privileged:
return []
if not enabled:
return []
if written:
return []
return [
"privileged restart path requires a successful audit write; "
"audit sink failed (fail closed, #665)"
]
# ── Incident descriptors ──────────────────────────────────────────────────────
def build_incident_descriptor(
*,
kind: str,
reasons: Sequence[str],
correlation_id: str | None = None,
requesting_session_id: str | None = None,
restart_class: str | None = None,
remote: str | None = None,
org: str | None = None,
repo: str | None = None,
proof_id: str | None = None,
at: str | None = None,
) -> dict[str, Any]:
"""Build a durable incident descriptor (no network)."""
titles = {
INCIDENT_FAILED_DRAIN: "Restart denied: drain proof failed the hard gate",
INCIDENT_BREAK_GLASS: "Break-glass MCP restart authorized",
INCIDENT_RECONCILE_UNRESOLVED: "Post-restart reconcile left unresolved work",
INCIDENT_UNGUARDED: "Unguarded MCP restart attempt detected",
}
title = titles.get(kind, f"MCP restart incident ({kind})")
redacted_reasons = [
gitea_audit.redact(str(r)) for r in reasons if str(r).strip()
]
return {
"kind": kind,
"title": title,
"labels": list(DEFAULT_INCIDENT_LABELS),
"reasons": redacted_reasons,
"correlation_id": correlation_id,
"requesting_session_id": requesting_session_id,
"restart_class": restart_class,
"remote": remote,
"org": org,
"repo": repo,
"proof_id": proof_id,
"at": at or _utc_now_iso(),
"source": "restart_audit#665",
}
def incident_body(descriptor: Mapping[str, Any]) -> str:
"""Render a redacted markdown body for a Gitea incident issue."""
reasons = descriptor.get("reasons") or []
reason_lines = "\n".join(f"- {gitea_audit.redact(str(r))}" for r in reasons) or (
"- (no reasons recorded)"
)
return "\n".join(
[
"<!-- mcp-restart-incident:v1 -->",
f"## MCP restart incident (`{descriptor.get('kind')}`)",
"",
f"**Correlation:** `{descriptor.get('correlation_id') or 'none'}`",
f"**Session:** `{descriptor.get('requesting_session_id') or 'none'}`",
f"**Class:** `{descriptor.get('restart_class') or 'none'}`",
f"**Scope:** `{descriptor.get('remote')}/{descriptor.get('org')}/"
f"{descriptor.get('repo')}`",
f"**At:** `{descriptor.get('at')}`",
f"**Proof id:** `{descriptor.get('proof_id') or 'none'}`",
"",
"### Reasons",
reason_lines,
"",
"### Operator next steps",
"- Treat this as durable follow-up work under the restart-governance umbrella (#655).",
"- Do not invent a second restart path; use sanctioned coordinator tools only.",
"- Raw secrets must never appear in this issue (already redacted).",
"",
f"_Source: {descriptor.get('source')}_",
]
)
def materialize_incident(
descriptor: Mapping[str, Any],
*,
create_issue_fn: CreateIssueFn | None,
dry_run: bool = False,
) -> dict[str, Any]:
"""Create a Gitea issue from *descriptor* when *create_issue_fn* is provided.
Returns a result dict with ``created`` / ``issue_number`` / ``dry_run`` /
``reasons``. Never raises.
"""
base: dict[str, Any] = {
"created": False,
"dry_run": bool(dry_run),
"issue_number": None,
"kind": descriptor.get("kind"),
"reasons": [],
"descriptor": dict(descriptor),
}
if dry_run:
base["reasons"] = ["dry-run only; no Gitea issue created"]
return base
if create_issue_fn is None:
base["reasons"] = [
"create_issue_fn not provided; incident descriptor retained only"
]
return base
try:
result = create_issue_fn(
title=str(descriptor.get("title") or "MCP restart incident"),
body=incident_body(descriptor),
labels=list(descriptor.get("labels") or DEFAULT_INCIDENT_LABELS),
org=descriptor.get("org"),
repo=descriptor.get("repo"),
)
number = None
if isinstance(result, dict):
number = result.get("number") or result.get("issue_number")
if number is not None:
base["created"] = True
base["issue_number"] = int(number)
base["reasons"] = [f"created incident issue #{int(number)}"]
else:
base["reasons"] = ["create_issue_fn returned no issue number"]
except Exception as exc: # noqa: BLE001 — never break restart path here
base["reasons"] = [
f"incident issue creation failed: {gitea_audit.redact(str(exc))}"
]
return base
def record_restart_lifecycle(
*,
event_type: str,
outcome: str,
correlation_id: str,
remote: str | None = None,
org: str | None = None,
repo: str | None = None,
requesting_session_id: str | None = None,
restart_class: str | None = None,
profile_name: str | None = None,
authenticated_username: str | None = None,
reasons: Sequence[str] | None = None,
details: Mapping[str, Any] | None = None,
privileged: bool = False,
create_incident: Mapping[str, Any] | None = None,
create_issue_fn: CreateIssueFn | None = None,
dry_run_incident: bool = False,
audit_path: str | None = None,
) -> dict[str, Any]:
"""Emit one restart audit event and optionally materialize an incident.
Returns ``{event, audit_written, deny_reasons, incident_result}``.
"""
event = build_restart_event(
event_type=event_type,
outcome=outcome,
correlation_id=correlation_id,
remote=remote,
org=org,
repo=repo,
requesting_session_id=requesting_session_id,
restart_class=restart_class,
profile_name=profile_name,
authenticated_username=authenticated_username,
reasons=reasons,
details=details,
)
written = emit_restart_event(event, path=audit_path)
deny = require_audit_or_deny(privileged=privileged, written=written)
incident_result = None
if create_incident is not None:
incident_result = materialize_incident(
create_incident,
create_issue_fn=create_issue_fn,
dry_run=dry_run_incident,
)
return {
"event": event,
"audit_written": written,
"deny_reasons": deny,
"incident_result": incident_result,
"correlation_id": correlation_id,
}
def incident_from_apply_gate(
*,
gate_payload: Mapping[str, Any],
break_glass: bool,
correlation_id: str,
requesting_session_id: str | None,
restart_class: str | None,
remote: str | None,
org: str | None,
repo: str | None,
) -> dict[str, Any] | None:
"""Choose an incident descriptor from an apply-gate payload, if required."""
reasons = list(gate_payload.get("reasons") or [])
proof_id = gate_payload.get("proof_id")
if break_glass:
return build_incident_descriptor(
kind=INCIDENT_BREAK_GLASS,
reasons=reasons
or ["break-glass restart path used; durable incident required (#665)"],
correlation_id=correlation_id,
requesting_session_id=requesting_session_id,
restart_class=restart_class,
remote=remote,
org=org,
repo=repo,
proof_id=proof_id if isinstance(proof_id, str) else None,
)
# Failed drain / deny path.
incident = gate_payload.get("incident")
if isinstance(incident, Mapping) and incident:
# Normalize gate-provided descriptor into our schema.
return build_incident_descriptor(
kind=INCIDENT_FAILED_DRAIN,
reasons=list(incident.get("reasons") or reasons),
correlation_id=correlation_id,
requesting_session_id=requesting_session_id
or incident.get("requesting_session_id"),
restart_class=restart_class,
remote=remote,
org=org,
repo=repo,
proof_id=incident.get("proof_id") or proof_id,
at=incident.get("at"),
)
if not gate_payload.get("allow") and not gate_payload.get("drain_gate_allow", True):
return build_incident_descriptor(
kind=INCIDENT_FAILED_DRAIN,
reasons=reasons or ["restart apply denied"],
correlation_id=correlation_id,
requesting_session_id=requesting_session_id,
restart_class=restart_class,
remote=remote,
org=org,
repo=repo,
proof_id=proof_id if isinstance(proof_id, str) else None,
)
return None
@@ -0,0 +1,453 @@
"""#897: stale-runtime / runtime-mode refusals must not look like permission denials.
Acceptance criteria (issue #897):
* Stale-runtime and runtime-mode refusals are typed distinctly from
profile-permission refusals (distinct ``blocker_kind``).
* A refusal caused by staleness or runtime mode never emits a
``permission_report`` and never names a permission the active profile holds.
* ``_permission_block_report`` verifies the active profile actually lacks the
operation before reporting it missing.
* A stale-runtime refusal reports reconnect-only recovery and never recommends
``gitea_activate_profile`` or an MCP session switch.
* The blocker payload states the observed heads (parity fields).
* Matrix across author / reviewer / merger / reconciler profiles.
* Regression: ``gitea_create_issue`` on a stale daemon under ``prgs-author``
never returns ``missing_permission: gitea.issue.create``.
"""
from __future__ import annotations
import os
import sys
import unittest
from unittest.mock import patch
sys.path.insert(0, str(__import__("pathlib").Path(__file__).resolve().parent.parent))
import gitea_config # noqa: E402
import gitea_mcp_server as mcp_server # noqa: E402
SHA_START = "7af40fb5ff7debd5e9165fe97d9c7c279358e175"
SHA_LIVE = "2f4dec832327513118f2fe92b74da25d124a01cb"
ROLE_MATRIX = (
(
"prgs-author",
"author",
"gitea.issue.create",
[
"gitea.read",
"gitea.issue.create",
"gitea.issue.comment",
"gitea.issue.close",
"gitea.branch.create",
"gitea.branch.push",
"gitea.pr.create",
"gitea.pr.comment",
"gitea.repo.commit",
],
["gitea.pr.approve", "gitea.pr.merge", "gitea.pr.request_changes"],
"gitea.pr.merge", # forbidden op for pure-permission case
),
(
"prgs-reviewer",
"reviewer",
"gitea.pr.review",
[
"gitea.read",
"gitea.pr.review",
"gitea.pr.approve",
"gitea.pr.request_changes",
"gitea.pr.comment",
"gitea.issue.comment",
],
["gitea.branch.push", "gitea.pr.create"],
"gitea.branch.push",
),
(
"prgs-merger",
"merger",
"gitea.pr.merge",
[
"gitea.read",
"gitea.pr.merge",
"gitea.pr.comment",
"gitea.issue.comment",
],
["gitea.pr.approve", "gitea.branch.push", "gitea.pr.create"],
"gitea.branch.push",
),
(
"prgs-reconciler",
"reconciler",
"gitea.branch.delete",
[
"gitea.read",
"gitea.branch.delete",
"gitea.pr.comment",
"gitea.issue.comment",
"gitea.pr.close",
"gitea.issue.close",
],
["gitea.pr.approve", "gitea.pr.merge"],
"gitea.pr.merge",
),
)
def _profile(name: str, role: str, allowed: list[str], forbidden: list[str]) -> dict:
return {
"profile_name": name,
"role": role,
"role_kind": role,
"allowed_operations": list(allowed),
"forbidden_operations": list(forbidden),
"identity": "test-user",
}
def _config(profiles: dict) -> dict:
return {
"version": 2,
"profiles": {
name: {
"role": p["role"],
"allowed_operations": p["allowed_operations"],
"forbidden_operations": p["forbidden_operations"],
}
for name, p in profiles.items()
},
"rules": {"allow_runtime_switching": True},
}
class Issue897Helpers(unittest.TestCase):
def test_classify_stale_reason_strings(self):
stale = (
f"live remote master is {SHA_LIVE[:12]} but the MCP server started "
f"at {SHA_START[:12]}; the daemon is stale relative to live master "
"-- restart/reconnect before mutating"
)
classified = mcp_server._classify_operation_gate_reasons([stale])
self.assertEqual(classified["stale_runtime"], [stale])
self.assertEqual(classified["permission"], [])
self.assertEqual(classified["runtime_mode"], [])
def test_classify_permission_reason(self):
reason = "profile is not allowed to gitea.pr.merge"
classified = mcp_server._classify_operation_gate_reasons([reason])
self.assertEqual(classified["permission"], [reason])
self.assertEqual(classified["stale_runtime"], [])
def test_classify_runtime_mode_reason(self):
reason = (
"runtime mode is 'dev-test' and the mutation targets the "
"production repository; dev/test runtimes must not mutate real "
"issues or PRs (ADR: stable control runtime vs dev runtime)"
)
classified = mcp_server._classify_operation_gate_reasons([reason])
self.assertEqual(classified["runtime_mode"], [reason])
self.assertEqual(classified["stale_runtime"], [])
class Issue897PermissionBlockReport(unittest.TestCase):
def test_holds_op_is_diagnostic_defect_not_missing_permission(self):
profile = _profile(
"prgs-author",
"author",
["gitea.read", "gitea.issue.create", "gitea.issue.comment"],
[],
)
with patch.object(mcp_server, "get_profile", return_value=profile), patch.object(
mcp_server.gitea_config, "load_config", return_value=_config({"prgs-author": profile})
), patch.object(
mcp_server.gitea_config, "is_runtime_switching_enabled", return_value=True
):
report = mcp_server._permission_block_report("gitea.issue.create")
self.assertTrue(report.get("diagnostic_defect"), report)
self.assertIsNone(report.get("missing_permission"), report)
action = (report.get("exact_safe_next_action") or "").lower()
# Must not *recommend* profile switching; mentioning the forbidden
# action in a "do not call" instruction is fine.
self.assertNotIn("call gitea_activate_profile with", action)
self.assertNotIn("switch to the author mcp session", action)
self.assertNotIn("switch to the reviewer mcp session", action)
self.assertIn("diagnostic defect", action)
def test_true_missing_permission_still_reports(self):
profile = _profile(
"prgs-author",
"author",
["gitea.read", "gitea.issue.create"],
["gitea.pr.merge"],
)
reviewer = _profile(
"prgs-reviewer",
"reviewer",
["gitea.read", "gitea.pr.merge", "gitea.pr.approve"],
[],
)
with patch.object(mcp_server, "get_profile", return_value=profile), patch.object(
mcp_server.gitea_config,
"load_config",
return_value=_config({"prgs-author": profile, "prgs-reviewer": reviewer}),
), patch.object(
mcp_server.gitea_config, "is_runtime_switching_enabled", return_value=True
):
report = mcp_server._permission_block_report("gitea.pr.merge")
self.assertFalse(report.get("diagnostic_defect"), report)
self.assertEqual(report.get("missing_permission"), "gitea.pr.merge")
self.assertIn("prgs-reviewer", report.get("matching_configured_profiles") or [])
class Issue897GateRefusalMatrix(unittest.TestCase):
def _stale_parity(self) -> dict:
return {
"in_parity": True,
"stale": False,
"restart_required": True,
"determinable": True,
"startup_head": SHA_START,
"current_head": SHA_START,
"daemon_start_head": SHA_START,
"local_head": SHA_START,
"live_remote_head": SHA_LIVE,
"live_known": True,
"live_stale": True,
"mutation_safe": False,
"reasons": [
f"live remote master is {SHA_LIVE[:12]} but the MCP server "
f"started at {SHA_START[:12]}; the daemon is stale relative "
"to live master -- restart/reconnect before mutating"
],
}
def test_stale_plus_permitted_op_all_roles(self):
for name, role, permitted_op, allowed, forbidden, _forbidden_op in ROLE_MATRIX:
with self.subTest(profile=name, op=permitted_op):
profile = _profile(name, role, allowed, forbidden)
parity = self._stale_parity()
with patch.object(mcp_server, "get_profile", return_value=profile), patch.object(
mcp_server, "_current_master_parity", return_value=parity
), patch.object(
mcp_server, "_master_parity_block", return_value=list(parity["reasons"])
), patch.object(
mcp_server, "_runtime_mode_block", return_value=[]
), patch.object(
mcp_server, "_ensure_matching_profile", return_value=None
), patch.object(
mcp_server.session_ctx,
"mutation_context_audit_fields",
return_value={"session_profile": name},
):
blocked = mcp_server._profile_permission_block(permitted_op)
self.assertIsNotNone(blocked, name)
assert blocked is not None
self.assertEqual(
blocked.get("blocker_kind"),
"runtime_reconnect_required",
blocked,
)
self.assertNotIn("permission_report", blocked, blocked)
self.assertTrue(blocked.get("restart_required"), blocked)
self.assertEqual(blocked.get("startup_head"), SHA_START, blocked)
self.assertEqual(blocked.get("live_remote_head"), SHA_LIVE, blocked)
action = (blocked.get("exact_safe_next_action") or "").lower()
self.assertIn("reconnect", action)
self.assertNotIn("call gitea_activate_profile with", action)
self.assertNotIn("switch to the author mcp session", action)
self.assertNotIn("switch to the reviewer mcp session", action)
def test_fresh_plus_forbidden_op_all_roles(self):
for name, role, _permitted, allowed, forbidden, forbidden_op in ROLE_MATRIX:
with self.subTest(profile=name, op=forbidden_op):
profile = _profile(name, role, allowed, forbidden)
with patch.object(mcp_server, "get_profile", return_value=profile), patch.object(
mcp_server, "_master_parity_block", return_value=[]
), patch.object(
mcp_server, "_runtime_mode_block", return_value=[]
), patch.object(
mcp_server, "_ensure_matching_profile", return_value=None
), patch.object(
mcp_server.session_ctx,
"mutation_context_audit_fields",
return_value={"session_profile": name},
), patch.object(
mcp_server.gitea_config,
"load_config",
return_value=_config({name: profile}),
), patch.object(
mcp_server.gitea_config, "is_runtime_switching_enabled", return_value=False
):
blocked = mcp_server._profile_permission_block(forbidden_op)
self.assertIsNotNone(blocked, name)
assert blocked is not None
self.assertEqual(blocked.get("blocker_kind"), "permission_denied", blocked)
self.assertIn("permission_report", blocked, blocked)
report = blocked["permission_report"]
self.assertEqual(report.get("missing_permission"), forbidden_op, report)
self.assertFalse(report.get("diagnostic_defect"), report)
# No runtime reconnect fields for pure permission denial
self.assertNotEqual(
blocked.get("blocker_kind"), "runtime_reconnect_required"
)
def test_stale_plus_forbidden_op_both_causes_separated(self):
for name, role, _permitted, allowed, forbidden, forbidden_op in ROLE_MATRIX:
with self.subTest(profile=name, op=forbidden_op):
profile = _profile(name, role, allowed, forbidden)
parity = self._stale_parity()
stale_reason = parity["reasons"][0]
with patch.object(mcp_server, "get_profile", return_value=profile), patch.object(
mcp_server, "_current_master_parity", return_value=parity
), patch.object(
mcp_server, "_master_parity_block", return_value=[stale_reason]
), patch.object(
mcp_server, "_runtime_mode_block", return_value=[]
), patch.object(
mcp_server, "_ensure_matching_profile", return_value=None
), patch.object(
mcp_server.session_ctx,
"mutation_context_audit_fields",
return_value={"session_profile": name},
):
# Gate collects both classes; force permission reason too.
with patch.object(
mcp_server,
"_profile_operation_gate",
return_value=[
stale_reason,
f"profile is not allowed to {forbidden_op}",
],
):
blocked = mcp_server._profile_permission_block(forbidden_op)
self.assertIsNotNone(blocked)
assert blocked is not None
self.assertEqual(
blocked.get("blocker_kind"), "runtime_reconnect_required", blocked
)
self.assertNotIn("permission_report", blocked, blocked)
self.assertIn("permission_block_reasons", blocked, blocked)
self.assertIn("stale_runtime_reasons", blocked, blocked)
classes = blocked.get("gate_reason_classes") or {}
self.assertTrue(classes.get("stale_runtime"), classes)
self.assertTrue(classes.get("permission"), classes)
def test_runtime_mode_block_no_permission_report(self):
profile = _profile(
"prgs-author",
"author",
["gitea.read", "gitea.issue.create"],
[],
)
runtime_reason = (
"runtime mode is 'dev-test' and the mutation targets the "
"production repository; dev/test runtimes must not mutate real "
"issues or PRs (ADR: stable control runtime vs dev runtime)"
)
with patch.object(mcp_server, "get_profile", return_value=profile), patch.object(
mcp_server, "_master_parity_block", return_value=[]
), patch.object(
mcp_server, "_runtime_mode_block", return_value=[runtime_reason]
), patch.object(
mcp_server, "_ensure_matching_profile", return_value=None
), patch.object(
mcp_server.session_ctx,
"mutation_context_audit_fields",
return_value={"session_profile": "prgs-author"},
):
blocked = mcp_server._profile_permission_block("gitea.issue.create")
self.assertIsNotNone(blocked)
assert blocked is not None
self.assertEqual(blocked.get("blocker_kind"), "runtime_mode_blocked", blocked)
self.assertNotIn("permission_report", blocked, blocked)
action = (blocked.get("exact_safe_next_action") or "").lower()
self.assertNotIn("call gitea_activate_profile with", action)
self.assertIn("stable control runtime", action)
class Issue897CreateIssueRegression(unittest.TestCase):
def test_create_issue_stale_daemon_never_missing_issue_create(self):
"""Regression AC: stale prgs-author create_issue must not claim missing create."""
profile = _profile(
"prgs-author",
"author",
[
"gitea.read",
"gitea.issue.create",
"gitea.issue.comment",
"gitea.branch.create",
"gitea.branch.push",
"gitea.pr.create",
"gitea.pr.comment",
"gitea.repo.commit",
],
[],
)
stale_reason = (
f"live remote master is {SHA_LIVE[:12]} but the MCP server started "
f"at {SHA_START[:12]}; the daemon is stale relative to live master "
"-- restart/reconnect before mutating"
)
parity = {
"in_parity": True,
"stale": False,
"restart_required": True,
"determinable": True,
"startup_head": SHA_START,
"current_head": SHA_START,
"daemon_start_head": SHA_START,
"local_head": SHA_START,
"live_remote_head": SHA_LIVE,
"live_known": True,
"live_stale": True,
"mutation_safe": False,
"reasons": [stale_reason],
}
with patch.object(mcp_server, "get_profile", return_value=profile), patch.object(
mcp_server, "_current_master_parity", return_value=parity
), patch.object(
mcp_server, "_master_parity_block", return_value=[stale_reason]
), patch.object(
mcp_server, "_runtime_mode_block", return_value=[]
), patch.object(
mcp_server, "_ensure_matching_profile", return_value=None
), patch.object(
mcp_server.session_ctx,
"mutation_context_audit_fields",
return_value={"session_profile": "prgs-author"},
), patch.object(
mcp_server, "_mutation_config_authority_block", return_value=None
), patch.object(
mcp_server, "_session_context_mutation_block", return_value=None
):
blocked = mcp_server._profile_permission_block(
"gitea.issue.create", remote="prgs"
)
self.assertIsNotNone(blocked)
assert blocked is not None
self.assertEqual(blocked.get("blocker_kind"), "runtime_reconnect_required")
self.assertNotIn("permission_report", blocked)
# Even if a caller still built a raw report, holds-check must not claim missing.
with patch.object(mcp_server, "get_profile", return_value=profile):
raw = mcp_server._permission_block_report("gitea.issue.create")
self.assertIsNone(raw.get("missing_permission"), raw)
self.assertNotEqual(raw.get("missing_permission"), "gitea.issue.create")
def test_permission_report_for_gate_reasons_skips_stale(self):
stale = (
f"live remote master is {SHA_LIVE[:12]} but the MCP server started "
f"at {SHA_START[:12]}; the daemon is stale relative to live master "
"-- restart/reconnect before mutating"
)
self.assertIsNone(
mcp_server._permission_report_for_gate_reasons(
"gitea.issue.comment", [stale]
)
)
if __name__ == "__main__":
unittest.main()
+342
View File
@@ -0,0 +1,342 @@
"""Tests for MCP restart lifecycle audit events and incidents (#665)."""
from __future__ import annotations
import json
import os
import tempfile
import unittest
from unittest.mock import patch
import gitea_audit
import restart_audit as ra
class TestEventSchema(unittest.TestCase):
def test_all_lifecycle_event_types_are_named(self):
expected = {
"mcp.restart.impact_preview",
"mcp.restart.drain_enter",
"mcp.restart.drain_exit",
"mcp.restart.drain_proof",
"mcp.restart.apply_gate",
"mcp.restart.break_glass",
"mcp.restart.post_restart_reconcile",
"mcp.restart.narrower_recovery",
"mcp.restart.unguarded_detected",
}
self.assertEqual(set(ra.RESTART_EVENT_TYPES), expected)
def test_build_restart_event_core_fields(self):
event = ra.build_restart_event(
event_type=ra.EVENT_IMPACT_PREVIEW,
outcome="safe",
correlation_id="rst-abc123",
remote="prgs",
org="Scaled-Tech-Consulting",
repo="Gitea-Tools",
requesting_session_id="sess-1",
restart_class="full_mcp_restart",
profile_name="prgs-author",
authenticated_username="bot",
reasons=["inventory complete"],
details={"allow_restart": True},
now="2026-07-25T12:00:00+00:00",
)
self.assertEqual(event["event_type"], ra.EVENT_IMPACT_PREVIEW)
self.assertEqual(event["action"], ra.EVENT_IMPACT_PREVIEW)
self.assertEqual(event["action_type"], "restart_lifecycle")
self.assertEqual(event["result"], "safe")
self.assertEqual(event["correlation_id"], "rst-abc123")
self.assertEqual(event["profile_name"], "prgs-author")
self.assertEqual(event["authenticated_username"], "bot")
meta = event["request_metadata"]
self.assertEqual(meta["event_family"], "mcp.restart")
self.assertEqual(meta["correlation_id"], "rst-abc123")
self.assertEqual(meta["restart_class"], "full_mcp_restart")
self.assertEqual(meta["details"]["allow_restart"], True)
def test_unknown_event_type_raises(self):
with self.assertRaises(ValueError) as ctx:
ra.build_restart_event(
event_type="mcp.restart.not_a_real_event",
outcome="x",
correlation_id="rst-1",
)
self.assertIn("unknown restart audit event_type", str(ctx.exception))
def test_reasons_and_details_are_redacted(self):
event = ra.build_restart_event(
event_type=ra.EVENT_APPLY_GATE,
outcome="deny",
correlation_id="rst-sec",
reasons=["token secret-xyz rejected", "ok"],
details={"token": "leak-token", "status": "denied"},
)
self.assertNotIn("secret-xyz", event.get("reason") or "")
meta = event["request_metadata"]
self.assertEqual(meta["details"]["token"], gitea_audit.REDACTED)
self.assertEqual(meta["details"]["status"], "denied")
for reason in meta["reasons"]:
self.assertNotIn("secret-xyz", reason)
def test_new_correlation_id_shape(self):
cid = ra.new_correlation_id()
self.assertTrue(cid.startswith("rst-"))
self.assertEqual(len(cid), len("rst-") + 16)
class TestEmitAndRequire(unittest.TestCase):
def test_emit_appends_json_line(self):
with tempfile.TemporaryDirectory() as d:
path = os.path.join(d, "audit.log")
event = ra.build_restart_event(
event_type=ra.EVENT_DRAIN_PROOF,
outcome="pass",
correlation_id="rst-write",
)
self.assertTrue(ra.emit_restart_event(event, path=path))
with open(path, encoding="utf-8") as fh:
lines = fh.read().splitlines()
self.assertEqual(len(lines), 1)
loaded = json.loads(lines[0])
self.assertEqual(loaded["event_type"], ra.EVENT_DRAIN_PROOF)
self.assertEqual(loaded["correlation_id"], "rst-write")
def test_emit_never_raises(self):
self.assertFalse(
ra.emit_restart_event({"action": "x"}, path="/no/such/dir/audit.log")
)
def test_require_audit_denies_privileged_when_write_fails_and_enabled(self):
deny = ra.require_audit_or_deny(
privileged=True, written=False, audit_enabled=True
)
self.assertEqual(len(deny), 1)
self.assertIn("fail closed", deny[0])
def test_require_audit_allows_when_audit_disabled(self):
# Rollout policy: enable audit before enforcing deny-on-audit-fail.
deny = ra.require_audit_or_deny(
privileged=True, written=False, audit_enabled=False
)
self.assertEqual(deny, [])
def test_require_audit_noop_for_non_privileged(self):
deny = ra.require_audit_or_deny(
privileged=False, written=False, audit_enabled=True
)
self.assertEqual(deny, [])
def test_require_audit_allows_when_written(self):
deny = ra.require_audit_or_deny(
privileged=True, written=True, audit_enabled=True
)
self.assertEqual(deny, [])
class TestIncidents(unittest.TestCase):
def test_break_glass_descriptor(self):
desc = ra.build_incident_descriptor(
kind=ra.INCIDENT_BREAK_GLASS,
reasons=["break-glass authorized"],
correlation_id="rst-bg",
requesting_session_id="s1",
restart_class="full_mcp_restart",
remote="prgs",
org="O",
repo="R",
)
self.assertEqual(desc["kind"], ra.INCIDENT_BREAK_GLASS)
self.assertIn("Break-glass", desc["title"])
self.assertIn("mcp-health", desc["labels"])
self.assertEqual(desc["source"], "restart_audit#665")
def test_incident_body_redacts_and_includes_correlation(self):
desc = ra.build_incident_descriptor(
kind=ra.INCIDENT_FAILED_DRAIN,
reasons=["token secret-xyz failed proof"],
correlation_id="rst-body",
remote="prgs",
org="O",
repo="R",
proof_id="proof-1",
)
body = ra.incident_body(desc)
self.assertIn("rst-body", body)
self.assertIn("proof-1", body)
self.assertIn("mcp-restart-incident:v1", body)
self.assertNotIn("secret-xyz", body)
def test_materialize_dry_run(self):
desc = ra.build_incident_descriptor(
kind=ra.INCIDENT_FAILED_DRAIN,
reasons=["denied"],
correlation_id="rst-dr",
)
result = ra.materialize_incident(desc, create_issue_fn=lambda **k: {}, dry_run=True)
self.assertFalse(result["created"])
self.assertTrue(result["dry_run"])
self.assertIn("dry-run", result["reasons"][0])
def test_materialize_without_create_fn(self):
desc = ra.build_incident_descriptor(
kind=ra.INCIDENT_FAILED_DRAIN,
reasons=["denied"],
correlation_id="rst-nfn",
)
result = ra.materialize_incident(desc, create_issue_fn=None)
self.assertFalse(result["created"])
self.assertIn("create_issue_fn not provided", result["reasons"][0])
def test_materialize_creates_issue(self):
created = {}
def _create(*, title, body, labels, org=None, repo=None, **_kw):
created["title"] = title
created["body"] = body
created["labels"] = labels
created["org"] = org
created["repo"] = repo
return {"number": 999}
desc = ra.build_incident_descriptor(
kind=ra.INCIDENT_BREAK_GLASS,
reasons=["break-glass"],
correlation_id="rst-create",
org="O",
repo="R",
)
result = ra.materialize_incident(desc, create_issue_fn=_create)
self.assertTrue(result["created"])
self.assertEqual(result["issue_number"], 999)
self.assertIn("Break-glass", created["title"])
self.assertIn("rst-create", created["body"])
self.assertEqual(created["org"], "O")
def test_materialize_never_raises_on_create_failure(self):
def _boom(**_kw):
raise RuntimeError("token secret-xyz network")
desc = ra.build_incident_descriptor(
kind=ra.INCIDENT_FAILED_DRAIN,
reasons=["x"],
correlation_id="rst-boom",
)
result = ra.materialize_incident(desc, create_issue_fn=_boom)
self.assertFalse(result["created"])
self.assertIn("failed", result["reasons"][0])
self.assertNotIn("secret-xyz", result["reasons"][0])
class TestIncidentFromApplyGate(unittest.TestCase):
def test_break_glass_always_incident(self):
desc = ra.incident_from_apply_gate(
gate_payload={"allow": True, "reasons": [], "proof_id": None},
break_glass=True,
correlation_id="rst-bg2",
requesting_session_id="s",
restart_class="full_mcp_restart",
remote="prgs",
org="O",
repo="R",
)
self.assertIsNotNone(desc)
self.assertEqual(desc["kind"], ra.INCIDENT_BREAK_GLASS)
def test_failed_drain_from_gate_incident(self):
desc = ra.incident_from_apply_gate(
gate_payload={
"allow": False,
"drain_gate_allow": False,
"reasons": ["proof expired"],
"incident": {
"reasons": ["proof expired"],
"proof_id": "p1",
},
},
break_glass=False,
correlation_id="rst-fd",
requesting_session_id="s",
restart_class="full_mcp_restart",
remote="prgs",
org="O",
repo="R",
)
self.assertIsNotNone(desc)
self.assertEqual(desc["kind"], ra.INCIDENT_FAILED_DRAIN)
self.assertEqual(desc["proof_id"], "p1")
def test_allow_without_break_glass_no_incident(self):
desc = ra.incident_from_apply_gate(
gate_payload={
"allow": True,
"drain_gate_allow": True,
"reasons": [],
},
break_glass=False,
correlation_id="rst-ok",
requesting_session_id="s",
restart_class="full_mcp_restart",
remote="prgs",
org="O",
repo="R",
)
self.assertIsNone(desc)
class TestRecordLifecycle(unittest.TestCase):
def test_record_emits_and_materializes(self):
created = []
def _create(**kwargs):
created.append(kwargs)
return {"number": 42}
with tempfile.TemporaryDirectory() as d:
path = os.path.join(d, "audit.log")
with patch.dict(os.environ, {"GITEA_AUDIT_LOG": path}, clear=False):
incident = ra.build_incident_descriptor(
kind=ra.INCIDENT_BREAK_GLASS,
reasons=["bg"],
correlation_id="rst-lc",
org="O",
repo="R",
)
out = ra.record_restart_lifecycle(
event_type=ra.EVENT_BREAK_GLASS,
outcome="break_glass",
correlation_id="rst-lc",
remote="prgs",
org="O",
repo="R",
privileged=True,
create_incident=incident,
create_issue_fn=_create,
audit_path=path,
)
self.assertTrue(out["audit_written"])
self.assertEqual(out["deny_reasons"], [])
self.assertTrue(out["incident_result"]["created"])
self.assertEqual(out["incident_result"]["issue_number"], 42)
self.assertEqual(len(created), 1)
def test_privileged_deny_when_audit_write_fails(self):
with patch.dict(
os.environ, {"GITEA_AUDIT_LOG": "/no/such/dir/a.log"}, clear=False
):
with patch("restart_audit.emit_restart_event", return_value=False):
with patch("gitea_audit.audit_enabled", return_value=True):
out = ra.record_restart_lifecycle(
event_type=ra.EVENT_APPLY_GATE,
outcome="deny",
correlation_id="rst-deny",
privileged=True,
audit_path="/no/such/dir/a.log",
)
self.assertFalse(out["audit_written"])
self.assertEqual(len(out["deny_reasons"]), 1)
if __name__ == "__main__":
unittest.main()
+739
View File
@@ -0,0 +1,739 @@
"""Tests for the Runtime and session view (Phase 1, #641).
Covers clean and stale session rendering, contamination marker surfacing,
worktree binding display, sanctioned recovery links (no pkill), nav/live
status, and the JSON API export.
Also pins the two invariants a reviewer found violated at head a81db754:
degraded ownership sections must render as *unknown* rather than as an
affirmative "none"/"unbound", and contamination payload text must be redacted
at the display boundary rather than trusted from the write-time denylist.
"""
from __future__ import annotations
import os
import sys
import unittest
from pathlib import Path
from unittest import mock
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
from tests.webui_testclient import TestClient
from webui.app import create_app
from webui.inventory import (
AUTHORITY_CONTROL_PLANE_DB,
AUTHORITY_FILESYSTEM,
InventorySection,
InventorySnapshot,
STATUS_DEGRADED,
STATUS_OK,
STATUS_UNAVAILABLE,
)
from webui.nav import NAV_GROUPS, STUB_PAGES, iter_nav_items
from webui.runtime_health import FileHash, RuntimeSnapshot
from webui.session_loader import (
ContaminationMarker,
SessionRow,
SessionViewSnapshot,
_build_session_rows,
_inspect_contamination,
load_session_view_snapshot,
snapshot_to_dict,
)
from webui.session_views import render_sessions_page
def _runtime(
*,
stale: str | None = None,
profile: str = "prgs-author",
role: str = "author",
) -> RuntimeSnapshot:
return RuntimeSnapshot(
project_id="gitea-tools",
repo_root="/tmp/repo",
remote="prgs",
host="gitea.prgs.cc",
profile_name=profile,
role_kind=role,
config_model="v2-contexts",
profile_mode="dynamic-profile",
profile_source="config file profile",
authenticated_username="jcwalker3",
identity_error=None,
repo_sha="a" * 40,
remote_master_sha="a" * 40,
commits_behind_master=0,
stale_runtime_warning=stale,
shell_health={"shell_use_allowed": True, "consecutive_spawn_failures": 0},
workflow_hashes=(
FileHash(label="SKILL.md", path="skills/llm-project-workflow/SKILL.md", sha256="abc"),
),
schema_hashes=(),
restart_guidance="docs/mcp-namespace-eof-recovery.md",
fetch_error=None,
)
def _inventory(
*,
sessions: tuple[dict, ...] = (),
leases: tuple[dict, ...] = (),
locks: tuple[dict, ...] = (),
worktrees: tuple[dict, ...] = (),
namespaces: tuple[dict, ...] = (),
statuses: dict[str, str] | None = None,
) -> InventorySnapshot:
"""Build a snapshot; ``statuses`` degrades named sections (default all ok)."""
status_of = statuses or {}
def _status(name: str) -> str:
return status_of.get(name, STATUS_OK)
sections = (
InventorySection(
name="sessions",
authority=AUTHORITY_CONTROL_PLANE_DB,
status=_status("sessions"),
items=sessions,
),
InventorySection(
name="leases",
authority=AUTHORITY_CONTROL_PLANE_DB,
status=_status("leases"),
items=leases,
),
InventorySection(
name="locks",
authority=AUTHORITY_FILESYSTEM,
status=_status("locks"),
items=locks,
),
InventorySection(
name="worktrees",
authority=AUTHORITY_FILESYSTEM,
status=_status("worktrees"),
items=worktrees,
),
InventorySection(
name="namespaces",
authority=AUTHORITY_FILESYSTEM,
status=_status("namespaces"),
items=namespaces
or (
{
"profile_name": "prgs-author",
"role": "author",
"mcp_namespace": "gitea-author",
"capability_summary": {
"can_author": True,
"can_review": False,
"can_merge": False,
},
"active": True,
},
),
reason="only the profile serving this web process is observable",
),
)
index = {section.name: section for section in sections}
return InventorySnapshot(
generated_at="2026-07-25T00:00:00+00:00",
sections=sections,
collisions=(),
correlations=(),
scan_ms=1.0,
_section_index=index,
)
def _clean_session() -> dict:
return {
"session_id": "prgs-author-111-clean",
"role": "author",
"profile": "prgs-author",
"namespace": "gitea-author",
"pid": 1111,
"pid_alive": True,
"status": "active",
"started_at": "2026-07-25T00:00:00Z",
"last_heartbeat_at": "2026-07-25T01:00:00Z",
}
def _stale_session() -> dict:
return {
"session_id": "prgs-author-222-stale",
"role": "author",
"profile": "prgs-author",
"namespace": "gitea-author",
"pid": 2222,
"pid_alive": False,
"status": "active",
"started_at": "2026-07-24T00:00:00Z",
"last_heartbeat_at": "2026-07-24T01:00:00Z",
}
class TestBuildSessionRows(unittest.TestCase):
def test_clean_session_has_no_stale_or_contamination_flags(self):
inventory = _inventory(
sessions=(_clean_session(),),
leases=(
{
"lease_id": "lease-clean",
"session_id": "prgs-author-111-clean",
"status": "active",
"expired": False,
"work_kind": "issue",
"work_number": 641,
},
),
locks=(
{
"issue_number": 641,
"branch_name": "feat/issue-641-runtime-session-view",
"worktree_path": "~/Development/Gitea-Tools/branches/feat-issue-641",
"live": True,
},
),
)
rows = _build_session_rows(inventory, contamination=())
self.assertEqual(len(rows), 1)
row = rows[0]
self.assertEqual(row.session_id, "prgs-author-111-clean")
self.assertEqual(row.role, "author")
self.assertEqual(row.namespace, "gitea-author")
self.assertEqual(row.pid_alive, True)
self.assertEqual(row.lease_ids, ("lease-clean",))
self.assertEqual(row.work_refs, ("issue#641",))
self.assertTrue(row.worktree_paths)
self.assertEqual(row.stale_flags, ())
self.assertEqual(row.contamination_flags, ())
def test_stale_session_flags_dead_pid(self):
inventory = _inventory(sessions=(_stale_session(),))
rows = _build_session_rows(inventory, contamination=())
self.assertEqual(rows[0].stale_flags, ("pid-dead",))
def test_contamination_marker_binds_to_session(self):
inventory = _inventory(sessions=(_clean_session(),))
marker = ContaminationMarker(
kind="runtime_recovery_contamination",
on_disk=True,
has_payload=True,
summary="manual daemon kill",
reason_class="manual_daemon_kill",
session_id="prgs-author-111-clean",
role="author",
command_summary="pkill -f mcp_server.py",
cleared=False,
)
rows = _build_session_rows(inventory, contamination=(marker,))
self.assertIn("runtime_recovery_contamination", rows[0].contamination_flags)
def test_process_wide_contamination_surfaces_on_all_sessions(self):
inventory = _inventory(sessions=(_clean_session(), _stale_session()))
marker = ContaminationMarker(
kind="stable_branch_contamination",
on_disk=True,
has_payload=True,
summary="direct master push attempt",
reason_class="stable_branch_push",
session_id=None,
cleared=False,
)
rows = _build_session_rows(inventory, contamination=(marker,))
self.assertEqual(len(rows), 2)
for row in rows:
self.assertTrue(
any("stable_branch_contamination" in f for f in row.contamination_flags)
)
class TestRenderSessionsPage(unittest.TestCase):
def _snapshot(
self,
*,
sessions: tuple[dict, ...],
contamination: tuple[ContaminationMarker, ...] = (),
stale_runtime: str | None = None,
) -> SessionViewSnapshot:
inventory = _inventory(
sessions=sessions,
leases=(
{
"lease_id": "lease-1",
"session_id": sessions[0]["session_id"] if sessions else "",
"status": "active",
"expired": False,
"work_kind": "issue",
"work_number": 641,
},
)
if sessions
else (),
locks=(
{
"issue_number": 641,
"worktree_path": "branches/feat-issue-641",
},
)
if sessions
else (),
worktrees=(
{
"rel_path": "branches/feat-issue-641",
"branch": "feat/issue-641-runtime-session-view",
"classification": "active_issue_work",
"registered_worktree": True,
"dirty": False,
},
),
)
rows = _build_session_rows(inventory, contamination)
return SessionViewSnapshot(
runtime=_runtime(stale=stale_runtime),
inventory=inventory,
sessions=rows,
contamination_markers=contamination,
)
def test_clean_session_render(self):
html = render_sessions_page(self._snapshot(sessions=(_clean_session(),)))
self.assertIn("Runtime and sessions", html)
self.assertIn("prgs-author-111-clean", html)
self.assertIn("gitea-author", html)
self.assertIn("branches/feat-issue-641", html)
self.assertIn("Sanctioned recovery", html)
self.assertIn("docs/mcp-namespace-eof-recovery.md", html)
# Recovery section must name reconnect and forbid manual kill.
recovery_idx = html.lower().find("sanctioned recovery")
self.assertGreaterEqual(recovery_idx, 0)
recovery = html[recovery_idx:].lower()
self.assertIn("reconnect", recovery)
self.assertIn("contamination", recovery)
self.assertIn("not recovery", recovery)
self.assertNotIn("run pkill", recovery)
self.assertNotIn("killall", recovery)
def test_stale_session_render(self):
html = render_sessions_page(self._snapshot(sessions=(_stale_session(),)))
self.assertIn("prgs-author-222-stale", html)
self.assertIn("pid-dead", html)
self.assertIn("badge-stale", html)
def test_contamination_render_is_not_silent(self):
marker = ContaminationMarker(
kind="runtime_recovery_contamination",
on_disk=True,
has_payload=True,
summary="manual kill",
reason_class="manual_daemon_kill",
session_id="prgs-author-111-clean",
command_summary="pkill -f mcp_server.py",
cleared=False,
)
html = render_sessions_page(
self._snapshot(sessions=(_clean_session(),), contamination=(marker,))
)
self.assertIn("Contamination markers", html)
self.assertIn("runtime_recovery_contamination", html)
self.assertIn("ACTIVE", html)
self.assertIn("badge-blocked", html)
def test_stale_runtime_banner(self):
html = render_sessions_page(
self._snapshot(
sessions=(_clean_session(),),
stale_runtime="server behind master by 3 commits",
)
)
self.assertIn("Stale runtime", html)
self.assertIn("server behind master", html)
class TestSessionLoaderComposition(unittest.TestCase):
def test_load_with_injected_sources(self):
inventory = _inventory(sessions=(_clean_session(), _stale_session()))
snap = load_session_view_snapshot(
load_runtime=lambda: _runtime(),
load_inventory=lambda: inventory,
inspect_contamination=lambda **_k: {
"on_disk": False,
"has_payload": False,
"summary": "absent",
},
load_contamination_payload=lambda **_k: None,
)
self.assertEqual(len(snap.sessions), 2)
self.assertEqual(snap.stale_session_count, 1)
self.assertEqual(snap.contaminated_session_count, 0)
data = snapshot_to_dict(snap)
self.assertEqual(data["view"], "runtime-sessions")
self.assertEqual(data["issue"], 641)
self.assertTrue(data["read_only"])
self.assertEqual(data["session_counts"]["total"], 2)
self.assertEqual(data["session_counts"]["stale"], 1)
self.assertIn("recovery_docs", data)
class TestSessionsRoutes(unittest.TestCase):
def setUp(self):
self.client = TestClient(create_app())
inventory = _inventory(
sessions=(_clean_session(), _stale_session()),
leases=(
{
"lease_id": "lease-x",
"session_id": "prgs-author-111-clean",
"status": "active",
"expired": False,
"work_kind": "issue",
"work_number": 641,
},
),
locks=(
{
"issue_number": 641,
"worktree_path": "branches/feat-issue-641",
},
),
worktrees=(
{
"rel_path": "branches/feat-issue-641",
"branch": "feat/issue-641-runtime-session-view",
"classification": "active_issue_work",
"registered_worktree": True,
"dirty": False,
},
),
)
rows = _build_session_rows(inventory, contamination=())
self.snapshot = SessionViewSnapshot(
runtime=_runtime(stale="stale for test"),
inventory=inventory,
sessions=rows,
contamination_markers=(),
)
self._patch = mock.patch(
"webui.app.load_session_view_snapshot",
return_value=self.snapshot,
)
self._patch.start()
def tearDown(self):
self._patch.stop()
def test_sessions_page_live(self):
response = self.client.get("/sessions")
self.assertEqual(response.status_code, 200)
self.assertIn("Runtime and sessions", response.text)
self.assertIn("prgs-author-111-clean", response.text)
self.assertIn("prgs-author-222-stale", response.text)
self.assertIn("pid-dead", response.text)
self.assertIn("Sanctioned recovery", response.text)
self.assertNotIn("Phase 1 shell placeholder", response.text)
self.assertNotIn("child issue of #425", response.text.lower())
def test_api_sessions_json(self):
for path in ("/api/sessions", "/api/v1/sessions"):
response = self.client.get(path)
self.assertEqual(response.status_code, 200, path)
data = response.json()
self.assertEqual(data["view"], "runtime-sessions")
self.assertEqual(data["session_counts"]["total"], 2)
self.assertEqual(data["session_counts"]["stale"], 1)
self.assertTrue(data["read_only"])
self.assertEqual(data["mutations"], [])
def test_nav_marks_sessions_live(self):
sessions_items = [
item for item in iter_nav_items() if item.href == "/sessions"
]
self.assertEqual(len(sessions_items), 1)
self.assertEqual(sessions_items[0].status, "live")
self.assertNotIn("/sessions", STUB_PAGES)
# Home page should not mark Sessions as stub.
home = self.client.get("/")
self.assertEqual(home.status_code, 200)
self.assertIn('href="/sessions"', home.text)
# Stub marker only appears next to remaining stub destinations.
self.assertNotIn(
'href="/sessions">Sessions</a> <span class="muted">(stub)</span>',
home.text,
)
class TestDegradedOwnershipAuthority(unittest.TestCase):
"""B1: a section that could not be read must never render as absence."""
def _snapshot(self, inventory: InventorySnapshot) -> SessionViewSnapshot:
return SessionViewSnapshot(
runtime=_runtime(),
inventory=inventory,
sessions=_build_session_rows(inventory, contamination=()),
contamination_markers=(),
)
def test_unavailable_leases_mark_row_authority_unproven(self):
inventory = _inventory(
sessions=(_clean_session(),),
statuses={"leases": STATUS_UNAVAILABLE},
)
row = _build_session_rows(inventory, contamination=())[0]
self.assertEqual(row.lease_ids, ())
self.assertEqual(row.lease_authority, STATUS_UNAVAILABLE)
# Worktree binding is correlated through lease work numbers, so it
# inherits the unreadable lease section.
self.assertEqual(row.worktree_authority, STATUS_UNAVAILABLE)
self.assertFalse(row.ownership_authority_complete)
def test_readable_locks_are_not_reported_unbound_when_leases_degrade(self):
# The narrow variant: locks hold a real worktree_path and read cleanly,
# but the lease section that supplies the correlating work number does
# not. The row must say unknown, not "unbound".
inventory = _inventory(
sessions=(_clean_session(),),
locks=(
{
"issue_number": 641,
"worktree_path": "branches/feat-issue-641",
},
),
statuses={"leases": STATUS_DEGRADED},
)
row = _build_session_rows(inventory, contamination=())[0]
self.assertEqual(row.worktree_paths, ())
self.assertEqual(row.worktree_authority, STATUS_DEGRADED)
self.assertFalse(row.ownership_authority_complete)
def test_degraded_render_says_unknown_not_none_or_unbound(self):
inventory = _inventory(
sessions=(_clean_session(),),
statuses={"leases": STATUS_UNAVAILABLE, "locks": STATUS_UNAVAILABLE},
)
html = render_sessions_page(self._snapshot(inventory))
self.assertIn("unknown (inventory unavailable)", html)
self.assertIn("authority unproven", html)
self.assertIn("Ownership authority incomplete", html)
# The affirmative-absence strings must be gone from the row entirely.
self.assertNotIn(">none<", html)
self.assertNotIn(">unbound<", html)
def test_clean_inventory_still_renders_affirmative_absence(self):
# Guards against over-correcting B1 into "everything is unknown".
inventory = _inventory(sessions=(_clean_session(),))
html = render_sessions_page(self._snapshot(inventory))
self.assertIn(">none<", html)
self.assertIn(">unbound<", html)
# The column legend mentions "unknown (inventory …)" as static copy, so
# assert on the per-row marker and the concrete statuses instead.
self.assertNotIn("authority unproven", html)
self.assertNotIn("unknown (inventory unavailable)", html)
self.assertNotIn("unknown (inventory degraded)", html)
self.assertNotIn("Ownership authority incomplete", html)
def test_json_export_carries_snapshot_and_per_row_authority(self):
inventory = _inventory(
sessions=(_clean_session(),),
statuses={"locks": STATUS_UNAVAILABLE},
)
data = snapshot_to_dict(self._snapshot(inventory))
self.assertFalse(data["ownership_authority_complete"])
self.assertEqual(
data["ownership_section_status"]["locks"], STATUS_UNAVAILABLE
)
self.assertEqual(data["ownership_section_status"]["leases"], STATUS_OK)
self.assertIn("unknown, not unowned", data["ownership_note"])
row = data["sessions"][0]
self.assertTrue(row["lease_authority_complete"])
self.assertFalse(row["worktree_authority_complete"])
self.assertEqual(row["worktree_authority"], STATUS_UNAVAILABLE)
self.assertFalse(row["ownership_authority_complete"])
self.assertIn("unknown, not unowned", row["ownership_note"])
def test_json_export_is_affirmative_when_every_source_reads(self):
inventory = _inventory(sessions=(_clean_session(),))
data = snapshot_to_dict(self._snapshot(inventory))
self.assertTrue(data["ownership_authority_complete"])
self.assertTrue(data["sessions"][0]["ownership_authority_complete"])
def test_missing_session_list_is_not_reported_as_no_sessions(self):
inventory = _inventory(statuses={"sessions": STATUS_UNAVAILABLE})
html = render_sessions_page(self._snapshot(inventory))
self.assertIn("could not be read", html)
self.assertIn("not evidence that no sessions exist", html)
def test_expired_lease_flags_row_as_stale(self):
inventory = _inventory(
sessions=(_clean_session(),),
leases=(
{
"lease_id": "lease-expired-1",
"session_id": "prgs-author-111-clean",
"status": "active",
"expired": True,
"work_kind": "issue",
"work_number": 641,
},
),
)
row = _build_session_rows(inventory, contamination=())[0]
self.assertIn("lease-expired", row.stale_flags)
self.assertIn("active-lease-past-expiry", row.stale_flags)
class TestContaminationRedaction(unittest.TestCase):
"""B2: marker payload text is redacted at the display boundary."""
def _marker(self, payload: dict) -> ContaminationMarker:
return _inspect_contamination(
"runtime_recovery_contamination",
remote="prgs",
inspect=lambda **_k: {
"on_disk": True,
"has_payload": True,
"summary": "",
},
load=lambda **_k: payload,
)
def test_home_paths_are_collapsed(self):
home = os.path.expanduser("~")
marker = self._marker(
{"command_summary": f"pkill -f {home}/Development/Gitea-Tools/x.py"}
)
self.assertNotIn(home, marker.command_summary)
self.assertIn("~/Development/Gitea-Tools/x.py", marker.command_summary)
def test_secrets_missed_by_the_write_time_denylist_are_redacted(self):
# Each of these was verified in review to survive
# stable_branch_push_guard.redact_command untouched.
cases = (
("curl -H 'X-Api-Key: SUPERSECRET123' https://example.invalid", "SUPERSECRET123"),
("cmd --password hunter2 origin master", "hunter2"),
("PRIVATE_KEY=abc123 python deploy.py", "abc123"),
("fetch https://user:[email protected]/x.git", "user:pw"),
)
for raw, secret in cases:
with self.subTest(raw=raw):
marker = self._marker({"command_summary": raw})
self.assertNotIn(secret, marker.command_summary)
self.assertIn("[redacted]", marker.command_summary)
def test_command_summary_is_redacted_not_removed(self):
# It is legitimate #630 evidence: the operator must still see which
# daemon was killed.
marker = self._marker(
{
"command_summary": "pkill -f gitea_mcp_server.py",
"reason_class": "manual_daemon_kill",
"session_id": "prgs-author-111-clean",
"role": "author",
}
)
self.assertIn("pkill -f gitea_mcp_server.py", marker.command_summary)
self.assertEqual(marker.reason_class, "manual_daemon_kill")
self.assertEqual(marker.session_id, "prgs-author-111-clean")
self.assertEqual(marker.role, "author")
def test_rendered_page_exposes_no_home_path_from_a_marker(self):
home = os.path.expanduser("~")
marker = self._marker(
{
"command_summary": f"pkill -f {home}/Development/Gitea-Tools/x.py",
"reason_class": "manual_daemon_kill",
}
)
inventory = _inventory(sessions=(_clean_session(),))
html = render_sessions_page(
SessionViewSnapshot(
runtime=_runtime(),
inventory=inventory,
sessions=_build_session_rows(inventory, (marker,)),
contamination_markers=(marker,),
)
)
self.assertIn("Contamination markers", html)
self.assertNotIn(home, html)
class TestSessionsPageEscaping(unittest.TestCase):
"""Hostile values from every rendered source stay inert (N2)."""
HOSTILE = '<script>alert("xss")</script>'
def test_hostile_session_and_marker_values_are_escaped(self):
session = dict(_clean_session())
session["session_id"] = f"sid-{self.HOSTILE}"
session["role"] = self.HOSTILE
session["profile"] = self.HOSTILE
session["namespace"] = self.HOSTILE
session["status"] = self.HOSTILE
inventory = _inventory(
sessions=(session,),
leases=(
{
"lease_id": self.HOSTILE,
"session_id": session["session_id"],
"status": "active",
"expired": False,
"work_kind": self.HOSTILE,
"work_number": 641,
},
),
locks=(
{
"issue_number": 641,
"worktree_path": self.HOSTILE,
},
),
)
marker = ContaminationMarker(
kind="runtime_recovery_contamination",
on_disk=True,
has_payload=True,
summary=self.HOSTILE,
reason_class=self.HOSTILE,
session_id=session["session_id"],
role=self.HOSTILE,
command_summary=self.HOSTILE,
cleared=False,
)
html = render_sessions_page(
SessionViewSnapshot(
runtime=_runtime(),
inventory=inventory,
sessions=_build_session_rows(inventory, (marker,)),
contamination_markers=(marker,),
)
)
self.assertNotIn("<script>", html)
self.assertNotIn('alert("xss")', html)
self.assertIn("&lt;script&gt;", html)
def test_hostile_values_in_a_degraded_render_are_escaped(self):
inventory = _inventory(
sessions=(dict(_clean_session(), session_id=f"sid-{self.HOSTILE}"),),
statuses={"leases": STATUS_UNAVAILABLE, "locks": STATUS_DEGRADED},
)
html = render_sessions_page(
SessionViewSnapshot(
runtime=_runtime(),
inventory=inventory,
sessions=_build_session_rows(inventory, contamination=()),
contamination_markers=(),
)
)
self.assertNotIn("<script>", html)
self.assertIn("&lt;script&gt;", html)
self.assertIn("unknown (inventory", html)
if __name__ == "__main__":
unittest.main()
+20
View File
@@ -48,6 +48,11 @@ from webui.worktree_scanner import load_hygiene_snapshot, snapshot_to_dict as wo
from webui.worktree_views import render_worktrees_page
from webui.runtime_health import load_runtime_snapshot, snapshot_to_dict as runtime_snapshot_to_dict
from webui.runtime_views import render_runtime_page
from webui.session_loader import (
load_session_view_snapshot,
snapshot_to_dict as session_view_snapshot_to_dict,
)
from webui.session_views import render_sessions_page
from webui.inventory import (
SECTION_NAMES as _INVENTORY_SECTIONS,
load_inventory_snapshot,
@@ -87,6 +92,7 @@ _LEGACY_PAGES = (
("/projects", "Projects", "registry and onboarding (#427)"),
("/prompts", "Prompts", "canonical workflow prompt library (#428)"),
("/runtime", "Runtime", "MCP health and stale-runtime detection (#430)"),
("/sessions", "Sessions", "runtime and session view (#641)"),
("/audit", "Audit", "final-report paste and validator preview (#431)"),
("/worktrees", "Worktrees", "branch hygiene dashboard (#432)"),
("/leases", "Leases", "collision and lease visibility (#433)"),
@@ -325,6 +331,17 @@ async def api_runtime(_request: Request) -> JSONResponse:
return JSONResponse(runtime_snapshot_to_dict(load_runtime_snapshot()))
async def sessions(_request: Request) -> HTMLResponse:
"""Runtime and session view (#641) — read-only composition of health + inventory."""
snapshot = load_session_view_snapshot()
return HTMLResponse(render_sessions_page(snapshot))
async def api_sessions(_request: Request) -> JSONResponse:
"""JSON export for the runtime/session view (#641)."""
return JSONResponse(session_view_snapshot_to_dict(load_session_view_snapshot()))
async def _parse_audit_form(request: Request) -> tuple[str, str | None]:
if request.method == "GET":
return "", None
@@ -764,6 +781,9 @@ def create_app(*, bind_host: str | None = None) -> Starlette:
Route("/api/prompts", api_prompts, methods=["GET"]),
Route("/runtime", runtime, methods=["GET"]),
Route("/api/runtime", api_runtime, methods=["GET"]),
Route("/sessions", sessions, methods=["GET"]),
Route("/api/sessions", api_sessions, methods=["GET"]),
Route("/api/v1/sessions", api_sessions, methods=["GET"]),
Route("/api/v1/timeline", api_v1_timeline, methods=["GET"]),
Route("/analytics", analytics, methods=["GET"]),
Route("/api/analytics", api_v1_analytics, methods=["GET"]),
+65
View File
@@ -76,6 +76,35 @@ _CREDENTIAL_KEY_RE = re.compile(
)
_REDACTED = "[redacted]"
#: Credential-shaped *name* as it appears inside a free-form command line. This
#: is deliberately broader than :data:`_CREDENTIAL_KEY_RE` — it also matches a
#: bare ``key`` component, so ``PRIVATE_KEY=`` is caught. Over-redacting a
#: displayed string is safe; under-redacting one is not.
_TEXT_CREDENTIAL_NAME = (
r"[A-Za-z0-9_.\-]*"
r"(?:token|secret|password|passwd|key|authorization|bearer|credential)"
r"[A-Za-z0-9_.\-]*"
)
#: A value following such a name: single-quoted, double-quoted, or bare. The
#: bare form stops at a quote so an enclosing quote survives the redaction.
_TEXT_CREDENTIAL_VALUE = r"'[^']*'|\"[^\"]*\"|[^\s'\"]+"
_TEXT_CREDENTIAL_FLAG_RE = re.compile(
rf"(?P<key>(?<![\w\-])--?{_TEXT_CREDENTIAL_NAME})"
rf"(?P<sep>[=\s]+)"
rf"(?P<value>{_TEXT_CREDENTIAL_VALUE})",
re.IGNORECASE,
)
_TEXT_CREDENTIAL_ASSIGN_RE = re.compile(
rf"(?P<key>(?<![\w\-]){_TEXT_CREDENTIAL_NAME})"
rf"(?P<sep>\s*[:=]\s*)"
rf"(?P<value>{_TEXT_CREDENTIAL_VALUE})",
re.IGNORECASE,
)
_TEXT_URL_USERINFO_RE = re.compile(
r"(?P<scheme>\b[A-Za-z][A-Za-z0-9+.\-]*://)[^\s/@]+@"
)
@dataclass(frozen=True)
class InventorySection:
@@ -225,6 +254,42 @@ def scrub(value: Any, *, key: str | None = None) -> Any:
return repr(value)
def collapse_home(text: str) -> str:
"""Collapse every ``$HOME`` occurrence *inside* a string, not just a prefix."""
home = os.path.expanduser("~")
if not home or home == "/":
return text
return text.replace(home, "~")
def scrub_text(value: Any) -> Any:
"""Redact a free-form text blob such as a recorded command line.
:func:`scrub` keys off structured field *names* and whole-value prefixes,
which is right for inventory records but blind to a secret embedded in the
middle of a sentence. This collapses ``$HOME`` and redacts credential-shaped
tokens and URL userinfo *anywhere* in the string, so operator-supplied text
rendered verbatim contamination ``command_summary`` (#630) above all — is
held to the same standard as every other field on the page.
Returns ``None`` unchanged so callers can keep "absent" distinct from "".
"""
if value is None:
return None
text = value if isinstance(value, str) else str(value)
text = collapse_home(text)
text = _TEXT_URL_USERINFO_RE.sub(
lambda m: f"{m.group('scheme')}{_REDACTED}@", text
)
text = _TEXT_CREDENTIAL_FLAG_RE.sub(
lambda m: f"{m.group('key')}{m.group('sep')}{_REDACTED}", text
)
text = _TEXT_CREDENTIAL_ASSIGN_RE.sub(
lambda m: f"{m.group('key')}{m.group('sep')}{_REDACTED}", text
)
return text
# ── control-plane database (read-only) ───────────────────────────────────────
+1 -6
View File
@@ -48,7 +48,7 @@ NAV_GROUPS: tuple[NavGroup, ...] = (
)),
NavGroup("Runtime/Sessions", (
NavItem("/runtime", "Runtime health"),
NavItem("/sessions", "Sessions", "stub"),
NavItem("/sessions", "Sessions"),
)),
NavGroup("Projects", (
NavItem("/projects", "Projects"),
@@ -76,11 +76,6 @@ NAV_GROUPS: tuple[NavGroup, ...] = (
# issues of epic #631. Each maps a path to (title, description). Routes are
# registered so nav links resolve to a graceful, read-only stub page.
STUB_PAGES: dict[str, tuple[str, str]] = {
"/sessions": (
"Sessions",
"Active session, capability, and role inventory. Backed by the unified "
"inventory API (#636) once it lands.",
),
"/inventory": (
"Inventory",
"Unified sessions, leases, locks, namespaces, and worktree inventory. "
+4 -2
View File
@@ -88,5 +88,7 @@ def render_runtime_page(snapshot: RuntimeSnapshot) -> str:
"<p class='muted'>MVP is read-only — restart MCP servers from your IDE/operator "
"workflow. Related issue: <code>#420</code>. Guidance: "
f"<code>{html.escape(snapshot.restart_guidance)}</code></p>"
"<p class='muted'>This page does not expose tokens or perform MCP restarts.</p>"
)
"<p class='muted'>This page does not expose tokens or perform MCP restarts. "
"Correlated sessions, worktree bindings, and contamination markers: "
"<a href='/sessions'>/sessions</a> (#641).</p>"
)
+517
View File
@@ -0,0 +1,517 @@
"""Compose runtime health + inventory into a sessions/runtime view (#641).
Phase 1 is read-only. It correlates namespaces, sessions, capabilities,
worktree bindings, lease ownership, stale flags, and contamination markers
when they are detectable on disk (#630 / #671). It never restarts, kills, or
takes over a session.
Sources:
* :mod:`webui.runtime_health` profile, role, stale runtime, shell health.
* :mod:`webui.inventory` sessions, leases, locks, worktrees, namespaces
and collision signals from the control-plane DB + filesystem.
* :mod:`mcp_session_state` durable contamination markers (inspect only).
Secrets are never read. Inventory-sourced values arrive already redacted by
:func:`webui.inventory.scrub`; free-form marker text this module loads itself is
put through :func:`webui.inventory.scrub_text`, which collapses ``$HOME`` and
redacts credential-shaped tokens *inside* a string rather than only at its start.
Ownership columns are authority-aware. A lease or lock section that could not be
read renders as ``unknown``, never as ``none`` or ``unbound``: a lease the reader
could not load is not an absent lease (see
:attr:`webui.inventory.InventorySnapshot.ownership_authority_complete`).
"""
from __future__ import annotations
from dataclasses import dataclass
from typing import Any, Callable
import mcp_session_state
from webui.inventory import (
OWNERSHIP_SECTIONS,
STATUS_OK,
InventorySection,
InventorySnapshot,
load_inventory_snapshot,
scrub_text,
snapshot_to_dict as inventory_snapshot_to_dict,
)
from webui.runtime_health import (
RuntimeSnapshot,
load_runtime_snapshot,
snapshot_to_dict as runtime_snapshot_to_dict,
)
# Sanctioned recovery pointers only — never pkill / killall (#630).
SANCTIONED_RECOVERY_DOCS: tuple[dict[str, str], ...] = (
{
"label": "MCP namespace EOF recovery (reconnect only)",
"path": "docs/mcp-namespace-eof-recovery.md",
"note": "IDE/client reconnect or operator-owned restart; never kill daemons.",
},
{
"label": "MCP namespace health",
"path": "docs/mcp-namespace-health.md",
"note": "client_namespace probe proves namespace health.",
},
{
"label": "Restart path inventory",
"path": "docs/mcp-restart-path-inventory.md",
"note": "Catalog of sanctioned reconnect/restart paths.",
},
{
"label": "Local web UI recovery",
"path": "docs/webui-local-dev.md",
"note": "Operator console start and documented recovery sequence.",
},
)
_CONTAMINATION_KINDS: tuple[str, ...] = (
mcp_session_state.KIND_RUNTIME_RECOVERY_CONTAMINATION,
mcp_session_state.KIND_STABLE_BRANCH_CONTAMINATION,
)
#: Status recorded on a row when the backing inventory section is absent
#: entirely — distinct from a section that reported itself degraded.
AUTHORITY_MISSING = "missing"
def _section_status(section: InventorySection | None) -> str:
"""Status of an ownership section, treating an absent section as missing."""
if section is None:
return AUTHORITY_MISSING
return section.status
def _combined_authority(*statuses: str) -> str:
"""Worst status of the sections a derived column depends on.
A column proved from two sections is only trustworthy when *both* read
cleanly, so the first non-``ok`` status wins.
"""
for status in statuses:
if status != STATUS_OK:
return status
return STATUS_OK
@dataclass(frozen=True)
class ContaminationMarker:
"""A detectable durable contamination marker (audit-safe summary)."""
kind: str
on_disk: bool
has_payload: bool
summary: str
reason_class: str | None = None
session_id: str | None = None
role: str | None = None
command_summary: str | None = None
cleared: bool = False
def to_dict(self) -> dict[str, Any]:
return {
"kind": self.kind,
"on_disk": self.on_disk,
"has_payload": self.has_payload,
"summary": self.summary,
"reason_class": self.reason_class,
"session_id": self.session_id,
"role": self.role,
"command_summary": self.command_summary,
"cleared": self.cleared,
"active": self.on_disk and self.has_payload and not self.cleared,
}
@dataclass(frozen=True)
class SessionRow:
"""One correlated session row for the sessions table."""
session_id: str
role: str | None
profile: str | None
namespace: str | None
pid: int | None
pid_alive: bool | None
status: str | None
started_at: str | None
last_heartbeat_at: str | None
lease_ids: tuple[str, ...] = ()
work_refs: tuple[str, ...] = ()
worktree_paths: tuple[str, ...] = ()
stale_flags: tuple[str, ...] = ()
contamination_flags: tuple[str, ...] = ()
#: Status of the section backing ``lease_ids``/``work_refs``. While this is
#: not ``ok`` those tuples mean "could not be read", never "none held".
lease_authority: str = STATUS_OK
#: Worst status across the sections backing ``worktree_paths`` (locks are
#: correlated through lease work numbers, so both must read cleanly).
worktree_authority: str = STATUS_OK
@property
def ownership_authority_complete(self) -> bool:
"""True only when this row's ownership columns are provable."""
return self.lease_authority == STATUS_OK and self.worktree_authority == STATUS_OK
def to_dict(self) -> dict[str, Any]:
return {
"session_id": self.session_id,
"role": self.role,
"profile": self.profile,
"namespace": self.namespace,
"pid": self.pid,
"pid_alive": self.pid_alive,
"status": self.status,
"started_at": self.started_at,
"last_heartbeat_at": self.last_heartbeat_at,
"lease_ids": list(self.lease_ids),
"work_refs": list(self.work_refs),
"worktree_paths": list(self.worktree_paths),
"stale_flags": list(self.stale_flags),
"contamination_flags": list(self.contamination_flags),
"is_stale": bool(self.stale_flags),
"is_contaminated": bool(self.contamination_flags),
"lease_authority": self.lease_authority,
"worktree_authority": self.worktree_authority,
"lease_authority_complete": self.lease_authority == STATUS_OK,
"worktree_authority_complete": self.worktree_authority == STATUS_OK,
"ownership_authority_complete": self.ownership_authority_complete,
"ownership_note": (
"Lease and worktree columns are proved from sections that read "
"cleanly."
if self.ownership_authority_complete
else "An ownership source could not be read; empty lease_ids or "
"worktree_paths on this row mean unknown, not unowned."
),
}
@dataclass(frozen=True)
class SessionViewSnapshot:
"""Composed runtime + session inventory view (#641)."""
runtime: RuntimeSnapshot
inventory: InventorySnapshot
sessions: tuple[SessionRow, ...]
contamination_markers: tuple[ContaminationMarker, ...]
recovery_docs: tuple[dict[str, str], ...] = SANCTIONED_RECOVERY_DOCS
fetch_error: str | None = None
@property
def stale_session_count(self) -> int:
return sum(1 for row in self.sessions if row.stale_flags)
@property
def contaminated_session_count(self) -> int:
return sum(1 for row in self.sessions if row.contamination_flags)
@property
def active_contamination(self) -> tuple[ContaminationMarker, ...]:
return tuple(m for m in self.contamination_markers if m.to_dict()["active"])
@property
def ownership_authority_complete(self) -> bool:
"""False while any ownership section is degraded, absent, or unavailable."""
return self.inventory.ownership_authority_complete
@property
def ownership_section_status(self) -> dict[str, str]:
"""Per-section status for the three ownership-bearing sections."""
return {
name: _section_status(self.inventory.section(name))
for name in OWNERSHIP_SECTIONS
}
def _inspect_contamination(
kind: str,
*,
remote: str | None,
inspect: Callable[..., dict[str, Any]] | None = None,
load: Callable[..., dict[str, Any] | None] | None = None,
) -> ContaminationMarker:
"""Inspect one contamination kind; never raises into the page render path."""
inspect_fn = inspect or mcp_session_state.inspect_state_envelope
load_fn = load or mcp_session_state.load_state
try:
envelope = inspect_fn(kind=kind, remote=remote)
except Exception as exc: # noqa: BLE001 — fail soft for the dashboard
return ContaminationMarker(
kind=kind,
on_disk=False,
has_payload=False,
summary=f"contamination inspect failed: {type(exc).__name__}",
)
reason_class = None
session_id = None
role = None
command_summary = None
cleared = False
summary = scrub_text(str(envelope.get("summary") or ""))
if envelope.get("on_disk") and envelope.get("has_payload"):
try:
payload = load_fn(kind=kind, remote=remote) or {}
except Exception: # noqa: BLE001
payload = {}
if isinstance(payload, dict):
reason_class = payload.get("reason_class")
session_id = payload.get("session_id")
role = payload.get("role")
command_summary = payload.get("command_summary") or payload.get("detail")
cleared = bool(payload.get("cleared_by_reconciler"))
if not summary:
summary = scrub_text(
f"{kind}: {reason_class or 'present'}"
+ (" (cleared)" if cleared else "")
)
# Marker payloads are operator-supplied free text and are the one thing on
# this page that does not arrive through inventory scrubbing. The write-time
# redactor is a narrow denylist, so redact again at the display boundary:
# it leaves $HOME paths, `-H 'X-Api-Key: …'`, `--password …`, and
# `PRIVATE_KEY=…` intact. The field itself stays — it is #630 evidence.
return ContaminationMarker(
kind=kind,
on_disk=bool(envelope.get("on_disk")),
has_payload=bool(envelope.get("has_payload")),
summary=summary or f"{kind}: not present",
reason_class=scrub_text(str(reason_class)) if reason_class else None,
session_id=scrub_text(str(session_id)) if session_id else None,
role=scrub_text(str(role)) if role else None,
command_summary=scrub_text(str(command_summary)) if command_summary else None,
cleared=cleared,
)
def _load_contamination_markers(
*,
remote: str | None,
inspect: Callable[..., dict[str, Any]] | None = None,
load: Callable[..., dict[str, Any] | None] | None = None,
) -> tuple[ContaminationMarker, ...]:
return tuple(
_inspect_contamination(kind, remote=remote, inspect=inspect, load=load)
for kind in _CONTAMINATION_KINDS
)
def _build_session_rows(
inventory: InventorySnapshot,
contamination: tuple[ContaminationMarker, ...],
) -> tuple[SessionRow, ...]:
sessions_section = inventory.section("sessions")
leases_section = inventory.section("leases")
locks_section = inventory.section("locks")
# Ownership columns may only assert absence when their source read cleanly.
lease_authority = _section_status(leases_section)
# Locks carry claimant profile/username, not control-plane session ids, so a
# worktree binding is correlated through lease work numbers: it depends on
# the locks *and* the leases section.
worktree_authority = _combined_authority(
lease_authority, _section_status(locks_section)
)
leases_by_session: dict[str, list[dict[str, Any]]] = {}
for lease in (leases_section.items if leases_section else ()):
sid = str(lease.get("session_id") or "")
if sid:
leases_by_session.setdefault(sid, []).append(lease)
active_markers = [m for m in contamination if m.to_dict()["active"]]
marker_session_ids = {
m.session_id for m in active_markers if m.session_id
}
rows: list[SessionRow] = []
for raw in sessions_section.items if sessions_section else ():
sid = str(raw.get("session_id") or "")
if not sid:
continue
session_leases = leases_by_session.get(sid, [])
lease_ids = tuple(
str(lease["lease_id"])
for lease in session_leases
if lease.get("lease_id")
)
work_refs: list[str] = []
work_numbers: list[int] = []
for lease in session_leases:
kind = lease.get("work_kind")
number = lease.get("work_number")
if kind and number is not None:
work_refs.append(f"{kind}#{number}")
try:
work_numbers.append(int(number))
except (TypeError, ValueError):
pass
worktree_paths: list[str] = []
for lock in (locks_section.items if locks_section else ()):
try:
issue_no = int(lock.get("issue_number"))
except (TypeError, ValueError):
continue
if issue_no in work_numbers and lock.get("worktree_path"):
worktree_paths.append(str(lock["worktree_path"]))
stale_flags: list[str] = []
if raw.get("pid_alive") is False:
stale_flags.append("pid-dead")
status = str(raw.get("status") or "").lower()
if status and status not in {"active", "alive", "running", "ok"}:
stale_flags.append(f"status:{status}")
for lease in session_leases:
if lease.get("expired") is True:
stale_flags.append("lease-expired")
if str(lease.get("status") or "").lower() == "active" and lease.get(
"expired"
) is True:
stale_flags.append("active-lease-past-expiry")
contamination_flags: list[str] = []
if sid in marker_session_ids:
for marker in active_markers:
if marker.session_id == sid:
contamination_flags.append(marker.kind)
# Process-wide contamination with no session binding still surfaces
# against every live session so it cannot be silent (#630).
for marker in active_markers:
if not marker.session_id and marker.kind not in contamination_flags:
contamination_flags.append(f"{marker.kind}:process-wide")
rows.append(
SessionRow(
session_id=sid,
role=raw.get("role"),
profile=raw.get("profile"),
namespace=raw.get("namespace"),
pid=raw.get("pid") if isinstance(raw.get("pid"), int) else None,
pid_alive=raw.get("pid_alive")
if isinstance(raw.get("pid_alive"), bool)
else None,
status=raw.get("status"),
started_at=raw.get("started_at"),
last_heartbeat_at=raw.get("last_heartbeat_at"),
lease_ids=lease_ids,
work_refs=tuple(work_refs),
worktree_paths=tuple(worktree_paths),
stale_flags=tuple(dict.fromkeys(stale_flags)),
contamination_flags=tuple(dict.fromkeys(contamination_flags)),
lease_authority=lease_authority,
worktree_authority=worktree_authority,
)
)
return tuple(rows)
def load_session_view_snapshot(
*,
load_runtime: Callable[..., RuntimeSnapshot] | None = None,
load_inventory: Callable[..., InventorySnapshot] | None = None,
inspect_contamination: Callable[..., dict[str, Any]] | None = None,
load_contamination_payload: Callable[..., dict[str, Any] | None] | None = None,
) -> SessionViewSnapshot:
"""Build the composed sessions/runtime view. Fail-soft on partial sources."""
runtime_loader = load_runtime or load_runtime_snapshot
inventory_loader = load_inventory or load_inventory_snapshot
fetch_error: str | None = None
try:
runtime = runtime_loader()
except Exception as exc: # noqa: BLE001
fetch_error = f"runtime snapshot failed: {type(exc).__name__}: {exc}"
# Minimal placeholder so the page still renders inventory + recovery.
from webui.runtime_health import RuntimeSnapshot as _RS
runtime = _RS(
project_id="unknown",
repo_root="",
remote="",
host="",
profile_name="unknown",
role_kind="unknown",
config_model="unknown",
profile_mode="unknown",
profile_source="unknown",
authenticated_username=None,
identity_error=str(exc),
repo_sha=None,
remote_master_sha=None,
commits_behind_master=None,
stale_runtime_warning=None,
shell_health={},
workflow_hashes=(),
schema_hashes=(),
restart_guidance="docs/mcp-namespace-eof-recovery.md",
fetch_error=str(exc),
)
try:
inventory = inventory_loader()
except Exception as exc: # noqa: BLE001
msg = f"inventory snapshot failed: {type(exc).__name__}: {exc}"
fetch_error = f"{fetch_error}; {msg}" if fetch_error else msg
inventory = load_inventory_snapshot(
db_path="/nonexistent-for-fail-soft",
lock_dir="/nonexistent-for-fail-soft",
)
remote = getattr(runtime, "remote", None)
contamination = _load_contamination_markers(
remote=remote,
inspect=inspect_contamination,
load=load_contamination_payload,
)
sessions = _build_session_rows(inventory, contamination)
return SessionViewSnapshot(
runtime=runtime,
inventory=inventory,
sessions=sessions,
contamination_markers=contamination,
fetch_error=fetch_error,
)
def snapshot_to_dict(snapshot: SessionViewSnapshot) -> dict[str, Any]:
"""JSON export for ``/api/sessions`` (read-only)."""
return {
"api_version": "v1",
"view": "runtime-sessions",
"issue": 641,
"fetch_error": snapshot.fetch_error,
"runtime": runtime_snapshot_to_dict(snapshot.runtime),
"inventory": inventory_snapshot_to_dict(snapshot.inventory),
"sessions": [row.to_dict() for row in snapshot.sessions],
"session_counts": {
"total": len(snapshot.sessions),
"stale": snapshot.stale_session_count,
"contaminated": snapshot.contaminated_session_count,
},
"ownership_authority_complete": snapshot.ownership_authority_complete,
"ownership_section_status": snapshot.ownership_section_status,
"ownership_note": (
"Every ownership source read cleanly; a session with no lease and no "
"worktree path genuinely holds neither."
if snapshot.ownership_authority_complete
else "An ownership source is degraded or unavailable. Empty lease_ids "
"and worktree_paths mean unknown, not unowned; consult each row's "
"lease_authority and worktree_authority."
),
"contamination_markers": [
marker.to_dict() for marker in snapshot.contamination_markers
],
"active_contamination": [
marker.to_dict() for marker in snapshot.active_contamination
],
"recovery_docs": [dict(doc) for doc in snapshot.recovery_docs],
"read_only": True,
"phase": 1,
"mutations": [],
}
+403
View File
@@ -0,0 +1,403 @@
"""HTML views for the Runtime and session view (Phase 1, #641).
Read-only composition of runtime health (#430) and inventory sessions /
namespaces / worktrees (#636). Surfaces stale and contamination indicators
when detectable. Recovery links point only at sanctioned reconnect/restart
docs never at manual process kill (#630).
"""
from __future__ import annotations
from html import escape
from typing import Sequence
from webui.inventory import STATUS_OK
from webui.layout import render_page
from webui.session_loader import (
ContaminationMarker,
SessionRow,
SessionViewSnapshot,
)
def _badge(text: str, css: str) -> str:
return f'<span class="badge {css}">{escape(text)}</span>'
def _flags(flags: Sequence[str], *, css: str) -> str:
if not flags:
return '<span class="muted">—</span>'
return " ".join(_badge(flag, css) for flag in flags)
def _unproven_cell(status: str) -> str:
"""Render an ownership column whose backing inventory section failed to read.
Never "none" and never "unbound": an unreadable source proves nothing about
ownership, and claiming otherwise is the exact failure the
``ownership_authority_complete`` invariant exists to prevent.
"""
return (
f'<span class="muted">unknown (inventory {escape(status)})</span><br>'
f'{_badge("authority unproven", "badge-health-degraded")}'
)
def _runtime_banner(snapshot: SessionViewSnapshot) -> str:
runtime = snapshot.runtime
stale = runtime.stale_runtime_warning
stale_html = ""
if stale:
stale_html = (
f'<div class="health-card health-stale" style="margin-top:0.75rem;">'
f"<strong>Stale runtime:</strong> {escape(stale)}</div>"
)
identity = runtime.authenticated_username or "unresolved"
if runtime.identity_error:
identity = f"unresolved ({runtime.identity_error})"
return f"""<div class="health-card">
<h3>Runtime context</h3>
<table class="detail">
<tr><th>Profile</th><td><code>{escape(runtime.profile_name)}</code></td></tr>
<tr><th>Role kind</th><td>{escape(runtime.role_kind)}</td></tr>
<tr><th>Identity</th><td>{escape(str(identity))}</td></tr>
<tr><th>Remote / host</th>
<td><code>{escape(runtime.remote)}</code> · <code>{escape(runtime.host)}</code></td>
</tr>
<tr><th>Local HEAD</th>
<td><code>{escape(runtime.repo_sha or "unknown")}</code></td>
</tr>
<tr><th>Remote master</th>
<td><code>{escape(runtime.remote_master_sha or "unknown")}</code></td>
</tr>
<tr><th>Commits behind</th>
<td>{escape(str(runtime.commits_behind_master if runtime.commits_behind_master is not None else "unknown"))}</td>
</tr>
</table>
<p class="muted">Full runtime detail: <a href="/runtime">/runtime</a> ·
Inventory API: <a href="/api/v1/inventory"><code>/api/v1/inventory</code></a></p>
{stale_html}
</div>"""
def _summary_bar(snapshot: SessionViewSnapshot) -> str:
total = len(snapshot.sessions)
stale = snapshot.stale_session_count
contaminated = snapshot.contaminated_session_count
active_markers = len(snapshot.active_contamination)
inv_status = snapshot.inventory.status
authority_complete = snapshot.ownership_authority_complete
authority_text = "complete" if authority_complete else "incomplete"
authority_css = "badge-health-ok" if authority_complete else "badge-health-degraded"
return f"""<div class="health-card" style="display:flex; flex-wrap:wrap; gap:1rem; align-items:center;">
<div><strong>Sessions:</strong> <span class="badge badge-health-ok">{total}</span></div>
<div><strong>Stale:</strong> <span class="badge badge-stale">{stale}</span></div>
<div><strong>Contaminated:</strong> <span class="badge badge-blocked">{contaminated}</span></div>
<div><strong>Active markers:</strong> <span class="badge badge-health-unproven">{active_markers}</span></div>
<div><strong>Inventory:</strong> <span class="badge badge-health-skipped">{escape(inv_status)}</span></div>
<div><strong>Ownership authority:</strong> <span class="badge {authority_css}">{authority_text}</span></div>
</div>"""
def _ownership_caveat(snapshot: SessionViewSnapshot) -> str:
"""Name the unreadable ownership sections, or render nothing when all read."""
if snapshot.ownership_authority_complete:
return ""
degraded = ", ".join(
f"{name}: {status}"
for name, status in snapshot.ownership_section_status.items()
if status != STATUS_OK
)
return (
'<div class="health-card health-stale">'
"<strong>Ownership authority incomplete:</strong> "
f"{escape(degraded)}. Columns marked <em>unknown</em> could not be read. "
"No session below may be treated as holding no lease or no worktree "
"binding — absence of evidence is not evidence of absence.</div>"
)
def _render_session_row(row: SessionRow) -> str:
pid = "" if row.pid is None else str(row.pid)
pid_alive = "" if row.pid_alive is None else ("alive" if row.pid_alive else "dead")
pid_css = (
"badge-health-ok"
if row.pid_alive is True
else ("badge-blocked" if row.pid_alive is False else "badge-health-skipped")
)
# An empty tuple only means "holds none" when its source read cleanly.
if row.lease_authority != STATUS_OK:
lease_cell = _unproven_cell(row.lease_authority)
else:
leases = (
", ".join(f"<code>{escape(lid)}</code>" for lid in row.lease_ids)
if row.lease_ids
else '<span class="muted">none</span>'
)
work = (
", ".join(escape(ref) for ref in row.work_refs)
if row.work_refs
else '<span class="muted">—</span>'
)
lease_cell = (
f'{leases}<div class="muted" style="font-size:0.82rem; '
f'margin-top:0.2rem;">{work}</div>'
)
if row.worktree_authority != STATUS_OK:
worktrees = _unproven_cell(row.worktree_authority)
else:
worktrees = (
"<br>".join(f"<code>{escape(path)}</code>" for path in row.worktree_paths)
if row.worktree_paths
else '<span class="muted">unbound</span>'
)
return f"""<tr>
<td><code>{escape(row.session_id)}</code></td>
<td>
<div><code>{escape(str(row.role or ""))}</code> / <code>{escape(str(row.profile or ""))}</code></div>
<div class="muted" style="font-size:0.82rem;">ns: <code>{escape(str(row.namespace or ""))}</code></div>
</td>
<td>
<code>{escape(pid)}</code>
{_badge(pid_alive, pid_css)}
</td>
<td>{escape(str(row.status or ""))}<div class="muted" style="font-size:0.82rem;">{escape(str(row.last_heartbeat_at or ""))}</div></td>
<td>{lease_cell}</td>
<td>{worktrees}</td>
<td>{_flags(row.stale_flags, css="badge-stale")}</td>
<td>{_flags(row.contamination_flags, css="badge-blocked")}</td>
</tr>"""
def _sessions_table(snapshot: SessionViewSnapshot) -> str:
rows: Sequence[SessionRow] = snapshot.sessions
if not rows:
sessions_status = snapshot.ownership_section_status.get("sessions", STATUS_OK)
if sessions_status != STATUS_OK:
return (
'<p class="muted">Session inventory is '
f"<strong>{escape(sessions_status)}</strong> — the session list "
"could not be read. This is not evidence that no sessions "
"exist.</p>"
)
return (
'<p class="muted">No control-plane sessions recorded. Inventory may '
"be unavailable, or no MCP workers have registered yet.</p>"
)
body = "".join(_render_session_row(row) for row in rows)
return f"""<table class="registry">
<thead>
<tr>
<th>Session</th>
<th>Role / profile / namespace</th>
<th>PID</th>
<th>Status</th>
<th>Leases / work</th>
<th>Worktree binding</th>
<th>Stale</th>
<th>Contamination</th>
</tr>
</thead>
<tbody>
{body}
</tbody>
</table>"""
def _namespaces_section(snapshot: SessionViewSnapshot) -> str:
section = snapshot.inventory.section("namespaces")
if section is None:
return (
'<div class="prompt-card"><h3>Namespaces</h3>'
'<p class="muted">Namespaces section not loaded.</p></div>'
)
if not section.ok:
return f"""<div class="prompt-card">
<h3>Namespaces {_badge(section.status, "badge-health-degraded")}</h3>
<p class="muted">{escape(section.reason or "unavailable")}</p>
</div>"""
rows = []
for item in section.items:
caps = item.get("capability_summary") or {}
cap_bits = ", ".join(
name for name, ok in sorted(caps.items()) if ok
) or "none"
rows.append(
"<tr>"
f"<td><code>{escape(str(item.get('mcp_namespace') or ''))}</code></td>"
f"<td><code>{escape(str(item.get('profile_name') or ''))}</code></td>"
f"<td>{escape(str(item.get('role') or ''))}</td>"
f"<td>{escape(cap_bits)}</td>"
f"<td>{'yes' if item.get('active') else 'no'}</td>"
"</tr>"
)
reason = (
f'<p class="muted">{escape(section.reason)}</p>'
if section.reason
else ""
)
return f"""<div class="prompt-card">
<h3>Namespaces / capabilities</h3>
{reason}
<table class="registry">
<thead>
<tr>
<th>Namespace</th>
<th>Profile</th>
<th>Role</th>
<th>Capabilities</th>
<th>Active in process</th>
</tr>
</thead>
<tbody>
{"".join(rows) if rows else '<tr><td colspan="5" class="muted">No namespace rows.</td></tr>'}
</tbody>
</table>
</div>"""
def _worktrees_section(snapshot: SessionViewSnapshot) -> str:
section = snapshot.inventory.section("worktrees")
if section is None:
return ""
if not section.ok and not section.items:
return f"""<div class="prompt-card">
<h3>Worktrees {_badge(section.status, "badge-health-degraded")}</h3>
<p class="muted">{escape(section.reason or "unavailable")}</p>
</div>"""
rows = []
for item in section.items[:50]:
rows.append(
"<tr>"
f"<td><code>{escape(str(item.get('rel_path') or item.get('path') or ''))}</code></td>"
f"<td><code>{escape(str(item.get('branch') or ''))}</code></td>"
f"<td>{escape(str(item.get('classification') or ''))}</td>"
f"<td>{'yes' if item.get('registered_worktree') else 'no'}</td>"
f"<td>{'dirty' if item.get('dirty') else 'clean'}</td>"
"</tr>"
)
more = ""
if len(section.items) > 50:
more = f'<p class="muted">Showing 50 of {len(section.items)}. Full list: <a href="/worktrees">/worktrees</a>.</p>'
return f"""<div class="prompt-card">
<h3>Worktree bindings</h3>
<p class="muted">Registered issue worktrees under <code>branches/</code>. Hygiene detail: <a href="/worktrees">/worktrees</a>.</p>
<table class="registry">
<thead>
<tr>
<th>Path</th>
<th>Branch</th>
<th>Classification</th>
<th>Registered</th>
<th>State</th>
</tr>
</thead>
<tbody>
{"".join(rows) if rows else '<tr><td colspan="5" class="muted">No worktrees recorded.</td></tr>'}
</tbody>
</table>
{more}
</div>"""
def _contamination_section(markers: Sequence[ContaminationMarker]) -> str:
if not markers:
return (
'<div class="prompt-card"><h3>Contamination markers</h3>'
'<p class="muted">No contamination kinds inspected.</p></div>'
)
rows = []
for marker in markers:
active = marker.to_dict()["active"]
status = "ACTIVE" if active else ("cleared" if marker.cleared else "absent")
css = "badge-blocked" if active else "badge-health-ok"
rows.append(
"<tr>"
f"<td><code>{escape(marker.kind)}</code></td>"
f"<td>{_badge(status, css)}</td>"
f"<td>{escape(marker.reason_class or '')}</td>"
f"<td><code>{escape(marker.session_id or '')}</code></td>"
f"<td>{escape(marker.command_summary or marker.summary)}</td>"
"</tr>"
)
return f"""<div class="prompt-card">
<h3>Contamination markers (#630 / #671)</h3>
<p class="muted">Durable markers only never silent when present. Clearance is reconciler-only.</p>
<table class="registry">
<thead>
<tr>
<th>Kind</th>
<th>State</th>
<th>Reason class</th>
<th>Session</th>
<th>Summary</th>
</tr>
</thead>
<tbody>
{"".join(rows)}
</tbody>
</table>
</div>"""
def _recovery_section(snapshot: SessionViewSnapshot) -> str:
items = []
for doc in snapshot.recovery_docs:
items.append(
"<li>"
f"<code>{escape(doc['path'])}</code> — "
f"<strong>{escape(doc['label'])}</strong>: {escape(doc['note'])}"
"</li>"
)
return f"""<div class="prompt-card">
<h3>Sanctioned recovery (read-only)</h3>
<p class="muted">This view does <strong>not</strong> restart, kill, or take over sessions.
Manual <code>pkill</code> / <code>kill</code> of MCP daemons is contamination (#630), not recovery.</p>
<ul class="reasons">
{"".join(items)}
<li>Prefer IDE/client reconnect (<code>/mcp reconnect</code>) or an operator-owned restart recorded in the restart inventory.</li>
</ul>
</div>"""
def render_sessions_page(snapshot: SessionViewSnapshot) -> str:
"""Render the full HTML body for the runtime/session view."""
error_block = ""
if snapshot.fetch_error:
error_block = (
f'<div class="health-card health-stale"><strong>Partial load:</strong> '
f"{escape(snapshot.fetch_error)}</div>"
)
if snapshot.runtime.fetch_error:
error_block += (
f'<div class="health-card health-stale"><strong>Runtime note:</strong> '
f"{escape(snapshot.runtime.fetch_error)}</div>"
)
body = f"""
{error_block}
{_runtime_banner(snapshot)}
{_summary_bar(snapshot)}
<div class="prompt-card">
<h3>Sessions</h3>
<p class="muted">Control-plane sessions correlated with leases and worktree bindings.
Stale and contamination flags are fail-soft: absence of a marker is not proof of cleanliness when inventory is degraded.
Lease and worktree columns read <em>unknown (inventory )</em> when their source could not be loaded.</p>
{_ownership_caveat(snapshot)}
{_sessions_table(snapshot)}
</div>
{_namespaces_section(snapshot)}
{_worktrees_section(snapshot)}
{_contamination_section(snapshot.contamination_markers)}
{_recovery_section(snapshot)}
"""
return render_page(title="Sessions", body_html=f"""<h2>Runtime and sessions</h2>
<p class="meta">Phase 1 read-only view (#641). Combines runtime health (#430) with
unified inventory sessions/namespaces/worktrees (#636). No restart or session-takeover controls.</p>
{body}""")
+4 -2
View File
@@ -278,8 +278,10 @@ def _recovery_card() -> str:
"controls arrive in Phase 2 (#642); until then recovery runs through "
"the sanctioned client reconnect / operator restart path.</p>"
"<ul class='reasons'>"
"<li><a href='/runtime'>Runtime and session view</a> — active profile, "
"workflow hashes, and shell health.</li>"
"<li><a href='/runtime'>Runtime health</a> — active profile, workflow "
"hashes, and shell health.</li>"
"<li><a href='/sessions'>Runtime and sessions</a> — namespaces, session "
"rows, worktree bindings, and contamination markers (#641).</li>"
"<li>Reconnect the MCP client from the IDE, then re-run the blocked "
"cycle. Never kill the daemon process manually: unmanaged kills are "
"recorded as runtime contamination (#630).</li>"