Author SHA1 Message Date
sysadminandClaude Opus 4.8 9bc021e9c0 Merge master into feat/issue-663-restart-classes (resolve #886 conflict)
Brings PR #886 up to date with master @ 2f4dec8323
(8 commits behind), resolving the single conflicted file.

Conflict: gitea_mcp_server.py, both hunks inside gitea_request_mcp_restart.
Both sides were purely additive to the same tool, so both are kept in full:

- Branch side (#663, restart classes): parameters restart_class,
  target_session_id, target_role, target_connector; payload keys
  controller_approval_authorized, requester_role, requester_permissions.
- Master side (#661 via PR #882, drain-proof hard gate): parameters
  drain_proof_json, request_break_glass; the explanatory comment describing
  the apply-path hard gate and break-glass authorization.

No behaviour from either side was dropped, reordered, or reimplemented. Every
parameter from both sides is already consumed by the auto-merged function body
(restart_class and the three target_* arguments flow into the coordinator call;
drain_proof_json and request_break_glass drive the dry_run=False hard gate), so
the union is the only resolution that keeps the merged function coherent.

Validation on the merged tree:

  python -m pytest tests/test_drain_proof.py tests/test_restart_classes.py \
    tests/test_restart_coordinator.py tests/test_mcp_restart_paths.py \
    tests/test_mcp_restart_governance_docs.py tests/test_webui_sanctioned_restart.py \
    tests/test_issue_662_post_restart_reconcile.py -q
  # 165 passed, 68 subtests passed

py_compile on gitea_mcp_server.py passes and no conflict markers remain.

Closes #663

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-07-25 01:16:09 -04:00
jcwalker3 930dc24632 Merge branch 'master' into feat/issue-663-restart-classes 2026-07-24 22:34:33 -05:00
jcwalker3 41622c5985 Merge branch 'master' into feat/issue-663-restart-classes 2026-07-24 21:27:15 -05:00
jcwalker3 301c78de20 Merge branch 'master' into feat/issue-663-restart-classes 2026-07-24 21:06:21 -05:00
sysadmin 714190e02a feat: enforce MCP restart class permissions (#663) 2026-07-24 18:10:22 -04:00
16 changed files with 737 additions and 1814 deletions
+53
View File
@@ -0,0 +1,53 @@
# MCP restart classes and blast-radius permissions (#663)
This is the machine-enforced class matrix used by
`restart_coordinator.RESTART_CLASS_POLICIES`. It implements the narrower-first
recovery ladder from #655 and the authorization policy from #656, using the
path inventory from #657 and the impact coordinator from #658. Product and
delivery lineage: vision #652 and roadmap #653.
Unknown class names are denied. The coordinator requires both the class
permission and an eligible request role. Approval gates are additional: a
caller cannot turn a request permission into execution authority.
| Restart class | Required permission | Expected blast radius | Drain requirement | Approval requirement | Audit requirement | Recovery behavior |
|---|---|---|---|---|---|---|
| `client_reconnect` | `mcp.reconnect.client` | none | none | self service | class, actor, client namespace, reason, outcome | Reconnect only the caller's client transport. No daemon or peer work changes. |
| `session_reconnect` | `mcp.reconnect.session` | low | requesting-session safe point | self service | class, actor, session, reason, outcome | Rebind identity, capability, and workspace state for one session. |
| `worker_restart` | `mcp.restart.worker.request` | low | target worker | controller approval + automated gates | class, actor, worker, approval, scoped drain, outcome | Restart one worker after its own leases and mutations drain. |
| `role_runtime_restart` | `mcp.restart.role_runtime.request` | medium | target role runtime | controller approval + automated gates | class, actor, role namespace, approval, scoped drain, outcome | Restart and re-probe one role runtime; unrelated roles remain available. |
| `connector_restart` | `mcp.restart.connector.request` | medium | target connector | controller approval + automated gates | class, actor, connector, approval, scoped drain, outcome | Restart one connector while unrelated runtimes remain available. |
| `configuration_reload` | `mcp.reload.configuration.request` | low | mutation quiesce | controller approval + automated gates | class, actor, configuration revision, approval, outcome | Gracefully reload configuration without replacing the daemon. |
| `rolling_mcp_restart` | `mcp.restart.rolling.request` | medium | one instance at a time | controller approval + automated gates | class, actor, instance order, approval, per-instance drains, outcome | Drain, restart, verify, and restore each instance before advancing. |
| `full_mcp_restart` | `mcp.restart.full.request` | high | all sessions and mutations | controller approval + automated gates | class, actor, full impact, approval, full drain proof, outcome | Replace the complete MCP runtime only after a verified full drain. |
| `host_restart` | `mcp.restart.host.request` | high | all host work | controller approval + infrastructure operator | class, actor, host/change or incident id, approval, full drain proof, outcome | Hand off to infrastructure ownership and reconcile every runtime afterward. |
## Drain boundary
Only `full_mcp_restart` and `host_restart` set `full_drain_required=true`.
Reconnects and configuration reloads do not disrupt peer sessions. Worker,
role-runtime, and connector restarts evaluate only their explicitly named
target. Rolling restart drains one instance at a time. Missing required target
scope denies the request rather than silently widening it to a full restart.
## Permission and approval boundary
Author, reviewer, merger, and reconciler roles may self-request reconnects and
request scoped worker/role/connector/reload recovery. They cannot request
rolling, full, or host restart classes. Controller/operator/admin roles may
request the broader classes, while execution remains operator/admin-owned.
Controller approval is independently required for every class above a session
reconnect. Host restart additionally requires infrastructure-operator proof.
The MCP request tool derives class permissions from its authenticated runtime
role. It does not accept caller-supplied permissions. Controller and operator
authorization are read from the already-running daemon environment, never
from a request argument.
## Audit and failure behavior
Every impact audit and every console restart/reload audit includes a
`restart_class` field. The impact audit also includes the exact
`required_permission`. Unknown classes, missing permissions, ineligible roles,
missing approval, missing scoped targets, and incomplete inventory all deny
fail closed. Manual process kills remain forbidden and contaminating (#630).
+12 -2
View File
@@ -12,6 +12,12 @@ restart/reload/kill paths). The **mutative apply** path — actually performing
restart — is a later child gated by a drain proof and is explicitly out of
scope here.
The coordinator now routes every request through the restart-class policy
matrix defined for #663. See
[`mcp-restart-classes.md`](./mcp-restart-classes.md) for permissions, expected
blast radius, scoped drain and approval requirements, audit fields, and
recovery behavior for all nine classes.
## Components
| Piece | Where | Responsibility |
@@ -77,7 +83,10 @@ authorization is present.
```text
gitea_request_mcp_restart(remote, host, org, repo,
dry_run=True, request_override=False,
session_id=None, limit=200)
session_id=None, limit=200,
restart_class="full_mcp_restart",
target_session_id=None, target_role=None,
target_connector=None)
```
Read-only, dry-run, and it **never restarts anything**. `apply_supported` is
@@ -87,7 +96,8 @@ apply is gated by a drain proof (a separate child).
## Audit
Every evaluation carries an `audit_record` (event, coordinator version, verdict,
allow decision, blast radius, counts, timestamp) so restart decisions are
restart class, required permission, allow decision, blast radius, counts,
timestamp) so restart decisions are
auditable. No secrets flow through the coordinator — session ids, pids, and
profiles are operational metadata only.
+6 -43
View File
@@ -77,9 +77,7 @@ status, onboarding checklist state, and the fail-closed error payloads (#635).
| `/api/actions/{id}/preview` | Mutation ledger preview (GET, read-only) |
| `/leases` | Lease and collision visibility (#433) |
| `/api/leases` | JSON lease/collision export |
| `/sessions` | Runtime and session view (#641) — health + inventory sessions/namespaces/worktrees |
| `/api/sessions` | JSON export for the runtime/session view |
| `/api/v1/sessions` | Versioned alias of `/api/sessions` |
| `/sessions` | Phase 1 shell stub — session inventory (backed by #636) |
| `/inventory` | Phase 1 shell stub — unified inventory (backed by #636) |
| `/timeline` | Phase 1 shell stub — workflow event timeline |
| `/policy` | Phase 1 shell stub — capability/role policy placeholder |
@@ -286,46 +284,11 @@ The header carries two read-only status badges — an **environment** badge
a **mode: read-only** badge — plus a **Docs** link to this document. No
privileged action controls are present in the Phase 1 shell.
Not-yet-implemented surfaces (`/inventory`, `/timeline`, `/policy`,
`/insights`) resolve to graceful read-only stub pages instead of 404s; their
backing views land in later child issues of #631 (the inventory surfaces are
backed by #636). Mutating methods on stub routes still fail closed with
`read-only-mvp`.
### Runtime and sessions (#641)
`/sessions` is a live Phase 1 read-only view that composes:
* runtime health from `#430` (profile, role, identity, master parity, stale warning)
* control-plane sessions / leases and filesystem locks / worktrees / namespaces from `#636`
* durable contamination markers when detectable (`#630` runtime recovery, `#671` stable-branch push)
It surfaces stale indicators (dead PID, expired lease) and never silences an
active contamination marker. Recovery links point only at sanctioned
reconnect/operator restart docs (`docs/mcp-namespace-eof-recovery.md`,
`docs/mcp-namespace-health.md`, `docs/mcp-restart-path-inventory.md`, this
document). The page does **not** restart, kill, or take over sessions; manual
`pkill` of MCP daemons is contamination, not recovery.
Honesty rules specific to this view:
* **Ownership columns never assert absence they cannot prove.** When the
`leases` or `locks` section is degraded or unavailable, the Leases and
Worktree-binding cells render `unknown (inventory <status>)` with an
*authority unproven* badge instead of `none` / `unbound`, and a caveat names
the unreadable sections. A worktree binding is correlated through lease work
numbers, so it is unproven when *either* section fails to read.
`/api/sessions` carries the same facts as `ownership_authority_complete`,
`ownership_section_status`, and per-row `lease_authority` /
`worktree_authority`, so a JSON consumer can tell "holds none" from "could
not be read".
* **Contamination text is redacted at the display boundary.** Marker payloads
(`command_summary`, `reason_class`, `session_id`, `role`) are
operator-supplied free text that does not arrive through inventory scrubbing,
so they pass through `webui.inventory.scrub_text`, which collapses `$HOME` and
redacts credential-shaped tokens and URL userinfo *anywhere* in the string.
The write-time redactor is a narrow denylist and is not relied on. The field
itself is kept — it is the `#630` evidence naming which daemon was killed.
Not-yet-implemented surfaces (`/sessions`, `/inventory`, `/timeline`,
`/policy`, `/insights`) resolve to graceful read-only stub pages instead of
404s; their backing views land in later child issues of #631 (the inventory
surfaces are backed by #636). Mutating methods on stub routes still fail closed
with `read-only-mvp`.
## System-health dashboard (#639)
+29 -1
View File
@@ -22343,12 +22343,17 @@ def gitea_request_mcp_restart(
request_override: bool = False,
session_id: str | None = None,
limit: int = 200,
restart_class: str = "full_mcp_restart",
target_session_id: str | None = None,
target_role: str | None = None,
target_connector: str | None = None,
drain_proof_json: str | None = None,
request_break_glass: bool = False,
) -> dict:
"""Evaluate a proposed MCP restart and return an impact preview (#658).
Central restart coordinator: gathers live control-plane state (sessions,
Central restart coordinator: resolves the requested restart class, gathers
live control-plane state (sessions,
leases/locks, in-flight issue/PR work, mutations, worktrees) and returns a
blast-radius impact report with a ``safe`` / ``unsafe`` / ``override``
verdict, so the console (#642/#652) and operators can see what a restart
@@ -22447,6 +22452,9 @@ def gitea_request_mcp_restart(
profile = get_profile()
profile_name = (profile.get("profile_name") or "").strip() or "session"
requester_role = (
profile.get("role_kind") or profile.get("role") or ""
).strip().lower()
sid = (session_id or "").strip() or f"{profile_name}-{os.getpid()}"
# Override authority is read from the environment only — a worker session
@@ -22456,6 +22464,15 @@ def gitea_request_mcp_restart(
(os.environ.get("GITEA_OPERATOR_RESTART_OVERRIDE_AUTHORIZATION") or "").strip()
)
operator_override = bool(request_override and operator_authorized)
controller_approved = bool(
(
os.environ.get("GITEA_CONTROLLER_RESTART_APPROVAL_AUTHORIZATION")
or ""
).strip()
)
requester_permissions = restart_coordinator.permissions_for_role(
requester_role
)
inventory = {
"sessions": sessions,
@@ -22470,6 +22487,14 @@ def gitea_request_mcp_restart(
operator_override=operator_override,
requesting_session_id=sid,
dry_run=True, # coordinator is always analysis-only (#658)
restart_class=restart_class,
requester_role=requester_role,
requester_permissions=requester_permissions,
controller_approved=controller_approved,
operator_authorized=operator_authorized,
target_session_id=target_session_id,
target_role=target_role,
target_connector=target_connector,
)
payload = report.as_dict()
@@ -22481,6 +22506,9 @@ def gitea_request_mcp_restart(
payload["requesting_session_id"] = sid
payload["operator_override_requested"] = bool(request_override)
payload["operator_override_authorized"] = operator_authorized
payload["controller_approval_authorized"] = controller_approved
payload["requester_role"] = requester_role
payload["requester_permissions"] = list(requester_permissions)
# Actual restart execution remains a further child; this tool never restarts
# a process. What #661 adds is the *hard gate*: an apply request (dry_run
# False) must present a valid, unexpired, clean drain proof, or it is denied
+371 -15
View File
@@ -28,11 +28,12 @@ from __future__ import annotations
from dataclasses import dataclass, field
from datetime import datetime, timezone
from enum import Enum
from typing import Any, Mapping, Sequence
import lease_lifecycle
COORDINATOR_VERSION = "1.0.0-issue-658"
COORDINATOR_VERSION = "1.1.0-issue-663"
# Restart verdicts. Exactly the three the acceptance criteria name.
VERDICT_SAFE = "safe"
@@ -54,6 +55,194 @@ LEASE_FRESHNESS_LIVE = "active"
DEFAULT_SESSION_HEARTBEAT_STALE_SECONDS = 900
class RestartClass(str, Enum):
"""The only restart/recovery classes accepted by the coordinator."""
CLIENT_RECONNECT = "client_reconnect"
SESSION_RECONNECT = "session_reconnect"
WORKER_RESTART = "worker_restart"
ROLE_RUNTIME_RESTART = "role_runtime_restart"
CONNECTOR_RESTART = "connector_restart"
CONFIGURATION_RELOAD = "configuration_reload"
ROLLING_MCP_RESTART = "rolling_mcp_restart"
FULL_MCP_RESTART = "full_mcp_restart"
HOST_RESTART = "host_restart"
@dataclass(frozen=True)
class RestartClassPolicy:
"""Least-privilege policy for one :class:`RestartClass`."""
restart_class: RestartClass
required_permission: str
expected_blast_radius: str
drain_requirement: str
full_drain_required: bool
approval_requirement: str
audit_requirement: str
recovery_behavior: str
request_roles: tuple[str, ...]
execution_roles: tuple[str, ...]
def as_dict(self) -> dict[str, Any]:
return {
"restart_class": self.restart_class.value,
"required_permission": self.required_permission,
"expected_blast_radius": self.expected_blast_radius,
"drain_requirement": self.drain_requirement,
"full_drain_required": self.full_drain_required,
"approval_requirement": self.approval_requirement,
"audit_requirement": self.audit_requirement,
"recovery_behavior": self.recovery_behavior,
"request_roles": list(self.request_roles),
"execution_roles": list(self.execution_roles),
}
WORKER_ROLES = ("author", "reviewer", "merger", "reconciler")
CONTROL_ROLES = ("controller", "operator", "admin")
ALL_REQUEST_ROLES = WORKER_ROLES + CONTROL_ROLES
RESTART_CLASS_POLICIES: dict[RestartClass, RestartClassPolicy] = {
RestartClass.CLIENT_RECONNECT: RestartClassPolicy(
RestartClass.CLIENT_RECONNECT,
"mcp.reconnect.client",
BLAST_NONE,
"none",
False,
"self_service",
"record class, actor, client namespace, reason, and outcome",
"Reconnect only the caller's client transport; no daemon or peer session changes.",
ALL_REQUEST_ROLES,
ALL_REQUEST_ROLES,
),
RestartClass.SESSION_RECONNECT: RestartClassPolicy(
RestartClass.SESSION_RECONNECT,
"mcp.reconnect.session",
BLAST_LOW,
"requesting_session_safe_point",
False,
"self_service",
"record class, actor, session id, reason, and outcome",
"Rebind identity, capability, and workspace state for one session.",
ALL_REQUEST_ROLES,
ALL_REQUEST_ROLES,
),
RestartClass.WORKER_RESTART: RestartClassPolicy(
RestartClass.WORKER_RESTART,
"mcp.restart.worker.request",
BLAST_LOW,
"target_worker",
False,
"controller_approval_and_automated_gates",
"record class, actor, target worker, approval, drain proof, and outcome",
"Restart one worker after its own lease and mutation scope is drained.",
ALL_REQUEST_ROLES,
("operator", "admin"),
),
RestartClass.ROLE_RUNTIME_RESTART: RestartClassPolicy(
RestartClass.ROLE_RUNTIME_RESTART,
"mcp.restart.role_runtime.request",
BLAST_MEDIUM,
"target_role_runtime",
False,
"controller_approval_and_automated_gates",
"record class, actor, role namespace, approval, drain proof, and outcome",
"Restart only the selected role runtime and then re-probe that namespace.",
ALL_REQUEST_ROLES,
("operator", "admin"),
),
RestartClass.CONNECTOR_RESTART: RestartClassPolicy(
RestartClass.CONNECTOR_RESTART,
"mcp.restart.connector.request",
BLAST_MEDIUM,
"target_connector",
False,
"controller_approval_and_automated_gates",
"record class, actor, connector id, approval, drain proof, and outcome",
"Restart one connector while unrelated role runtimes remain available.",
ALL_REQUEST_ROLES,
("operator", "admin"),
),
RestartClass.CONFIGURATION_RELOAD: RestartClassPolicy(
RestartClass.CONFIGURATION_RELOAD,
"mcp.reload.configuration.request",
BLAST_LOW,
"mutation_quiesce",
False,
"controller_approval_and_automated_gates",
"record class, actor, configuration revision, approval, and outcome",
"Gracefully reload configuration without replacing the daemon process.",
ALL_REQUEST_ROLES,
("operator", "admin"),
),
RestartClass.ROLLING_MCP_RESTART: RestartClassPolicy(
RestartClass.ROLLING_MCP_RESTART,
"mcp.restart.rolling.request",
BLAST_MEDIUM,
"one_instance_at_a_time",
False,
"controller_approval_and_automated_gates",
"record class, actor, instance order, approval, per-instance drains, and outcome",
"Drain, restart, verify, and restore one instance before advancing to the next.",
CONTROL_ROLES,
("operator", "admin"),
),
RestartClass.FULL_MCP_RESTART: RestartClassPolicy(
RestartClass.FULL_MCP_RESTART,
"mcp.restart.full.request",
BLAST_HIGH,
"all_sessions_and_mutations",
True,
"controller_approval_and_automated_gates",
"record class, actor, full impact report, approval, drain proof, and outcome",
"Stop and restore the complete MCP runtime only after a verified full drain.",
CONTROL_ROLES,
("operator", "admin"),
),
RestartClass.HOST_RESTART: RestartClassPolicy(
RestartClass.HOST_RESTART,
"mcp.restart.host.request",
BLAST_HIGH,
"all_host_work",
True,
"controller_approval_plus_infrastructure_operator",
"record class, actor, host, incident or change id, approval, drain proof, and outcome",
"Hand off to infrastructure ownership; reconcile every runtime after the host returns.",
("controller", "operator", "admin"),
("operator", "admin"),
),
}
def resolve_restart_class(value: RestartClass | str) -> RestartClass:
"""Resolve a restart class or fail closed for an unknown value."""
if isinstance(value, RestartClass):
return value
try:
return RestartClass(str(value).strip())
except ValueError as exc:
raise ValueError(f"unknown restart class {value!r}; deny (fail closed)") from exc
def restart_class_policy(value: RestartClass | str) -> RestartClassPolicy:
"""Return the canonical policy for *value*."""
return RESTART_CLASS_POLICIES[resolve_restart_class(value)]
def permissions_for_role(role: str | None) -> tuple[str, ...]:
"""Return request permissions granted to a workflow role by this policy."""
normalized = str(role or "").strip().lower()
return tuple(
policy.required_permission
for policy in RESTART_CLASS_POLICIES.values()
if normalized in policy.request_roles
)
def _utc_now() -> datetime:
return datetime.now(timezone.utc)
@@ -75,6 +264,7 @@ class SessionImpact:
heartbeat_stale: bool
is_requester: bool
live: bool
connector: str | None = None
def as_dict(self) -> dict[str, Any]:
return {
@@ -87,6 +277,7 @@ class SessionImpact:
"heartbeat_stale": self.heartbeat_stale,
"is_requester": self.is_requester,
"live": self.live,
"connector": self.connector,
}
@@ -105,6 +296,7 @@ class LeaseImpact:
disruptive: bool
is_mutation: bool
is_critical_section: bool
connector: str | None = None
def as_dict(self) -> dict[str, Any]:
return {
@@ -119,6 +311,7 @@ class LeaseImpact:
"disruptive": self.disruptive,
"is_mutation": self.is_mutation,
"is_critical_section": self.is_critical_section,
"connector": self.connector,
}
@@ -127,6 +320,13 @@ class RestartImpactReport:
"""Impact preview DTO returned to the console / operator (#642/#652)."""
coordinator_version: str
restart_class: str
restart_policy: dict[str, Any]
policy_enforced: bool
permission_authorized: bool
role_authorized: bool
approval_satisfied: bool
authorization_reasons: list[str]
evaluated_at: str
dry_run: bool
restart_performed: bool
@@ -153,6 +353,13 @@ class RestartImpactReport:
def as_dict(self) -> dict[str, Any]:
return {
"coordinator_version": self.coordinator_version,
"restart_class": self.restart_class,
"restart_policy": dict(self.restart_policy),
"policy_enforced": self.policy_enforced,
"permission_authorized": self.permission_authorized,
"role_authorized": self.role_authorized,
"approval_satisfied": self.approval_satisfied,
"authorization_reasons": list(self.authorization_reasons),
"evaluated_at": self.evaluated_at,
"dry_run": self.dry_run,
"restart_performed": self.restart_performed,
@@ -206,6 +413,7 @@ def _classify_session(
requesting_session_id and session_id == requesting_session_id
),
live=live,
connector=(str(row.get("connector") or "").strip() or None),
)
@@ -258,6 +466,7 @@ def _classify_lease(row: Mapping[str, Any]) -> LeaseImpact:
disruptive=disruptive,
is_mutation=is_mutation,
is_critical_section=disruptive,
connector=(str(row.get("connector") or "").strip() or None),
)
@@ -279,6 +488,14 @@ def evaluate_restart_impact(
requesting_session_id: str | None = None,
dry_run: bool = True,
session_heartbeat_stale_seconds: int = DEFAULT_SESSION_HEARTBEAT_STALE_SECONDS,
restart_class: RestartClass | str | None = None,
requester_role: str | None = None,
requester_permissions: Sequence[str] | None = None,
controller_approved: bool = False,
operator_authorized: bool = False,
target_session_id: str | None = None,
target_role: str | None = None,
target_connector: str | None = None,
) -> RestartImpactReport:
"""Evaluate a proposed MCP restart and return an impact preview.
@@ -301,6 +518,61 @@ def evaluate_restart_impact(
"""
moment = now or _utc_now()
reasons: list[str] = []
authorization_reasons: list[str] = []
# ``None`` preserves the pre-#663 impact-only API for callers that have not
# yet been migrated. All MCP requests pass an explicit class and therefore
# take the fail-closed policy path.
policy_enforced = restart_class is not None
try:
resolved_class = resolve_restart_class(
restart_class or RestartClass.FULL_MCP_RESTART
)
policy = RESTART_CLASS_POLICIES[resolved_class]
unknown_class = False
except ValueError as exc:
resolved_class = None
policy = None
unknown_class = True
authorization_reasons.append(str(exc))
normalized_role = str(requester_role or "").strip().lower()
granted = {str(p).strip() for p in (requester_permissions or ())}
if policy_enforced and policy is not None:
permission_authorized = policy.required_permission in granted
role_authorized = normalized_role in policy.request_roles
if not permission_authorized:
authorization_reasons.append(
f"missing required permission {policy.required_permission!r}"
)
if not role_authorized:
authorization_reasons.append(
f"role {normalized_role or 'unknown'!r} may not request "
f"{policy.restart_class.value}"
)
elif unknown_class:
permission_authorized = False
role_authorized = False
else:
permission_authorized = True
role_authorized = True
if policy_enforced and policy is not None:
approval = policy.approval_requirement
if approval == "self_service":
approval_satisfied = True
elif approval == "controller_approval_plus_infrastructure_operator":
approval_satisfied = bool(controller_approved and operator_authorized)
else:
approval_satisfied = bool(controller_approved)
if not approval_satisfied:
authorization_reasons.append(
f"approval requirement not satisfied: {approval}"
)
elif unknown_class:
approval_satisfied = False
else:
approval_satisfied = True
inventory_complete = bool(inventory.get("inventory_complete", False))
incomplete_reasons = [str(r) for r in (inventory.get("incomplete_reasons") or [])]
@@ -323,15 +595,67 @@ def evaluate_restart_impact(
]
lease_impacts = [_classify_lease(l) for l in leases_raw]
# Only *other* live sessions and live leases constitute blast radius: a
# restart that would kill only the requesting session with no other work in
# flight is safe.
other_live_sessions = [
s for s in session_impacts if s.live and not s.is_requester
# Route impact through the selected class. Narrow classes never inherit a
# full-runtime drain merely because unrelated work exists.
target_complete = True
if resolved_class in {
RestartClass.CLIENT_RECONNECT,
RestartClass.SESSION_RECONNECT,
RestartClass.CONFIGURATION_RELOAD,
}:
scoped_sessions: list[SessionImpact] = []
scoped_leases: list[LeaseImpact] = []
elif resolved_class == RestartClass.WORKER_RESTART:
selected_session = (target_session_id or "").strip()
target_complete = bool(selected_session)
scoped_sessions = [
s for s in session_impacts if s.session_id == selected_session
]
disruptive_leases = [l for l in lease_impacts if l.disruptive]
critical_sections = [l for l in lease_impacts if l.is_critical_section]
mutations = [l for l in lease_impacts if l.is_mutation]
scoped_leases = [
l for l in lease_impacts if l.session_id == selected_session
]
elif resolved_class == RestartClass.ROLE_RUNTIME_RESTART:
selected_role = (target_role or "").strip().lower()
target_complete = bool(selected_role)
scoped_sessions = [
s for s in session_impacts if str(s.role or "").lower() == selected_role
]
scoped_leases = [
l for l in lease_impacts if str(l.role or "").lower() == selected_role
]
elif resolved_class == RestartClass.CONNECTOR_RESTART:
selected_connector = (target_connector or "").strip()
target_complete = bool(selected_connector)
scoped_sessions = [
s for s in session_impacts if s.connector == selected_connector
]
scoped_leases = [
l for l in lease_impacts if l.connector == selected_connector
]
else:
scoped_sessions = list(session_impacts)
scoped_leases = list(lease_impacts)
if policy_enforced and not target_complete:
authorization_reasons.append(
f"target required for {resolved_class.value if resolved_class else 'unknown class'}"
)
other_live_sessions = [
s for s in scoped_sessions if s.live and not s.is_requester
]
disruptive_leases = [l for l in scoped_leases if l.disruptive]
critical_sections = [l for l in scoped_leases if l.is_critical_section]
mutations = [l for l in scoped_leases if l.is_mutation]
terminal_lock_in_scope = (
terminal_lock
if resolved_class
not in {
RestartClass.CLIENT_RECONNECT,
RestartClass.SESSION_RECONNECT,
}
else None
)
affected_issues = sorted(
{
@@ -348,9 +672,24 @@ def evaluate_restart_impact(
}
)
disruptive = bool(disruptive_leases or other_live_sessions or terminal_lock)
disruptive = bool(
disruptive_leases or other_live_sessions or terminal_lock_in_scope
)
if not inventory_complete:
authorization_ok = bool(
not unknown_class
and permission_authorized
and role_authorized
and approval_satisfied
and target_complete
)
if policy_enforced and not authorization_ok:
verdict = VERDICT_UNSAFE
allow_restart = False
reasons.append("restart class authorization denied (fail closed)")
reasons.extend(authorization_reasons)
elif not inventory_complete:
verdict = VERDICT_UNSAFE
allow_restart = False
reasons.append(
@@ -381,7 +720,7 @@ def evaluate_restart_impact(
f"{len(critical_sections)} critical section(s) in flight "
"(active lease with a live owner)"
)
if terminal_lock:
if terminal_lock_in_scope:
reasons.append("active terminal (merge) lock present")
override_would_allow = bool(inventory_complete and disruptive)
@@ -411,6 +750,12 @@ def evaluate_restart_impact(
audit_record = {
"event": "restart_impact_evaluated",
"coordinator_version": COORDINATOR_VERSION,
"restart_class": (
resolved_class.value if resolved_class else str(restart_class or "")
),
"required_permission": (
policy.required_permission if policy is not None else None
),
"evaluated_at": moment.isoformat(),
"dry_run": dry_run,
"operator_override": bool(operator_override),
@@ -424,6 +769,15 @@ def evaluate_restart_impact(
return RestartImpactReport(
coordinator_version=COORDINATOR_VERSION,
restart_class=(
resolved_class.value if resolved_class else str(restart_class or "")
),
restart_policy=policy.as_dict() if policy is not None else {},
policy_enforced=policy_enforced,
permission_authorized=permission_authorized,
role_authorized=role_authorized,
approval_satisfied=approval_satisfied,
authorization_reasons=authorization_reasons,
evaluated_at=moment.isoformat(),
dry_run=dry_run,
restart_performed=False,
@@ -440,9 +794,11 @@ def evaluate_restart_impact(
affected_issues=affected_issues,
affected_prs=affected_prs,
mutations=mutations,
terminal_lock=dict(terminal_lock)
if isinstance(terminal_lock, Mapping)
else terminal_lock,
terminal_lock=(
dict(terminal_lock_in_scope)
if isinstance(terminal_lock_in_scope, Mapping)
else terminal_lock_in_scope
),
ack_state=ack_state,
prior_recovery_attempts=prior_recovery_attempts,
counts=counts,
+232
View File
@@ -0,0 +1,232 @@
"""Permission, drain, routing, and audit matrix for restart classes (#663)."""
from __future__ import annotations
import os
from datetime import datetime, timezone
import restart_coordinator as rc
NOW = datetime(2026, 7, 24, 20, 0, tzinfo=timezone.utc)
def _inventory() -> dict:
return {
"inventory_complete": True,
"sessions": [
{
"session_id": "requester",
"role": "author",
"profile": "prgs-author",
"pid": os.getpid(),
"status": "active",
"last_heartbeat_at": NOW.isoformat(),
},
{
"session_id": "reviewer",
"role": "reviewer",
"profile": "prgs-reviewer",
"pid": os.getpid(),
"status": "active",
"last_heartbeat_at": NOW.isoformat(),
},
],
"leases": [
{
"lease_id": "review-lease",
"session_id": "reviewer",
"role": "reviewer",
"phase": "reviewing",
"work_kind": "pr",
"work_number": 900,
"worktree_path": "/tmp/review-900",
"freshness": {"freshness": "active"},
}
],
}
def _evaluate(
restart_class: rc.RestartClass,
*,
role: str = "controller",
permissions: tuple[str, ...] | None = None,
approved: bool = True,
operator: bool = True,
**targets,
):
return rc.evaluate_restart_impact(
_inventory(),
now=NOW,
requesting_session_id="requester",
restart_class=restart_class,
requester_role=role,
requester_permissions=(
permissions if permissions is not None
else rc.permissions_for_role(role)
),
controller_approved=approved,
operator_authorized=operator,
**targets,
)
def test_policy_table_covers_exactly_all_nine_classes():
assert set(rc.RESTART_CLASS_POLICIES) == set(rc.RestartClass)
assert len(rc.RESTART_CLASS_POLICIES) == 9
for restart_class, policy in rc.RESTART_CLASS_POLICIES.items():
assert policy.restart_class is restart_class
assert policy.required_permission
assert policy.expected_blast_radius in {
rc.BLAST_NONE, rc.BLAST_LOW, rc.BLAST_MEDIUM, rc.BLAST_HIGH
}
assert policy.drain_requirement
assert policy.approval_requirement
assert policy.audit_requirement
assert policy.recovery_behavior
def test_permission_matrix_allows_each_class_with_exact_permission():
targets = {
rc.RestartClass.WORKER_RESTART: {"target_session_id": "reviewer"},
rc.RestartClass.ROLE_RUNTIME_RESTART: {"target_role": "reviewer"},
rc.RestartClass.CONNECTOR_RESTART: {"target_connector": "github"},
}
for restart_class, policy in rc.RESTART_CLASS_POLICIES.items():
report = _evaluate(
restart_class,
permissions=(policy.required_permission,),
**targets.get(restart_class, {}),
)
assert report.permission_authorized, restart_class
assert report.role_authorized, restart_class
assert report.approval_satisfied, restart_class
assert report.audit_record["restart_class"] == restart_class.value
assert (
report.audit_record["required_permission"]
== policy.required_permission
)
def test_missing_or_nearby_permission_denies():
report = _evaluate(
rc.RestartClass.ROLE_RUNTIME_RESTART,
permissions=("mcp.restart.worker.request",),
target_role="reviewer",
)
assert report.verdict == rc.VERDICT_UNSAFE
assert not report.allow_restart
assert not report.permission_authorized
assert any("missing required permission" in r for r in report.reasons)
def test_unknown_restart_class_denies_fail_closed():
report = rc.evaluate_restart_impact(
_inventory(),
now=NOW,
restart_class="surprise_reboot",
requester_role="admin",
requester_permissions=("mcp.restart.host.request",),
controller_approved=True,
operator_authorized=True,
)
assert report.verdict == rc.VERDICT_UNSAFE
assert not report.allow_restart
assert report.restart_policy == {}
assert any("unknown restart class" in r for r in report.reasons)
def test_worker_roles_cannot_request_full_or_host_restart():
for role in rc.WORKER_ROLES:
granted = rc.permissions_for_role(role)
assert "mcp.restart.full.request" not in granted
assert "mcp.restart.host.request" not in granted
report = _evaluate(
rc.RestartClass.FULL_MCP_RESTART,
role=role,
permissions=granted,
)
assert not report.role_authorized
assert not report.allow_restart
def test_controller_approval_is_independent_of_permission():
report = _evaluate(
rc.RestartClass.WORKER_RESTART,
approved=False,
target_session_id="reviewer",
)
assert report.permission_authorized
assert not report.approval_satisfied
assert not report.allow_restart
def test_narrow_classes_do_not_inherit_full_drain_or_peer_lease_block():
for restart_class in (
rc.RestartClass.CLIENT_RECONNECT,
rc.RestartClass.SESSION_RECONNECT,
rc.RestartClass.CONFIGURATION_RELOAD,
):
report = _evaluate(restart_class)
assert not report.restart_policy["full_drain_required"]
assert report.counts["leases_disruptive"] == 0
assert report.counts["sessions_live_other"] == 0
assert report.counts["critical_sections"] == 0
assert report.counts["mutations"] == 0
assert report.allow_restart, (restart_class, report.reasons)
def test_client_reconnect_does_not_wait_for_unrelated_terminal_lock():
inventory = _inventory()
inventory["terminal_lock"] = {"terminal_pr": 901}
report = rc.evaluate_restart_impact(
inventory,
now=NOW,
requesting_session_id="requester",
restart_class=rc.RestartClass.CLIENT_RECONNECT,
requester_role="author",
requester_permissions=rc.permissions_for_role("author"),
)
assert report.allow_restart
assert report.terminal_lock is None
def test_scoped_restart_only_counts_named_target():
report = _evaluate(
rc.RestartClass.ROLE_RUNTIME_RESTART,
target_role="author",
)
assert report.counts["leases_disruptive"] == 0
assert report.affected_prs == []
assert report.allow_restart
reviewer = _evaluate(
rc.RestartClass.ROLE_RUNTIME_RESTART,
target_role="reviewer",
)
assert reviewer.counts["leases_disruptive"] == 1
assert reviewer.affected_prs == [900]
assert not reviewer.allow_restart
def test_missing_scoped_target_denies_instead_of_widening():
for restart_class in (
rc.RestartClass.WORKER_RESTART,
rc.RestartClass.ROLE_RUNTIME_RESTART,
rc.RestartClass.CONNECTOR_RESTART,
):
report = _evaluate(restart_class)
assert not report.allow_restart
assert any("target required" in r for r in report.reasons)
def test_only_full_and_host_classes_require_full_drain():
requiring_full = {
restart_class
for restart_class, policy in rc.RESTART_CLASS_POLICIES.items()
if policy.full_drain_required
}
assert requiring_full == {
rc.RestartClass.FULL_MCP_RESTART,
rc.RestartClass.HOST_RESTART,
}
+6
View File
@@ -444,6 +444,12 @@ class TestAuditEmission(unittest.TestCase):
)
self.assertEqual(record["target"]["namespace"], NAMESPACE)
self.assertEqual(record["target"]["mode"], "restart")
self.assertEqual(
record["target"]["restart_class"], "role_runtime_restart"
)
self.assertEqual(
record["metadata"]["restart_class"], "role_runtime_restart"
)
self.assertEqual(record["result"], console_audit.RESULT_ALLOWED)
self.assertEqual(record["actor"]["subject"], "[email protected]")
self.assertFalse(record["metadata"]["process_kill_executed"])
-739
View File
@@ -1,739 +0,0 @@
"""Tests for the Runtime and session view (Phase 1, #641).
Covers clean and stale session rendering, contamination marker surfacing,
worktree binding display, sanctioned recovery links (no pkill), nav/live
status, and the JSON API export.
Also pins the two invariants a reviewer found violated at head a81db754:
degraded ownership sections must render as *unknown* rather than as an
affirmative "none"/"unbound", and contamination payload text must be redacted
at the display boundary rather than trusted from the write-time denylist.
"""
from __future__ import annotations
import os
import sys
import unittest
from pathlib import Path
from unittest import mock
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
from tests.webui_testclient import TestClient
from webui.app import create_app
from webui.inventory import (
AUTHORITY_CONTROL_PLANE_DB,
AUTHORITY_FILESYSTEM,
InventorySection,
InventorySnapshot,
STATUS_DEGRADED,
STATUS_OK,
STATUS_UNAVAILABLE,
)
from webui.nav import NAV_GROUPS, STUB_PAGES, iter_nav_items
from webui.runtime_health import FileHash, RuntimeSnapshot
from webui.session_loader import (
ContaminationMarker,
SessionRow,
SessionViewSnapshot,
_build_session_rows,
_inspect_contamination,
load_session_view_snapshot,
snapshot_to_dict,
)
from webui.session_views import render_sessions_page
def _runtime(
*,
stale: str | None = None,
profile: str = "prgs-author",
role: str = "author",
) -> RuntimeSnapshot:
return RuntimeSnapshot(
project_id="gitea-tools",
repo_root="/tmp/repo",
remote="prgs",
host="gitea.prgs.cc",
profile_name=profile,
role_kind=role,
config_model="v2-contexts",
profile_mode="dynamic-profile",
profile_source="config file profile",
authenticated_username="jcwalker3",
identity_error=None,
repo_sha="a" * 40,
remote_master_sha="a" * 40,
commits_behind_master=0,
stale_runtime_warning=stale,
shell_health={"shell_use_allowed": True, "consecutive_spawn_failures": 0},
workflow_hashes=(
FileHash(label="SKILL.md", path="skills/llm-project-workflow/SKILL.md", sha256="abc"),
),
schema_hashes=(),
restart_guidance="docs/mcp-namespace-eof-recovery.md",
fetch_error=None,
)
def _inventory(
*,
sessions: tuple[dict, ...] = (),
leases: tuple[dict, ...] = (),
locks: tuple[dict, ...] = (),
worktrees: tuple[dict, ...] = (),
namespaces: tuple[dict, ...] = (),
statuses: dict[str, str] | None = None,
) -> InventorySnapshot:
"""Build a snapshot; ``statuses`` degrades named sections (default all ok)."""
status_of = statuses or {}
def _status(name: str) -> str:
return status_of.get(name, STATUS_OK)
sections = (
InventorySection(
name="sessions",
authority=AUTHORITY_CONTROL_PLANE_DB,
status=_status("sessions"),
items=sessions,
),
InventorySection(
name="leases",
authority=AUTHORITY_CONTROL_PLANE_DB,
status=_status("leases"),
items=leases,
),
InventorySection(
name="locks",
authority=AUTHORITY_FILESYSTEM,
status=_status("locks"),
items=locks,
),
InventorySection(
name="worktrees",
authority=AUTHORITY_FILESYSTEM,
status=_status("worktrees"),
items=worktrees,
),
InventorySection(
name="namespaces",
authority=AUTHORITY_FILESYSTEM,
status=_status("namespaces"),
items=namespaces
or (
{
"profile_name": "prgs-author",
"role": "author",
"mcp_namespace": "gitea-author",
"capability_summary": {
"can_author": True,
"can_review": False,
"can_merge": False,
},
"active": True,
},
),
reason="only the profile serving this web process is observable",
),
)
index = {section.name: section for section in sections}
return InventorySnapshot(
generated_at="2026-07-25T00:00:00+00:00",
sections=sections,
collisions=(),
correlations=(),
scan_ms=1.0,
_section_index=index,
)
def _clean_session() -> dict:
return {
"session_id": "prgs-author-111-clean",
"role": "author",
"profile": "prgs-author",
"namespace": "gitea-author",
"pid": 1111,
"pid_alive": True,
"status": "active",
"started_at": "2026-07-25T00:00:00Z",
"last_heartbeat_at": "2026-07-25T01:00:00Z",
}
def _stale_session() -> dict:
return {
"session_id": "prgs-author-222-stale",
"role": "author",
"profile": "prgs-author",
"namespace": "gitea-author",
"pid": 2222,
"pid_alive": False,
"status": "active",
"started_at": "2026-07-24T00:00:00Z",
"last_heartbeat_at": "2026-07-24T01:00:00Z",
}
class TestBuildSessionRows(unittest.TestCase):
def test_clean_session_has_no_stale_or_contamination_flags(self):
inventory = _inventory(
sessions=(_clean_session(),),
leases=(
{
"lease_id": "lease-clean",
"session_id": "prgs-author-111-clean",
"status": "active",
"expired": False,
"work_kind": "issue",
"work_number": 641,
},
),
locks=(
{
"issue_number": 641,
"branch_name": "feat/issue-641-runtime-session-view",
"worktree_path": "~/Development/Gitea-Tools/branches/feat-issue-641",
"live": True,
},
),
)
rows = _build_session_rows(inventory, contamination=())
self.assertEqual(len(rows), 1)
row = rows[0]
self.assertEqual(row.session_id, "prgs-author-111-clean")
self.assertEqual(row.role, "author")
self.assertEqual(row.namespace, "gitea-author")
self.assertEqual(row.pid_alive, True)
self.assertEqual(row.lease_ids, ("lease-clean",))
self.assertEqual(row.work_refs, ("issue#641",))
self.assertTrue(row.worktree_paths)
self.assertEqual(row.stale_flags, ())
self.assertEqual(row.contamination_flags, ())
def test_stale_session_flags_dead_pid(self):
inventory = _inventory(sessions=(_stale_session(),))
rows = _build_session_rows(inventory, contamination=())
self.assertEqual(rows[0].stale_flags, ("pid-dead",))
def test_contamination_marker_binds_to_session(self):
inventory = _inventory(sessions=(_clean_session(),))
marker = ContaminationMarker(
kind="runtime_recovery_contamination",
on_disk=True,
has_payload=True,
summary="manual daemon kill",
reason_class="manual_daemon_kill",
session_id="prgs-author-111-clean",
role="author",
command_summary="pkill -f mcp_server.py",
cleared=False,
)
rows = _build_session_rows(inventory, contamination=(marker,))
self.assertIn("runtime_recovery_contamination", rows[0].contamination_flags)
def test_process_wide_contamination_surfaces_on_all_sessions(self):
inventory = _inventory(sessions=(_clean_session(), _stale_session()))
marker = ContaminationMarker(
kind="stable_branch_contamination",
on_disk=True,
has_payload=True,
summary="direct master push attempt",
reason_class="stable_branch_push",
session_id=None,
cleared=False,
)
rows = _build_session_rows(inventory, contamination=(marker,))
self.assertEqual(len(rows), 2)
for row in rows:
self.assertTrue(
any("stable_branch_contamination" in f for f in row.contamination_flags)
)
class TestRenderSessionsPage(unittest.TestCase):
def _snapshot(
self,
*,
sessions: tuple[dict, ...],
contamination: tuple[ContaminationMarker, ...] = (),
stale_runtime: str | None = None,
) -> SessionViewSnapshot:
inventory = _inventory(
sessions=sessions,
leases=(
{
"lease_id": "lease-1",
"session_id": sessions[0]["session_id"] if sessions else "",
"status": "active",
"expired": False,
"work_kind": "issue",
"work_number": 641,
},
)
if sessions
else (),
locks=(
{
"issue_number": 641,
"worktree_path": "branches/feat-issue-641",
},
)
if sessions
else (),
worktrees=(
{
"rel_path": "branches/feat-issue-641",
"branch": "feat/issue-641-runtime-session-view",
"classification": "active_issue_work",
"registered_worktree": True,
"dirty": False,
},
),
)
rows = _build_session_rows(inventory, contamination)
return SessionViewSnapshot(
runtime=_runtime(stale=stale_runtime),
inventory=inventory,
sessions=rows,
contamination_markers=contamination,
)
def test_clean_session_render(self):
html = render_sessions_page(self._snapshot(sessions=(_clean_session(),)))
self.assertIn("Runtime and sessions", html)
self.assertIn("prgs-author-111-clean", html)
self.assertIn("gitea-author", html)
self.assertIn("branches/feat-issue-641", html)
self.assertIn("Sanctioned recovery", html)
self.assertIn("docs/mcp-namespace-eof-recovery.md", html)
# Recovery section must name reconnect and forbid manual kill.
recovery_idx = html.lower().find("sanctioned recovery")
self.assertGreaterEqual(recovery_idx, 0)
recovery = html[recovery_idx:].lower()
self.assertIn("reconnect", recovery)
self.assertIn("contamination", recovery)
self.assertIn("not recovery", recovery)
self.assertNotIn("run pkill", recovery)
self.assertNotIn("killall", recovery)
def test_stale_session_render(self):
html = render_sessions_page(self._snapshot(sessions=(_stale_session(),)))
self.assertIn("prgs-author-222-stale", html)
self.assertIn("pid-dead", html)
self.assertIn("badge-stale", html)
def test_contamination_render_is_not_silent(self):
marker = ContaminationMarker(
kind="runtime_recovery_contamination",
on_disk=True,
has_payload=True,
summary="manual kill",
reason_class="manual_daemon_kill",
session_id="prgs-author-111-clean",
command_summary="pkill -f mcp_server.py",
cleared=False,
)
html = render_sessions_page(
self._snapshot(sessions=(_clean_session(),), contamination=(marker,))
)
self.assertIn("Contamination markers", html)
self.assertIn("runtime_recovery_contamination", html)
self.assertIn("ACTIVE", html)
self.assertIn("badge-blocked", html)
def test_stale_runtime_banner(self):
html = render_sessions_page(
self._snapshot(
sessions=(_clean_session(),),
stale_runtime="server behind master by 3 commits",
)
)
self.assertIn("Stale runtime", html)
self.assertIn("server behind master", html)
class TestSessionLoaderComposition(unittest.TestCase):
def test_load_with_injected_sources(self):
inventory = _inventory(sessions=(_clean_session(), _stale_session()))
snap = load_session_view_snapshot(
load_runtime=lambda: _runtime(),
load_inventory=lambda: inventory,
inspect_contamination=lambda **_k: {
"on_disk": False,
"has_payload": False,
"summary": "absent",
},
load_contamination_payload=lambda **_k: None,
)
self.assertEqual(len(snap.sessions), 2)
self.assertEqual(snap.stale_session_count, 1)
self.assertEqual(snap.contaminated_session_count, 0)
data = snapshot_to_dict(snap)
self.assertEqual(data["view"], "runtime-sessions")
self.assertEqual(data["issue"], 641)
self.assertTrue(data["read_only"])
self.assertEqual(data["session_counts"]["total"], 2)
self.assertEqual(data["session_counts"]["stale"], 1)
self.assertIn("recovery_docs", data)
class TestSessionsRoutes(unittest.TestCase):
def setUp(self):
self.client = TestClient(create_app())
inventory = _inventory(
sessions=(_clean_session(), _stale_session()),
leases=(
{
"lease_id": "lease-x",
"session_id": "prgs-author-111-clean",
"status": "active",
"expired": False,
"work_kind": "issue",
"work_number": 641,
},
),
locks=(
{
"issue_number": 641,
"worktree_path": "branches/feat-issue-641",
},
),
worktrees=(
{
"rel_path": "branches/feat-issue-641",
"branch": "feat/issue-641-runtime-session-view",
"classification": "active_issue_work",
"registered_worktree": True,
"dirty": False,
},
),
)
rows = _build_session_rows(inventory, contamination=())
self.snapshot = SessionViewSnapshot(
runtime=_runtime(stale="stale for test"),
inventory=inventory,
sessions=rows,
contamination_markers=(),
)
self._patch = mock.patch(
"webui.app.load_session_view_snapshot",
return_value=self.snapshot,
)
self._patch.start()
def tearDown(self):
self._patch.stop()
def test_sessions_page_live(self):
response = self.client.get("/sessions")
self.assertEqual(response.status_code, 200)
self.assertIn("Runtime and sessions", response.text)
self.assertIn("prgs-author-111-clean", response.text)
self.assertIn("prgs-author-222-stale", response.text)
self.assertIn("pid-dead", response.text)
self.assertIn("Sanctioned recovery", response.text)
self.assertNotIn("Phase 1 shell placeholder", response.text)
self.assertNotIn("child issue of #425", response.text.lower())
def test_api_sessions_json(self):
for path in ("/api/sessions", "/api/v1/sessions"):
response = self.client.get(path)
self.assertEqual(response.status_code, 200, path)
data = response.json()
self.assertEqual(data["view"], "runtime-sessions")
self.assertEqual(data["session_counts"]["total"], 2)
self.assertEqual(data["session_counts"]["stale"], 1)
self.assertTrue(data["read_only"])
self.assertEqual(data["mutations"], [])
def test_nav_marks_sessions_live(self):
sessions_items = [
item for item in iter_nav_items() if item.href == "/sessions"
]
self.assertEqual(len(sessions_items), 1)
self.assertEqual(sessions_items[0].status, "live")
self.assertNotIn("/sessions", STUB_PAGES)
# Home page should not mark Sessions as stub.
home = self.client.get("/")
self.assertEqual(home.status_code, 200)
self.assertIn('href="/sessions"', home.text)
# Stub marker only appears next to remaining stub destinations.
self.assertNotIn(
'href="/sessions">Sessions</a> <span class="muted">(stub)</span>',
home.text,
)
class TestDegradedOwnershipAuthority(unittest.TestCase):
"""B1: a section that could not be read must never render as absence."""
def _snapshot(self, inventory: InventorySnapshot) -> SessionViewSnapshot:
return SessionViewSnapshot(
runtime=_runtime(),
inventory=inventory,
sessions=_build_session_rows(inventory, contamination=()),
contamination_markers=(),
)
def test_unavailable_leases_mark_row_authority_unproven(self):
inventory = _inventory(
sessions=(_clean_session(),),
statuses={"leases": STATUS_UNAVAILABLE},
)
row = _build_session_rows(inventory, contamination=())[0]
self.assertEqual(row.lease_ids, ())
self.assertEqual(row.lease_authority, STATUS_UNAVAILABLE)
# Worktree binding is correlated through lease work numbers, so it
# inherits the unreadable lease section.
self.assertEqual(row.worktree_authority, STATUS_UNAVAILABLE)
self.assertFalse(row.ownership_authority_complete)
def test_readable_locks_are_not_reported_unbound_when_leases_degrade(self):
# The narrow variant: locks hold a real worktree_path and read cleanly,
# but the lease section that supplies the correlating work number does
# not. The row must say unknown, not "unbound".
inventory = _inventory(
sessions=(_clean_session(),),
locks=(
{
"issue_number": 641,
"worktree_path": "branches/feat-issue-641",
},
),
statuses={"leases": STATUS_DEGRADED},
)
row = _build_session_rows(inventory, contamination=())[0]
self.assertEqual(row.worktree_paths, ())
self.assertEqual(row.worktree_authority, STATUS_DEGRADED)
self.assertFalse(row.ownership_authority_complete)
def test_degraded_render_says_unknown_not_none_or_unbound(self):
inventory = _inventory(
sessions=(_clean_session(),),
statuses={"leases": STATUS_UNAVAILABLE, "locks": STATUS_UNAVAILABLE},
)
html = render_sessions_page(self._snapshot(inventory))
self.assertIn("unknown (inventory unavailable)", html)
self.assertIn("authority unproven", html)
self.assertIn("Ownership authority incomplete", html)
# The affirmative-absence strings must be gone from the row entirely.
self.assertNotIn(">none<", html)
self.assertNotIn(">unbound<", html)
def test_clean_inventory_still_renders_affirmative_absence(self):
# Guards against over-correcting B1 into "everything is unknown".
inventory = _inventory(sessions=(_clean_session(),))
html = render_sessions_page(self._snapshot(inventory))
self.assertIn(">none<", html)
self.assertIn(">unbound<", html)
# The column legend mentions "unknown (inventory …)" as static copy, so
# assert on the per-row marker and the concrete statuses instead.
self.assertNotIn("authority unproven", html)
self.assertNotIn("unknown (inventory unavailable)", html)
self.assertNotIn("unknown (inventory degraded)", html)
self.assertNotIn("Ownership authority incomplete", html)
def test_json_export_carries_snapshot_and_per_row_authority(self):
inventory = _inventory(
sessions=(_clean_session(),),
statuses={"locks": STATUS_UNAVAILABLE},
)
data = snapshot_to_dict(self._snapshot(inventory))
self.assertFalse(data["ownership_authority_complete"])
self.assertEqual(
data["ownership_section_status"]["locks"], STATUS_UNAVAILABLE
)
self.assertEqual(data["ownership_section_status"]["leases"], STATUS_OK)
self.assertIn("unknown, not unowned", data["ownership_note"])
row = data["sessions"][0]
self.assertTrue(row["lease_authority_complete"])
self.assertFalse(row["worktree_authority_complete"])
self.assertEqual(row["worktree_authority"], STATUS_UNAVAILABLE)
self.assertFalse(row["ownership_authority_complete"])
self.assertIn("unknown, not unowned", row["ownership_note"])
def test_json_export_is_affirmative_when_every_source_reads(self):
inventory = _inventory(sessions=(_clean_session(),))
data = snapshot_to_dict(self._snapshot(inventory))
self.assertTrue(data["ownership_authority_complete"])
self.assertTrue(data["sessions"][0]["ownership_authority_complete"])
def test_missing_session_list_is_not_reported_as_no_sessions(self):
inventory = _inventory(statuses={"sessions": STATUS_UNAVAILABLE})
html = render_sessions_page(self._snapshot(inventory))
self.assertIn("could not be read", html)
self.assertIn("not evidence that no sessions exist", html)
def test_expired_lease_flags_row_as_stale(self):
inventory = _inventory(
sessions=(_clean_session(),),
leases=(
{
"lease_id": "lease-expired-1",
"session_id": "prgs-author-111-clean",
"status": "active",
"expired": True,
"work_kind": "issue",
"work_number": 641,
},
),
)
row = _build_session_rows(inventory, contamination=())[0]
self.assertIn("lease-expired", row.stale_flags)
self.assertIn("active-lease-past-expiry", row.stale_flags)
class TestContaminationRedaction(unittest.TestCase):
"""B2: marker payload text is redacted at the display boundary."""
def _marker(self, payload: dict) -> ContaminationMarker:
return _inspect_contamination(
"runtime_recovery_contamination",
remote="prgs",
inspect=lambda **_k: {
"on_disk": True,
"has_payload": True,
"summary": "",
},
load=lambda **_k: payload,
)
def test_home_paths_are_collapsed(self):
home = os.path.expanduser("~")
marker = self._marker(
{"command_summary": f"pkill -f {home}/Development/Gitea-Tools/x.py"}
)
self.assertNotIn(home, marker.command_summary)
self.assertIn("~/Development/Gitea-Tools/x.py", marker.command_summary)
def test_secrets_missed_by_the_write_time_denylist_are_redacted(self):
# Each of these was verified in review to survive
# stable_branch_push_guard.redact_command untouched.
cases = (
("curl -H 'X-Api-Key: SUPERSECRET123' https://example.invalid", "SUPERSECRET123"),
("cmd --password hunter2 origin master", "hunter2"),
("PRIVATE_KEY=abc123 python deploy.py", "abc123"),
("fetch https://user:[email protected]/x.git", "user:pw"),
)
for raw, secret in cases:
with self.subTest(raw=raw):
marker = self._marker({"command_summary": raw})
self.assertNotIn(secret, marker.command_summary)
self.assertIn("[redacted]", marker.command_summary)
def test_command_summary_is_redacted_not_removed(self):
# It is legitimate #630 evidence: the operator must still see which
# daemon was killed.
marker = self._marker(
{
"command_summary": "pkill -f gitea_mcp_server.py",
"reason_class": "manual_daemon_kill",
"session_id": "prgs-author-111-clean",
"role": "author",
}
)
self.assertIn("pkill -f gitea_mcp_server.py", marker.command_summary)
self.assertEqual(marker.reason_class, "manual_daemon_kill")
self.assertEqual(marker.session_id, "prgs-author-111-clean")
self.assertEqual(marker.role, "author")
def test_rendered_page_exposes_no_home_path_from_a_marker(self):
home = os.path.expanduser("~")
marker = self._marker(
{
"command_summary": f"pkill -f {home}/Development/Gitea-Tools/x.py",
"reason_class": "manual_daemon_kill",
}
)
inventory = _inventory(sessions=(_clean_session(),))
html = render_sessions_page(
SessionViewSnapshot(
runtime=_runtime(),
inventory=inventory,
sessions=_build_session_rows(inventory, (marker,)),
contamination_markers=(marker,),
)
)
self.assertIn("Contamination markers", html)
self.assertNotIn(home, html)
class TestSessionsPageEscaping(unittest.TestCase):
"""Hostile values from every rendered source stay inert (N2)."""
HOSTILE = '<script>alert("xss")</script>'
def test_hostile_session_and_marker_values_are_escaped(self):
session = dict(_clean_session())
session["session_id"] = f"sid-{self.HOSTILE}"
session["role"] = self.HOSTILE
session["profile"] = self.HOSTILE
session["namespace"] = self.HOSTILE
session["status"] = self.HOSTILE
inventory = _inventory(
sessions=(session,),
leases=(
{
"lease_id": self.HOSTILE,
"session_id": session["session_id"],
"status": "active",
"expired": False,
"work_kind": self.HOSTILE,
"work_number": 641,
},
),
locks=(
{
"issue_number": 641,
"worktree_path": self.HOSTILE,
},
),
)
marker = ContaminationMarker(
kind="runtime_recovery_contamination",
on_disk=True,
has_payload=True,
summary=self.HOSTILE,
reason_class=self.HOSTILE,
session_id=session["session_id"],
role=self.HOSTILE,
command_summary=self.HOSTILE,
cleared=False,
)
html = render_sessions_page(
SessionViewSnapshot(
runtime=_runtime(),
inventory=inventory,
sessions=_build_session_rows(inventory, (marker,)),
contamination_markers=(marker,),
)
)
self.assertNotIn("<script>", html)
self.assertNotIn('alert("xss")', html)
self.assertIn("&lt;script&gt;", html)
def test_hostile_values_in_a_degraded_render_are_escaped(self):
inventory = _inventory(
sessions=(dict(_clean_session(), session_id=f"sid-{self.HOSTILE}"),),
statuses={"leases": STATUS_UNAVAILABLE, "locks": STATUS_DEGRADED},
)
html = render_sessions_page(
SessionViewSnapshot(
runtime=_runtime(),
inventory=inventory,
sessions=_build_session_rows(inventory, contamination=()),
contamination_markers=(),
)
)
self.assertNotIn("<script>", html)
self.assertIn("&lt;script&gt;", html)
self.assertIn("unknown (inventory", html)
if __name__ == "__main__":
unittest.main()
-20
View File
@@ -48,11 +48,6 @@ from webui.worktree_scanner import load_hygiene_snapshot, snapshot_to_dict as wo
from webui.worktree_views import render_worktrees_page
from webui.runtime_health import load_runtime_snapshot, snapshot_to_dict as runtime_snapshot_to_dict
from webui.runtime_views import render_runtime_page
from webui.session_loader import (
load_session_view_snapshot,
snapshot_to_dict as session_view_snapshot_to_dict,
)
from webui.session_views import render_sessions_page
from webui.inventory import (
SECTION_NAMES as _INVENTORY_SECTIONS,
load_inventory_snapshot,
@@ -92,7 +87,6 @@ _LEGACY_PAGES = (
("/projects", "Projects", "registry and onboarding (#427)"),
("/prompts", "Prompts", "canonical workflow prompt library (#428)"),
("/runtime", "Runtime", "MCP health and stale-runtime detection (#430)"),
("/sessions", "Sessions", "runtime and session view (#641)"),
("/audit", "Audit", "final-report paste and validator preview (#431)"),
("/worktrees", "Worktrees", "branch hygiene dashboard (#432)"),
("/leases", "Leases", "collision and lease visibility (#433)"),
@@ -331,17 +325,6 @@ async def api_runtime(_request: Request) -> JSONResponse:
return JSONResponse(runtime_snapshot_to_dict(load_runtime_snapshot()))
async def sessions(_request: Request) -> HTMLResponse:
"""Runtime and session view (#641) — read-only composition of health + inventory."""
snapshot = load_session_view_snapshot()
return HTMLResponse(render_sessions_page(snapshot))
async def api_sessions(_request: Request) -> JSONResponse:
"""JSON export for the runtime/session view (#641)."""
return JSONResponse(session_view_snapshot_to_dict(load_session_view_snapshot()))
async def _parse_audit_form(request: Request) -> tuple[str, str | None]:
if request.method == "GET":
return "", None
@@ -781,9 +764,6 @@ def create_app(*, bind_host: str | None = None) -> Starlette:
Route("/api/prompts", api_prompts, methods=["GET"]),
Route("/runtime", runtime, methods=["GET"]),
Route("/api/runtime", api_runtime, methods=["GET"]),
Route("/sessions", sessions, methods=["GET"]),
Route("/api/sessions", api_sessions, methods=["GET"]),
Route("/api/v1/sessions", api_sessions, methods=["GET"]),
Route("/api/v1/timeline", api_v1_timeline, methods=["GET"]),
Route("/analytics", analytics, methods=["GET"]),
Route("/api/analytics", api_v1_analytics, methods=["GET"]),
-65
View File
@@ -76,35 +76,6 @@ _CREDENTIAL_KEY_RE = re.compile(
)
_REDACTED = "[redacted]"
#: Credential-shaped *name* as it appears inside a free-form command line. This
#: is deliberately broader than :data:`_CREDENTIAL_KEY_RE` — it also matches a
#: bare ``key`` component, so ``PRIVATE_KEY=`` is caught. Over-redacting a
#: displayed string is safe; under-redacting one is not.
_TEXT_CREDENTIAL_NAME = (
r"[A-Za-z0-9_.\-]*"
r"(?:token|secret|password|passwd|key|authorization|bearer|credential)"
r"[A-Za-z0-9_.\-]*"
)
#: A value following such a name: single-quoted, double-quoted, or bare. The
#: bare form stops at a quote so an enclosing quote survives the redaction.
_TEXT_CREDENTIAL_VALUE = r"'[^']*'|\"[^\"]*\"|[^\s'\"]+"
_TEXT_CREDENTIAL_FLAG_RE = re.compile(
rf"(?P<key>(?<![\w\-])--?{_TEXT_CREDENTIAL_NAME})"
rf"(?P<sep>[=\s]+)"
rf"(?P<value>{_TEXT_CREDENTIAL_VALUE})",
re.IGNORECASE,
)
_TEXT_CREDENTIAL_ASSIGN_RE = re.compile(
rf"(?P<key>(?<![\w\-]){_TEXT_CREDENTIAL_NAME})"
rf"(?P<sep>\s*[:=]\s*)"
rf"(?P<value>{_TEXT_CREDENTIAL_VALUE})",
re.IGNORECASE,
)
_TEXT_URL_USERINFO_RE = re.compile(
r"(?P<scheme>\b[A-Za-z][A-Za-z0-9+.\-]*://)[^\s/@]+@"
)
@dataclass(frozen=True)
class InventorySection:
@@ -254,42 +225,6 @@ def scrub(value: Any, *, key: str | None = None) -> Any:
return repr(value)
def collapse_home(text: str) -> str:
"""Collapse every ``$HOME`` occurrence *inside* a string, not just a prefix."""
home = os.path.expanduser("~")
if not home or home == "/":
return text
return text.replace(home, "~")
def scrub_text(value: Any) -> Any:
"""Redact a free-form text blob such as a recorded command line.
:func:`scrub` keys off structured field *names* and whole-value prefixes,
which is right for inventory records but blind to a secret embedded in the
middle of a sentence. This collapses ``$HOME`` and redacts credential-shaped
tokens and URL userinfo *anywhere* in the string, so operator-supplied text
rendered verbatim — contamination ``command_summary`` (#630) above all — is
held to the same standard as every other field on the page.
Returns ``None`` unchanged so callers can keep "absent" distinct from "".
"""
if value is None:
return None
text = value if isinstance(value, str) else str(value)
text = collapse_home(text)
text = _TEXT_URL_USERINFO_RE.sub(
lambda m: f"{m.group('scheme')}{_REDACTED}@", text
)
text = _TEXT_CREDENTIAL_FLAG_RE.sub(
lambda m: f"{m.group('key')}{m.group('sep')}{_REDACTED}", text
)
text = _TEXT_CREDENTIAL_ASSIGN_RE.sub(
lambda m: f"{m.group('key')}{m.group('sep')}{_REDACTED}", text
)
return text
# ── control-plane database (read-only) ───────────────────────────────────────
+6 -1
View File
@@ -48,7 +48,7 @@ NAV_GROUPS: tuple[NavGroup, ...] = (
)),
NavGroup("Runtime/Sessions", (
NavItem("/runtime", "Runtime health"),
NavItem("/sessions", "Sessions"),
NavItem("/sessions", "Sessions", "stub"),
)),
NavGroup("Projects", (
NavItem("/projects", "Projects"),
@@ -76,6 +76,11 @@ NAV_GROUPS: tuple[NavGroup, ...] = (
# issues of epic #631. Each maps a path to (title, description). Routes are
# registered so nav links resolve to a graceful, read-only stub page.
STUB_PAGES: dict[str, tuple[str, str]] = {
"/sessions": (
"Sessions",
"Active session, capability, and role inventory. Backed by the unified "
"inventory API (#636) once it lands.",
),
"/inventory": (
"Inventory",
"Unified sessions, leases, locks, namespaces, and worktree inventory. "
+1 -3
View File
@@ -88,7 +88,5 @@ def render_runtime_page(snapshot: RuntimeSnapshot) -> str:
"<p class='muted'>MVP is read-only — restart MCP servers from your IDE/operator "
"workflow. Related issue: <code>#420</code>. Guidance: "
f"<code>{html.escape(snapshot.restart_guidance)}</code></p>"
"<p class='muted'>This page does not expose tokens or perform MCP restarts. "
"Correlated sessions, worktree bindings, and contamination markers: "
"<a href='/sessions'>/sessions</a> (#641).</p>"
"<p class='muted'>This page does not expose tokens or perform MCP restarts.</p>"
)
+19 -1
View File
@@ -38,6 +38,7 @@ from dataclasses import asdict, dataclass
from typing import Any
import mcp_namespace_health
import restart_coordinator
import runtime_recovery_guard
from webui import console_audit, console_authz
@@ -99,6 +100,14 @@ def _clean(value: Any) -> str:
return str(value or "").strip()
def restart_class_for_mode(mode: str) -> str:
"""Map the existing namespace controls onto the #663 class taxonomy."""
if _clean(mode) == MODE_RELOAD:
return restart_coordinator.RestartClass.CONFIGURATION_RELOAD.value
return restart_coordinator.RestartClass.ROLE_RUNTIME_RESTART.value
# --- Mutation ledger --------------------------------------------------------
@@ -256,6 +265,7 @@ def build_restart_preview(
return {
"action_id": action_id,
"restart_class": restart_class_for_mode(md),
"namespace": ns,
"mode": md,
"scope_valid": scope_error is None,
@@ -309,6 +319,7 @@ def assess_restart_request(
"reason_code": reason_code,
"detail": detail,
"action_id": action_id,
"restart_class": restart_class_for_mode(md),
"namespace": ns,
"mode": md,
"preview": preview,
@@ -393,6 +404,7 @@ def assess_restart_request(
"process."
),
"action_id": action_id,
"restart_class": restart_class_for_mode(md),
"namespace": ns,
"mode": md,
"preview": preview,
@@ -441,7 +453,11 @@ def execute_restart(
else console_audit.RESULT_DENIED
),
principal=principal,
target={"namespace": assessment["namespace"], "mode": assessment["mode"]},
target={
"namespace": assessment["namespace"],
"mode": assessment["mode"],
"restart_class": assessment["restart_class"],
},
reason_code=assessment["reason_code"],
detail=assessment["detail"],
request_id=request_id,
@@ -450,6 +466,7 @@ def execute_restart(
"gates_passed": assessment["gates_passed"],
"process_kill_executed": False,
"post_restart_verification_required": True,
"restart_class": assessment["restart_class"],
},
)
@@ -463,6 +480,7 @@ def execute_restart(
"namespace": assessment["namespace"],
"mode": assessment["mode"],
"action_id": action_id,
"restart_class": assessment["restart_class"],
"process_kill_executed": False,
"host_hook": assessment["preview"]["restart_hook"],
"next_action": (
-517
View File
@@ -1,517 +0,0 @@
"""Compose runtime health + inventory into a sessions/runtime view (#641).
Phase 1 is read-only. It correlates namespaces, sessions, capabilities,
worktree bindings, lease ownership, stale flags, and contamination markers
when they are detectable on disk (#630 / #671). It never restarts, kills, or
takes over a session.
Sources:
* :mod:`webui.runtime_health` — profile, role, stale runtime, shell health.
* :mod:`webui.inventory` — sessions, leases, locks, worktrees, namespaces
and collision signals from the control-plane DB + filesystem.
* :mod:`mcp_session_state` — durable contamination markers (inspect only).
Secrets are never read. Inventory-sourced values arrive already redacted by
:func:`webui.inventory.scrub`; free-form marker text this module loads itself is
put through :func:`webui.inventory.scrub_text`, which collapses ``$HOME`` and
redacts credential-shaped tokens *inside* a string rather than only at its start.
Ownership columns are authority-aware. A lease or lock section that could not be
read renders as ``unknown``, never as ``none`` or ``unbound``: a lease the reader
could not load is not an absent lease (see
:attr:`webui.inventory.InventorySnapshot.ownership_authority_complete`).
"""
from __future__ import annotations
from dataclasses import dataclass
from typing import Any, Callable
import mcp_session_state
from webui.inventory import (
OWNERSHIP_SECTIONS,
STATUS_OK,
InventorySection,
InventorySnapshot,
load_inventory_snapshot,
scrub_text,
snapshot_to_dict as inventory_snapshot_to_dict,
)
from webui.runtime_health import (
RuntimeSnapshot,
load_runtime_snapshot,
snapshot_to_dict as runtime_snapshot_to_dict,
)
# Sanctioned recovery pointers only — never pkill / killall (#630).
SANCTIONED_RECOVERY_DOCS: tuple[dict[str, str], ...] = (
{
"label": "MCP namespace EOF recovery (reconnect only)",
"path": "docs/mcp-namespace-eof-recovery.md",
"note": "IDE/client reconnect or operator-owned restart; never kill daemons.",
},
{
"label": "MCP namespace health",
"path": "docs/mcp-namespace-health.md",
"note": "client_namespace probe proves namespace health.",
},
{
"label": "Restart path inventory",
"path": "docs/mcp-restart-path-inventory.md",
"note": "Catalog of sanctioned reconnect/restart paths.",
},
{
"label": "Local web UI recovery",
"path": "docs/webui-local-dev.md",
"note": "Operator console start and documented recovery sequence.",
},
)
_CONTAMINATION_KINDS: tuple[str, ...] = (
mcp_session_state.KIND_RUNTIME_RECOVERY_CONTAMINATION,
mcp_session_state.KIND_STABLE_BRANCH_CONTAMINATION,
)
#: Status recorded on a row when the backing inventory section is absent
#: entirely — distinct from a section that reported itself degraded.
AUTHORITY_MISSING = "missing"
def _section_status(section: InventorySection | None) -> str:
"""Status of an ownership section, treating an absent section as missing."""
if section is None:
return AUTHORITY_MISSING
return section.status
def _combined_authority(*statuses: str) -> str:
"""Worst status of the sections a derived column depends on.
A column proved from two sections is only trustworthy when *both* read
cleanly, so the first non-``ok`` status wins.
"""
for status in statuses:
if status != STATUS_OK:
return status
return STATUS_OK
@dataclass(frozen=True)
class ContaminationMarker:
"""A detectable durable contamination marker (audit-safe summary)."""
kind: str
on_disk: bool
has_payload: bool
summary: str
reason_class: str | None = None
session_id: str | None = None
role: str | None = None
command_summary: str | None = None
cleared: bool = False
def to_dict(self) -> dict[str, Any]:
return {
"kind": self.kind,
"on_disk": self.on_disk,
"has_payload": self.has_payload,
"summary": self.summary,
"reason_class": self.reason_class,
"session_id": self.session_id,
"role": self.role,
"command_summary": self.command_summary,
"cleared": self.cleared,
"active": self.on_disk and self.has_payload and not self.cleared,
}
@dataclass(frozen=True)
class SessionRow:
"""One correlated session row for the sessions table."""
session_id: str
role: str | None
profile: str | None
namespace: str | None
pid: int | None
pid_alive: bool | None
status: str | None
started_at: str | None
last_heartbeat_at: str | None
lease_ids: tuple[str, ...] = ()
work_refs: tuple[str, ...] = ()
worktree_paths: tuple[str, ...] = ()
stale_flags: tuple[str, ...] = ()
contamination_flags: tuple[str, ...] = ()
#: Status of the section backing ``lease_ids``/``work_refs``. While this is
#: not ``ok`` those tuples mean "could not be read", never "none held".
lease_authority: str = STATUS_OK
#: Worst status across the sections backing ``worktree_paths`` (locks are
#: correlated through lease work numbers, so both must read cleanly).
worktree_authority: str = STATUS_OK
@property
def ownership_authority_complete(self) -> bool:
"""True only when this row's ownership columns are provable."""
return self.lease_authority == STATUS_OK and self.worktree_authority == STATUS_OK
def to_dict(self) -> dict[str, Any]:
return {
"session_id": self.session_id,
"role": self.role,
"profile": self.profile,
"namespace": self.namespace,
"pid": self.pid,
"pid_alive": self.pid_alive,
"status": self.status,
"started_at": self.started_at,
"last_heartbeat_at": self.last_heartbeat_at,
"lease_ids": list(self.lease_ids),
"work_refs": list(self.work_refs),
"worktree_paths": list(self.worktree_paths),
"stale_flags": list(self.stale_flags),
"contamination_flags": list(self.contamination_flags),
"is_stale": bool(self.stale_flags),
"is_contaminated": bool(self.contamination_flags),
"lease_authority": self.lease_authority,
"worktree_authority": self.worktree_authority,
"lease_authority_complete": self.lease_authority == STATUS_OK,
"worktree_authority_complete": self.worktree_authority == STATUS_OK,
"ownership_authority_complete": self.ownership_authority_complete,
"ownership_note": (
"Lease and worktree columns are proved from sections that read "
"cleanly."
if self.ownership_authority_complete
else "An ownership source could not be read; empty lease_ids or "
"worktree_paths on this row mean unknown, not unowned."
),
}
@dataclass(frozen=True)
class SessionViewSnapshot:
"""Composed runtime + session inventory view (#641)."""
runtime: RuntimeSnapshot
inventory: InventorySnapshot
sessions: tuple[SessionRow, ...]
contamination_markers: tuple[ContaminationMarker, ...]
recovery_docs: tuple[dict[str, str], ...] = SANCTIONED_RECOVERY_DOCS
fetch_error: str | None = None
@property
def stale_session_count(self) -> int:
return sum(1 for row in self.sessions if row.stale_flags)
@property
def contaminated_session_count(self) -> int:
return sum(1 for row in self.sessions if row.contamination_flags)
@property
def active_contamination(self) -> tuple[ContaminationMarker, ...]:
return tuple(m for m in self.contamination_markers if m.to_dict()["active"])
@property
def ownership_authority_complete(self) -> bool:
"""False while any ownership section is degraded, absent, or unavailable."""
return self.inventory.ownership_authority_complete
@property
def ownership_section_status(self) -> dict[str, str]:
"""Per-section status for the three ownership-bearing sections."""
return {
name: _section_status(self.inventory.section(name))
for name in OWNERSHIP_SECTIONS
}
def _inspect_contamination(
kind: str,
*,
remote: str | None,
inspect: Callable[..., dict[str, Any]] | None = None,
load: Callable[..., dict[str, Any] | None] | None = None,
) -> ContaminationMarker:
"""Inspect one contamination kind; never raises into the page render path."""
inspect_fn = inspect or mcp_session_state.inspect_state_envelope
load_fn = load or mcp_session_state.load_state
try:
envelope = inspect_fn(kind=kind, remote=remote)
except Exception as exc: # noqa: BLE001 — fail soft for the dashboard
return ContaminationMarker(
kind=kind,
on_disk=False,
has_payload=False,
summary=f"contamination inspect failed: {type(exc).__name__}",
)
reason_class = None
session_id = None
role = None
command_summary = None
cleared = False
summary = scrub_text(str(envelope.get("summary") or ""))
if envelope.get("on_disk") and envelope.get("has_payload"):
try:
payload = load_fn(kind=kind, remote=remote) or {}
except Exception: # noqa: BLE001
payload = {}
if isinstance(payload, dict):
reason_class = payload.get("reason_class")
session_id = payload.get("session_id")
role = payload.get("role")
command_summary = payload.get("command_summary") or payload.get("detail")
cleared = bool(payload.get("cleared_by_reconciler"))
if not summary:
summary = scrub_text(
f"{kind}: {reason_class or 'present'}"
+ (" (cleared)" if cleared else "")
)
# Marker payloads are operator-supplied free text and are the one thing on
# this page that does not arrive through inventory scrubbing. The write-time
# redactor is a narrow denylist, so redact again at the display boundary:
# it leaves $HOME paths, `-H 'X-Api-Key: …'`, `--password …`, and
# `PRIVATE_KEY=…` intact. The field itself stays — it is #630 evidence.
return ContaminationMarker(
kind=kind,
on_disk=bool(envelope.get("on_disk")),
has_payload=bool(envelope.get("has_payload")),
summary=summary or f"{kind}: not present",
reason_class=scrub_text(str(reason_class)) if reason_class else None,
session_id=scrub_text(str(session_id)) if session_id else None,
role=scrub_text(str(role)) if role else None,
command_summary=scrub_text(str(command_summary)) if command_summary else None,
cleared=cleared,
)
def _load_contamination_markers(
*,
remote: str | None,
inspect: Callable[..., dict[str, Any]] | None = None,
load: Callable[..., dict[str, Any] | None] | None = None,
) -> tuple[ContaminationMarker, ...]:
return tuple(
_inspect_contamination(kind, remote=remote, inspect=inspect, load=load)
for kind in _CONTAMINATION_KINDS
)
def _build_session_rows(
inventory: InventorySnapshot,
contamination: tuple[ContaminationMarker, ...],
) -> tuple[SessionRow, ...]:
sessions_section = inventory.section("sessions")
leases_section = inventory.section("leases")
locks_section = inventory.section("locks")
# Ownership columns may only assert absence when their source read cleanly.
lease_authority = _section_status(leases_section)
# Locks carry claimant profile/username, not control-plane session ids, so a
# worktree binding is correlated through lease work numbers: it depends on
# the locks *and* the leases section.
worktree_authority = _combined_authority(
lease_authority, _section_status(locks_section)
)
leases_by_session: dict[str, list[dict[str, Any]]] = {}
for lease in (leases_section.items if leases_section else ()):
sid = str(lease.get("session_id") or "")
if sid:
leases_by_session.setdefault(sid, []).append(lease)
active_markers = [m for m in contamination if m.to_dict()["active"]]
marker_session_ids = {
m.session_id for m in active_markers if m.session_id
}
rows: list[SessionRow] = []
for raw in sessions_section.items if sessions_section else ():
sid = str(raw.get("session_id") or "")
if not sid:
continue
session_leases = leases_by_session.get(sid, [])
lease_ids = tuple(
str(lease["lease_id"])
for lease in session_leases
if lease.get("lease_id")
)
work_refs: list[str] = []
work_numbers: list[int] = []
for lease in session_leases:
kind = lease.get("work_kind")
number = lease.get("work_number")
if kind and number is not None:
work_refs.append(f"{kind}#{number}")
try:
work_numbers.append(int(number))
except (TypeError, ValueError):
pass
worktree_paths: list[str] = []
for lock in (locks_section.items if locks_section else ()):
try:
issue_no = int(lock.get("issue_number"))
except (TypeError, ValueError):
continue
if issue_no in work_numbers and lock.get("worktree_path"):
worktree_paths.append(str(lock["worktree_path"]))
stale_flags: list[str] = []
if raw.get("pid_alive") is False:
stale_flags.append("pid-dead")
status = str(raw.get("status") or "").lower()
if status and status not in {"active", "alive", "running", "ok"}:
stale_flags.append(f"status:{status}")
for lease in session_leases:
if lease.get("expired") is True:
stale_flags.append("lease-expired")
if str(lease.get("status") or "").lower() == "active" and lease.get(
"expired"
) is True:
stale_flags.append("active-lease-past-expiry")
contamination_flags: list[str] = []
if sid in marker_session_ids:
for marker in active_markers:
if marker.session_id == sid:
contamination_flags.append(marker.kind)
# Process-wide contamination with no session binding still surfaces
# against every live session so it cannot be silent (#630).
for marker in active_markers:
if not marker.session_id and marker.kind not in contamination_flags:
contamination_flags.append(f"{marker.kind}:process-wide")
rows.append(
SessionRow(
session_id=sid,
role=raw.get("role"),
profile=raw.get("profile"),
namespace=raw.get("namespace"),
pid=raw.get("pid") if isinstance(raw.get("pid"), int) else None,
pid_alive=raw.get("pid_alive")
if isinstance(raw.get("pid_alive"), bool)
else None,
status=raw.get("status"),
started_at=raw.get("started_at"),
last_heartbeat_at=raw.get("last_heartbeat_at"),
lease_ids=lease_ids,
work_refs=tuple(work_refs),
worktree_paths=tuple(worktree_paths),
stale_flags=tuple(dict.fromkeys(stale_flags)),
contamination_flags=tuple(dict.fromkeys(contamination_flags)),
lease_authority=lease_authority,
worktree_authority=worktree_authority,
)
)
return tuple(rows)
def load_session_view_snapshot(
*,
load_runtime: Callable[..., RuntimeSnapshot] | None = None,
load_inventory: Callable[..., InventorySnapshot] | None = None,
inspect_contamination: Callable[..., dict[str, Any]] | None = None,
load_contamination_payload: Callable[..., dict[str, Any] | None] | None = None,
) -> SessionViewSnapshot:
"""Build the composed sessions/runtime view. Fail-soft on partial sources."""
runtime_loader = load_runtime or load_runtime_snapshot
inventory_loader = load_inventory or load_inventory_snapshot
fetch_error: str | None = None
try:
runtime = runtime_loader()
except Exception as exc: # noqa: BLE001
fetch_error = f"runtime snapshot failed: {type(exc).__name__}: {exc}"
# Minimal placeholder so the page still renders inventory + recovery.
from webui.runtime_health import RuntimeSnapshot as _RS
runtime = _RS(
project_id="unknown",
repo_root="",
remote="",
host="",
profile_name="unknown",
role_kind="unknown",
config_model="unknown",
profile_mode="unknown",
profile_source="unknown",
authenticated_username=None,
identity_error=str(exc),
repo_sha=None,
remote_master_sha=None,
commits_behind_master=None,
stale_runtime_warning=None,
shell_health={},
workflow_hashes=(),
schema_hashes=(),
restart_guidance="docs/mcp-namespace-eof-recovery.md",
fetch_error=str(exc),
)
try:
inventory = inventory_loader()
except Exception as exc: # noqa: BLE001
msg = f"inventory snapshot failed: {type(exc).__name__}: {exc}"
fetch_error = f"{fetch_error}; {msg}" if fetch_error else msg
inventory = load_inventory_snapshot(
db_path="/nonexistent-for-fail-soft",
lock_dir="/nonexistent-for-fail-soft",
)
remote = getattr(runtime, "remote", None)
contamination = _load_contamination_markers(
remote=remote,
inspect=inspect_contamination,
load=load_contamination_payload,
)
sessions = _build_session_rows(inventory, contamination)
return SessionViewSnapshot(
runtime=runtime,
inventory=inventory,
sessions=sessions,
contamination_markers=contamination,
fetch_error=fetch_error,
)
def snapshot_to_dict(snapshot: SessionViewSnapshot) -> dict[str, Any]:
"""JSON export for ``/api/sessions`` (read-only)."""
return {
"api_version": "v1",
"view": "runtime-sessions",
"issue": 641,
"fetch_error": snapshot.fetch_error,
"runtime": runtime_snapshot_to_dict(snapshot.runtime),
"inventory": inventory_snapshot_to_dict(snapshot.inventory),
"sessions": [row.to_dict() for row in snapshot.sessions],
"session_counts": {
"total": len(snapshot.sessions),
"stale": snapshot.stale_session_count,
"contaminated": snapshot.contaminated_session_count,
},
"ownership_authority_complete": snapshot.ownership_authority_complete,
"ownership_section_status": snapshot.ownership_section_status,
"ownership_note": (
"Every ownership source read cleanly; a session with no lease and no "
"worktree path genuinely holds neither."
if snapshot.ownership_authority_complete
else "An ownership source is degraded or unavailable. Empty lease_ids "
"and worktree_paths mean unknown, not unowned; consult each row's "
"lease_authority and worktree_authority."
),
"contamination_markers": [
marker.to_dict() for marker in snapshot.contamination_markers
],
"active_contamination": [
marker.to_dict() for marker in snapshot.active_contamination
],
"recovery_docs": [dict(doc) for doc in snapshot.recovery_docs],
"read_only": True,
"phase": 1,
"mutations": [],
}
-403
View File
@@ -1,403 +0,0 @@
"""HTML views for the Runtime and session view (Phase 1, #641).
Read-only composition of runtime health (#430) and inventory sessions /
namespaces / worktrees (#636). Surfaces stale and contamination indicators
when detectable. Recovery links point only at sanctioned reconnect/restart
docs — never at manual process kill (#630).
"""
from __future__ import annotations
from html import escape
from typing import Sequence
from webui.inventory import STATUS_OK
from webui.layout import render_page
from webui.session_loader import (
ContaminationMarker,
SessionRow,
SessionViewSnapshot,
)
def _badge(text: str, css: str) -> str:
return f'<span class="badge {css}">{escape(text)}</span>'
def _flags(flags: Sequence[str], *, css: str) -> str:
if not flags:
return '<span class="muted">—</span>'
return " ".join(_badge(flag, css) for flag in flags)
def _unproven_cell(status: str) -> str:
"""Render an ownership column whose backing inventory section failed to read.
Never "none" and never "unbound": an unreadable source proves nothing about
ownership, and claiming otherwise is the exact failure the
``ownership_authority_complete`` invariant exists to prevent.
"""
return (
f'<span class="muted">unknown (inventory {escape(status)})</span><br>'
f'{_badge("authority unproven", "badge-health-degraded")}'
)
def _runtime_banner(snapshot: SessionViewSnapshot) -> str:
runtime = snapshot.runtime
stale = runtime.stale_runtime_warning
stale_html = ""
if stale:
stale_html = (
f'<div class="health-card health-stale" style="margin-top:0.75rem;">'
f"<strong>Stale runtime:</strong> {escape(stale)}</div>"
)
identity = runtime.authenticated_username or "unresolved"
if runtime.identity_error:
identity = f"unresolved ({runtime.identity_error})"
return f"""<div class="health-card">
<h3>Runtime context</h3>
<table class="detail">
<tr><th>Profile</th><td><code>{escape(runtime.profile_name)}</code></td></tr>
<tr><th>Role kind</th><td>{escape(runtime.role_kind)}</td></tr>
<tr><th>Identity</th><td>{escape(str(identity))}</td></tr>
<tr><th>Remote / host</th>
<td><code>{escape(runtime.remote)}</code> · <code>{escape(runtime.host)}</code></td>
</tr>
<tr><th>Local HEAD</th>
<td><code>{escape(runtime.repo_sha or "unknown")}</code></td>
</tr>
<tr><th>Remote master</th>
<td><code>{escape(runtime.remote_master_sha or "unknown")}</code></td>
</tr>
<tr><th>Commits behind</th>
<td>{escape(str(runtime.commits_behind_master if runtime.commits_behind_master is not None else "unknown"))}</td>
</tr>
</table>
<p class="muted">Full runtime detail: <a href="/runtime">/runtime</a> ·
Inventory API: <a href="/api/v1/inventory"><code>/api/v1/inventory</code></a></p>
{stale_html}
</div>"""
def _summary_bar(snapshot: SessionViewSnapshot) -> str:
total = len(snapshot.sessions)
stale = snapshot.stale_session_count
contaminated = snapshot.contaminated_session_count
active_markers = len(snapshot.active_contamination)
inv_status = snapshot.inventory.status
authority_complete = snapshot.ownership_authority_complete
authority_text = "complete" if authority_complete else "incomplete"
authority_css = "badge-health-ok" if authority_complete else "badge-health-degraded"
return f"""<div class="health-card" style="display:flex; flex-wrap:wrap; gap:1rem; align-items:center;">
<div><strong>Sessions:</strong> <span class="badge badge-health-ok">{total}</span></div>
<div><strong>Stale:</strong> <span class="badge badge-stale">{stale}</span></div>
<div><strong>Contaminated:</strong> <span class="badge badge-blocked">{contaminated}</span></div>
<div><strong>Active markers:</strong> <span class="badge badge-health-unproven">{active_markers}</span></div>
<div><strong>Inventory:</strong> <span class="badge badge-health-skipped">{escape(inv_status)}</span></div>
<div><strong>Ownership authority:</strong> <span class="badge {authority_css}">{authority_text}</span></div>
</div>"""
def _ownership_caveat(snapshot: SessionViewSnapshot) -> str:
"""Name the unreadable ownership sections, or render nothing when all read."""
if snapshot.ownership_authority_complete:
return ""
degraded = ", ".join(
f"{name}: {status}"
for name, status in snapshot.ownership_section_status.items()
if status != STATUS_OK
)
return (
'<div class="health-card health-stale">'
"<strong>Ownership authority incomplete:</strong> "
f"{escape(degraded)}. Columns marked <em>unknown</em> could not be read. "
"No session below may be treated as holding no lease or no worktree "
"binding — absence of evidence is not evidence of absence.</div>"
)
def _render_session_row(row: SessionRow) -> str:
pid = "" if row.pid is None else str(row.pid)
pid_alive = "" if row.pid_alive is None else ("alive" if row.pid_alive else "dead")
pid_css = (
"badge-health-ok"
if row.pid_alive is True
else ("badge-blocked" if row.pid_alive is False else "badge-health-skipped")
)
# An empty tuple only means "holds none" when its source read cleanly.
if row.lease_authority != STATUS_OK:
lease_cell = _unproven_cell(row.lease_authority)
else:
leases = (
", ".join(f"<code>{escape(lid)}</code>" for lid in row.lease_ids)
if row.lease_ids
else '<span class="muted">none</span>'
)
work = (
", ".join(escape(ref) for ref in row.work_refs)
if row.work_refs
else '<span class="muted">—</span>'
)
lease_cell = (
f'{leases}<div class="muted" style="font-size:0.82rem; '
f'margin-top:0.2rem;">{work}</div>'
)
if row.worktree_authority != STATUS_OK:
worktrees = _unproven_cell(row.worktree_authority)
else:
worktrees = (
"<br>".join(f"<code>{escape(path)}</code>" for path in row.worktree_paths)
if row.worktree_paths
else '<span class="muted">unbound</span>'
)
return f"""<tr>
<td><code>{escape(row.session_id)}</code></td>
<td>
<div><code>{escape(str(row.role or ""))}</code> / <code>{escape(str(row.profile or ""))}</code></div>
<div class="muted" style="font-size:0.82rem;">ns: <code>{escape(str(row.namespace or ""))}</code></div>
</td>
<td>
<code>{escape(pid)}</code>
{_badge(pid_alive, pid_css)}
</td>
<td>{escape(str(row.status or ""))}<div class="muted" style="font-size:0.82rem;">{escape(str(row.last_heartbeat_at or ""))}</div></td>
<td>{lease_cell}</td>
<td>{worktrees}</td>
<td>{_flags(row.stale_flags, css="badge-stale")}</td>
<td>{_flags(row.contamination_flags, css="badge-blocked")}</td>
</tr>"""
def _sessions_table(snapshot: SessionViewSnapshot) -> str:
rows: Sequence[SessionRow] = snapshot.sessions
if not rows:
sessions_status = snapshot.ownership_section_status.get("sessions", STATUS_OK)
if sessions_status != STATUS_OK:
return (
'<p class="muted">Session inventory is '
f"<strong>{escape(sessions_status)}</strong> — the session list "
"could not be read. This is not evidence that no sessions "
"exist.</p>"
)
return (
'<p class="muted">No control-plane sessions recorded. Inventory may '
"be unavailable, or no MCP workers have registered yet.</p>"
)
body = "".join(_render_session_row(row) for row in rows)
return f"""<table class="registry">
<thead>
<tr>
<th>Session</th>
<th>Role / profile / namespace</th>
<th>PID</th>
<th>Status</th>
<th>Leases / work</th>
<th>Worktree binding</th>
<th>Stale</th>
<th>Contamination</th>
</tr>
</thead>
<tbody>
{body}
</tbody>
</table>"""
def _namespaces_section(snapshot: SessionViewSnapshot) -> str:
section = snapshot.inventory.section("namespaces")
if section is None:
return (
'<div class="prompt-card"><h3>Namespaces</h3>'
'<p class="muted">Namespaces section not loaded.</p></div>'
)
if not section.ok:
return f"""<div class="prompt-card">
<h3>Namespaces {_badge(section.status, "badge-health-degraded")}</h3>
<p class="muted">{escape(section.reason or "unavailable")}</p>
</div>"""
rows = []
for item in section.items:
caps = item.get("capability_summary") or {}
cap_bits = ", ".join(
name for name, ok in sorted(caps.items()) if ok
) or "none"
rows.append(
"<tr>"
f"<td><code>{escape(str(item.get('mcp_namespace') or ''))}</code></td>"
f"<td><code>{escape(str(item.get('profile_name') or ''))}</code></td>"
f"<td>{escape(str(item.get('role') or ''))}</td>"
f"<td>{escape(cap_bits)}</td>"
f"<td>{'yes' if item.get('active') else 'no'}</td>"
"</tr>"
)
reason = (
f'<p class="muted">{escape(section.reason)}</p>'
if section.reason
else ""
)
return f"""<div class="prompt-card">
<h3>Namespaces / capabilities</h3>
{reason}
<table class="registry">
<thead>
<tr>
<th>Namespace</th>
<th>Profile</th>
<th>Role</th>
<th>Capabilities</th>
<th>Active in process</th>
</tr>
</thead>
<tbody>
{"".join(rows) if rows else '<tr><td colspan="5" class="muted">No namespace rows.</td></tr>'}
</tbody>
</table>
</div>"""
def _worktrees_section(snapshot: SessionViewSnapshot) -> str:
section = snapshot.inventory.section("worktrees")
if section is None:
return ""
if not section.ok and not section.items:
return f"""<div class="prompt-card">
<h3>Worktrees {_badge(section.status, "badge-health-degraded")}</h3>
<p class="muted">{escape(section.reason or "unavailable")}</p>
</div>"""
rows = []
for item in section.items[:50]:
rows.append(
"<tr>"
f"<td><code>{escape(str(item.get('rel_path') or item.get('path') or ''))}</code></td>"
f"<td><code>{escape(str(item.get('branch') or ''))}</code></td>"
f"<td>{escape(str(item.get('classification') or ''))}</td>"
f"<td>{'yes' if item.get('registered_worktree') else 'no'}</td>"
f"<td>{'dirty' if item.get('dirty') else 'clean'}</td>"
"</tr>"
)
more = ""
if len(section.items) > 50:
more = f'<p class="muted">Showing 50 of {len(section.items)}. Full list: <a href="/worktrees">/worktrees</a>.</p>'
return f"""<div class="prompt-card">
<h3>Worktree bindings</h3>
<p class="muted">Registered issue worktrees under <code>branches/</code>. Hygiene detail: <a href="/worktrees">/worktrees</a>.</p>
<table class="registry">
<thead>
<tr>
<th>Path</th>
<th>Branch</th>
<th>Classification</th>
<th>Registered</th>
<th>State</th>
</tr>
</thead>
<tbody>
{"".join(rows) if rows else '<tr><td colspan="5" class="muted">No worktrees recorded.</td></tr>'}
</tbody>
</table>
{more}
</div>"""
def _contamination_section(markers: Sequence[ContaminationMarker]) -> str:
if not markers:
return (
'<div class="prompt-card"><h3>Contamination markers</h3>'
'<p class="muted">No contamination kinds inspected.</p></div>'
)
rows = []
for marker in markers:
active = marker.to_dict()["active"]
status = "ACTIVE" if active else ("cleared" if marker.cleared else "absent")
css = "badge-blocked" if active else "badge-health-ok"
rows.append(
"<tr>"
f"<td><code>{escape(marker.kind)}</code></td>"
f"<td>{_badge(status, css)}</td>"
f"<td>{escape(marker.reason_class or '')}</td>"
f"<td><code>{escape(marker.session_id or '')}</code></td>"
f"<td>{escape(marker.command_summary or marker.summary)}</td>"
"</tr>"
)
return f"""<div class="prompt-card">
<h3>Contamination markers (#630 / #671)</h3>
<p class="muted">Durable markers only — never silent when present. Clearance is reconciler-only.</p>
<table class="registry">
<thead>
<tr>
<th>Kind</th>
<th>State</th>
<th>Reason class</th>
<th>Session</th>
<th>Summary</th>
</tr>
</thead>
<tbody>
{"".join(rows)}
</tbody>
</table>
</div>"""
def _recovery_section(snapshot: SessionViewSnapshot) -> str:
items = []
for doc in snapshot.recovery_docs:
items.append(
"<li>"
f"<code>{escape(doc['path'])}</code> — "
f"<strong>{escape(doc['label'])}</strong>: {escape(doc['note'])}"
"</li>"
)
return f"""<div class="prompt-card">
<h3>Sanctioned recovery (read-only)</h3>
<p class="muted">This view does <strong>not</strong> restart, kill, or take over sessions.
Manual <code>pkill</code> / <code>kill</code> of MCP daemons is contamination (#630), not recovery.</p>
<ul class="reasons">
{"".join(items)}
<li>Prefer IDE/client reconnect (<code>/mcp reconnect</code>) or an operator-owned restart recorded in the restart inventory.</li>
</ul>
</div>"""
def render_sessions_page(snapshot: SessionViewSnapshot) -> str:
"""Render the full HTML body for the runtime/session view."""
error_block = ""
if snapshot.fetch_error:
error_block = (
f'<div class="health-card health-stale"><strong>Partial load:</strong> '
f"{escape(snapshot.fetch_error)}</div>"
)
if snapshot.runtime.fetch_error:
error_block += (
f'<div class="health-card health-stale"><strong>Runtime note:</strong> '
f"{escape(snapshot.runtime.fetch_error)}</div>"
)
body = f"""
{error_block}
{_runtime_banner(snapshot)}
{_summary_bar(snapshot)}
<div class="prompt-card">
<h3>Sessions</h3>
<p class="muted">Control-plane sessions correlated with leases and worktree bindings.
Stale and contamination flags are fail-soft: absence of a marker is not proof of cleanliness when inventory is degraded.
Lease and worktree columns read <em>unknown (inventory …)</em> when their source could not be loaded.</p>
{_ownership_caveat(snapshot)}
{_sessions_table(snapshot)}
</div>
{_namespaces_section(snapshot)}
{_worktrees_section(snapshot)}
{_contamination_section(snapshot.contamination_markers)}
{_recovery_section(snapshot)}
"""
return render_page(title="Sessions", body_html=f"""<h2>Runtime and sessions</h2>
<p class="meta">Phase 1 read-only view (#641). Combines runtime health (#430) with
unified inventory sessions/namespaces/worktrees (#636). No restart or session-takeover controls.</p>
{body}""")
+2 -4
View File
@@ -278,10 +278,8 @@ def _recovery_card() -> str:
"controls arrive in Phase 2 (#642); until then recovery runs through "
"the sanctioned client reconnect / operator restart path.</p>"
"<ul class='reasons'>"
"<li><a href='/runtime'>Runtime health</a> — active profile, workflow "
"hashes, and shell health.</li>"
"<li><a href='/sessions'>Runtime and sessions</a> — namespaces, session "
"rows, worktree bindings, and contamination markers (#641).</li>"
"<li><a href='/runtime'>Runtime and session view</a> — active profile, "
"workflow hashes, and shell health.</li>"
"<li>Reconnect the MCP client from the IDE, then re-run the blocked "
"cycle. Never kill the daemon process manually: unmanaged kills are "
"recorded as runtime contamination (#630).</li>"