Compare commits

..
Author SHA1 Message Date
jcwalker3 c4d089f931 Merge branch 'master' into feat/issue-636-inventory-api 2026-07-24 06:55:53 -05:00
sysadmin 5e935dffb4 Merge pull request 'feat(webui): system-health dashboard (Closes #639)' (#862) from feat/issue-639-webui-system-health-dashboard into master 2026-07-24 06:28:46 -05:00
jcwalker3 baf3a474df Merge branch 'master' into feat/issue-636-inventory-api 2026-07-24 06:16:42 -05:00
jcwalker3 82464f4054 Merge branch 'master' into feat/issue-639-webui-system-health-dashboard 2026-07-24 06:16:12 -05:00
sysadmin c33c69b3f3 Merge pull request 'fix: conflict-fix lease lifecycle chain termination and TTL handling (#842)' (#846) from fix/issue-842-conflict-fix-lease-lifecycle into master 2026-07-24 06:11:09 -05:00
sysadminandClaude Opus 4.8 784369cc25 merge(master): sync PR #838 with master after #875
No content conflicts. Master advanced with the #658 restart coordinator;
inventory API surfaces auto-merge cleanly.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-07-24 07:07:32 -04:00
sysadmin 67b4889984 Merge pull request 'feat(mcp-health): MCP restart coordinator and impact analysis (Closes #658)' (#875) from feat/issue-658-mcp-restart-coordinator into master 2026-07-24 06:04:14 -05:00
sysadmin e0b87a0ae5 Merge remote-tracking branch 'prgs/master' into feat/issue-636-inventory-api 2026-07-24 07:01:45 -04:00
sysadmin 34173e079c Merge branch 'master' of https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools into feat/issue-636-inventory-api
# Conflicts:
#	docs/webui-local-dev.md
#	webui/app.py
2026-07-24 07:01:32 -04:00
jcwalker3 fd558ce5d8 Merge branch 'master' into feat/issue-658-mcp-restart-coordinator 2026-07-24 05:54:50 -05:00
sysadmin ae1161524d Merge pull request 'fix(author): bootstrap recovery for dirty orphaned issue worktrees (#860)' (#861) from fix/issue-860-dirty-orphan-worktree-recovery into master 2026-07-24 05:52:24 -05:00
jcwalker3 d456a763fa Merge branch 'master' into feat/issue-658-mcp-restart-coordinator 2026-07-24 05:48:25 -05:00
jcwalker3 78e3befbbb Merge branch 'master' into feat/issue-658-mcp-restart-coordinator 2026-07-24 02:15:32 -05:00
jcwalker3 d7e69fbe77 Merge branch 'master' into feat/issue-639-webui-system-health-dashboard 2026-07-24 02:07:53 -05:00
jcwalker3 467e35504c Merge branch 'master' into fix/issue-842-conflict-fix-lease-lifecycle 2026-07-24 02:06:45 -05:00
sysadminandClaude Opus 4.8 2d0d8a682b feat(mcp-health): add MCP restart coordinator and impact analysis (Closes #658)
Child of umbrella #655 (governed MCP restart coordination); builds on the
#657 restart-path inventory. Adds a central coordinator that evaluates live
control-plane state before a restart and returns a blast-radius impact
preview, so operators and the web console (#642/#652) can see what a restart
would disrupt before concurrent LLM work is destroyed.

Changes
- restart_coordinator.py (new) — pure classification: inventory -> impact
  report DTO (RestartImpactReport/SessionImpact/LeaseImpact). Verdicts:
  safe / unsafe / override. Never restarts anything; fails closed on an
  incomplete inventory.
- control_plane_db.py — additive ControlPlaneDB.list_sessions() read-only
  session inventory (the process-level unit a restart kills).
- gitea_mcp_server.py — new dry-run MCP tool gitea_request_mcp_restart:
  gathers sessions/leases/terminal-lock from the #613 DB, calls the
  coordinator, returns the report. Override authority is read from the
  environment, never self-asserted (#630/#710 F1 pattern). Apply is gated
  by a later drain proof (non-goal here).
- docs/mcp-restart-coordinator.md + docs/mcp-restart-impact-sample.json — doc
  and a real dry-run sample report.
- docs/mcp-tool-inventory.md — register the new tool (inventory sync).
- tests/test_restart_coordinator.py (new) — 15 tests: multi-session fixtures,
  deny-when-critical-section-open, fail-closed deny, override, terminal lock,
  stale heartbeat, JSON-serializable DTO, list_sessions.

Tests: pytest tests/test_restart_coordinator.py -> 15 passed. Full suite:
13 failed / 4753 passed; all 13 reproduce identically on clean master
@ef14622 (0 regressions). The residual test_issue_781 doc-registry failure
is a pre-existing baseline gap for gitea_rebind_dirty_same_claimant_author_session
(merged #864, undocumented on master) — out of scope for #658.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-07-24 02:14:27 -04:00
sysadmin f80e3b33b0 Merge remote-tracking branch 'prgs/master' into feat/issue-639-webui-system-health-dashboard 2026-07-23 21:14:50 -04:00
jcwalker3 3a0d9e24ea Merge branch 'master' into fix/issue-842-conflict-fix-lease-lifecycle 2026-07-23 19:52:51 -05:00
jcwalker3 dc0a05e5e9 Merge branch 'master' into fix/issue-842-conflict-fix-lease-lifecycle 2026-07-23 19:13:22 -05:00
sysadminandClaude Opus 4.8 edd5f813b2 Merge master into feat/issue-639-webui-system-health-dashboard
Resolve the #638 shell landing against the #639 dashboard:

- webui/layout.py: drop the flat NAV_ITEMS tuple in favor of master's
  grouped NAV_GROUPS nav-config module.
- webui/nav.py: register /system-health as a live item in the Health
  group, satisfying issue #639 AC5 through the canonical nav source.
- docs/webui-local-dev.md: keep both additive sections (#638 shell and
  #639 dashboard).
- tests/test_webui_system_health_dashboard.py: assert the nav entry via
  iter_nav_items() instead of the removed NAV_ITEMS tuple.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-07-23 19:56:57 -04:00
jcwalker3 8b34f9da0a Merge branch 'master' into fix/issue-842-conflict-fix-lease-lifecycle 2026-07-23 16:21:17 -05:00
sysadminandClaude Opus 4.8 ecda200180 feat(webui): system-health dashboard (Closes #639)
Phase 1 child of the Web Console epic #631. Adds the operator-facing
system-health dashboard on top of the read-only system-health API landed
by #634, so runtime problems are visible on a surface instead of being
discovered late through failed LLM sessions.

- webui/system_health_views.py (new): renders the SystemHealthSnapshot as
  readiness, stale-runtime parity, version/uptime, dependency, MCP
  namespace, probe-error, and recovery cards.
- webui/app.py: GET /system-health, sharing load_system_health() with the
  JSON API so page and API cannot disagree. ?deep=1 behaves as on the API.
- webui/layout.py: nav entry and health card/badge styles.
- tests/test_webui_system_health_dashboard.py (new, 26 cases).
- docs/webui-local-dev.md: route, field authority, and redaction split.

Readiness honesty is preserved from the API: ready and readiness_complete
render separately, a probe that did not run is listed under "Not probed"
rather than counted healthy, and mutation safety is never claimed when the
runtime is stale or parity is indeterminate.

Redaction is split by field kind. Free text (probe details, reasons, probe
errors) passes through system_health.redact. Structured fields (commit
SHAs, probe names, statuses, timestamps) are HTML-escaped only: redact's
opaque-token rule matches any run of 32 or more characters, so routing a
40-character git SHA through it rendered "[redacted]" and blanked the
parity evidence the page exists to show.

Non-goals honored: no restart or reload controls (Phase 2, #642), no
manual process-kill guidance (#630). Read-only throughout.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-07-23 16:08:36 -04:00
jcwalker3 04d9df559e Merge branch 'master' into fix/issue-842-conflict-fix-lease-lifecycle 2026-07-23 12:30:59 -05:00
sysadmin 8a63476787 fix: conflict-fix lease lifecycle chain termination and TTL handling (Closes #847, Refs #842) 2026-07-23 04:46:30 -04:00
sysadminandClaude Opus 4.8 b7a63a5579 feat(webui): unified session/lease/lock/worktree inventory API (Closes #636)
Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-07-22 21:13:54 -04:00
17 changed files with 3645 additions and 2 deletions
+29
View File
@@ -599,6 +599,35 @@ class ControlPlaneDB:
(_ts(), session_id), (_ts(), session_id),
) )
def list_sessions(
self,
*,
statuses: Sequence[str] | None = None,
limit: int = 500,
) -> list[dict[str, Any]]:
"""List session rows for restart / impact analysis (#658).
Read-only. Sessions are the process-level unit an MCP restart
disrupts, so the restart coordinator inventories them to compute blast
radius. Optional ``statuses`` filter (e.g. ``('active',)``) narrows to
live rows. Never returns secrets — only operational metadata.
"""
clauses: list[str] = []
params: list[Any] = []
if statuses:
placeholders = ", ".join("?" for _ in statuses)
clauses.append(f"status IN ({placeholders})")
params.extend(statuses)
where = ("WHERE " + " AND ".join(clauses)) if clauses else ""
sql = (
f"SELECT * FROM sessions {where} "
"ORDER BY last_heartbeat_at DESC LIMIT ?"
)
params.append(max(1, int(limit)))
with self._tx(immediate=False) as conn:
rows = conn.execute(sql, params).fetchall()
return [dict(r) for r in rows]
# ── work items ──────────────────────────────────────────────────────── # ── work items ────────────────────────────────────────────────────────
def upsert_work_item( def upsert_work_item(
+95
View File
@@ -0,0 +1,95 @@
# MCP restart coordinator and impact analysis (#658)
Before any sanctioned MCP restart, a central coordinator evaluates the live
control-plane state and produces an **impact preview** so operators and the web
console (#642 / #652) can see the blast radius *before* concurrent LLM work is
disrupted. Uncoordinated restarts destroy in-flight author/reviewer/merger work
and give operators no way to see what they are about to break.
This lands the coordinator + impact DTO + a dry-run MCP tool. It is the single
sanctioned entry point for restart evaluation post-#657 (which inventoried the
restart/reload/kill paths). The **mutative apply** path — actually performing a
restart — is a later child gated by a drain proof and is explicitly out of
scope here.
## Components
| Piece | Where | Responsibility |
|-------|-------|----------------|
| `restart_coordinator.evaluate_restart_impact` | `restart_coordinator.py` | Pure classification: inventory → impact report DTO. No I/O, no restart. |
| `RestartImpactReport` / `SessionImpact` / `LeaseImpact` | `restart_coordinator.py` | Console-facing DTO (`.as_dict()` is JSON-serializable). |
| `ControlPlaneDB.list_sessions` | `control_plane_db.py` | Read-only session inventory (the process-level unit a restart kills). |
| `gitea_request_mcp_restart` | `gitea_mcp_server.py` | MCP tool: gathers inventory from the #613 DB, calls the coordinator, returns the report. Dry-run only. |
## Dimensions evaluated
The coordinator classifies the inventory across the dimensions #658 requires:
- **Sessions** — every active MCP session; a restart terminates all of them.
Liveness = `status == active` **and** the owner pid is alive **and** the
heartbeat is fresh (default window 15 min). Dead/stale sessions do not count
toward blast radius.
- **Leases / locks** — control-plane leases joined with work items and their
freshness (`lease_lifecycle.classify_lease_freshness`). Only `active` (live
owner) leases are *disruptive*; expired / released / dead-process leases never
withhold a restart.
- **Issue / PR work** — the issues and PRs behind disruptive leases.
- **Mutations / critical sections** — a live lease carrying an author worktree
or a mutating phase (`implementing`, `publishing`, `merging`, …) is a
critical section a restart must not sever.
- **Terminal (merge) lock** — an active terminal lock always makes a restart
unsafe.
- **Prior recovery attempts** — narrower recovery already tried (e.g. sanctioned
client reconnects) is echoed so the operator sees the escalation history.
## Verdict
Exactly three verdicts, matching the acceptance criteria:
| Verdict | `allow_restart` | Meaning |
|---------|-----------------|---------|
| `safe` | `true` | No other live sessions, no live leases, no terminal lock. |
| `unsafe` | `false` | Live work would be disrupted and no operator override is present — **or** the inventory could not be completed (fail closed). |
| `override` | `true` | Live work present, but an operator override accepts the blast radius. |
`override_would_allow` tells the console whether an override path exists for the
current state. `blast_radius` is a `none` / `low` / `medium` / `high` severity
band derived from the affected session and work counts.
### Fail closed
If the control-plane inventory cannot be completed (DB unavailable, a listing
failed), `inventory_complete` is `false` and the verdict is `unsafe` / deny. An
incomplete evaluation must never green-light a restart.
### Operator override authority
Override authority is read from the environment variable
`GITEA_OPERATOR_RESTART_OVERRIDE_AUTHORIZATION` and **never** from a tool
argument. A worker session cannot set an environment variable on an
already-running daemon, so override cannot be self-asserted (same pattern as the
#630 daemon-maintenance authorization). The `request_override` tool argument only
expresses caller intent; it takes effect solely when the environment
authorization is present.
## The tool
```text
gitea_request_mcp_restart(remote, host, org, repo,
dry_run=True, request_override=False,
session_id=None, limit=200)
```
Read-only, dry-run, and it **never restarts anything**. `apply_supported` is
always `false`; passing `dry_run=False` performs no restart and reports that
apply is gated by a drain proof (a separate child).
## Audit
Every evaluation carries an `audit_record` (event, coordinator version, verdict,
allow decision, blast radius, counts, timestamp) so restart decisions are
auditable. No secrets flow through the coordinator — session ids, pids, and
profiles are operational metadata only.
A representative dry-run report is in
[`mcp-restart-impact-sample.json`](./mcp-restart-impact-sample.json).
+148
View File
@@ -0,0 +1,148 @@
{
"coordinator_version": "1.0.0-issue-658",
"evaluated_at": "2026-07-24T06:00:00+00:00",
"dry_run": true,
"restart_performed": false,
"inventory_complete": true,
"incomplete_reasons": [],
"verdict": "unsafe",
"allow_restart": false,
"override_would_allow": true,
"operator_override": false,
"blast_radius": "high",
"reasons": [
"live work would be disrupted; restart denied without operator override",
"1 critical section(s) in flight (active lease with a live owner)"
],
"affected_sessions": [
{
"session_id": "prgs-author-30988-d6f43c25",
"role": "author",
"profile": "prgs-author",
"pid": 1,
"status": "active",
"alive": true,
"heartbeat_stale": false,
"is_requester": false,
"live": true
},
{
"session_id": "prgs-reviewer-4157-0ce9",
"role": "reviewer",
"profile": "prgs-reviewer",
"pid": 1,
"status": "active",
"alive": true,
"heartbeat_stale": false,
"is_requester": true,
"live": true
}
],
"affected_leases": [
{
"lease_id": "lease-abc",
"session_id": "prgs-author-30988-d6f43c25",
"role": "author",
"phase": "implementing",
"freshness": "active",
"work_kind": "issue",
"work_number": 658,
"worktree_path": "/repo/branches/feat-issue-658",
"disruptive": true,
"is_mutation": true,
"is_critical_section": true
},
{
"lease_id": "lease-dead",
"session_id": "prgs-author-91485",
"role": "author",
"phase": "allocated",
"freshness": "stale_dead_process",
"work_kind": "issue",
"work_number": 651,
"worktree_path": null,
"disruptive": false,
"is_mutation": false,
"is_critical_section": false
}
],
"critical_sections": [
{
"lease_id": "lease-abc",
"session_id": "prgs-author-30988-d6f43c25",
"role": "author",
"phase": "implementing",
"freshness": "active",
"work_kind": "issue",
"work_number": 658,
"worktree_path": "/repo/branches/feat-issue-658",
"disruptive": true,
"is_mutation": true,
"is_critical_section": true
}
],
"affected_issues": [
658
],
"affected_prs": [],
"mutations": [
{
"lease_id": "lease-abc",
"session_id": "prgs-author-30988-d6f43c25",
"role": "author",
"phase": "implementing",
"freshness": "active",
"work_kind": "issue",
"work_number": 658,
"worktree_path": "/repo/branches/feat-issue-658",
"disruptive": true,
"is_mutation": true,
"is_critical_section": true
}
],
"terminal_lock": null,
"ack_state": {
"prgs-author-30988-d6f43c25": "pending"
},
"prior_recovery_attempts": [
{
"kind": "client_reconnect",
"at": "2026-07-24T06:00:00+00:00",
"outcome": "insufficient"
}
],
"counts": {
"sessions_total": 2,
"sessions_live_other": 1,
"leases_total": 2,
"leases_disruptive": 1,
"critical_sections": 1,
"mutations": 1,
"affected_issues": 1,
"affected_prs": 0,
"prior_recovery_attempts": 1
},
"audit_record": {
"event": "restart_impact_evaluated",
"coordinator_version": "1.0.0-issue-658",
"evaluated_at": "2026-07-24T06:00:00+00:00",
"dry_run": true,
"operator_override": false,
"requesting_session_id": "prgs-reviewer-4157-0ce9",
"inventory_complete": true,
"verdict": "unsafe",
"allow_restart": false,
"blast_radius": "high",
"counts": {
"sessions_total": 2,
"sessions_live_other": 1,
"leases_total": 2,
"leases_disruptive": 1,
"critical_sections": 1,
"mutations": 1,
"affected_issues": 1,
"affected_prs": 0,
"prior_recovery_attempts": 1
}
}
}
+1
View File
@@ -135,6 +135,7 @@ that gates each call, not which tools exist.
- `gitea_release_merger_pr_lease` - `gitea_release_merger_pr_lease`
- `gitea_release_reviewer_pr_lease` - `gitea_release_reviewer_pr_lease`
- `gitea_release_workflow_lease` - `gitea_release_workflow_lease`
- `gitea_request_mcp_restart`
- `gitea_resolve_task_capability` - `gitea_resolve_task_capability`
- `gitea_resume_review_draft` - `gitea_resume_review_draft`
- `gitea_review_pr` - `gitea_review_pr`
+79
View File
@@ -54,6 +54,7 @@ status, onboarding checklist state, and the fail-closed error payloads (#635).
| `/` | Home / operator overview | | `/` | Home / operator overview |
| `/health` | JSON liveness (`status`, `service`, `mode`, `timestamp`, `uptime_seconds`) | | `/health` | JSON liveness (`status`, `service`, `mode`, `timestamp`, `uptime_seconds`) |
| `/api/v1/system/health` | Structured read-only system health (#634) | | `/api/v1/system/health` | Structured read-only system health (#634) |
| `/system-health` | System-health dashboard — readiness, version/uptime, dependencies, MCP namespaces, stale-runtime parity (#639) |
| `/queue` | Live PR and issue queue dashboard (#429) | | `/queue` | Live PR and issue queue dashboard (#429) |
| `/api/queue` | JSON queue export with pagination metadata | | `/api/queue` | JSON queue export with pagination metadata |
| `/projects` | Project registry list with status and onboarding progress (#427, #635) | | `/projects` | Project registry list with status and onboarding progress (#427, #635) |
@@ -258,6 +259,37 @@ Not-yet-implemented surfaces (`/sessions`, `/inventory`, `/timeline`,
surfaces are backed by #636). Mutating methods on stub routes still fail closed surfaces are backed by #636). Mutating methods on stub routes still fail closed
with `read-only-mvp`. with `read-only-mvp`.
## System-health dashboard (#639)
`/system-health` renders the same snapshot the `/api/v1/system/health` API
returns, so the page and the API can never disagree. Cards: overall readiness,
stale-runtime parity, version and uptime, dependency probes, MCP namespaces,
probe errors (only when present), and recovery pointers. `?deep=1` opts into
the network probe exactly as the API does; the plain page load stays cheap.
Field authority and honesty rules:
* `ready` and `readiness_complete` are shown separately. A snapshot whose
required probes never ran is not the same as one that ran them and passed,
and the page never collapses the two into an unproven green.
* A probe that did not run appears under **Not probed**, never as healthy.
* `stale_runtime.mutation_safe` is displayed verbatim from the API. When the
runtime is stale, or when parity is indeterminate, the page warns and does
not claim mutation safety.
* MCP namespaces are reported `unproven`: the web process runs outside the
IDE-managed MCP client and cannot prove that path (#543).
Redaction is split by field kind. Free text — probe details, readiness and
parity reasons, probe errors — passes through `system_health.redact`.
Structured fields — commit SHAs, probe names, statuses, timestamps — are
HTML-escaped only, because `redact`'s opaque-token rule matches any run of 32
or more characters and would otherwise blank every 40-character git SHA, which
is precisely the evidence the parity view exists to show.
The dashboard is read-only: no restart, reload, or process-kill control. Those
arrive in Phase 2 (#642). Recovery guidance points at the sanctioned client
reconnect / operator restart path — never a manual daemon kill (#630).
## Deployment boundary (#435) ## Deployment boundary (#435)
MVP serves on loopback by default. Binding `0.0.0.0` or `::` is **refused** MVP serves on loopback by default. Binding `0.0.0.0` or `::` is **refused**
@@ -317,6 +349,53 @@ health, workflow/schema SHA-256 hashes, and stale-runtime warnings when the
checkout is behind merged safety-gate changes. Restart guidance links to #420; checkout is behind merged safety-gate changes. Restart guidance links to #420;
no tokens or MCP restart actions are exposed. no tokens or MCP restart actions are exposed.
## Inventory API (#636)
`GET /api/v1/inventory` returns one versioned, read-only snapshot that unifies
what the lease (#433), worktree (#432), and runtime (#430) MVP views each show
separately, so traffic-control and recovery consumers read the same source.
`GET /api/v1/inventory/{section}` returns a single section under the identical
schema (`sessions`, `leases`, `locks`, `worktrees`, `namespaces`); an unknown
section is a `404` with `error: unknown_section`. Both routes are `GET`-only.
Each section carries its own `status` (`ok` / `degraded` / `unavailable`), a
`reason` when not `ok`, and a `scan_ms`. A subsystem that cannot be read
degrades to a reasoned section; it never raises and never emits an empty list
that would read as "nothing is there".
### Field authority
Every section names where its rows came from; authorities are never blended.
| Section | Authority | Source |
|---|---|---|
| `sessions` | `control_plane_db` | #613 control-plane DB (`mode=ro`), authoritative for exclusive ownership (#600/#601) |
| `leases` | `control_plane_db` | #613 control-plane DB; degrades if the `work_items` table is absent |
| `locks` | `filesystem` | durable per-issue lock files (`issue_lock_store`) |
| `worktrees` | `filesystem` | registered git worktrees via the #432 hygiene scanner |
| `namespaces` | `filesystem` | the active profile serving this web process (others are not enumerable) |
The payload restates this map under `field_authority` for machine consumers.
### Ownership safety
`ownership_authority_complete` is true only when every ownership-bearing section
(`sessions`, `leases`, `locks`) read cleanly. While it is false, nothing is
reported as unowned and no collision is asserted from a degraded source —
absence of evidence is reported as absence of evidence, never as free work.
`collisions` surfaces detectable conflicts, each with a `kind` and `severity`:
`lock-without-worktree`, `duplicate-live-lock`, `live-lock-dead-owner` (unexpired
lease, dead pid — a #753 recovery candidate that would read as live to a naive
timestamp check), `stale-lock-dead-owner`, `expired-lock-live-owner` (the
#635/#760 daemon-pid deadlock), `concurrent-active-lease`, `active-lease-past-expiry`,
and `orphan-lease`. Collisions are emitted only from sections that read cleanly.
The control-plane DB is opened through a `mode=ro` URI so a read never creates
or migrates it; paths are collapsed against `$HOME`, URLs lose userinfo and
query strings, and credential-shaped values are redacted at the boundary. Lease
steal/release and worktree deletion are Phase 2+ and have no representation here.
## Workflow-event timeline (#637) ## Workflow-event timeline (#637)
`GET /api/v1/timeline` is a read-only, versioned aggregation of workflow `GET /api/v1/timeline` is a read-only, versioned aggregation of workflow
+151
View File
@@ -2040,6 +2040,7 @@ import dependency_graph # noqa: E402 # #784 durable dependency edges
import control_plane_db # noqa: E402 import control_plane_db # noqa: E402
import lease_lifecycle # noqa: E402 import lease_lifecycle # noqa: E402
import workflow_dashboard # noqa: E402 # #605 live queue/lease dashboard import workflow_dashboard # noqa: E402 # #605 live queue/lease dashboard
import restart_coordinator # noqa: E402 # #658 MCP restart coordinator/impact
import incident_bridge # noqa: E402 import incident_bridge # noqa: E402
import sentry_observability # noqa: E402 (#606 optional Sentry observability) import sentry_observability # noqa: E402 (#606 optional Sentry observability)
import sentry_incident_bridge # noqa: E402 (#607 Sentry→Gitea incident bridge) import sentry_incident_bridge # noqa: E402 (#607 Sentry→Gitea incident bridge)
@@ -22050,6 +22051,156 @@ def gitea_workflow_dashboard(
return payload return payload
@mcp.tool()
def gitea_request_mcp_restart(
remote: str = "dadeschools",
host: str | None = None,
org: str | None = None,
repo: str | None = None,
dry_run: bool = True,
request_override: bool = False,
session_id: str | None = None,
limit: int = 200,
) -> dict:
"""Evaluate a proposed MCP restart and return an impact preview (#658).
Central restart coordinator: gathers live control-plane state (sessions,
leases/locks, in-flight issue/PR work, mutations, worktrees) and returns a
blast-radius impact report with a ``safe`` / ``unsafe`` / ``override``
verdict, so the console (#642/#652) and operators can see what a restart
would disrupt *before* any concurrent LLM work is destroyed.
This tool is **dry-run and never restarts anything.** The mutative apply
path is a separate child gated by a drain proof (non-goal here); calling
with ``dry_run=False`` still performs no restart and reports that apply is
not yet available.
Operator override authority is read from the process environment
(``GITEA_OPERATOR_RESTART_OVERRIDE_AUTHORIZATION``), never self-asserted by
the requesting session: ``request_override`` only expresses caller intent
and takes effect solely when that environment authorization is present.
Fails closed: if the control-plane inventory cannot be completed, the
verdict is ``unsafe`` / deny (an incomplete evaluation must never green-light
a restart).
"""
read_block = _profile_operation_gate("gitea.read")
if read_block:
return {
"success": False,
"read_only": True,
"dry_run": True,
"restart_performed": False,
"reasons": read_block,
"permission_report": _permission_block_report("gitea.read"),
}
try:
h, o, r = _resolve(remote, host, org, repo)
except ValueError as exc:
return {
"success": False,
"read_only": True,
"dry_run": True,
"restart_performed": False,
"reasons": [str(exc)],
}
inventory_complete = True
incomplete_reasons: list[str] = []
sessions: list[dict] = []
leases: list[dict] = []
terminal_lock: dict | None = None
db, db_errs = _control_plane_db_or_error()
if db is None:
inventory_complete = False
incomplete_reasons.extend(
db_errs or ["control-plane DB unavailable; cannot evaluate restart"]
)
else:
try:
sessions = db.list_sessions(statuses=("active",), limit=max(1, int(limit)))
except Exception as exc: # noqa: BLE001
inventory_complete = False
incomplete_reasons.append(
f"session inventory failed: {_redact(str(exc))}"
)
try:
lease_result = lease_lifecycle.list_active_leases(
db,
remote=remote if remote in REMOTES else remote,
org=o,
repo=r,
role=None,
include_non_active=False,
limit=max(1, int(limit)),
)
leases = list(lease_result.get("leases") or [])
except Exception as exc: # noqa: BLE001
inventory_complete = False
incomplete_reasons.append(
f"lease inventory failed: {_redact(str(exc))}"
)
try:
terminal = db.get_active_terminal_lock(
remote=remote if remote in REMOTES else remote,
org=o,
repo=r,
)
if terminal:
terminal_lock = dict(terminal)
except Exception as exc: # noqa: BLE001
inventory_complete = False
incomplete_reasons.append(
f"terminal lock lookup failed: {_redact(str(exc))}"
)
profile = get_profile()
profile_name = (profile.get("profile_name") or "").strip() or "session"
sid = (session_id or "").strip() or f"{profile_name}-{os.getpid()}"
# Override authority is read from the environment only — a worker session
# cannot set an env var for an already-running daemon, so it cannot be
# self-asserted the way a tool argument could (#630/#710 F1 pattern).
operator_authorized = bool(
(os.environ.get("GITEA_OPERATOR_RESTART_OVERRIDE_AUTHORIZATION") or "").strip()
)
operator_override = bool(request_override and operator_authorized)
inventory = {
"sessions": sessions,
"leases": leases,
"terminal_lock": terminal_lock,
"inventory_complete": inventory_complete,
"incomplete_reasons": incomplete_reasons,
}
report = restart_coordinator.evaluate_restart_impact(
inventory,
operator_override=operator_override,
requesting_session_id=sid,
dry_run=True, # coordinator is always analysis-only (#658)
)
payload = report.as_dict()
payload["success"] = True
payload["read_only"] = True
payload["remote"] = remote
payload["org"] = o
payload["repo"] = r
payload["requesting_session_id"] = sid
payload["operator_override_requested"] = bool(request_override)
payload["operator_override_authorized"] = operator_authorized
payload["apply_supported"] = False
if not dry_run:
payload["reasons"] = list(payload.get("reasons") or []) + [
"apply requested but not supported: sanctioned restart apply is "
"gated by a drain proof (separate child); no restart performed (#658)"
]
return payload
@mcp.tool() @mcp.tool()
def gitea_inspect_workflow_lease( def gitea_inspect_workflow_lease(
lease_id: str, lease_id: str,
+51 -2
View File
@@ -228,25 +228,74 @@ def find_active_reviewer_lease(
return None return None
def _conflict_fix_chain_key(lease: dict) -> tuple | None:
"""Identity of the lease chain a conflict-fix marker belongs to (#842).
Keyed by PR number, profile, head_before, and branch. Returns None when any
required component (pr_number, profile, head_before) is missing or malformed.
"""
raw = lease.get("raw_fields") or {}
pr_number = lease.get("pr_number")
profile = (lease.get("profile") or "").strip().lower()
head_before = lease.get("head_before")
branch = (lease.get("branch") or raw.get("branch") or "").strip()
if not (pr_number and profile and head_before):
return None
return (pr_number, profile, head_before, branch)
def _conflict_fix_chain_matches(key1: tuple, key2: tuple) -> bool:
"""True when two conflict-fix chain keys refer to the same lease chain."""
pr1, profile1, head1, branch1 = key1
pr2, profile2, head2, branch2 = key2
if pr1 != pr2 or profile1 != profile2 or head1 != head2:
return False
if branch1 and branch2 and branch1 != branch2:
return False
return True
def _conflict_fix_chain_terminated_after(entries: list[dict], index: int) -> bool:
"""True when a later marker terminates the conflict-fix chain of ``entries[index]``.
Append-only newest-wins: a terminal marker (phase=released/blocked/done)
ends only its matching claim chain (#842).
"""
key = _conflict_fix_chain_key(entries[index])
if key is None:
return False
for later in entries[index + 1:]:
phase = (later.get("phase") or "").strip().lower()
if phase not in _TERMINAL_CONFLICT_FIX_PHASES:
continue
later_key = _conflict_fix_chain_key(later)
if later_key and _conflict_fix_chain_matches(key, later_key):
return True
return False
def find_active_conflict_fix_lease( def find_active_conflict_fix_lease(
comments: list[dict], comments: list[dict],
*, *,
pr_number: int, pr_number: int,
now: datetime | None = None, now: datetime | None = None,
) -> dict[str, Any] | None: ) -> dict[str, Any] | None:
"""Return the newest unexpired conflict-fix lease for *pr_number*, if any.""" """Return the newest unexpired, non-terminated conflict-fix lease for *pr_number*, if any."""
now = now or datetime.now(timezone.utc) now = now or datetime.now(timezone.utc)
candidates = [ candidates = [
entry for entry in _comment_entries(comments, pr_number=pr_number) entry for entry in _comment_entries(comments, pr_number=pr_number)
if entry.get("lease_kind") == "conflict_fix" if entry.get("lease_kind") == "conflict_fix"
] ]
for lease in reversed(candidates): for index in range(len(candidates) - 1, -1, -1):
lease = candidates[index]
if _lease_expired(lease, now=now): if _lease_expired(lease, now=now):
continue continue
phase = (lease.get("phase") or "").strip().lower() phase = (lease.get("phase") or "").strip().lower()
if phase in _TERMINAL_CONFLICT_FIX_PHASES: if phase in _TERMINAL_CONFLICT_FIX_PHASES:
continue continue
if phase in _ACTIVE_CONFLICT_FIX_PHASES or phase: if phase in _ACTIVE_CONFLICT_FIX_PHASES or phase:
if _conflict_fix_chain_terminated_after(candidates, index):
continue
return lease return lease
return None return None
+451
View File
@@ -0,0 +1,451 @@
"""MCP restart coordinator and impact analysis (#658).
Before any sanctioned MCP restart, a central coordinator must evaluate the
live control-plane state — active sessions, leases/locks, in-flight issue/PR
work, mutations, worktrees, and recovery history — and produce an *impact
preview* so operators (and the web console, #642/#652) can see the blast
radius **before** concurrent LLM work is disrupted.
Design rules (mirrors the read-only posture of ``workflow_dashboard`` /
``lease_lifecycle``):
* **Pure classification.** :func:`evaluate_restart_impact` takes an already
gathered inventory and returns a structured report. It never touches the
network, the filesystem, or a live process, so multi-session fixtures can
drive every branch in unit tests. The coordinator *never restarts anything*;
a mutative apply path is a later child gated by a drain proof (non-goal here).
* **Fail closed.** If the inventory is not explicitly complete, the verdict is
``unsafe`` / deny — an incomplete evaluation must never green-light a restart.
* **No secrets.** Session ids, pids, and profiles are operational metadata, not
credentials; nothing secret flows through this module.
The single sanctioned entry point post-#657 is the MCP tool
``gitea_request_mcp_restart`` (dry-run by default), which gathers the inventory
from the #613 control-plane DB and calls :func:`evaluate_restart_impact`.
"""
from __future__ import annotations
from dataclasses import dataclass, field
from datetime import datetime, timezone
from typing import Any, Mapping, Sequence
import lease_lifecycle
COORDINATOR_VERSION = "1.0.0-issue-658"
# Restart verdicts. Exactly the three the acceptance criteria name.
VERDICT_SAFE = "safe"
VERDICT_UNSAFE = "unsafe"
VERDICT_OVERRIDE = "override"
# Blast-radius severity bands.
BLAST_NONE = "none"
BLAST_LOW = "low"
BLAST_MEDIUM = "medium"
BLAST_HIGH = "high"
# A live lease with a live owner process is treated as active in-flight work.
LEASE_FRESHNESS_LIVE = "active"
# Default staleness window for a session heartbeat (seconds). A session whose
# last heartbeat is older than this is not counted as live even if its row is
# still marked ``active`` — it is assumed dead/detached.
DEFAULT_SESSION_HEARTBEAT_STALE_SECONDS = 900
def _utc_now() -> datetime:
return datetime.now(timezone.utc)
def _parse_ts(value: str | None) -> datetime | None:
return lease_lifecycle._parse_ts(value)
@dataclass(frozen=True)
class SessionImpact:
"""One MCP session a restart would terminate."""
session_id: str
role: str | None
profile: str | None
pid: int | None
status: str | None
alive: bool | None
heartbeat_stale: bool
is_requester: bool
live: bool
def as_dict(self) -> dict[str, Any]:
return {
"session_id": self.session_id,
"role": self.role,
"profile": self.profile,
"pid": self.pid,
"status": self.status,
"alive": self.alive,
"heartbeat_stale": self.heartbeat_stale,
"is_requester": self.is_requester,
"live": self.live,
}
@dataclass(frozen=True)
class LeaseImpact:
"""One control-plane lease a restart would disrupt."""
lease_id: str | None
session_id: str | None
role: str | None
phase: str | None
freshness: str | None
work_kind: str | None
work_number: int | None
worktree_path: str | None
disruptive: bool
is_mutation: bool
is_critical_section: bool
def as_dict(self) -> dict[str, Any]:
return {
"lease_id": self.lease_id,
"session_id": self.session_id,
"role": self.role,
"phase": self.phase,
"freshness": self.freshness,
"work_kind": self.work_kind,
"work_number": self.work_number,
"worktree_path": self.worktree_path,
"disruptive": self.disruptive,
"is_mutation": self.is_mutation,
"is_critical_section": self.is_critical_section,
}
@dataclass(frozen=True)
class RestartImpactReport:
"""Impact preview DTO returned to the console / operator (#642/#652)."""
coordinator_version: str
evaluated_at: str
dry_run: bool
restart_performed: bool
inventory_complete: bool
verdict: str
allow_restart: bool
override_would_allow: bool
operator_override: bool
blast_radius: str
reasons: list[str]
affected_sessions: list[SessionImpact]
affected_leases: list[LeaseImpact]
critical_sections: list[LeaseImpact]
affected_issues: list[int]
affected_prs: list[int]
mutations: list[LeaseImpact]
terminal_lock: dict[str, Any] | None
ack_state: dict[str, str]
prior_recovery_attempts: list[dict[str, Any]]
counts: dict[str, int]
audit_record: dict[str, Any]
incomplete_reasons: list[str] = field(default_factory=list)
def as_dict(self) -> dict[str, Any]:
return {
"coordinator_version": self.coordinator_version,
"evaluated_at": self.evaluated_at,
"dry_run": self.dry_run,
"restart_performed": self.restart_performed,
"inventory_complete": self.inventory_complete,
"incomplete_reasons": list(self.incomplete_reasons),
"verdict": self.verdict,
"allow_restart": self.allow_restart,
"override_would_allow": self.override_would_allow,
"operator_override": self.operator_override,
"blast_radius": self.blast_radius,
"reasons": list(self.reasons),
"affected_sessions": [s.as_dict() for s in self.affected_sessions],
"affected_leases": [l.as_dict() for l in self.affected_leases],
"critical_sections": [l.as_dict() for l in self.critical_sections],
"affected_issues": list(self.affected_issues),
"affected_prs": list(self.affected_prs),
"mutations": [l.as_dict() for l in self.mutations],
"terminal_lock": self.terminal_lock,
"ack_state": dict(self.ack_state),
"prior_recovery_attempts": list(self.prior_recovery_attempts),
"counts": dict(self.counts),
"audit_record": dict(self.audit_record),
}
def _classify_session(
row: Mapping[str, Any],
*,
now: datetime,
requesting_session_id: str | None,
heartbeat_stale_seconds: int,
) -> SessionImpact:
session_id = str(row.get("session_id") or "")
pid = row.get("pid")
status = (row.get("status") or "").strip().lower() or None
alive = lease_lifecycle.is_process_alive(pid) if pid is not None else None
hb = _parse_ts(row.get("last_heartbeat_at"))
heartbeat_stale = bool(
hb is not None and (now - hb).total_seconds() > heartbeat_stale_seconds
)
live = bool(status == "active" and alive is not False and not heartbeat_stale)
return SessionImpact(
session_id=session_id,
role=row.get("role"),
profile=row.get("profile"),
pid=pid,
status=status,
alive=alive,
heartbeat_stale=heartbeat_stale,
is_requester=bool(
requesting_session_id and session_id == requesting_session_id
),
live=live,
)
# Lease phases that represent an active mutation in flight (as opposed to a
# mere allocation/claim with no work committed yet). An active lease in any of
# these phases is a critical section a restart must not sever.
_MUTATING_PHASES = frozenset(
{
"implementing",
"publishing",
"pushing",
"committing",
"reviewing",
"merging",
"reconciling",
"conflict_fix",
}
)
def _classify_lease(row: Mapping[str, Any]) -> LeaseImpact:
freshness_obj = row.get("freshness")
if isinstance(freshness_obj, Mapping):
freshness = str(freshness_obj.get("freshness") or "").strip().lower() or None
else:
freshness = str(freshness_obj or "").strip().lower() or None
phase = (row.get("phase") or "").strip().lower() or None
worktree = row.get("worktree_path")
disruptive = freshness == LEASE_FRESHNESS_LIVE
# A live lease is a mutation-in-flight if it carries an author worktree or
# its phase names a mutating step. All disruptive leases are critical
# sections a restart would sever regardless.
is_mutation = bool(
disruptive and (bool(worktree) or (phase in _MUTATING_PHASES))
)
number = row.get("work_number")
try:
number = int(number) if number is not None else None
except (TypeError, ValueError):
number = None
return LeaseImpact(
lease_id=row.get("lease_id"),
session_id=row.get("session_id"),
role=row.get("role"),
phase=phase,
freshness=freshness,
work_kind=(str(row.get("work_kind") or "").strip().lower() or None),
work_number=number,
worktree_path=worktree,
disruptive=disruptive,
is_mutation=is_mutation,
is_critical_section=disruptive,
)
def _blast_radius(*, session_count: int, work_count: int, mutation_count: int) -> str:
if mutation_count > 0 or work_count >= 3 or session_count >= 3:
return BLAST_HIGH
if work_count > 0 or session_count == 2:
return BLAST_MEDIUM
if session_count == 1:
return BLAST_LOW
return BLAST_NONE
def evaluate_restart_impact(
inventory: Mapping[str, Any],
*,
now: datetime | None = None,
operator_override: bool = False,
requesting_session_id: str | None = None,
dry_run: bool = True,
session_heartbeat_stale_seconds: int = DEFAULT_SESSION_HEARTBEAT_STALE_SECONDS,
) -> RestartImpactReport:
"""Evaluate a proposed MCP restart and return an impact preview.
``inventory`` is a mapping with:
* ``sessions`` — session rows (session_id, role, profile, pid, status,
last_heartbeat_at).
* ``leases`` — control-plane lease rows, each ideally carrying an enriched
``freshness`` dict (as :func:`lease_lifecycle.list_active_leases` returns);
a bare string freshness is also accepted.
* ``terminal_lock`` — the active terminal (merge) lock, if any.
* ``prior_recovery_attempts`` — narrower recovery attempts already tried
(e.g. sanctioned reconnects) so the operator sees escalation history.
* ``inventory_complete`` — bool. **Must** be explicitly True; a missing or
falsy value forces a deny (fail closed).
* ``incomplete_reasons`` — optional reasons the inventory is incomplete.
The coordinator never restarts anything: ``restart_performed`` is always
False and the mutative apply path is a later drain-gated child.
"""
moment = now or _utc_now()
reasons: list[str] = []
inventory_complete = bool(inventory.get("inventory_complete", False))
incomplete_reasons = [str(r) for r in (inventory.get("incomplete_reasons") or [])]
sessions_raw: Sequence[Mapping[str, Any]] = inventory.get("sessions") or []
leases_raw: Sequence[Mapping[str, Any]] = inventory.get("leases") or []
terminal_lock = inventory.get("terminal_lock") or None
prior_recovery_attempts = [
dict(a) for a in (inventory.get("prior_recovery_attempts") or [])
]
session_impacts = [
_classify_session(
s,
now=moment,
requesting_session_id=requesting_session_id,
heartbeat_stale_seconds=session_heartbeat_stale_seconds,
)
for s in sessions_raw
]
lease_impacts = [_classify_lease(l) for l in leases_raw]
# Only *other* live sessions and live leases constitute blast radius: a
# restart that would kill only the requesting session with no other work in
# flight is safe.
other_live_sessions = [
s for s in session_impacts if s.live and not s.is_requester
]
disruptive_leases = [l for l in lease_impacts if l.disruptive]
critical_sections = [l for l in lease_impacts if l.is_critical_section]
mutations = [l for l in lease_impacts if l.is_mutation]
affected_issues = sorted(
{
l.work_number
for l in disruptive_leases
if l.work_kind == "issue" and l.work_number is not None
}
)
affected_prs = sorted(
{
l.work_number
for l in disruptive_leases
if l.work_kind == "pr" and l.work_number is not None
}
)
disruptive = bool(disruptive_leases or other_live_sessions or terminal_lock)
if not inventory_complete:
verdict = VERDICT_UNSAFE
allow_restart = False
reasons.append(
"inventory incomplete: restart evaluation cannot confirm blast "
"radius — deny (fail closed, #658)"
)
reasons.extend(incomplete_reasons)
elif not disruptive:
verdict = VERDICT_SAFE
allow_restart = True
reasons.append("no other live sessions, live leases, or terminal lock")
elif operator_override:
verdict = VERDICT_OVERRIDE
allow_restart = True
reasons.append(
"live work present; operator override accepts the blast radius"
)
else:
verdict = VERDICT_UNSAFE
allow_restart = False
reasons.append(
"live work would be disrupted; restart denied without operator "
"override"
)
if critical_sections and inventory_complete:
reasons.append(
f"{len(critical_sections)} critical section(s) in flight "
"(active lease with a live owner)"
)
if terminal_lock:
reasons.append("active terminal (merge) lock present")
override_would_allow = bool(inventory_complete and disruptive)
blast_radius = _blast_radius(
session_count=len(other_live_sessions),
work_count=len(affected_issues) + len(affected_prs),
mutation_count=len(mutations),
)
# Acknowledgement is a later child (drain protocol); expose per-session
# placeholders so the console can render the ack column now.
ack_state = {s.session_id: "pending" for s in other_live_sessions}
counts = {
"sessions_total": len(session_impacts),
"sessions_live_other": len(other_live_sessions),
"leases_total": len(lease_impacts),
"leases_disruptive": len(disruptive_leases),
"critical_sections": len(critical_sections),
"mutations": len(mutations),
"affected_issues": len(affected_issues),
"affected_prs": len(affected_prs),
"prior_recovery_attempts": len(prior_recovery_attempts),
}
audit_record = {
"event": "restart_impact_evaluated",
"coordinator_version": COORDINATOR_VERSION,
"evaluated_at": moment.isoformat(),
"dry_run": dry_run,
"operator_override": bool(operator_override),
"requesting_session_id": requesting_session_id,
"inventory_complete": inventory_complete,
"verdict": verdict,
"allow_restart": allow_restart,
"blast_radius": blast_radius,
"counts": counts,
}
return RestartImpactReport(
coordinator_version=COORDINATOR_VERSION,
evaluated_at=moment.isoformat(),
dry_run=dry_run,
restart_performed=False,
inventory_complete=inventory_complete,
verdict=verdict,
allow_restart=allow_restart,
override_would_allow=override_would_allow,
operator_override=bool(operator_override),
blast_radius=blast_radius,
reasons=reasons,
affected_sessions=session_impacts,
affected_leases=lease_impacts,
critical_sections=critical_sections,
affected_issues=affected_issues,
affected_prs=affected_prs,
mutations=mutations,
terminal_lock=dict(terminal_lock)
if isinstance(terminal_lock, Mapping)
else terminal_lock,
ack_state=ack_state,
prior_recovery_attempts=prior_recovery_attempts,
counts=counts,
audit_record=audit_record,
incomplete_reasons=incomplete_reasons,
)
+153
View File
@@ -19,6 +19,7 @@ from pr_work_lease import ( # noqa: E402
assess_reviewer_mutation_blocked, assess_reviewer_mutation_blocked,
assess_reviewer_stale_head_final_report, assess_reviewer_stale_head_final_report,
format_conflict_fix_lease_body, format_conflict_fix_lease_body,
find_active_conflict_fix_lease,
parse_conflict_fix_lease_comment, parse_conflict_fix_lease_comment,
parse_reviewer_lease_comment, parse_reviewer_lease_comment,
) )
@@ -203,5 +204,157 @@ class TestFormatLease(unittest.TestCase):
self.assertEqual(parsed["pr_number"], 376) self.assertEqual(parsed["pr_number"], 376)
class TestConflictFixLeaseLifecycle(unittest.TestCase):
def test_claim_followed_by_matching_release(self):
claim_body = _conflict_fix_body(phase="claimed", worktree="branches/fix-376")
expires = (NOW + timedelta(minutes=60)).isoformat().replace("+00:00", "Z")
release_body = "\n".join([
CONFLICT_FIX_LEASE_MARKER,
"pr: #376",
"branch: feat/fix-376",
"worktree: branches/fix-376",
"profile: prgs-author",
"phase: released",
f"head_before: {HEAD_A}",
f"head_after: {HEAD_B}",
f"expires_at: {expires}",
])
comments = [{"body": claim_body}, {"body": release_body}]
lease = find_active_conflict_fix_lease(comments, pr_number=376, now=NOW)
self.assertIsNone(lease)
def test_expired_claim_without_release(self):
past_expires = (NOW - timedelta(minutes=10)).isoformat().replace("+00:00", "Z")
claim_body = "\n".join([
CONFLICT_FIX_LEASE_MARKER,
"pr: #376",
"phase: claimed",
f"head_before: {HEAD_A}",
f"expires_at: {past_expires}",
"profile: prgs-author",
])
comments = [{"body": claim_body}]
lease = find_active_conflict_fix_lease(comments, pr_number=376, now=NOW)
self.assertIsNone(lease)
def test_mismatched_release_different_head(self):
claim_body = _conflict_fix_body(phase="claimed", worktree="branches/fix-376")
expires = (NOW + timedelta(minutes=60)).isoformat().replace("+00:00", "Z")
release_body = "\n".join([
CONFLICT_FIX_LEASE_MARKER,
"pr: #376",
"profile: prgs-author",
"phase: released",
f"head_before: {HEAD_B}",
f"expires_at: {expires}",
])
comments = [{"body": claim_body}, {"body": release_body}]
lease = find_active_conflict_fix_lease(comments, pr_number=376, now=NOW)
self.assertIsNotNone(lease)
self.assertEqual(lease["phase"], "claimed")
def test_mismatched_release_different_branch(self):
claim_body = "\n".join([
CONFLICT_FIX_LEASE_MARKER,
"pr: #376",
"branch: feat/branch-A",
"phase: claimed",
f"head_before: {HEAD_A}",
f"expires_at: {(NOW + timedelta(minutes=60)).isoformat().replace('+00:00', 'Z')}",
"profile: prgs-author",
])
release_body = "\n".join([
CONFLICT_FIX_LEASE_MARKER,
"pr: #376",
"branch: feat/branch-B",
"phase: released",
f"head_before: {HEAD_A}",
f"expires_at: {(NOW + timedelta(minutes=60)).isoformat().replace('+00:00', 'Z')}",
"profile: prgs-author",
])
comments = [{"body": claim_body}, {"body": release_body}]
lease = find_active_conflict_fix_lease(comments, pr_number=376, now=NOW)
self.assertIsNotNone(lease)
self.assertEqual(lease["phase"], "claimed")
def test_release_followed_by_newer_claim(self):
claim_1 = _conflict_fix_body(phase="claimed", worktree="branches/fix-376")
expires = (NOW + timedelta(minutes=60)).isoformat().replace("+00:00", "Z")
release_1 = "\n".join([
CONFLICT_FIX_LEASE_MARKER,
"pr: #376",
"profile: prgs-author",
"phase: released",
f"head_before: {HEAD_A}",
f"head_after: {HEAD_B}",
f"expires_at: {expires}",
])
claim_2 = "\n".join([
CONFLICT_FIX_LEASE_MARKER,
"pr: #376",
"profile: prgs-author",
"phase: claimed",
f"head_before: {HEAD_B}",
f"expires_at: {expires}",
])
comments = [{"body": claim_1}, {"body": release_1}, {"body": claim_2}]
lease = find_active_conflict_fix_lease(comments, pr_number=376, now=NOW)
self.assertIsNotNone(lease)
self.assertEqual(lease["head_before"], HEAD_B)
def test_malformed_or_ambiguous_markers(self):
malformed_release = "\n".join([
CONFLICT_FIX_LEASE_MARKER,
"pr: #376",
"phase: released",
# missing head_before and profile
])
claim_body = _conflict_fix_body(phase="claimed")
comments = [{"body": claim_body}, {"body": malformed_release}]
lease = find_active_conflict_fix_lease(comments, pr_number=376, now=NOW)
self.assertIsNotNone(lease)
def test_pr818_historical_sequence(self):
comment_14696 = "\n".join([
"<!-- mcp-conflict-fix-lease:v1 -->",
"pr: #818",
"branch: feat/issue-638-webui-app-shell-phase1",
"worktree: /Users/jasonwalker/Development/Gitea-Tools/branches/issue-638-webui-app-shell-phase1",
"profile: prgs-author",
"session_id: unknown",
"phase: claimed",
"head_before: 08061b7b8aebdd099a37d1abf5dafcf38e4fd3fb",
"expires_at: 2026-07-23T07:12:13Z",
"reviewer_active: no",
])
comment_14730 = "\n".join([
"<!-- mcp-conflict-fix-lease:v1 -->",
"pr: #818",
"branch: feat/issue-638-webui-app-shell-phase1",
"worktree: /Users/jasonwalker/Development/Gitea-Tools/branches/issue-638-webui-app-shell-phase1",
"profile: prgs-author",
"session_id: prgs-author-61241-e5129c60",
"phase: released",
"head_before: 08061b7b8aebdd099a37d1abf5dafcf38e4fd3fb",
"head_after: 64b6eb5d5402663098de5ded3b0617cc3b3df98f",
"expires_at: 2026-07-23T06:05:00Z",
"reviewer_active: no",
])
comments = [{"body": comment_14696}, {"body": comment_14730}]
check_now = datetime(2026, 7, 23, 6, 30, tzinfo=timezone.utc)
lease = find_active_conflict_fix_lease(comments, pr_number=818, now=check_now)
self.assertIsNone(lease)
reviewer_gate = assess_reviewer_mutation_blocked(
pr_number=818,
comments=comments,
reviewed_head_sha="64b6eb5d5402663098de5ded3b0617cc3b3df98f",
live_head_sha="64b6eb5d5402663098de5ded3b0617cc3b3df98f",
mutation="approve",
now=check_now,
)
self.assertTrue(reviewer_gate["mutation_allowed"])
if __name__ == "__main__": if __name__ == "__main__":
unittest.main() unittest.main()
+340
View File
@@ -0,0 +1,340 @@
"""Tests for the MCP restart coordinator and impact analysis (#658).
Multi-session fixtures exercise every verdict branch: safe, unsafe (live work),
override, and the fail-closed deny on incomplete inventory. Also covers the
critical-section deny path and the new ``ControlPlaneDB.list_sessions``.
"""
from __future__ import annotations
import os
import tempfile
import unittest
from datetime import datetime, timedelta, timezone
import restart_coordinator as rc
from control_plane_db import ControlPlaneDB
NOW = datetime(2026, 7, 24, 6, 0, 0, tzinfo=timezone.utc)
def _ts(dt: datetime) -> str:
return dt.isoformat()
def _live_pid() -> int:
return os.getpid()
def _dead_pid() -> int:
# A pid that is essentially never alive. os.kill(0) on it raises
# ProcessLookupError → is_process_alive False.
return 2_000_000_000
def _session(session_id, *, pid, status="active", heartbeat=None, role="author"):
return {
"session_id": session_id,
"role": role,
"profile": "prgs-author",
"pid": pid,
"status": status,
"last_heartbeat_at": _ts(heartbeat or NOW),
}
def _lease(
lease_id,
*,
session_id,
freshness,
kind="issue",
number=658,
phase="allocated",
worktree=None,
role="author",
):
return {
"lease_id": lease_id,
"session_id": session_id,
"role": role,
"phase": phase,
"work_kind": kind,
"work_number": number,
"worktree_path": worktree,
"freshness": {"freshness": freshness},
}
class EvaluateRestartImpactTest(unittest.TestCase):
def test_incomplete_inventory_denies_fail_closed(self) -> None:
report = rc.evaluate_restart_impact(
{"inventory_complete": False, "incomplete_reasons": ["db down"]},
now=NOW,
)
self.assertEqual(report.verdict, rc.VERDICT_UNSAFE)
self.assertFalse(report.allow_restart)
self.assertFalse(report.restart_performed)
self.assertIn("db down", report.incomplete_reasons)
self.assertTrue(
any("fail closed" in reasoning for reasoning in report.reasons)
)
def test_missing_completeness_flag_denies(self) -> None:
# No inventory_complete key at all → treated as incomplete.
report = rc.evaluate_restart_impact({}, now=NOW)
self.assertEqual(report.verdict, rc.VERDICT_UNSAFE)
self.assertFalse(report.allow_restart)
def test_no_other_work_is_safe(self) -> None:
report = rc.evaluate_restart_impact(
{
"inventory_complete": True,
"sessions": [_session("requester", pid=_live_pid())],
"leases": [],
},
now=NOW,
requesting_session_id="requester",
)
self.assertEqual(report.verdict, rc.VERDICT_SAFE)
self.assertTrue(report.allow_restart)
self.assertEqual(report.blast_radius, rc.BLAST_NONE)
self.assertEqual(report.affected_issues, [])
def test_dead_foreign_session_and_lease_are_not_disruptive(self) -> None:
report = rc.evaluate_restart_impact(
{
"inventory_complete": True,
"sessions": [
_session("requester", pid=_live_pid()),
_session("dead", pid=_dead_pid()),
],
"leases": [
_lease("l-dead", session_id="dead", freshness="stale_dead_process")
],
},
now=NOW,
requesting_session_id="requester",
)
self.assertEqual(report.verdict, rc.VERDICT_SAFE)
self.assertTrue(report.allow_restart)
self.assertEqual(report.counts["leases_disruptive"], 0)
self.assertEqual(report.counts["sessions_live_other"], 0)
def test_live_foreign_lease_denies_without_override(self) -> None:
report = rc.evaluate_restart_impact(
{
"inventory_complete": True,
"sessions": [
_session("requester", pid=_live_pid()),
_session("worker", pid=_live_pid()),
],
"leases": [
_lease(
"l1",
session_id="worker",
freshness="active",
worktree="/tmp/wt-658",
phase="implementing",
)
],
},
now=NOW,
requesting_session_id="requester",
)
self.assertEqual(report.verdict, rc.VERDICT_UNSAFE)
self.assertFalse(report.allow_restart)
# Critical section detected: active lease with a live owner.
self.assertEqual(len(report.critical_sections), 1)
self.assertEqual(report.affected_issues, [658])
self.assertEqual(report.counts["mutations"], 1)
self.assertTrue(report.override_would_allow)
self.assertEqual(report.blast_radius, rc.BLAST_HIGH)
# Placeholder ack state for the affected session.
self.assertEqual(report.ack_state.get("worker"), "pending")
def test_operator_override_allows_despite_live_work(self) -> None:
inv = {
"inventory_complete": True,
"sessions": [
_session("requester", pid=_live_pid()),
_session("worker", pid=_live_pid()),
],
"leases": [_lease("l1", session_id="worker", freshness="active")],
}
report = rc.evaluate_restart_impact(
inv,
now=NOW,
requesting_session_id="requester",
operator_override=True,
)
self.assertEqual(report.verdict, rc.VERDICT_OVERRIDE)
self.assertTrue(report.allow_restart)
self.assertFalse(report.restart_performed)
def test_deny_when_critical_section_open(self) -> None:
# A single live author lease in a mutating phase is a critical section
# that must deny an un-overridden restart.
report = rc.evaluate_restart_impact(
{
"inventory_complete": True,
"sessions": [_session("worker", pid=_live_pid())],
"leases": [
_lease(
"l1",
session_id="worker",
freshness="active",
phase="merging",
kind="pr",
number=900,
)
],
},
now=NOW,
requesting_session_id="requester",
)
self.assertEqual(report.verdict, rc.VERDICT_UNSAFE)
self.assertFalse(report.allow_restart)
self.assertEqual(report.affected_prs, [900])
self.assertEqual(len(report.critical_sections), 1)
def test_terminal_lock_makes_restart_unsafe(self) -> None:
report = rc.evaluate_restart_impact(
{
"inventory_complete": True,
"sessions": [_session("requester", pid=_live_pid())],
"leases": [],
"terminal_lock": {"terminal_pr": 812},
},
now=NOW,
requesting_session_id="requester",
)
self.assertEqual(report.verdict, rc.VERDICT_UNSAFE)
self.assertFalse(report.allow_restart)
self.assertIsNotNone(report.terminal_lock)
self.assertTrue(
any("terminal" in reasoning for reasoning in report.reasons)
)
def test_other_live_session_without_lease_is_disruptive(self) -> None:
report = rc.evaluate_restart_impact(
{
"inventory_complete": True,
"sessions": [
_session("requester", pid=_live_pid()),
_session("idle-but-live", pid=_live_pid()),
],
"leases": [],
},
now=NOW,
requesting_session_id="requester",
)
self.assertEqual(report.verdict, rc.VERDICT_UNSAFE)
self.assertEqual(report.counts["sessions_live_other"], 1)
def test_stale_heartbeat_session_not_counted_live(self) -> None:
stale = NOW - timedelta(hours=2)
report = rc.evaluate_restart_impact(
{
"inventory_complete": True,
"sessions": [
_session("requester", pid=_live_pid()),
_session("stale", pid=_live_pid(), heartbeat=stale),
],
"leases": [],
},
now=NOW,
requesting_session_id="requester",
)
self.assertEqual(report.verdict, rc.VERDICT_SAFE)
self.assertEqual(report.counts["sessions_live_other"], 0)
def test_prior_recovery_attempts_echoed(self) -> None:
report = rc.evaluate_restart_impact(
{
"inventory_complete": True,
"sessions": [_session("requester", pid=_live_pid())],
"leases": [],
"prior_recovery_attempts": [
{"kind": "client_reconnect", "at": _ts(NOW)}
],
},
now=NOW,
requesting_session_id="requester",
)
self.assertEqual(len(report.prior_recovery_attempts), 1)
self.assertEqual(report.counts["prior_recovery_attempts"], 1)
def test_bare_string_freshness_accepted(self) -> None:
lease = _lease("l1", session_id="worker", freshness="active")
lease["freshness"] = "active" # bare string, not a dict
report = rc.evaluate_restart_impact(
{
"inventory_complete": True,
"sessions": [_session("worker", pid=_live_pid())],
"leases": [lease],
},
now=NOW,
requesting_session_id="requester",
)
self.assertEqual(report.counts["leases_disruptive"], 1)
def test_as_dict_is_serializable_dto(self) -> None:
import json
report = rc.evaluate_restart_impact(
{
"inventory_complete": True,
"sessions": [_session("requester", pid=_live_pid())],
"leases": [],
},
now=NOW,
requesting_session_id="requester",
)
payload = report.as_dict()
# Round-trips through JSON — safe for the console DTO.
encoded = json.dumps(payload)
decoded = json.loads(encoded)
self.assertEqual(decoded["verdict"], rc.VERDICT_SAFE)
self.assertIn("audit_record", decoded)
self.assertEqual(decoded["audit_record"]["event"], "restart_impact_evaluated")
self.assertFalse(decoded["restart_performed"])
self.assertIn("coordinator_version", decoded)
class ListSessionsTest(unittest.TestCase):
def setUp(self) -> None:
self._tmp = tempfile.TemporaryDirectory()
self.db = ControlPlaneDB(os.path.join(self._tmp.name, "cp.sqlite3"))
def tearDown(self) -> None:
self._tmp.cleanup()
def test_list_sessions_filters_by_status(self) -> None:
self.db.upsert_session(session_id="a", role="author", pid=1, status="active")
self.db.upsert_session(session_id="b", role="author", pid=2, status="ended")
active = self.db.list_sessions(statuses=("active",))
ids = {row["session_id"] for row in active}
self.assertEqual(ids, {"a"})
every = self.db.list_sessions()
self.assertEqual({row["session_id"] for row in every}, {"a", "b"})
def test_list_sessions_feeds_coordinator(self) -> None:
self.db.upsert_session(
session_id="requester", role="author", pid=os.getpid(), status="active"
)
report = rc.evaluate_restart_impact(
{
"inventory_complete": True,
"sessions": self.db.list_sessions(statuses=("active",)),
"leases": [],
},
now=NOW,
requesting_session_id="requester",
)
self.assertEqual(report.counts["sessions_total"], 1)
if __name__ == "__main__": # pragma: no cover
unittest.main()
+468
View File
@@ -0,0 +1,468 @@
"""Tests for the unified web-console inventory API (#636).
Covers the four cases the issue names — empty, populated, partial failure, and
the no-false-unowned invariant — plus redaction, collision detection, the
resource-split routes, and read-only guarantees against a real control-plane
database and real durable lock files.
"""
import json
import os
import sys
import tempfile
import unittest
from datetime import datetime, timedelta, timezone
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
from starlette.testclient import TestClient
import control_plane_db
from webui.app import create_app
from webui import inventory
def _iso(dt: datetime) -> str:
return dt.astimezone(timezone.utc).isoformat()
def _write_lock(lock_dir: str, name: str, payload: dict) -> str:
path = os.path.join(lock_dir, name)
with open(path, "w", encoding="utf-8") as handle:
json.dump(payload, handle)
return path
def _live_lock_payload(
*,
issue_number: int,
branch: str,
worktree_path: str,
pid: int,
username: str = "jcwalker3",
profile: str = "prgs-author",
) -> dict:
now = datetime.now(timezone.utc)
future = now + timedelta(hours=2)
return {
"branch_name": branch,
"issue_number": issue_number,
"org": "Scaled-Tech-Consulting",
"repo": "Gitea-Tools",
"remote": "prgs",
"pid": pid,
"session_pid": pid,
"lock_generation": 1,
"worktree_path": worktree_path,
"claimant": {"username": username, "profile": profile},
"work_lease": {
"branch": branch,
"issue_number": issue_number,
"operation_type": "author_issue_work",
"created_at": _iso(now),
"expires_at": _iso(future),
"last_heartbeat_at": _iso(now),
"claimant": {"username": username, "profile": profile},
},
}
class _FixtureMixin(unittest.TestCase):
def setUp(self) -> None:
self._tmp = tempfile.TemporaryDirectory()
self.tmp = self._tmp.name
self.lock_dir = os.path.join(self.tmp, "locks")
os.makedirs(self.lock_dir, mode=0o700)
self.db_path = os.path.join(self.tmp, "control_plane.db")
self.addCleanup(self._tmp.cleanup)
def _seed_db(self) -> control_plane_db.ControlPlaneDB:
db = control_plane_db.ControlPlaneDB(self.db_path)
db.upsert_session(
session_id="prgs-author-1",
role="author",
profile="prgs-author",
namespace="gitea-author",
pid=os.getpid(),
)
db.upsert_work_item(
remote="prgs",
org="Scaled-Tech-Consulting",
repo="Gitea-Tools",
kind="issue",
number=636,
)
db.assign_and_lease(
session_id="prgs-author-1",
role="author",
remote="prgs",
org="Scaled-Tech-Consulting",
repo="Gitea-Tools",
kind="issue",
number=636,
)
return db
class TestRedaction(unittest.TestCase):
def test_redact_path_collapses_home(self):
home = os.path.expanduser("~")
self.assertEqual(
inventory.redact_path(f"{home}/Development/Gitea-Tools"),
"~/Development/Gitea-Tools",
)
def test_redact_url_strips_userinfo_and_query(self):
self.assertEqual(
inventory.redact_url("https://user:[email protected]/api?token=abc"),
"https://gitea.prgs.cc/api",
)
def test_scrub_drops_credential_keys(self):
scrubbed = inventory.scrub(
{"token": "abc123", "api_key": "k", "profile": "prgs-author"}
)
self.assertEqual(scrubbed["token"], "[redacted]")
self.assertEqual(scrubbed["api_key"], "[redacted]")
self.assertEqual(scrubbed["profile"], "prgs-author")
def test_scrub_is_recursive_and_never_raises(self):
class Weird:
def __repr__(self) -> str:
return "weird-obj"
out = inventory.scrub({"nested": [{"password": "p", "obj": Weird()}]})
self.assertEqual(out["nested"][0]["password"], "[redacted]")
self.assertEqual(out["nested"][0]["obj"], "weird-obj")
class TestEmptyInventory(_FixtureMixin):
def test_empty_db_and_locks_degrade_without_raising(self):
# No DB file, no locks: sessions/leases unavailable, locks ok+empty.
snap = inventory.load_inventory_snapshot(
db_path=self.db_path,
lock_dir=self.lock_dir,
load_hygiene=lambda: _StubHygiene(entries=()),
)
sessions = snap.section("sessions")
leases = snap.section("leases")
locks = snap.section("locks")
self.assertEqual(sessions.status, inventory.STATUS_UNAVAILABLE)
self.assertEqual(leases.status, inventory.STATUS_UNAVAILABLE)
self.assertEqual(locks.status, inventory.STATUS_OK)
self.assertEqual(len(locks.items), 0)
# Ownership authority is incomplete because the DB is missing.
self.assertFalse(snap.ownership_authority_complete)
self.assertEqual(snap.collisions, ())
def test_empty_db_present_but_unpopulated(self):
control_plane_db.ControlPlaneDB(self.db_path) # creates schema, no rows
snap = inventory.load_inventory_snapshot(
db_path=self.db_path,
lock_dir=self.lock_dir,
load_hygiene=lambda: _StubHygiene(entries=()),
)
self.assertEqual(snap.section("sessions").status, inventory.STATUS_OK)
self.assertEqual(len(snap.section("sessions").items), 0)
self.assertEqual(snap.section("leases").status, inventory.STATUS_OK)
self.assertTrue(snap.ownership_authority_complete)
class TestPopulatedInventory(_FixtureMixin):
def test_sections_populated_and_correlated(self):
self._seed_db()
wt = f"{self.tmp}/branches/issue-636-inventory-api"
_write_lock(
self.lock_dir,
"prgs-Scaled-Tech-Consulting-Gitea-Tools-636.json",
_live_lock_payload(
issue_number=636,
branch="feat/issue-636-inventory-api",
worktree_path=wt,
pid=os.getpid(),
),
)
hygiene = _StubHygiene(
entries=(
_StubEntry(
rel_path="branches/issue-636-inventory-api",
branch="feat/issue-636-inventory-api",
classification="active-issue",
),
)
)
snap = inventory.load_inventory_snapshot(
db_path=self.db_path,
lock_dir=self.lock_dir,
load_hygiene=lambda: hygiene,
)
self.assertTrue(snap.ownership_authority_complete)
self.assertEqual(len(snap.section("sessions").items), 1)
self.assertEqual(len(snap.section("leases").items), 1)
self.assertEqual(len(snap.section("locks").items), 1)
self.assertEqual(len(snap.section("worktrees").items), 1)
# The lease, lock, and worktree for #636 correlate onto one row.
row = next(r for r in snap.correlations if r["issue_number"] == 636)
self.assertEqual(row["branch"], "feat/issue-636-inventory-api")
self.assertTrue(row["lock_live"])
self.assertEqual(row["worktree_classification"], "active-issue")
self.assertEqual(len(row["lease_ids"]), 1)
# No collision: live lock, live pid, matching worktree.
self.assertEqual(snap.collisions, ())
def test_serialized_payload_declares_field_authority(self):
self._seed_db()
snap = inventory.load_inventory_snapshot(
db_path=self.db_path,
lock_dir=self.lock_dir,
load_hygiene=lambda: _StubHygiene(entries=()),
)
payload = inventory.snapshot_to_dict(snap)
self.assertEqual(payload["api_version"], "v1")
self.assertEqual(payload["schema_version"], 1)
self.assertEqual(payload["field_authority"]["sessions"], "control_plane_db")
self.assertEqual(payload["field_authority"]["locks"], "filesystem")
self.assertIn("sessions", payload["sections"])
class TestPartialFailure(_FixtureMixin):
def test_worktree_scan_failure_degrades_only_that_section(self):
self._seed_db()
def _boom():
raise RuntimeError("git worktree list exploded")
snap = inventory.load_inventory_snapshot(
db_path=self.db_path,
lock_dir=self.lock_dir,
load_hygiene=_boom,
)
self.assertEqual(
snap.section("worktrees").status, inventory.STATUS_UNAVAILABLE
)
self.assertIn("exploded", snap.section("worktrees").reason)
# DB-backed sections still healthy.
self.assertEqual(snap.section("sessions").status, inventory.STATUS_OK)
self.assertIn("worktrees", snap.degraded_sections)
def test_degraded_ownership_suppresses_unowned_claim(self):
# DB absent → sessions/leases unavailable → ownership incomplete even
# though a lock exists and could look "unclaimed" by the DB alone.
_write_lock(
self.lock_dir,
"prgs-Scaled-Tech-Consulting-Gitea-Tools-636.json",
_live_lock_payload(
issue_number=636,
branch="feat/issue-636-inventory-api",
worktree_path=f"{self.tmp}/wt",
pid=os.getpid(),
),
)
snap = inventory.load_inventory_snapshot(
db_path=self.db_path,
lock_dir=self.lock_dir,
load_hygiene=lambda: _StubHygiene(entries=()),
)
self.assertFalse(snap.ownership_authority_complete)
payload = inventory.snapshot_to_dict(snap)
self.assertIn("may be treated as unowned", payload["ownership_note"])
class TestCollisionDetection(_FixtureMixin):
def test_live_lock_dead_owner_flagged(self):
_write_lock(
self.lock_dir,
"prgs-Scaled-Tech-Consulting-Gitea-Tools-700.json",
_live_lock_payload(
issue_number=700,
branch="feat/issue-700-x",
worktree_path=f"{self.tmp}/wt700",
pid=999_999_999, # not a running pid
),
)
control_plane_db.ControlPlaneDB(self.db_path) # empty but present
snap = inventory.load_inventory_snapshot(
db_path=self.db_path,
lock_dir=self.lock_dir,
load_hygiene=lambda: _StubHygiene(entries=()),
)
kinds = {c.kind for c in snap.collisions}
self.assertIn("live-lock-dead-owner", kinds)
# Also lock-without-worktree, since no worktree carries the branch.
self.assertIn("lock-without-worktree", kinds)
def test_duplicate_live_lock_on_same_branch(self):
for issue in (800, 801):
_write_lock(
self.lock_dir,
f"prgs-Scaled-Tech-Consulting-Gitea-Tools-{issue}.json",
_live_lock_payload(
issue_number=issue,
branch="feat/issue-800-shared",
worktree_path=f"{self.tmp}/wt{issue}",
pid=os.getpid(),
),
)
control_plane_db.ControlPlaneDB(self.db_path)
snap = inventory.load_inventory_snapshot(
db_path=self.db_path,
lock_dir=self.lock_dir,
load_hygiene=lambda: _StubHygiene(entries=()),
)
self.assertIn(
"duplicate-live-lock", {c.kind for c in snap.collisions}
)
def test_no_collision_when_sections_degraded(self):
# locks ok but worktrees unavailable → lock-without-worktree must NOT
# be asserted (a missing scan is not a missing worktree).
_write_lock(
self.lock_dir,
"prgs-Scaled-Tech-Consulting-Gitea-Tools-636.json",
_live_lock_payload(
issue_number=636,
branch="feat/issue-636-inventory-api",
worktree_path=f"{self.tmp}/wt",
pid=os.getpid(),
),
)
control_plane_db.ControlPlaneDB(self.db_path)
def _boom():
raise RuntimeError("scan down")
snap = inventory.load_inventory_snapshot(
db_path=self.db_path,
lock_dir=self.lock_dir,
load_hygiene=_boom,
)
self.assertNotIn(
"lock-without-worktree", {c.kind for c in snap.collisions}
)
class TestSectionInclude(_FixtureMixin):
def test_include_restricts_scanned_sections(self):
self._seed_db()
snap = inventory.load_inventory_snapshot(
db_path=self.db_path,
lock_dir=self.lock_dir,
include=("locks",),
)
self.assertIsNotNone(snap.section("locks"))
self.assertIsNone(snap.section("sessions"))
self.assertIsNone(snap.section("worktrees"))
class TestRoutes(_FixtureMixin):
def setUp(self) -> None:
super().setUp()
# Point the loaders at the fixture DB and lock dir via env, and stub
# the worktree scan so the route does not shell out to git.
self._prev_env = {
"GITEA_CONTROL_PLANE_DB": os.environ.get("GITEA_CONTROL_PLANE_DB"),
"GITEA_ISSUE_LOCK_DIR": os.environ.get("GITEA_ISSUE_LOCK_DIR"),
"WEBUI_TEST_OFFLINE": os.environ.get("WEBUI_TEST_OFFLINE"),
}
os.environ["GITEA_CONTROL_PLANE_DB"] = self.db_path
os.environ["GITEA_ISSUE_LOCK_DIR"] = self.lock_dir
os.environ["WEBUI_TEST_OFFLINE"] = "1"
self._seed_db()
self.client = TestClient(create_app())
def tearDown(self) -> None:
for key, value in self._prev_env.items():
if value is None:
os.environ.pop(key, None)
else:
os.environ[key] = value
def test_inventory_route_returns_versioned_payload(self):
resp = self.client.get("/api/v1/inventory")
self.assertEqual(resp.status_code, 200)
body = resp.json()
self.assertEqual(body["api_version"], "v1")
self.assertIn("sessions", body["sections"])
self.assertIn("field_authority", body)
def test_section_route_restricts_and_labels(self):
resp = self.client.get("/api/v1/inventory/locks")
self.assertEqual(resp.status_code, 200)
body = resp.json()
self.assertEqual(body["requested_section"], "locks")
self.assertIn("locks", body["sections"])
self.assertNotIn("sessions", body["sections"])
def test_unknown_section_is_404(self):
resp = self.client.get("/api/v1/inventory/bogus")
self.assertEqual(resp.status_code, 404)
self.assertEqual(resp.json()["error"], "unknown_section")
def test_inventory_route_rejects_post(self):
resp = self.client.post("/api/v1/inventory")
self.assertEqual(resp.status_code, 405)
class TestReadOnly(_FixtureMixin):
def test_snapshot_does_not_create_db_file(self):
missing = os.path.join(self.tmp, "does-not-exist.db")
inventory.load_inventory_snapshot(
db_path=missing,
lock_dir=self.lock_dir,
load_hygiene=lambda: _StubHygiene(entries=()),
)
self.assertFalse(os.path.exists(missing))
def test_readonly_connection_refuses_write(self):
self._seed_db()
conn = inventory._open_readonly(self.db_path)
try:
with self.assertRaises(Exception):
conn.execute(
"INSERT INTO sessions(session_id, role, started_at, "
"last_heartbeat_at, status) VALUES ('x','author',"
"'t','t','active')"
)
conn.commit()
finally:
conn.close()
# ── lightweight stand-ins for the #432 hygiene snapshot ──────────────────────
class _StubEntry:
def __init__(
self,
*,
rel_path: str,
branch: str | None = None,
classification: str = "stale-clean",
head_sha: str | None = "abc123",
dirty_tracked: int = 0,
dirty_untracked: bool = False,
detached: bool = False,
registered_worktree: bool = True,
notes: str = "",
) -> None:
self.rel_path = rel_path
self.folder_name = rel_path.split("/", 1)[-1]
self.branch = branch
self.classification = classification
self.head_sha = head_sha
self.dirty_tracked = dirty_tracked
self.dirty_untracked = dirty_untracked
self.detached = detached
self.registered_worktree = registered_worktree
self.notes = notes
class _StubHygiene:
def __init__(self, *, entries=(), scan_error=None) -> None:
self.entries = tuple(entries)
self.scan_error = scan_error
if __name__ == "__main__":
unittest.main()
+345
View File
@@ -0,0 +1,345 @@
"""Tests for the system-health dashboard view (#639).
Covers the acceptance criteria directly: the page renders the health DTO
fields (AC1), degraded dependencies are visible (AC2), stale runtime is warned
prominently and never rendered as mutation-safe (AC3), healthy and degraded
fixtures both render (AC4), and the shell carries a nav entry (AC5).
"""
import sys
import unittest
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
from starlette.testclient import TestClient
from webui.app import create_app
from webui.deployment_boundary import scan_text_for_client_secrets
from webui.layout import render_page
from webui.nav import iter_nav_items
from webui.system_health import (
STATUS_DEGRADED,
STATUS_DOWN,
STATUS_OK,
STATUS_SKIPPED,
STATUS_UNPROVEN,
DependencyProbe,
StaleRuntime,
SystemHealthSnapshot,
VersionInfo,
)
from webui.system_health_views import render_system_health_page
DASHBOARD_PATH = "/system-health"
def _version(*, known: bool = True) -> VersionInfo:
return VersionInfo(
git_sha="1c455b6ec0f9cb761fe6248de68c17e061fb5ecd" if known else None,
git_describe="v0.4.1-12-g1c455b6" if known else None,
control_plane_schema_version=4 if known else None,
python_version="3.13.1",
known=known,
)
def _parity(*, stale: bool = False, determinable: bool = True) -> StaleRuntime:
if stale:
return StaleRuntime(
daemon_head="aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
checkout_head="bbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbb",
remote_head="cccccccccccccccccccccccccccccccccccccccc",
stale=True,
determinable=True,
mutation_safe=False,
reasons=("runtime, checkout, and remote commits disagree",),
)
if not determinable:
return StaleRuntime(
daemon_head=None,
checkout_head=None,
remote_head=None,
stale=False,
determinable=False,
mutation_safe=False,
reasons=("local checkout HEAD could not be read",),
)
return StaleRuntime(
daemon_head="1c455b6ec0f9cb761fe6248de68c17e061fb5ecd",
checkout_head="1c455b6ec0f9cb761fe6248de68c17e061fb5ecd",
remote_head="1c455b6ec0f9cb761fe6248de68c17e061fb5ecd",
stale=False,
determinable=True,
mutation_safe=True,
reasons=(),
)
def _snapshot(
*,
status: str = STATUS_OK,
ready: bool = True,
readiness_complete: bool = True,
readiness_reasons: tuple[str, ...] = (),
dependencies: tuple[DependencyProbe, ...] | None = None,
parity: StaleRuntime | None = None,
namespaces: tuple[dict, ...] = (),
probe_errors: tuple[str, ...] = (),
version_known: bool = True,
) -> SystemHealthSnapshot:
if dependencies is None:
dependencies = (
DependencyProbe(
name="control_plane_db",
kind="sqlite",
status=STATUS_OK,
detail="schema version 4",
required=True,
latency_ms=1.25,
metadata={"schema_version": 4},
),
)
return SystemHealthSnapshot(
status=status,
ready=ready,
readiness_complete=readiness_complete,
readiness_reasons=readiness_reasons,
service="mcp-control-plane-webui",
mode="read-only",
version=_version(known=version_known),
started_at="2026-07-23T19:50:47+00:00",
uptime_seconds=3661.5,
timestamp="2026-07-23T20:51:48+00:00",
deep_probes_requested=False,
dependencies=dependencies,
mcp_namespaces=namespaces,
stale_runtime=parity if parity is not None else _parity(),
probe_errors=probe_errors,
)
class TestHealthyRender(unittest.TestCase):
"""AC1 / AC4 — every health DTO field reaches the page."""
def setUp(self):
self.html = render_system_health_page(_snapshot())
def test_readiness_fields_render(self):
self.assertIn("System health", self.html)
self.assertIn("Ready", self.html)
self.assertIn("mcp-control-plane-webui", self.html)
self.assertIn("read-only", self.html)
self.assertIn("2026-07-23T20:51:48+00:00", self.html)
def test_version_and_uptime_render(self):
self.assertIn("1c455b6ec0f9cb761fe6248de68c17e061fb5ecd", self.html)
self.assertIn("v0.4.1-12-g1c455b6", self.html)
self.assertIn("3.13.1", self.html)
self.assertIn("3661.500s", self.html)
self.assertIn("1.02h", self.html)
def test_dependency_row_renders_with_latency(self):
self.assertIn("control_plane_db", self.html)
self.assertIn("sqlite", self.html)
self.assertIn("schema version 4", self.html)
self.assertIn("1.2 ms", self.html)
def test_healthy_page_shows_no_stale_warning(self):
self.assertNotIn("Stale runtime:", self.html)
self.assertNotIn("Staleness", self.html)
def test_unknown_version_is_labelled_not_faked(self):
html = render_system_health_page(_snapshot(version_known=False))
self.assertIn("unknown", html)
self.assertIn("unresolved", html)
class TestDegradedRender(unittest.TestCase):
"""AC2 — a degraded or unrun dependency is visible, not swallowed."""
def setUp(self):
self.deps = (
DependencyProbe(
name="control_plane_db",
kind="sqlite",
status=STATUS_OK,
detail="schema version 4",
required=True,
latency_ms=0.9,
),
DependencyProbe(
name="repository",
kind="git",
status=STATUS_DOWN,
detail="repository root is not a git checkout",
required=True,
latency_ms=4.0,
),
DependencyProbe(
name="gitea",
kind="http",
status=STATUS_SKIPPED,
detail="deep probe not requested",
required=False,
),
)
self.html = render_system_health_page(
_snapshot(
status=STATUS_DEGRADED,
ready=False,
readiness_complete=False,
readiness_reasons=("required dependency 'repository' is down",),
dependencies=self.deps,
)
)
def test_degraded_banner_names_the_dependency(self):
self.assertIn("Degraded dependencies:", self.html)
self.assertIn("repository", self.html)
def test_not_run_probe_is_reported_separately(self):
self.assertIn("Not probed:", self.html)
self.assertIn("gitea", self.html)
self.assertIn("not counted", self.html)
def test_not_ready_headline_and_reason(self):
self.assertIn("Not ready", self.html)
self.assertIn("required dependency &#x27;repository&#x27; is down", self.html)
def test_degraded_status_badge_present(self):
self.assertIn("badge-health-degraded", self.html)
self.assertIn("badge-health-down", self.html)
def test_ready_but_incomplete_is_not_shown_as_plain_ready(self):
html = render_system_health_page(
_snapshot(ready=True, readiness_complete=False)
)
self.assertIn("Ready (incomplete evidence)", html)
class TestStaleRuntimeWarning(unittest.TestCase):
"""AC3 — staleness is prominent and never claims mutation safety."""
def test_stale_runtime_warns_and_denies_mutation_safety(self):
html = render_system_health_page(_snapshot(parity=_parity(stale=True)))
self.assertIn("Stale runtime:", html)
self.assertIn("do not treat this runtime as mutation-safe", html)
self.assertIn("<tr><th>Mutation safe</th><td>False</td></tr>", html)
def test_indeterminate_parity_is_not_reported_safe(self):
html = render_system_health_page(
_snapshot(parity=_parity(determinable=False))
)
self.assertIn("Staleness", html)
self.assertIn("<tr><th>Mutation safe</th><td>False</td></tr>", html)
self.assertIn("<tr><th>Determinable</th><td>False</td></tr>", html)
def test_healthy_parity_reports_mutation_safe_true(self):
html = render_system_health_page(_snapshot())
self.assertIn("<tr><th>Mutation safe</th><td>True</td></tr>", html)
class TestNamespacesAndErrors(unittest.TestCase):
def test_unproven_namespace_rows_render(self):
html = render_system_health_page(
_snapshot(
namespaces=(
{
"namespace": "gitea-author",
"required_tool": "gitea_lock_issue",
"status": STATUS_UNPROVEN,
"ide_namespace_proven": False,
"reason": "the web console cannot invoke the IDE-managed MCP client",
},
)
)
)
self.assertIn("gitea-author", html)
self.assertIn("gitea_lock_issue", html)
self.assertIn("badge-health-unproven", html)
def test_no_namespaces_degrades_gracefully(self):
html = render_system_health_page(_snapshot(namespaces=()))
self.assertIn("No MCP namespaces are declared.", html)
def test_probe_errors_render_when_present(self):
html = render_system_health_page(
_snapshot(probe_errors=("probe raised: disk offline",))
)
self.assertIn("Probe errors", html)
self.assertIn("disk offline", html)
def test_probe_error_card_absent_when_clean(self):
self.assertNotIn("Probe errors", render_system_health_page(_snapshot()))
class TestReadOnlyAndRedaction(unittest.TestCase):
def test_no_restart_or_kill_controls(self):
html = render_system_health_page(_snapshot())
self.assertNotIn("<button", html)
self.assertNotIn("<form", html)
self.assertNotIn("pkill", html)
self.assertIn("read-only", html)
def test_recovery_points_at_sanctioned_path(self):
html = render_system_health_page(_snapshot())
self.assertIn("Reconnect the MCP client", html)
self.assertIn("Never kill the daemon process manually", html)
def test_secret_shaped_detail_is_redacted(self):
leaky = DependencyProbe(
name="gitea",
kind="http",
status=STATUS_DOWN,
detail="auth failed for token=ghp_ABCDEFGHIJKLMNOPQRSTUVWXYZ0123456789",
required=False,
latency_ms=12.0,
)
html = render_system_health_page(_snapshot(dependencies=(leaky,)))
self.assertNotIn("ghp_ABCDEFGHIJKLMNOPQRSTUVWXYZ0123456789", html)
def test_html_in_detail_is_escaped(self):
hostile = DependencyProbe(
name="repository",
kind="git",
status=STATUS_DOWN,
detail="<script>alert(1)</script>",
required=True,
)
html = render_system_health_page(_snapshot(dependencies=(hostile,)))
self.assertNotIn("<script>", html)
self.assertIn("&lt;script&gt;", html)
class TestNavAndRoute(unittest.TestCase):
"""AC5 — the shell links the dashboard, and the route serves it."""
def setUp(self):
self.client = TestClient(create_app())
def test_nav_contains_system_health(self):
self.assertIn(
(DASHBOARD_PATH, "System health"),
[(item.href, item.label) for item in iter_nav_items()],
)
def test_rendered_shell_links_dashboard(self):
page = render_page(title="Home", body_html="<p>x</p>")
self.assertIn(f'href="{DASHBOARD_PATH}"', page)
def test_route_renders_dashboard(self):
response = self.client.get(DASHBOARD_PATH)
self.assertEqual(response.status_code, 200)
self.assertIn("System health", response.text)
self.assertIn("Stale-runtime parity", response.text)
def test_route_is_read_only(self):
self.assertEqual(self.client.post(DASHBOARD_PATH).status_code, 405)
def test_live_page_leaks_no_client_secret(self):
findings = scan_text_for_client_secrets(self.client.get(DASHBOARD_PATH).text)
self.assertEqual(findings, [])
if __name__ == "__main__": # pragma: no cover
unittest.main()
+55
View File
@@ -46,6 +46,11 @@ from webui.worktree_scanner import load_hygiene_snapshot, snapshot_to_dict as wo
from webui.worktree_views import render_worktrees_page from webui.worktree_views import render_worktrees_page
from webui.runtime_health import load_runtime_snapshot, snapshot_to_dict as runtime_snapshot_to_dict from webui.runtime_health import load_runtime_snapshot, snapshot_to_dict as runtime_snapshot_to_dict
from webui.runtime_views import render_runtime_page from webui.runtime_views import render_runtime_page
from webui.inventory import (
SECTION_NAMES as _INVENTORY_SECTIONS,
load_inventory_snapshot,
snapshot_to_dict as inventory_snapshot_to_dict,
)
from webui.timeline import load_timeline, snapshot_to_dict as timeline_snapshot_to_dict from webui.timeline import load_timeline, snapshot_to_dict as timeline_snapshot_to_dict
from webui.system_health import ( from webui.system_health import (
API_PATH as SYSTEM_HEALTH_API_PATH, API_PATH as SYSTEM_HEALTH_API_PATH,
@@ -53,6 +58,7 @@ from webui.system_health import (
process_uptime, process_uptime,
snapshot_to_dict as system_health_to_dict, snapshot_to_dict as system_health_to_dict,
) )
from webui.system_health_views import render_system_health_page
_READ_ONLY_METHODS = frozenset({"GET", "HEAD", "OPTIONS"}) _READ_ONLY_METHODS = frozenset({"GET", "HEAD", "OPTIONS"})
_AUDIT_MUTATION_PATHS = frozenset({"/audit", "/api/audit"}) _AUDIT_MUTATION_PATHS = frozenset({"/audit", "/api/audit"})
@@ -161,6 +167,24 @@ async def api_system_health(request: Request) -> JSONResponse:
return JSONResponse(payload, status_code=200 if snapshot.ready else 503) return JSONResponse(payload, status_code=200 if snapshot.ready else 503)
async def system_health(request: Request) -> HTMLResponse:
"""Read-only system-health dashboard (#639).
Shares the #634 snapshot loader with the JSON API so the page can never
disagree with it. `?deep=1` opts into the network probe exactly as the API
does; the default page load stays cheap. The response is always 200: this
is an operator view that must render the degraded state, not withhold it.
"""
deep = _truthy_flag(request.query_params.get("deep"))
snapshot = load_system_health(deep=deep)
return HTMLResponse(
render_page(
title="System health",
body_html=render_system_health_page(snapshot),
)
)
async def queue(_request: Request) -> HTMLResponse: async def queue(_request: Request) -> HTMLResponse:
snapshot = load_queue_snapshot() snapshot = load_queue_snapshot()
return HTMLResponse(render_page(title="Queue", body_html=render_queue_page(snapshot))) return HTMLResponse(render_page(title="Queue", body_html=render_queue_page(snapshot)))
@@ -441,6 +465,30 @@ async def api_action_attempt(request: Request) -> JSONResponse:
return JSONResponse(result, status_code=status) return JSONResponse(result, status_code=status)
async def api_inventory(_request: Request) -> JSONResponse:
"""Unified read-only session/lease/lock/worktree inventory (#636)."""
snapshot = load_inventory_snapshot()
return JSONResponse(inventory_snapshot_to_dict(snapshot))
async def api_inventory_section(request: Request) -> JSONResponse:
"""Resource-split view: one inventory section under the shared schema."""
section = request.path_params["section"]
if section not in _INVENTORY_SECTIONS:
return JSONResponse(
{
"error": "unknown_section",
"detail": f"no inventory section named {section!r}",
"available": sorted(_INVENTORY_SECTIONS),
},
status_code=404,
)
snapshot = load_inventory_snapshot(include=(section,))
payload = inventory_snapshot_to_dict(snapshot)
payload["requested_section"] = section
return JSONResponse(payload)
async def api_console_security_model(_request: Request) -> JSONResponse: async def api_console_security_model(_request: Request) -> JSONResponse:
"""Read-only publication of the #633 authorization/redaction/audit model.""" """Read-only publication of the #633 authorization/redaction/audit model."""
return JSONResponse({ return JSONResponse({
@@ -571,6 +619,7 @@ def create_app(*, bind_host: str | None = None) -> Starlette:
Route("/", home, methods=["GET"]), Route("/", home, methods=["GET"]),
Route("/health", health, methods=["GET"]), Route("/health", health, methods=["GET"]),
Route(SYSTEM_HEALTH_API_PATH, api_system_health, methods=["GET"]), Route(SYSTEM_HEALTH_API_PATH, api_system_health, methods=["GET"]),
Route("/system-health", system_health, methods=["GET"]),
Route("/queue", queue, methods=["GET"]), Route("/queue", queue, methods=["GET"]),
Route("/api/queue", api_queue, methods=["GET"]), Route("/api/queue", api_queue, methods=["GET"]),
Route("/projects", projects, methods=["GET"]), Route("/projects", projects, methods=["GET"]),
@@ -606,6 +655,12 @@ def create_app(*, bind_host: str | None = None) -> Starlette:
methods=["POST"], methods=["POST"],
), ),
Route("/api/leases", api_leases, methods=["GET"]), Route("/api/leases", api_leases, methods=["GET"]),
Route("/api/v1/inventory", api_inventory, methods=["GET"]),
Route(
"/api/v1/inventory/{section}",
api_inventory_section,
methods=["GET"],
),
Route( Route(
"/api/console/security-model", "/api/console/security-model",
api_console_security_model, api_console_security_model,
+952
View File
@@ -0,0 +1,952 @@
"""Unified session/lease/lock/worktree inventory for the web console (#636).
Leases (#433), worktrees (#432), and runtime (#430) each ship their own MVP
view, each with its own shape and its own idea of what "owned" means. A
traffic-control or recovery operator has to read all three and correlate them
by hand, which is exactly the step that goes wrong under collision pressure.
This module aggregates them into one versioned, read-only snapshot so the
console, and any worker asking "what is safe to do next", read the same
inventory from the same authority.
Field authority is explicit and never blended. Every section declares where its
rows came from:
* ``control_plane_db`` — the #613 substrate: sessions, leases, assignments.
Authoritative for *exclusive ownership* (#600/#601).
* ``filesystem`` — durable per-issue lock files (:mod:`issue_lock_store`) and
registered git worktrees. Authoritative for *what exists on this machine*.
* ``gitea`` — remote issue/PR state, reached only through existing loaders.
Safety invariants:
* **Read-only.** The control-plane database is opened through a ``mode=ro``
URI. :class:`control_plane_db.ControlPlaneDB` creates directories and runs
migrations in its constructor, which an inventory read must never do, so this
module talks to sqlite directly rather than through that class.
* **Fail-soft, never fail-silent.** A subsystem that cannot be read degrades to
a section carrying ``status`` and ``reason``. It never raises, and it never
produces an empty list that reads like "nothing is there".
* **Never invent active ownership.** This is the invariant that matters most.
A degraded ownership source sets ``ownership_authority_complete`` false, and
while that flag is false no work item is reported unowned and no collision is
asserted. Absence of evidence is reported as absence of evidence.
* **Redaction at the boundary.** Absolute paths are collapsed against the home
directory, URLs lose userinfo and query strings, and no credential-shaped
value is emitted. No session token exists in these sources and none is read.
Phase 1 is read-only. Lease steal/release and worktree deletion are Phase 2+
and deliberately have no representation here, not even a disabled one.
"""
from __future__ import annotations
import os
import re
import sqlite3
import time
from dataclasses import dataclass, field
from datetime import datetime, timezone
from typing import Any, Callable
from urllib.parse import urlparse
import control_plane_db
import issue_lock_store
SCHEMA_VERSION = 1
API_VERSION = "v1"
#: Sections whose absence would make an ownership claim unprovable. If any of
#: these is not ``ok``, the snapshot refuses to describe anything as unowned.
OWNERSHIP_SECTIONS = ("sessions", "leases", "locks")
SECTION_NAMES = ("sessions", "leases", "locks", "worktrees", "namespaces")
STATUS_OK = "ok"
STATUS_DEGRADED = "degraded"
STATUS_UNAVAILABLE = "unavailable"
AUTHORITY_CONTROL_PLANE_DB = "control_plane_db"
AUTHORITY_FILESYSTEM = "filesystem"
AUTHORITY_GITEA = "gitea"
_CREDENTIAL_KEY_RE = re.compile(
r"(token|secret|password|passwd|api[_-]?key|authorization|bearer|credential)",
re.IGNORECASE,
)
_REDACTED = "[redacted]"
@dataclass(frozen=True)
class InventorySection:
"""One subsystem's contribution, with its authority and health."""
name: str
authority: str
status: str
items: tuple[dict[str, Any], ...] = ()
reason: str | None = None
scan_ms: float | None = None
@property
def ok(self) -> bool:
return self.status == STATUS_OK
def to_dict(self) -> dict[str, Any]:
return {
"name": self.name,
"authority": self.authority,
"status": self.status,
"count": len(self.items),
"reason": self.reason,
"scan_ms": self.scan_ms,
"items": [dict(item) for item in self.items],
}
@dataclass(frozen=True)
class CollisionSignal:
"""A detected conflict between two ownership records."""
kind: str
message: str
severity: str = "warning"
issue_number: int | None = None
branch: str | None = None
worktree_path: str | None = None
session_ids: tuple[str, ...] = ()
def to_dict(self) -> dict[str, Any]:
return {
"kind": self.kind,
"severity": self.severity,
"message": self.message,
"issue_number": self.issue_number,
"branch": self.branch,
"worktree_path": self.worktree_path,
"session_ids": list(self.session_ids),
}
@dataclass(frozen=True)
class InventorySnapshot:
"""Versioned aggregate of every inventory section."""
generated_at: str
sections: tuple[InventorySection, ...]
collisions: tuple[CollisionSignal, ...] = ()
correlations: tuple[dict[str, Any], ...] = ()
schema_version: int = SCHEMA_VERSION
api_version: str = API_VERSION
scan_ms: float | None = None
_section_index: dict[str, InventorySection] = field(
default_factory=dict, repr=False, compare=False
)
def section(self, name: str) -> InventorySection | None:
return self._section_index.get(name)
@property
def degraded_sections(self) -> tuple[str, ...]:
return tuple(s.name for s in self.sections if not s.ok)
@property
def ownership_authority_complete(self) -> bool:
"""True only when every ownership-bearing section read cleanly.
While this is false the snapshot must not describe any work item as
unowned: a lease the reader could not load is not an absent lease.
"""
for name in OWNERSHIP_SECTIONS:
section = self._section_index.get(name)
if section is None or not section.ok:
return False
return True
@property
def status(self) -> str:
if all(s.ok for s in self.sections):
return STATUS_OK
return STATUS_DEGRADED
# ── redaction ────────────────────────────────────────────────────────────────
def redact_path(path: str | None) -> str | None:
"""Collapse an absolute path against ``$HOME`` for browser display."""
if not path:
return path
text = str(path)
home = os.path.expanduser("~")
if home and home != "/" and text.startswith(home):
return "~" + text[len(home) :]
return text
def redact_url(value: str | None) -> str | None:
"""Strip userinfo and query string from a URL."""
if not value:
return value
text = str(value)
try:
parsed = urlparse(text)
except ValueError:
return _REDACTED
if not parsed.scheme or not parsed.netloc:
return text
netloc = parsed.hostname or ""
if parsed.port:
netloc = f"{netloc}:{parsed.port}"
rebuilt = f"{parsed.scheme}://{netloc}{parsed.path}"
return rebuilt.rstrip("/") or rebuilt
def scrub(value: Any, *, key: str | None = None) -> Any:
"""Recursively drop credential-shaped values and redact paths/URLs.
Never raises: an unexpected object degrades to its ``repr`` rather than
propagating out of a read-only view.
"""
if key and _CREDENTIAL_KEY_RE.search(key):
return _REDACTED
if isinstance(value, dict):
return {str(k): scrub(v, key=str(k)) for k, v in value.items()}
if isinstance(value, (list, tuple)):
return [scrub(v, key=key) for v in value]
if isinstance(value, str):
if value.startswith(("http://", "https://")):
return redact_url(value)
if value.startswith("/") or value.startswith("~"):
return redact_path(value)
return value
if isinstance(value, (int, float, bool)) or value is None:
return value
return repr(value)
# ── control-plane database (read-only) ───────────────────────────────────────
def _open_readonly(db_path: str) -> sqlite3.Connection:
"""Open the control-plane DB without creating or migrating anything."""
conn = sqlite3.connect(f"file:{db_path}?mode=ro", uri=True, timeout=5)
conn.row_factory = sqlite3.Row
return conn
def _table_names(conn: sqlite3.Connection) -> set[str]:
rows = conn.execute(
"SELECT name FROM sqlite_master WHERE type = 'table'"
).fetchall()
return {str(row[0]) for row in rows}
def _load_cp_db_sections(
*,
db_path: str | None = None,
limit: int = 200,
) -> tuple[InventorySection, InventorySection]:
"""Return the ``sessions`` and ``leases`` sections from the #613 DB."""
path = (db_path or control_plane_db.default_db_path()).strip()
def _both_unavailable(reason: str) -> tuple[InventorySection, InventorySection]:
return (
InventorySection(
name="sessions",
authority=AUTHORITY_CONTROL_PLANE_DB,
status=STATUS_UNAVAILABLE,
reason=reason,
),
InventorySection(
name="leases",
authority=AUTHORITY_CONTROL_PLANE_DB,
status=STATUS_UNAVAILABLE,
reason=reason,
),
)
if not path:
return _both_unavailable("control-plane database path is not configured")
if not os.path.exists(path):
return _both_unavailable(
f"control-plane database not present at {redact_path(path)}; "
"no session or lease authority available"
)
started = time.perf_counter()
try:
conn = _open_readonly(path)
except sqlite3.Error as exc:
return _both_unavailable(f"control-plane database could not be opened: {exc}")
try:
tables = _table_names(conn)
if "sessions" not in tables or "leases" not in tables:
missing = sorted({"sessions", "leases"} - tables)
return _both_unavailable(
"control-plane database is missing required tables: "
+ ", ".join(missing)
)
session_rows = [
dict(row)
for row in conn.execute(
"SELECT session_id, role, profile, namespace, pid, started_at,"
" last_heartbeat_at, status FROM sessions"
" ORDER BY last_heartbeat_at DESC LIMIT ?",
(max(1, int(limit)),),
).fetchall()
]
has_work_items = "work_items" in tables
if has_work_items:
lease_sql = (
"SELECT l.lease_id, l.session_id, l.role, l.phase, l.status,"
" l.expires_at, w.remote, w.org, w.repo, w.kind AS work_kind,"
" w.number AS work_number, w.state AS work_state,"
" s.pid AS session_pid, s.profile AS session_profile,"
" s.namespace AS session_namespace, s.status AS session_status"
" FROM leases l"
" JOIN work_items w ON w.work_item_id = l.work_item_id"
" LEFT JOIN sessions s ON s.session_id = l.session_id"
" ORDER BY l.expires_at DESC LIMIT ?"
)
else:
lease_sql = (
"SELECT l.lease_id, l.session_id, l.role, l.phase, l.status,"
" l.expires_at FROM leases l"
" ORDER BY l.expires_at DESC LIMIT ?"
)
lease_rows = [
dict(row)
for row in conn.execute(lease_sql, (max(1, int(limit)),)).fetchall()
]
except sqlite3.Error as exc:
return _both_unavailable(f"control-plane database read failed: {exc}")
finally:
conn.close()
elapsed = (time.perf_counter() - started) * 1000.0
now = datetime.now(timezone.utc)
sessions = tuple(
scrub(
{
"session_id": row.get("session_id"),
"role": row.get("role"),
"profile": row.get("profile"),
"namespace": row.get("namespace"),
"pid": row.get("pid"),
"pid_alive": issue_lock_store.is_process_alive(row.get("pid")),
"started_at": row.get("started_at"),
"last_heartbeat_at": row.get("last_heartbeat_at"),
"status": row.get("status"),
}
)
for row in session_rows
)
leases = tuple(
scrub(
{
"lease_id": row.get("lease_id"),
"session_id": row.get("session_id"),
"role": row.get("role"),
"phase": row.get("phase"),
"status": row.get("status"),
"expires_at": row.get("expires_at"),
"expired": _is_expired(row.get("expires_at"), now=now),
"remote": row.get("remote"),
"org": row.get("org"),
"repo": row.get("repo"),
"work_kind": row.get("work_kind"),
"work_number": row.get("work_number"),
"work_state": row.get("work_state"),
"session_pid": row.get("session_pid"),
"session_profile": row.get("session_profile"),
"session_namespace": row.get("session_namespace"),
"session_status": row.get("session_status"),
}
)
for row in lease_rows
)
degraded_reason = (
None
if has_work_items
else "work_items table absent; lease rows carry no work linkage"
)
lease_status = STATUS_OK if has_work_items else STATUS_DEGRADED
return (
InventorySection(
name="sessions",
authority=AUTHORITY_CONTROL_PLANE_DB,
status=STATUS_OK,
items=sessions,
scan_ms=round(elapsed, 3),
),
InventorySection(
name="leases",
authority=AUTHORITY_CONTROL_PLANE_DB,
status=lease_status,
items=leases,
reason=degraded_reason,
scan_ms=round(elapsed, 3),
),
)
def _is_expired(expires_at: str | None, *, now: datetime) -> bool | None:
if not expires_at:
return None
text = str(expires_at).strip().replace("Z", "+00:00")
try:
parsed = datetime.fromisoformat(text)
except ValueError:
return None
if parsed.tzinfo is None:
parsed = parsed.replace(tzinfo=timezone.utc)
return parsed <= now
# ── durable issue locks (filesystem) ─────────────────────────────────────────
def _load_locks_section(*, lock_dir: str | None = None) -> InventorySection:
started = time.perf_counter()
try:
paths = issue_lock_store.iter_lock_files(lock_dir)
except OSError as exc:
return InventorySection(
name="locks",
authority=AUTHORITY_FILESYSTEM,
status=STATUS_UNAVAILABLE,
reason=f"issue lock directory could not be listed: {exc}",
)
items: list[dict[str, Any]] = []
unreadable = 0
for path in paths:
try:
record = issue_lock_store.read_lock_file(path)
except (OSError, ValueError):
unreadable += 1
continue
if not record:
unreadable += 1
continue
try:
freshness = issue_lock_store.assess_lock_freshness(record)
except Exception: # noqa: BLE001 — a read-only view never raises
freshness = {"status": "unknown", "live": False, "stale": False}
claimant = record.get("claimant") or (
(record.get("work_lease") or {}).get("claimant") or {}
)
items.append(
scrub(
{
"issue_number": record.get("issue_number"),
"branch_name": record.get("branch_name"),
"remote": record.get("remote"),
"org": record.get("org"),
"repo": record.get("repo"),
"worktree_path": record.get("worktree_path"),
"pid": record.get("session_pid") or record.get("pid"),
"pid_alive": issue_lock_store.is_process_alive(
record.get("session_pid") or record.get("pid")
),
"claimant_username": (claimant or {}).get("username"),
"claimant_profile": (claimant or {}).get("profile"),
"lock_generation": record.get("lock_generation"),
"freshness_status": freshness.get("status"),
"live": bool(freshness.get("live")),
"stale": bool(freshness.get("stale")),
"freshness_reason": freshness.get("reason"),
"lock_path": record.get("lock_file_path") or path,
}
)
)
elapsed = (time.perf_counter() - started) * 1000.0
reason = (
f"{unreadable} lock file(s) were unreadable and are not represented"
if unreadable
else None
)
return InventorySection(
name="locks",
authority=AUTHORITY_FILESYSTEM,
status=STATUS_DEGRADED if unreadable else STATUS_OK,
items=tuple(items),
reason=reason,
scan_ms=round(elapsed, 3),
)
# ── worktrees (filesystem, via the #432 scanner) ─────────────────────────────
def _load_worktrees_section(
*, load_hygiene: Callable[[], Any] | None = None
) -> InventorySection:
started = time.perf_counter()
try:
loader = load_hygiene
if loader is None:
from webui.worktree_scanner import load_hygiene_snapshot
loader = load_hygiene_snapshot
snapshot = loader()
except Exception as exc: # noqa: BLE001 — fail soft, never fail the request
return InventorySection(
name="worktrees",
authority=AUTHORITY_FILESYSTEM,
status=STATUS_UNAVAILABLE,
reason=f"worktree scan failed: {exc}",
)
items = tuple(
scrub(
{
"rel_path": entry.rel_path,
"folder_name": entry.folder_name,
"classification": entry.classification,
"branch": entry.branch,
"head_sha": entry.head_sha,
"dirty_tracked": entry.dirty_tracked,
"dirty_untracked": entry.dirty_untracked,
"detached": entry.detached,
"registered_worktree": entry.registered_worktree,
"notes": entry.notes,
}
)
for entry in snapshot.entries
)
scan_error = getattr(snapshot, "scan_error", None)
elapsed = (time.perf_counter() - started) * 1000.0
return InventorySection(
name="worktrees",
authority=AUTHORITY_FILESYSTEM,
status=STATUS_DEGRADED if scan_error else STATUS_OK,
items=items,
reason=scan_error,
scan_ms=round(elapsed, 3),
)
# ── namespaces / capability summary ──────────────────────────────────────────
def _load_namespaces_section() -> InventorySection:
started = time.perf_counter()
try:
from gitea_auth import get_profile
profile = get_profile() or {}
except Exception as exc: # noqa: BLE001
return InventorySection(
name="namespaces",
authority=AUTHORITY_FILESYSTEM,
status=STATUS_UNAVAILABLE,
reason=f"active profile could not be resolved: {exc}",
)
allowed = list(profile.get("allowed_operations") or [])
forbidden = list(profile.get("forbidden_operations") or [])
profile_name = str(profile.get("profile_name") or "")
namespace = None
try:
import role_namespace_gate
namespace = role_namespace_gate.infer_mcp_namespace(profile_name)
except Exception: # noqa: BLE001 — namespace inference is advisory
namespace = None
item = scrub(
{
"profile_name": profile_name,
"role": profile.get("role"),
"mcp_namespace": namespace,
"allowed_operations": sorted(allowed),
"forbidden_operations": sorted(forbidden),
"capability_summary": {
"can_author": "gitea.pr.create" in allowed,
"can_review": "gitea.pr.approve" in allowed,
"can_merge": "gitea.pr.merge" in allowed,
"can_close_pr": "gitea.pr.close" in allowed,
},
"active": True,
}
)
elapsed = (time.perf_counter() - started) * 1000.0
return InventorySection(
name="namespaces",
authority=AUTHORITY_FILESYSTEM,
status=STATUS_OK,
items=(item,),
reason=(
"only the profile serving this web process is observable; other "
"namespaces are not enumerable from here"
),
scan_ms=round(elapsed, 3),
)
# ── correlation and collision detection ──────────────────────────────────────
def _issue_from_branch(branch: str | None) -> int | None:
match = re.search(r"issue-(\d+)", str(branch or ""), re.IGNORECASE)
return int(match.group(1)) if match else None
def correlate(
*,
leases: InventorySection,
locks: InventorySection,
worktrees: InventorySection,
sessions: InventorySection,
) -> tuple[tuple[dict[str, Any], ...], tuple[CollisionSignal, ...]]:
"""Join lease owner ↔ lock ↔ worktree ↔ namespace where evidence allows.
Correlation rows are emitted from whatever sections did load. Collision
signals are only emitted from sections that are ``ok``: a conflict inferred
from a partially-read source would be a false accusation.
"""
correlations: list[dict[str, Any]] = []
collisions: list[CollisionSignal] = []
worktree_by_branch: dict[str, dict[str, Any]] = {}
for entry in worktrees.items:
branch = (entry.get("branch") or "").strip()
if branch:
worktree_by_branch.setdefault(branch, entry)
session_by_id = {
str(s.get("session_id")): s for s in sessions.items if s.get("session_id")
}
# Lock-centred rows: a durable lock names an issue, a branch, and a worktree.
for lock in locks.items:
branch = (lock.get("branch_name") or "").strip()
worktree = worktree_by_branch.get(branch)
matching_leases = [
lease
for lease in leases.items
if lease.get("work_kind") == "issue"
and lease.get("work_number") == lock.get("issue_number")
]
correlations.append(
{
"issue_number": lock.get("issue_number"),
"branch": branch or None,
"lock_live": bool(lock.get("live")),
"lock_claimant": lock.get("claimant_profile"),
"lock_pid": lock.get("pid"),
"lock_pid_alive": lock.get("pid_alive"),
"worktree_rel_path": (worktree or {}).get("rel_path"),
"worktree_classification": (worktree or {}).get("classification"),
"worktree_registered": (worktree or {}).get("registered_worktree"),
"lease_ids": [
lease.get("lease_id")
for lease in matching_leases
if lease.get("lease_id")
],
"lease_sessions": [
lease.get("session_id")
for lease in matching_leases
if lease.get("session_id")
],
}
)
if locks.ok and worktrees.ok:
# A claim whose lease window is still open but has no registered
# worktree is an anomaly regardless of whether its pid is alive; a
# fully time-expired lease is on its way out and is not flagged.
if (
lock.get("freshness_status") != "expired"
and branch
and worktree is None
):
collisions.append(
CollisionSignal(
kind="lock-without-worktree",
severity="warning",
issue_number=lock.get("issue_number"),
branch=branch,
worktree_path=lock.get("worktree_path"),
message=(
f"Live lock on issue #{lock.get('issue_number')} names "
f"branch {branch!r} but no registered worktree carries "
"that branch (#404)"
),
)
)
if locks.ok:
# A lock whose recorded pid is gone is held by nobody: a clean #753
# dead-session recovery candidate. Subclassify by the lease window,
# because the two cases need different operator urgency. When the
# window is still open the lock would read as live to a naive
# timestamp check even though the owner is dead — the more dangerous
# case — so it is flagged distinctly from a fully time-expired lease.
if lock.get("pid_alive") is False and lock.get("stale"):
if lock.get("freshness_status") == "expired":
collisions.append(
CollisionSignal(
kind="stale-lock-dead-owner",
severity="warning",
issue_number=lock.get("issue_number"),
branch=branch or None,
message=(
f"Lock on issue #{lock.get('issue_number')} is stale "
f"and its recorded pid {lock.get('pid')} is not running "
"(#753 dead-session recovery candidate)"
),
)
)
else:
collisions.append(
CollisionSignal(
kind="live-lock-dead-owner",
severity="warning",
issue_number=lock.get("issue_number"),
branch=branch or None,
message=(
f"Lock on issue #{lock.get('issue_number')} has an "
"unexpired lease but its recorded pid "
f"{lock.get('pid')} is not running; it would read as "
"live to a timestamp check (#753 dead-session recovery "
"candidate)"
),
)
)
# The #635 trap: the lease has expired but the recorded pid is a
# still-running daemon, so neither dead-pid reclaim nor exact-owner
# renewal applies. This is the collision an operator must see.
elif (
lock.get("freshness_status") == "expired"
and lock.get("pid_alive") is True
):
collisions.append(
CollisionSignal(
kind="expired-lock-live-owner",
severity="error",
issue_number=lock.get("issue_number"),
branch=branch or None,
message=(
f"Lock on issue #{lock.get('issue_number')} has an expired "
f"lease but its recorded pid {lock.get('pid')} is still "
"running (daemon-pid deadlock; needs an operator decision, "
"#635/#760)"
),
)
)
# Two live locks on one branch, or two active leases on one work item.
if locks.ok:
by_branch: dict[str, list[dict[str, Any]]] = {}
for lock in locks.items:
if not lock.get("live"):
continue
branch = (lock.get("branch_name") or "").strip()
if branch:
by_branch.setdefault(branch, []).append(lock)
for branch, entries in sorted(by_branch.items()):
if len(entries) > 1:
collisions.append(
CollisionSignal(
kind="duplicate-live-lock",
severity="error",
branch=branch,
message=(
f"{len(entries)} live locks name branch {branch!r}: "
"issues "
+ ", ".join(
f"#{e.get('issue_number')}" for e in entries
)
),
)
)
if leases.ok:
by_work: dict[tuple[str, int], list[dict[str, Any]]] = {}
for lease in leases.items:
if str(lease.get("status") or "").lower() != "active":
continue
kind = str(lease.get("work_kind") or "").strip().lower()
number = lease.get("work_number")
if not kind or number is None:
continue
by_work.setdefault((kind, int(number)), []).append(lease)
for (kind, number), entries in sorted(by_work.items()):
sessions_held = {
str(e.get("session_id")) for e in entries if e.get("session_id")
}
if len(sessions_held) > 1:
collisions.append(
CollisionSignal(
kind="concurrent-active-lease",
severity="error",
issue_number=number if kind == "issue" else None,
session_ids=tuple(sorted(sessions_held)),
message=(
f"{len(sessions_held)} sessions hold an active lease on "
f"{kind} #{number}"
),
)
)
for entry in entries:
if entry.get("expired") is True:
collisions.append(
CollisionSignal(
kind="active-lease-past-expiry",
severity="warning",
issue_number=number if kind == "issue" else None,
session_ids=(
(str(entry.get("session_id")),)
if entry.get("session_id")
else ()
),
message=(
f"Lease {entry.get('lease_id')} on {kind} #{number} "
"is still marked active past its expiry"
),
)
)
# A lease whose owning session is gone is an orphan, not free work.
if leases.ok and sessions.ok:
for lease in leases.items:
if str(lease.get("status") or "").lower() != "active":
continue
session_id = str(lease.get("session_id") or "")
if session_id and session_id not in session_by_id:
collisions.append(
CollisionSignal(
kind="orphan-lease",
severity="error",
session_ids=(session_id,),
message=(
f"Active lease {lease.get('lease_id')} names session "
f"{session_id}, which has no session record"
),
)
)
return tuple(correlations), tuple(collisions)
# ── snapshot assembly ────────────────────────────────────────────────────────
def load_inventory_snapshot(
*,
db_path: str | None = None,
lock_dir: str | None = None,
load_hygiene: Callable[[], Any] | None = None,
include: tuple[str, ...] | None = None,
) -> InventorySnapshot:
"""Build the unified inventory snapshot.
Every section is loaded independently and fails soft. *include* restricts
the sections that are scanned; omitted sections are simply absent rather
than reported as empty, so a resource-split request cannot be mistaken for
a whole-inventory answer.
"""
started = time.perf_counter()
wanted = tuple(include) if include else SECTION_NAMES
sections: list[InventorySection] = []
sessions_section: InventorySection | None = None
leases_section: InventorySection | None = None
if "sessions" in wanted or "leases" in wanted:
sessions_section, leases_section = _load_cp_db_sections(db_path=db_path)
if "sessions" in wanted:
sections.append(sessions_section)
if "leases" in wanted:
sections.append(leases_section)
locks_section = (
_load_locks_section(lock_dir=lock_dir)
if "locks" in wanted
else _empty_section("locks", AUTHORITY_FILESYSTEM)
)
if "locks" in wanted:
sections.append(locks_section)
worktrees_section = (
_load_worktrees_section(load_hygiene=load_hygiene)
if "worktrees" in wanted
else _empty_section("worktrees", AUTHORITY_FILESYSTEM)
)
if "worktrees" in wanted:
sections.append(worktrees_section)
if "namespaces" in wanted:
sections.append(_load_namespaces_section())
correlations, collisions = correlate(
leases=leases_section or _empty_section("leases", AUTHORITY_CONTROL_PLANE_DB),
locks=locks_section,
worktrees=worktrees_section,
sessions=sessions_section
or _empty_section("sessions", AUTHORITY_CONTROL_PLANE_DB),
)
elapsed = (time.perf_counter() - started) * 1000.0
index = {section.name: section for section in sections}
return InventorySnapshot(
generated_at=datetime.now(timezone.utc).isoformat(),
sections=tuple(sections),
collisions=collisions,
correlations=correlations,
scan_ms=round(elapsed, 3),
_section_index=index,
)
def _empty_section(name: str, authority: str) -> InventorySection:
"""A section that was not requested — never a claim that it is empty."""
return InventorySection(
name=name,
authority=authority,
status=STATUS_UNAVAILABLE,
reason="section not requested in this scan",
)
def snapshot_to_dict(snapshot: InventorySnapshot) -> dict[str, Any]:
"""Serialize the snapshot for the versioned API."""
return {
"api_version": snapshot.api_version,
"schema_version": snapshot.schema_version,
"generated_at": snapshot.generated_at,
"status": snapshot.status,
"scan_ms": snapshot.scan_ms,
"ownership_authority_complete": snapshot.ownership_authority_complete,
"ownership_note": (
"Every ownership source read cleanly; an item absent from leases "
"and locks is genuinely unclaimed."
if snapshot.ownership_authority_complete
else "One or more ownership sources are degraded; nothing in this "
"snapshot may be treated as unowned. Collisions are reported only "
"from sections that read cleanly."
),
"degraded_sections": list(snapshot.degraded_sections),
"field_authority": {
"sessions": AUTHORITY_CONTROL_PLANE_DB,
"leases": AUTHORITY_CONTROL_PLANE_DB,
"locks": AUTHORITY_FILESYSTEM,
"worktrees": AUTHORITY_FILESYSTEM,
"namespaces": AUTHORITY_FILESYSTEM,
},
"sections": {section.name: section.to_dict() for section in snapshot.sections},
"correlations": [dict(row) for row in snapshot.correlations],
"collisions": [signal.to_dict() for signal in snapshot.collisions],
}
+19
View File
@@ -236,6 +236,25 @@ def render_page(*, title: str, body_html: str, extra_head: str = "") -> str:
.badge-in-review {{ color: #9ec8f0; border-color: #3d5f7a; }} .badge-in-review {{ color: #9ec8f0; border-color: #3d5f7a; }}
.badge-duplicate {{ color: #e0c27a; border-color: #6b5730; }} .badge-duplicate {{ color: #e0c27a; border-color: #6b5730; }}
.badge-stale {{ color: #c9b8e8; border-color: #5a4a78; }} .badge-stale {{ color: #c9b8e8; border-color: #5a4a78; }}
.badge-health-ok {{ color: #8fd19e; border-color: #3d6b4a; }}
.badge-health-degraded {{ color: #e0c27a; border-color: #6b5730; }}
.badge-health-down {{ color: #f0a8a8; border-color: #7a3b3b; }}
.badge-health-skipped {{ color: var(--muted); }}
.badge-health-unproven {{ color: #c9b8e8; border-color: #5a4a78; }}
.health-card {{
margin: 1.25rem 0;
padding: 0.85rem 1rem 1rem;
border: 1px solid var(--border);
border-radius: 8px;
background: var(--surface);
}}
.health-card h3 {{ margin: 0 0 0.5rem; font-size: 1.05rem; }}
.health-card h4 {{ margin: 1rem 0 0.35rem; font-size: 0.92rem; color: var(--muted); }}
.health-headline {{ color: var(--text); font-size: 1rem; margin: 0 0 0.5rem; }}
.health-degraded {{ border-left-color: #e0c27a; }}
.health-stale {{ border-left-color: #f0a8a8; }}
ul.reasons {{ margin: 0.35rem 0; padding-left: 1.15rem; color: var(--muted); font-size: 0.9rem; }}
ul.reasons li {{ margin-bottom: 0.3rem; }}
</style> </style>
{extra_head} {extra_head}
</head> </head>
+1
View File
@@ -38,6 +38,7 @@ class NavGroup:
NAV_GROUPS: tuple[NavGroup, ...] = ( NAV_GROUPS: tuple[NavGroup, ...] = (
NavGroup("Health", ( NavGroup("Health", (
NavItem("/health", "Liveness"), NavItem("/health", "Liveness"),
NavItem("/system-health", "System health"),
)), )),
NavGroup("Traffic", ( NavGroup("Traffic", (
NavItem("/queue", "Queue"), NavItem("/queue", "Queue"),
+307
View File
@@ -0,0 +1,307 @@
"""HTML views for the system-health dashboard (#639).
Renders the read-only :class:`~webui.system_health.SystemHealthSnapshot`
produced by the Phase 1 system-health API (#634). The page offers no restart,
reload, or process-kill control: those are Phase 2 work, and manual process
kills are the contamination path #630 exists to prevent.
Every free-text field passes through :func:`webui.system_health.redact` before
it reaches HTML, so a probe detail that captured a token or a credentialed URL
cannot leak through the dashboard even though the API redacts it already.
"""
from __future__ import annotations
import html
from webui.system_health import (
STATUS_DEGRADED,
STATUS_DOWN,
STATUS_OK,
STATUS_SKIPPED,
STATUS_UNPROVEN,
DependencyProbe,
SystemHealthSnapshot,
redact,
)
_STATUS_BADGE_CLASS = {
STATUS_OK: "badge-health-ok",
STATUS_DEGRADED: "badge-health-degraded",
STATUS_DOWN: "badge-health-down",
STATUS_SKIPPED: "badge-health-skipped",
STATUS_UNPROVEN: "badge-health-unproven",
}
def _safe(value: object) -> str:
"""Escape free text for HTML after redacting anything secret-shaped.
Use this for every value that can carry arbitrary text probe details,
reasons, probe errors because those are where a credential could ride
along.
"""
return html.escape(redact(str(value)))
def _esc(value: object) -> str:
"""Escape a structured field for HTML without redacting it.
Commit SHAs, probe names, statuses, and timestamps are enumerated or
machine-generated, never credential-bearing. They must not go through
:func:`redact`: its opaque-token rule matches any 32-plus-character run,
so a 40-character git SHA would render as ``[redacted]`` and the parity
view the one thing an operator reads this page for would be blank.
"""
return html.escape(str(value))
def _status_badge(status: str) -> str:
css = _STATUS_BADGE_CLASS.get(status, "badge-health-unproven")
return f'<span class="badge {css}">{_esc(status)}</span>'
def _reason_list(reasons: tuple[str, ...], *, empty: str) -> str:
if not reasons:
return f"<p class='muted'>{html.escape(empty)}</p>"
items = "".join(f"<li>{_safe(reason)}</li>" for reason in reasons)
return f"<ul class='reasons'>{items}</ul>"
def _readiness_card(snapshot: SystemHealthSnapshot) -> str:
"""Overall readiness.
``ready`` and ``readiness_complete`` are shown separately on purpose: a
snapshot whose required probes never ran is not the same as one that ran
them and passed, and collapsing the two would render an unproven green.
"""
if snapshot.ready and snapshot.readiness_complete:
headline = "Ready"
elif snapshot.ready:
headline = "Ready (incomplete evidence)"
else:
headline = "Not ready"
return (
"<section class='health-card'>"
f"<h3>Readiness {_status_badge(snapshot.status)}</h3>"
f"<p class='health-headline'>{html.escape(headline)}</p>"
"<table class='detail'>"
f"<tr><th>Service</th><td><code>{_esc(snapshot.service)}</code></td></tr>"
f"<tr><th>Mode</th><td>{_esc(snapshot.mode)}</td></tr>"
f"<tr><th>Ready</th><td>{_esc(snapshot.ready)}</td></tr>"
"<tr><th>Readiness evidence complete</th>"
f"<td>{_esc(snapshot.readiness_complete)}</td></tr>"
"<tr><th>Deep probes requested</th>"
f"<td>{_esc(snapshot.deep_probes_requested)}</td></tr>"
f"<tr><th>Observed at</th><td><code>{_esc(snapshot.timestamp)}</code></td></tr>"
"</table>"
"<h4>Readiness reasons</h4>"
f"{_reason_list(snapshot.readiness_reasons, empty='No readiness objections recorded.')}"
"</section>"
)
def _version_card(snapshot: SystemHealthSnapshot) -> str:
version = snapshot.version
uptime_hours = snapshot.uptime_seconds / 3600.0
known = (
"resolved"
if version.known
else "unresolved — version fields could not be read from the checkout"
)
schema = version.control_plane_schema_version
return (
"<section class='health-card'>"
"<h3>Version and uptime</h3>"
"<table class='detail'>"
f"<tr><th>Git SHA</th><td><code>{_esc(version.git_sha or 'unknown')}</code></td></tr>"
"<tr><th>Git describe</th>"
f"<td><code>{_esc(version.git_describe or 'unknown')}</code></td></tr>"
"<tr><th>Control-plane schema</th>"
f"<td>{_esc(schema if schema is not None else 'unknown')}</td></tr>"
f"<tr><th>Python</th><td><code>{_esc(version.python_version)}</code></td></tr>"
f"<tr><th>Version status</th><td>{html.escape(known)}</td></tr>"
f"<tr><th>Started at</th><td><code>{_esc(snapshot.started_at)}</code></td></tr>"
"<tr><th>Uptime</th>"
f"<td>{snapshot.uptime_seconds:.3f}s ({uptime_hours:.2f}h)</td></tr>"
"</table>"
"</section>"
)
def _dependency_rows(probes: tuple[DependencyProbe, ...]) -> str:
if not probes:
return "<p class='muted'>No dependency probes were reported.</p>"
rows = []
for probe in probes:
latency = (
f"{probe.latency_ms:.1f} ms" if probe.latency_ms is not None else "n/a"
)
rows.append(
"<tr>"
f"<td><code>{_esc(probe.name)}</code></td>"
f"<td>{_esc(probe.kind)}</td>"
f"<td>{_status_badge(probe.status)}</td>"
f"<td>{_esc('required' if probe.required else 'optional')}</td>"
f"<td>{html.escape(latency)}</td>"
f"<td>{_safe(probe.detail)}</td>"
"</tr>"
)
return (
"<table class='registry'><thead><tr>"
"<th>Dependency</th><th>Kind</th><th>Status</th><th>Requirement</th>"
"<th>Latency</th><th>Detail</th>"
"</tr></thead><tbody>"
f"{''.join(rows)}</tbody></table>"
)
def _dependency_card(snapshot: SystemHealthSnapshot) -> str:
degraded = [probe for probe in snapshot.dependencies if probe.ran and not probe.healthy]
not_run = [probe for probe in snapshot.dependencies if not probe.ran]
banner = ""
if degraded:
names = ", ".join(sorted(probe.name for probe in degraded))
banner += (
"<div class='stub health-degraded'><p><strong>Degraded dependencies:</strong> "
f"{_esc(names)}</p></div>"
)
if not_run:
names = ", ".join(sorted(probe.name for probe in not_run))
banner += (
"<div class='stub'><p><strong>Not probed:</strong> "
f"{_esc(names)} — these contribute no evidence and are not counted "
"as healthy.</p></div>"
)
return (
"<section class='health-card'>"
"<h3>Dependencies</h3>"
f"{banner}"
f"{_dependency_rows(snapshot.dependencies)}"
"<p class='muted'>Details are redacted at the API boundary and again "
"before rendering; credentials are never displayed.</p>"
"</section>"
)
def _namespace_card(snapshot: SystemHealthSnapshot) -> str:
if not snapshot.mcp_namespaces:
body = "<p class='muted'>No MCP namespaces are declared.</p>"
else:
rows = []
for entry in snapshot.mcp_namespaces:
rows.append(
"<tr>"
f"<td><code>{_esc(entry.get('namespace'))}</code></td>"
f"<td><code>{_esc(entry.get('required_tool'))}</code></td>"
f"<td>{_status_badge(str(entry.get('status') or STATUS_UNPROVEN))}</td>"
f"<td>{_esc(entry.get('ide_namespace_proven'))}</td>"
f"<td>{_safe(entry.get('reason'))}</td>"
"</tr>"
)
body = (
"<table class='registry'><thead><tr>"
"<th>Namespace</th><th>Required tool</th><th>Status</th>"
"<th>IDE-proven</th><th>Reason</th>"
"</tr></thead><tbody>"
f"{''.join(rows)}</tbody></table>"
)
return (
"<section class='health-card'>"
"<h3>MCP namespaces</h3>"
f"{body}"
"<p class='muted'>The web process runs outside the IDE-managed MCP "
"client, so namespace health is reported as unproven rather than "
"guessed (#543).</p>"
"</section>"
)
def _stale_runtime_card(snapshot: SystemHealthSnapshot) -> str:
stale = snapshot.stale_runtime
if stale.stale:
warning = (
"<div class='stub health-stale'><p><strong>Stale runtime:</strong> "
"the running code, the checkout, and the remote-tracking commit "
"disagree. Capability gates may be evaluating obsolete code — "
"do not treat this runtime as mutation-safe.</p></div>"
)
elif not stale.determinable:
warning = (
"<div class='stub health-stale'><p><strong>Staleness "
"indeterminate:</strong> parity could not be proven, so this "
"runtime is not reported as mutation-safe.</p></div>"
)
else:
warning = ""
return (
"<section class='health-card'>"
"<h3>Stale-runtime parity</h3>"
f"{warning}"
"<table class='detail'>"
"<tr><th>Daemon head</th>"
f"<td><code>{_esc(stale.daemon_head or 'unknown')}</code></td></tr>"
"<tr><th>Checkout head</th>"
f"<td><code>{_esc(stale.checkout_head or 'unknown')}</code></td></tr>"
"<tr><th>Remote head</th>"
f"<td><code>{_esc(stale.remote_head or 'unknown')}</code></td></tr>"
f"<tr><th>Stale</th><td>{_esc(stale.stale)}</td></tr>"
f"<tr><th>Determinable</th><td>{_esc(stale.determinable)}</td></tr>"
f"<tr><th>Mutation safe</th><td>{_esc(stale.mutation_safe)}</td></tr>"
"</table>"
f"{_reason_list(stale.reasons, empty='Runtime, checkout, and remote agree.')}"
"</section>"
)
def _probe_error_card(snapshot: SystemHealthSnapshot) -> str:
if not snapshot.probe_errors:
return ""
return (
"<section class='health-card'>"
"<h3>Probe errors</h3>"
f"{_reason_list(snapshot.probe_errors, empty='')}"
"</section>"
)
def _recovery_card() -> str:
"""Sanctioned recovery pointers only — never a manual process kill (#630)."""
return (
"<section class='health-card'>"
"<h3>Recovery</h3>"
"<p class='muted'>This dashboard is read-only. Restart and reload "
"controls arrive in Phase 2 (#642); until then recovery runs through "
"the sanctioned client reconnect / operator restart path.</p>"
"<ul class='reasons'>"
"<li><a href='/runtime'>Runtime and session view</a> — active profile, "
"workflow hashes, and shell health.</li>"
"<li>Reconnect the MCP client from the IDE, then re-run the blocked "
"cycle. Never kill the daemon process manually: unmanaged kills are "
"recorded as runtime contamination (#630).</li>"
"<li>See <code>docs/webui-local-dev.md</code> for the documented "
"recovery sequence.</li>"
"</ul>"
"</section>"
)
def render_system_health_page(snapshot: SystemHealthSnapshot) -> str:
"""Render the full system-health dashboard body."""
return (
"<h2>System health</h2>"
"<p class='meta'>Read-only view of the Phase 1 system-health API "
"(<code>/api/v1/system/health</code>). Reload this page to refresh; "
"nothing here polls or mutates on your behalf.</p>"
f"{_readiness_card(snapshot)}"
f"{_stale_runtime_card(snapshot)}"
f"{_version_card(snapshot)}"
f"{_dependency_card(snapshot)}"
f"{_namespace_card(snapshot)}"
f"{_probe_error_card(snapshot)}"
f"{_recovery_card()}"
)