Compare commits

..
Author SHA1 Message Date
sysadminandGrok 4.5 461e1dac78 feat(mcp): enforce scoped recovery playbook before full restart (Closes #669)
Add recovery_playbook.py with the narrow-to-broad recovery ladder, symptom
routing, attempt-log helpers, and escalation metrics. Wire the attempt-log
gate into restart_coordinator so rolling/full/host restarts require prior
insufficient narrower attempts (or break-glass). Document the ladder and
update gitea_request_mcp_restart for prior_recovery_attempts_json.

Co-Authored-By: Grok 4.5 <[email protected]>
2026-07-25 17:35:11 -04:00
sysadmin 9c69bfcd80 Merge pull request 'feat(webui): Gitea issue and PR linkage console (Closes #645)' (#904) from feat/issue-645-linkage-console into master 2026-07-25 16:32:21 -05:00
sysadminandClaude Opus 4.8 211890f361 feat(webui): Gitea issue and PR linkage console (Closes #645)
Add a Phase 3 read-only console that resolves issue↔PR linkage with
evidence (closes keyword, branch marker, body mention), surfaces the
latest canonical handoff for a focused thread, and deep-links to Gitea
only under the admin reveal opt-in.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-07-25 16:40:07 -04:00
18 changed files with 2788 additions and 1489 deletions
+94
View File
@@ -0,0 +1,94 @@
# MCP scoped recovery playbook (#669)
**Parent:** [#655](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/655)
**Vision / roadmap:** [#652](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/652) · [#653](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/653)
**Class matrix:** [#663](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/663) · `docs/mcp-restart-classes.md`
**Coordinator:** [#658](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/658) · `restart_coordinator.py`
**Audit lineage:** [#665](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/665)
## Decision
Full-server MCP reset is a **last resort**. Prefer the narrowest recovery that
can clear the symptom. The coordinator **refuses** `rolling_mcp_restart`,
`full_mcp_restart`, and `host_restart` unless:
1. The inventory carries a prior **attempt log** of at least one *insufficient*
narrower recovery, **or**
2. **Break-glass** is authorized
(`request_break_glass` + `GITEA_BREAKGLASS_RESTART_AUTHORIZATION`).
Break-glass still never bypasses the #663 class matrix (role/permission).
## Ladder (narrow → broad)
| Rank | Action | Self-service | Implementation / delegation |
|---:|---|---|---|
| 0 | `client_reconnect` | yes | Host auto-reconnect / client reconnect · [#584](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/584) · `docs/mcp-namespace-eof-recovery.md` |
| 1 | `capability_refresh` | yes | `gitea_resolve_task_capability` + `gitea_whoami` · [#610](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/610) · [#685](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/685) |
| 2 | `session_reconnect` | yes | Runtime rebind + explicit `worktree_path` · [#543](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/543) · [#618](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/618) |
| 3 | `configuration_reload` | no | Class `configuration_reload` · console reload · [#642](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/642) |
| 4 | `lease_recovery` | no | Lock/lease recovery paths · [#702](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/702) · [#753](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/753) · [#790](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/790) |
| 5 | `worker_restart` | no | Class `worker_restart` · #663 |
| 6 | `role_runtime_restart` | no | Class `role_runtime_restart` · console restart · #642/#663 |
| 7 | `connector_restart` | no | Class `connector_restart` · #663 |
| 8 | `rolling_mcp_restart` | no | Class `rolling_mcp_restart` · design [#668](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/668) · **attempt log required** |
| 9 | `full_mcp_restart` | no | Class `full_mcp_restart` · **attempt log required** |
| 10 | `host_restart` | no | Class `host_restart` · **attempt log required** |
Machine-readable source of truth: `recovery_playbook.RECOVERY_LADDER` and
`recovery_playbook.ladder_document()`.
## Attempt log shape
Each prior attempt is a mapping:
```json
{
"action": "client_reconnect",
"outcome": "insufficient",
"reason": "transport still closed after IDE reconnect",
"actor": "prgs-controller-12345",
"recorded_at": "2026-07-25T21:00:00+00:00"
}
```
Outcomes that count toward escalation: `failed`, `insufficient`, `denied`,
`unresolved`, `timeout`, `error`.
Pass attempts into the coordinator via inventory
`prior_recovery_attempts` or the MCP tool argument
`prior_recovery_attempts_json` on `gitea_request_mcp_restart`.
Helper: `recovery_playbook.build_attempt_record(...)`.
## Symptom → first rung
`recovery_playbook.recommend_actions(symptoms=[...])` maps symptoms such as
`transport_eof`, `stale_capability`, `stale_lease`, `daemon_corrupt` to the
narrowest recommended action, then walks the ladder. Soft recommendations
never replace the hard gate on broad restarts.
## Enforcement points
1. **`recovery_playbook.assess_escalation`** — pure gate.
2. **`restart_coordinator.evaluate_restart_impact`** — when `restart_class` is
set (policy-enforced path), broad classes require the gate; report fields
`attempt_log_satisfied`, `playbook_escalation`, `break_glass`.
3. **`gitea_request_mcp_restart`** — accepts attempt JSON and env-authorized
break-glass; never restarts a process.
## Metrics
`recovery_playbook.recovery_metrics(attempts)` reports the fraction of
successful recoveries that avoided full/host restart
(`fraction_avoided_full_restart`).
## Non-goals
* HA multi-instance execution (#668 design only here).
* Normalizing `pkill` (#630 contamination stays forbidden).
* Silent mutation of leases or processes from the playbook itself.
## Manual process kills
Remain forbidden and contaminating (#630). The playbook never recommends them.
+15 -5
View File
@@ -95,12 +95,18 @@ gitea_request_mcp_restart(remote, host, org, repo,
target_session_id=None, target_role=None,
target_connector=None,
drain_proof_json=None,
request_break_glass=False)
request_break_glass=False,
prior_recovery_attempts_json=None)
```
It **never restarts anything**: `apply_supported` is always `false` and
`restart_performed` is always `false`.
`prior_recovery_attempts_json` (#669) is an optional JSON array of prior
narrow recovery attempts. Rolling / full / host classes require at least one
*insufficient* narrower attempt (or authorized break-glass). See
`docs/mcp-recovery-playbook.md`.
### Dry-run versus apply
| Call | Behavior |
@@ -112,9 +118,10 @@ It **never restarts anything**: `apply_supported` is always `false` and
An apply requires **both** authorizations, and they are independent:
1. **Restart-class authorization** (#663) — the requester's role and permissions
must allow the requested class, the class's approval requirement must be
satisfied, and any target-scoped class must name its target. Failing any of
1. **Restart-class authorization** (#663 / #669) — the requester's role and
permissions must allow the requested class, the class's approval requirement
must be satisfied, any target-scoped class must name its target, and broad
classes must satisfy the recovery-playbook attempt-log gate. Failing any of
these makes `allow_restart` `false`.
2. **Drain-proof gate** (#661) — a valid, unexpired, clean proof bound to the
current impact fingerprint, or an authorized break-glass.
@@ -127,7 +134,10 @@ the authorization that produced it.
### Break-glass
Break-glass bypasses the **drain proof only** — never the restart-class matrix.
Break-glass bypasses the **drain proof only** — never the restart-class matrix
(role/permission). Separately, authorized break-glass also satisfies the #669
attempt-log requirement for broad restarts (rolling/full/host), because that
gate is not a class-matrix permission check.
It is honoured solely when `request_break_glass` is set *and* the environment
carries `GITEA_BREAKGLASS_RESTART_AUTHORIZATION`; like operator override, the
tool argument expresses caller intent and cannot be self-asserted by a worker
+64
View File
@@ -80,6 +80,8 @@ status, onboarding checklist state, and the fail-closed error payloads (#635).
| `/sessions` | Runtime and session view (#641) — health + inventory sessions/namespaces/worktrees |
| `/api/sessions` | JSON export for the runtime/session view |
| `/api/v1/sessions` | Versioned alias of `/api/sessions` |
| `/gitea` | Gitea issue↔PR linkage console (#645) — both directions, with the evidence for each edge |
| `/api/v1/gitea/linkage` | JSON linkage export; `502` when the read could not be answered |
| `/inventory` | Phase 1 shell stub — unified inventory (backed by #636) |
| `/timeline` | Phase 1 shell stub — workflow event timeline |
| `/policy` | Phase 1 shell stub — capability/role policy placeholder |
@@ -327,6 +329,68 @@ Honesty rules specific to this view:
The write-time redactor is a narrow denylist and is not relied on. The field
itself is kept — it is the `#630` evidence naming which daemon was killed.
## Gitea issue/PR linkage (#645)
`/gitea` is the Phase 3 read-only linkage console: which PR carries which issue,
which issues are claimed by more than one PR, and what the latest Canonical
Thread Handoff on a thread said. Gitea remains the source of truth — this
surface reads it and never writes to it. There is no issue/PR editor, no review,
and no merge control.
Query parameters (all optional):
| Parameter | Meaning |
|-----------|---------|
| `project` | Registry project id to scope the read (default: first registry entry) |
| `state` | `open` (default) or `all`; `all` widens the window to merged/closed items, where a landed edge lives |
| `issue=N` / `pr=N` | Focus one thread and load *its* latest canonical handoff |
`GET /api/v1/gitea/linkage` returns the same model as JSON
(`schema_version: 1`). It answers `502` when the read could not be answered, so
an automated consumer cannot mistake a fail-closed payload for "no links exist".
The HTML page always answers `200` and renders the reason instead — an operator
view must show why a read failed rather than withhold the page.
### How an edge is found
Each edge carries the evidence that produced it, strongest first:
| Evidence | Meaning |
|----------|---------|
| `closes_keyword` | The PR title or body declares `closes/fixes/resolves #N`. Gitea itself acts on this keyword. |
| `branch_marker` | The PR head branch carries the canonical `(fix\|feat\|docs\|chore)/issue-N-…` marker minted by the issue lock. |
| `body_reference` | The PR body mentions `#N` with no closing keyword. A mention is not a claim to close. |
Only closing and branch-marker edges populate the **issue → PR** direction: a
bare mention is a cross-link, and counting it as ownership would invent
contested issues out of ordinary references. The mention stays visible on the
**PR → issue** side, labelled as such. A PR whose two strongest edges tie is
flagged `ambiguous`; an issue claimed by two PRs is flagged `contested`.
### Honesty rules specific to this view
* **A partial read never reads as an absence.** Linkage is a claim about the
loaded window only. When pagination did not complete, every empty edge cell
renders `none found (partial inventory)` rather than `none`, and the JSON
carries `inventory_complete: false` plus per-row `links_authoritative: false`.
* **A failed read renders no table at all.** Missing credentials, an unknown
project, or a fetch error produce `ok: false` with a reason. An empty linkage
table would assert that no issue is linked to any PR, which such a read is not
in a position to claim.
* **Handoffs are loaded, never assumed.** CTH comments are thread-scoped, so
only the focused issue or PR has its comments fetched. Every other row reports
`not_loaded` with the reason; a thread whose comments *were* loaded and carried
no CTH says exactly that. A comment-source failure degrades the handoff alone —
the linkage tables still render.
* **Unrecognised handoff headings are reported, not republished.** A `## CTH:`
heading outside `CTH_TYPES` renders as `unrecognized`.
* **Redaction precedes display.** Titles, labels, handoff fields, and error
reasons pass through `webui.console_redaction` before serialization, and the
page HTML-escapes everything it renders.
* **Deep links are opt-in.** A link out to the Gitea web UI appears only when
`GITEA_MCP_REVEAL_ENDPOINTS=1` is set server-side, matching how the MCP tools
gate URL exposure. Item numbers stay usable without it.
## System-health dashboard (#639)
`/system-health` renders the same snapshot the `/api/v1/system/health` API
-102
View File
@@ -1,102 +0,0 @@
# Web Console: restart status, impact preview, and approval state (#667)
Phase 1 of the console restart surface. It consumes the #655 coordinator
substrate and displays it. It performs no restart, reload, drain, approval, or
process action, and it registers no write endpoint.
Issue #667's rollout is explicit — *status views first, write approval after the
backend gates are green* — and this change delivers only the status half.
## Surfaces
| Path | Method | Purpose |
|------|--------|---------|
| `/runtime/restart` | GET | Restart status page |
| `/api/v1/system/restart/status` | GET | Same snapshot as JSON |
Both accept an optional `restart_class` query parameter (default
`full_mcp_restart`). An unrecognised class is not an error: the coordinator
resolves it as unknown and fails closed, and the page shows the resulting deny.
Neither path accepts `POST`; a write attempt returns `405`, and a test asserts
it.
## What it shows
* **Impact preview (#658)** — verdict, blast radius, affected sessions, leases,
critical sections, mutations, and the counts behind them, evaluated
`dry_run=True` against live control-plane state.
* **Drain proof (#661)** — verification of a supplied proof: valid, clean,
expired, tampered, and the reasons behind a refusal.
* **Post-restart reconcile (#662)** — the most recent completion proof, its
overall status, and which dimensions still require follow-up.
* **Restart classes (#663)** — the least-privilege matrix, with *you may
request* and *you may execute* computed for the viewing role rather than for a
generic operator.
* **Approval controls (#633)** — the authorization state of
`system.restart_namespace` and `system.reload_namespace`.
* **Break-glass (#664)** — declared and marked unavailable; see below.
## Three rules this surface holds itself to
A status page that is wrong is worse than one that is missing, because an
operator acts on it. Three properties are enforced by tests, and each was
verified by reverting the guard and watching a test fail.
### An unreadable source reports unavailable, never green
Every source carries its own `SourceStatus`. Nothing substitutes a default,
placeholder, or self-comparison for a reading that failed. An unreadable
control-plane database yields `inventory_complete: false`, which the coordinator
itself turns into a fail-closed verdict, and the page says the blast radius is
unknown rather than showing an empty affected-sessions table.
An absent drain proof is reported as absent — not as a pass. The #661 gate
authorizes a restart only against a valid, unexpired, clean proof, so no proof
is precisely the state that gate denies on.
### Authorization is asked the way execution would ask it
Every probe passes `for_execution=True`.
Asked without it, an admin is `allowed` for `system.restart_namespace`. On a
control surface that reads as a live button. Asked the way an execution attempt
would ask, the same principal is refused `phase_not_active`, because the console
is in Phase 1 and the action is Phase 2. This surface reports the second answer.
`execution_enabled` is therefore `false` for every action and every role today,
and a test asserts that across the whole role matrix.
### The control-plane database is opened read-only
`ControlPlaneDB()` creates directories and runs migrations on construction — a
write. This surface never constructs one. It opens the sqlite file with
`mode=ro`, exactly as `webui/inventory.py` does, and treats a missing file as
missing authority rather than as an empty inventory.
The test that protects this points at a path inside a directory that already
exists, so a read-write `connect` would really create the file. A nested
missing-directory path would have passed for the wrong reason.
## Break-glass is declared, not offered
The break-glass workflow (#664) is not available on this branch's base. The
panel is rendered to operator-class roles as **unavailable**, naming the issue
that tracks it. It is not silently omitted, because an operator who has been
told a governance path exists needs to see that it is not wired here; and it is
not rendered as a control, because there is nothing behind it.
Unprivileged viewers see only a note that the surface is operator-class.
## Redaction and escaping
Every interpolated value passes through `_esc` (`html.escape(..., quote=True)`).
Free-form text and anything that can carry a filesystem path additionally passes
through `webui.inventory.scrub_text`, which redacts credential-shaped tokens
inside a string rather than only at its start. The impact payload is passed
through `webui.inventory.scrub` before rendering.
## Linkage
Parent #655 · extends #642 · consumes #658, #661, #662, #663 · RBAC #633 ·
console #631 · vision #652 · roadmap #653 · break-glass #664.
+35 -9
View File
@@ -22575,8 +22575,9 @@ def gitea_request_mcp_restart(
target_connector: str | None = None,
drain_proof_json: str | None = None,
request_break_glass: bool = False,
prior_recovery_attempts_json: str | None = None,
) -> dict:
"""Evaluate a proposed MCP restart and return an impact preview (#658).
"""Evaluate a proposed MCP restart and return an impact preview (#658/#669).
Central restart coordinator: resolves the requested restart class, gathers
live control-plane state (sessions,
@@ -22600,10 +22601,16 @@ def gitea_request_mcp_restart(
independent the drain gate proves the blast radius was drained and knows
nothing about whether this requester may request this class so a class the
matrix denied never reports an authorized apply. Break-glass bypasses the
drain proof only; it never bypasses the class matrix. ``apply_gate`` carries
drain proof and, when env-authorized, the #669 attempt-log requirement for
broad restarts; it never bypasses the class matrix. ``apply_gate`` carries
``drain_gate_allow`` and ``restart_class_authorized`` so a denial is
attributable to the authorization that produced it.
``prior_recovery_attempts_json`` (#669) is an optional JSON array of prior
narrow recovery attempts ``{action, outcome, reason, ...}``. Rolling / full
/ host restart classes require at least one *insufficient* narrower attempt
unless break-glass is authorized.
Operator override authority is read from the process environment
(``GITEA_OPERATOR_RESTART_OVERRIDE_AUTHORIZATION``), never self-asserted by
the requesting session: ``request_override`` only expresses caller intent
@@ -22709,12 +22716,37 @@ def gitea_request_mcp_restart(
requester_role
)
prior_recovery_attempts: list[dict] = []
if prior_recovery_attempts_json:
try:
parsed_attempts = json.loads(prior_recovery_attempts_json)
if isinstance(parsed_attempts, list):
prior_recovery_attempts = [
dict(a) for a in parsed_attempts if isinstance(a, dict)
]
else:
incomplete_reasons.append(
"prior_recovery_attempts_json must be a JSON array (#669)"
)
inventory_complete = False
except (ValueError, TypeError) as exc:
incomplete_reasons.append(
f"invalid prior_recovery_attempts_json: {_redact(str(exc))}"
)
inventory_complete = False
break_glass_authorized = bool(
(os.environ.get("GITEA_BREAKGLASS_RESTART_AUTHORIZATION") or "").strip()
)
break_glass = bool(request_break_glass and break_glass_authorized)
inventory = {
"sessions": sessions,
"leases": leases,
"terminal_lock": terminal_lock,
"inventory_complete": inventory_complete,
"incomplete_reasons": incomplete_reasons,
"prior_recovery_attempts": prior_recovery_attempts,
}
report = restart_coordinator.evaluate_restart_impact(
@@ -22730,6 +22762,7 @@ def gitea_request_mcp_restart(
target_session_id=target_session_id,
target_role=target_role,
target_connector=target_connector,
break_glass=break_glass,
)
payload = report.as_dict()
@@ -22762,13 +22795,6 @@ def gitea_request_mcp_restart(
except (ValueError, TypeError) as exc:
proof_parse_error = f"invalid drain_proof_json: {_redact(str(exc))}"
break_glass_authorized = bool(
(
os.environ.get("GITEA_BREAKGLASS_RESTART_AUTHORIZATION") or ""
).strip()
)
break_glass = bool(request_break_glass and break_glass_authorized)
expected_fp = drain_proof.impact_fingerprint(report.as_dict())
gate = drain_proof.gate_apply_restart(
proof=proof_obj,
+583
View File
@@ -0,0 +1,583 @@
"""Scoped MCP recovery playbook (#669).
Operational recovery must prefer the *narrowest* action that can fix the
symptom. Full MCP / host restarts are last-resort rungs on a documented
ladder; the coordinator refuses those rungs unless a prior attempt log
shows narrower recoveries already failed (or break-glass is authorized).
This module is pure classification and recommendation:
* No network, filesystem, or process I/O.
* Never restarts anything.
* Narrow recovery *execution* is delegated to existing tools/docs (linked
per rung) — the playbook records which rung to try next and whether
escalation to a broad restart is allowed.
Design lineage: umbrella #655, class matrix #663, coordinator #658,
auto-reconnect #584, stale-runtime #610, contamination #630, audit #665.
Vision #652 / roadmap #653.
"""
from __future__ import annotations
from dataclasses import dataclass, field
from datetime import datetime, timezone
from enum import Enum
from typing import Any, Mapping, Sequence
PLAYBOOK_VERSION = "1.0.0-issue-669"
# Attempt outcomes that count as "tried and insufficient" for escalation.
INSUFFICIENT_OUTCOMES = frozenset(
{
"failed",
"insufficient",
"denied",
"unresolved",
"timeout",
"error",
}
)
# Break-glass / operator override still records that the ladder was skipped.
OUTCOME_BREAK_GLASS = "break_glass"
OUTCOME_SUCCESS = "success"
OUTCOME_SKIPPED = "skipped"
class RecoveryAction(str, Enum):
"""Ordered recovery ladder (narrow → broad)."""
CLIENT_RECONNECT = "client_reconnect"
CAPABILITY_REFRESH = "capability_refresh"
SESSION_RECONNECT = "session_reconnect"
CONFIGURATION_RELOAD = "configuration_reload"
LEASE_RECOVERY = "lease_recovery"
WORKER_RESTART = "worker_restart"
ROLE_RUNTIME_RESTART = "role_runtime_restart"
CONNECTOR_RESTART = "connector_restart"
ROLLING_MCP_RESTART = "rolling_mcp_restart"
FULL_MCP_RESTART = "full_mcp_restart"
HOST_RESTART = "host_restart"
# Classes that require a prior narrow-attempt log (unless break-glass).
BROAD_RESTART_ACTIONS: frozenset[RecoveryAction] = frozenset(
{
RecoveryAction.ROLLING_MCP_RESTART,
RecoveryAction.FULL_MCP_RESTART,
RecoveryAction.HOST_RESTART,
}
)
# Map #663 restart_class strings onto playbook actions.
RESTART_CLASS_TO_ACTION: dict[str, RecoveryAction] = {
"client_reconnect": RecoveryAction.CLIENT_RECONNECT,
"session_reconnect": RecoveryAction.SESSION_RECONNECT,
"configuration_reload": RecoveryAction.CONFIGURATION_RELOAD,
"worker_restart": RecoveryAction.WORKER_RESTART,
"role_runtime_restart": RecoveryAction.ROLE_RUNTIME_RESTART,
"connector_restart": RecoveryAction.CONNECTOR_RESTART,
"rolling_mcp_restart": RecoveryAction.ROLLING_MCP_RESTART,
"full_mcp_restart": RecoveryAction.FULL_MCP_RESTART,
"host_restart": RecoveryAction.HOST_RESTART,
}
@dataclass(frozen=True)
class RecoveryRung:
"""One rung on the recovery ladder."""
action: RecoveryAction
rank: int
summary: str
# Existing implementation or explicit delegation target.
implementation: str
issue_links: tuple[str, ...]
self_service: bool
# Restart-class permission when this rung is requested via coordinator.
restart_class: str | None = None
def as_dict(self) -> dict[str, Any]:
return {
"action": self.action.value,
"rank": self.rank,
"summary": self.summary,
"implementation": self.implementation,
"issue_links": list(self.issue_links),
"self_service": self.self_service,
"restart_class": self.restart_class,
}
# Canonical ladder. Rank 0 is narrowest.
RECOVERY_LADDER: tuple[RecoveryRung, ...] = (
RecoveryRung(
RecoveryAction.CLIENT_RECONNECT,
0,
"Reconnect the IDE/client MCP transport (EOF / transport flap).",
"Host auto-reconnect or explicit client reconnect; "
"docs/mcp-namespace-eof-recovery.md",
("#584", "#655"),
True,
"client_reconnect",
),
RecoveryRung(
RecoveryAction.CAPABILITY_REFRESH,
1,
"Re-resolve task capability and clear stale permission context.",
"Delegated: gitea_resolve_task_capability + gitea_whoami "
"(no process change).",
("#610", "#685", "#655"),
True,
None,
),
RecoveryRung(
RecoveryAction.SESSION_RECONNECT,
2,
"Rebind identity, workspace, and namespace for one session.",
"Delegated: gitea_get_runtime_context + explicit worktree_path "
"rebind (#618); docs/mcp-namespace-health.md",
("#543", "#618", "#655"),
True,
"session_reconnect",
),
RecoveryRung(
RecoveryAction.CONFIGURATION_RELOAD,
3,
"Gracefully reload configuration without replacing the daemon.",
"restart_coordinator class configuration_reload; console "
"system.reload_namespace (#642).",
("#642", "#663", "#655"),
False,
"configuration_reload",
),
RecoveryRung(
RecoveryAction.LEASE_RECOVERY,
4,
"Recover or rebind stale leases/locks without a process restart.",
"Delegated: issue lock recovery / lease lifecycle paths "
"(#702, #753, #790).",
("#702", "#753", "#790", "#655"),
False,
None,
),
RecoveryRung(
RecoveryAction.WORKER_RESTART,
5,
"Restart one worker after its own lease and mutation scope drains.",
"restart_coordinator class worker_restart (#663).",
("#663", "#655"),
False,
"worker_restart",
),
RecoveryRung(
RecoveryAction.ROLE_RUNTIME_RESTART,
6,
"Restart one role runtime and re-probe that namespace only.",
"restart_coordinator class role_runtime_restart; console "
"system.restart_namespace (#642).",
("#642", "#663", "#655"),
False,
"role_runtime_restart",
),
RecoveryRung(
RecoveryAction.CONNECTOR_RESTART,
7,
"Restart one connector while unrelated runtimes stay available.",
"restart_coordinator class connector_restart (#663).",
("#663", "#655"),
False,
"connector_restart",
),
RecoveryRung(
RecoveryAction.ROLLING_MCP_RESTART,
8,
"Drain/restart/verify one instance at a time (HA path).",
"restart_coordinator class rolling_mcp_restart; design #668.",
("#668", "#663", "#655"),
False,
"rolling_mcp_restart",
),
RecoveryRung(
RecoveryAction.FULL_MCP_RESTART,
9,
"Full stable-control MCP process restart after verified full drain.",
"restart_coordinator class full_mcp_restart; requires attempt log "
"unless break-glass (#669).",
("#658", "#661", "#663", "#669", "#655"),
False,
"full_mcp_restart",
),
RecoveryRung(
RecoveryAction.HOST_RESTART,
10,
"Host/infrastructure restart — broadest last-resort action.",
"restart_coordinator class host_restart; operator-owned.",
("#663", "#669", "#655"),
False,
"host_restart",
),
)
_LADDER_BY_ACTION: dict[RecoveryAction, RecoveryRung] = {
rung.action: rung for rung in RECOVERY_LADDER
}
# Symptom tokens → preferred first rung (decision tree, #663 lineage).
SYMPTOM_TO_FIRST_ACTION: dict[str, RecoveryAction] = {
"transport_eof": RecoveryAction.CLIENT_RECONNECT,
"client_closing_eof": RecoveryAction.CLIENT_RECONNECT,
"transport_flap": RecoveryAction.CLIENT_RECONNECT,
"namespace_disconnected": RecoveryAction.CLIENT_RECONNECT,
"stale_capability": RecoveryAction.CAPABILITY_REFRESH,
"permission_stale": RecoveryAction.CAPABILITY_REFRESH,
"runtime_reconnect_required": RecoveryAction.CAPABILITY_REFRESH,
"stale_runtime": RecoveryAction.SESSION_RECONNECT,
"worktree_unbound": RecoveryAction.SESSION_RECONNECT,
"namespace_unhealthy": RecoveryAction.SESSION_RECONNECT,
"config_drift": RecoveryAction.CONFIGURATION_RELOAD,
"profile_misbound": RecoveryAction.CONFIGURATION_RELOAD,
"stale_lease": RecoveryAction.LEASE_RECOVERY,
"dead_pid_lock": RecoveryAction.LEASE_RECOVERY,
"orphan_worktree": RecoveryAction.LEASE_RECOVERY,
"single_worker_stuck": RecoveryAction.WORKER_RESTART,
"role_runtime_dead": RecoveryAction.ROLE_RUNTIME_RESTART,
"connector_dead": RecoveryAction.CONNECTOR_RESTART,
"ha_instance_unhealthy": RecoveryAction.ROLLING_MCP_RESTART,
"daemon_corrupt": RecoveryAction.FULL_MCP_RESTART,
"full_process_deadlock": RecoveryAction.FULL_MCP_RESTART,
"host_unresponsive": RecoveryAction.HOST_RESTART,
}
def _utc_now() -> datetime:
return datetime.now(timezone.utc)
def resolve_action(value: RecoveryAction | str) -> RecoveryAction:
"""Resolve a recovery action or fail closed for unknown values."""
if isinstance(value, RecoveryAction):
return value
text = str(value or "").strip()
# Accept #663 restart_class aliases.
if text in RESTART_CLASS_TO_ACTION:
return RESTART_CLASS_TO_ACTION[text]
try:
return RecoveryAction(text)
except ValueError as exc:
raise ValueError(
f"unknown recovery action {value!r}; deny (fail closed, #669)"
) from exc
def ladder_rank(action: RecoveryAction | str) -> int:
resolved = resolve_action(action)
return _LADDER_BY_ACTION[resolved].rank
def rung_for(action: RecoveryAction | str) -> RecoveryRung:
return _LADDER_BY_ACTION[resolve_action(action)]
def normalize_attempt(raw: Mapping[str, Any]) -> dict[str, Any] | None:
"""Normalize one prior-recovery attempt record; return None if unusable."""
if not isinstance(raw, Mapping):
return None
action_raw = raw.get("action") or raw.get("recovery_action") or raw.get(
"restart_class"
)
if not action_raw:
return None
try:
action = resolve_action(str(action_raw))
except ValueError:
return None
outcome = str(
raw.get("outcome") or raw.get("status") or raw.get("result") or ""
).strip().lower()
if not outcome:
return None
recorded_at = raw.get("recorded_at") or raw.get("at") or raw.get("timestamp")
reason = str(raw.get("reason") or raw.get("detail") or "").strip()
actor = str(raw.get("actor") or raw.get("session_id") or "").strip()
return {
"action": action.value,
"outcome": outcome,
"reason": reason,
"actor": actor,
"recorded_at": recorded_at,
"rank": ladder_rank(action),
"raw": dict(raw),
}
def normalize_attempt_log(
attempts: Sequence[Mapping[str, Any]] | None,
) -> list[dict[str, Any]]:
"""Return usable attempt records in ladder order."""
out: list[dict[str, Any]] = []
for raw in attempts or ():
norm = normalize_attempt(raw)
if norm is not None:
out.append(norm)
out.sort(key=lambda a: (a["rank"], str(a.get("recorded_at") or "")))
return out
def narrower_insufficient_attempts(
attempts: Sequence[Mapping[str, Any]] | None,
*,
requested: RecoveryAction | str,
) -> list[dict[str, Any]]:
"""Return prior attempts narrower than *requested* that were insufficient."""
target_rank = ladder_rank(requested)
usable = []
for attempt in normalize_attempt_log(attempts):
if attempt["rank"] >= target_rank:
continue
if attempt["outcome"] in INSUFFICIENT_OUTCOMES:
usable.append(attempt)
return usable
@dataclass(frozen=True)
class EscalationAssessment:
"""Whether a requested broad recovery may proceed given the attempt log."""
requested_action: str
allowed: bool
require_attempt_log: bool
break_glass: bool
reasons: list[str] = field(default_factory=list)
qualifying_attempts: list[dict[str, Any]] = field(default_factory=list)
recommended_next: list[dict[str, Any]] = field(default_factory=list)
playbook_version: str = PLAYBOOK_VERSION
def as_dict(self) -> dict[str, Any]:
return {
"playbook_version": self.playbook_version,
"requested_action": self.requested_action,
"allowed": self.allowed,
"require_attempt_log": self.require_attempt_log,
"break_glass": self.break_glass,
"reasons": list(self.reasons),
"qualifying_attempts": list(self.qualifying_attempts),
"recommended_next": list(self.recommended_next),
}
def assess_escalation(
requested: RecoveryAction | str,
*,
prior_recovery_attempts: Sequence[Mapping[str, Any]] | None = None,
break_glass: bool = False,
) -> EscalationAssessment:
"""Gate broad restarts on a prior narrow-attempt log (#669 AC3).
Narrow / mid-ladder actions do not require a prior attempt log.
``full_mcp_restart``, ``host_restart``, and ``rolling_mcp_restart``
require at least one *insufficient* narrower attempt unless
``break_glass`` is true.
"""
action = resolve_action(requested)
require_log = action in BROAD_RESTART_ACTIONS
reasons: list[str] = []
qualifying = narrower_insufficient_attempts(
prior_recovery_attempts, requested=action
)
if not require_log:
return EscalationAssessment(
requested_action=action.value,
allowed=True,
require_attempt_log=False,
break_glass=bool(break_glass),
reasons=["narrow recovery; attempt log not required"],
qualifying_attempts=qualifying,
recommended_next=[],
)
if break_glass:
return EscalationAssessment(
requested_action=action.value,
allowed=True,
require_attempt_log=True,
break_glass=True,
reasons=[
"break-glass authorized; broad restart permitted without "
"narrow-attempt log (#669)"
],
qualifying_attempts=qualifying,
recommended_next=[],
)
if qualifying:
return EscalationAssessment(
requested_action=action.value,
allowed=True,
require_attempt_log=True,
break_glass=False,
reasons=[
f"{len(qualifying)} narrower recovery attempt(s) recorded as "
"insufficient; escalation permitted"
],
qualifying_attempts=qualifying,
recommended_next=[],
)
# Deny: recommend the next untried narrow rung(s).
recommended = recommend_actions(
symptoms=(),
prior_recovery_attempts=prior_recovery_attempts,
max_actions=3,
)
reasons.append(
f"{action.value} requires a prior attempt log of insufficient "
"narrower recoveries (or break-glass); none found — deny (fail "
"closed, #669)"
)
return EscalationAssessment(
requested_action=action.value,
allowed=False,
require_attempt_log=True,
break_glass=False,
reasons=reasons,
qualifying_attempts=[],
recommended_next=recommended.get("recommended_actions") or [],
)
def recommend_actions(
*,
symptoms: Sequence[str] = (),
prior_recovery_attempts: Sequence[Mapping[str, Any]] | None = None,
max_actions: int = 5,
) -> dict[str, Any]:
"""Return ordered recommended recovery actions for the given symptoms.
Soft mode (rollout): recommendations only — callers decide whether to
hard-gate. Hard mode for broad restarts is :func:`assess_escalation`.
"""
attempted_success = {
a["action"]
for a in normalize_attempt_log(prior_recovery_attempts)
if a["outcome"] == OUTCOME_SUCCESS
}
attempted_any = {
a["action"] for a in normalize_attempt_log(prior_recovery_attempts)
}
first_actions: list[RecoveryAction] = []
for symptom in symptoms:
key = str(symptom or "").strip().lower().replace(" ", "_").replace("-", "_")
mapped = SYMPTOM_TO_FIRST_ACTION.get(key)
if mapped is not None and mapped not in first_actions:
first_actions.append(mapped)
# Default entry: client reconnect then walk the ladder.
if not first_actions:
first_actions = [RecoveryAction.CLIENT_RECONNECT]
recommended: list[dict[str, Any]] = []
seen: set[str] = set()
min_rank = min(ladder_rank(a) for a in first_actions)
for rung in RECOVERY_LADDER:
if rung.rank < min_rank:
continue
if rung.action.value in attempted_success:
continue
if rung.action.value in seen:
continue
# Prefer rungs not yet attempted; still list previously-failed ones
# only if nothing else remains.
entry = rung.as_dict()
entry["already_attempted"] = rung.action.value in attempted_any
recommended.append(entry)
seen.add(rung.action.value)
if len(recommended) >= max(1, int(max_actions)):
break
return {
"playbook_version": PLAYBOOK_VERSION,
"symptoms": [str(s) for s in symptoms],
"recommended_actions": recommended,
"ladder": [r.as_dict() for r in RECOVERY_LADDER],
"read_only": True,
"hard_gate_note": (
"Broad restarts (rolling/full/host) still require "
"assess_escalation / coordinator attempt-log enforcement."
),
}
def build_attempt_record(
action: RecoveryAction | str,
*,
outcome: str,
reason: str = "",
actor: str = "",
recorded_at: str | None = None,
extra: Mapping[str, Any] | None = None,
) -> dict[str, Any]:
"""Build a durable-shaped attempt log entry for inventory/audit (#665)."""
resolved = resolve_action(action)
record = {
"action": resolved.value,
"outcome": str(outcome or "").strip().lower(),
"reason": str(reason or "").strip(),
"actor": str(actor or "").strip(),
"recorded_at": recorded_at or _utc_now().isoformat(),
"rank": ladder_rank(resolved),
"playbook_version": PLAYBOOK_VERSION,
}
if extra:
record["extra"] = dict(extra)
return record
def recovery_metrics(
attempts: Sequence[Mapping[str, Any]] | None,
) -> dict[str, Any]:
"""Compute the fraction of recoveries that avoided full/host restart.
A recovery *episode* is approximated as one attempt with
``outcome=success``. Successes on non-broad rungs count as avoided full
restart; successes on full/host count as full-restart recoveries.
"""
norms = normalize_attempt_log(attempts)
successes = [a for a in norms if a["outcome"] == OUTCOME_SUCCESS]
broad_success = [
a
for a in successes
if resolve_action(a["action"])
in {RecoveryAction.FULL_MCP_RESTART, RecoveryAction.HOST_RESTART}
]
avoided = [a for a in successes if a not in broad_success]
total = len(successes)
fraction_avoided = (len(avoided) / total) if total else None
return {
"playbook_version": PLAYBOOK_VERSION,
"attempts_total": len(norms),
"successes_total": total,
"successes_avoided_full_restart": len(avoided),
"successes_full_or_host_restart": len(broad_success),
"fraction_avoided_full_restart": fraction_avoided,
"insufficient_attempts": sum(
1 for a in norms if a["outcome"] in INSUFFICIENT_OUTCOMES
),
}
def ladder_document() -> dict[str, Any]:
"""Machine-readable ladder for docs/tools inventory."""
return {
"playbook_version": PLAYBOOK_VERSION,
"parent_issues": ["#655", "#652", "#653"],
"enforcement_issue": "#669",
"ladder": [r.as_dict() for r in RECOVERY_LADDER],
"broad_restart_actions": [a.value for a in sorted(BROAD_RESTART_ACTIONS, key=lambda x: x.value)],
"insufficient_outcomes": sorted(INSUFFICIENT_OUTCOMES),
"symptom_map": {k: v.value for k, v in sorted(SYMPTOM_TO_FIRST_ACTION.items())},
}
+46 -2
View File
@@ -1,4 +1,4 @@
"""MCP restart coordinator and impact analysis (#658).
"""MCP restart coordinator and impact analysis (#658 / #669).
Before any sanctioned MCP restart, a central coordinator must evaluate the
live control-plane state — active sessions, leases/locks, in-flight issue/PR
@@ -16,6 +16,9 @@ Design rules (mirrors the read-only posture of ``workflow_dashboard`` /
a mutative apply path is a later child gated by a drain proof (non-goal here).
* **Fail closed.** If the inventory is not explicitly complete, the verdict is
``unsafe`` / deny — an incomplete evaluation must never green-light a restart.
* **Narrow-first (#669).** Broad classes (rolling / full / host) require a
prior attempt log of insufficient narrower recoveries unless break-glass is
authorized. See :mod:`recovery_playbook`.
* **No secrets.** Session ids, pids, and profiles are operational metadata, not
credentials; nothing secret flows through this module.
@@ -32,8 +35,9 @@ from enum import Enum
from typing import Any, Mapping, Sequence
import lease_lifecycle
import recovery_playbook
COORDINATOR_VERSION = "1.1.0-issue-663"
COORDINATOR_VERSION = "1.2.0-issue-669"
# Restart verdicts. Exactly the three the acceptance criteria name.
VERDICT_SAFE = "safe"
@@ -349,6 +353,10 @@ class RestartImpactReport:
counts: dict[str, int]
audit_record: dict[str, Any]
incomplete_reasons: list[str] = field(default_factory=list)
# #669 playbook escalation gate (attempt-log enforcement).
playbook_escalation: dict[str, Any] = field(default_factory=dict)
attempt_log_satisfied: bool = True
break_glass: bool = False
def as_dict(self) -> dict[str, Any]:
return {
@@ -382,6 +390,9 @@ class RestartImpactReport:
"prior_recovery_attempts": list(self.prior_recovery_attempts),
"counts": dict(self.counts),
"audit_record": dict(self.audit_record),
"playbook_escalation": dict(self.playbook_escalation),
"attempt_log_satisfied": self.attempt_log_satisfied,
"break_glass": self.break_glass,
}
@@ -496,6 +507,7 @@ def evaluate_restart_impact(
target_session_id: str | None = None,
target_role: str | None = None,
target_connector: str | None = None,
break_glass: bool = False,
) -> RestartImpactReport:
"""Evaluate a proposed MCP restart and return an impact preview.
@@ -584,6 +596,30 @@ def evaluate_restart_impact(
dict(a) for a in (inventory.get("prior_recovery_attempts") or [])
]
# #669: broad restarts require a prior narrow-attempt log unless break-glass.
playbook_escalation: dict[str, Any] = {}
attempt_log_satisfied = True
if policy_enforced and resolved_class is not None:
try:
escalation = recovery_playbook.assess_escalation(
resolved_class.value,
prior_recovery_attempts=prior_recovery_attempts,
break_glass=bool(break_glass),
)
playbook_escalation = escalation.as_dict()
attempt_log_satisfied = bool(escalation.allowed)
if not attempt_log_satisfied:
authorization_reasons.extend(list(escalation.reasons))
except ValueError as exc:
# Unknown mapping should never happen for enum values; fail closed.
attempt_log_satisfied = False
playbook_escalation = {
"allowed": False,
"reasons": [str(exc)],
"playbook_version": recovery_playbook.PLAYBOOK_VERSION,
}
authorization_reasons.append(str(exc))
session_impacts = [
_classify_session(
s,
@@ -682,6 +718,7 @@ def evaluate_restart_impact(
and role_authorized
and approval_satisfied
and target_complete
and attempt_log_satisfied
)
if policy_enforced and not authorization_ok:
@@ -745,6 +782,7 @@ def evaluate_restart_impact(
"affected_issues": len(affected_issues),
"affected_prs": len(affected_prs),
"prior_recovery_attempts": len(prior_recovery_attempts),
"attempt_log_satisfied": attempt_log_satisfied,
}
audit_record = {
@@ -765,6 +803,9 @@ def evaluate_restart_impact(
"allow_restart": allow_restart,
"blast_radius": blast_radius,
"counts": counts,
"attempt_log_satisfied": attempt_log_satisfied,
"break_glass": bool(break_glass),
"playbook_version": recovery_playbook.PLAYBOOK_VERSION,
}
return RestartImpactReport(
@@ -804,4 +845,7 @@ def evaluate_restart_impact(
counts=counts,
audit_record=audit_record,
incomplete_reasons=incomplete_reasons,
playbook_escalation=playbook_escalation,
attempt_log_satisfied=attempt_log_satisfied,
break_glass=bool(break_glass),
)
@@ -35,6 +35,17 @@ BREAK_GLASS_ENV = "GITEA_BREAKGLASS_RESTART_AUTHORIZATION"
QUIET_SESSIONS: list[dict] = []
QUIET_LEASES: list[dict] = []
# #669: broad restarts need a prior narrow-attempt log (unless break-glass).
PRIOR_NARROW_ATTEMPTS_JSON = json.dumps(
[
{
"action": "client_reconnect",
"outcome": "insufficient",
"reason": "still flapping after reconnect",
}
]
)
class _FakeDB:
"""Minimal control-plane DB stand-in for the restart inventory."""
@@ -128,6 +139,7 @@ class TestConjunction(_RestartToolHarness):
preview = self._call(
role="operator",
restart_class="full_mcp_restart",
prior_recovery_attempts_json=PRIOR_NARROW_ATTEMPTS_JSON,
env={CONTROLLER_APPROVAL_ENV: "operator-approved"},
)
self.assertTrue(preview["allow_restart"],
@@ -136,6 +148,7 @@ class TestConjunction(_RestartToolHarness):
result = self._call(
role="operator",
restart_class="full_mcp_restart",
prior_recovery_attempts_json=PRIOR_NARROW_ATTEMPTS_JSON,
dry_run=False,
drain_proof_json=self._clean_proof_for(preview),
env={CONTROLLER_APPROVAL_ENV: "operator-approved"},
@@ -342,6 +355,7 @@ class TestExistingPathsStillWork(_RestartToolHarness):
result = self._call(
role="operator",
restart_class="full_mcp_restart",
prior_recovery_attempts_json=PRIOR_NARROW_ATTEMPTS_JSON,
dry_run=False,
env={CONTROLLER_APPROVAL_ENV: "operator-approved"},
)
+217
View File
@@ -0,0 +1,217 @@
"""Unit tests for the scoped recovery playbook (#669)."""
from __future__ import annotations
import recovery_playbook as rp
import restart_coordinator as rc
def test_ladder_covers_eleven_ordered_rungs():
ranks = [r.rank for r in rp.RECOVERY_LADDER]
assert ranks == list(range(len(rp.RECOVERY_LADDER)))
assert len(rp.RECOVERY_LADDER) == 11
assert rp.RECOVERY_LADDER[0].action is rp.RecoveryAction.CLIENT_RECONNECT
assert rp.RECOVERY_LADDER[-1].action is rp.RecoveryAction.HOST_RESTART
def test_ladder_document_links_parent_issues():
doc = rp.ladder_document()
assert "#655" in doc["parent_issues"]
assert "#652" in doc["parent_issues"]
assert "#653" in doc["parent_issues"]
assert doc["enforcement_issue"] == "#669"
assert "full_mcp_restart" in doc["broad_restart_actions"]
def test_recommend_transport_eof_starts_at_client_reconnect():
plan = rp.recommend_actions(symptoms=["transport_eof"])
assert plan["recommended_actions"][0]["action"] == "client_reconnect"
assert plan["recommended_actions"][0]["issue_links"]
def test_recommend_skips_successful_prior_attempts():
attempts = [
rp.build_attempt_record(
"client_reconnect", outcome="success", reason="reconnected"
)
]
plan = rp.recommend_actions(
symptoms=["transport_eof"], prior_recovery_attempts=attempts
)
actions = [a["action"] for a in plan["recommended_actions"]]
assert "client_reconnect" not in actions
assert actions[0] == "capability_refresh"
def test_escalation_denied_without_attempt_log():
result = rp.assess_escalation("full_mcp_restart", prior_recovery_attempts=[])
assert result.allowed is False
assert result.require_attempt_log is True
assert any("#669" in r for r in result.reasons)
assert result.recommended_next # soft recommendations still provided
def test_escalation_allowed_after_insufficient_narrower():
attempts = [
rp.build_attempt_record(
"client_reconnect",
outcome="insufficient",
reason="still flapping",
),
rp.build_attempt_record(
"session_reconnect",
outcome="failed",
reason="namespace still dead",
),
]
result = rp.assess_escalation(
"full_mcp_restart", prior_recovery_attempts=attempts
)
assert result.allowed is True
assert len(result.qualifying_attempts) == 2
def test_escalation_break_glass_bypasses_attempt_log():
result = rp.assess_escalation(
"host_restart", prior_recovery_attempts=[], break_glass=True
)
assert result.allowed is True
assert result.break_glass is True
def test_narrow_action_does_not_require_attempt_log():
result = rp.assess_escalation(
"client_reconnect", prior_recovery_attempts=[]
)
assert result.allowed is True
assert result.require_attempt_log is False
def test_same_rank_attempt_does_not_qualify_for_escalation():
attempts = [
rp.build_attempt_record(
"full_mcp_restart", outcome="failed", reason="already failed full"
)
]
result = rp.assess_escalation(
"full_mcp_restart", prior_recovery_attempts=attempts
)
assert result.allowed is False
def test_success_outcome_does_not_qualify_for_escalation():
attempts = [
rp.build_attempt_record(
"client_reconnect", outcome="success", reason="fixed"
)
]
result = rp.assess_escalation(
"full_mcp_restart", prior_recovery_attempts=attempts
)
assert result.allowed is False
def test_recovery_metrics_fraction_avoided():
attempts = [
rp.build_attempt_record("client_reconnect", outcome="success"),
rp.build_attempt_record("session_reconnect", outcome="success"),
rp.build_attempt_record("full_mcp_restart", outcome="success"),
]
metrics = rp.recovery_metrics(attempts)
assert metrics["successes_total"] == 3
assert metrics["successes_avoided_full_restart"] == 2
assert metrics["successes_full_or_host_restart"] == 1
assert abs(metrics["fraction_avoided_full_restart"] - (2 / 3)) < 1e-9
def test_coordinator_denies_full_restart_without_attempt_log():
inv = {
"inventory_complete": True,
"sessions": [],
"leases": [],
"prior_recovery_attempts": [],
}
report = rc.evaluate_restart_impact(
inv,
restart_class=rc.RestartClass.FULL_MCP_RESTART,
requester_role="controller",
requester_permissions=rc.permissions_for_role("controller"),
controller_approved=True,
operator_authorized=True,
)
assert report.allow_restart is False
assert report.attempt_log_satisfied is False
assert report.verdict == rc.VERDICT_UNSAFE
blob = " ".join(report.reasons + report.authorization_reasons)
assert "#669" in blob or "attempt log" in blob
def test_coordinator_allows_full_restart_with_attempt_log():
inv = {
"inventory_complete": True,
"sessions": [],
"leases": [],
"prior_recovery_attempts": [
{
"action": "client_reconnect",
"outcome": "insufficient",
"reason": "still broken",
}
],
}
report = rc.evaluate_restart_impact(
inv,
restart_class=rc.RestartClass.FULL_MCP_RESTART,
requester_role="controller",
requester_permissions=rc.permissions_for_role("controller"),
controller_approved=True,
operator_authorized=True,
)
assert report.attempt_log_satisfied is True
assert report.allow_restart is True
assert report.verdict == rc.VERDICT_SAFE
def test_coordinator_break_glass_allows_without_log():
inv = {
"inventory_complete": True,
"sessions": [],
"leases": [],
"prior_recovery_attempts": [],
}
report = rc.evaluate_restart_impact(
inv,
restart_class=rc.RestartClass.FULL_MCP_RESTART,
requester_role="controller",
requester_permissions=rc.permissions_for_role("controller"),
controller_approved=True,
operator_authorized=True,
break_glass=True,
)
assert report.break_glass is True
assert report.attempt_log_satisfied is True
assert report.allow_restart is True
def test_coordinator_client_reconnect_unaffected():
inv = {
"inventory_complete": True,
"sessions": [],
"leases": [],
"prior_recovery_attempts": [],
}
report = rc.evaluate_restart_impact(
inv,
restart_class=rc.RestartClass.CLIENT_RECONNECT,
requester_role="author",
requester_permissions=rc.permissions_for_role("author"),
)
assert report.attempt_log_satisfied is True
assert report.allow_restart is True
def test_restart_class_alias_accepted():
assert (
rp.resolve_action("full_mcp_restart")
is rp.RecoveryAction.FULL_MCP_RESTART
)
+511
View File
@@ -0,0 +1,511 @@
"""Tests for the Gitea issue↔PR linkage console (#645, Phase 3).
Covers the acceptance criteria of the issue:
* AC1 — issue↔PR linkage is visible for the selected project/repo, in both
directions, with the evidence that produced each edge.
* AC2 — the latest canonical handoff (CTH) is summarized for a focused thread.
* AC3 — an external Gitea link appears only under the admin reveal opt-in.
* AC4 — every case is driven by mocked Gitea payloads; no network.
Plus the invariants this console must not violate: a partial or failed read is
never rendered as "no link exists", an unfetched thread is never rendered as
"no handoff", redaction happens before display, and the surface stays read-only.
"""
from __future__ import annotations
import json
import os
import sys
import unittest
from pathlib import Path
from unittest import mock
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
from tests.webui_testclient import TestClient
from canonical_thread_handoff import format_cth_body
from webui.app import create_app
from webui.linkage_loader import (
EVIDENCE_BRANCH,
EVIDENCE_CLOSES,
EVIDENCE_REFERENCE,
HANDOFF_LOADED,
HANDOFF_NOT_LOADED,
HANDOFF_UNAVAILABLE,
LinkageSnapshot,
load_linkage_snapshot,
resolve_linkage,
resolve_pr_links,
snapshot_to_dict,
summarize_handoff,
)
from webui.linkage_views import render_linkage_page
from webui.nav import nav_hrefs
from webui.queue_loader import PaginationMeta
def _pagination(*, complete: bool = True, count: int = 0) -> PaginationMeta:
return PaginationMeta(
page=1,
per_page=50,
returned_count=count,
has_more=not complete,
is_final_page=complete,
inventory_complete=complete,
pages_fetched=1,
)
def _pr(
number: int,
*,
title: str = "",
body: str = "",
head: str = "",
state: str = "open",
labels: tuple[str, ...] = (),
) -> dict:
return {
"number": number,
"title": title or f"pr {number}",
"body": body,
"state": state,
"head": {"ref": head},
"labels": [{"name": name} for name in labels],
}
def _issue(
number: int,
*,
title: str = "",
state: str = "open",
labels: tuple[str, ...] = (),
) -> dict:
return {
"number": number,
"title": title or f"issue {number}",
"state": state,
"labels": [{"name": name} for name in labels],
}
def _fetcher(items: list[dict], *, complete: bool = True):
def _fetch(*_args, **_kwargs):
return items, _pagination(complete=complete, count=len(items))
return _fetch
def _load(
issues: list[dict],
prs: list[dict],
*,
complete: bool = True,
**kwargs,
) -> LinkageSnapshot:
return load_linkage_snapshot(
fetch_prs=_fetcher(prs, complete=complete),
fetch_issues=_fetcher(issues, complete=complete),
**kwargs,
)
def _cth(comment_id: int, *, created_at: str, status: str, next_owner: str) -> dict:
return {
"id": comment_id,
"created_at": created_at,
"user": {"login": "jcwalker3"},
"body": format_cth_body(
cth_type="Author Handoff",
status=status,
next_owner=next_owner,
current_blocker="none",
decision="implemented",
proof="full suite green",
next_action="review PR",
ready_to_paste_prompt="Review PR #902 now.",
),
}
class TestLinkageEvidence(unittest.TestCase):
"""AC1 — every edge records how it was found, and keeps all candidates."""
def test_closes_keyword_in_body_is_strongest_evidence(self):
links = resolve_pr_links(_pr(902, body="Closes #643"))
self.assertEqual([link.issue_number for link in links], [643])
self.assertEqual(links[0].evidence, (EVIDENCE_CLOSES,))
self.assertTrue(links[0].closes)
def test_closes_keyword_in_title_counts(self):
links = resolve_pr_links(_pr(902, title="feat(webui): preview (Closes #643)"))
self.assertEqual(links[0].evidence, (EVIDENCE_CLOSES,))
def test_canonical_branch_marker_links_without_a_keyword(self):
links = resolve_pr_links(_pr(902, head="feat/issue-643-request-preview"))
self.assertEqual([link.issue_number for link in links], [643])
self.assertEqual(links[0].evidence, (EVIDENCE_BRANCH,))
self.assertFalse(links[0].closes)
def test_non_canonical_branch_is_not_treated_as_a_marker(self):
self.assertEqual(resolve_pr_links(_pr(902, head="issue-643-preview")), ())
def test_bare_mention_is_recorded_as_the_weakest_evidence(self):
links = resolve_pr_links(_pr(902, body="context in #643"))
self.assertEqual(links[0].evidence, (EVIDENCE_REFERENCE,))
self.assertFalse(links[0].closes)
def test_several_evidence_kinds_merge_onto_one_edge(self):
links = resolve_pr_links(
_pr(902, body="Closes #643 — see #643", head="feat/issue-643-preview")
)
self.assertEqual(len(links), 1)
self.assertEqual(
links[0].evidence,
(EVIDENCE_CLOSES, EVIDENCE_BRANCH, EVIDENCE_REFERENCE),
)
def test_stronger_evidence_sorts_first(self):
links = resolve_pr_links(_pr(902, body="Closes #643, related #700"))
self.assertEqual([link.issue_number for link in links], [643, 700])
def test_self_reference_is_not_linkage(self):
links = resolve_pr_links(_pr(902, body="supersedes #902"))
self.assertEqual(links, ())
def test_every_candidate_is_kept_never_collapsed_to_a_guess(self):
links = resolve_pr_links(_pr(902, body="Closes #643\nCloses #644"))
self.assertEqual([link.issue_number for link in links], [643, 644])
class TestLinkageIndex(unittest.TestCase):
def test_issue_direction_ignores_mention_only_edges(self):
index = resolve_linkage([_pr(902, body="context in #643")])
self.assertIsNone(index.issue_prs.get(643))
self.assertEqual(index.pr_links[902][0].evidence, (EVIDENCE_REFERENCE,))
def test_contested_issue_is_reported_when_two_prs_claim_it(self):
index = resolve_linkage(
[_pr(902, body="Closes #643"), _pr(903, head="feat/issue-643-again")]
)
self.assertEqual(index.contested_issues(), (643,))
self.assertEqual(index.issue_prs[643], (902, 903))
def test_single_claim_is_not_contested(self):
index = resolve_linkage([_pr(902, body="Closes #643")])
self.assertEqual(index.contested_issues(), ())
def test_ambiguous_when_two_issues_tie_at_the_strongest_evidence(self):
index = resolve_linkage([_pr(902, body="Closes #643\nCloses #644")])
self.assertTrue(index.ambiguous(902))
def test_weaker_candidate_alongside_a_stronger_one_is_not_ambiguous(self):
index = resolve_linkage([_pr(902, body="Closes #643, see #700")])
self.assertFalse(index.ambiguous(902))
self.assertEqual(index.primary_issue(902).issue_number, 643)
def test_malformed_pr_row_is_skipped_not_raised_on(self):
index = resolve_linkage([{"title": "no number"}, _pr(902, body="Closes #643")])
self.assertEqual(sorted(index.pr_links), [902])
class TestLinkageSnapshot(unittest.TestCase):
"""AC1 — linkage is visible per project/repo, in both directions."""
def test_both_directions_are_populated(self):
snapshot = _load([_issue(643)], [_pr(902, body="Closes #643")])
self.assertTrue(snapshot.ok)
self.assertEqual([node.number for node in snapshot.issues], [643])
self.assertEqual(snapshot.issues[0].linked_prs, (902,))
self.assertEqual(snapshot.prs[0].links[0].issue_number, 643)
def test_repo_scope_comes_from_the_registry_project(self):
snapshot = _load([], [])
self.assertIn("/", snapshot.repo_label)
self.assertTrue(snapshot.project_id)
def test_unknown_project_fails_closed_with_a_reason(self):
snapshot = _load([_issue(643)], [], project_id="no-such-project")
self.assertFalse(snapshot.ok)
self.assertIn("not found in registry", snapshot.fetch_error)
self.assertEqual(snapshot.issues, ())
def test_orphan_pr_is_identifiable(self):
snapshot = _load([], [_pr(902), _pr(903, body="Closes #643")])
self.assertEqual([node.number for node in snapshot.orphan_prs], [902])
def test_state_scope_defaults_to_open_and_is_reported(self):
self.assertEqual(_load([], []).state_scope, "open")
self.assertEqual(_load([], [], state="all").state_scope, "all")
def test_unsupported_state_falls_back_to_open(self):
self.assertEqual(_load([], [], state="../etc").state_scope, "open")
def test_state_is_passed_through_to_the_fetchers(self):
seen: list[str] = []
def _fetch(*_args, **kwargs):
seen.append(kwargs.get("state", ""))
return [], _pagination()
load_linkage_snapshot(state="all", fetch_prs=_fetch, fetch_issues=_fetch)
self.assertEqual(seen, ["all", "all"])
class TestPartialInventoryIsNotAnAbsenceClaim(unittest.TestCase):
"""An empty edge list from a partial read must never read as 'no link'."""
def test_incomplete_pagination_marks_links_non_authoritative(self):
snapshot = _load([_issue(643)], [], complete=False)
self.assertFalse(snapshot.inventory_complete)
self.assertFalse(snapshot.issues[0].links_authoritative)
def test_complete_pagination_marks_links_authoritative(self):
snapshot = _load([_issue(643)], [], complete=True)
self.assertTrue(snapshot.inventory_complete)
self.assertTrue(snapshot.issues[0].links_authoritative)
def test_partial_window_renders_a_qualified_empty_cell(self):
html = render_linkage_page(_load([_issue(643)], [], complete=False))
self.assertIn("none found (partial inventory)", html)
def test_complete_window_renders_a_plain_none(self):
html = render_linkage_page(_load([_issue(643)], [], complete=True))
self.assertNotIn("partial inventory", html)
self.assertIn(">none<", html)
def test_missing_credentials_fail_closed_without_a_table(self):
with mock.patch(
"webui.linkage_loader._offline_test_mode", return_value=False
), mock.patch("webui.linkage_loader.get_auth_header", return_value=""):
snapshot = load_linkage_snapshot()
self.assertFalse(snapshot.ok)
self.assertIn("credentials unavailable", snapshot.fetch_error)
html = render_linkage_page(snapshot)
self.assertIn("Linkage unavailable", html)
self.assertNotIn("Issues → pull requests", html)
def test_fetch_failure_is_reported_not_raised(self):
def _boom(*_args, **_kwargs):
raise RuntimeError("gitea 502")
snapshot = load_linkage_snapshot(fetch_prs=_boom, fetch_issues=_boom)
self.assertFalse(snapshot.ok)
self.assertIn("Gitea fetch failed", snapshot.fetch_error)
class TestHandoffSummary(unittest.TestCase):
"""AC2 — the latest canonical handoff is summarized for a focused thread."""
def test_latest_cth_wins(self):
summary = summarize_handoff([
_cth(1, created_at="2026-07-24T10:00:00Z", status="in progress",
next_owner="author"),
_cth(2, created_at="2026-07-25T10:00:00Z", status="PR-open",
next_owner="reviewer"),
])
self.assertEqual(summary.comment_id, 2)
self.assertEqual(summary.status, "PR-open")
self.assertEqual(summary.next_owner, "reviewer")
self.assertTrue(summary.cth_type_known)
def test_thread_without_a_cth_summarizes_to_none(self):
self.assertIsNone(summarize_handoff([{"id": 1, "body": "ordinary comment"}]))
def test_unknown_heading_is_reported_not_republished(self):
summary = summarize_handoff([
{
"id": 5,
"created_at": "2026-07-25T10:00:00Z",
"user": {"login": "someone"},
"body": "<!-- cth:v1 -->\n## CTH: Totally Made Up\n\nStatus: odd\n",
}
])
self.assertFalse(summary.cth_type_known)
self.assertEqual(summary.cth_type, "unrecognized")
self.assertNotIn("Totally Made Up", json.dumps(summary.to_dict()))
def test_focused_pr_loads_its_handoff(self):
snapshot = _load(
[_issue(643)],
[_pr(902, body="Closes #643")],
pr=902,
comment_source=lambda kind, number: [
_cth(2, created_at="2026-07-25T10:00:00Z", status="PR-open",
next_owner="reviewer")
],
)
self.assertEqual(snapshot.handoff_status.state, HANDOFF_LOADED)
self.assertEqual(snapshot.focus, ("pr", 902))
self.assertEqual(snapshot.prs[0].handoff.status, "PR-open")
def test_unfocused_rows_report_not_loaded_never_none(self):
snapshot = _load(
[_issue(643)],
[_pr(902, body="Closes #643"), _pr(903)],
pr=902,
comment_source=lambda kind, number: [],
)
other = next(node for node in snapshot.prs if node.number == 903)
self.assertIsNone(other.handoff)
self.assertEqual(other.handoff_status.state, HANDOFF_NOT_LOADED)
self.assertIn("not loaded", render_linkage_page(snapshot))
def test_no_focus_means_no_thread_is_claimed_handoff_free(self):
snapshot = _load([_issue(643)], [])
self.assertEqual(snapshot.handoff_status.state, HANDOFF_NOT_LOADED)
self.assertIn("thread-scoped", snapshot.handoff_status.reason)
def test_comment_source_failure_degrades_only_the_handoff(self):
def _boom(_kind, _number):
raise RuntimeError("comments 500")
snapshot = _load(
[_issue(643)], [_pr(902, body="Closes #643")], pr=902, comment_source=_boom
)
self.assertTrue(snapshot.ok)
self.assertEqual(snapshot.handoff_status.state, HANDOFF_UNAVAILABLE)
self.assertEqual(snapshot.issues[0].linked_prs, (902,))
self.assertIn("unavailable", render_linkage_page(snapshot))
def test_loaded_thread_with_no_cth_says_so_explicitly(self):
snapshot = _load(
[_issue(643)],
[_pr(902, body="Closes #643")],
pr=902,
comment_source=lambda kind, number: [{"id": 1, "body": "hi"}],
)
self.assertIn(
"no Canonical Thread Handoff comment found", render_linkage_page(snapshot)
)
class TestDeepLinks(unittest.TestCase):
"""AC3 — an external Gitea link is emitted only when permitted."""
def test_deep_links_are_withheld_by_default(self):
with mock.patch.dict(os.environ, {"GITEA_MCP_REVEAL_ENDPOINTS": ""}):
snapshot = _load([_issue(643)], [])
html = render_linkage_page(snapshot)
self.assertFalse(snapshot.deep_links_enabled)
self.assertIsNone(snapshot.issues[0].deep_link)
self.assertIn("Gitea deep links are withheld", html)
def test_reveal_opt_in_emits_the_link(self):
with mock.patch.dict(os.environ, {"GITEA_MCP_REVEAL_ENDPOINTS": "1"}):
snapshot = _load([_issue(643)], [_pr(902, body="Closes #643")])
html = render_linkage_page(snapshot)
self.assertTrue(snapshot.deep_links_enabled)
self.assertIn("/issues/643", snapshot.issues[0].deep_link)
self.assertIn("/pulls/902", snapshot.prs[0].deep_link)
self.assertIn(f'href="{snapshot.issues[0].deep_link}"', html)
class TestRedactionBoundary(unittest.TestCase):
def test_secret_shaped_title_is_redacted_before_display(self):
snapshot = _load(
[_issue(643, title="token=ghp_thisisnotarealsecretvalue0001")], []
)
payload = json.dumps(snapshot_to_dict(snapshot))
self.assertNotIn("ghp_thisisnotarealsecretvalue0001", payload)
self.assertNotIn(
"ghp_thisisnotarealsecretvalue0001", render_linkage_page(snapshot)
)
def test_handoff_fields_are_redacted(self):
comment = _cth(
2, created_at="2026-07-25T10:00:00Z", status="ok", next_owner="reviewer"
)
comment["body"] += "\nDecision: password=hunter2hunter2\n"
snapshot = _load(
[_issue(643)],
[_pr(902, body="Closes #643")],
pr=902,
comment_source=lambda kind, number: [comment],
)
self.assertNotIn("hunter2hunter2", json.dumps(snapshot_to_dict(snapshot)))
self.assertNotIn("hunter2hunter2", render_linkage_page(snapshot))
def test_html_escapes_markup_in_a_title(self):
snapshot = _load([_issue(643, title="<script>alert(1)</script>")], [])
html = render_linkage_page(snapshot)
self.assertNotIn("<script>alert(1)</script>", html)
self.assertIn("&lt;script&gt;", html)
class TestLinkageRoutes(unittest.TestCase):
def setUp(self):
self.snapshot = _load(
[_issue(643, labels=("status:ready",))],
[_pr(902, body="Closes #643", labels=("status:pr-open",))],
)
self.client = TestClient(create_app())
def test_page_renders_both_tables(self):
with mock.patch("webui.app.load_linkage_snapshot", return_value=self.snapshot):
response = self.client.get("/gitea")
self.assertEqual(response.status_code, 200)
self.assertIn("Issues → pull requests", response.text)
self.assertIn("Pull requests → issues", response.text)
self.assertIn("#643", response.text)
def test_api_exports_the_same_model(self):
with mock.patch("webui.app.load_linkage_snapshot", return_value=self.snapshot):
response = self.client.get("/api/v1/gitea/linkage")
self.assertEqual(response.status_code, 200)
payload = response.json()
self.assertTrue(payload["ok"])
self.assertEqual(payload["issues"][0]["linked_prs"], [902])
self.assertEqual(payload["prs"][0]["links"][0]["issue_number"], 643)
self.assertEqual(payload["schema_version"], 1)
def test_api_declares_the_evidence_vocabulary(self):
with mock.patch("webui.app.load_linkage_snapshot", return_value=self.snapshot):
payload = self.client.get("/api/v1/gitea/linkage").json()
names = {entry["name"] for entry in payload["evidence_kinds"]}
self.assertEqual(names, {EVIDENCE_CLOSES, EVIDENCE_BRANCH, EVIDENCE_REFERENCE})
def test_api_fails_closed_with_a_non_200_when_the_read_failed(self):
failed = _load([], [], project_id="no-such-project")
with mock.patch("webui.app.load_linkage_snapshot", return_value=failed):
response = self.client.get("/api/v1/gitea/linkage")
self.assertEqual(response.status_code, 502)
self.assertFalse(response.json()["ok"])
def test_page_still_renders_when_the_read_failed(self):
failed = _load([], [], project_id="no-such-project")
with mock.patch("webui.app.load_linkage_snapshot", return_value=failed):
response = self.client.get("/gitea")
self.assertEqual(response.status_code, 200)
self.assertIn("Linkage unavailable", response.text)
def test_query_parameters_reach_the_loader(self):
with mock.patch(
"webui.app.load_linkage_snapshot", return_value=self.snapshot
) as loader:
self.client.get("/gitea?project=gitea-tools&state=all&pr=902")
loader.assert_called_once()
args, kwargs = loader.call_args
self.assertEqual(args[0], "gitea-tools")
self.assertEqual(kwargs["state"], "all")
self.assertEqual(kwargs["pr"], 902)
self.assertIsNone(kwargs["issue"])
def test_surface_stays_read_only(self):
for path in ("/gitea", "/api/v1/gitea/linkage"):
with self.subTest(path=path):
self.assertEqual(self.client.post(path).status_code, 405)
def test_nav_exposes_the_linkage_page_as_live(self):
self.assertIn("/gitea", nav_hrefs())
home = self.client.get("/").text
self.assertIn('href="/gitea"', home)
self.assertIn(">Gitea<", home)
if __name__ == "__main__":
unittest.main()
-452
View File
@@ -1,452 +0,0 @@
"""Read-only restart console: views, gates, and honesty rules (#667).
The console consumes the #655 substrate. These tests hold it to the three
properties that make a status surface trustworthy:
* an unreadable source is reported unavailable, never rendered as green;
* authorization is probed the way execution would probe it, so an allow is
never shown for something that could not run;
* the surface performs no mutation, including no write to the control-plane DB.
"""
from __future__ import annotations
import os
import sqlite3
import sys
import tempfile
import unittest
from datetime import datetime, timedelta, timezone
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
from starlette.testclient import TestClient # noqa: E402
import restart_coordinator # noqa: E402
from webui import console_authz, restart_console, restart_views # noqa: E402
from webui.app import create_app # noqa: E402
NOW = datetime(2026, 7, 25, 21, 0, 0, tzinfo=timezone.utc)
def _principal(role: str) -> console_authz.Principal:
return console_authz.Principal(
subject="[email protected]",
role=role,
identity_source=console_authz.IDENTITY_LOCAL_DEV,
authenticated=True,
)
def _inventory(*, complete: bool = True, sessions=(), leases=()):
def _read(**_kwargs):
return {
"sessions": list(sessions),
"leases": list(leases),
"terminal_lock": None,
"prior_recovery_attempts": [],
"inventory_complete": complete,
"incomplete_reasons": (
[] if complete else ["fixture: inventory withheld"]
),
}
return _read
def _live_session(session_id: str = "prgs-author-1234-abcd") -> dict:
return {
"session_id": session_id,
"role": "author",
"profile": "prgs-author",
"pid": os.getpid(),
"status": "active",
"last_heartbeat_at": (NOW - timedelta(seconds=30)).isoformat(),
}
def drain_proof_fixture() -> dict:
"""A structurally complete but unsigned drain proof."""
return {
"version": "drain-proof/v1",
"proof_id": "deadbeef" * 8,
"clean": True,
"issued_at": (NOW - timedelta(minutes=1)).isoformat(),
"expires_at": (NOW + timedelta(minutes=5)).isoformat(),
"requesting_session_id": "s-live",
"impact_fingerprint": "f" * 64,
"checks": [],
"failed_checks": [],
}
class RestartClassMatrixTest(unittest.TestCase):
def test_every_policy_class_is_rendered(self) -> None:
views = restart_console.build_restart_class_views("operator")
self.assertEqual(len(views), len(restart_coordinator.RESTART_CLASS_POLICIES))
def test_viewer_capability_is_role_scoped_not_generic(self) -> None:
"""A worker role must not be shown as able to request a full restart."""
author = {
v.restart_class: v
for v in restart_console.build_restart_class_views("author")
}
operator = {
v.restart_class: v
for v in restart_console.build_restart_class_views("operator")
}
full = restart_coordinator.RestartClass.FULL_MCP_RESTART.value
self.assertFalse(author[full].viewer_may_request)
self.assertFalse(author[full].viewer_may_execute)
self.assertTrue(operator[full].viewer_may_request)
self.assertTrue(operator[full].viewer_may_execute)
def test_unknown_role_may_do_nothing(self) -> None:
views = restart_console.build_restart_class_views("not-a-role")
self.assertTrue(all(not v.viewer_may_request for v in views))
self.assertTrue(all(not v.viewer_may_execute for v in views))
class AuthorizationProbeTest(unittest.TestCase):
def test_probe_asks_for_execution_so_phase_gate_is_reported(self) -> None:
"""An admin clears the role bar and still cannot execute in Phase 1.
This is the case that distinguishes the two probes. Asked without
``for_execution`` an admin is *allowed* for ``system.restart_namespace``,
which on a control surface reads as a live button. Asked the way
execution asks, the same principal is refused ``phase_not_active``. The
console must report the second answer.
"""
by_id = {
a.action_id: a
for a in restart_console.build_action_authorizations(
_principal(console_authz.ADMIN)
)
}
restart = by_id["system.restart_namespace"]
self.assertFalse(restart.execution_enabled)
self.assertEqual(restart.reason_code, console_authz.DENY_PHASE_NOT_ACTIVE)
permissive = console_authz.authorize(
"system.restart_namespace", _principal(console_authz.ADMIN)
)
self.assertTrue(
permissive.allowed,
"guard precondition: without for_execution an admin is allowed, "
"which is exactly why the console must not probe that way",
)
def test_operator_is_refused_the_admin_only_restart_action(self) -> None:
"""Role refusal precedes the phase gate and is reported as such."""
by_id = {
a.action_id: a
for a in restart_console.build_action_authorizations(
_principal(console_authz.OPERATOR)
)
}
self.assertEqual(
by_id["system.restart_namespace"].reason_code,
console_authz.DENY_INSUFFICIENT_ROLE,
)
def test_anonymous_is_denied_unauthenticated(self) -> None:
by_id = {
a.action_id: a for a in restart_console.build_action_authorizations(None)
}
self.assertEqual(
by_id["system.restart_namespace"].reason_code,
console_authz.DENY_UNAUTHENTICATED,
)
def test_no_authorization_ever_reports_execution_enabled(self) -> None:
for role in (
console_authz.VIEWER,
console_authz.OPERATOR,
console_authz.CONTROLLER,
console_authz.ADMIN,
):
for auth in restart_console.build_action_authorizations(_principal(role)):
self.assertFalse(
auth.execution_enabled,
f"{role} reported execution_enabled for {auth.action_id}",
)
class ImpactPreviewTest(unittest.TestCase):
def test_impact_renders_from_coordinator_dto(self) -> None:
impact, source = restart_console.load_impact_report(
principal=_principal(console_authz.OPERATOR),
read_inventory=_inventory(sessions=[_live_session()]),
now=NOW,
)
self.assertTrue(source.available)
self.assertIsNotNone(impact)
self.assertEqual(
impact["restart_class"],
restart_coordinator.RestartClass.FULL_MCP_RESTART.value,
)
self.assertIn("verdict", impact)
self.assertFalse(impact["restart_performed"])
self.assertTrue(impact["dry_run"])
def test_incomplete_inventory_is_surfaced_and_denies(self) -> None:
impact, source = restart_console.load_impact_report(
principal=_principal(console_authz.OPERATOR),
read_inventory=_inventory(complete=False),
now=NOW,
)
self.assertFalse(impact["inventory_complete"])
self.assertFalse(impact["allow_restart"])
self.assertTrue(source.detail, "incomplete inventory must explain itself")
def test_inventory_reader_failure_is_unavailable_not_empty(self) -> None:
"""A reader that raises must not be rendered as 'no sessions affected'."""
def _boom(**_kwargs):
raise RuntimeError("control-plane unreachable")
impact, source = restart_console.load_impact_report(
principal=_principal(console_authz.OPERATOR),
read_inventory=_boom,
now=NOW,
)
self.assertIsNone(impact)
self.assertFalse(source.available)
self.assertIn("control-plane unreachable", source.detail)
class ControlPlaneReadTest(unittest.TestCase):
def test_missing_database_is_incomplete_not_empty(self) -> None:
inventory = restart_console.read_control_plane_inventory(
db_path="/nonexistent/control-plane.sqlite3"
)
self.assertFalse(inventory["inventory_complete"])
self.assertEqual(inventory["sessions"], [])
self.assertTrue(inventory["incomplete_reasons"])
def test_reader_never_creates_the_database(self) -> None:
"""Reading status must not bring a control-plane DB into existence.
The path deliberately sits in a directory that already exists: a
read-write ``sqlite3.connect`` would happily create the file there, so
this fails if the reader ever stops opening the database ``mode=ro``.
A nested-missing-directory path would pass for the wrong reason,
because sqlite cannot create the parent directory either way.
"""
with tempfile.TemporaryDirectory() as tmp:
path = os.path.join(tmp, "control_plane.sqlite3")
self.assertTrue(os.path.isdir(os.path.dirname(path)))
inventory = restart_console.read_control_plane_inventory(db_path=path)
self.assertFalse(
os.path.exists(path),
"reading restart status created a control-plane database",
)
self.assertFalse(inventory["inventory_complete"])
def test_reads_active_sessions_from_a_real_database(self) -> None:
with tempfile.TemporaryDirectory() as tmp:
path = os.path.join(tmp, "cp.sqlite3")
conn = sqlite3.connect(path)
conn.execute(
"CREATE TABLE sessions (session_id TEXT, role TEXT, profile TEXT,"
" pid INTEGER, status TEXT, last_heartbeat_at TEXT)"
)
conn.execute(
"CREATE TABLE work_items (work_item_id INTEGER, kind TEXT,"
" number INTEGER)"
)
conn.execute(
"CREATE TABLE leases (lease_id TEXT, session_id TEXT, role TEXT,"
" phase TEXT, status TEXT, worktree_path TEXT,"
" work_item_id INTEGER, expires_at TEXT)"
)
conn.execute(
"INSERT INTO sessions VALUES (?,?,?,?,?,?)",
("s-live", "author", "prgs-author", 4242, "active", NOW.isoformat()),
)
conn.execute(
"INSERT INTO sessions VALUES (?,?,?,?,?,?)",
("s-done", "author", "prgs-author", 11, "closed", NOW.isoformat()),
)
conn.execute("INSERT INTO work_items VALUES (1, 'issue', 667)")
conn.execute(
"INSERT INTO leases VALUES (?,?,?,?,?,?,?,?)",
(
"l-1",
"s-live",
"author",
"allocated",
"active",
None,
1,
NOW.isoformat(),
),
)
conn.commit()
conn.close()
inventory = restart_console.read_control_plane_inventory(db_path=path)
self.assertTrue(inventory["inventory_complete"])
self.assertEqual([s["session_id"] for s in inventory["sessions"]], ["s-live"])
self.assertEqual(inventory["leases"][0]["work_number"], 667)
class DrainAndReconcileTest(unittest.TestCase):
def test_absent_drain_proof_is_not_a_pass(self) -> None:
drain, source = restart_console.load_drain_status(proof=None, now=NOW)
self.assertIsNone(drain)
self.assertFalse(source.available)
self.assertIn("denies", source.detail)
def test_tampered_drain_proof_is_reported_invalid(self) -> None:
proof = drain_proof_fixture()
proof["clean"] = True
proof["proof_id"] = "0" * 64
drain, source = restart_console.load_drain_status(proof=proof, now=NOW)
self.assertTrue(source.available)
self.assertFalse(drain["valid"])
def test_absent_reconcile_proof_is_unavailable(self) -> None:
reconcile, source = restart_console.load_reconcile_status(load_proof=None)
self.assertIsNone(reconcile)
self.assertFalse(source.available)
def test_reconcile_proof_is_rendered_when_supplied(self) -> None:
payload = {
"overall_status": "degraded",
"mode": "log_only",
"resolved_count": 3,
"unresolved_count": 2,
"items": [
{
"dimension": "leases",
"status": "unresolved",
"summary": "2 orphaned leases",
"follow_up_required": True,
}
],
}
reconcile, source = restart_console.load_reconcile_status(
load_proof=lambda: payload
)
self.assertTrue(source.available)
self.assertEqual(reconcile["unresolved_count"], 2)
class RenderingTest(unittest.TestCase):
def _snapshot(self, **kwargs):
params = {
"principal": _principal(console_authz.OPERATOR),
"read_inventory": _inventory(sessions=[_live_session()]),
"now": NOW,
}
params.update(kwargs)
return restart_console.load_restart_console_snapshot(**params)
def test_page_renders_every_section(self) -> None:
html = restart_views.render_restart_console_page(self._snapshot())
for heading in (
"Impact preview",
"Drain proof",
"Post-restart reconcile",
"Restart classes",
"Approval controls",
"Break-glass",
):
self.assertIn(heading, html)
def test_hostile_session_id_is_escaped(self) -> None:
hostile = "<script>alert('x')</script>"
html = restart_views.render_restart_console_page(
self._snapshot(read_inventory=_inventory(sessions=[_live_session(hostile)]))
)
self.assertNotIn("<script>alert", html)
self.assertIn("&lt;script&gt;", html)
def test_unavailable_impact_says_unsafe_rather_than_clean(self) -> None:
def _boom(**_kwargs):
raise RuntimeError("nope")
snapshot = self._snapshot(read_inventory=_boom)
html = restart_views.render_restart_console_page(snapshot)
self.assertIn("blast radius of a restart is unknown", html)
self.assertIn("unavailable", html)
def test_break_glass_is_hidden_from_unprivileged_viewers(self) -> None:
viewer_html = restart_views.render_restart_console_page(
self._snapshot(principal=_principal(console_authz.VIEWER))
)
self.assertIn("visible to operator-class", viewer_html)
self.assertNotIn(
f"#{restart_console.BREAK_GLASS_ISSUE}", viewer_html
)
def test_break_glass_shown_to_operator_is_marked_unavailable(self) -> None:
html = restart_views.render_restart_console_page(self._snapshot())
self.assertIn("unavailable", html)
self.assertIn(f"#{restart_console.BREAK_GLASS_ISSUE}", html)
def test_snapshot_always_declares_itself_read_only(self) -> None:
self.assertTrue(self._snapshot().read_only)
class RestartConsoleRouteTest(unittest.TestCase):
def setUp(self) -> None:
self.client = TestClient(create_app())
def test_page_route_renders(self) -> None:
res = self.client.get("/runtime/restart")
self.assertEqual(res.status_code, 200)
self.assertIn("Restart status and impact", res.text)
def test_api_route_exports_snapshot(self) -> None:
res = self.client.get("/api/v1/system/restart/status")
self.assertEqual(res.status_code, 200)
payload = res.json()
self.assertTrue(payload["read_only"])
self.assertEqual(payload["links"]["issue"], 667)
self.assertEqual(
len(payload["restart_classes"]),
len(restart_coordinator.RESTART_CLASS_POLICIES),
)
def test_restart_class_is_selectable(self) -> None:
res = self.client.get(
"/api/v1/system/restart/status?restart_class=client_reconnect"
)
self.assertEqual(res.status_code, 200)
self.assertEqual(res.json()["impact"]["restart_class"], "client_reconnect")
def test_unknown_restart_class_fails_closed(self) -> None:
res = self.client.get(
"/api/v1/system/restart/status?restart_class=obliterate-everything"
)
self.assertEqual(res.status_code, 200)
impact = res.json()["impact"]
self.assertFalse(impact["allow_restart"])
def test_anonymous_api_reader_gets_no_execution_grant(self) -> None:
payload = self.client.get("/api/v1/system/restart/status").json()
self.assertFalse(payload["break_glass"]["available"])
for auth in payload["authorizations"]:
self.assertFalse(auth["execution_enabled"])
def test_route_is_registered_in_nav(self) -> None:
from webui.nav import nav_hrefs
self.assertIn("/runtime/restart", nav_hrefs())
def test_no_write_method_is_exposed(self) -> None:
"""The surface is read-only: nothing accepts a POST."""
for path in ("/runtime/restart", "/api/v1/system/restart/status"):
self.assertEqual(self.client.post(path).status_code, 405, path)
if __name__ == "__main__":
unittest.main()
+42 -37
View File
@@ -47,15 +47,17 @@ from webui.traffic_views import render_traffic_page
from webui.worktree_scanner import load_hygiene_snapshot, snapshot_to_dict as worktree_snapshot_to_dict
from webui.worktree_views import render_worktrees_page
from webui.runtime_health import load_runtime_snapshot, snapshot_to_dict as runtime_snapshot_to_dict
import restart_coordinator
from webui.restart_console import load_restart_console_snapshot
from webui.restart_views import render_restart_console_page
from webui.runtime_views import render_runtime_page
from webui.session_loader import (
load_session_view_snapshot,
snapshot_to_dict as session_view_snapshot_to_dict,
)
from webui.session_views import render_sessions_page
from webui.linkage_loader import (
load_linkage_snapshot,
snapshot_to_dict as linkage_snapshot_to_dict,
)
from webui.linkage_views import render_linkage_page
from webui.inventory import (
SECTION_NAMES as _INVENTORY_SECTIONS,
load_inventory_snapshot,
@@ -334,33 +336,6 @@ async def api_runtime(_request: Request) -> JSONResponse:
return JSONResponse(runtime_snapshot_to_dict(load_runtime_snapshot()))
def _restart_console_snapshot(request: Request):
"""Build the read-only restart snapshot for the requesting principal (#667)."""
principal = resolve_principal(request.headers)
restart_class = (
request.query_params.get("restart_class")
or restart_coordinator.RestartClass.FULL_MCP_RESTART.value
)
return load_restart_console_snapshot(
principal=principal, restart_class=restart_class
)
async def restart_console_page(request: Request) -> HTMLResponse:
"""Restart status, impact preview, and approval state (#667). Read-only."""
snapshot = _restart_console_snapshot(request)
return HTMLResponse(
render_page(
title="Restart", body_html=render_restart_console_page(snapshot)
)
)
async def api_restart_status(request: Request) -> JSONResponse:
"""JSON export of the read-only restart console snapshot (#667)."""
return JSONResponse(_restart_console_snapshot(request).as_dict())
async def sessions(_request: Request) -> HTMLResponse:
"""Runtime and session view (#641) — read-only composition of health + inventory."""
snapshot = load_session_view_snapshot()
@@ -372,6 +347,41 @@ async def api_sessions(_request: Request) -> JSONResponse:
return JSONResponse(session_view_snapshot_to_dict(load_session_view_snapshot()))
def _linkage_snapshot(request: Request):
"""Load one linkage snapshot from the request's scope and focus parameters."""
return load_linkage_snapshot(
request.query_params.get("project") or None,
state=request.query_params.get("state"),
issue=_query_int(request, "issue"),
pr=_query_int(request, "pr"),
)
async def gitea_linkage(request: Request) -> HTMLResponse:
"""Gitea issue↔PR linkage console (#645) — read-only.
Always 200, including on a failed read: this is an operator view, and it
must render *why* linkage could not be loaded rather than withhold the page.
The snapshot itself carries ``ok=False`` and the page refuses to draw a
linkage table it cannot stand behind.
"""
return HTMLResponse(render_linkage_page(_linkage_snapshot(request)))
async def api_v1_gitea_linkage(request: Request) -> JSONResponse:
"""JSON export of the issue↔PR linkage model (#645).
Unlike the HTML view, the API answers with 502 when the snapshot could not
be loaded, so an automated consumer cannot read a fail-closed payload as a
successful "no links exist" result.
"""
snapshot = _linkage_snapshot(request)
return JSONResponse(
linkage_snapshot_to_dict(snapshot),
status_code=200 if snapshot.ok else 502,
)
async def _parse_audit_form(request: Request) -> tuple[str, str | None]:
if request.method == "GET":
return "", None
@@ -811,17 +821,12 @@ def create_app(*, bind_host: str | None = None) -> Starlette:
Route("/api/prompts", api_prompts, methods=["GET"]),
Route("/runtime", runtime, methods=["GET"]),
Route("/api/runtime", api_runtime, methods=["GET"]),
# #667 read-only restart status / impact preview / approval state.
Route("/runtime/restart", restart_console_page, methods=["GET"]),
Route(
"/api/v1/system/restart/status",
api_restart_status,
methods=["GET"],
),
Route("/sessions", sessions, methods=["GET"]),
Route("/api/sessions", api_sessions, methods=["GET"]),
Route("/api/v1/sessions", api_sessions, methods=["GET"]),
Route("/api/v1/timeline", api_v1_timeline, methods=["GET"]),
Route("/gitea", gitea_linkage, methods=["GET"]),
Route("/api/v1/gitea/linkage", api_v1_gitea_linkage, methods=["GET"]),
Route("/analytics", analytics, methods=["GET"]),
Route("/api/analytics", api_v1_analytics, methods=["GET"]),
Route("/api/v1/analytics", api_v1_analytics, methods=["GET"]),
+771
View File
@@ -0,0 +1,771 @@
"""Gitea issue↔PR linkage model for the console (#645, Phase 3).
Operators lose context between an issue and the PR that closes it: which PR
carries which issue, whether two PRs claim the same issue, and what the latest
canonical handoff on that thread said. The evidence exists in Gitea, but only
as free text scattered across PR titles, bodies, and branch names.
This module resolves that linkage into one read-only model:
* :func:`resolve_linkage` is a pure function from raw Gitea issue/PR payloads to
a :class:`LinkageIndex`. It records *how* each edge was found (a ``Closes #N``
keyword, the canonical ``feat/issue-N-…`` branch marker, or a bare ``#N``
body reference) and never collapses several candidates into one silent guess.
* :func:`load_linkage_snapshot` scopes that index to a registry project and
optionally attaches the latest Canonical Thread Handoff (CTH) summary for one
focused issue or PR.
Design rules, matching the rest of the console:
- **Read-only.** Gitea is read through the shared authenticated helpers. No
endpoint here mutates anything, and no write action is registered.
- **Qualified absence.** Linkage is a claim about a *loaded* window of Gitea.
When pagination did not complete, when credentials were unavailable, or when
only open items were fetched, the snapshot says so and every "no linked PR"
is marked non-authoritative. An empty edge list from a partial read is not
evidence that no link exists.
- **Handoff is loaded, never assumed.** CTH comments are thread-scoped, so they
are fetched only for an explicitly focused issue or PR. Every other row
reports ``not_loaded`` rather than rendering as "no handoff".
- **Redaction at the boundary.** Titles, labels, handoff fields, and error
reasons are free text from Gitea and cross :mod:`webui.console_redaction`
before they leave this module.
- **Deep links are opt-in.** A link to the Gitea web UI is emitted only under
the ``GITEA_MCP_REVEAL_ENDPOINTS`` admin opt-in, exactly as the MCP tools
gate their own URL exposure.
Non-goals (from the issue): no issue/PR editor, no browser review or merge, no
reimplementation of Gitea search.
"""
from __future__ import annotations
import os
import re
from dataclasses import dataclass
from typing import Any, Callable, Iterable, Sequence
from gitea_auth import api_fetch_page, get_auth_header, gitea_url, repo_api_url
from webui import console_redaction
from webui.project_registry import ProjectRecord, load_registry
from webui.queue_loader import (
PaginationMeta,
_fetch_issues,
_fetch_prs,
_host_from_url,
)
#: Version of the serialized linkage contract. Bump on any breaking change.
LINKAGE_SCHEMA_VERSION = 1
# --- Linkage evidence -------------------------------------------------------
# Ordered strongest to weakest. The strength ordering is what makes an
# ambiguous PR detectable: two candidates at the same strength are a genuine
# ambiguity, while a weaker candidate alongside a stronger one is not.
EVIDENCE_CLOSES = "closes_keyword"
EVIDENCE_BRANCH = "branch_marker"
EVIDENCE_REFERENCE = "body_reference"
EVIDENCE_ORDER: tuple[str, ...] = (
EVIDENCE_CLOSES,
EVIDENCE_BRANCH,
EVIDENCE_REFERENCE,
)
_EVIDENCE_RANK = {name: rank for rank, name in enumerate(EVIDENCE_ORDER)}
EVIDENCE_DESCRIPTIONS: dict[str, str] = {
EVIDENCE_CLOSES: (
"the PR title or body declares 'closes/fixes/resolves #N' — Gitea itself "
"acts on this keyword, so it is the strongest available evidence"
),
EVIDENCE_BRANCH: (
"the PR head branch carries the canonical issue marker "
"'(fix|feat|docs|chore)/issue-N-…' minted by the issue lock"
),
EVIDENCE_REFERENCE: (
"the PR body mentions '#N' without a closing keyword; a mention is not "
"a claim that the PR closes that issue"
),
}
_CLOSES_RE = re.compile(r"(?:closes|fixes|resolves)\s+#(\d+)", re.IGNORECASE)
_REFERENCE_RE = re.compile(r"#(\d+)")
_BRANCH_MARKER_RE = re.compile(
r"^(?:fix|feat|docs|chore)/issue-(\d+)(?:[-/]|$)", re.IGNORECASE
)
# Handoff-source states. ``not_loaded`` is deliberately distinct from "none
# found": a row whose comments were never fetched proves nothing about whether
# a handoff exists on that thread.
HANDOFF_NOT_LOADED = "not_loaded"
HANDOFF_LOADED = "loaded"
HANDOFF_UNAVAILABLE = "unavailable"
# Which item states were fetched. Linkage claims are scoped to this window.
STATE_OPEN = "open"
STATE_ALL = "all"
_SUPPORTED_STATES = (STATE_OPEN, STATE_ALL)
def _redact(value: Any) -> Any:
"""Redact one free-text field, failing closed to the placeholder."""
if value is None:
return None
return console_redaction.redact_text(str(value))
def deep_links_enabled(env: dict[str, str] | None = None) -> bool:
"""Whether Gitea web-UI deep links may be emitted (admin/debug opt-in)."""
source = env if env is not None else os.environ
return (source.get("GITEA_MCP_REVEAL_ENDPOINTS") or "").strip().lower() in {
"1",
"true",
"yes",
"on",
}
def _deep_link(host: str, org: str, repo: str, kind: str, number: int) -> str | None:
"""Build a Gitea web link for one item, or None when reveal is not enabled."""
if not deep_links_enabled() or not (host and org and repo):
return None
segment = "pulls" if kind == "pr" else "issues"
try:
return gitea_url(host, f"/{org}/{repo}/{segment}/{int(number)}")
except Exception:
return None
# --- Pure linkage resolution -------------------------------------------------
@dataclass(frozen=True)
class IssueLink:
"""One resolved edge from a PR to an issue, with the evidence that found it."""
issue_number: int
evidence: tuple[str, ...]
@property
def strength(self) -> int:
"""Rank of the strongest evidence backing this edge (lower is stronger)."""
return min(
(_EVIDENCE_RANK.get(name, len(EVIDENCE_ORDER)) for name in self.evidence),
default=len(EVIDENCE_ORDER),
)
@property
def closes(self) -> bool:
"""True only when the PR *declares* it closes the issue."""
return EVIDENCE_CLOSES in self.evidence
def to_dict(self) -> dict[str, Any]:
return {
"issue_number": self.issue_number,
"evidence": list(self.evidence),
"closes": self.closes,
}
def _sorted_links(links: Iterable[IssueLink]) -> tuple[IssueLink, ...]:
return tuple(sorted(links, key=lambda link: (link.strength, link.issue_number)))
def resolve_pr_links(pr: dict[str, Any]) -> tuple[IssueLink, ...]:
"""Resolve every issue a PR points at, strongest evidence first.
Every candidate is kept. Collapsing to a single "linked issue" is what makes
a mislinked or double-claimed PR invisible, so the caller decides what to do
with several candidates rather than being handed one guess.
"""
found: dict[int, set[str]] = {}
def _add(number: Any, evidence: str) -> None:
try:
issue_number = int(number)
except (TypeError, ValueError):
return
if issue_number <= 0:
return
found.setdefault(issue_number, set()).add(evidence)
title = str(pr.get("title") or "")
body = str(pr.get("body") or "")
for text in (title, body):
for match in _CLOSES_RE.finditer(text):
_add(match.group(1), EVIDENCE_CLOSES)
head_ref = str((pr.get("head") or {}).get("ref") or "")
branch_match = _BRANCH_MARKER_RE.match(head_ref.strip())
if branch_match:
_add(branch_match.group(1), EVIDENCE_BRANCH)
# The ``#N`` inside "Closes #N" is the *same* textual occurrence as the
# closing keyword, not a second, independent mention. Blanking the closing
# phrases first keeps "mention" meaning what the legend says it means: a
# reference the PR made without claiming to close anything.
for match in _REFERENCE_RE.finditer(_CLOSES_RE.sub(" ", body)):
_add(match.group(1), EVIDENCE_REFERENCE)
# A PR's own number appearing in its body is self-reference, not linkage.
try:
found.pop(int(pr.get("number")), None)
except (TypeError, ValueError):
pass
return _sorted_links(
IssueLink(
issue_number=number,
evidence=tuple(name for name in EVIDENCE_ORDER if name in evidence),
)
for number, evidence in found.items()
)
@dataclass(frozen=True)
class LinkageIndex:
"""Resolved linkage over one loaded window of issues and PRs."""
pr_links: dict[int, tuple[IssueLink, ...]]
issue_prs: dict[int, tuple[int, ...]]
def primary_issue(self, pr_number: int) -> IssueLink | None:
"""The strongest edge for a PR, or None when it points at no issue."""
links = self.pr_links.get(int(pr_number)) or ()
return links[0] if links else None
def ambiguous(self, pr_number: int) -> bool:
"""True when two or more issues tie at the PR's strongest evidence."""
links = self.pr_links.get(int(pr_number)) or ()
if len(links) < 2:
return False
best = links[0].strength
return sum(1 for link in links if link.strength == best) > 1
def contested_issues(self) -> tuple[int, ...]:
"""Issues claimed by more than one PR in the loaded window."""
return tuple(
number for number, prs in sorted(self.issue_prs.items()) if len(prs) > 1
)
def resolve_linkage(prs: Sequence[dict[str, Any]]) -> LinkageIndex:
"""Build the bidirectional linkage index for a loaded window of PRs.
Only *closing* and *branch-marker* edges populate the issue→PR direction: a
bare ``#N`` mention is a reference, and treating it as "this PR is the work
for issue N" would invent contested issues out of ordinary cross-links. The
weaker edge stays visible on the PR→issue side, where it is labelled.
"""
pr_links: dict[int, tuple[IssueLink, ...]] = {}
issue_prs: dict[int, list[int]] = {}
for pr in prs or []:
try:
pr_number = int(pr["number"])
except (KeyError, TypeError, ValueError):
continue
links = resolve_pr_links(pr)
pr_links[pr_number] = links
for link in links:
if link.evidence == (EVIDENCE_REFERENCE,):
continue
bucket = issue_prs.setdefault(link.issue_number, [])
if pr_number not in bucket:
bucket.append(pr_number)
return LinkageIndex(
pr_links=pr_links,
issue_prs={number: tuple(sorted(items)) for number, items in issue_prs.items()},
)
# --- Canonical handoff summary ----------------------------------------------
@dataclass(frozen=True)
class HandoffSummary:
"""The latest CTH comment on one thread, redacted for display."""
comment_id: int | None
created_at: str | None
author: str | None
cth_type: str
cth_type_known: bool
status: str | None
next_owner: str | None
current_blocker: str | None
decision: str | None
next_action: str | None
def to_dict(self) -> dict[str, Any]:
return {
"comment_id": self.comment_id,
"created_at": self.created_at,
"author": self.author,
"cth_type": self.cth_type,
"cth_type_known": self.cth_type_known,
"status": self.status,
"next_owner": self.next_owner,
"current_blocker": self.current_blocker,
"decision": self.decision,
"next_action": self.next_action,
}
@dataclass(frozen=True)
class HandoffStatus:
"""Why a thread's handoff summary is present, absent, or unknown."""
state: str
reason: str | None = None
target: str | None = None
@property
def loaded(self) -> bool:
return self.state == HANDOFF_LOADED
def to_dict(self) -> dict[str, Any]:
return {"state": self.state, "reason": self.reason, "target": self.target}
def summarize_handoff(comments: Sequence[dict[str, Any]]) -> HandoffSummary | None:
"""Summarize the newest CTH comment in *comments*, or None when there is none.
Every field is redacted before it is returned: a handoff body is operator
free text that regularly quotes commands, and it is rendered verbatim on the
page this feeds.
"""
from canonical_thread_handoff import find_latest_cth, is_known_cth_type
try:
latest = find_latest_cth(list(comments or []))
except Exception:
return None
if not latest:
return None
fields = latest.get("fields") or {}
cth_type = str(latest.get("cth_type") or "").strip()
known = is_known_cth_type(cth_type)
try:
comment_id: int | None = int(latest.get("comment_id"))
except (TypeError, ValueError):
comment_id = None
return HandoffSummary(
comment_id=comment_id,
created_at=_redact(latest.get("created_at")),
author=_redact(latest.get("author")),
# An unrecognised heading is reported as such rather than republished:
# the heading is free text, and CTH_TYPES is the only authority for what
# a handoff type may be.
cth_type=cth_type if known else "unrecognized",
cth_type_known=known,
status=_redact(fields.get("status")),
next_owner=_redact(fields.get("next owner")),
current_blocker=_redact(fields.get("current blocker")),
decision=_redact(fields.get("decision")),
next_action=_redact(fields.get("next action")),
)
CommentSource = Callable[[str, int], list[dict[str, Any]]]
def build_comment_source(host: str, org: str, repo: str) -> CommentSource | None:
"""Build an authenticated ``(kind, number) -> comments`` fetcher, or None.
Returns None when the console is running in offline test mode or when no
credential is available for *host*, so the caller reports the handoff source
as unavailable instead of as an empty thread.
"""
if _offline_test_mode() or not (host and org and repo):
return None
auth = get_auth_header(host)
if not auth:
return None
def _fetch(kind: str, number: int) -> list[dict[str, Any]]:
segment = "pulls" if kind == "pr" else "issues"
url = f"{repo_api_url(host, org, repo)}/{segment}/{int(number)}/comments"
comments: list[dict[str, Any]] = []
page = 1
while page <= 20:
raw, meta = api_fetch_page(url, auth, page=page, limit=50)
comments.extend(raw)
if bool(meta["is_final_page"]):
break
page += 1
return comments
return _fetch
# --- Snapshot ----------------------------------------------------------------
@dataclass(frozen=True)
class LinkageNode:
"""One issue or PR row with its resolved links and display metadata."""
kind: str
number: int
title: str
state: str
labels: tuple[str, ...] = ()
links: tuple[IssueLink, ...] = ()
linked_prs: tuple[int, ...] = ()
ambiguous: bool = False
contested: bool = False
deep_link: str | None = None
handoff: HandoffSummary | None = None
handoff_status: HandoffStatus = HandoffStatus(HANDOFF_NOT_LOADED)
links_authoritative: bool = True
def to_dict(self) -> dict[str, Any]:
return {
"kind": self.kind,
"number": self.number,
"title": self.title,
"state": self.state,
"labels": list(self.labels),
"links": [link.to_dict() for link in self.links],
"linked_prs": list(self.linked_prs),
"ambiguous": self.ambiguous,
"contested": self.contested,
"deep_link": self.deep_link,
"links_authoritative": self.links_authoritative,
"handoff": self.handoff.to_dict() if self.handoff else None,
"handoff_status": self.handoff_status.to_dict(),
}
@dataclass(frozen=True)
class LinkageSnapshot:
"""One answered linkage query over a scoped window of a Gitea repo."""
ok: bool
project_id: str
repo_label: str
host: str
state_scope: str
issues: tuple[LinkageNode, ...] = ()
prs: tuple[LinkageNode, ...] = ()
contested_issues: tuple[int, ...] = ()
focus: tuple[str, int] | None = None
inventory_complete: bool = False
deep_links_enabled: bool = False
handoff_status: HandoffStatus = HandoffStatus(HANDOFF_NOT_LOADED)
fetch_error: str | None = None
@property
def orphan_prs(self) -> tuple[LinkageNode, ...]:
"""PRs in the loaded window that point at no issue at all."""
return tuple(node for node in self.prs if not node.links)
def to_dict(self) -> dict[str, Any]:
return {
"ok": self.ok,
"schema_version": LINKAGE_SCHEMA_VERSION,
"project_id": self.project_id,
"repo": self.repo_label,
"state_scope": self.state_scope,
"inventory_complete": self.inventory_complete,
"deep_links_enabled": self.deep_links_enabled,
"focus": (
None
if self.focus is None
else {"kind": self.focus[0], "number": self.focus[1]}
),
"handoff_source": self.handoff_status.to_dict(),
"fetch_error": self.fetch_error,
"contested_issues": list(self.contested_issues),
"issues": [node.to_dict() for node in self.issues],
"prs": [node.to_dict() for node in self.prs],
"evidence_kinds": [
{"name": name, "description": EVIDENCE_DESCRIPTIONS[name]}
for name in EVIDENCE_ORDER
],
}
def snapshot_to_dict(snapshot: LinkageSnapshot) -> dict[str, Any]:
"""JSON-serializable export for ``/api/v1/gitea/linkage``."""
return snapshot.to_dict()
def _offline_test_mode() -> bool:
return (os.environ.get("WEBUI_TEST_OFFLINE") or "").strip().lower() in {
"1",
"true",
"yes",
}
def _labels_of(item: dict[str, Any]) -> tuple[str, ...]:
return tuple(
str(_redact(label.get("name")))
for label in (item.get("labels") or [])
if label.get("name")
)
def _failed_snapshot(
*,
project_id: str,
repo_label: str,
host: str,
state_scope: str,
reason: str,
) -> LinkageSnapshot:
"""A read that could not be answered. Never an empty-and-healthy snapshot."""
return LinkageSnapshot(
ok=False,
project_id=project_id,
repo_label=repo_label,
host=host,
state_scope=state_scope,
inventory_complete=False,
deep_links_enabled=deep_links_enabled(),
handoff_status=HandoffStatus(
HANDOFF_UNAVAILABLE, reason="linkage inventory could not be loaded"
),
fetch_error=str(_redact(reason)),
)
def _resolve_project(project_id: str | None) -> ProjectRecord | None:
registry = load_registry()
if project_id:
for entry in registry.projects:
if entry.id == project_id:
return entry
return None
return registry.projects[0] if registry.projects else None
def _normalize_state(state: str | None) -> str:
text = (state or STATE_OPEN).strip().lower()
return text if text in _SUPPORTED_STATES else STATE_OPEN
def load_linkage_snapshot(
project_id: str | None = None,
*,
state: str | None = None,
issue: int | None = None,
pr: int | None = None,
fetch_prs: Callable[..., tuple[list[dict], PaginationMeta]] | None = None,
fetch_issues: Callable[..., tuple[list[dict], PaginationMeta]] | None = None,
comment_source: CommentSource | None = None,
) -> LinkageSnapshot:
"""Load issue↔PR linkage for a registry project.
``issue``/``pr`` focus one thread: the focused row is the only one whose
Canonical Thread Handoff comments are fetched, because handoff comments are
thread-scoped and loading them for a whole queue would be one request per
row. Every unfocused row reports its handoff as ``not_loaded``.
"""
state_scope = _normalize_state(state)
try:
project = _resolve_project(project_id)
except Exception as exc: # registry invalid — fail closed with the reason
return _failed_snapshot(
project_id=project_id or "",
repo_label="",
host="",
state_scope=state_scope,
reason=f"project registry unavailable: {exc}",
)
if project is None:
return _failed_snapshot(
project_id=project_id or "",
repo_label="",
host="",
state_scope=state_scope,
reason=(
f"project {project_id!r} not found in registry"
if project_id
else "no projects registered"
),
)
host = _host_from_url(project.remote_host)
repo_label = f"{project.gitea_owner}/{project.repo_name}"
offline_test = _offline_test_mode()
def _empty_fetch(*_args, **_kwargs):
return [], PaginationMeta(
page=1,
per_page=50,
returned_count=0,
has_more=False,
is_final_page=True,
# An offline stub loaded nothing; claiming a complete inventory here
# would let the page assert that no issue has a linked PR.
inventory_complete=False,
pages_fetched=0,
)
pr_fetch = fetch_prs or (_empty_fetch if offline_test else _fetch_prs)
issue_fetch = fetch_issues or (_empty_fetch if offline_test else _fetch_issues)
using_live_fetch = not offline_test and (fetch_prs is None or fetch_issues is None)
auth = get_auth_header(host) if using_live_fetch else "test-auth"
if using_live_fetch and not auth:
return _failed_snapshot(
project_id=project.id,
repo_label=repo_label,
host=host,
state_scope=state_scope,
reason=(
f"Gitea credentials unavailable for {host}; linkage cannot be "
"loaded (fail closed — not rendering an empty linkage table)"
),
)
try:
raw_prs, pr_pagination = pr_fetch(
host, project.gitea_owner, project.repo_name, auth, state=state_scope
)
raw_issues, issue_pagination = issue_fetch(
host, project.gitea_owner, project.repo_name, auth, state=state_scope
)
except Exception as exc: # noqa: BLE001 — operator-visible fetch failure
return _failed_snapshot(
project_id=project.id,
repo_label=repo_label,
host=host,
state_scope=state_scope,
reason=f"Gitea fetch failed: {exc}",
)
inventory_complete = bool(
getattr(pr_pagination, "inventory_complete", False)
and getattr(issue_pagination, "inventory_complete", False)
)
index = resolve_linkage(raw_prs)
contested = index.contested_issues()
focus: tuple[str, int] | None = None
if pr is not None:
focus = ("pr", int(pr))
elif issue is not None:
focus = ("issue", int(issue))
unfocused_reason = (
"canonical handoff comments are thread-scoped; focus one issue or PR "
"to load its latest handoff"
)
handoff_status = HandoffStatus(HANDOFF_NOT_LOADED, reason=unfocused_reason)
focus_handoff: HandoffSummary | None = None
if focus is not None:
source = comment_source
if source is None and not offline_test:
source = build_comment_source(host, project.gitea_owner, project.repo_name)
target = f"{focus[0]}#{focus[1]}"
if source is None:
handoff_status = HandoffStatus(
HANDOFF_UNAVAILABLE,
reason="no authenticated comment source available for this read",
target=target,
)
else:
try:
focus_handoff = summarize_handoff(source(focus[0], focus[1]) or [])
handoff_status = HandoffStatus(HANDOFF_LOADED, target=target)
except Exception as exc: # fail soft: degrade this source only
handoff_status = HandoffStatus(
HANDOFF_UNAVAILABLE,
reason=str(_redact(f"handoff fetch failed: {exc}")),
target=target,
)
def _node_handoff(
kind: str, number: int
) -> tuple[HandoffSummary | None, HandoffStatus]:
"""Attach the handoff only to the focused row; qualify every other row."""
if focus == (kind, number):
return (focus_handoff, handoff_status)
return (
None,
HandoffStatus(
HANDOFF_NOT_LOADED,
reason=unfocused_reason if focus is None else "not the focused thread",
),
)
def _number_of(raw: dict[str, Any]) -> int | None:
try:
return int(raw["number"])
except (KeyError, TypeError, ValueError):
return None
def _sort_key(raw: dict[str, Any]) -> int:
number = _number_of(raw)
return -1 if number is None else number
issue_nodes: list[LinkageNode] = []
for raw in sorted(raw_issues or [], key=_sort_key, reverse=True):
number = _number_of(raw)
if number is None:
continue
node_handoff, node_status = _node_handoff("issue", number)
linked_prs = index.issue_prs.get(number, ())
issue_nodes.append(
LinkageNode(
kind="issue",
number=number,
title=str(_redact(raw.get("title")) or ""),
state=str(raw.get("state") or ""),
labels=_labels_of(raw),
linked_prs=linked_prs,
contested=len(linked_prs) > 1,
deep_link=_deep_link(
host, project.gitea_owner, project.repo_name, "issue", number
),
handoff=node_handoff,
handoff_status=node_status,
links_authoritative=inventory_complete,
)
)
pr_nodes: list[LinkageNode] = []
for raw in sorted(raw_prs or [], key=_sort_key, reverse=True):
number = _number_of(raw)
if number is None:
continue
node_handoff, node_status = _node_handoff("pr", number)
links = index.pr_links.get(number, ())
pr_nodes.append(
LinkageNode(
kind="pr",
number=number,
title=str(_redact(raw.get("title")) or ""),
state=str(raw.get("state") or ""),
labels=_labels_of(raw),
links=links,
ambiguous=index.ambiguous(number),
contested=any(link.issue_number in contested for link in links),
deep_link=_deep_link(
host, project.gitea_owner, project.repo_name, "pr", number
),
handoff=node_handoff,
handoff_status=node_status,
links_authoritative=inventory_complete,
)
)
return LinkageSnapshot(
ok=True,
project_id=project.id,
repo_label=repo_label,
host=host,
state_scope=state_scope,
issues=tuple(issue_nodes),
prs=tuple(pr_nodes),
contested_issues=contested,
focus=focus,
inventory_complete=inventory_complete,
deep_links_enabled=deep_links_enabled(),
handoff_status=handoff_status,
)
+364
View File
@@ -0,0 +1,364 @@
"""HTML views for the Gitea issue↔PR linkage console (#645, Phase 3).
Read-only renderer over :mod:`webui.linkage_loader`. The page's job is to make
three things impossible to misread:
* **why** an edge exists — every link carries its evidence badge, so a bare
``#N`` mention never looks like a closing claim;
* **what was not loaded** — a partial inventory, an unfocused thread, or an
unavailable handoff source renders as an explicit qualifier, never as an
affirmative "none";
* **that nothing here mutates** — there is no review, merge, or edit control,
and the deep link out to Gitea appears only under the admin reveal opt-in.
"""
from __future__ import annotations
from html import escape
from typing import Sequence
from webui.layout import render_page
from webui.linkage_loader import (
EVIDENCE_BRANCH,
EVIDENCE_CLOSES,
EVIDENCE_DESCRIPTIONS,
EVIDENCE_ORDER,
EVIDENCE_REFERENCE,
HANDOFF_LOADED,
HANDOFF_NOT_LOADED,
HandoffSummary,
LinkageNode,
LinkageSnapshot,
)
_EVIDENCE_CSS = {
EVIDENCE_CLOSES: "badge-health-ok",
EVIDENCE_BRANCH: "badge-health-skipped",
EVIDENCE_REFERENCE: "badge-health-unproven",
}
_EVIDENCE_LABEL = {
EVIDENCE_CLOSES: "closes",
EVIDENCE_BRANCH: "branch",
EVIDENCE_REFERENCE: "mention",
}
def _badge(text: str, css: str) -> str:
return f'<span class="badge {css}">{escape(text)}</span>'
def _labels(names: Sequence[str]) -> str:
if not names:
return '<span class="muted">—</span>'
return " ".join(_badge(name, "badge-health-skipped") for name in names)
def _ref(node: LinkageNode) -> str:
"""Render an item reference, hyperlinked only when deep links are revealed."""
label = f"#{node.number}"
if node.deep_link:
return f'<a href="{escape(node.deep_link)}"><code>{escape(label)}</code></a>'
return f"<code>{escape(label)}</code>"
def _scope_card(snapshot: LinkageSnapshot) -> str:
focus = (
"none"
if snapshot.focus is None
else f"{snapshot.focus[0]}#{snapshot.focus[1]}"
)
completeness = (
_badge("complete", "badge-health-ok")
if snapshot.inventory_complete
else _badge("partial", "badge-health-degraded")
)
links_note = (
"Every linkage edge below is a claim about this loaded window only."
if snapshot.inventory_complete
else (
"Pagination did not complete for this window, so an empty link list "
"means <em>none found in what was loaded</em> — not that no link exists."
)
)
deep_links = (
_badge("enabled", "badge-health-ok")
if snapshot.deep_links_enabled
else _badge("hidden", "badge-health-skipped")
)
return f"""<div class="health-card">
<h3>Scope</h3>
<table class="detail">
<tr><th>Project</th><td><code>{escape(snapshot.project_id or "")}</code></td></tr>
<tr><th>Repository</th><td><code>{escape(snapshot.repo_label or "")}</code></td></tr>
<tr><th>Item state</th><td><code>{escape(snapshot.state_scope)}</code></td></tr>
<tr><th>Focused thread</th><td><code>{escape(focus)}</code></td></tr>
<tr><th>Inventory</th><td>{completeness}</td></tr>
<tr><th>Gitea deep links</th><td>{deep_links}</td></tr>
</table>
<p class="muted">{links_note}</p>
</div>"""
def _error_card(snapshot: LinkageSnapshot) -> str:
if snapshot.ok and not snapshot.fetch_error:
return ""
return (
'<div class="health-card health-stale"><strong>Linkage unavailable:</strong> '
f"{escape(snapshot.fetch_error or 'the linkage read did not complete')}. "
"No linkage table is rendered: an empty table would read as "
"<em>no issue is linked to any PR</em>, which this read cannot claim."
"</div>"
)
def _contested_card(snapshot: LinkageSnapshot) -> str:
if not snapshot.contested_issues:
return ""
refs = ", ".join(f"<code>#{number}</code>" for number in snapshot.contested_issues)
return (
'<div class="health-card health-stale">'
f"<strong>Contested issues:</strong> {refs}. More than one PR in this "
"window claims each of these — duplicate work or a superseded PR. "
"Resolution stays in Gitea and the workflow; this console only reports it."
"</div>"
)
def _evidence_badges(evidence: Sequence[str]) -> str:
return " ".join(
_badge(
_EVIDENCE_LABEL.get(name, name),
_EVIDENCE_CSS.get(name, "badge-health-skipped"),
)
for name in EVIDENCE_ORDER
if name in evidence
)
def _handoff_inline(handoff: HandoffSummary) -> str:
type_css = "badge-health-ok" if handoff.cth_type_known else "badge-health-degraded"
return (
f'{_badge(handoff.cth_type or "", type_css)}'
f'<div class="muted" style="font-size:0.82rem;">'
f'{escape(handoff.status or "")}{escape(handoff.next_owner or "")}</div>'
)
def _handoff_cell(node: LinkageNode) -> str:
"""Render the handoff column, distinguishing 'none found' from 'not loaded'."""
status = node.handoff_status
if status.state == HANDOFF_LOADED:
if node.handoff is None:
return '<span class="muted">no canonical handoff on this thread</span>'
return _handoff_inline(node.handoff)
if status.state == HANDOFF_NOT_LOADED:
return (
f'{_badge("not loaded", "badge-health-skipped")}'
f'<div class="muted" style="font-size:0.82rem;">'
f'{escape(status.reason or "")}</div>'
)
return (
f'{_badge("unavailable", "badge-health-degraded")}'
f'<div class="muted" style="font-size:0.82rem;">'
f'{escape(status.reason or "")}</div>'
)
def _issue_rows(snapshot: LinkageSnapshot) -> str:
rows = []
for node in snapshot.issues:
if node.linked_prs:
linked = ", ".join(f"<code>#{number}</code>" for number in node.linked_prs)
if node.contested:
linked += " " + _badge("contested", "badge-blocked")
elif node.links_authoritative:
linked = '<span class="muted">none</span>'
else:
# The distinction an operator needs: nothing found in a window that
# was not fully loaded is not the same as nothing existing.
linked = '<span class="muted">none found (partial inventory)</span>'
rows.append(
"<tr>"
f"<td>{_ref(node)}</td>"
f"<td>{escape(node.title)}</td>"
f"<td>{escape(node.state or '')}</td>"
f"<td>{_labels(node.labels)}</td>"
f"<td>{linked}</td>"
f"<td>{_handoff_cell(node)}</td>"
"</tr>"
)
if not rows:
return '<tr><td colspan="6" class="muted">No issues in the loaded window.</td></tr>'
return "".join(rows)
def _pr_rows(snapshot: LinkageSnapshot) -> str:
rows = []
for node in snapshot.prs:
if node.links:
linked = "".join(
f"<div><code>#{link.issue_number}</code> "
f"{_evidence_badges(link.evidence)}</div>"
for link in node.links
)
if node.ambiguous:
linked += _badge("ambiguous", "badge-blocked")
if node.contested:
linked += " " + _badge("contested", "badge-blocked")
elif node.links_authoritative:
linked = '<span class="muted">no issue reference</span>'
else:
linked = '<span class="muted">none found (partial inventory)</span>'
rows.append(
"<tr>"
f"<td>{_ref(node)}</td>"
f"<td>{escape(node.title)}</td>"
f"<td>{escape(node.state or '')}</td>"
f"<td>{_labels(node.labels)}</td>"
f"<td>{linked}</td>"
f"<td>{_handoff_cell(node)}</td>"
"</tr>"
)
if not rows:
return (
'<tr><td colspan="6" class="muted">No pull requests in the loaded '
"window.</td></tr>"
)
return "".join(rows)
def _focus_card(snapshot: LinkageSnapshot) -> str:
"""Render the focused thread's latest canonical handoff, when one was loaded."""
if snapshot.focus is None:
return f"""<div class="prompt-card">
<h3>Canonical handoff</h3>
<p class="muted">{escape(snapshot.handoff_status.reason or "")}
Add <code>?issue=N</code> or <code>?pr=N</code> to load the latest
Canonical Thread Handoff for one thread.</p>
</div>"""
kind, number = snapshot.focus
target = f"{kind} #{number}"
if not snapshot.handoff_status.loaded:
return f"""<div class="prompt-card">
<h3>Canonical handoff — {escape(target)}</h3>
<p class="muted">{_badge("unavailable", "badge-health-degraded")}
{escape(snapshot.handoff_status.reason or "handoff source did not run")}.
This is not evidence that the thread carries no handoff.</p>
</div>"""
handoff = next(
(
node.handoff
for node in (snapshot.issues + snapshot.prs)
if node.kind == kind and node.number == number and node.handoff
),
None,
)
if handoff is None:
return f"""<div class="prompt-card">
<h3>Canonical handoff — {escape(target)}</h3>
<p class="muted">Comments loaded; no Canonical Thread Handoff comment found on
this thread.</p>
</div>"""
type_css = "badge-health-ok" if handoff.cth_type_known else "badge-health-degraded"
unknown_note = (
""
if handoff.cth_type_known
else (
'<p class="muted">The comment\'s heading is not a declared CTH type, '
"so it is reported as unrecognized rather than republished.</p>"
)
)
return f"""<div class="prompt-card">
<h3>Canonical handoff — {escape(target)}</h3>
<p class="meta">{_badge(handoff.cth_type or "", type_css)}
by <code>{escape(handoff.author or "unknown")}</code>
at <code>{escape(handoff.created_at or "unknown")}</code></p>
{unknown_note}
<table class="detail">
<tr><th>Status</th><td>{escape(handoff.status or "")}</td></tr>
<tr><th>Next owner</th><td>{escape(handoff.next_owner or "")}</td></tr>
<tr><th>Current blocker</th><td>{escape(handoff.current_blocker or "")}</td></tr>
<tr><th>Decision</th><td>{escape(handoff.decision or "")}</td></tr>
<tr><th>Next action</th><td>{escape(handoff.next_action or "")}</td></tr>
</table>
<p class="muted">Full event history:
<a href="/api/v1/timeline?{escape(kind)}={number}"><code>/api/v1/timeline</code></a></p>
</div>"""
def _legend_card(snapshot: LinkageSnapshot) -> str:
items = "".join(
f"<li>{_badge(_EVIDENCE_LABEL[name], _EVIDENCE_CSS[name])}"
f"{escape(EVIDENCE_DESCRIPTIONS[name])}</li>"
for name in EVIDENCE_ORDER
)
reveal_note = (
"Gitea deep links are shown because the "
"<code>GITEA_MCP_REVEAL_ENDPOINTS</code> admin opt-in is set."
if snapshot.deep_links_enabled
else (
"Gitea deep links are withheld. Set "
"<code>GITEA_MCP_REVEAL_ENDPOINTS=1</code> server-side to reveal "
"them; item numbers stay usable without them."
)
)
return f"""<div class="prompt-card">
<h3>How an edge was found</h3>
<ul class="reasons">{items}</ul>
<p class="muted">{reveal_note}</p>
<p class="muted">Read-only surface: no issue or PR editing, no review, and no merge.
JSON export: <a href="/api/v1/gitea/linkage"><code>/api/v1/gitea/linkage</code></a></p>
</div>"""
def render_linkage_page(snapshot: LinkageSnapshot) -> str:
"""Render the full HTML page for the Gitea linkage console."""
if not snapshot.ok:
return render_page(
title="Gitea linkage",
body_html=f"""<h2>Gitea issue and PR linkage</h2>
<p class="meta">Phase 3 read-only linkage console (#645).</p>
{_error_card(snapshot)}
{_scope_card(snapshot)}""",
)
body = f"""<h2>Gitea issue and PR linkage</h2>
<p class="meta">Phase 3 read-only linkage console (#645). Gitea remains the source of
truth; this page reads it and never writes to it.</p>
{_scope_card(snapshot)}
{_contested_card(snapshot)}
<div class="prompt-card">
<h3>Issues → pull requests</h3>
<table class="registry">
<thead>
<tr>
<th>Issue</th><th>Title</th><th>State</th><th>Labels</th>
<th>Linked PRs</th><th>Latest handoff</th>
</tr>
</thead>
<tbody>{_issue_rows(snapshot)}</tbody>
</table>
</div>
<div class="prompt-card">
<h3>Pull requests → issues</h3>
<table class="registry">
<thead>
<tr>
<th>PR</th><th>Title</th><th>State</th><th>Labels</th>
<th>Linked issues</th><th>Latest handoff</th>
</tr>
</thead>
<tbody>{_pr_rows(snapshot)}</tbody>
</table>
</div>
{_focus_card(snapshot)}
{_legend_card(snapshot)}
"""
return render_page(title="Gitea linkage", body_html=body)
+5 -2
View File
@@ -6,7 +6,8 @@ destination is a GET view or a Phase 1 placeholder. No mutation links.
Nav groups follow the #631 Phase 1 information architecture: Health, Traffic,
Runtime/Sessions, Projects, Inventory, Timeline, Policy (placeholder), and
Insights (placeholder). Later-phase surfaces are declared as ``stub`` items and
Insights (placeholder), joined by the Phase 3 Gitea linkage group (#645).
Later-phase surfaces are declared as ``stub`` items and
backed by ``STUB_PAGES`` so their nav links resolve to a graceful placeholder
instead of a 404.
"""
@@ -48,7 +49,6 @@ NAV_GROUPS: tuple[NavGroup, ...] = (
)),
NavGroup("Runtime/Sessions", (
NavItem("/runtime", "Runtime health"),
NavItem("/runtime/restart", "Restart status"),
NavItem("/sessions", "Sessions"),
)),
NavGroup("Projects", (
@@ -61,6 +61,9 @@ NAV_GROUPS: tuple[NavGroup, ...] = (
NavGroup("Timeline", (
NavItem("/timeline", "Timeline", "stub"),
)),
NavGroup("Gitea", (
NavItem("/gitea", "Issue/PR linkage"),
)),
NavGroup("Policy", (
NavItem("/policy", "Policy", "stub"),
NavItem("/prompts", "Prompts"),
+27 -2
View File
@@ -211,6 +211,19 @@ def _pagination_from_pages(
)
_SUPPORTED_FETCH_STATES = ("open", "closed", "all")
def _safe_state(state: str | None) -> str:
"""Constrain a caller-supplied item state before it reaches a query string.
The value is interpolated into the Gitea URL, so an unrecognised state falls
back to ``open`` rather than being passed through.
"""
text = (state or "open").strip().lower()
return text if text in _SUPPORTED_FETCH_STATES else "open"
def _fetch_prs(
host: str,
org: str,
@@ -218,8 +231,15 @@ def _fetch_prs(
auth: str,
*,
per_page: int = 50,
state: str = "open",
) -> tuple[list[dict], PaginationMeta]:
url = f"{repo_api_url(host, org, repo)}/pulls?state=open"
"""Fetch PRs in *state* (``open``, ``closed``, or ``all``).
The queue dashboard only ever wants the open window, so ``open`` stays the
default. The linkage console (#645) widens it, because a landed issue↔PR
edge lives on a merged PR.
"""
url = f"{repo_api_url(host, org, repo)}/pulls?state={_safe_state(state)}"
all_raw: list[dict] = []
pages_fetched = 0
is_final = False
@@ -248,8 +268,13 @@ def _fetch_issues(
auth: str,
*,
per_page: int = 50,
state: str = "open",
) -> tuple[list[dict], PaginationMeta]:
url = f"{repo_api_url(host, org, repo)}/issues?state=open&type=issues"
"""Fetch issues in *state* (``open``, ``closed``, or ``all``); see :func:`_fetch_prs`."""
url = (
f"{repo_api_url(host, org, repo)}/issues"
f"?state={_safe_state(state)}&type=issues"
)
all_raw: list[dict] = []
page = 1
pages_fetched = 0
-579
View File
@@ -1,579 +0,0 @@
"""Read-only restart status, impact preview, and approval state (#667).
Phase 1 of the console restart surface. It *consumes* the #655 coordinator
substrate and renders it; it never restarts, reloads, drains, approves, or kills
anything. There is no apply path in this module, so there is no execution gate
here to arm incorrectly the only writes the console could perform are the ones
it does not implement.
Sources, each independently fail-soft and each reported with its own
:class:`SourceStatus`:
* :mod:`restart_coordinator` restart-class policy matrix (#663) and the
blast-radius impact report (#658).
* :mod:`drain_proof` drain checklist and gate verdict (#661), verified
read-only against a caller-supplied proof.
* :mod:`post_restart_reconcile` post-restart completion proof (#662).
* :mod:`webui.console_authz` role authorization for the approval controls
(#633).
Three rules this module holds itself to, because a status surface that lies is
worse than one that is absent:
**A source that could not be read is reported unavailable, never green.** No
default, placeholder, or self-comparison is substituted for a reading that
failed. An unreadable control-plane DB yields ``inventory_complete=False``,
which the coordinator itself turns into a fail-closed verdict.
**Authorization is asked the way execution would ask it.** Every authorization
probe passes ``for_execution=True``, so the console reports whether the action
could actually run rather than the weaker "this principal is the right role".
While the console is in Phase 1 that answer is ``phase_not_active`` for every
phase-2 action, and the surface says so plainly instead of showing an allow.
**The database is opened read-only.** ``ControlPlaneDB()`` creates directories
and runs migrations on construction, which is a write; this module opens the
sqlite file with ``mode=ro`` exactly as :mod:`webui.inventory` does, and treats
a missing file as missing authority rather than an empty inventory.
"""
from __future__ import annotations
import os
import sqlite3
from dataclasses import dataclass, field
from datetime import datetime, timezone
from typing import Any, Callable, Mapping
import control_plane_db
import drain_proof
import restart_coordinator
from webui import console_authz
from webui.inventory import redact_path, scrub
# --- Source status ----------------------------------------------------------
STATUS_OK = "ok"
STATUS_UNAVAILABLE = "unavailable"
#: Console actions whose authorization state this surface reports. Both are
#: pre-existing #642 actions; this module adds no new console action because it
#: performs no console action.
REPORTED_ACTIONS: tuple[str, ...] = (
"system.restart_namespace",
"system.reload_namespace",
)
#: The break-glass workflow (#664) is not consumed here. It is declared so the
#: surface is honest about the gap rather than silently omitting a governance
#: path the operator has been told exists.
BREAK_GLASS_ISSUE = 664
BREAK_GLASS_PENDING_REASON = (
"The break-glass workflow (#664) is not yet available on this branch's "
"base; no break-glass control is offered and none is implied."
)
@dataclass(frozen=True)
class SourceStatus:
"""Whether one backing source could be read, and why not when it could not."""
name: str
status: str
detail: str = ""
@property
def available(self) -> bool:
return self.status == STATUS_OK
def as_dict(self) -> dict[str, Any]:
return {
"name": self.name,
"status": self.status,
"available": self.available,
"detail": self.detail,
}
@dataclass(frozen=True)
class RestartClassView:
"""One row of the #663 restart-class matrix, scoped to the viewer's role."""
restart_class: str
required_permission: str
expected_blast_radius: str
drain_requirement: str
full_drain_required: bool
approval_requirement: str
request_roles: tuple[str, ...]
execution_roles: tuple[str, ...]
viewer_may_request: bool
viewer_may_execute: bool
def as_dict(self) -> dict[str, Any]:
return {
"restart_class": self.restart_class,
"required_permission": self.required_permission,
"expected_blast_radius": self.expected_blast_radius,
"drain_requirement": self.drain_requirement,
"full_drain_required": self.full_drain_required,
"approval_requirement": self.approval_requirement,
"request_roles": list(self.request_roles),
"execution_roles": list(self.execution_roles),
"viewer_may_request": self.viewer_may_request,
"viewer_may_execute": self.viewer_may_execute,
}
@dataclass(frozen=True)
class ActionAuthorization:
"""Authorization state for one console action, asked as execution would."""
action_id: str
summary: str
required_role: str
allowed: bool
execution_enabled: bool
reason_code: str
detail: str
def as_dict(self) -> dict[str, Any]:
return {
"action_id": self.action_id,
"summary": self.summary,
"required_role": self.required_role,
"allowed": self.allowed,
"execution_enabled": self.execution_enabled,
"reason_code": self.reason_code,
"detail": self.detail,
}
@dataclass(frozen=True)
class BreakGlassSurface:
"""Declared-but-unavailable break-glass panel (#664 is not on this base)."""
available: bool
issue: int
reason: str
viewer_is_privileged: bool
def as_dict(self) -> dict[str, Any]:
return {
"available": self.available,
"issue": self.issue,
"reason": self.reason,
"viewer_is_privileged": self.viewer_is_privileged,
}
@dataclass(frozen=True)
class RestartConsoleSnapshot:
"""Everything the read-only restart console renders."""
generated_at: str
viewer_role: str
viewer_authenticated: bool
read_only: bool
impact: dict[str, Any] | None
impact_source: SourceStatus
drain: dict[str, Any] | None
drain_source: SourceStatus
reconcile: dict[str, Any] | None
reconcile_source: SourceStatus
restart_classes: tuple[RestartClassView, ...]
authorizations: tuple[ActionAuthorization, ...]
break_glass: BreakGlassSurface
notes: tuple[str, ...] = field(default_factory=tuple)
def as_dict(self) -> dict[str, Any]:
return {
"generated_at": self.generated_at,
"viewer_role": self.viewer_role,
"viewer_authenticated": self.viewer_authenticated,
"read_only": self.read_only,
"impact": self.impact,
"impact_source": self.impact_source.as_dict(),
"drain": self.drain,
"drain_source": self.drain_source.as_dict(),
"reconcile": self.reconcile,
"reconcile_source": self.reconcile_source.as_dict(),
"restart_classes": [c.as_dict() for c in self.restart_classes],
"authorizations": [a.as_dict() for a in self.authorizations],
"break_glass": self.break_glass.as_dict(),
"notes": list(self.notes),
"links": {
"issue": 667,
"extends": 642,
"umbrella": 655,
"coordinator": 658,
"drain_proof": 661,
"reconcile": 662,
"restart_classes": 663,
"break_glass": BREAK_GLASS_ISSUE,
"vision": 652,
"roadmap": 653,
},
}
def _utc_now() -> datetime:
return datetime.now(timezone.utc)
# --- Control-plane inventory (read-only) ------------------------------------
def read_control_plane_inventory(
*,
db_path: str | None = None,
limit: int = 200,
) -> dict[str, Any]:
"""Read sessions and leases for an impact evaluation, read-only.
Returns the inventory mapping
:func:`restart_coordinator.evaluate_restart_impact` expects.
``inventory_complete`` is True only when every read succeeded, so a partial
read denies rather than under-reporting the blast radius.
The database is never created, migrated, or written: a missing file means
the console has no session authority, which is not the same as there being
no sessions.
"""
path = (db_path or control_plane_db.default_db_path() or "").strip()
incomplete: list[str] = []
def _incomplete(reason: str) -> dict[str, Any]:
return {
"sessions": [],
"leases": [],
"terminal_lock": None,
"prior_recovery_attempts": [],
"inventory_complete": False,
"incomplete_reasons": [reason],
}
if not path:
return _incomplete("control-plane database path is not configured")
if not os.path.exists(path):
return _incomplete(
f"control-plane database not present at {redact_path(path)}; "
"no session or lease authority available"
)
try:
conn = sqlite3.connect(f"file:{path}?mode=ro", uri=True, timeout=5)
conn.row_factory = sqlite3.Row
except sqlite3.Error as exc:
return _incomplete(f"control-plane database could not be opened: {exc}")
sessions: list[dict[str, Any]] = []
leases: list[dict[str, Any]] = []
capped = max(1, int(limit))
try:
tables = {
str(row[0])
for row in conn.execute(
"SELECT name FROM sqlite_master WHERE type = 'table'"
).fetchall()
}
if "sessions" not in tables:
incomplete.append("control-plane database has no sessions table")
else:
sessions = [
dict(row)
for row in conn.execute(
"SELECT session_id, role, profile, pid, status,"
" last_heartbeat_at FROM sessions"
" WHERE status = 'active'"
" ORDER BY last_heartbeat_at DESC LIMIT ?",
(capped,),
).fetchall()
]
if "leases" not in tables:
incomplete.append("control-plane database has no leases table")
elif "work_items" not in tables:
incomplete.append(
"control-plane database has no work_items table; lease work "
"identity cannot be resolved"
)
else:
leases = [
dict(row)
for row in conn.execute(
"SELECT l.lease_id, l.session_id, l.role, l.phase,"
" l.status AS freshness, l.worktree_path,"
" w.kind AS work_kind, w.number AS work_number"
" FROM leases l"
" JOIN work_items w ON w.work_item_id = l.work_item_id"
" WHERE l.status = 'active'"
" ORDER BY l.expires_at DESC LIMIT ?",
(capped,),
).fetchall()
]
except sqlite3.Error as exc:
return _incomplete(f"control-plane database read failed: {exc}")
finally:
conn.close()
return {
"sessions": sessions,
"leases": leases,
"terminal_lock": None,
"prior_recovery_attempts": [],
"inventory_complete": not incomplete,
"incomplete_reasons": incomplete,
}
# --- Composition ------------------------------------------------------------
def build_restart_class_views(viewer_role: str | None) -> tuple[RestartClassView, ...]:
"""Render the #663 class matrix, marking what this viewer may request."""
normalized = str(viewer_role or "").strip().lower()
views: list[RestartClassView] = []
for policy in restart_coordinator.RESTART_CLASS_POLICIES.values():
views.append(
RestartClassView(
restart_class=policy.restart_class.value,
required_permission=policy.required_permission,
expected_blast_radius=policy.expected_blast_radius,
drain_requirement=policy.drain_requirement,
full_drain_required=policy.full_drain_required,
approval_requirement=policy.approval_requirement,
request_roles=tuple(policy.request_roles),
execution_roles=tuple(policy.execution_roles),
viewer_may_request=normalized in policy.request_roles,
viewer_may_execute=normalized in policy.execution_roles,
)
)
return tuple(views)
def build_action_authorizations(
principal: console_authz.Principal | None,
) -> tuple[ActionAuthorization, ...]:
"""Authorization state for the approval controls, asked as execution.
``for_execution=True`` is deliberate. Asking without it answers "is this
principal senior enough", which is not the question an operator looking at a
control needs answered; asking with it answers "would this run", and while
the console is in Phase 1 the honest answer is no.
"""
results: list[ActionAuthorization] = []
for action_id in REPORTED_ACTIONS:
action = console_authz.get_action(action_id)
decision = console_authz.authorize(action_id, principal, for_execution=True)
results.append(
ActionAuthorization(
action_id=action_id,
summary=action.summary if action else "",
required_role=(
action.minimum_role if action else console_authz.OPERATOR
),
allowed=bool(decision.allowed),
execution_enabled=bool(decision.execution_enabled),
reason_code=str(decision.reason_code or ""),
detail=str(decision.detail or ""),
)
)
return tuple(results)
def viewer_is_privileged(principal: console_authz.Principal | None) -> bool:
"""True when the viewer holds at least the operator role."""
who = principal if principal is not None else console_authz.ANONYMOUS
if not who.authenticated:
return False
return who.rank >= console_authz.ROLE_ORDER.index(console_authz.OPERATOR)
def load_impact_report(
*,
principal: console_authz.Principal | None = None,
restart_class: str = restart_coordinator.RestartClass.FULL_MCP_RESTART.value,
db_path: str | None = None,
limit: int = 200,
read_inventory: Callable[..., Mapping[str, Any]] | None = None,
now: datetime | None = None,
) -> tuple[dict[str, Any] | None, SourceStatus]:
"""Evaluate the blast radius for *restart_class*, always dry-run."""
reader = read_inventory or read_control_plane_inventory
try:
inventory = dict(reader(db_path=db_path, limit=limit))
except Exception as exc: # noqa: BLE001
return None, SourceStatus(
"impact",
STATUS_UNAVAILABLE,
f"control-plane inventory failed: {type(exc).__name__}: {exc}",
)
who = principal if principal is not None else console_authz.ANONYMOUS
viewer_role = str(who.role or "").strip().lower()
try:
report = restart_coordinator.evaluate_restart_impact(
inventory,
now=now,
dry_run=True,
restart_class=restart_class,
requester_role=viewer_role,
requester_permissions=restart_coordinator.permissions_for_role(
viewer_role
),
)
except Exception as exc: # noqa: BLE001
return None, SourceStatus(
"impact",
STATUS_UNAVAILABLE,
f"impact evaluation failed: {type(exc).__name__}: {exc}",
)
payload = scrub(report.as_dict())
detail = ""
if not report.inventory_complete:
detail = "; ".join(report.incomplete_reasons) or "inventory incomplete"
return payload, SourceStatus("impact", STATUS_OK, detail)
def load_drain_status(
*,
proof: Mapping[str, Any] | None = None,
now: datetime | None = None,
expected_impact_fingerprint: str | None = None,
) -> tuple[dict[str, Any] | None, SourceStatus]:
"""Verify a supplied drain proof read-only and report the verdict.
No proof supplied is not a failure and not a pass: it is reported as the
absence of a proof, which is exactly what the #661 gate would deny on.
"""
if proof is None:
return None, SourceStatus(
"drain",
STATUS_UNAVAILABLE,
"no drain proof supplied; the #661 gate denies a restart without a "
"valid unexpired clean proof",
)
try:
verified = drain_proof.verify_drain_proof(
proof,
now=now,
expected_impact_fingerprint=expected_impact_fingerprint,
)
except Exception as exc: # noqa: BLE001
return None, SourceStatus(
"drain",
STATUS_UNAVAILABLE,
f"drain proof verification failed: {type(exc).__name__}: {exc}",
)
return scrub(verified.as_dict()), SourceStatus("drain", STATUS_OK)
def load_reconcile_status(
*,
load_proof: Callable[[], Any] | None = None,
) -> tuple[dict[str, Any] | None, SourceStatus]:
"""Report the most recent post-restart completion proof (#662)."""
if load_proof is None:
return None, SourceStatus(
"reconcile",
STATUS_UNAVAILABLE,
"no post-restart completion proof source is wired into this view",
)
try:
proof = load_proof()
except Exception as exc: # noqa: BLE001
return None, SourceStatus(
"reconcile",
STATUS_UNAVAILABLE,
f"reconcile proof unavailable: {type(exc).__name__}: {exc}",
)
if proof is None:
return None, SourceStatus(
"reconcile",
STATUS_UNAVAILABLE,
"no post-restart reconcile has been recorded",
)
payload = proof.as_dict() if hasattr(proof, "as_dict") else dict(proof)
return scrub(payload), SourceStatus("reconcile", STATUS_OK)
def load_restart_console_snapshot(
*,
principal: console_authz.Principal | None = None,
restart_class: str = restart_coordinator.RestartClass.FULL_MCP_RESTART.value,
db_path: str | None = None,
limit: int = 200,
drain_proof_payload: Mapping[str, Any] | None = None,
read_inventory: Callable[..., Mapping[str, Any]] | None = None,
load_reconcile_proof: Callable[[], Any] | None = None,
now: datetime | None = None,
) -> RestartConsoleSnapshot:
"""Compose the read-only restart console snapshot."""
who = principal if principal is not None else console_authz.ANONYMOUS
moment = now or _utc_now()
impact, impact_source = load_impact_report(
principal=who,
restart_class=restart_class,
db_path=db_path,
limit=limit,
read_inventory=read_inventory,
now=moment,
)
fingerprint = None
if impact is not None:
try:
fingerprint = drain_proof.impact_fingerprint(impact)
except Exception: # noqa: BLE001
fingerprint = None
drain, drain_source = load_drain_status(
proof=drain_proof_payload,
now=moment,
expected_impact_fingerprint=fingerprint,
)
reconcile, reconcile_source = load_reconcile_status(
load_proof=load_reconcile_proof
)
notes: list[str] = [
"This surface is read-only: it evaluates and displays, and performs no "
"restart, reload, drain, approval, or process action.",
]
if not impact_source.available:
notes.append(
"Impact preview unavailable — a restart decision must not be made "
"from this page while the blast radius is unknown."
)
return RestartConsoleSnapshot(
generated_at=moment.isoformat(),
viewer_role=str(who.role or "anonymous"),
viewer_authenticated=bool(who.authenticated),
read_only=True,
impact=impact,
impact_source=impact_source,
drain=drain,
drain_source=drain_source,
reconcile=reconcile,
reconcile_source=reconcile_source,
restart_classes=build_restart_class_views(who.role),
authorizations=build_action_authorizations(who),
break_glass=BreakGlassSurface(
available=False,
issue=BREAK_GLASS_ISSUE,
reason=BREAK_GLASS_PENDING_REASON,
viewer_is_privileged=viewer_is_privileged(who),
),
notes=tuple(notes),
)
-299
View File
@@ -1,299 +0,0 @@
"""HTML views for the read-only restart console (#667).
Every interpolated value passes through :func:`_esc`. Values that can carry a
filesystem path or free-form operator text additionally pass through
:func:`webui.inventory.scrub_text`, which redacts credential-shaped tokens
*inside* a string rather than only at its start.
The page renders state and never offers a control that would mutate anything:
the approval and break-glass panels report authorization and availability, and
there is no form, button, or endpoint behind them.
"""
from __future__ import annotations
import html
from webui.inventory import scrub_text
from webui.restart_console import RestartConsoleSnapshot, SourceStatus
def _esc(value: object) -> str:
"""Escape any value for HTML text or a quoted attribute."""
if value is None:
return ""
return html.escape(str(value), quote=True)
def _esc_text(value: object) -> str:
"""Escape free-form text after redacting secrets embedded inside it."""
if value is None:
return ""
return _esc(scrub_text(str(value)))
def _bool_badge(
value: bool, *, true_label: str = "yes", false_label: str = "no"
) -> str:
css = "badge-ok" if value else "badge-blocked"
label = true_label if value else false_label
return f'<span class="badge {css}">{_esc(label)}</span>'
def _source_badge(source: SourceStatus) -> str:
css = "badge-ok" if source.available else "badge-blocked"
badge = f'<span class="badge {css}">{_esc(source.status)}</span>'
if source.detail:
badge += f' <span class="muted">{_esc_text(source.detail)}</span>'
return badge
def _notes_block(snapshot: RestartConsoleSnapshot) -> str:
if not snapshot.notes:
return ""
items = "".join(f"<li>{_esc_text(note)}</li>" for note in snapshot.notes)
return f"<ul class='reasons'>{items}</ul>"
def _impact_section(snapshot: RestartConsoleSnapshot) -> str:
head = (
"<section class='health-card'>"
f"<h3>Impact preview {_source_badge(snapshot.impact_source)}</h3>"
)
impact = snapshot.impact
if impact is None:
return (
head
+ "<p class='muted'>No impact preview is available, so the blast "
"radius of a restart is unknown. Treat this as unsafe.</p></section>"
)
counts = impact.get("counts") or {}
verdict = str(impact.get("verdict") or "unknown")
verdict_css = "badge-ok" if verdict == "safe" else "badge-blocked"
rows = "".join(
f"<tr><th>{_esc(key.replace('_', ' '))}</th><td>{_esc(value)}</td></tr>"
for key, value in sorted(counts.items())
)
reasons = "".join(
f"<li>{_esc_text(reason)}</li>" for reason in (impact.get("reasons") or [])
)
incomplete = ""
if not impact.get("inventory_complete", False):
detail = "; ".join(str(r) for r in (impact.get("incomplete_reasons") or []))
incomplete = (
"<p class='error'><strong>Inventory incomplete:</strong> "
f"{_esc_text(detail or 'unspecified')}. The coordinator fails "
"closed on an incomplete inventory.</p>"
)
sessions = impact.get("affected_sessions") or []
session_rows = "".join(
"<tr>"
f"<td><code>{_esc(s.get('session_id'))}</code></td>"
f"<td>{_esc(s.get('role'))}</td>"
f"<td>{_esc(s.get('pid'))}</td>"
f"<td>{_bool_badge(bool(s.get('live')), true_label='live', false_label='idle')}</td>"
f"<td>{_bool_badge(not s.get('heartbeat_stale'), true_label='fresh', false_label='stale')}</td>"
"</tr>"
for s in sessions[:50]
)
session_table = (
"<h4>Sessions a restart would terminate</h4>"
"<div class='table-scroll'><table class='registry'><thead><tr>"
"<th>Session</th><th>Role</th><th>PID</th><th>State</th>"
"<th>Heartbeat</th></tr></thead><tbody>"
f"{session_rows}</tbody></table></div>"
if session_rows
else "<p class='muted'>No affected sessions reported.</p>"
)
truncated = (
f"<p class='muted'>Showing the first 50 of {_esc(len(sessions))} "
"affected sessions.</p>"
if len(sessions) > 50
else ""
)
return (
head
+ "<p class='health-headline'>Verdict "
f"<span class='badge {verdict_css}'>{_esc(verdict)}</span> · "
f"blast radius <code>{_esc(impact.get('blast_radius'))}</code> · "
f"class <code>{_esc(impact.get('restart_class'))}</code></p>"
+ incomplete
+ (f"<ul class='reasons'>{reasons}</ul>" if reasons else "")
+ (f"<table class='registry'><tbody>{rows}</tbody></table>" if rows else "")
+ session_table
+ truncated
+ "</section>"
)
def _drain_section(snapshot: RestartConsoleSnapshot) -> str:
head = (
"<section class='health-card'>"
f"<h3>Drain proof {_source_badge(snapshot.drain_source)}</h3>"
)
drain = snapshot.drain
if drain is None:
return (
head
+ "<p class='muted'>No drain proof has been presented to this view. "
"The #661 gate authorizes a restart only against a valid, unexpired, "
"clean proof, so the absence of one is a denial, not a pass.</p>"
"</section>"
)
reasons = "".join(
f"<li>{_esc_text(reason)}</li>" for reason in (drain.get("reasons") or [])
)
return (
head
+ "<table class='registry'><tbody>"
f"<tr><th>Valid</th><td>{_bool_badge(bool(drain.get('valid')))}</td></tr>"
f"<tr><th>Clean</th><td>{_bool_badge(bool(drain.get('clean')))}</td></tr>"
f"<tr><th>Expired</th><td>{_bool_badge(not drain.get('expired'), true_label='no', false_label='yes')}</td></tr>"
f"<tr><th>Tampered</th><td>{_bool_badge(not drain.get('tampered'), true_label='no', false_label='yes')}</td></tr>"
f"<tr><th>Proof id</th><td><code>{_esc(drain.get('proof_id'))}</code></td></tr>"
"</tbody></table>"
+ (f"<ul class='reasons'>{reasons}</ul>" if reasons else "")
+ "</section>"
)
def _reconcile_section(snapshot: RestartConsoleSnapshot) -> str:
head = (
"<section class='health-card'>"
f"<h3>Post-restart reconcile {_source_badge(snapshot.reconcile_source)}</h3>"
)
proof = snapshot.reconcile
if proof is None:
return (
head
+ "<p class='muted'>No post-restart completion proof is recorded. "
"Until one is, the last restart's recovery state is unproven.</p>"
"</section>"
)
items = "".join(
"<tr>"
f"<td>{_esc(item.get('dimension'))}</td>"
f"<td>{_esc(item.get('status'))}</td>"
f"<td>{_esc_text(item.get('summary'))}</td>"
f"<td>{_bool_badge(not item.get('follow_up_required'), true_label='no', false_label='yes')}</td>"
"</tr>"
for item in (proof.get("items") or [])
)
return (
head
+ "<p class='health-headline'>Status "
f"<code>{_esc(proof.get('overall_status'))}</code> · mode "
f"<code>{_esc(proof.get('mode'))}</code> · resolved "
f"{_esc(proof.get('resolved_count'))} · unresolved "
f"{_esc(proof.get('unresolved_count'))}</p>"
+ (
"<div class='table-scroll'><table class='registry'><thead><tr>"
"<th>Dimension</th><th>Status</th><th>Summary</th>"
"<th>Follow-up required</th></tr></thead><tbody>"
f"{items}</tbody></table></div>"
if items
else "<p class='muted'>No reconcile dimensions reported.</p>"
)
+ "</section>"
)
def _class_matrix_section(snapshot: RestartConsoleSnapshot) -> str:
rows = "".join(
"<tr>"
f"<td><code>{_esc(view.restart_class)}</code></td>"
f"<td><code>{_esc(view.required_permission)}</code></td>"
f"<td>{_esc(view.expected_blast_radius)}</td>"
f"<td>{_esc(view.drain_requirement)}</td>"
f"<td>{_esc(view.approval_requirement)}</td>"
f"<td>{_bool_badge(view.viewer_may_request)}</td>"
f"<td>{_bool_badge(view.viewer_may_execute)}</td>"
"</tr>"
for view in snapshot.restart_classes
)
return (
"<section class='health-card'>"
"<h3>Restart classes</h3>"
"<p class='muted'>The least-privilege matrix each restart request is "
"resolved against. &ldquo;You may request&rdquo; and &ldquo;you may "
"execute&rdquo; are computed for the current viewer role, not for a "
"generic operator.</p>"
"<div class='table-scroll'><table class='registry'><thead><tr>"
"<th>Class</th><th>Permission</th><th>Blast radius</th>"
"<th>Drain</th><th>Approval</th><th>You may request</th>"
"<th>You may execute</th></tr></thead><tbody>"
f"{rows}</tbody></table></div>"
"</section>"
)
def _approval_section(snapshot: RestartConsoleSnapshot) -> str:
rows = "".join(
"<tr>"
f"<td><code>{_esc(a.action_id)}</code></td>"
f"<td>{_esc(a.required_role)}</td>"
f"<td>{_bool_badge(a.allowed)}</td>"
f"<td>{_bool_badge(a.execution_enabled)}</td>"
f"<td><code>{_esc(a.reason_code)}</code></td>"
f"<td>{_esc_text(a.detail)}</td>"
"</tr>"
for a in snapshot.authorizations
)
return (
"<section class='health-card'>"
"<h3>Approval controls</h3>"
"<p class='muted'>Authorization is probed the way execution would probe "
"it, so &ldquo;execution enabled&rdquo; answers whether the action would "
"actually run — not merely whether this role outranks the requirement. "
"No control on this page performs the action.</p>"
"<div class='table-scroll'><table class='registry'><thead><tr>"
"<th>Action</th><th>Required role</th><th>Authorized</th>"
"<th>Execution enabled</th><th>Reason</th><th>Detail</th>"
"</tr></thead><tbody>"
f"{rows}</tbody></table></div>"
"</section>"
)
def _break_glass_section(snapshot: RestartConsoleSnapshot) -> str:
bg = snapshot.break_glass
if not bg.viewer_is_privileged:
return (
"<section class='health-card'>"
"<h3>Break-glass</h3>"
"<p class='muted'>Break-glass status is visible to operator-class "
"roles only. Your role does not carry that authority, so no "
"emergency surface is shown.</p>"
"</section>"
)
return (
"<section class='health-card'>"
"<h3>Break-glass "
f"{_bool_badge(bg.available, true_label='available', false_label='unavailable')}"
"</h3>"
f"<p class='muted'>{_esc_text(bg.reason)}</p>"
f"<p class='meta'>Tracked by issue #{_esc(bg.issue)}.</p>"
"</section>"
)
def render_restart_console_page(snapshot: RestartConsoleSnapshot) -> str:
"""Render the whole read-only restart console body."""
return (
"<h2>Restart status and impact</h2>"
f"<p class='meta'>Generated <code>{_esc(snapshot.generated_at)}</code> · "
f"viewer role <code>{_esc(snapshot.viewer_role)}</code> · "
f"authenticated {_bool_badge(snapshot.viewer_authenticated)} · "
f"read-only {_bool_badge(snapshot.read_only)}</p>"
+ _notes_block(snapshot)
+ _impact_section(snapshot)
+ _drain_section(snapshot)
+ _reconcile_section(snapshot)
+ _class_matrix_section(snapshot)
+ _approval_section(snapshot)
+ _break_glass_section(snapshot)
)