Compare commits

..
Author SHA1 Message Date
sysadmin f0c6255d7d Merge remote-tracking branch 'prgs/master' into feat/issue-659-maintenance-drain-mode 2026-07-25 18:40:57 -04:00
sysadmin d7ad2838ec Merge pull request 'docs(incident): retroactive audit for direct-to-master commit 2fa97c26 (#670)' (#915) from fix/issue-670-direct-master-incident into master 2026-07-25 17:40:23 -05:00
sysadmin c6d68dbc7b Merge pull request 'feat(webui): notifications and human-attention routing (#648)' (#905) from feat/issue-648-notifications-console into master 2026-07-25 17:40:03 -05:00
sysadmin c83a10d7c2 Merge remote-tracking branch 'prgs/master' into feat/issue-648-notifications-console 2026-07-25 18:37:17 -04:00
sysadmin ca22c326a4 Merge pull request 'docs(architecture): ADR for high-availability and rolling-restart MCP architecture (Closes #668)' (#912) from docs/issue-668-mcp-ha-rolling-restart into master 2026-07-25 17:34:31 -05:00
sysadmin 3bbe6df6c7 Merge pull request 'feat(webui): add read-only console restart status and impact controls (Closes #667)' (#911) from feat/issue-667-console-restart-controls into master 2026-07-25 17:34:12 -05:00
jcwalker3 71031c812e Merge branch 'master' into feat/issue-648-notifications-console 2026-07-25 17:29:57 -05:00
jcwalker3 e43ddd3cbe docs(incident): retroactive audit for direct-to-master commit 2fa97c26 (#670) 2026-07-25 17:29:34 -05:00
sysadmin a64ba08e27 fix(webui): address #905 REQUEST_CHANGES on notifications classifier
B1: classify_attention_event uses structured flags/category only — never
substring-match human-authored title/summary for escalation.

B2: make notification ids unique across probe_errors and collisions
(include loop index / kind).

B3: do not assign probe_errors to fetch_error (avoids false Fetch Warning
and double-reporting).

Regression tests cover all three blockers.

Refs #648
2026-07-25 18:26:46 -04:00
sysadmin 26f54851d1 Merge pull request 'feat(mcp): enforce scoped recovery playbook before full restart (Closes #669)' (#914) from feat/issue-669-scoped-component-recovery into master 2026-07-25 17:14:21 -05:00
sysadmin f02a2dc030 Merge pull request 'feat(webui): Web Console stale-runtime recovery & reconciliation controls (Phase 2) (#644)' (#903) from feat/issue-644-console-recovery into master 2026-07-25 17:14:06 -05:00
sysadminandClaude Opus 4.8 77d808e7d4 merge(master): resolve PR #907 allocator drain vs side_effect_free
Keep #659 maintenance-drain assignment stop and #643 side_effect_free
apply guard; session registration remains gated by side_effect_free.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-07-25 18:12:03 -04:00
sysadminandClaude Opus 4.8 bb8c3a537b merge(master): resolve PR #905 conflicts with requests/linkage
Keep notifications (#648) routes and nav alongside master requests (#643)
and other base updates.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-07-25 18:11:06 -04:00
sysadminandClaude Opus 4.8 04ae3532cc merge(master): resolve #903 conflicts with #643 request initiation
Keep #644 recovery console actions and #643 initiate_workflow side by side
in console_authz and authz audit docs after merging latest master.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-07-25 18:03:04 -04:00
sysadmin 7bb5ff4719 Merge pull request 'feat(webui): request preview, authorization, and workflow initiation (Closes #643)' (#902) from feat/issue-643-request-preview-initiate into master 2026-07-25 17:02:17 -05:00
sysadmin 5b7ceefa9a Merge remote-tracking branch 'prgs/master' into feat/issue-644-console-recovery 2026-07-25 18:00:31 -04:00
sysadminandClaude Opus 4.8 4a2fae8495 fix(webui): arm the recovery gates and make the playbooks reach the process (#644)
Reviewer REQUEST_CHANGES on PR #903 at head 1c88b87 raised five blockers, all
reproduced by executing that head. The shared shape: a write path that declared
itself gated, audited, and verified, but never armed the gate, mutated a copy of
the state it claimed to fix, and then verified against that same copy.

B1 - the apply path never asked the execution gate.
execute_recovery_playbook called console_authz.authorize with the default
for_execution=False, and the phase branch only fires when it is True. ACTIVE_PHASE
is 1 and every new action is phase 2, so an operator executed a phase-2 write
through POST /api/v1/system/recovery/apply while build_recovery_preview reported
execution_enabled false. The call now passes for_execution=True and surfaces the
phase_not_active refusal. Preview reports the same decision under
execution_authorization / execution_blocked_reason instead of a hardcoded False it
could not explain.

B2 - both env playbooks mutated a discarded copy and verified against it.
source_env = dict(os.environ) meant clear_stale_binding and rebind_session_worktree
never touched the running process, and verify_post_recovery(env=source_env)
re-diagnosed the same copy, confirming a change that had not happened. Mutations
now target the live mapping (apply_recovery's sanctioned env=None -> os.environ
path, #702 AC2) and verification re-reads state rather than the mutated input.
binding_before / binding_after / binding_changed are returned, and a playbook that
changed nothing reports performed: false. verify_post_recovery no longer reads an
unverified_inherited binding as clean, because unproven is not clean.

B3 - the reconcile playbook called a function that does not exist.
merged_cleanup_reconcile.reconcile_merged_cleanups is absent from that module and a
bare except turned the AttributeError into a generic failure, so the playbook could
never succeed. It now calls gitea_mcp_server.gitea_reconcile_merged_cleanups, the
real orchestrator, imported lazily; failures carry error_type. task_capability_map
declared gitea.pr.close for reconcile_cleanups while the entry point gates on
gitea.read; the two authority statements are reconciled to the one that is enforced.

B4 - the #630 contamination integration could not block.
assess_contamination_gate was fed marker=None, which short-circuits to block: False
on its first statement; the task passed was a console action id outside
CONTAMINATION_GATED_TASKS; and the result was read through a "contaminated" key the
gate never returns, making STATUS_BLOCKED_CONTAMINATION unreachable. The live marker
now comes from the #641 session inventory reader, the gated task key
console_recovery_apply is added to CONTAMINATION_GATED_TASKS, every read uses the
"block" key the gate actually returns, and the marker is forwarded to
sanctioned_restart.execute_restart so a restart cannot launder a contaminated
runtime. The reconciler cleanup playbook stays exempt as the designated remedy.

B5 - the parity baseline was captured from the head it was compared against.
capture_startup_parity(root, head=checkout_head) stores the head verbatim, so
in_parity was structurally incapable of being false, and live_remote_head was never
passed. The baseline is now the daemon start head that assess_stale_runtime already
returns, and the #610 live-remote dimension is restored.

Also: _recovery_card was the one renderer in system_health_views.py interpolating
without _esc(), and it is where a marker's operator-supplied command_summary lands
once B4 is wired; it now escapes, including the except branch. Docs no longer claim
apply enforces master parity or that verify asserts clean: true, and the absolute
file:///Users/... links are relative.

Tests: the two that asserted the defects as intended are inverted -
test_api_recovery_apply_with_dev_auth asserted the phase-gate bypass, and the rebind
test asserted the input echoed back. Added coverage per blocker, including a no-op
detection test that fails when a playbook reports success without changing anything,
the previously untested reconcile playbook, contamination block and remedy-exemption
tests, and parity baseline/live-remote tests. Both new guards were mutation-verified:
disarming for_execution fails 2 tests, restoring the env copy fails 2 tests.

Validation: WEBUI_TEST_OFFLINE=1 ../../venv/bin/python -m pytest tests/ -q from
branches/feat-issue-644 gives 27 failed / 5291 passed / 6 skipped / 953 subtests;
the same command from branches/baseline-master-76f293e at 76f293eb28 gives
28 failed / 5262 passed / 6 skipped / 926 subtests. Suites run one at a time.
comm of the sorted FAILED lines shows no new signature at the head. The single
absent signature, test_workspace_guard_alignment.py::
TestRuntimeContextGuardAlignment::test_declared_branches_worktree_passes_when_mcp_root_differs,
is suite-order dependent: that file passes 9/9 in isolation at both revisions.

Closes #644

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-07-25 17:42:49 -04:00
sysadminandGrok 4.5 461e1dac78 feat(mcp): enforce scoped recovery playbook before full restart (Closes #669)
Add recovery_playbook.py with the narrow-to-broad recovery ladder, symptom
routing, attempt-log helpers, and escalation metrics. Wire the attempt-log
gate into restart_coordinator so rolling/full/host restarts require prior
insufficient narrower attempts (or break-glass). Document the ladder and
update gitea_request_mcp_restart for prior_recovery_attempts_json.

Co-Authored-By: Grok 4.5 <[email protected]>
2026-07-25 17:35:11 -04:00
sysadmin 9c69bfcd80 Merge pull request 'feat(webui): Gitea issue and PR linkage console (Closes #645)' (#904) from feat/issue-645-linkage-console into master 2026-07-25 16:32:21 -05:00
sysadmin 6010f4295b docs(architecture): add ADR for HA and rolling-restart MCP architecture (Closes #668) 2026-07-25 17:27:50 -04:00
sysadminandClaude Opus 4.8 e91b94db56 feat(mcp): implement graceful maintenance-drain mode (Closes #659)
Add durable per-repo drain state, allocator assignment stop, mutation
deferral with a safety allowlist, observable status, and capability-gated
enter/exit tools. Drain proof/restart gate remain #661.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-07-25 17:02:12 -04:00
sysadmin 4f06d30e07 feat(webui): implement notifications and human-attention routing (#648) 2026-07-25 16:41:16 -04:00
sysadminandClaude Opus 4.8 211890f361 feat(webui): Gitea issue and PR linkage console (Closes #645)
Add a Phase 3 read-only console that resolves issue↔PR linkage with
evidence (closes keyword, branch marker, body mention), surfaces the
latest canonical handoff for a focused thread, and deep-links to Gitea
only under the admin reveal opt-in.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-07-25 16:40:07 -04:00
sysadmin 1c88b87ec5 feat(webui): implement Phase 2 recovery controls & playbooks (#644) 2026-07-25 08:19:06 -04:00
jcwalker3 b993ad1c64 Merge branch 'master' into feat/issue-643-request-preview-initiate 2026-07-25 07:13:21 -05:00
jcwalker3 d0006e9f71 Merge branch 'master' into feat/issue-643-request-preview-initiate 2026-07-25 06:42:41 -05:00
jcwalker3 6da68fffb8 Merge branch 'master' into feat/issue-643-request-preview-initiate 2026-07-25 02:44:58 -05:00
sysadminandClaude Opus 5 53ce1b1a5e fix(webui): compensate stray allocations and keep preview side-effect free
Addresses both blockers from the PR #902 review (review 589) for issue #643.

B1 — an allocator-created assignment could be orphaned and reported as no
mutation.

apply_request re-previews, then re-runs the allocator with apply=True. The CAS
fingerprint hashes only {kind, number} plus exclusions, so a competing lease
taken on the requested unit inside the window leaves the fingerprint identical:
the pin passes, the selection loop skips the now-claimed unit and commits an
assignment on the *next* one, and _selection_matches then fails on egress. The
old code returned mutation_performed False with that lease still committed and
owned by a synthetic session nothing heartbeats. The window contains a second
full load_queue_snapshot(), so it is seconds wide, and foreign sessions acting
on this repo concurrently are an observed condition.

The egress mismatch now releases the assignment the allocator created before
refusing. When the release succeeds the refusal reports mutation_performed
False and a compensation record; when it fails the response carries
mutation_performed True, an explicit orphaned_assignment, and a
gitea_release_workflow_lease reclaim action, because claiming nothing changed
while a lease is live is the defect rather than a report of it.

The same state was reachable through _run_allocator's bare except Exception:
allocate_next_work only catches InvalidWorkKindError, LeaseRequiredError and
ControlPlaneError, so anything raised after assign_and_lease committed arrived
as "no result" with a durable lease. A None result on the apply path now sweeps
and releases whatever this flow's session owns. That sweep is only possible
because of the B2 fix below — the session id is now stable across the flow, so
the lease is findable.

B2 — the "read-only" preview wrote to the control-plane DB.

allocate_next_work called db.upsert_session and db.expire_stale_leases
unconditionally, before the apply branch was consulted, and default_allocator
minted a fresh webui-request-<hex> per call. Every preview therefore appended a
never-reused session row and mutated global lease state while the payload said
dry_run True / mutation_performed False, driven by an operator refreshing a
form. One apply wrote two rows and bound the lease to the second, which is why
B1's orphan had no reclaimable owner.

allocate_next_work gains a keyword-only side_effect_free flag, default False so
every existing caller is byte-for-byte unchanged. Under the flag both writes are
suppressed and expired leases are instead filtered out of the claim map in
memory, which reaches the same selection the sweep would have produced without
persisting anything; a claim whose expiry cannot be parsed is kept, since an
unreadable expiry is not evidence that work is free. side_effect_free with
apply=True fails closed rather than silently reserving. request_service mints
one session id per request flow and threads it through both the dry-run and the
apply, and the dry-run now routes through the side-effect-free path.

Coverage.

test_allocator_drift_on_apply_is_not_read_as_an_assignment asserted the defect —
it built the orphan state and then required mutation_performed to be False, which
a leak satisfies. It now requires the compensating release. Added: release-failure
surfacing a reclaim action, an assignment with no lease id, the post-commit
exception route, and session-id identity across the flow. The two areas the review
named as having zero coverage now have it: default_allocator past its two
fail-closed early returns (side_effect_free routing, session-id pass-through and
minting, scope and fingerprint propagation) and default_claims_source (scoped
read, and an unreadable substrate denying rather than reading as "nothing
claimed"). allocator_service gains side-effect-free tests against a real
temp-file DB plus the expiry-filter unit tests.

Every new guard was mutation-tested by reverting it one at a time: B1
compensation removed → 3 failures; post-commit sweep removed → 1; session id
re-minted per call → 2; side_effect_free ignored → 3; in-memory expiry filter
removed → 1.

Verification: WEBUI_TEST_OFFLINE=1 python -m pytest tests/ -q from this
branches/ worktree gives 23 failed, 5263 passed, 6 skipped, 899 subtests. The
sorted FAILED set is identical to the reviewer's clean-master baseline at
2f4dec83 (23 failed, 5190 passed) — no new, changed, or disappeared failure —
and 21 tests were added over the reviewed head's 5242. Zero conflict markers;
py_compile passes; git diff --check clean.

Refs #643, PR #902

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01V6xFqovhbArPv61j9KCGkL
2026-07-25 03:32:24 -04:00
sysadminandClaude Opus 4.8 433f66add8 feat(webui): request preview, authorization, and workflow initiation (Closes #643)
Operators had to paste a role prompt into a terminal to start work, and
nothing enforced that the allocator had been consulted first, so two sessions
could reach for the same issue and each believe it was theirs. This adds a
request surface: a desired role, an issue or PR, and a stated intent, answered
by an authorization decision and - on confirmation - an exclusive assignment
from the allocator.

Preview (POST /api/v1/requests/preview, and the /requests form) runs five
checks and reports authorize/deny with a reason for each: console
authorization, capability resolution for the desired role, lease availability,
whether the allocator would independently select this work unit, and head
pinning for PR work. It is read-only - it calls the allocator with apply=false
and writes only an audit line. An unauthorized principal never reaches the
allocator or the control-plane DB, so a denial cannot enumerate the queue.

Initiation (POST /api/v1/requests/apply) never assigns the requested item
directly. It runs a dry-run first and proceeds only when the allocator would
independently pick that exact work unit, carrying the dry-run's
candidate_set_fingerprint as a CAS pin; otherwise it returns wait or blocked
and mutates nothing. An active claim on the work unit rejects a duplicate
assign before one is attempted. A returned assignment carries a handoff block
naming the required profile, namespace, and the actions that stay forbidden.

Authorization reuses the #633 model rather than adding a second one. The new
initiate_workflow action is operator-class because its outcome is a claim, not
a Gitea verdict: requesting reviewer or merger work reserves that work but
grants no right to approve or merge. Execution is gated by a new per-action
execution_env_flag (WEBUI_REQUESTS_EXECUTION), deliberately in place of raising
ACTIVE_PHASE - a phase bump would enable execution for every phase-2 action at
once, including ones whose execution path is not implemented. Actions that
declare no flag are unchanged and still report execution_enabled false.

Every preview and apply emits a console audit record correlated to the
resulting assignment by correlation.request_id.

Fail-closed throughout: an unreadable control-plane DB, an incomplete queue
inventory (#758), an allocator that raises, an unpinned PR head, a moved PR
head, and an unconfirmed apply all deny without mutating.

Files:
- webui/request_service.py (new) - request model, preview, initiation
- webui/request_views.py (new) - form and preview rendering, escaped
- tests/test_webui_request_initiation.py (new) - 52 tests
- webui/console_authz.py - initiate_workflow action, execution_wired()
- webui/app.py - /requests, /api/v1/requests/preview, /api/v1/requests/apply
- webui/nav.py - Requests nav entry
- webui/traffic_loader.py - public candidates_from_queue_snapshot alias
- docs/webui-requests.md (new), docs/webui-authz-audit.md

Validation: full suite on this branch 5242 passed, 6 skipped, 899 subtests, 23
failed. Clean master baseline at 2f4dec83 in an equivalent branches/ worktree:
5190 passed, 6 skipped, 867 subtests, the same 23 tests failed. The branch adds
52 passing tests and introduces no new full-suite failure signature.

Closes #643

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-07-25 01:47:38 -04:00
40 changed files with 10048 additions and 73 deletions
+131 -16
View File
@@ -23,8 +23,10 @@ import json
import os
import uuid
from dataclasses import dataclass, field
from datetime import datetime, timezone
from typing import Any, Mapping, Sequence
import maintenance_drain
from control_plane_db import (
ControlPlaneDB,
ControlPlaneError,
@@ -738,6 +740,46 @@ def normalize_exclude_issue_numbers(
return sorted(out)
def _claim_expires_at(claim: Any) -> datetime | None:
"""Parse a claim's ``expires_at``, or ``None`` when it is absent/malformed."""
if not isinstance(claim, Mapping):
return None
text = str(claim.get("expires_at") or "").strip()
if not text:
return None
if text.endswith("Z"):
text = text[:-1] + "+00:00"
try:
parsed = datetime.fromisoformat(text)
except ValueError:
return None
if parsed.tzinfo is None:
parsed = parsed.replace(tzinfo=timezone.utc)
return parsed.astimezone(timezone.utc)
def _drop_expired_claims(
claims: Mapping[tuple[str, int], dict[str, Any]],
*,
now: datetime | None = None,
) -> dict[tuple[str, int], dict[str, Any]]:
"""Claims minus those whose lease has already expired (#643).
The read-only mirror of ``expire_stale_leases``: the sweep marks such rows
``expired`` so they stop being returned as claims, and this reaches the same
view without writing. A claim with no parseable ``expires_at`` is **kept** —
an unreadable expiry is not evidence that work is free.
"""
moment = now or datetime.now(timezone.utc)
kept: dict[tuple[str, int], dict[str, Any]] = {}
for key, claim in (claims or {}).items():
expires_at = _claim_expires_at(claim)
if expires_at is not None and expires_at <= moment:
continue
kept[key] = claim
return kept
def candidate_set_fingerprint(
candidates: Sequence[WorkCandidate],
*,
@@ -826,12 +868,22 @@ def allocate_next_work(
exclude_issue_numbers: Sequence[int] | None = None,
expected_candidate_set_fingerprint: str | None = None,
allocation_mode: str | None = None,
side_effect_free: bool = False,
) -> dict[str, Any]:
"""Select and optionally reserve the next work unit via control-plane DB.
*apply=False* (default): dry-run selection only — no lease/assignment.
*apply=True*: atomic ``assign_and_lease`` for the selected candidate.
*side_effect_free* (#643): a dry run that writes **nothing** to the
control-plane DB. A plain ``apply=False`` still registered a session row and
swept stale leases globally, so a caller advertising a read-only preview was
mutating on every call. Under this flag both writes are suppressed and stale
leases are instead filtered out of the claim map in memory, which yields the
same selection the sweep would have produced without persisting anything.
Incompatible with *apply* — the combination fails closed rather than
silently reserving.
*allocation_mode* (#840): ``cross_role`` (default for controller) inspects
the complete queue and returns one authoritative selection naming the
required downstream role/profile/action. ``role_scoped`` keeps prior
@@ -885,41 +937,98 @@ def allocate_next_work(
"allocation_mode": (allocation_mode or "").strip() or None,
}
session_id = (session_id or "").strip() or f"alloc-{uuid.uuid4().hex[:12]}"
# #659 AC2: while maintenance drain is active, no new work is assigned —
# for dry-run and apply alike, so a preview can never be read as evidence
# that work was assignable during the drain. Checked before session
# registration so a drained allocator leaves no new state behind.
try:
db.upsert_session(
session_id=session_id,
role=role_norm,
profile=profile_name,
pid=os.getpid(),
controller_instance_id=controller_instance_id,
)
except Exception as exc: # noqa: BLE001 — surface structured
drain_record = db.read_maintenance_drain(remote=remote, org=org, repo=repo)
except Exception as exc: # noqa: BLE001 — unreadable drain state fails closed
return {
"success": False,
"outcome": OUTCOME_NO_SAFE,
"reasons": [
f"failed to register session in control-plane DB: {exc} "
"(fail closed, #613)"
f"maintenance-drain state lookup failed: {exc} (fail closed, #659)"
],
"skipped": [],
"assignment": None,
"substrate": "control_plane_db",
}
# Expire stale leases globally before selection.
try:
db.expire_stale_leases()
except Exception as exc: # noqa: BLE001
drain_decision = maintenance_drain.classify_assignment(drain_record)
if not drain_decision["assignment_allowed"]:
return {
"success": True,
"outcome": OUTCOME_WAIT,
"apply": apply,
"role": role_norm,
"allocation_mode": mode,
"remote": remote,
"org": org,
"repo": repo,
"selected": None,
"reasons": list(drain_decision["reasons"]),
"reason_code": drain_decision["reason_code"],
"skipped": [],
"assignment": None,
"substrate": "control_plane_db",
"maintenance_drain": maintenance_drain.status_payload(
drain_record, remote=remote, org=org, repo=repo
),
}
# A side-effect-free run may never reserve: reserving is a write, and the
# flag is the caller's assertion that this call writes nothing (#643).
if side_effect_free and apply:
return {
"success": False,
"outcome": OUTCOME_NO_SAFE,
"reasons": [f"lease expiry failed: {exc} (fail closed)"],
"apply": True,
"reasons": [
"side_effect_free is incompatible with apply=True; an "
"assignment is a write (fail closed, #643)"
],
"skipped": [],
"assignment": None,
"substrate": "control_plane_db",
}
session_id = (session_id or "").strip() or f"alloc-{uuid.uuid4().hex[:12]}"
if not side_effect_free:
try:
db.upsert_session(
session_id=session_id,
role=role_norm,
profile=profile_name,
pid=os.getpid(),
controller_instance_id=controller_instance_id,
)
except Exception as exc: # noqa: BLE001 — surface structured
return {
"success": False,
"outcome": OUTCOME_NO_SAFE,
"reasons": [
f"failed to register session in control-plane DB: {exc} "
"(fail closed, #613)"
],
"skipped": [],
"assignment": None,
"substrate": "control_plane_db",
}
# Expire stale leases globally before selection.
try:
db.expire_stale_leases()
except Exception as exc: # noqa: BLE001
return {
"success": False,
"outcome": OUTCOME_NO_SAFE,
"reasons": [f"lease expiry failed: {exc} (fail closed)"],
"skipped": [],
"assignment": None,
"substrate": "control_plane_db",
}
terminal = None
try:
terminal = db.get_active_terminal_lock(remote=remote, org=org, repo=repo)
@@ -953,6 +1062,12 @@ def allocate_next_work(
"assignment": None,
"substrate": "control_plane_db",
}
if side_effect_free:
# ``list_active_claims`` filters on status alone, so without the
# global sweep an already-expired lease would still read as a live
# claim and the preview would report work as taken that is free.
# Drop those in memory: same view the sweep produces, no write.
claims = _drop_expired_claims(claims)
try:
exclude_nums = normalize_exclude_issue_numbers(exclude_issue_numbers)
+194 -1
View File
@@ -31,8 +31,9 @@ from typing import Any, Iterator, Sequence
import dependency_graph
import gitea_audit
import maintenance_drain
SCHEMA_VERSION = 5
SCHEMA_VERSION = 6
# Assignable work kinds only — raw monitoring incidents are never work items.
WORK_KINDS = frozenset({"issue", "pr"})
@@ -239,6 +240,31 @@ CREATE INDEX IF NOT EXISTS idx_session_checkpoints_session
CREATE INDEX IF NOT EXISTS idx_session_checkpoints_work
ON session_checkpoints(remote, org, repo, work_kind, work_number);
-- Graceful maintenance-drain state (#659). One current row per repository
-- scope — drain is a *state*, not a history, so entering and exiting update
-- the same row and every transition is audited to ``events``. Creating the
-- table is the v5->v6 migration: additive, idempotent, and it never touches
-- prior tables. ``state`` is CHECK-constrained so an unknown value can never
-- be written and later read as "not draining".
CREATE TABLE IF NOT EXISTS maintenance_drain (
drain_id TEXT PRIMARY KEY,
remote TEXT NOT NULL,
org TEXT NOT NULL,
repo TEXT NOT NULL,
state TEXT NOT NULL DEFAULT 'inactive'
CHECK (state IN ('inactive', 'draining')),
reason TEXT NOT NULL DEFAULT '',
requested_by TEXT NOT NULL DEFAULT '',
requested_by_profile TEXT NOT NULL DEFAULT '',
session_id TEXT NOT NULL DEFAULT '',
entered_at TEXT NOT NULL DEFAULT '',
exited_at TEXT NOT NULL DEFAULT '',
drain_schema_version INTEGER NOT NULL DEFAULT 6,
created_at TEXT NOT NULL,
updated_at TEXT NOT NULL,
UNIQUE (remote, org, repo)
);
-- Model usage, token cost, latency, and performance events (#651)
CREATE TABLE IF NOT EXISTS usage_events (
usage_id INTEGER PRIMARY KEY AUTOINCREMENT,
@@ -3025,3 +3051,170 @@ class ControlPlaneDB:
"live_lease_id": None if live_lease_id is None else str(live_lease_id),
"reconcile_action": "reconcile_required" if stale else "safe_to_resume",
}
# ── Maintenance drain (#659) ─────────────────────────────────────────────
@staticmethod
def _maintenance_drain_row(row: sqlite3.Row | None) -> dict[str, Any] | None:
"""Convert a ``maintenance_drain`` row to a plain record."""
if row is None:
return None
return {key: row[key] for key in row.keys()}
def read_maintenance_drain(
self, *, remote: str, org: str, repo: str
) -> dict[str, Any] | None:
"""Return the current drain record for a scope, or None if never set.
None and a stored ``inactive`` row mean the same thing to callers —
``maintenance_drain.is_draining`` treats both as not draining — so the
read never has to invent a record to answer the gate.
"""
with self._tx(immediate=False) as conn:
row = conn.execute(
"""
SELECT * FROM maintenance_drain
WHERE remote = ? AND org = ? AND repo = ?
""",
(str(remote or ""), str(org or ""), str(repo or "")),
).fetchone()
return self._maintenance_drain_row(row)
def set_maintenance_drain(
self,
*,
remote: str,
org: str,
repo: str,
state: str,
reason: str = "",
requested_by: str = "",
requested_by_profile: str = "",
session_id: str = "",
) -> dict[str, Any]:
"""Enter or exit maintenance drain for one repository scope (AC1).
The state transition is audited to ``events`` — entering and exiting
are exactly the moments an operator has to be able to reconstruct
later. Re-entering an already-draining scope is idempotent: it refreshes
the reason/owner metadata, keeps the original ``entered_at``, and
records no duplicate transition event.
Capability authorization happens above this layer (the drain tasks
carry a non-``gitea.*`` permission in the task capability map); the DB
records who asked and why, and never grants the right itself.
"""
state_norm = maintenance_drain.normalize_state(state)
raw = {
"reason": str(reason or ""),
"requested_by": str(requested_by or ""),
"requested_by_profile": str(requested_by_profile or ""),
"session_id": str(session_id or ""),
}
clean = gitea_audit.redact(raw)
remote_s, org_s, repo_s = str(remote or ""), str(org or ""), str(repo or "")
now_s = _ts()
with self._tx() as conn:
existing = conn.execute(
"""
SELECT * FROM maintenance_drain
WHERE remote = ? AND org = ? AND repo = ?
""",
(remote_s, org_s, repo_s),
).fetchone()
prior_state = (
maintenance_drain.normalize_state(existing["state"])
if existing is not None
else maintenance_drain.STATE_INACTIVE
)
transitioned = prior_state != state_norm
prior_entered = (
str(existing["entered_at"] or "") if existing is not None else ""
)
prior_exited = (
str(existing["exited_at"] or "") if existing is not None else ""
)
if state_norm == maintenance_drain.STATE_DRAINING:
# A re-entry keeps the original entry time (the drain never
# stopped); a fresh entry stamps now and clears the old exit.
entered_at = prior_entered if (not transitioned and prior_entered) else now_s
exited_at = ""
else:
entered_at = prior_entered
exited_at = now_s if (transitioned or not prior_exited) else prior_exited
if existing is None:
drain_id = uuid.uuid4().hex
conn.execute(
"""
INSERT INTO maintenance_drain(
drain_id, remote, org, repo, state, reason,
requested_by, requested_by_profile, session_id,
entered_at, exited_at, drain_schema_version,
created_at, updated_at
) VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)
""",
(
drain_id, remote_s, org_s, repo_s, state_norm,
clean["reason"], clean["requested_by"],
clean["requested_by_profile"], clean["session_id"],
entered_at, exited_at,
maintenance_drain.DRAIN_SCHEMA_VERSION, now_s, now_s,
),
)
else:
drain_id = str(existing["drain_id"])
conn.execute(
"""
UPDATE maintenance_drain
SET state = ?, reason = ?, requested_by = ?,
requested_by_profile = ?, session_id = ?,
entered_at = ?, exited_at = ?,
drain_schema_version = ?, updated_at = ?
WHERE drain_id = ?
""",
(
state_norm, clean["reason"], clean["requested_by"],
clean["requested_by_profile"], clean["session_id"],
entered_at, exited_at,
maintenance_drain.DRAIN_SCHEMA_VERSION, now_s, drain_id,
),
)
if transitioned:
event_type = (
"maintenance_drain_enter"
if state_norm == maintenance_drain.STATE_DRAINING
else "maintenance_drain_exit"
)
conn.execute(
"""
INSERT INTO events(work_item_id, event_type, message, created_at)
VALUES (NULL, ?, ?, ?)
""",
(
event_type,
f"drain {drain_id} scope {remote_s}/{org_s}/{repo_s} "
f"{prior_state} -> {state_norm} by "
f"{clean['requested_by'] or '(unknown)'} "
f"({clean['requested_by_profile'] or 'no profile'}); "
f"reason: {clean['reason'] or '(none)'}",
now_s,
),
)
row = conn.execute(
"SELECT * FROM maintenance_drain WHERE drain_id = ?", (drain_id,)
).fetchone()
record = self._maintenance_drain_row(row) or {}
return {
"record": record,
"drain_id": drain_id,
"state": state_norm,
"prior_state": prior_state,
"transitioned": transitioned,
}
+167
View File
@@ -0,0 +1,167 @@
# ADR: High-availability and rolling-restart architecture for Gitea MCP control plane
- **Status:** Proposed (Design ADR under [#668](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/668))
- **Date:** 2026-07-25
- **Tracking Issue:** [#668](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/668)
- **Policy Version:** `mcp-ha-rolling-restart/v1`
- **Related:**
- Parent: [#655](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/655) — Governed MCP restart coordination and zero-disruption recovery
- Governance Policy: [#656](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/656) / `docs/architecture/mcp-restart-governance.md`
- Control-Plane DB Substrate: [#613](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/613) / `docs/architecture/control-plane-db-substrate.md`
- Runtime Policy: [#615](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/615) / `docs/architecture/mcp-stable-control-runtime-policy-adr.md`
- Product Vision: [#652](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/652) (Phase 5 Maturity)
- Delivery Roadmap: [#653](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/653)
---
## 1. Context & Problem Statement
The Gitea MCP server operates as the authoritative **control plane** for managing issues, Pull Requests, code mutations, formal reviews, and workflow reconciliations. Under single-process governance ([#656](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/656)), process restarts are strictly controlled using pre-flight checks, drain phases, and operator approvals.
However, a single-instance control plane inherently presents fundamental constraints:
1. **Downtime during updates:** Even a perfectly executed single-process drain requires a window where incoming client requests must be paused or rejected while the server binary or python environment reloads.
2. **Single point of failure:** Infrastructure issues, process crashes, or unhandled host-level terminations immediately disconnect active LLM sessions and leave transient workflows incomplete.
3. **Multi-agent concurrency bottlenecks:** High volumes of concurrent multi-LLM tasks put all lock management, lease allocation, and Gitea API interactions through a single process event loop.
To achieve true zero-disruption operation and seamless rolling deployments without stopping active work, the system requires a high-availability (HA), multi-instance MCP architecture.
---
## 2. Architectural Principles & Non-Goals
### 2.1 Core Architectural Principles
* **Gitea as Canonical Work SoT:** Gitea remains the ultimate System of Record (SoT) for issue states, pull requests, labels, and audit comments. The MCP control plane does not duplicate domain entities.
* **Control-Plane DB as Multi-Instance State Substrate:** The control-plane SQLite/durable database ([#613](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/613)) acts as the single source of truth for workflow leases, session tokens, assignment records, and lock fences across all MCP nodes.
* **Stateless Worker Nodes:** MCP role server processes (`gitea-author`, `gitea-reviewer`, `gitea-merger`, `gitea-reconciler`, `gitea-controller`) maintain no unique in-memory state; any node can handle any request given a valid session resume token.
* **Fail-Closed Split-Brain Defense:** In any network partition or quorum loss scenario, nodes must fail closed rather than risk double-mutations or conflicting Gitea states.
### 2.2 Non-Goals
* **Replacing Gitea:** We do not replace Gitea issue/PR tracking with an independent database.
* **Immediate Multi-Node Cluster Execution in v1:** This ADR defines the target architecture and phased roadmap; immediate implementation occurs incrementally post-[#655] v1.
---
## 3. High-Availability & Rolling-Restart Architecture
### 3.1 Architecture Overview
```
+----------------------------+
| LLM Clients / IDE Sessions |
+--------------+-------------+
|
v
+----------------------------+
| HA Proxy / Router |
| (Health-based & Affinity) |
+------+--------------+------+
| |
+--------------+ +--------------+
v v
+--------------------+ +--------------------+
| MCP Instance Node A| | MCP Instance Node B|
| (Version N) | | (Version N+1) |
+---------+----------+ +---------+----------+
| |
+----------------------+----------------------+
|
v
+----------------------------+
| Control-Plane DB Substrate|
| (Shared Lease & Locks) |
+--------------+-------------+
|
v
+----------------------------+
| Gitea API |
+----------------------------+
```
---
### 3.2 Key System Components
#### A. Multiple MCP Instance Cohorts
* The control plane runs across $N \ge 2$ redundant process nodes.
* Dual-namespace deployment allows running the old version (Node A) alongside a updated version (Node B) during rolling upgrades.
#### B. Shared Durable Session Storage & Resume Tokens
* Session context, preflight verification proofs, and capability resolution states are stored in the shared control-plane database.
* Client requests carry an explicit `session_id` and `resume_token`. If an MCP instance restarts or a request routes to a different instance, the target node validates the token against the database without requiring full session re-initialization.
#### C. Shared Lease Authority & Fencing Counters
* Workflow leases (`gitea_allocate_next_work`, `gitea_adopt_workflow_lease`) use monotonic fencing tokens (`lease_generation_id`).
* When Node B acquires or renews a lease, it increments the generation counter. Any delayed or out-of-order write attempt from Node A using an older generation token is rejected by database constraints.
#### D. Leader Election & Coordinated Drain
* Node clusters elect a primary coordinator node for administrative background tasks (such as stale lease cleanup or incident Watchdogs).
* During a rolling deployment:
1. Node B (new version) is launched and registers as healthy.
2. Router directs new session creations to Node B.
3. Node A enters `MAINTENANCE_DRAIN` status ([#659](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/659)), completing in-flight mutations while refusing new tasks.
4. Once all active sessions migrate or complete, Node A shuts down cleanly.
#### E. Idempotent Mutations & Failover Safety
* All state-changing tool executions (PR creation, review submission, merge operations, label changes) carry a deterministic `idempotency_key`.
* If a network connection flaps or a node fails mid-mutation, the re-issued request with the same `idempotency_key` is recognized by the control-plane substrate, returning the existing recorded result without repeating side effects on Gitea.
#### F. Schema Version Compatibility
* Database migrations follow non-breaking additive patterns.
* During rolling upgrades where Node A (Version $N$) and Node B (Version $N+1$) run concurrently, both versions operate against the shared schema without structural conflicts.
---
## 4. Split-Brain & Failure Behavior
### 4.1 Split-Brain Risk Scenarios & Mitigation
| Scenario | Risk | Mitigation Strategy |
|---|---|---|
| **Network Partition between Nodes** | Both Node A and Node B attempt to process operations for the same issue/PR. | **Generation Fencing:** Lease renewal requires updating the DB generation counter. The node isolated from the DB fails closed immediately. |
| **Stale Node Recovery** | Node A recovers after a long pause and executes a queued mutation. | **Lease Expiry & TTL Fencing:** Transactions verify that `expires_at > NOW()` within the atomic SQLite transaction boundaries. |
| **Database Connection Loss** | Node loses access to shared control-plane DB substrate. | **Strict Fail-Closed:** The node immediately marks all task capabilities as `blocked` and rejects mutation tools until DB connectivity is re-established. |
---
## 5. Phased Implementation Milestones
```mermaid
flowchart TD
M1[Milestone 1: Shared Control-Plane DB Schema & Resume Tokens] --> M2[Milestone 2: Idempotent Mutation Layer]
M2 --> M3[Milestone 3: Health Routing & Standby Failover]
M3 --> M4[Milestone 4: Active-Active Rolling Deployment & Auto-Drain]
```
### Milestone 1: Shared Control-Plane DB Schema & Resume Tokens (Post-#655)
* Extend [#613] Control-Plane DB schema to store multi-instance node heartbeat records and session resume tokens.
* Enable session lookup across instances via `session_id`.
### Milestone 2: Idempotent Mutation Layer & Lease Fencing
* Add mandatory `idempotency_key` tracking to all Gitea mutation tools.
* Implement monotonic lease fencing counters in `gitea_allocate_next_work` and `gitea_adopt_workflow_lease`.
### Milestone 3: Health-Based Routing & Active-Passive Standby
* Introduce lightweight proxy/router capable of checking node health endpoints.
* Implement active-standby failover where standby node automatically assumes work if active node fails health checks.
### Milestone 4: Active-Active Horizontal Deployment & Rolling Upgrade Automation
* Enable true active-active multi-instance execution.
* Integrate automated zero-downtime rolling upgrades coordinated with `gitea_request_mcp_restart` maintenance drain.
---
## 6. Observability & Audit Requirements
High-availability control plane operations must expose clear telemetry and audit trails:
* **Node Registry Telemetry:** Active nodes, version numbers, uptime, and heartbeat timestamps reported via `gitea_get_runtime_context`.
* **Lease Fencing Metrics:** Tracking lease acquire latency, fence rejection counts, and lease handoff durations.
* **Failover & Re-route Audit Logs:** Durable logging of session migrations between nodes, drain initiation, and process retirement events.
---
## 7. Tradeoffs & Accepted Risks
* **Increased Architectural Complexity:** Moving from a single process to a multi-instance control plane requires robust DB locking, proxy routing, and migration governance.
* **Database Dependency:** The control-plane database substrate becomes a critical shared dependency for multi-node deployments. High availability for the underlying SQLite file system / DB must be guaranteed.
@@ -0,0 +1,83 @@
# Incident #670: bare direct-to-master commit `2fa97c26` (retroactive audit)
Status: verified; disposition recommendation: **accept as-is, no revert** (final
disposition owned by controller per issue #670).
## Summary
Commit `2fa97c26fbda555a1a83930ca5fdcea9d8e47b50`
(`fix(mcp): load dotenv relative to project root`) landed on `prgs/master`
as a single-parent commit with no PR wrapper and no review record, bypassing
the sanctioned issue → branch → PR → review → merge workflow. It was
discovered during the PR #654 post-merge audit. PR #654 itself merged
cleanly via the Gitea API and did **not** introduce this commit.
## Verification evidence (acceptance criteria 13)
- **AC1 — present on `prgs/master`: yes.**
`git merge-base --is-ancestor 2fa97c26fbda555a1a83930ca5fdcea9d8e47b50 prgs/master` → true.
- **AC2 — no PR or review record: confirmed.**
The commit is a single-parent, non-merge commit sitting directly on
first-parent master between the #629 merge (`5ab5fe85`) and the #654
merge (`ec903b0d`). A PR landing on master produces a merge commit (or a
PR-linked head); neither exists here. The controller audit at issue-create
time also found no PR wrapper and no review record for this SHA.
- **AC3 — changed files and diff summary: confirmed.**
`gitea_auth.py | 5 +++--` (+3/2). Single parent
`5ab5fe8583c07134d55dadf09381aecb67df246e`. The change moves
`PROJECT_ROOT` derivation above `load_dotenv()` and loads
`.env` relative to the project root instead of the process CWD.
## AC4 — why no immediate revert
- The dotenv fix is intentional and required for correct runtime behavior:
without it, `load_dotenv()` resolves `.env` against the process working
directory, which breaks MCP server launches whose CWD is not the project
root.
- The change is small (+3/2), self-contained in `gitea_auth.py`, and has
been running on master without incident since 2026-07-10.
- Reverting would re-introduce a real bug to remove a provenance defect —
the wrong trade. Provenance is repaired retroactively by this document,
issue #670, and the hardening landed under #671.
- If the controller later judges the change unsafe, a separate
revert/repair issue is the sanctioned path (issue #670, recommended
disposition option 4).
## AC5 — workflow-hardening linkage
Prevention already landed: **issue #671** (closed)
*“Block direct pushes to stable branches from MCP workflow sessions”*,
implemented by commit `5933d87647656643a67a50331c4c7b06ea751dad`
(`feat(guard): block direct stable-branch pushes from MCP workflow sessions`).
Shipped guardrails include:
- `gitea_record_stable_branch_push_attempt` — classifies proposed commands
for direct stable-branch push intent (`git push <remote> master`,
refspecs, `HEAD:master`, `--force`, dry-run intent, `:master` delete),
plus root/control-checkout local commits not carried by an issue branch,
and writes a durable `stable_branch_contamination` marker.
- `gitea_audit_stable_branch_contamination` — reconciler-only audit/clear
path; a contaminated worker session cannot self-clear.
- Review/merge/close/completion mutations fail closed while a
contamination marker is active.
## AC6 — PR #654 was not the source
- `2fa97c26` is the **first parent** of the #654 merge commit
`ec903b0d619e7a27d24aed272a890f4e5d381411`; it predates the #654 merge.
- First-parent history `5ab5fe8..ec903b0`:
`2fa97c2 fix(mcp): load dotenv relative to project root` followed by
`ec903b0 Merge pull request 'feat: lifecycle role/hazard labels ... (#603)' (#654)`.
- The #654 merger audit confirmed `ec903b0d` was a valid Gitea-API merge,
the `git push prgs master` attempt during that run was a no-op, and the
net change `2fa97c2..ec903b0` contained only the reviewed #603
lifecycle-label files.
- Conclusion: #654 merged reviewed content only; the unauthorized-path
defect is solely the earlier bare commit `2fa97c26`.
## Explicit non-actions (unchanged by this audit)
- No revert of `2fa97c26`.
- No force-push or history rewrite.
- No master mutation from the audit session.
+45
View File
@@ -0,0 +1,45 @@
# MCP maintenance-drain mode (#659)
Graceful **maintenance drain** stops new work assignment and defers non-allowlisted
mutations so sessions can finish critical handoffs and checkpoint before a
restart. It is **not** a restart authorization: the drain *proof* and apply gate
remain #661.
## State
Per repository scope (`remote`/`org`/`repo`) in the control-plane DB table
`maintenance_drain` (schema v6):
| State | Meaning |
|-------|---------|
| `inactive` | Normal operation (also: no row) |
| `draining` | Assignment stopped; non-allowlisted mutations deferred |
Enter/exit transitions are audited as `maintenance_drain_enter` /
`maintenance_drain_exit` events.
## Tools
| Tool | Permission | Effect |
|------|------------|--------|
| `gitea_maintenance_drain_status` | `gitea.read` | Observe drain (every session) |
| `gitea_enter_maintenance_drain` | `runtime.maintenance_drain` | Enter drain (capability-gated) |
| `gitea_exit_maintenance_drain` | `runtime.maintenance_drain` | Exit drain |
`runtime.maintenance_drain` is intentionally **not** a `gitea.*` op, so ordinary
author profiles cannot enter drain by accident.
## Enforcement
1. **Allocator** (`allocate_next_work`): while draining, returns `outcome=wait`
with `reason_code=maintenance_drain_assignment_stopped` for dry-run and apply.
2. **Mutation preflight** (`verify_preflight_purity`): non-allowlisted mutation
tasks raise `MaintenanceDrainError` with a typed next action.
3. **Allowlist** (safety only): heartbeats, lease release/abandon, session
checkpoints, enter/exit drain. Reads always work.
## Restart relationship
Drain mode prepares the blast radius. Restart apply still requires a clean
`DrainProof` (#661) or authorized break-glass. Status payloads never claim
restart permission.
+94
View File
@@ -0,0 +1,94 @@
# MCP scoped recovery playbook (#669)
**Parent:** [#655](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/655)
**Vision / roadmap:** [#652](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/652) · [#653](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/653)
**Class matrix:** [#663](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/663) · `docs/mcp-restart-classes.md`
**Coordinator:** [#658](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/658) · `restart_coordinator.py`
**Audit lineage:** [#665](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/665)
## Decision
Full-server MCP reset is a **last resort**. Prefer the narrowest recovery that
can clear the symptom. The coordinator **refuses** `rolling_mcp_restart`,
`full_mcp_restart`, and `host_restart` unless:
1. The inventory carries a prior **attempt log** of at least one *insufficient*
narrower recovery, **or**
2. **Break-glass** is authorized
(`request_break_glass` + `GITEA_BREAKGLASS_RESTART_AUTHORIZATION`).
Break-glass still never bypasses the #663 class matrix (role/permission).
## Ladder (narrow → broad)
| Rank | Action | Self-service | Implementation / delegation |
|---:|---|---|---|
| 0 | `client_reconnect` | yes | Host auto-reconnect / client reconnect · [#584](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/584) · `docs/mcp-namespace-eof-recovery.md` |
| 1 | `capability_refresh` | yes | `gitea_resolve_task_capability` + `gitea_whoami` · [#610](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/610) · [#685](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/685) |
| 2 | `session_reconnect` | yes | Runtime rebind + explicit `worktree_path` · [#543](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/543) · [#618](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/618) |
| 3 | `configuration_reload` | no | Class `configuration_reload` · console reload · [#642](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/642) |
| 4 | `lease_recovery` | no | Lock/lease recovery paths · [#702](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/702) · [#753](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/753) · [#790](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/790) |
| 5 | `worker_restart` | no | Class `worker_restart` · #663 |
| 6 | `role_runtime_restart` | no | Class `role_runtime_restart` · console restart · #642/#663 |
| 7 | `connector_restart` | no | Class `connector_restart` · #663 |
| 8 | `rolling_mcp_restart` | no | Class `rolling_mcp_restart` · design [#668](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/668) · **attempt log required** |
| 9 | `full_mcp_restart` | no | Class `full_mcp_restart` · **attempt log required** |
| 10 | `host_restart` | no | Class `host_restart` · **attempt log required** |
Machine-readable source of truth: `recovery_playbook.RECOVERY_LADDER` and
`recovery_playbook.ladder_document()`.
## Attempt log shape
Each prior attempt is a mapping:
```json
{
"action": "client_reconnect",
"outcome": "insufficient",
"reason": "transport still closed after IDE reconnect",
"actor": "prgs-controller-12345",
"recorded_at": "2026-07-25T21:00:00+00:00"
}
```
Outcomes that count toward escalation: `failed`, `insufficient`, `denied`,
`unresolved`, `timeout`, `error`.
Pass attempts into the coordinator via inventory
`prior_recovery_attempts` or the MCP tool argument
`prior_recovery_attempts_json` on `gitea_request_mcp_restart`.
Helper: `recovery_playbook.build_attempt_record(...)`.
## Symptom → first rung
`recovery_playbook.recommend_actions(symptoms=[...])` maps symptoms such as
`transport_eof`, `stale_capability`, `stale_lease`, `daemon_corrupt` to the
narrowest recommended action, then walks the ladder. Soft recommendations
never replace the hard gate on broad restarts.
## Enforcement points
1. **`recovery_playbook.assess_escalation`** — pure gate.
2. **`restart_coordinator.evaluate_restart_impact`** — when `restart_class` is
set (policy-enforced path), broad classes require the gate; report fields
`attempt_log_satisfied`, `playbook_escalation`, `break_glass`.
3. **`gitea_request_mcp_restart`** — accepts attempt JSON and env-authorized
break-glass; never restarts a process.
## Metrics
`recovery_playbook.recovery_metrics(attempts)` reports the fraction of
successful recoveries that avoided full/host restart
(`fraction_avoided_full_restart`).
## Non-goals
* HA multi-instance execution (#668 design only here).
* Normalizing `pkill` (#630 contamination stays forbidden).
* Silent mutation of leases or processes from the playbook itself.
## Manual process kills
Remain forbidden and contaminating (#630). The playbook never recommends them.
+15 -5
View File
@@ -95,12 +95,18 @@ gitea_request_mcp_restart(remote, host, org, repo,
target_session_id=None, target_role=None,
target_connector=None,
drain_proof_json=None,
request_break_glass=False)
request_break_glass=False,
prior_recovery_attempts_json=None)
```
It **never restarts anything**: `apply_supported` is always `false` and
`restart_performed` is always `false`.
`prior_recovery_attempts_json` (#669) is an optional JSON array of prior
narrow recovery attempts. Rolling / full / host classes require at least one
*insufficient* narrower attempt (or authorized break-glass). See
`docs/mcp-recovery-playbook.md`.
### Dry-run versus apply
| Call | Behavior |
@@ -112,9 +118,10 @@ It **never restarts anything**: `apply_supported` is always `false` and
An apply requires **both** authorizations, and they are independent:
1. **Restart-class authorization** (#663) — the requester's role and permissions
must allow the requested class, the class's approval requirement must be
satisfied, and any target-scoped class must name its target. Failing any of
1. **Restart-class authorization** (#663 / #669) — the requester's role and
permissions must allow the requested class, the class's approval requirement
must be satisfied, any target-scoped class must name its target, and broad
classes must satisfy the recovery-playbook attempt-log gate. Failing any of
these makes `allow_restart` `false`.
2. **Drain-proof gate** (#661) — a valid, unexpired, clean proof bound to the
current impact fingerprint, or an authorized break-glass.
@@ -127,7 +134,10 @@ the authorization that produced it.
### Break-glass
Break-glass bypasses the **drain proof only** — never the restart-class matrix.
Break-glass bypasses the **drain proof only** — never the restart-class matrix
(role/permission). Separately, authorized break-glass also satisfies the #669
attempt-log requirement for broad restarts (rolling/full/host), because that
gate is not a class-matrix permission check.
It is honoured solely when `request_break_glass` is set *and* the environment
carries `GITEA_BREAKGLASS_RESTART_AUTHORIZATION`; like operator override, the
tool argument expresses caller intent and cannot be self-asserted by a worker
+64
View File
@@ -0,0 +1,64 @@
# Sanctioned Recovery Playbooks & Controls (Phase 2 #644)
## Overview
Stale runtimes, worktree binding mismatches, and un-reconciled merged branches previously required expert manual shell recovery. Manual process kills (`pkill -f mcp_server.py`) are strictly forbidden and classified as runtime contamination ([#630](sanctioned-restart-controls.md)).
Phase 2 introduces **sanctioned recovery playbooks and controls** into the Web Console:
- **Diagnose**: Surface stale runtimes, worktree binding errors, contamination markers, and worktree anomalies via health & inventory APIs.
- **Preview**: Render mutation ledgers and exact confirmation phrases for recovery playbooks.
- **Confirm & Apply**: Execute sanctioned recovery actions through gated, audited paths.
- **Verify**: Revalidate control-plane state post-recovery before claiming clean status.
---
## Recovery Playbook Taxonomy
| Playbook ID | Action ID | Minimum Role | Target / Scope | Description |
|---|---|---|---|---|
| `clear_stale_binding` | `system.clear_stale_binding` | Operator | Active worktree binding | Clear provably missing or superseded `GITEA_ACTIVE_WORKTREE` binding ([#702](../stale_binding_recovery.py)). |
| `rebind_session_worktree` | `system.rebind_session_worktree` | Operator | Session worktree | Rebind or synchronize session worktree to verified lease worktree ([#864](../dirty_same_claimant_session_rebind.py)). |
| `reconcile_cleanups` | `system.reconcile_cleanups` | Controller | Worktree hygiene | Execute reconciler cleanup preview and apply for merged/superseded PR branches. |
| `sanctioned_restart` | `system.restart_namespace` | Admin | MCP Namespace | Restart MCP daemon gracefully via host supervisor ([#642](sanctioned-restart-controls.md)). |
---
## Wizard Workflow (Diagnose &rarr; Preview &rarr; Confirm &rarr; Verify)
### 1. Diagnose (`GET /api/v1/system/recovery/diagnose`)
Runs control-plane diagnostics:
- **Stale Runtime**: Mismatch between running daemon HEAD, local checkout HEAD, and remote-tracking HEAD.
- **Worktree Binding**: Missing path (`provably_stale_missing_path`), unverified inherited binding (`unverified_inherited`), or superseded binding (`superseded_by_session_lease`).
- **Contamination**: Checks for live contamination markers from unmanaged process kills.
- **Worktree Anomalies**: Scans `branches/` directory for un-reconciled cleanups or missing preserved worktrees.
Returns `RecoveryDiagnosis` with eligible playbooks.
### 2. Preview (`POST /api/v1/system/recovery/preview`)
Takes `playbook_id` and optional `target`/`params`.
Returns:
- **Mutation Ledger**: Step-by-step sequence of actions.
- **Confirmation Phrase**: Exact phrase required to authorize execution (e.g., `confirm clear_stale_binding`).
- **Authorization Decision**: RBAC check against the operator's principal.
### 3. Apply (`POST /api/v1/system/recovery/apply`)
Requires `playbook_id` and matching `confirmation` phrase. Gates run in this order, and each fails closed before anything is mutated:
1. **RBAC and execution phase** (`console_authz.authorize(..., for_execution=True)`). The phase branch only applies when `for_execution` is set. While `ACTIVE_PHASE` is `1`, every phase-2 recovery action is refused with `phase_not_active`, so no recovery playbook writes yet. Preview reports the same decision under `execution_authorization` / `execution_blocked_reason`.
2. **Confirmation phrase** (`confirmation_matches`).
3. **Contamination rules** ([#630](sanctioned-restart-controls.md)): the live marker is read from the session inventory and assessed under the gated task key `console_recovery_apply`. A contaminated runtime must be cleared through the reconciler cleanup playbook, which is the one playbook exempted from this gate because it is the designated remedy. The marker is also forwarded to `sanctioned_restart.execute_restart`, so a restart cannot launder a contaminated runtime.
Apply then executes the sanctioned recovery logic against the **live** process environment — not a copy — and records an audit entry in `console_audit`. A playbook that leaves the binding unchanged reports `performed: false`; `binding_before`, `binding_after`, and `binding_changed` are returned so a no-op cannot read as success.
Apply does **not** enforce master parity. Parity is reported by Diagnose ([#610](../master_parity_gate.py)) as evidence for the operator; it is not a precondition of this endpoint.
### 4. Verify (`POST /api/v1/system/recovery/verify`)
Re-evaluates control-plane diagnostics post-recovery and **reports** `clean`, `stale_runtime_clean`, `binding_clean`, `binding_classification`, and `contamination_clean`. It reports; it does not assert or block. State is read fresh rather than from the mapping a mutation just wrote. An `unverified_inherited` binding is reported as not clean, because unproven is not clean.
---
## Safety & Governance Principles
1. **No Manual `pkill`**: Direct process killing remains forbidden and is recorded as contamination.
2. **Auditability**: Every recovery preview and execution is logged in the console audit trail.
3. **Master Parity & Dual Control**: High-privilege recovery actions require controller/admin roles and explicit confirmation phrases.
+44 -9
View File
@@ -94,6 +94,10 @@ already define, and a regression test asserts each mapping matches.
| `record_analytics_usage` | operator | gated_write | `runtime.record_analytics_usage` | Yes | No | No | 2 |
| `system.reload_namespace` | controller | privileged | `runtime.reload_namespace` | Yes | No | No | 2 |
| `system.restart_namespace` | admin | destructive | `runtime.restart_namespace` | Yes | **Yes** | **Yes** | 2 |
| `system.clear_stale_binding` | operator | gated_write | `gitea.read` | Yes | No | No | 2 |
| `system.rebind_session_worktree` | operator | gated_write | `gitea.read` | Yes | No | No | 2 |
| `system.reconcile_cleanups` | controller | privileged | `gitea.pr.close` | Yes | No | No | 2 |
| `initiate_workflow` | operator | gated_write | `gitea.read` | Yes | No | No | 2 |
**Dual control** means the acting principal may not be the sole authority: a
second distinct principal must confirm. **Break-glass** means the action is
@@ -112,6 +116,12 @@ by the console — both hand off to a host supervisor, and neither exposes a raw
process kill. See
[`sanctioned-restart-controls.md`](sanctioned-restart-controls.md) (#642).
`initiate_workflow` (#643) is operator-class because its outcome is a *claim*,
not a Gitea verdict. Requesting reviewer or merger work reserves that work
through the allocator; it does not grant the right to approve or merge, which
stays with the MCP role profile and its own capability gates. See
[`webui-requests.md`](webui-requests.md).
### Authorization decision
`authorize(action_id, principal, for_execution=False)` returns a decision
@@ -126,9 +136,24 @@ record and **denies by default**. The deny reasons are closed and enumerated:
| `phase_not_active` | Execution requested for an action whose phase is not open. |
| `allowed_preview_only` | Authorized — preview only, execution still disabled. |
There is no implicit allow branch. Even the allow result reports
`execution_enabled: false` while the console is in Phase 1, so no caller can
read an allow as permission to mutate.
There is no implicit allow branch.
`execution_enabled` on the decision reports whether the action has a live
execution path at all, and is computed by `execution_wired(action)`. There are
exactly two ways to be wired:
1. the action's `phase` is at or below `ACTIVE_PHASE`; or
2. the action declares an `execution_env_flag` **and** that variable is set.
Every action that declares no flag therefore reports `execution_enabled: false`
while the console is in Phase 1, so no caller can read an allow as permission
to mutate. The per-action flag exists because raising `ACTIVE_PHASE` would
enable execution for every action of that phase at once, including ones whose
execution path is not implemented. One implemented action goes live on its own
flag instead of dragging its unimplemented phase-mates with it.
`initiate_workflow` is the only action that currently declares a flag
(`WEBUI_REQUESTS_EXECUTION`), and it stays denied until an operator sets it.
## Secret redaction
@@ -235,13 +260,22 @@ second one. The integration points are already wired and observable:
instead of adding a parallel check.
- **`GET /api/console/security-model`** publishes the RBAC matrix, redaction
policy, and audit policy as JSON for operators and tests.
- **`POST /api/v1/requests/preview` and `.../apply`** (#643) are the first
actions to use this model for a real execution path. Preview always returns a
decision and an audited `previewed` record; apply requires `confirm=true`,
emits `succeeded` or `denied`, and reserves work only through the allocator.
See [`webui-requests.md`](webui-requests.md).
To open Phase 2, a child issue must: raise `ACTIVE_PHASE`, implement the
confirmation and dual-control flow the matrix already declares, emit a
`succeeded` or `failed` record alongside the `gitea_audit` mutation record, and
keep `viewer` unable to reach any of it. Turning on execution without the
confirmation flow contradicts a declared requirement and is a review failure,
not a shortcut.
A Phase 2 action must: use `execution_wired` rather than a private enable flag,
implement the confirmation and dual-control flow the matrix already declares,
emit a `succeeded` or `failed` record alongside the `gitea_audit` mutation
record, and keep `viewer` unable to reach any of it. Turning on execution
without the confirmation flow contradicts a declared requirement and is a
review failure, not a shortcut.
Raising `ACTIVE_PHASE` remains the way to open a whole phase at once, and is
deliberately *not* what #643 did: an action-scoped opt-in cannot enable an
action whose execution path nobody wrote.
## Local-dev mode
@@ -294,6 +328,7 @@ Until Phase 2 wires it, probe protection rests on network placement alone, as
| `WEBUI_ROLE_MAP` | unset | JSON subject → role map |
| `WEBUI_REQUIRE_PROBE_AUTH` | unset | Require auth for non-public probes |
| `WEBUI_CONSOLE_AUDIT_LOG` | unset | Append-only audit sink path |
| `WEBUI_REQUESTS_EXECUTION` | unset | Opt in to `initiate_workflow` execution (#643) |
All are read server-side only. None is ever rendered into a page or returned by
an API.
+64
View File
@@ -80,6 +80,8 @@ status, onboarding checklist state, and the fail-closed error payloads (#635).
| `/sessions` | Runtime and session view (#641) — health + inventory sessions/namespaces/worktrees |
| `/api/sessions` | JSON export for the runtime/session view |
| `/api/v1/sessions` | Versioned alias of `/api/sessions` |
| `/gitea` | Gitea issue↔PR linkage console (#645) — both directions, with the evidence for each edge |
| `/api/v1/gitea/linkage` | JSON linkage export; `502` when the read could not be answered |
| `/inventory` | Phase 1 shell stub — unified inventory (backed by #636) |
| `/timeline` | Phase 1 shell stub — workflow event timeline |
| `/policy` | Phase 1 shell stub — capability/role policy placeholder |
@@ -327,6 +329,68 @@ Honesty rules specific to this view:
The write-time redactor is a narrow denylist and is not relied on. The field
itself is kept — it is the `#630` evidence naming which daemon was killed.
## Gitea issue/PR linkage (#645)
`/gitea` is the Phase 3 read-only linkage console: which PR carries which issue,
which issues are claimed by more than one PR, and what the latest Canonical
Thread Handoff on a thread said. Gitea remains the source of truth — this
surface reads it and never writes to it. There is no issue/PR editor, no review,
and no merge control.
Query parameters (all optional):
| Parameter | Meaning |
|-----------|---------|
| `project` | Registry project id to scope the read (default: first registry entry) |
| `state` | `open` (default) or `all`; `all` widens the window to merged/closed items, where a landed edge lives |
| `issue=N` / `pr=N` | Focus one thread and load *its* latest canonical handoff |
`GET /api/v1/gitea/linkage` returns the same model as JSON
(`schema_version: 1`). It answers `502` when the read could not be answered, so
an automated consumer cannot mistake a fail-closed payload for "no links exist".
The HTML page always answers `200` and renders the reason instead — an operator
view must show why a read failed rather than withhold the page.
### How an edge is found
Each edge carries the evidence that produced it, strongest first:
| Evidence | Meaning |
|----------|---------|
| `closes_keyword` | The PR title or body declares `closes/fixes/resolves #N`. Gitea itself acts on this keyword. |
| `branch_marker` | The PR head branch carries the canonical `(fix\|feat\|docs\|chore)/issue-N-…` marker minted by the issue lock. |
| `body_reference` | The PR body mentions `#N` with no closing keyword. A mention is not a claim to close. |
Only closing and branch-marker edges populate the **issue → PR** direction: a
bare mention is a cross-link, and counting it as ownership would invent
contested issues out of ordinary references. The mention stays visible on the
**PR → issue** side, labelled as such. A PR whose two strongest edges tie is
flagged `ambiguous`; an issue claimed by two PRs is flagged `contested`.
### Honesty rules specific to this view
* **A partial read never reads as an absence.** Linkage is a claim about the
loaded window only. When pagination did not complete, every empty edge cell
renders `none found (partial inventory)` rather than `none`, and the JSON
carries `inventory_complete: false` plus per-row `links_authoritative: false`.
* **A failed read renders no table at all.** Missing credentials, an unknown
project, or a fetch error produce `ok: false` with a reason. An empty linkage
table would assert that no issue is linked to any PR, which such a read is not
in a position to claim.
* **Handoffs are loaded, never assumed.** CTH comments are thread-scoped, so
only the focused issue or PR has its comments fetched. Every other row reports
`not_loaded` with the reason; a thread whose comments *were* loaded and carried
no CTH says exactly that. A comment-source failure degrades the handoff alone —
the linkage tables still render.
* **Unrecognised handoff headings are reported, not republished.** A `## CTH:`
heading outside `CTH_TYPES` renders as `unrecognized`.
* **Redaction precedes display.** Titles, labels, handoff fields, and error
reasons pass through `webui.console_redaction` before serialization, and the
page HTML-escapes everything it renders.
* **Deep links are opt-in.** A link out to the Gitea web UI appears only when
`GITEA_MCP_REVEAL_ENDPOINTS=1` is set server-side, matching how the MCP tools
gate URL exposure. Item numbers stay usable without it.
## System-health dashboard (#639)
`/system-health` renders the same snapshot the `/api/v1/system/health` API
+81
View File
@@ -0,0 +1,81 @@
# Web Console: Notifications & Human-Attention Routing (#648)
- **Status:** Phase 3 Live
- **Tracking Issue:** [#648](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/648)
- **Parent Epic:** [#631](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/631)
- **Attention Boundary Reference:** [#628](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/628)
---
## 1. Overview
The **Notifications & Human-Attention Console** (`/notifications`, `/api/v1/notifications`) provides intelligent event classification and human-attention routing for autonomous workflow operations.
To prevent alert fatigue while ensuring critical escalation boundaries are never missed, events are classified into three distinct **Attention Classes**:
1. **`human-required`** (Urgent Escalation Boundary):
- Items requiring immediate human intervention or business decisions.
- Triggers: Auth failures, hard stops, irrecoverable state, decision locks, failed report validations, critical probe errors.
- Display: Highlighted in red (`badge-blocked`) with a `HUMAN REQUIRED` badge.
2. **`operator`** (Operational Inbox):
- Items requiring controller or operator review/triage during routine execution.
- Triggers: Blocked PRs (merge conflicts), stale leases, duplicate PRs on issues, unassigned ready work.
- Display: Displayed in orange/yellow (`badge-claimed`).
3. **`routine`** (Background Workflow Transitions):
- Normal, healthy workflow transitions and state progressions.
- Triggers: Active PRs/issues in standard state, clean branch creation, routine heartbeats.
- Display: Filtered out of default inbox views to eliminate notification spam; viewable on demand via the "Routine" or "All" tab.
---
## 2. API Endpoints
### `GET /api/v1/notifications`
*Compatibility Alias:* `GET /api/notifications`
#### Query Parameters:
- `project_id` (optional): Filter notifications by project ID.
- `attention_class` (optional): `inbox` (default: human-required + operator), `human-required`, `operator`, `routine`, `all`.
#### Example JSON Response:
```json
{
"project_id": "gitea-tools",
"repo_label": "Scaled-Tech-Consulting/Gitea-Tools",
"human_required_count": 0,
"operator_count": 2,
"routine_count": 5,
"total_count": 7,
"fetch_error": null,
"inbox_items": [
{
"id": "notif-pr-block-742",
"attention_class": "operator",
"category": "blocker",
"title": "Blocked PR #742",
"summary": "PR #742 requires merge conflict resolution.",
"work_kind": "pr",
"work_number": 742,
"project_id": "gitea-tools",
"repo_label": "Scaled-Tech-Consulting/Gitea-Tools",
"created_at": "2026-07-25T16:39:47Z",
"deep_link": "/traffic",
"requires_human": false,
"extra": {}
}
],
"all_items": [...]
}
```
---
## 3. UI Navigation
- Access via the **Traffic** navigation menu: **Traffic → Notifications**.
- The main view displays:
- **Metrics Summary Bar**: Highlighting counts for Human Required, Operator Inbox, and Routine items.
- **Attention Filter Tabs**: Toggle between Inbox (Human + Operator), Human Required, Operator, Routine, and All.
- **Structured Event Table**: Displays category, title, summary, work item links, and timestamps.
+160
View File
@@ -0,0 +1,160 @@
# Web console requests: intent preview and workflow initiation (#643)
**Phase 2. Preview is always live and always read-only. Initiation is wired but
denied until an operator opts in.**
Before this surface, starting role work meant pasting a prompt into a terminal
and trusting the operator to have checked the allocator first. Nothing enforced
that check, so two sessions could reach for the same issue and each believe it
was theirs. This page replaces the paste with a *request*: a desired role, an
issue or PR, and a stated intent, answered by an authorization decision and —
on confirmation — an exclusive assignment from the allocator.
| Concern | Module |
|---------|--------|
| Request model, preview, initiation | `webui/request_service.py` |
| Form and preview rendering | `webui/request_views.py` |
| Authorization | `webui/console_authz.py` (`initiate_workflow`) |
| Audit | `webui/console_audit.py` |
| Ownership substrate | `allocator_service.py` + `control_plane_db.py` |
## Surfaces
| Path | Method | Purpose |
|------|--------|---------|
| `/requests` | GET | Request form |
| `/requests` | POST | Render an intent preview. **Never assigns.** |
| `/api/v1/requests/preview` | POST | Intent preview as JSON |
| `/api/v1/requests/apply` | POST | Initiate — confirmed, audited, allocator-owned |
The HTML form has no initiate button on purpose. Initiating requires a
confirmed POST to `/api/v1/requests/apply`, so a stray form submission cannot
reserve work as a side effect.
## The request
```json
{
"desired_role": "author",
"work_kind": "issue",
"work_number": 643,
"intent_summary": "implement request preview and initiation",
"remote": "prgs",
"org": "Scaled-Tech-Consulting",
"repo": "Gitea-Tools",
"expected_head_sha": null
}
```
`desired_role` is one of `author`, `reviewer`, `merger`, `reconciler`,
`controller`. `work_kind` is `issue` or `pr`. `remote`/`org`/`repo` default to
the first project in the registry when omitted; when neither the request nor
the registry resolves them, the request is rejected rather than pointed at some
other repository. `intent_summary` is required — it is what the audit record
states as the reason — and is truncated to 500 characters.
Parsing rejects rather than corrects. An unknown role, an unknown work kind, a
non-positive number, or a missing intent each return `400` with a `reason_code`
and the offending `field`.
## Preview
Five checks, each with its own verdict, reason code, and detail:
| Check | Passes when |
|-------|-------------|
| `authorization` | The console principal holds `operator` or above |
| `capability` | The desired role maps to a declared profile and MCP namespace |
| `lease_availability` | No active claim holds the work unit |
| `next_safe_action` | The allocator would independently select this exact work unit |
| `head_pin` | PR work resolves to a head SHA, and a supplied SHA still matches |
A preview also returns the role's `allowed_actions` and `prohibited_actions`
(from `allocator_service.ROLE_ACTIONS`), the `required_profile` and
`required_namespace` the work must run under, and a `correlation_id` that ties
the preview to its audit record and to any assignment that follows.
Preview is read-only in the strict sense: it calls the allocator with
`apply=false` and writes nothing but an audit line. An unauthorized principal
never reaches the allocator or the control-plane DB at all, so a denial cannot
be used to enumerate the queue.
## Initiation
`POST /api/v1/requests/apply` refuses in this order, and every refusal returns
before any assignment is attempted:
| Condition | Outcome | Status |
|-----------|---------|--------|
| Unparseable request | `invalid_request` | 400 |
| Not authorized, or execution not wired | `denied` | 403 |
| `confirm` not set | `denied` / `confirmation_required` | 409 |
| Work unit already claimed | `blocked` / `duplicate_assignment` | 409 |
| Allocator would select other work | `wait` / `not_next_safe_work` | 409 |
| Allocator declines on apply | `blocked` or `wait` | 409 |
| Evidence unavailable | `wait` / `evidence_unavailable` | 503 |
| Assigned | `assigned_work` | 201 |
A success returns the assignment plus a `handoff` block naming the profile, the
namespace, and the actions that stay forbidden — enough for the operator to
continue in the right MCP namespace without guessing.
### Why apply runs the allocator twice
The allocator is the only source of exclusive ownership (#600 / #613), and it
selects work; it does not take orders. So `apply` runs a dry-run first and
proceeds only when the allocator would independently pick the requested work
unit. If it would not, the request reports `wait` and mutates nothing.
A request is therefore a *confirmation* of the allocator's decision, never an
override of it. The apply call carries the dry-run's
`candidate_set_fingerprint` as a CAS pin (#776), so a queue that changed
between the two calls fails closed rather than assigning against a stale view.
The result is checked again on the way out: an assignment naming a different
work unit is not read as success.
### Fail-closed defaults
- An unreadable control-plane DB denies. It is never treated as "nothing holds
this work unit".
- An incomplete queue inventory denies (#758). Ranking a partial candidate set
can select the wrong work.
- An allocator that raises denies.
- PR work with no resolvable head SHA denies; a supplied SHA that no longer
matches denies with `head_moved`.
## Enabling initiation
Execution is wired off. Set `WEBUI_REQUESTS_EXECUTION=1` to enable it for the
`initiate_workflow` action only — see
[`webui-authz-audit.md`](webui-authz-audit.md) for why this is an
action-scoped flag rather than a phase bump. With the variable unset, `apply`
returns `403` with `reason_code: unauthorized` no matter who asks.
Enabling execution does **not** enable approvals or merges. Those are phase 3
console actions and remain forbidden in every path here; the console reserves
work and hands off, and the MCP role profile enforces what that role may then
do.
## Audit
Every preview and every apply emits a console audit record (schema in
[`webui-authz-audit.md`](webui-authz-audit.md)):
| Event | `result` |
|-------|----------|
| Preview | `previewed` |
| Refusal at any stage | `denied` |
| Assignment created | `succeeded` |
`correlation.request_id` carries the request's `correlation_id`, and a
successful record's `metadata` carries `assignment_id` and `lease_id`, so an
assignment can be traced back to the intent that produced it. The operator's
`intent_summary` travels in `metadata` and passes through the standard
redaction pass before persistence like every other field.
## Non-goals
- No browser-initiated approve or merge, in this phase or any other.
- No bypass of allocator exclusive ownership; no self-selection of work.
- No auto-start from raw monitoring incidents (#612 stays downstream).
+275 -9
View File
@@ -1472,6 +1472,9 @@ def verify_preflight_purity(
# contaminated by manual MCP daemon process killing (reconciler-exempt).
_enforce_runtime_recovery_contamination_gate(task, remote)
# #659 AC3: defer non-allowlisted mutations while maintenance drain is active.
_enforce_maintenance_drain_gate(task, remote=remote, org=org, repo=repo)
ctx = _resolve_namespace_mutation_context(worktree_path)
workspace = ctx["workspace_path"]
canonical_root = ctx["canonical_repo_root"]
@@ -2004,6 +2007,55 @@ def _enforce_runtime_recovery_contamination_gate(
)
def _enforce_maintenance_drain_gate(
task: str | None,
remote: str | None = None,
org: str | None = None,
repo: str | None = None,
) -> None:
"""#659 AC3: defer non-allowlisted mutations while drain is active.
The single mutation chokepoint already used by every gated task, so drain
coverage cannot drift per-tool. Allowlisted safety operations (heartbeat,
release/abandon, checkpoint, drain exit) pass through so an in-flight
session can still finish and hand off; everything else is deferred with a
typed blocker. Unreadable drain state fails closed a drain that cannot be
read is not evidence that no drain is running.
"""
if _preflight_in_test_mode() and not os.environ.get(
"GITEA_TEST_FORCE_MAINTENANCE_DRAIN"
):
return
if maintenance_drain.is_allowlisted_task(task):
return
try:
_h, o, r = _resolve(remote, None, org, repo)
except Exception: # noqa: BLE001 — scope resolution is best-effort here
o, r = (org or ""), (repo or "")
db, errs = _control_plane_db_or_error()
if db is None:
raise RuntimeError(
"maintenance-drain state could not be read: "
f"{'; '.join(errs) or 'control-plane DB unavailable'} (fail closed, #659)"
)
try:
record = db.read_maintenance_drain(remote=remote or "", org=o, repo=r)
except Exception as exc: # noqa: BLE001
raise RuntimeError(
f"maintenance-drain state could not be read: {_redact(str(exc))} "
"(fail closed, #659)"
) from exc
decision = maintenance_drain.classify_mutation(task, record)
if not decision["allowed"]:
raise maintenance_drain.MaintenanceDrainError(
maintenance_drain.format_drain_block_error(decision),
decision=decision,
)
def _enforce_stable_branch_contamination_gate(
task: str | None,
remote: str | None = None,
@@ -2064,6 +2116,7 @@ import allocator_service # noqa: E402
import allocator_dependencies # noqa: E402
import dependency_graph # noqa: E402 # #784 durable dependency edges
import control_plane_db # noqa: E402
import maintenance_drain # noqa: E402 # #659 graceful maintenance-drain mode
import lease_lifecycle # noqa: E402
import lease_policy # noqa: E402
import workflow_dashboard # noqa: E402 # #605 live queue/lease dashboard
@@ -22559,6 +22612,193 @@ def gitea_workflow_dashboard(
return payload
@mcp.tool()
def gitea_maintenance_drain_status(
remote: str = "dadeschools",
host: str | None = None,
org: str | None = None,
repo: str | None = None,
) -> dict:
"""Read-only: current maintenance-drain state for a repository scope (#659 AC4).
Every session must be able to observe drain so it can stop creating new work
and finish only allowlisted safety operations. Never mutates; never restarts.
"""
read_block = _profile_operation_gate("gitea.read")
if read_block:
return {
"success": False,
"read_only": True,
"reasons": read_block,
"permission_report": _permission_block_report("gitea.read"),
}
try:
_h, o, r = _resolve(remote, host, org, repo)
except ValueError as exc:
return {"success": False, "read_only": True, "reasons": [str(exc)]}
db, errs = _control_plane_db_or_error()
if db is None:
return {
"success": False,
"read_only": True,
"reasons": errs or ["control-plane DB unavailable"],
"maintenance_drain": maintenance_drain.status_payload(
None, remote=remote, org=o, repo=r
),
}
try:
record = db.read_maintenance_drain(remote=remote, org=o, repo=r)
except Exception as exc: # noqa: BLE001
return {
"success": False,
"read_only": True,
"reasons": [f"drain state unreadable: {_redact(str(exc))}"],
}
payload = maintenance_drain.status_payload(
record, remote=remote, org=o, repo=r
)
return {"success": True, "read_only": True, "maintenance_drain": payload}
@mcp.tool()
def gitea_enter_maintenance_drain(
reason: str = "",
remote: str = "dadeschools",
host: str | None = None,
org: str | None = None,
repo: str | None = None,
session_id: str | None = None,
) -> dict:
"""Enter graceful maintenance-drain mode for a repository scope (#659 AC1).
Stops new assignment and defers non-allowlisted mutations until exit. Requires
``runtime.maintenance_drain`` (controller/lifecycle capability not granted
by ordinary Gitea author profiles). Audited in the control-plane event log.
"""
cap_block = _profile_operation_gate("runtime.maintenance_drain")
if cap_block:
return {
"success": False,
"performed": False,
"reasons": cap_block,
"permission_report": _permission_block_report(
"runtime.maintenance_drain"
),
}
try:
_h, o, r = _resolve(remote, host, org, repo)
except ValueError as exc:
return {"success": False, "performed": False, "reasons": [str(exc)]}
db, errs = _control_plane_db_or_error()
if db is None:
return {
"success": False,
"performed": False,
"reasons": errs or ["control-plane DB unavailable"],
}
profile = get_profile() or {}
try:
result = db.set_maintenance_drain(
remote=remote,
org=o,
repo=r,
state=maintenance_drain.STATE_DRAINING,
reason=reason or "operator-entered maintenance drain",
requested_by=str(
(profile.get("identity") or {}).get("username")
or profile.get("expected_username")
or ""
),
requested_by_profile=str(profile.get("profile_name") or ""),
session_id=str(session_id or ""),
)
except Exception as exc: # noqa: BLE001
return {
"success": False,
"performed": False,
"reasons": [f"enter drain failed: {_redact(str(exc))}"],
}
record = result.get("record") or {}
return {
"success": True,
"performed": True,
"transitioned": bool(result.get("transitioned")),
"state": result.get("state"),
"prior_state": result.get("prior_state"),
"drain_id": result.get("drain_id"),
"maintenance_drain": maintenance_drain.status_payload(
record, remote=remote, org=o, repo=r
),
}
@mcp.tool()
def gitea_exit_maintenance_drain(
reason: str = "",
remote: str = "dadeschools",
host: str | None = None,
org: str | None = None,
repo: str | None = None,
session_id: str | None = None,
) -> dict:
"""Exit graceful maintenance-drain mode (#659 AC1). Restores assignment and mutations."""
cap_block = _profile_operation_gate("runtime.maintenance_drain")
if cap_block:
return {
"success": False,
"performed": False,
"reasons": cap_block,
"permission_report": _permission_block_report(
"runtime.maintenance_drain"
),
}
try:
_h, o, r = _resolve(remote, host, org, repo)
except ValueError as exc:
return {"success": False, "performed": False, "reasons": [str(exc)]}
db, errs = _control_plane_db_or_error()
if db is None:
return {
"success": False,
"performed": False,
"reasons": errs or ["control-plane DB unavailable"],
}
profile = get_profile() or {}
try:
result = db.set_maintenance_drain(
remote=remote,
org=o,
repo=r,
state=maintenance_drain.STATE_INACTIVE,
reason=reason or "operator-exited maintenance drain",
requested_by=str(
(profile.get("identity") or {}).get("username")
or profile.get("expected_username")
or ""
),
requested_by_profile=str(profile.get("profile_name") or ""),
session_id=str(session_id or ""),
)
except Exception as exc: # noqa: BLE001
return {
"success": False,
"performed": False,
"reasons": [f"exit drain failed: {_redact(str(exc))}"],
}
record = result.get("record") or {}
return {
"success": True,
"performed": True,
"transitioned": bool(result.get("transitioned")),
"state": result.get("state"),
"prior_state": result.get("prior_state"),
"drain_id": result.get("drain_id"),
"maintenance_drain": maintenance_drain.status_payload(
record, remote=remote, org=o, repo=r
),
}
@mcp.tool()
def gitea_request_mcp_restart(
remote: str = "dadeschools",
@@ -22575,8 +22815,9 @@ def gitea_request_mcp_restart(
target_connector: str | None = None,
drain_proof_json: str | None = None,
request_break_glass: bool = False,
prior_recovery_attempts_json: str | None = None,
) -> dict:
"""Evaluate a proposed MCP restart and return an impact preview (#658).
"""Evaluate a proposed MCP restart and return an impact preview (#658/#669).
Central restart coordinator: resolves the requested restart class, gathers
live control-plane state (sessions,
@@ -22600,10 +22841,16 @@ def gitea_request_mcp_restart(
independent the drain gate proves the blast radius was drained and knows
nothing about whether this requester may request this class so a class the
matrix denied never reports an authorized apply. Break-glass bypasses the
drain proof only; it never bypasses the class matrix. ``apply_gate`` carries
drain proof and, when env-authorized, the #669 attempt-log requirement for
broad restarts; it never bypasses the class matrix. ``apply_gate`` carries
``drain_gate_allow`` and ``restart_class_authorized`` so a denial is
attributable to the authorization that produced it.
``prior_recovery_attempts_json`` (#669) is an optional JSON array of prior
narrow recovery attempts ``{action, outcome, reason, ...}``. Rolling / full
/ host restart classes require at least one *insufficient* narrower attempt
unless break-glass is authorized.
Operator override authority is read from the process environment
(``GITEA_OPERATOR_RESTART_OVERRIDE_AUTHORIZATION``), never self-asserted by
the requesting session: ``request_override`` only expresses caller intent
@@ -22709,12 +22956,37 @@ def gitea_request_mcp_restart(
requester_role
)
prior_recovery_attempts: list[dict] = []
if prior_recovery_attempts_json:
try:
parsed_attempts = json.loads(prior_recovery_attempts_json)
if isinstance(parsed_attempts, list):
prior_recovery_attempts = [
dict(a) for a in parsed_attempts if isinstance(a, dict)
]
else:
incomplete_reasons.append(
"prior_recovery_attempts_json must be a JSON array (#669)"
)
inventory_complete = False
except (ValueError, TypeError) as exc:
incomplete_reasons.append(
f"invalid prior_recovery_attempts_json: {_redact(str(exc))}"
)
inventory_complete = False
break_glass_authorized = bool(
(os.environ.get("GITEA_BREAKGLASS_RESTART_AUTHORIZATION") or "").strip()
)
break_glass = bool(request_break_glass and break_glass_authorized)
inventory = {
"sessions": sessions,
"leases": leases,
"terminal_lock": terminal_lock,
"inventory_complete": inventory_complete,
"incomplete_reasons": incomplete_reasons,
"prior_recovery_attempts": prior_recovery_attempts,
}
report = restart_coordinator.evaluate_restart_impact(
@@ -22730,6 +23002,7 @@ def gitea_request_mcp_restart(
target_session_id=target_session_id,
target_role=target_role,
target_connector=target_connector,
break_glass=break_glass,
)
payload = report.as_dict()
@@ -22762,13 +23035,6 @@ def gitea_request_mcp_restart(
except (ValueError, TypeError) as exc:
proof_parse_error = f"invalid drain_proof_json: {_redact(str(exc))}"
break_glass_authorized = bool(
(
os.environ.get("GITEA_BREAKGLASS_RESTART_AUTHORIZATION") or ""
).strip()
)
break_glass = bool(request_break_glass and break_glass_authorized)
expected_fp = drain_proof.impact_fingerprint(report.as_dict())
gate = drain_proof.gate_apply_restart(
proof=proof_obj,
+281
View File
@@ -0,0 +1,281 @@
"""Graceful MCP maintenance-drain mode (#659).
Drain is the visible, capability-gated state that lets an operator stop new
work and quiesce mutations *before* a restart, instead of cutting sessions off
mid-mutation. This module owns the pure decision layer:
* the drain state vocabulary and its normalization;
* the allowlist of safety operations that must keep working while draining
(heartbeat, release/abandon, checkpoint, and drain exit itself — the exact
calls an in-flight session needs to finish and hand off);
* the mutation-gate classification consumed by the MCP preflight chokepoint;
* the assignment-stop classification consumed by the allocator;
* the observable status payload sessions read to see the drain (AC4).
Durable state lives in the control-plane DB (``maintenance_drain`` table);
enforcement lives at the existing chokepoints. Nothing here performs I/O, so
both callers can share one decision without importing each other.
Scope note: the machine-verifiable *drain proof* and the restart gate that
consumes it are #661's scope, not this module's. Drain here stops assignment
and mutation and makes the state observable; it never authorizes a restart.
"""
from __future__ import annotations
from typing import Any, Mapping
# ── State vocabulary ──────────────────────────────────────────────────────────
STATE_INACTIVE = "inactive"
STATE_DRAINING = "draining"
DRAIN_STATES = frozenset({STATE_INACTIVE, STATE_DRAINING})
# Typed blocker code surfaced to clients (never a bare string at call sites).
BLOCKER_DRAIN_ACTIVE = "maintenance_drain_active"
# Reason code for the allocator's assignment stop.
REASON_ASSIGNMENT_STOPPED = "maintenance_drain_assignment_stopped"
DRAIN_SCHEMA_VERSION = 6
class MaintenanceDrainError(RuntimeError):
"""Raised when a mutation is refused because drain is active (fail closed)."""
def __init__(self, message: str, *, decision: Mapping[str, Any] | None = None):
super().__init__(message)
self.decision = dict(decision or {})
self.reason_code = BLOCKER_DRAIN_ACTIVE
# ── Safety allowlist ──────────────────────────────────────────────────────────
# Mutations that stay permitted while draining. Every entry is a *quiesce*
# operation: it either proves an in-flight task is still alive, hands its claim
# back, records the durable state a restart needs, or ends the drain. Nothing
# that creates new work, new branches, new PRs, or new review/merge verdicts is
# on this list — that is the whole point of the drain.
ALLOWLISTED_DRAIN_TASKS: frozenset[str] = frozenset(
{
# Liveness of work already in flight.
"heartbeat_issue_lock",
"heartbeat_reviewer_pr_lease",
"post_heartbeat",
# Handing claims back so nothing is stranded across the restart.
"release_workflow_lease",
"release_reviewer_pr_lease",
"release_merger_pr_lease",
"abandon_workflow_lease",
# Durable recovery state (#660) must be writable *during* drain.
"write_session_checkpoint",
"checkpoint_session",
# The drain controls themselves — exit must never be self-blocked.
"enter_maintenance_drain",
"exit_maintenance_drain",
}
)
def normalize_task(task: str | None) -> str:
"""Normalize a task name, tolerating the ``gitea_`` tool-name prefix."""
name = str(task or "").strip()
if name.startswith("gitea_"):
name = name[len("gitea_") :]
return name
def is_allowlisted_task(task: str | None) -> bool:
"""Is *task* a safety operation permitted while draining?"""
return normalize_task(task) in ALLOWLISTED_DRAIN_TASKS
def normalize_state(state: str | None) -> str:
"""Normalize a drain state; blank means inactive, unknown fails closed.
Blank normalizes to ``inactive`` (no drain record = not draining), but an
unrecognized non-blank value raises: silently treating ``"drainig"`` as
inactive would disable the gate.
"""
value = str(state or "").strip().lower()
if not value:
return STATE_INACTIVE
if value not in DRAIN_STATES:
raise MaintenanceDrainError(
f"unknown maintenance-drain state {value!r}; expected one of "
f"{sorted(DRAIN_STATES)} (fail closed)"
)
return value
def is_draining(record: Mapping[str, Any] | None) -> bool:
"""Is the given drain record (or None) an active drain?"""
if not record:
return False
return normalize_state(record.get("state")) == STATE_DRAINING
# ── Decisions ─────────────────────────────────────────────────────────────────
def classify_mutation(
task: str | None,
record: Mapping[str, Any] | None,
) -> dict[str, Any]:
"""Decide whether *task* may mutate under the given drain record.
Returns a decision dict with ``allowed``/``deferred`` and, when refused, a
typed ``reason_code`` plus the one exact next action the caller may take.
Deferred (not failed): the operation is legal again after drain exits, so
the caller is told to wait rather than to retry a different way.
"""
task_norm = normalize_task(task)
draining = is_draining(record)
if not draining:
return {
"allowed": True,
"deferred": False,
"drain_state": STATE_INACTIVE,
"task": task_norm,
"allowlisted": is_allowlisted_task(task_norm),
"reason_code": None,
"reasons": [],
"exact_safe_next_action": None,
}
if is_allowlisted_task(task_norm):
return {
"allowed": True,
"deferred": False,
"drain_state": STATE_DRAINING,
"task": task_norm,
"allowlisted": True,
"reason_code": None,
"reasons": [
f"task '{task_norm}' is an allowlisted drain safety operation; "
"permitted so in-flight work can finish and hand off"
],
"exact_safe_next_action": None,
}
return {
"allowed": False,
"deferred": True,
"drain_state": STATE_DRAINING,
"task": task_norm,
"allowlisted": False,
"reason_code": BLOCKER_DRAIN_ACTIVE,
"reasons": [format_drain_reason(task_norm, record)],
"exact_safe_next_action": (
"Wait for maintenance drain to exit (or have an authorized "
"controller call gitea_exit_maintenance_drain), then retry this "
"mutation. Reads and gitea_maintenance_drain_status stay available."
),
}
def classify_assignment(record: Mapping[str, Any] | None) -> dict[str, Any]:
"""Decide whether the allocator may assign new work (AC2)."""
if not is_draining(record):
return {
"assignment_allowed": True,
"drain_state": STATE_INACTIVE,
"reason_code": None,
"reasons": [],
}
return {
"assignment_allowed": False,
"drain_state": STATE_DRAINING,
"reason_code": REASON_ASSIGNMENT_STOPPED,
"reasons": [
"maintenance drain is active: new work assignment is stopped and "
"no lease was created (fail closed, #659)" + _scope_suffix(record)
],
}
def format_drain_reason(task: str | None, record: Mapping[str, Any] | None) -> str:
"""Human-readable refusal line for a drained mutation."""
task_norm = normalize_task(task) or "(unnamed task)"
return (
f"maintenance drain is active: mutation '{task_norm}' is deferred; only "
"allowlisted drain safety operations "
f"({', '.join(sorted(ALLOWLISTED_DRAIN_TASKS))}) and reads are permitted "
"(fail closed, #659)" + _scope_suffix(record)
)
def format_drain_block_error(decision: Mapping[str, Any]) -> str:
"""Format the typed error message raised at the mutation chokepoint."""
reasons = list(decision.get("reasons") or [])
head = reasons[0] if reasons else "maintenance drain is active (fail closed)"
action = decision.get("exact_safe_next_action")
return f"{head}. Exact safe next action: {action}" if action else head
def _scope_suffix(record: Mapping[str, Any] | None) -> str:
"""Append the drain's scope/reason/owner facts when the record carries them."""
if not record:
return ""
bits: list[str] = []
scope = "/".join(
str(record.get(key) or "") for key in ("remote", "org", "repo")
).strip("/")
if scope:
bits.append(f"scope {scope}")
if record.get("reason"):
bits.append(f"reason: {record['reason']}")
if record.get("requested_by"):
bits.append(f"entered by {record['requested_by']}")
if record.get("entered_at"):
bits.append(f"at {record['entered_at']}")
return f" ({'; '.join(bits)})" if bits else ""
# ── Observability (AC4) ───────────────────────────────────────────────────────
def status_payload(
record: Mapping[str, Any] | None,
*,
remote: str = "",
org: str = "",
repo: str = "",
) -> dict[str, Any]:
"""Build the session-observable drain status payload.
Always answers, including when no drain record exists: an absent record is
a definitive "not draining", not an unknown.
"""
draining = is_draining(record)
rec: Mapping[str, Any] = record or {}
return {
"drain_state": STATE_DRAINING if draining else STATE_INACTIVE,
"draining": draining,
"remote": str(rec.get("remote") or "") or remote,
"org": str(rec.get("org") or "") or org,
"repo": str(rec.get("repo") or "") or repo,
"reason": str(rec.get("reason") or ""),
"requested_by": str(rec.get("requested_by") or ""),
"requested_by_profile": str(rec.get("requested_by_profile") or ""),
"session_id": str(rec.get("session_id") or ""),
"entered_at": str(rec.get("entered_at") or ""),
"exited_at": str(rec.get("exited_at") or ""),
"assignment_stopped": draining,
"mutations_deferred": draining,
"allowlisted_tasks": sorted(ALLOWLISTED_DRAIN_TASKS),
"reads_permitted": True,
"record_present": bool(record),
"schema_version": DRAIN_SCHEMA_VERSION,
"drain_proof_scope": (
"drain proof and the restart gate that consumes it are #661 scope; "
"this status never authorizes a restart"
),
"safe_next_action": (
"Wait for drain to exit before retrying deferred mutations; "
"allowlisted safety operations and reads remain available."
if draining
else "None; maintenance drain is not active."
),
}
+583
View File
@@ -0,0 +1,583 @@
"""Scoped MCP recovery playbook (#669).
Operational recovery must prefer the *narrowest* action that can fix the
symptom. Full MCP / host restarts are last-resort rungs on a documented
ladder; the coordinator refuses those rungs unless a prior attempt log
shows narrower recoveries already failed (or break-glass is authorized).
This module is pure classification and recommendation:
* No network, filesystem, or process I/O.
* Never restarts anything.
* Narrow recovery *execution* is delegated to existing tools/docs (linked
per rung) — the playbook records which rung to try next and whether
escalation to a broad restart is allowed.
Design lineage: umbrella #655, class matrix #663, coordinator #658,
auto-reconnect #584, stale-runtime #610, contamination #630, audit #665.
Vision #652 / roadmap #653.
"""
from __future__ import annotations
from dataclasses import dataclass, field
from datetime import datetime, timezone
from enum import Enum
from typing import Any, Mapping, Sequence
PLAYBOOK_VERSION = "1.0.0-issue-669"
# Attempt outcomes that count as "tried and insufficient" for escalation.
INSUFFICIENT_OUTCOMES = frozenset(
{
"failed",
"insufficient",
"denied",
"unresolved",
"timeout",
"error",
}
)
# Break-glass / operator override still records that the ladder was skipped.
OUTCOME_BREAK_GLASS = "break_glass"
OUTCOME_SUCCESS = "success"
OUTCOME_SKIPPED = "skipped"
class RecoveryAction(str, Enum):
"""Ordered recovery ladder (narrow → broad)."""
CLIENT_RECONNECT = "client_reconnect"
CAPABILITY_REFRESH = "capability_refresh"
SESSION_RECONNECT = "session_reconnect"
CONFIGURATION_RELOAD = "configuration_reload"
LEASE_RECOVERY = "lease_recovery"
WORKER_RESTART = "worker_restart"
ROLE_RUNTIME_RESTART = "role_runtime_restart"
CONNECTOR_RESTART = "connector_restart"
ROLLING_MCP_RESTART = "rolling_mcp_restart"
FULL_MCP_RESTART = "full_mcp_restart"
HOST_RESTART = "host_restart"
# Classes that require a prior narrow-attempt log (unless break-glass).
BROAD_RESTART_ACTIONS: frozenset[RecoveryAction] = frozenset(
{
RecoveryAction.ROLLING_MCP_RESTART,
RecoveryAction.FULL_MCP_RESTART,
RecoveryAction.HOST_RESTART,
}
)
# Map #663 restart_class strings onto playbook actions.
RESTART_CLASS_TO_ACTION: dict[str, RecoveryAction] = {
"client_reconnect": RecoveryAction.CLIENT_RECONNECT,
"session_reconnect": RecoveryAction.SESSION_RECONNECT,
"configuration_reload": RecoveryAction.CONFIGURATION_RELOAD,
"worker_restart": RecoveryAction.WORKER_RESTART,
"role_runtime_restart": RecoveryAction.ROLE_RUNTIME_RESTART,
"connector_restart": RecoveryAction.CONNECTOR_RESTART,
"rolling_mcp_restart": RecoveryAction.ROLLING_MCP_RESTART,
"full_mcp_restart": RecoveryAction.FULL_MCP_RESTART,
"host_restart": RecoveryAction.HOST_RESTART,
}
@dataclass(frozen=True)
class RecoveryRung:
"""One rung on the recovery ladder."""
action: RecoveryAction
rank: int
summary: str
# Existing implementation or explicit delegation target.
implementation: str
issue_links: tuple[str, ...]
self_service: bool
# Restart-class permission when this rung is requested via coordinator.
restart_class: str | None = None
def as_dict(self) -> dict[str, Any]:
return {
"action": self.action.value,
"rank": self.rank,
"summary": self.summary,
"implementation": self.implementation,
"issue_links": list(self.issue_links),
"self_service": self.self_service,
"restart_class": self.restart_class,
}
# Canonical ladder. Rank 0 is narrowest.
RECOVERY_LADDER: tuple[RecoveryRung, ...] = (
RecoveryRung(
RecoveryAction.CLIENT_RECONNECT,
0,
"Reconnect the IDE/client MCP transport (EOF / transport flap).",
"Host auto-reconnect or explicit client reconnect; "
"docs/mcp-namespace-eof-recovery.md",
("#584", "#655"),
True,
"client_reconnect",
),
RecoveryRung(
RecoveryAction.CAPABILITY_REFRESH,
1,
"Re-resolve task capability and clear stale permission context.",
"Delegated: gitea_resolve_task_capability + gitea_whoami "
"(no process change).",
("#610", "#685", "#655"),
True,
None,
),
RecoveryRung(
RecoveryAction.SESSION_RECONNECT,
2,
"Rebind identity, workspace, and namespace for one session.",
"Delegated: gitea_get_runtime_context + explicit worktree_path "
"rebind (#618); docs/mcp-namespace-health.md",
("#543", "#618", "#655"),
True,
"session_reconnect",
),
RecoveryRung(
RecoveryAction.CONFIGURATION_RELOAD,
3,
"Gracefully reload configuration without replacing the daemon.",
"restart_coordinator class configuration_reload; console "
"system.reload_namespace (#642).",
("#642", "#663", "#655"),
False,
"configuration_reload",
),
RecoveryRung(
RecoveryAction.LEASE_RECOVERY,
4,
"Recover or rebind stale leases/locks without a process restart.",
"Delegated: issue lock recovery / lease lifecycle paths "
"(#702, #753, #790).",
("#702", "#753", "#790", "#655"),
False,
None,
),
RecoveryRung(
RecoveryAction.WORKER_RESTART,
5,
"Restart one worker after its own lease and mutation scope drains.",
"restart_coordinator class worker_restart (#663).",
("#663", "#655"),
False,
"worker_restart",
),
RecoveryRung(
RecoveryAction.ROLE_RUNTIME_RESTART,
6,
"Restart one role runtime and re-probe that namespace only.",
"restart_coordinator class role_runtime_restart; console "
"system.restart_namespace (#642).",
("#642", "#663", "#655"),
False,
"role_runtime_restart",
),
RecoveryRung(
RecoveryAction.CONNECTOR_RESTART,
7,
"Restart one connector while unrelated runtimes stay available.",
"restart_coordinator class connector_restart (#663).",
("#663", "#655"),
False,
"connector_restart",
),
RecoveryRung(
RecoveryAction.ROLLING_MCP_RESTART,
8,
"Drain/restart/verify one instance at a time (HA path).",
"restart_coordinator class rolling_mcp_restart; design #668.",
("#668", "#663", "#655"),
False,
"rolling_mcp_restart",
),
RecoveryRung(
RecoveryAction.FULL_MCP_RESTART,
9,
"Full stable-control MCP process restart after verified full drain.",
"restart_coordinator class full_mcp_restart; requires attempt log "
"unless break-glass (#669).",
("#658", "#661", "#663", "#669", "#655"),
False,
"full_mcp_restart",
),
RecoveryRung(
RecoveryAction.HOST_RESTART,
10,
"Host/infrastructure restart — broadest last-resort action.",
"restart_coordinator class host_restart; operator-owned.",
("#663", "#669", "#655"),
False,
"host_restart",
),
)
_LADDER_BY_ACTION: dict[RecoveryAction, RecoveryRung] = {
rung.action: rung for rung in RECOVERY_LADDER
}
# Symptom tokens → preferred first rung (decision tree, #663 lineage).
SYMPTOM_TO_FIRST_ACTION: dict[str, RecoveryAction] = {
"transport_eof": RecoveryAction.CLIENT_RECONNECT,
"client_closing_eof": RecoveryAction.CLIENT_RECONNECT,
"transport_flap": RecoveryAction.CLIENT_RECONNECT,
"namespace_disconnected": RecoveryAction.CLIENT_RECONNECT,
"stale_capability": RecoveryAction.CAPABILITY_REFRESH,
"permission_stale": RecoveryAction.CAPABILITY_REFRESH,
"runtime_reconnect_required": RecoveryAction.CAPABILITY_REFRESH,
"stale_runtime": RecoveryAction.SESSION_RECONNECT,
"worktree_unbound": RecoveryAction.SESSION_RECONNECT,
"namespace_unhealthy": RecoveryAction.SESSION_RECONNECT,
"config_drift": RecoveryAction.CONFIGURATION_RELOAD,
"profile_misbound": RecoveryAction.CONFIGURATION_RELOAD,
"stale_lease": RecoveryAction.LEASE_RECOVERY,
"dead_pid_lock": RecoveryAction.LEASE_RECOVERY,
"orphan_worktree": RecoveryAction.LEASE_RECOVERY,
"single_worker_stuck": RecoveryAction.WORKER_RESTART,
"role_runtime_dead": RecoveryAction.ROLE_RUNTIME_RESTART,
"connector_dead": RecoveryAction.CONNECTOR_RESTART,
"ha_instance_unhealthy": RecoveryAction.ROLLING_MCP_RESTART,
"daemon_corrupt": RecoveryAction.FULL_MCP_RESTART,
"full_process_deadlock": RecoveryAction.FULL_MCP_RESTART,
"host_unresponsive": RecoveryAction.HOST_RESTART,
}
def _utc_now() -> datetime:
return datetime.now(timezone.utc)
def resolve_action(value: RecoveryAction | str) -> RecoveryAction:
"""Resolve a recovery action or fail closed for unknown values."""
if isinstance(value, RecoveryAction):
return value
text = str(value or "").strip()
# Accept #663 restart_class aliases.
if text in RESTART_CLASS_TO_ACTION:
return RESTART_CLASS_TO_ACTION[text]
try:
return RecoveryAction(text)
except ValueError as exc:
raise ValueError(
f"unknown recovery action {value!r}; deny (fail closed, #669)"
) from exc
def ladder_rank(action: RecoveryAction | str) -> int:
resolved = resolve_action(action)
return _LADDER_BY_ACTION[resolved].rank
def rung_for(action: RecoveryAction | str) -> RecoveryRung:
return _LADDER_BY_ACTION[resolve_action(action)]
def normalize_attempt(raw: Mapping[str, Any]) -> dict[str, Any] | None:
"""Normalize one prior-recovery attempt record; return None if unusable."""
if not isinstance(raw, Mapping):
return None
action_raw = raw.get("action") or raw.get("recovery_action") or raw.get(
"restart_class"
)
if not action_raw:
return None
try:
action = resolve_action(str(action_raw))
except ValueError:
return None
outcome = str(
raw.get("outcome") or raw.get("status") or raw.get("result") or ""
).strip().lower()
if not outcome:
return None
recorded_at = raw.get("recorded_at") or raw.get("at") or raw.get("timestamp")
reason = str(raw.get("reason") or raw.get("detail") or "").strip()
actor = str(raw.get("actor") or raw.get("session_id") or "").strip()
return {
"action": action.value,
"outcome": outcome,
"reason": reason,
"actor": actor,
"recorded_at": recorded_at,
"rank": ladder_rank(action),
"raw": dict(raw),
}
def normalize_attempt_log(
attempts: Sequence[Mapping[str, Any]] | None,
) -> list[dict[str, Any]]:
"""Return usable attempt records in ladder order."""
out: list[dict[str, Any]] = []
for raw in attempts or ():
norm = normalize_attempt(raw)
if norm is not None:
out.append(norm)
out.sort(key=lambda a: (a["rank"], str(a.get("recorded_at") or "")))
return out
def narrower_insufficient_attempts(
attempts: Sequence[Mapping[str, Any]] | None,
*,
requested: RecoveryAction | str,
) -> list[dict[str, Any]]:
"""Return prior attempts narrower than *requested* that were insufficient."""
target_rank = ladder_rank(requested)
usable = []
for attempt in normalize_attempt_log(attempts):
if attempt["rank"] >= target_rank:
continue
if attempt["outcome"] in INSUFFICIENT_OUTCOMES:
usable.append(attempt)
return usable
@dataclass(frozen=True)
class EscalationAssessment:
"""Whether a requested broad recovery may proceed given the attempt log."""
requested_action: str
allowed: bool
require_attempt_log: bool
break_glass: bool
reasons: list[str] = field(default_factory=list)
qualifying_attempts: list[dict[str, Any]] = field(default_factory=list)
recommended_next: list[dict[str, Any]] = field(default_factory=list)
playbook_version: str = PLAYBOOK_VERSION
def as_dict(self) -> dict[str, Any]:
return {
"playbook_version": self.playbook_version,
"requested_action": self.requested_action,
"allowed": self.allowed,
"require_attempt_log": self.require_attempt_log,
"break_glass": self.break_glass,
"reasons": list(self.reasons),
"qualifying_attempts": list(self.qualifying_attempts),
"recommended_next": list(self.recommended_next),
}
def assess_escalation(
requested: RecoveryAction | str,
*,
prior_recovery_attempts: Sequence[Mapping[str, Any]] | None = None,
break_glass: bool = False,
) -> EscalationAssessment:
"""Gate broad restarts on a prior narrow-attempt log (#669 AC3).
Narrow / mid-ladder actions do not require a prior attempt log.
``full_mcp_restart``, ``host_restart``, and ``rolling_mcp_restart``
require at least one *insufficient* narrower attempt unless
``break_glass`` is true.
"""
action = resolve_action(requested)
require_log = action in BROAD_RESTART_ACTIONS
reasons: list[str] = []
qualifying = narrower_insufficient_attempts(
prior_recovery_attempts, requested=action
)
if not require_log:
return EscalationAssessment(
requested_action=action.value,
allowed=True,
require_attempt_log=False,
break_glass=bool(break_glass),
reasons=["narrow recovery; attempt log not required"],
qualifying_attempts=qualifying,
recommended_next=[],
)
if break_glass:
return EscalationAssessment(
requested_action=action.value,
allowed=True,
require_attempt_log=True,
break_glass=True,
reasons=[
"break-glass authorized; broad restart permitted without "
"narrow-attempt log (#669)"
],
qualifying_attempts=qualifying,
recommended_next=[],
)
if qualifying:
return EscalationAssessment(
requested_action=action.value,
allowed=True,
require_attempt_log=True,
break_glass=False,
reasons=[
f"{len(qualifying)} narrower recovery attempt(s) recorded as "
"insufficient; escalation permitted"
],
qualifying_attempts=qualifying,
recommended_next=[],
)
# Deny: recommend the next untried narrow rung(s).
recommended = recommend_actions(
symptoms=(),
prior_recovery_attempts=prior_recovery_attempts,
max_actions=3,
)
reasons.append(
f"{action.value} requires a prior attempt log of insufficient "
"narrower recoveries (or break-glass); none found — deny (fail "
"closed, #669)"
)
return EscalationAssessment(
requested_action=action.value,
allowed=False,
require_attempt_log=True,
break_glass=False,
reasons=reasons,
qualifying_attempts=[],
recommended_next=recommended.get("recommended_actions") or [],
)
def recommend_actions(
*,
symptoms: Sequence[str] = (),
prior_recovery_attempts: Sequence[Mapping[str, Any]] | None = None,
max_actions: int = 5,
) -> dict[str, Any]:
"""Return ordered recommended recovery actions for the given symptoms.
Soft mode (rollout): recommendations only — callers decide whether to
hard-gate. Hard mode for broad restarts is :func:`assess_escalation`.
"""
attempted_success = {
a["action"]
for a in normalize_attempt_log(prior_recovery_attempts)
if a["outcome"] == OUTCOME_SUCCESS
}
attempted_any = {
a["action"] for a in normalize_attempt_log(prior_recovery_attempts)
}
first_actions: list[RecoveryAction] = []
for symptom in symptoms:
key = str(symptom or "").strip().lower().replace(" ", "_").replace("-", "_")
mapped = SYMPTOM_TO_FIRST_ACTION.get(key)
if mapped is not None and mapped not in first_actions:
first_actions.append(mapped)
# Default entry: client reconnect then walk the ladder.
if not first_actions:
first_actions = [RecoveryAction.CLIENT_RECONNECT]
recommended: list[dict[str, Any]] = []
seen: set[str] = set()
min_rank = min(ladder_rank(a) for a in first_actions)
for rung in RECOVERY_LADDER:
if rung.rank < min_rank:
continue
if rung.action.value in attempted_success:
continue
if rung.action.value in seen:
continue
# Prefer rungs not yet attempted; still list previously-failed ones
# only if nothing else remains.
entry = rung.as_dict()
entry["already_attempted"] = rung.action.value in attempted_any
recommended.append(entry)
seen.add(rung.action.value)
if len(recommended) >= max(1, int(max_actions)):
break
return {
"playbook_version": PLAYBOOK_VERSION,
"symptoms": [str(s) for s in symptoms],
"recommended_actions": recommended,
"ladder": [r.as_dict() for r in RECOVERY_LADDER],
"read_only": True,
"hard_gate_note": (
"Broad restarts (rolling/full/host) still require "
"assess_escalation / coordinator attempt-log enforcement."
),
}
def build_attempt_record(
action: RecoveryAction | str,
*,
outcome: str,
reason: str = "",
actor: str = "",
recorded_at: str | None = None,
extra: Mapping[str, Any] | None = None,
) -> dict[str, Any]:
"""Build a durable-shaped attempt log entry for inventory/audit (#665)."""
resolved = resolve_action(action)
record = {
"action": resolved.value,
"outcome": str(outcome or "").strip().lower(),
"reason": str(reason or "").strip(),
"actor": str(actor or "").strip(),
"recorded_at": recorded_at or _utc_now().isoformat(),
"rank": ladder_rank(resolved),
"playbook_version": PLAYBOOK_VERSION,
}
if extra:
record["extra"] = dict(extra)
return record
def recovery_metrics(
attempts: Sequence[Mapping[str, Any]] | None,
) -> dict[str, Any]:
"""Compute the fraction of recoveries that avoided full/host restart.
A recovery *episode* is approximated as one attempt with
``outcome=success``. Successes on non-broad rungs count as avoided full
restart; successes on full/host count as full-restart recoveries.
"""
norms = normalize_attempt_log(attempts)
successes = [a for a in norms if a["outcome"] == OUTCOME_SUCCESS]
broad_success = [
a
for a in successes
if resolve_action(a["action"])
in {RecoveryAction.FULL_MCP_RESTART, RecoveryAction.HOST_RESTART}
]
avoided = [a for a in successes if a not in broad_success]
total = len(successes)
fraction_avoided = (len(avoided) / total) if total else None
return {
"playbook_version": PLAYBOOK_VERSION,
"attempts_total": len(norms),
"successes_total": total,
"successes_avoided_full_restart": len(avoided),
"successes_full_or_host_restart": len(broad_success),
"fraction_avoided_full_restart": fraction_avoided,
"insufficient_attempts": sum(
1 for a in norms if a["outcome"] in INSUFFICIENT_OUTCOMES
),
}
def ladder_document() -> dict[str, Any]:
"""Machine-readable ladder for docs/tools inventory."""
return {
"playbook_version": PLAYBOOK_VERSION,
"parent_issues": ["#655", "#652", "#653"],
"enforcement_issue": "#669",
"ladder": [r.as_dict() for r in RECOVERY_LADDER],
"broad_restart_actions": [a.value for a in sorted(BROAD_RESTART_ACTIONS, key=lambda x: x.value)],
"insufficient_outcomes": sorted(INSUFFICIENT_OUTCOMES),
"symptom_map": {k: v.value for k, v in sorted(SYMPTOM_TO_FIRST_ACTION.items())},
}
+46 -2
View File
@@ -1,4 +1,4 @@
"""MCP restart coordinator and impact analysis (#658).
"""MCP restart coordinator and impact analysis (#658 / #669).
Before any sanctioned MCP restart, a central coordinator must evaluate the
live control-plane state — active sessions, leases/locks, in-flight issue/PR
@@ -16,6 +16,9 @@ Design rules (mirrors the read-only posture of ``workflow_dashboard`` /
a mutative apply path is a later child gated by a drain proof (non-goal here).
* **Fail closed.** If the inventory is not explicitly complete, the verdict is
``unsafe`` / deny — an incomplete evaluation must never green-light a restart.
* **Narrow-first (#669).** Broad classes (rolling / full / host) require a
prior attempt log of insufficient narrower recoveries unless break-glass is
authorized. See :mod:`recovery_playbook`.
* **No secrets.** Session ids, pids, and profiles are operational metadata, not
credentials; nothing secret flows through this module.
@@ -32,8 +35,9 @@ from enum import Enum
from typing import Any, Mapping, Sequence
import lease_lifecycle
import recovery_playbook
COORDINATOR_VERSION = "1.1.0-issue-663"
COORDINATOR_VERSION = "1.2.0-issue-669"
# Restart verdicts. Exactly the three the acceptance criteria name.
VERDICT_SAFE = "safe"
@@ -349,6 +353,10 @@ class RestartImpactReport:
counts: dict[str, int]
audit_record: dict[str, Any]
incomplete_reasons: list[str] = field(default_factory=list)
# #669 playbook escalation gate (attempt-log enforcement).
playbook_escalation: dict[str, Any] = field(default_factory=dict)
attempt_log_satisfied: bool = True
break_glass: bool = False
def as_dict(self) -> dict[str, Any]:
return {
@@ -382,6 +390,9 @@ class RestartImpactReport:
"prior_recovery_attempts": list(self.prior_recovery_attempts),
"counts": dict(self.counts),
"audit_record": dict(self.audit_record),
"playbook_escalation": dict(self.playbook_escalation),
"attempt_log_satisfied": self.attempt_log_satisfied,
"break_glass": self.break_glass,
}
@@ -496,6 +507,7 @@ def evaluate_restart_impact(
target_session_id: str | None = None,
target_role: str | None = None,
target_connector: str | None = None,
break_glass: bool = False,
) -> RestartImpactReport:
"""Evaluate a proposed MCP restart and return an impact preview.
@@ -584,6 +596,30 @@ def evaluate_restart_impact(
dict(a) for a in (inventory.get("prior_recovery_attempts") or [])
]
# #669: broad restarts require a prior narrow-attempt log unless break-glass.
playbook_escalation: dict[str, Any] = {}
attempt_log_satisfied = True
if policy_enforced and resolved_class is not None:
try:
escalation = recovery_playbook.assess_escalation(
resolved_class.value,
prior_recovery_attempts=prior_recovery_attempts,
break_glass=bool(break_glass),
)
playbook_escalation = escalation.as_dict()
attempt_log_satisfied = bool(escalation.allowed)
if not attempt_log_satisfied:
authorization_reasons.extend(list(escalation.reasons))
except ValueError as exc:
# Unknown mapping should never happen for enum values; fail closed.
attempt_log_satisfied = False
playbook_escalation = {
"allowed": False,
"reasons": [str(exc)],
"playbook_version": recovery_playbook.PLAYBOOK_VERSION,
}
authorization_reasons.append(str(exc))
session_impacts = [
_classify_session(
s,
@@ -682,6 +718,7 @@ def evaluate_restart_impact(
and role_authorized
and approval_satisfied
and target_complete
and attempt_log_satisfied
)
if policy_enforced and not authorization_ok:
@@ -745,6 +782,7 @@ def evaluate_restart_impact(
"affected_issues": len(affected_issues),
"affected_prs": len(affected_prs),
"prior_recovery_attempts": len(prior_recovery_attempts),
"attempt_log_satisfied": attempt_log_satisfied,
}
audit_record = {
@@ -765,6 +803,9 @@ def evaluate_restart_impact(
"allow_restart": allow_restart,
"blast_radius": blast_radius,
"counts": counts,
"attempt_log_satisfied": attempt_log_satisfied,
"break_glass": bool(break_glass),
"playbook_version": recovery_playbook.PLAYBOOK_VERSION,
}
return RestartImpactReport(
@@ -804,4 +845,7 @@ def evaluate_restart_impact(
counts=counts,
audit_record=audit_record,
incomplete_reasons=incomplete_reasons,
playbook_escalation=playbook_escalation,
attempt_log_satisfied=attempt_log_satisfied,
break_glass=bool(break_glass),
)
+5
View File
@@ -63,6 +63,11 @@ CONTAMINATION_GATED_TASKS = frozenset({
"merge_pr",
"delete_branch",
"complete_issue",
# Web console recovery playbooks that write (#644). These mutate runtime
# binding and process state, so a live contamination marker must block them
# exactly as it blocks the Gitea-side mutations above. The reconciler
# cleanup playbook is the designated remedy and is exempted by its caller.
"console_recovery_apply",
})
CONTAMINATION_KIND = "stable_branch_push"
+35
View File
@@ -142,6 +142,23 @@ TASK_CAPABILITY_MAP: dict[str, dict[str, str]] = {
"permission": "gitea.read",
"role": "author",
},
# #644: Phase 2 Web Console recovery tasks.
"clear_stale_binding": {
"permission": "gitea.read",
"role": "author",
},
"rebind_session_worktree": {
"permission": "gitea.read",
"role": "author",
},
# The console playbook orchestrates gitea_reconcile_merged_cleanups, whose
# own gate is gitea.read (matching the existing reconcile_merged_cleanups
# entry). Declaring a stricter permission here stated a second, conflicting
# authority for one operation.
"reconcile_cleanups": {
"permission": "gitea.read",
"role": "reconciler",
},
# PR synchronization lifecycle: assess is read-only (any role with gitea.read);
# update-by-merge is author-only and mutates the PR head via Gitea API.
"assess_pr_sync_status": {
@@ -397,6 +414,24 @@ TASK_CAPABILITY_MAP: dict[str, dict[str, str]] = {
"role": "controller",
},
# #659 maintenance drain. Same reasoning as the lifecycle controls above:
# entering/exiting drain quiesces a whole namespace, so it carries a
# non-``gitea.*`` permission that no configured Gitea profile satisfies by
# accident (AC1 — capability-gated and audited). Reading drain state is
# ordinary read authority: every session must be able to see the drain (AC4).
"enter_maintenance_drain": {
"permission": "runtime.maintenance_drain",
"role": "controller",
},
"exit_maintenance_drain": {
"permission": "runtime.maintenance_drain",
"role": "controller",
},
"maintenance_drain_status": {
"permission": "gitea.read",
"role": "author",
},
# #601 first-class lease lifecycle — inspect/list need read; mutations gate on
# ownership in the control-plane DB (not a separate Gitea write permission).
"list_workflow_leases": {
+158
View File
@@ -7,6 +7,7 @@ import tempfile
import threading
import unittest
from concurrent.futures import ThreadPoolExecutor, as_completed
from datetime import datetime, timezone
from allocator_service import (
OUTCOME_ASSIGNED,
@@ -15,6 +16,7 @@ from allocator_service import (
OUTCOME_PREVIEW,
OUTCOME_WAIT,
WorkCandidate,
_drop_expired_claims,
allocate_next_work,
candidate_from_dict,
classify_skip,
@@ -362,5 +364,161 @@ class AllocatorServiceTest(unittest.TestCase):
self.assertIn("unavailable", res["reasons"][0].lower())
class SideEffectFreeAllocationTest(unittest.TestCase):
"""``side_effect_free`` dry runs write nothing to the control plane (#643).
A plain ``apply=False`` still called ``upsert_session`` and
``expire_stale_leases`` before the apply branch was consulted, so a caller
advertising a read-only preview mutated on every call one unreferenced
session row per preview, plus a global lease sweep.
"""
def setUp(self) -> None:
self._tmp = tempfile.TemporaryDirectory()
self.db = ControlPlaneDB(os.path.join(self._tmp.name, "cp.sqlite3"))
def tearDown(self) -> None:
self._tmp.cleanup()
def _alloc(self, **kwargs):
defaults = dict(
db=self.db,
session_id="s-preview",
role="author",
remote="prgs",
org="org",
repo="repo",
candidates=[
WorkCandidate(kind="issue", number=643, labels=("status:ready",))
],
apply=False,
profile_name="prgs-author",
username="jcwalker3",
)
defaults.update(kwargs)
return allocate_next_work(**defaults)
def _session_ids(self) -> set[str]:
return {str(r.get("session_id")) for r in self.db.list_sessions()}
def test_side_effect_free_preview_writes_no_session_row(self):
before = self._session_ids()
result = self._alloc(side_effect_free=True)
self.assertEqual(result["outcome"], OUTCOME_PREVIEW)
self.assertEqual(self._session_ids(), before)
self.assertNotIn("s-preview", self._session_ids())
def test_plain_dry_run_still_registers_a_session(self):
# The default is unchanged for every existing caller.
self._alloc()
self.assertIn("s-preview", self._session_ids())
def test_repeated_previews_do_not_accumulate_rows(self):
for index in range(5):
self._alloc(side_effect_free=True, session_id=f"s-{index}")
self.assertEqual(self._session_ids(), set())
def test_side_effect_free_does_not_sweep_stale_leases(self):
self.db.upsert_session(session_id="owner", role="author", pid=1)
assigned = self.db.assign_and_lease(
session_id="owner",
role="author",
remote="prgs",
org="org",
repo="repo",
kind="issue",
number=999,
lease_ttl_seconds=-60, # already expired
)
self.assertEqual(assigned.outcome, "assigned")
self._alloc(side_effect_free=True)
# The expired row is still 'active' in the DB: nothing swept it.
statuses = {
r["lease_id"]: r["status"]
for r in self.db.list_leases(
remote="prgs", org="org", repo="repo",
statuses=("active", "expired"),
)
}
self.assertEqual(statuses.get(assigned.lease_id), "active")
def test_expired_claims_are_filtered_in_memory_so_work_stays_selectable(self):
"""The read-only mirror of the sweep: expired claims must not block."""
self.db.upsert_session(session_id="owner", role="author", pid=1)
self.db.assign_and_lease(
session_id="owner",
role="author",
remote="prgs",
org="org",
repo="repo",
kind="issue",
number=643,
lease_ttl_seconds=-60, # expired: must not withhold #643
)
result = self._alloc(side_effect_free=True)
self.assertEqual(result["outcome"], OUTCOME_PREVIEW)
self.assertEqual(result["selected"]["number"], 643)
def test_a_live_claim_still_withholds_the_work(self):
self.db.upsert_session(session_id="owner", role="author", pid=1)
self.db.assign_and_lease(
session_id="owner",
role="author",
remote="prgs",
org="org",
repo="repo",
kind="issue",
number=643,
lease_ttl_seconds=3600,
)
result = self._alloc(side_effect_free=True)
self.assertNotEqual(result["outcome"], OUTCOME_ASSIGNED)
self.assertNotEqual((result.get("selected") or {}).get("number"), 643)
def test_side_effect_free_with_apply_fails_closed(self):
result = self._alloc(side_effect_free=True, apply=True)
self.assertFalse(result["success"])
self.assertEqual(result["outcome"], OUTCOME_NO_SAFE)
self.assertIsNone(result["assignment"])
self.assertIn("incompatible with apply", result["reasons"][0])
# And it reserved nothing.
self.assertEqual(
self.db.list_leases(remote="prgs", org="org", repo="repo"), []
)
class DropExpiredClaimsTest(unittest.TestCase):
"""The in-memory expiry filter behind side-effect-free previews (#643)."""
def test_unparseable_expiry_is_kept_rather_than_assumed_free(self):
claims = {
("issue", 1): {"lease_id": "l1", "expires_at": "not-a-date"},
("issue", 2): {"lease_id": "l2"},
("issue", 3): {"lease_id": "l3", "expires_at": None},
}
self.assertEqual(_drop_expired_claims(claims), claims)
def test_expired_dropped_and_future_kept(self):
now = datetime(2026, 7, 25, 12, 0, tzinfo=timezone.utc)
claims = {
("issue", 1): {"expires_at": "2026-07-25T11:59:59+00:00"},
("issue", 2): {"expires_at": "2026-07-25T12:00:01+00:00"},
("issue", 3): {"expires_at": "2026-07-25T12:00:00+00:00"}, # boundary
}
kept = _drop_expired_claims(claims, now=now)
self.assertEqual(set(kept), {("issue", 2)})
def test_naive_and_zulu_timestamps_are_treated_as_utc(self):
now = datetime(2026, 7, 25, 12, 0, tzinfo=timezone.utc)
claims = {
("issue", 1): {"expires_at": "2026-07-25T11:00:00"}, # naive, past
("issue", 2): {"expires_at": "2026-07-25T13:00:00Z"}, # zulu, future
}
kept = _drop_expired_claims(claims, now=now)
self.assertEqual(set(kept), {("issue", 2)})
if __name__ == "__main__":
unittest.main()
@@ -35,6 +35,17 @@ BREAK_GLASS_ENV = "GITEA_BREAKGLASS_RESTART_AUTHORIZATION"
QUIET_SESSIONS: list[dict] = []
QUIET_LEASES: list[dict] = []
# #669: broad restarts need a prior narrow-attempt log (unless break-glass).
PRIOR_NARROW_ATTEMPTS_JSON = json.dumps(
[
{
"action": "client_reconnect",
"outcome": "insufficient",
"reason": "still flapping after reconnect",
}
]
)
class _FakeDB:
"""Minimal control-plane DB stand-in for the restart inventory."""
@@ -128,6 +139,7 @@ class TestConjunction(_RestartToolHarness):
preview = self._call(
role="operator",
restart_class="full_mcp_restart",
prior_recovery_attempts_json=PRIOR_NARROW_ATTEMPTS_JSON,
env={CONTROLLER_APPROVAL_ENV: "operator-approved"},
)
self.assertTrue(preview["allow_restart"],
@@ -136,6 +148,7 @@ class TestConjunction(_RestartToolHarness):
result = self._call(
role="operator",
restart_class="full_mcp_restart",
prior_recovery_attempts_json=PRIOR_NARROW_ATTEMPTS_JSON,
dry_run=False,
drain_proof_json=self._clean_proof_for(preview),
env={CONTROLLER_APPROVAL_ENV: "operator-approved"},
@@ -342,6 +355,7 @@ class TestExistingPathsStillWork(_RestartToolHarness):
result = self._call(
role="operator",
restart_class="full_mcp_restart",
prior_recovery_attempts_json=PRIOR_NARROW_ATTEMPTS_JSON,
dry_run=False,
env={CONTROLLER_APPROVAL_ENV: "operator-approved"},
)
+237
View File
@@ -0,0 +1,237 @@
"""Tests for graceful MCP maintenance-drain mode (#659).
Acceptance coverage:
1. Enter/exit is durable and audited (DB substrate).
2. New work assignment stops during drain (allocator WAIT).
3. Mutations deferred except allowlisted safety ops.
4. Sessions can observe drain state.
5. Fail-closed on unreadable drain state.
"""
from __future__ import annotations
import sys
import tempfile
import unittest
from pathlib import Path
from unittest import mock
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
import maintenance_drain
from control_plane_db import ControlPlaneDB
from allocator_service import WorkCandidate, allocate_next_work, OUTCOME_WAIT
class TestDrainDecisions(unittest.TestCase):
def test_inactive_allows_mutations_and_assignment(self):
decision = maintenance_drain.classify_mutation("create_pr", None)
self.assertTrue(decision["allowed"])
self.assertFalse(decision["deferred"])
assign = maintenance_drain.classify_assignment(None)
self.assertTrue(assign["assignment_allowed"])
def test_draining_defers_non_allowlisted_mutation(self):
record = {"state": "draining", "remote": "prgs", "org": "o", "repo": "r"}
decision = maintenance_drain.classify_mutation("create_pr", record)
self.assertFalse(decision["allowed"])
self.assertTrue(decision["deferred"])
self.assertEqual(decision["reason_code"], maintenance_drain.BLOCKER_DRAIN_ACTIVE)
self.assertIn("create_pr", decision["reasons"][0])
def test_allowlisted_safety_ops_pass_during_drain(self):
record = {"state": "draining"}
for task in (
"heartbeat_issue_lock",
"gitea_release_reviewer_pr_lease",
"write_session_checkpoint",
"exit_maintenance_drain",
):
with self.subTest(task=task):
decision = maintenance_drain.classify_mutation(task, record)
self.assertTrue(decision["allowed"], decision)
def test_assignment_stopped_during_drain(self):
record = {"state": "draining", "reason": "upgrade"}
decision = maintenance_drain.classify_assignment(record)
self.assertFalse(decision["assignment_allowed"])
self.assertEqual(
decision["reason_code"], maintenance_drain.REASON_ASSIGNMENT_STOPPED
)
def test_unknown_state_fails_closed(self):
with self.assertRaises(maintenance_drain.MaintenanceDrainError):
maintenance_drain.normalize_state("drainig")
def test_status_payload_always_answers(self):
inactive = maintenance_drain.status_payload(None, remote="prgs", org="o", repo="r")
self.assertFalse(inactive["draining"])
self.assertTrue(inactive["reads_permitted"])
active = maintenance_drain.status_payload(
{"state": "draining", "reason": "reboot", "requested_by": "ops"},
remote="prgs",
org="o",
repo="r",
)
self.assertTrue(active["draining"])
self.assertTrue(active["assignment_stopped"])
self.assertTrue(active["mutations_deferred"])
self.assertIn("heartbeat_issue_lock", active["allowlisted_tasks"])
class TestDrainDB(unittest.TestCase):
def setUp(self):
self._tmp = tempfile.TemporaryDirectory()
self.addCleanup(self._tmp.cleanup)
self.db = ControlPlaneDB(db_path=str(Path(self._tmp.name) / "cp.sqlite3"))
def test_enter_exit_idempotent_and_audited(self):
first = self.db.set_maintenance_drain(
remote="prgs",
org="Scaled-Tech-Consulting",
repo="Gitea-Tools",
state="draining",
reason="planned restart",
requested_by="sysadmin",
requested_by_profile="prgs-controller",
session_id="s1",
)
self.assertTrue(first["transitioned"])
self.assertEqual(first["state"], "draining")
self.assertTrue(maintenance_drain.is_draining(first["record"]))
again = self.db.set_maintenance_drain(
remote="prgs",
org="Scaled-Tech-Consulting",
repo="Gitea-Tools",
state="draining",
reason="still draining",
requested_by="sysadmin",
requested_by_profile="prgs-controller",
session_id="s1",
)
self.assertFalse(again["transitioned"])
self.assertEqual(again["record"]["entered_at"], first["record"]["entered_at"])
exited = self.db.set_maintenance_drain(
remote="prgs",
org="Scaled-Tech-Consulting",
repo="Gitea-Tools",
state="inactive",
reason="done",
requested_by="sysadmin",
requested_by_profile="prgs-controller",
session_id="s1",
)
self.assertTrue(exited["transitioned"])
self.assertFalse(maintenance_drain.is_draining(exited["record"]))
self.assertTrue(exited["record"]["exited_at"])
# Events recorded for transitions only (enter + exit).
with self.db._tx(immediate=False) as conn:
rows = conn.execute(
"SELECT event_type FROM events WHERE event_type LIKE 'maintenance_drain_%' "
"ORDER BY event_id"
).fetchall()
types = [r[0] for r in rows]
self.assertEqual(types, ["maintenance_drain_enter", "maintenance_drain_exit"])
def test_read_missing_is_none_not_error(self):
self.assertIsNone(
self.db.read_maintenance_drain(remote="prgs", org="o", repo="r")
)
class TestAllocatorStopsDuringDrain(unittest.TestCase):
def setUp(self):
self._tmp = tempfile.TemporaryDirectory()
self.addCleanup(self._tmp.cleanup)
self.db = ControlPlaneDB(db_path=str(Path(self._tmp.name) / "cp.sqlite3"))
def test_allocate_returns_wait_while_draining(self):
self.db.set_maintenance_drain(
remote="prgs",
org="Scaled-Tech-Consulting",
repo="Gitea-Tools",
state="draining",
reason="test",
requested_by="tester",
)
candidates = [
WorkCandidate(
kind="issue",
number=659,
title="drain",
labels=("status:ready",),
priority=20,
)
]
result = allocate_next_work(
self.db,
role="author",
session_id="test-session",
remote="prgs",
org="Scaled-Tech-Consulting",
repo="Gitea-Tools",
apply=False,
candidates=candidates,
username="jcwalker3",
profile_name="prgs-author",
)
self.assertEqual(result["outcome"], OUTCOME_WAIT)
self.assertIsNone(result.get("selected"))
self.assertEqual(
result.get("reason_code"),
maintenance_drain.REASON_ASSIGNMENT_STOPPED,
)
self.assertTrue(result["maintenance_drain"]["draining"])
def test_allocate_works_when_inactive(self):
candidates = [
WorkCandidate(
kind="issue",
number=659,
title="drain",
labels=("status:ready",),
priority=20,
)
]
result = allocate_next_work(
self.db,
role="author",
session_id="test-session-2",
remote="prgs",
org="Scaled-Tech-Consulting",
repo="Gitea-Tools",
apply=False,
candidates=candidates,
username="jcwalker3",
profile_name="prgs-author",
)
self.assertNotEqual(
result.get("reason_code"),
maintenance_drain.REASON_ASSIGNMENT_STOPPED,
)
class TestCapabilityMap(unittest.TestCase):
def test_drain_tasks_mapped(self):
import task_capability_map as tcm
self.assertEqual(
tcm.required_permission("enter_maintenance_drain"),
"runtime.maintenance_drain",
)
self.assertEqual(
tcm.required_permission("exit_maintenance_drain"),
"runtime.maintenance_drain",
)
self.assertEqual(
tcm.required_permission("maintenance_drain_status"),
"gitea.read",
)
if __name__ == "__main__":
unittest.main()
+217
View File
@@ -0,0 +1,217 @@
"""Unit tests for the scoped recovery playbook (#669)."""
from __future__ import annotations
import recovery_playbook as rp
import restart_coordinator as rc
def test_ladder_covers_eleven_ordered_rungs():
ranks = [r.rank for r in rp.RECOVERY_LADDER]
assert ranks == list(range(len(rp.RECOVERY_LADDER)))
assert len(rp.RECOVERY_LADDER) == 11
assert rp.RECOVERY_LADDER[0].action is rp.RecoveryAction.CLIENT_RECONNECT
assert rp.RECOVERY_LADDER[-1].action is rp.RecoveryAction.HOST_RESTART
def test_ladder_document_links_parent_issues():
doc = rp.ladder_document()
assert "#655" in doc["parent_issues"]
assert "#652" in doc["parent_issues"]
assert "#653" in doc["parent_issues"]
assert doc["enforcement_issue"] == "#669"
assert "full_mcp_restart" in doc["broad_restart_actions"]
def test_recommend_transport_eof_starts_at_client_reconnect():
plan = rp.recommend_actions(symptoms=["transport_eof"])
assert plan["recommended_actions"][0]["action"] == "client_reconnect"
assert plan["recommended_actions"][0]["issue_links"]
def test_recommend_skips_successful_prior_attempts():
attempts = [
rp.build_attempt_record(
"client_reconnect", outcome="success", reason="reconnected"
)
]
plan = rp.recommend_actions(
symptoms=["transport_eof"], prior_recovery_attempts=attempts
)
actions = [a["action"] for a in plan["recommended_actions"]]
assert "client_reconnect" not in actions
assert actions[0] == "capability_refresh"
def test_escalation_denied_without_attempt_log():
result = rp.assess_escalation("full_mcp_restart", prior_recovery_attempts=[])
assert result.allowed is False
assert result.require_attempt_log is True
assert any("#669" in r for r in result.reasons)
assert result.recommended_next # soft recommendations still provided
def test_escalation_allowed_after_insufficient_narrower():
attempts = [
rp.build_attempt_record(
"client_reconnect",
outcome="insufficient",
reason="still flapping",
),
rp.build_attempt_record(
"session_reconnect",
outcome="failed",
reason="namespace still dead",
),
]
result = rp.assess_escalation(
"full_mcp_restart", prior_recovery_attempts=attempts
)
assert result.allowed is True
assert len(result.qualifying_attempts) == 2
def test_escalation_break_glass_bypasses_attempt_log():
result = rp.assess_escalation(
"host_restart", prior_recovery_attempts=[], break_glass=True
)
assert result.allowed is True
assert result.break_glass is True
def test_narrow_action_does_not_require_attempt_log():
result = rp.assess_escalation(
"client_reconnect", prior_recovery_attempts=[]
)
assert result.allowed is True
assert result.require_attempt_log is False
def test_same_rank_attempt_does_not_qualify_for_escalation():
attempts = [
rp.build_attempt_record(
"full_mcp_restart", outcome="failed", reason="already failed full"
)
]
result = rp.assess_escalation(
"full_mcp_restart", prior_recovery_attempts=attempts
)
assert result.allowed is False
def test_success_outcome_does_not_qualify_for_escalation():
attempts = [
rp.build_attempt_record(
"client_reconnect", outcome="success", reason="fixed"
)
]
result = rp.assess_escalation(
"full_mcp_restart", prior_recovery_attempts=attempts
)
assert result.allowed is False
def test_recovery_metrics_fraction_avoided():
attempts = [
rp.build_attempt_record("client_reconnect", outcome="success"),
rp.build_attempt_record("session_reconnect", outcome="success"),
rp.build_attempt_record("full_mcp_restart", outcome="success"),
]
metrics = rp.recovery_metrics(attempts)
assert metrics["successes_total"] == 3
assert metrics["successes_avoided_full_restart"] == 2
assert metrics["successes_full_or_host_restart"] == 1
assert abs(metrics["fraction_avoided_full_restart"] - (2 / 3)) < 1e-9
def test_coordinator_denies_full_restart_without_attempt_log():
inv = {
"inventory_complete": True,
"sessions": [],
"leases": [],
"prior_recovery_attempts": [],
}
report = rc.evaluate_restart_impact(
inv,
restart_class=rc.RestartClass.FULL_MCP_RESTART,
requester_role="controller",
requester_permissions=rc.permissions_for_role("controller"),
controller_approved=True,
operator_authorized=True,
)
assert report.allow_restart is False
assert report.attempt_log_satisfied is False
assert report.verdict == rc.VERDICT_UNSAFE
blob = " ".join(report.reasons + report.authorization_reasons)
assert "#669" in blob or "attempt log" in blob
def test_coordinator_allows_full_restart_with_attempt_log():
inv = {
"inventory_complete": True,
"sessions": [],
"leases": [],
"prior_recovery_attempts": [
{
"action": "client_reconnect",
"outcome": "insufficient",
"reason": "still broken",
}
],
}
report = rc.evaluate_restart_impact(
inv,
restart_class=rc.RestartClass.FULL_MCP_RESTART,
requester_role="controller",
requester_permissions=rc.permissions_for_role("controller"),
controller_approved=True,
operator_authorized=True,
)
assert report.attempt_log_satisfied is True
assert report.allow_restart is True
assert report.verdict == rc.VERDICT_SAFE
def test_coordinator_break_glass_allows_without_log():
inv = {
"inventory_complete": True,
"sessions": [],
"leases": [],
"prior_recovery_attempts": [],
}
report = rc.evaluate_restart_impact(
inv,
restart_class=rc.RestartClass.FULL_MCP_RESTART,
requester_role="controller",
requester_permissions=rc.permissions_for_role("controller"),
controller_approved=True,
operator_authorized=True,
break_glass=True,
)
assert report.break_glass is True
assert report.attempt_log_satisfied is True
assert report.allow_restart is True
def test_coordinator_client_reconnect_unaffected():
inv = {
"inventory_complete": True,
"sessions": [],
"leases": [],
"prior_recovery_attempts": [],
}
report = rc.evaluate_restart_impact(
inv,
restart_class=rc.RestartClass.CLIENT_RECONNECT,
requester_role="author",
requester_permissions=rc.permissions_for_role("author"),
)
assert report.attempt_log_satisfied is True
assert report.allow_restart is True
def test_restart_class_alias_accepted():
assert (
rp.resolve_action("full_mcp_restart")
is rp.RecoveryAction.FULL_MCP_RESTART
)
+506
View File
@@ -0,0 +1,506 @@
"""Unit and integration tests for Phase 2 Web Console recovery controls (#644)."""
from __future__ import annotations
import os
import sys
import types
import unittest
from unittest.mock import patch
from starlette.testclient import TestClient
import merged_cleanup_reconcile
import runtime_recovery_guard
import stable_branch_push_guard
import stale_binding_recovery
from webui import console_authz, console_recovery, system_health
from webui.app import create_app
class TestConsoleRecovery(unittest.TestCase):
def test_diagnose_recovery_healthy(self) -> None:
diag = console_recovery.diagnose_recovery()
self.assertIn(diag.status, {console_recovery.STATUS_HEALTHY, console_recovery.STATUS_ACTION_REQUIRED})
self.assertIsInstance(diag.playbooks, tuple)
self.assertGreaterEqual(len(diag.playbooks), 4)
playbook_ids = {pb.playbook_id for pb in diag.playbooks}
self.assertIn(console_recovery.PLAYBOOK_CLEAR_STALE_BINDING, playbook_ids)
self.assertIn(console_recovery.PLAYBOOK_REBIND_SESSION, playbook_ids)
self.assertIn(console_recovery.PLAYBOOK_RECONCILE_CLEANUPS, playbook_ids)
self.assertIn(console_recovery.PLAYBOOK_SANCTIONED_RESTART, playbook_ids)
def test_confirmation_phrase_generation_and_matching(self) -> None:
phrase = console_recovery.confirmation_phrase("clear_stale_binding")
self.assertEqual(phrase, "confirm clear_stale_binding")
self.assertTrue(console_recovery.confirmation_matches("clear_stale_binding", "confirm clear_stale_binding"))
self.assertFalse(console_recovery.confirmation_matches("clear_stale_binding", "wrong phrase"))
phrase_target = console_recovery.confirmation_phrase("sanctioned_restart", "gitea-author")
self.assertEqual(phrase_target, "confirm sanctioned_restart gitea-author")
self.assertTrue(console_recovery.confirmation_matches("sanctioned_restart", "confirm sanctioned_restart gitea-author", "gitea-author"))
def test_build_recovery_preview(self) -> None:
principal = console_authz.Principal("[email protected]", console_authz.OPERATOR, console_authz.IDENTITY_LOCAL_DEV, True)
preview = console_recovery.build_recovery_preview(
playbook_id=console_recovery.PLAYBOOK_CLEAR_STALE_BINDING,
target="test-worktree",
principal=principal,
)
self.assertEqual(preview["playbook_id"], console_recovery.PLAYBOOK_CLEAR_STALE_BINDING)
self.assertEqual(preview["action_id"], console_recovery.ACTION_CLEAR_STALE_BINDING)
self.assertEqual(preview["confirmation_phrase"], "confirm clear_stale_binding test-worktree")
self.assertTrue(len(preview["mutation_ledger"]) >= 3)
self.assertTrue(preview["authorization"]["allowed"])
def test_build_recovery_preview_unknown_playbook(self) -> None:
preview = console_recovery.build_recovery_preview("unknown_playbook")
self.assertFalse(preview.get("allowed"))
self.assertEqual(preview.get("error"), "unknown_playbook")
def test_execute_recovery_playbook_confirmation_mismatch(self) -> None:
# Authorization is checked before confirmation, so the phase gate has to
# pass for this test to reach the branch it is about.
principal = console_authz.Principal("[email protected]", console_authz.OPERATOR, console_authz.IDENTITY_LOCAL_DEV, True)
with self._phase_two_enabled():
result = console_recovery.execute_recovery_playbook(
playbook_id=console_recovery.PLAYBOOK_CLEAR_STALE_BINDING,
confirmation="invalid confirmation",
principal=principal,
)
self.assertFalse(result["success"])
self.assertFalse(result["allowed"])
self.assertEqual(result["error"], "confirmation_mismatch")
def test_execute_recovery_playbook_unauthorized(self) -> None:
# Anonymous principal has viewer role -> should be denied
result = console_recovery.execute_recovery_playbook(
playbook_id=console_recovery.PLAYBOOK_CLEAR_STALE_BINDING,
confirmation="confirm clear_stale_binding",
principal=console_authz.ANONYMOUS,
)
self.assertFalse(result["success"])
self.assertFalse(result["allowed"])
self.assertEqual(result["error"], console_authz.DENY_UNAUTHENTICATED)
def test_execute_refuses_phase_two_write_while_console_is_phase_one(self) -> None:
"""B1: the apply path must arm the phase gate, not skip it.
``authorize`` only applies the phase branch when ``for_execution=True``.
The apply path used the default, so an operator executed a phase-2 write
while ``ACTIVE_PHASE`` was 1.
"""
self.assertGreater(
console_authz.get_action(console_recovery.ACTION_CLEAR_STALE_BINDING).phase,
console_authz.ACTIVE_PHASE,
"fixture assumes the recovery actions are ahead of the active phase",
)
principal = console_authz.Principal(
"[email protected]", console_authz.OPERATOR, console_authz.IDENTITY_LOCAL_DEV, True
)
phrase = console_recovery.confirmation_phrase(
console_recovery.PLAYBOOK_CLEAR_STALE_BINDING
)
result = console_recovery.execute_recovery_playbook(
playbook_id=console_recovery.PLAYBOOK_CLEAR_STALE_BINDING,
confirmation=phrase,
principal=principal,
)
self.assertFalse(result["success"])
self.assertFalse(result["allowed"])
self.assertEqual(result["error"], console_authz.DENY_PHASE_NOT_ACTIVE)
def test_preview_execution_enabled_matches_the_execution_decision(self) -> None:
"""B1: preview must not report a bare False it cannot explain."""
principal = console_authz.Principal(
"[email protected]", console_authz.OPERATOR, console_authz.IDENTITY_LOCAL_DEV, True
)
preview = console_recovery.build_recovery_preview(
playbook_id=console_recovery.PLAYBOOK_REBIND_SESSION,
target="branches/feat-issue-644",
principal=principal,
)
self.assertFalse(preview["execution_enabled"])
self.assertEqual(
preview["execution_blocked_reason"], console_authz.DENY_PHASE_NOT_ACTIVE
)
self.assertFalse(preview["execution_authorization"]["allowed"])
# The preview (non-execution) decision still allows, by role.
self.assertTrue(preview["authorization"]["allowed"])
def _phase_two_enabled(self):
"""Raise ACTIVE_PHASE so the execution branches are reachable in tests."""
return patch.object(console_authz, "ACTIVE_PHASE", 2)
def _operator(self) -> console_authz.Principal:
return console_authz.Principal(
"[email protected]", console_authz.OPERATOR, console_authz.IDENTITY_LOCAL_DEV, True
)
def test_rebind_mutates_the_live_environment_not_a_copy(self) -> None:
"""B2: the playbook must change the mapping it claims to have changed."""
live_env = {stale_binding_recovery.ACTIVE_WORKTREE_ENV: "branches/stale-old"}
phrase = console_recovery.confirmation_phrase(
console_recovery.PLAYBOOK_REBIND_SESSION, "branches/feat-issue-644"
)
with self._phase_two_enabled():
result = console_recovery.execute_recovery_playbook(
playbook_id=console_recovery.PLAYBOOK_REBIND_SESSION,
confirmation=phrase,
target="branches/feat-issue-644",
principal=self._operator(),
env=live_env,
)
self.assertTrue(result["success"])
self.assertEqual(
live_env[stale_binding_recovery.ACTIVE_WORKTREE_ENV],
"branches/feat-issue-644",
"rebind reported success without changing the caller's environment",
)
self.assertTrue(result["applied_result"]["binding_changed"])
self.assertEqual(result["applied_result"]["binding_before"], "branches/stale-old")
self.assertEqual(
result["applied_result"]["binding_after"], "branches/feat-issue-644"
)
def test_clear_stale_binding_reports_failure_when_nothing_changed(self) -> None:
"""B2: a no-op recovery must never be reported as success."""
live_env: dict[str, str] = {}
phrase = console_recovery.confirmation_phrase(
console_recovery.PLAYBOOK_CLEAR_STALE_BINDING
)
with self._phase_two_enabled():
result = console_recovery.execute_recovery_playbook(
playbook_id=console_recovery.PLAYBOOK_CLEAR_STALE_BINDING,
confirmation=phrase,
principal=self._operator(),
env=live_env,
)
self.assertFalse(
result["success"],
"a clear that changed no binding must not report success",
)
self.assertFalse(result["applied_result"]["binding_changed"])
def test_clear_stale_binding_clears_the_live_binding(self) -> None:
"""B2: the sanctioned clear must reach the caller's environment."""
missing = "/nonexistent/branches/deleted-worktree"
live_env = {stale_binding_recovery.ACTIVE_WORKTREE_ENV: missing}
phrase = console_recovery.confirmation_phrase(
console_recovery.PLAYBOOK_CLEAR_STALE_BINDING
)
with self._phase_two_enabled():
result = console_recovery.execute_recovery_playbook(
playbook_id=console_recovery.PLAYBOOK_CLEAR_STALE_BINDING,
confirmation=phrase,
principal=self._operator(),
env=live_env,
)
if result["success"]:
self.assertNotIn(stale_binding_recovery.ACTIVE_WORKTREE_ENV, live_env)
self.assertEqual(result["applied_result"]["binding_before"], missing)
self.assertIsNone(result["applied_result"]["binding_after"])
else:
# Fail closed is acceptable; reporting a clear that did not happen
# is not. This is the invariant the blocker was about.
self.assertFalse(result["applied_result"]["binding_changed"])
self.assertEqual(
live_env.get(stale_binding_recovery.ACTIVE_WORKTREE_ENV), missing
)
def test_reconcile_playbook_calls_an_entry_point_that_exists(self) -> None:
"""B3: the previous call named a function absent from the module."""
phrase = console_recovery.confirmation_phrase(
console_recovery.PLAYBOOK_RECONCILE_CLEANUPS
)
fake_server = types.SimpleNamespace(
gitea_reconcile_merged_cleanups=lambda **kwargs: {
"success": True,
"entries": [{"issue_number": 100}],
}
)
with self._phase_two_enabled(), patch.dict(
sys.modules, {"gitea_mcp_server": fake_server}
):
result = console_recovery.execute_recovery_playbook(
playbook_id=console_recovery.PLAYBOOK_RECONCILE_CLEANUPS,
confirmation=phrase,
principal=console_authz.Principal(
"[email protected]",
console_authz.ADMIN,
console_authz.IDENTITY_LOCAL_DEV,
True,
),
)
self.assertTrue(result["success"], result.get("applied_result"))
self.assertNotIn("error_type", result["applied_result"])
self.assertEqual(result["applied_result"]["reconciled_count"], 1)
def test_reconcile_entry_point_exists_on_the_real_module(self) -> None:
"""B3 regression: guard the symbol itself, not just the call shape."""
import gitea_mcp_server
self.assertTrue(
hasattr(gitea_mcp_server, "gitea_reconcile_merged_cleanups"),
"console recovery depends on this reconciler entry point",
)
self.assertFalse(
hasattr(merged_cleanup_reconcile, "reconcile_merged_cleanups"),
"if this module grows the orchestrator, point the playbook back at it",
)
def test_contamination_gate_blocks_a_writing_playbook(self) -> None:
"""B4: a live marker plus a gated task key must actually block."""
marker = {
"kind": "manual_daemon_kill",
"reason_class": "manual_daemon_kill",
"command_summary": "pkill -f gitea_mcp_server",
"active": True,
}
phrase = console_recovery.confirmation_phrase(
console_recovery.PLAYBOOK_REBIND_SESSION, "branches/feat-issue-644"
)
live_env = {stale_binding_recovery.ACTIVE_WORKTREE_ENV: "branches/stale-old"}
with self._phase_two_enabled(), patch.object(
console_recovery, "load_active_contamination_marker", return_value=marker
):
result = console_recovery.execute_recovery_playbook(
playbook_id=console_recovery.PLAYBOOK_REBIND_SESSION,
confirmation=phrase,
target="branches/feat-issue-644",
principal=self._operator(),
env=live_env,
)
self.assertFalse(result["success"])
self.assertEqual(result["error"], "contaminated_runtime")
self.assertEqual(
live_env[stale_binding_recovery.ACTIVE_WORKTREE_ENV],
"branches/stale-old",
"a blocked playbook must not have mutated anything",
)
def test_contamination_gate_exempts_the_reconciler_remedy(self) -> None:
"""B4: the designated remedy must stay reachable while contaminated."""
marker = {
"kind": "manual_daemon_kill",
"reason_class": "manual_daemon_kill",
"command_summary": "pkill -f gitea_mcp_server",
"active": True,
}
phrase = console_recovery.confirmation_phrase(
console_recovery.PLAYBOOK_RECONCILE_CLEANUPS
)
fake_server = types.SimpleNamespace(
gitea_reconcile_merged_cleanups=lambda **kwargs: {
"success": True,
"entries": [],
}
)
with self._phase_two_enabled(), patch.object(
console_recovery, "load_active_contamination_marker", return_value=marker
), patch.dict(sys.modules, {"gitea_mcp_server": fake_server}):
result = console_recovery.execute_recovery_playbook(
playbook_id=console_recovery.PLAYBOOK_RECONCILE_CLEANUPS,
confirmation=phrase,
principal=console_authz.Principal(
"[email protected]",
console_authz.ADMIN,
console_authz.IDENTITY_LOCAL_DEV,
True,
),
)
self.assertNotEqual(result.get("error"), "contaminated_runtime")
def test_gated_task_key_is_actually_gated(self) -> None:
"""B4: the console action id was never a member of the gated set."""
self.assertIn(
console_recovery.CONTAMINATION_GATED_TASK,
stable_branch_push_guard.CONTAMINATION_GATED_TASKS,
)
self.assertNotIn(
console_recovery.ACTION_CLEAR_STALE_BINDING,
stable_branch_push_guard.CONTAMINATION_GATED_TASKS,
)
def test_diagnosis_reads_the_key_the_gate_returns(self) -> None:
"""B4: ``contaminated`` is a key assess_contamination_gate never returns."""
gate = runtime_recovery_guard.assess_contamination_gate(
None, task=console_recovery.CONTAMINATION_GATED_TASK, actual_role="operator"
)
self.assertNotIn("contaminated", gate)
self.assertIn("block", gate)
def test_contaminated_runtime_is_reported_unclean(self) -> None:
"""B4: verify_post_recovery reported contamination_clean unconditionally."""
marker = {
"kind": "manual_daemon_kill",
"reason_class": "manual_daemon_kill",
"command_summary": "pkill -f gitea_mcp_server",
"active": True,
}
with patch.object(
console_recovery, "load_active_contamination_marker", return_value=marker
):
verification = console_recovery.verify_post_recovery()
diag = console_recovery.diagnose_recovery()
self.assertFalse(verification["contamination_clean"])
self.assertFalse(verification["clean"])
self.assertEqual(diag.status, console_recovery.STATUS_BLOCKED_CONTAMINATION)
def test_master_parity_baseline_is_not_the_head_it_is_compared_against(self) -> None:
"""B5: capture_startup_parity was fed the head it was then compared to."""
stale = system_health.StaleRuntime(
daemon_head="a" * 40,
checkout_head="b" * 40,
remote_head="b" * 40,
stale=True,
determinable=True,
mutation_safe=False,
reasons=("daemon is behind the checkout",),
)
with patch.object(system_health, "assess_stale_runtime", return_value=stale):
diag = console_recovery.diagnose_recovery()
parity = diag.master_parity
self.assertEqual(parity["startup_head"], "a" * 40)
self.assertEqual(parity["current_head"], "b" * 40)
self.assertNotEqual(parity["startup_head"], parity["current_head"])
self.assertFalse(parity["in_parity"])
def test_master_parity_carries_the_live_remote_dimension(self) -> None:
"""B5: live_remote_head was never passed, dropping the #610 dimension."""
stale = system_health.StaleRuntime(
daemon_head="c" * 40,
checkout_head="c" * 40,
remote_head="d" * 40,
stale=False,
determinable=True,
mutation_safe=False,
reasons=(),
)
with patch.object(system_health, "assess_stale_runtime", return_value=stale):
diag = console_recovery.diagnose_recovery()
self.assertEqual(diag.master_parity.get("live_remote_head"), "d" * 40)
def test_verify_post_recovery(self) -> None:
verification = console_recovery.verify_post_recovery()
self.assertIn("clean", verification)
self.assertIn("status", verification)
self.assertIn("reasons", verification)
def test_unverified_inherited_binding_is_not_reported_clean(self) -> None:
"""B2: ``not clear_eligible`` also read clean for unproven bindings."""
binding = {
"classification": stale_binding_recovery.CLASSIFICATION_UNVERIFIED_INHERITED,
"clear_eligible": False,
}
diag = console_recovery.diagnose_recovery()
patched = console_recovery.RecoveryDiagnosis(
status=diag.status,
clean=diag.clean,
stale_runtime=diag.stale_runtime,
master_parity=diag.master_parity,
stale_binding=binding,
contamination=diag.contamination,
worktree_anomalies=diag.worktree_anomalies,
playbooks=diag.playbooks,
reasons=diag.reasons,
)
with patch.object(console_recovery, "diagnose_recovery", return_value=patched):
verification = console_recovery.verify_post_recovery()
self.assertFalse(verification["binding_clean"])
self.assertEqual(
verification["binding_classification"],
stale_binding_recovery.CLASSIFICATION_UNVERIFIED_INHERITED,
)
class TestConsoleRecoveryApi(unittest.TestCase):
def setUp(self) -> None:
self.app = create_app()
self.client = TestClient(self.app)
def test_api_recovery_diagnose(self) -> None:
res = self.client.get("/api/v1/system/recovery/diagnose")
self.assertEqual(res.status_code, 200)
data = res.json()
self.assertIn("status", data)
self.assertIn("clean", data)
self.assertIn("playbooks", data)
self.assertTrue(len(data["playbooks"]) >= 4)
def test_api_recovery_preview(self) -> None:
res = self.client.post(
"/api/v1/system/recovery/preview",
json={"playbook_id": "clear_stale_binding", "target": "active"},
)
self.assertEqual(res.status_code, 200)
data = res.json()
self.assertEqual(data["playbook_id"], "clear_stale_binding")
self.assertEqual(data["confirmation_phrase"], "confirm clear_stale_binding active")
self.assertIn("mutation_ledger", data)
def test_api_recovery_apply_denied_without_auth(self) -> None:
res = self.client.post(
"/api/v1/system/recovery/apply",
json={"playbook_id": "clear_stale_binding", "confirmation": "confirm clear_stale_binding"},
)
self.assertEqual(res.status_code, 400)
data = res.json()
self.assertFalse(data["success"])
self.assertFalse(data["allowed"])
def test_api_recovery_apply_refuses_phase_two_write_with_dev_auth(self) -> None:
"""B1: this previously asserted the phase-gate bypass as intended.
An authenticated operator posting a valid confirmation still must not
execute a phase-2 write while the console is in phase 1. The refusal is
the contract; a 200 here means the gate is not armed.
"""
env = {
"WEBUI_AUTH_MODE": "local_dev",
"WEBUI_DEV_SUBJECT": "[email protected]",
"WEBUI_DEV_ROLE": "operator",
}
before = os.environ.get("GITEA_ACTIVE_WORKTREE")
with patch.dict(os.environ, env):
res = self.client.post(
"/api/v1/system/recovery/apply",
json={
"playbook_id": "rebind_session_worktree",
"target": "branches/feat-issue-644",
"confirmation": "confirm rebind_session_worktree branches/feat-issue-644",
},
)
self.assertEqual(res.status_code, 400)
data = res.json()
self.assertFalse(data["success"])
self.assertFalse(data["allowed"])
self.assertEqual(data["error"], console_authz.DENY_PHASE_NOT_ACTIVE)
self.assertEqual(
os.environ.get("GITEA_ACTIVE_WORKTREE"),
before,
"a refused apply must not have rebound the live process environment",
)
def test_api_recovery_preview_reports_why_execution_is_disabled(self) -> None:
res = self.client.post(
"/api/v1/system/recovery/preview",
json={"playbook_id": "rebind_session_worktree", "target": "active"},
)
self.assertEqual(res.status_code, 200)
data = res.json()
self.assertFalse(data["execution_enabled"])
self.assertIn("execution_authorization", data)
def test_api_recovery_verify(self) -> None:
res = self.client.get("/api/v1/system/recovery/verify")
self.assertEqual(res.status_code, 200)
data = res.json()
self.assertIn("clean", data)
self.assertIn("status", data)
if __name__ == "__main__":
unittest.main()
+511
View File
@@ -0,0 +1,511 @@
"""Tests for the Gitea issue↔PR linkage console (#645, Phase 3).
Covers the acceptance criteria of the issue:
* AC1 issuePR linkage is visible for the selected project/repo, in both
directions, with the evidence that produced each edge.
* AC2 the latest canonical handoff (CTH) is summarized for a focused thread.
* AC3 an external Gitea link appears only under the admin reveal opt-in.
* AC4 every case is driven by mocked Gitea payloads; no network.
Plus the invariants this console must not violate: a partial or failed read is
never rendered as "no link exists", an unfetched thread is never rendered as
"no handoff", redaction happens before display, and the surface stays read-only.
"""
from __future__ import annotations
import json
import os
import sys
import unittest
from pathlib import Path
from unittest import mock
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
from tests.webui_testclient import TestClient
from canonical_thread_handoff import format_cth_body
from webui.app import create_app
from webui.linkage_loader import (
EVIDENCE_BRANCH,
EVIDENCE_CLOSES,
EVIDENCE_REFERENCE,
HANDOFF_LOADED,
HANDOFF_NOT_LOADED,
HANDOFF_UNAVAILABLE,
LinkageSnapshot,
load_linkage_snapshot,
resolve_linkage,
resolve_pr_links,
snapshot_to_dict,
summarize_handoff,
)
from webui.linkage_views import render_linkage_page
from webui.nav import nav_hrefs
from webui.queue_loader import PaginationMeta
def _pagination(*, complete: bool = True, count: int = 0) -> PaginationMeta:
return PaginationMeta(
page=1,
per_page=50,
returned_count=count,
has_more=not complete,
is_final_page=complete,
inventory_complete=complete,
pages_fetched=1,
)
def _pr(
number: int,
*,
title: str = "",
body: str = "",
head: str = "",
state: str = "open",
labels: tuple[str, ...] = (),
) -> dict:
return {
"number": number,
"title": title or f"pr {number}",
"body": body,
"state": state,
"head": {"ref": head},
"labels": [{"name": name} for name in labels],
}
def _issue(
number: int,
*,
title: str = "",
state: str = "open",
labels: tuple[str, ...] = (),
) -> dict:
return {
"number": number,
"title": title or f"issue {number}",
"state": state,
"labels": [{"name": name} for name in labels],
}
def _fetcher(items: list[dict], *, complete: bool = True):
def _fetch(*_args, **_kwargs):
return items, _pagination(complete=complete, count=len(items))
return _fetch
def _load(
issues: list[dict],
prs: list[dict],
*,
complete: bool = True,
**kwargs,
) -> LinkageSnapshot:
return load_linkage_snapshot(
fetch_prs=_fetcher(prs, complete=complete),
fetch_issues=_fetcher(issues, complete=complete),
**kwargs,
)
def _cth(comment_id: int, *, created_at: str, status: str, next_owner: str) -> dict:
return {
"id": comment_id,
"created_at": created_at,
"user": {"login": "jcwalker3"},
"body": format_cth_body(
cth_type="Author Handoff",
status=status,
next_owner=next_owner,
current_blocker="none",
decision="implemented",
proof="full suite green",
next_action="review PR",
ready_to_paste_prompt="Review PR #902 now.",
),
}
class TestLinkageEvidence(unittest.TestCase):
"""AC1 — every edge records how it was found, and keeps all candidates."""
def test_closes_keyword_in_body_is_strongest_evidence(self):
links = resolve_pr_links(_pr(902, body="Closes #643"))
self.assertEqual([link.issue_number for link in links], [643])
self.assertEqual(links[0].evidence, (EVIDENCE_CLOSES,))
self.assertTrue(links[0].closes)
def test_closes_keyword_in_title_counts(self):
links = resolve_pr_links(_pr(902, title="feat(webui): preview (Closes #643)"))
self.assertEqual(links[0].evidence, (EVIDENCE_CLOSES,))
def test_canonical_branch_marker_links_without_a_keyword(self):
links = resolve_pr_links(_pr(902, head="feat/issue-643-request-preview"))
self.assertEqual([link.issue_number for link in links], [643])
self.assertEqual(links[0].evidence, (EVIDENCE_BRANCH,))
self.assertFalse(links[0].closes)
def test_non_canonical_branch_is_not_treated_as_a_marker(self):
self.assertEqual(resolve_pr_links(_pr(902, head="issue-643-preview")), ())
def test_bare_mention_is_recorded_as_the_weakest_evidence(self):
links = resolve_pr_links(_pr(902, body="context in #643"))
self.assertEqual(links[0].evidence, (EVIDENCE_REFERENCE,))
self.assertFalse(links[0].closes)
def test_several_evidence_kinds_merge_onto_one_edge(self):
links = resolve_pr_links(
_pr(902, body="Closes #643 — see #643", head="feat/issue-643-preview")
)
self.assertEqual(len(links), 1)
self.assertEqual(
links[0].evidence,
(EVIDENCE_CLOSES, EVIDENCE_BRANCH, EVIDENCE_REFERENCE),
)
def test_stronger_evidence_sorts_first(self):
links = resolve_pr_links(_pr(902, body="Closes #643, related #700"))
self.assertEqual([link.issue_number for link in links], [643, 700])
def test_self_reference_is_not_linkage(self):
links = resolve_pr_links(_pr(902, body="supersedes #902"))
self.assertEqual(links, ())
def test_every_candidate_is_kept_never_collapsed_to_a_guess(self):
links = resolve_pr_links(_pr(902, body="Closes #643\nCloses #644"))
self.assertEqual([link.issue_number for link in links], [643, 644])
class TestLinkageIndex(unittest.TestCase):
def test_issue_direction_ignores_mention_only_edges(self):
index = resolve_linkage([_pr(902, body="context in #643")])
self.assertIsNone(index.issue_prs.get(643))
self.assertEqual(index.pr_links[902][0].evidence, (EVIDENCE_REFERENCE,))
def test_contested_issue_is_reported_when_two_prs_claim_it(self):
index = resolve_linkage(
[_pr(902, body="Closes #643"), _pr(903, head="feat/issue-643-again")]
)
self.assertEqual(index.contested_issues(), (643,))
self.assertEqual(index.issue_prs[643], (902, 903))
def test_single_claim_is_not_contested(self):
index = resolve_linkage([_pr(902, body="Closes #643")])
self.assertEqual(index.contested_issues(), ())
def test_ambiguous_when_two_issues_tie_at_the_strongest_evidence(self):
index = resolve_linkage([_pr(902, body="Closes #643\nCloses #644")])
self.assertTrue(index.ambiguous(902))
def test_weaker_candidate_alongside_a_stronger_one_is_not_ambiguous(self):
index = resolve_linkage([_pr(902, body="Closes #643, see #700")])
self.assertFalse(index.ambiguous(902))
self.assertEqual(index.primary_issue(902).issue_number, 643)
def test_malformed_pr_row_is_skipped_not_raised_on(self):
index = resolve_linkage([{"title": "no number"}, _pr(902, body="Closes #643")])
self.assertEqual(sorted(index.pr_links), [902])
class TestLinkageSnapshot(unittest.TestCase):
"""AC1 — linkage is visible per project/repo, in both directions."""
def test_both_directions_are_populated(self):
snapshot = _load([_issue(643)], [_pr(902, body="Closes #643")])
self.assertTrue(snapshot.ok)
self.assertEqual([node.number for node in snapshot.issues], [643])
self.assertEqual(snapshot.issues[0].linked_prs, (902,))
self.assertEqual(snapshot.prs[0].links[0].issue_number, 643)
def test_repo_scope_comes_from_the_registry_project(self):
snapshot = _load([], [])
self.assertIn("/", snapshot.repo_label)
self.assertTrue(snapshot.project_id)
def test_unknown_project_fails_closed_with_a_reason(self):
snapshot = _load([_issue(643)], [], project_id="no-such-project")
self.assertFalse(snapshot.ok)
self.assertIn("not found in registry", snapshot.fetch_error)
self.assertEqual(snapshot.issues, ())
def test_orphan_pr_is_identifiable(self):
snapshot = _load([], [_pr(902), _pr(903, body="Closes #643")])
self.assertEqual([node.number for node in snapshot.orphan_prs], [902])
def test_state_scope_defaults_to_open_and_is_reported(self):
self.assertEqual(_load([], []).state_scope, "open")
self.assertEqual(_load([], [], state="all").state_scope, "all")
def test_unsupported_state_falls_back_to_open(self):
self.assertEqual(_load([], [], state="../etc").state_scope, "open")
def test_state_is_passed_through_to_the_fetchers(self):
seen: list[str] = []
def _fetch(*_args, **kwargs):
seen.append(kwargs.get("state", ""))
return [], _pagination()
load_linkage_snapshot(state="all", fetch_prs=_fetch, fetch_issues=_fetch)
self.assertEqual(seen, ["all", "all"])
class TestPartialInventoryIsNotAnAbsenceClaim(unittest.TestCase):
"""An empty edge list from a partial read must never read as 'no link'."""
def test_incomplete_pagination_marks_links_non_authoritative(self):
snapshot = _load([_issue(643)], [], complete=False)
self.assertFalse(snapshot.inventory_complete)
self.assertFalse(snapshot.issues[0].links_authoritative)
def test_complete_pagination_marks_links_authoritative(self):
snapshot = _load([_issue(643)], [], complete=True)
self.assertTrue(snapshot.inventory_complete)
self.assertTrue(snapshot.issues[0].links_authoritative)
def test_partial_window_renders_a_qualified_empty_cell(self):
html = render_linkage_page(_load([_issue(643)], [], complete=False))
self.assertIn("none found (partial inventory)", html)
def test_complete_window_renders_a_plain_none(self):
html = render_linkage_page(_load([_issue(643)], [], complete=True))
self.assertNotIn("partial inventory", html)
self.assertIn(">none<", html)
def test_missing_credentials_fail_closed_without_a_table(self):
with mock.patch(
"webui.linkage_loader._offline_test_mode", return_value=False
), mock.patch("webui.linkage_loader.get_auth_header", return_value=""):
snapshot = load_linkage_snapshot()
self.assertFalse(snapshot.ok)
self.assertIn("credentials unavailable", snapshot.fetch_error)
html = render_linkage_page(snapshot)
self.assertIn("Linkage unavailable", html)
self.assertNotIn("Issues → pull requests", html)
def test_fetch_failure_is_reported_not_raised(self):
def _boom(*_args, **_kwargs):
raise RuntimeError("gitea 502")
snapshot = load_linkage_snapshot(fetch_prs=_boom, fetch_issues=_boom)
self.assertFalse(snapshot.ok)
self.assertIn("Gitea fetch failed", snapshot.fetch_error)
class TestHandoffSummary(unittest.TestCase):
"""AC2 — the latest canonical handoff is summarized for a focused thread."""
def test_latest_cth_wins(self):
summary = summarize_handoff([
_cth(1, created_at="2026-07-24T10:00:00Z", status="in progress",
next_owner="author"),
_cth(2, created_at="2026-07-25T10:00:00Z", status="PR-open",
next_owner="reviewer"),
])
self.assertEqual(summary.comment_id, 2)
self.assertEqual(summary.status, "PR-open")
self.assertEqual(summary.next_owner, "reviewer")
self.assertTrue(summary.cth_type_known)
def test_thread_without_a_cth_summarizes_to_none(self):
self.assertIsNone(summarize_handoff([{"id": 1, "body": "ordinary comment"}]))
def test_unknown_heading_is_reported_not_republished(self):
summary = summarize_handoff([
{
"id": 5,
"created_at": "2026-07-25T10:00:00Z",
"user": {"login": "someone"},
"body": "<!-- cth:v1 -->\n## CTH: Totally Made Up\n\nStatus: odd\n",
}
])
self.assertFalse(summary.cth_type_known)
self.assertEqual(summary.cth_type, "unrecognized")
self.assertNotIn("Totally Made Up", json.dumps(summary.to_dict()))
def test_focused_pr_loads_its_handoff(self):
snapshot = _load(
[_issue(643)],
[_pr(902, body="Closes #643")],
pr=902,
comment_source=lambda kind, number: [
_cth(2, created_at="2026-07-25T10:00:00Z", status="PR-open",
next_owner="reviewer")
],
)
self.assertEqual(snapshot.handoff_status.state, HANDOFF_LOADED)
self.assertEqual(snapshot.focus, ("pr", 902))
self.assertEqual(snapshot.prs[0].handoff.status, "PR-open")
def test_unfocused_rows_report_not_loaded_never_none(self):
snapshot = _load(
[_issue(643)],
[_pr(902, body="Closes #643"), _pr(903)],
pr=902,
comment_source=lambda kind, number: [],
)
other = next(node for node in snapshot.prs if node.number == 903)
self.assertIsNone(other.handoff)
self.assertEqual(other.handoff_status.state, HANDOFF_NOT_LOADED)
self.assertIn("not loaded", render_linkage_page(snapshot))
def test_no_focus_means_no_thread_is_claimed_handoff_free(self):
snapshot = _load([_issue(643)], [])
self.assertEqual(snapshot.handoff_status.state, HANDOFF_NOT_LOADED)
self.assertIn("thread-scoped", snapshot.handoff_status.reason)
def test_comment_source_failure_degrades_only_the_handoff(self):
def _boom(_kind, _number):
raise RuntimeError("comments 500")
snapshot = _load(
[_issue(643)], [_pr(902, body="Closes #643")], pr=902, comment_source=_boom
)
self.assertTrue(snapshot.ok)
self.assertEqual(snapshot.handoff_status.state, HANDOFF_UNAVAILABLE)
self.assertEqual(snapshot.issues[0].linked_prs, (902,))
self.assertIn("unavailable", render_linkage_page(snapshot))
def test_loaded_thread_with_no_cth_says_so_explicitly(self):
snapshot = _load(
[_issue(643)],
[_pr(902, body="Closes #643")],
pr=902,
comment_source=lambda kind, number: [{"id": 1, "body": "hi"}],
)
self.assertIn(
"no Canonical Thread Handoff comment found", render_linkage_page(snapshot)
)
class TestDeepLinks(unittest.TestCase):
"""AC3 — an external Gitea link is emitted only when permitted."""
def test_deep_links_are_withheld_by_default(self):
with mock.patch.dict(os.environ, {"GITEA_MCP_REVEAL_ENDPOINTS": ""}):
snapshot = _load([_issue(643)], [])
html = render_linkage_page(snapshot)
self.assertFalse(snapshot.deep_links_enabled)
self.assertIsNone(snapshot.issues[0].deep_link)
self.assertIn("Gitea deep links are withheld", html)
def test_reveal_opt_in_emits_the_link(self):
with mock.patch.dict(os.environ, {"GITEA_MCP_REVEAL_ENDPOINTS": "1"}):
snapshot = _load([_issue(643)], [_pr(902, body="Closes #643")])
html = render_linkage_page(snapshot)
self.assertTrue(snapshot.deep_links_enabled)
self.assertIn("/issues/643", snapshot.issues[0].deep_link)
self.assertIn("/pulls/902", snapshot.prs[0].deep_link)
self.assertIn(f'href="{snapshot.issues[0].deep_link}"', html)
class TestRedactionBoundary(unittest.TestCase):
def test_secret_shaped_title_is_redacted_before_display(self):
snapshot = _load(
[_issue(643, title="token=ghp_thisisnotarealsecretvalue0001")], []
)
payload = json.dumps(snapshot_to_dict(snapshot))
self.assertNotIn("ghp_thisisnotarealsecretvalue0001", payload)
self.assertNotIn(
"ghp_thisisnotarealsecretvalue0001", render_linkage_page(snapshot)
)
def test_handoff_fields_are_redacted(self):
comment = _cth(
2, created_at="2026-07-25T10:00:00Z", status="ok", next_owner="reviewer"
)
comment["body"] += "\nDecision: password=hunter2hunter2\n"
snapshot = _load(
[_issue(643)],
[_pr(902, body="Closes #643")],
pr=902,
comment_source=lambda kind, number: [comment],
)
self.assertNotIn("hunter2hunter2", json.dumps(snapshot_to_dict(snapshot)))
self.assertNotIn("hunter2hunter2", render_linkage_page(snapshot))
def test_html_escapes_markup_in_a_title(self):
snapshot = _load([_issue(643, title="<script>alert(1)</script>")], [])
html = render_linkage_page(snapshot)
self.assertNotIn("<script>alert(1)</script>", html)
self.assertIn("&lt;script&gt;", html)
class TestLinkageRoutes(unittest.TestCase):
def setUp(self):
self.snapshot = _load(
[_issue(643, labels=("status:ready",))],
[_pr(902, body="Closes #643", labels=("status:pr-open",))],
)
self.client = TestClient(create_app())
def test_page_renders_both_tables(self):
with mock.patch("webui.app.load_linkage_snapshot", return_value=self.snapshot):
response = self.client.get("/gitea")
self.assertEqual(response.status_code, 200)
self.assertIn("Issues → pull requests", response.text)
self.assertIn("Pull requests → issues", response.text)
self.assertIn("#643", response.text)
def test_api_exports_the_same_model(self):
with mock.patch("webui.app.load_linkage_snapshot", return_value=self.snapshot):
response = self.client.get("/api/v1/gitea/linkage")
self.assertEqual(response.status_code, 200)
payload = response.json()
self.assertTrue(payload["ok"])
self.assertEqual(payload["issues"][0]["linked_prs"], [902])
self.assertEqual(payload["prs"][0]["links"][0]["issue_number"], 643)
self.assertEqual(payload["schema_version"], 1)
def test_api_declares_the_evidence_vocabulary(self):
with mock.patch("webui.app.load_linkage_snapshot", return_value=self.snapshot):
payload = self.client.get("/api/v1/gitea/linkage").json()
names = {entry["name"] for entry in payload["evidence_kinds"]}
self.assertEqual(names, {EVIDENCE_CLOSES, EVIDENCE_BRANCH, EVIDENCE_REFERENCE})
def test_api_fails_closed_with_a_non_200_when_the_read_failed(self):
failed = _load([], [], project_id="no-such-project")
with mock.patch("webui.app.load_linkage_snapshot", return_value=failed):
response = self.client.get("/api/v1/gitea/linkage")
self.assertEqual(response.status_code, 502)
self.assertFalse(response.json()["ok"])
def test_page_still_renders_when_the_read_failed(self):
failed = _load([], [], project_id="no-such-project")
with mock.patch("webui.app.load_linkage_snapshot", return_value=failed):
response = self.client.get("/gitea")
self.assertEqual(response.status_code, 200)
self.assertIn("Linkage unavailable", response.text)
def test_query_parameters_reach_the_loader(self):
with mock.patch(
"webui.app.load_linkage_snapshot", return_value=self.snapshot
) as loader:
self.client.get("/gitea?project=gitea-tools&state=all&pr=902")
loader.assert_called_once()
args, kwargs = loader.call_args
self.assertEqual(args[0], "gitea-tools")
self.assertEqual(kwargs["state"], "all")
self.assertEqual(kwargs["pr"], 902)
self.assertIsNone(kwargs["issue"])
def test_surface_stays_read_only(self):
for path in ("/gitea", "/api/v1/gitea/linkage"):
with self.subTest(path=path):
self.assertEqual(self.client.post(path).status_code, 405)
def test_nav_exposes_the_linkage_page_as_live(self):
self.assertIn("/gitea", nav_hrefs())
home = self.client.get("/").text
self.assertIn('href="/gitea"', home)
self.assertIn(">Gitea<", home)
if __name__ == "__main__":
unittest.main()
+465
View File
@@ -0,0 +1,465 @@
"""Unit tests for Phase 3 Notifications and Human-Attention Console (#648)."""
from __future__ import annotations
import pytest
from starlette.testclient import TestClient
from webui.app import create_app
from webui.notifications import (
ATTENTION_HUMAN_REQUIRED,
ATTENTION_OPERATOR,
ATTENTION_ROUTINE,
CATEGORY_AUTH,
CATEGORY_BLOCKER,
CATEGORY_LEASE,
CATEGORY_SYSTEM,
CATEGORY_VALIDATION,
CATEGORY_WORKFLOW,
NotificationItem,
NotificationSnapshot,
classify_attention_event,
load_notifications_snapshot,
snapshot_to_dict,
)
from webui.notification_views import render_notifications_page
from webui.project_registry import load_registry
from webui.queue_loader import QueueItem, QueueSnapshot
from webui.lease_loader import CollisionWarning, LeaseSnapshot
from webui.system_health import DependencyProbe, SystemHealthSnapshot, VersionInfo, StaleRuntime
def test_classify_attention_event_rules():
# 1. Critical escalation boundaries -> human-required
att_cls, req_human = classify_attention_event(
CATEGORY_AUTH, "Auth error", "Unauthorized access attempt", is_auth_failure=True
)
assert att_cls == ATTENTION_HUMAN_REQUIRED
assert req_human is True
att_cls, req_human = classify_attention_event(
CATEGORY_SYSTEM, "Hard stop", "Hard stop triggered", is_hard_stop=True
)
assert att_cls == ATTENTION_HUMAN_REQUIRED
assert req_human is True
att_cls, req_human = classify_attention_event(
CATEGORY_VALIDATION, "Validation Error", "Report validation failed", is_validation_failure=True
)
assert att_cls == ATTENTION_HUMAN_REQUIRED
assert req_human is True
# 2. Operational issues -> operator
att_cls, req_human = classify_attention_event(
CATEGORY_BLOCKER, "PR Blocked", "Merge conflict detected", is_blocker=True
)
assert att_cls == ATTENTION_OPERATOR
assert req_human is False
att_cls, req_human = classify_attention_event(
CATEGORY_LEASE, "Lease Expired", "Session lease expired", is_stale=True
)
assert att_cls == ATTENTION_OPERATOR
assert req_human is False
# 3. Routine workflow transitions -> routine
att_cls, req_human = classify_attention_event(
CATEGORY_WORKFLOW, "PR Active", "PR in review"
)
assert att_cls == ATTENTION_ROUTINE
assert req_human is False
def test_notification_snapshot_aggregation():
reg = load_registry()
proj_id = reg.projects[0].id if reg.projects else "gitea-tools"
mock_queue = QueueSnapshot(
project_id=proj_id,
repo_label="org/repo",
prs=(
QueueItem(
number=101,
title="Blocked PR",
badges=("blocked",),
extra={},
),
QueueItem(
number=102,
title="Normal PR",
badges=("in-review",),
extra={},
),
),
issues=(),
pr_pagination=None,
issue_pagination=None,
)
mock_leases = LeaseSnapshot(
project_id=proj_id,
repo_label="org/repo",
issue_lock=None,
claim_inventory={},
reviewer_leases=(
{
"pr_number": 101,
"status": "expired",
"is_expired": True,
},
),
duplicate_prs=(
CollisionWarning(
kind="duplicate_pr",
message="Multiple open PRs for issue #101",
issue_number=101,
pr_numbers=(101, 103),
),
),
duplicate_branches=(),
collision_history=(),
fetch_error=None,
)
mock_version = VersionInfo(
git_sha="abc1234",
git_describe="v1.0.0",
control_plane_schema_version=1,
python_version="3.11",
known=True,
)
mock_stale = StaleRuntime(
daemon_head="abc1234",
checkout_head="abc1234",
remote_head="abc1234",
stale=False,
determinable=True,
mutation_safe=True,
reasons=(),
)
mock_health = SystemHealthSnapshot(
status="degraded",
ready=False,
readiness_complete=True,
readiness_reasons=("Auth failure",),
service="webui",
mode="test",
version=mock_version,
started_at="2026-07-25T00:00:00Z",
uptime_seconds=100.0,
timestamp="2026-07-25T00:00:00Z",
deep_probes_requested=True,
dependencies=(
DependencyProbe(
name="auth_service",
kind="auth",
status="unauthorized",
detail="Token expired",
required=True,
),
),
mcp_namespaces=(),
stale_runtime=mock_stale,
probe_errors=(),
)
snapshot = load_notifications_snapshot(
proj_id,
load_queue=lambda _id: mock_queue,
load_leases=lambda **_kwargs: mock_leases,
load_health=lambda **_kwargs: mock_health,
)
assert snapshot.project_id == proj_id
assert snapshot.total_count == 5
assert snapshot.human_required_count >= 1 # auth probe failure
assert snapshot.operator_count >= 3 # blocked PR + expired lease + duplicate PR collision
assert snapshot.routine_count >= 1 # normal PR
# Inbox items should include operator and human-required items only
inbox_classes = {item.attention_class for item in snapshot.inbox_items}
assert ATTENTION_ROUTINE not in inbox_classes
assert ATTENTION_OPERATOR in inbox_classes
assert ATTENTION_HUMAN_REQUIRED in inbox_classes
def test_snapshot_to_dict_and_redaction():
item = NotificationItem(
id="notif-1",
attention_class=ATTENTION_HUMAN_REQUIRED,
category=CATEGORY_AUTH,
title="Auth Error",
summary="Failed auth header: Bearer secret_token_12345",
work_kind="system",
work_number=None,
project_id="test-proj",
repo_label="org/repo",
created_at="2026-07-25T16:00:00Z",
requires_human=True,
)
snap = NotificationSnapshot(
project_id="test-proj",
repo_label="org/repo",
items=(item,),
human_required_count=1,
operator_count=0,
routine_count=0,
total_count=1,
)
data = snapshot_to_dict(snap)
assert data["project_id"] == "test-proj"
assert data["human_required_count"] == 1
assert len(data["inbox_items"]) == 1
# Redaction test
summary = data["inbox_items"][0]["summary"]
assert "secret_token_12345" not in summary
assert "<redacted>" in summary or "Bearer" in summary
def test_notifications_html_views():
item = NotificationItem(
id="notif-1",
attention_class=ATTENTION_HUMAN_REQUIRED,
category=CATEGORY_AUTH,
title="Critical Auth Failure",
summary="Auth failure details",
work_kind="issue",
work_number=42,
project_id="test-proj",
repo_label="org/repo",
created_at="2026-07-25T16:00:00Z",
requires_human=True,
)
snap = NotificationSnapshot(
project_id="test-proj",
repo_label="org/repo",
items=(item,),
human_required_count=1,
operator_count=0,
routine_count=0,
total_count=1,
)
html = render_notifications_page(snap, filter_class="inbox")
assert "Notifications &amp; Attention Inbox" in html or "Notifications & Attention Inbox" in html
assert "Critical Auth Failure" in html
assert "HUMAN REQUIRED" in html
assert "Human Required" in html
def test_notifications_app_routes():
app = create_app()
client = TestClient(app)
# 1. HTML Route
res = client.get("/notifications")
assert res.status_code == 200
assert "Notifications" in res.text
assert "Attention Inbox" in res.text
# 2. API Route /api/v1/notifications
res_api = client.get("/api/v1/notifications")
assert res_api.status_code == 200
json_data = res_api.json()
assert "human_required_count" in json_data
assert "operator_count" in json_data
assert "routine_count" in json_data
assert "inbox_items" in json_data
# 3. Compatibility Alias /api/notifications
res_alias = client.get("/api/notifications")
assert res_alias.status_code == 200
assert res_alias.json()["project_id"] == json_data["project_id"]
def test_classify_ignores_human_authored_title_and_summary_keywords():
"""B1: keywords in human-authored titles must not escalate routine work (#905)."""
# Routine transition whose title/summary mention critical-boundary words
att_cls, req_human = classify_attention_event(
CATEGORY_WORKFLOW,
"record irrecoverable decision lock provenance",
"PR #999 'record irrecoverable decision lock provenance' is in routine state in-review.",
)
assert att_cls == ATTENTION_ROUTINE
assert req_human is False
att_cls, req_human = classify_attention_event(
CATEGORY_WORKFLOW,
"fix unauthorized token path",
"Issue #1 'fix unauthorized token path' state: claimed. hard stop docs only.",
)
assert att_cls == ATTENTION_ROUTINE
assert req_human is False
# Structured flags still escalate (machine-driven)
att_cls, req_human = classify_attention_event(
CATEGORY_SYSTEM,
"anything",
"anything with hard stop in text",
is_hard_stop=True,
)
assert att_cls == ATTENTION_HUMAN_REQUIRED
assert req_human is True
def test_notification_ids_are_unique_across_probe_errors_and_collisions():
"""B2: published notification ids must be unique within a snapshot (#905)."""
reg = load_registry()
proj_id = reg.projects[0].id if reg.projects else "gitea-tools"
mock_queue = QueueSnapshot(
project_id=proj_id,
repo_label="org/repo",
prs=(),
issues=(),
pr_pagination=None,
issue_pagination=None,
)
mock_leases = LeaseSnapshot(
project_id=proj_id,
repo_label="org/repo",
issue_lock=None,
claim_inventory={},
reviewer_leases=(),
duplicate_prs=(
CollisionWarning(
kind="duplicate_pr",
message="Multiple open PRs for issue #10",
issue_number=10,
pr_numbers=(10, 11),
),
CollisionWarning(
kind="duplicate_branch",
message="Another collision without issue",
issue_number=None,
pr_numbers=(12, 13),
),
CollisionWarning(
kind="duplicate_pr",
message="Second issue collision",
issue_number=10,
pr_numbers=(14, 15),
),
),
duplicate_branches=(),
collision_history=(),
fetch_error=None,
)
mock_version = VersionInfo(
git_sha="abc1234",
git_describe="v1.0.0",
control_plane_schema_version=1,
python_version="3.11",
known=True,
)
mock_stale = StaleRuntime(
daemon_head="abc1234",
checkout_head="abc1234",
remote_head="abc1234",
stale=False,
determinable=True,
mutation_safe=True,
reasons=(),
)
mock_health = SystemHealthSnapshot(
status="degraded",
ready=False,
readiness_complete=True,
readiness_reasons=(),
service="webui",
mode="test",
version=mock_version,
started_at="2026-07-25T00:00:00Z",
uptime_seconds=100.0,
timestamp="2026-07-25T00:00:00Z",
deep_probes_requested=True,
dependencies=(),
mcp_namespaces=(),
stale_runtime=mock_stale,
probe_errors=("error alpha", "error beta"),
)
snapshot = load_notifications_snapshot(
proj_id,
load_queue=lambda _id: mock_queue,
load_leases=lambda **_kwargs: mock_leases,
load_health=lambda **_kwargs: mock_health,
)
ids = [item.id for item in snapshot.items]
assert len(ids) == len(set(ids)), f"duplicate notification ids: {ids}"
assert any(i.startswith(f"notif-sys-err-{proj_id}-") for i in ids)
assert any(i.startswith("notif-collision-") for i in ids)
def test_probe_errors_do_not_set_fetch_error():
"""B3: probe_errors must not be reported as fetch_error (#905)."""
reg = load_registry()
proj_id = reg.projects[0].id if reg.projects else "gitea-tools"
mock_queue = QueueSnapshot(
project_id=proj_id,
repo_label="org/repo",
prs=(),
issues=(),
pr_pagination=None,
issue_pagination=None,
fetch_error=None,
)
mock_leases = LeaseSnapshot(
project_id=proj_id,
repo_label="org/repo",
issue_lock=None,
claim_inventory={},
reviewer_leases=(),
duplicate_prs=(),
duplicate_branches=(),
collision_history=(),
fetch_error=None,
)
mock_version = VersionInfo(
git_sha="abc1234",
git_describe="v1.0.0",
control_plane_schema_version=1,
python_version="3.11",
known=True,
)
mock_stale = StaleRuntime(
daemon_head="abc1234",
checkout_head="abc1234",
remote_head="abc1234",
stale=False,
determinable=True,
mutation_safe=True,
reasons=(),
)
mock_health = SystemHealthSnapshot(
status="degraded",
ready=False,
readiness_complete=True,
readiness_reasons=(),
service="webui",
mode="test",
version=mock_version,
started_at="2026-07-25T00:00:00Z",
uptime_seconds=100.0,
timestamp="2026-07-25T00:00:00Z",
deep_probes_requested=True,
dependencies=(),
mcp_namespaces=(),
stale_runtime=mock_stale,
probe_errors=("probe blew up",),
)
snapshot = load_notifications_snapshot(
proj_id,
load_queue=lambda _id: mock_queue,
load_leases=lambda **_kwargs: mock_leases,
load_health=lambda **_kwargs: mock_health,
)
assert snapshot.fetch_error is None
# probe errors still appear as items
assert any("probe blew up" in item.summary for item in snapshot.items)
File diff suppressed because it is too large Load Diff
+265
View File
@@ -56,6 +56,11 @@ from webui.session_loader import (
snapshot_to_dict as session_view_snapshot_to_dict,
)
from webui.session_views import render_sessions_page
from webui.linkage_loader import (
load_linkage_snapshot,
snapshot_to_dict as linkage_snapshot_to_dict,
)
from webui.linkage_views import render_linkage_page
from webui.inventory import (
SECTION_NAMES as _INVENTORY_SECTIONS,
load_inventory_snapshot,
@@ -75,6 +80,13 @@ from webui.system_health import (
snapshot_to_dict as system_health_to_dict,
)
from webui.system_health_views import render_system_health_page
from webui.notifications import (
load_notifications_snapshot,
snapshot_to_dict as notifications_snapshot_to_dict,
)
from webui.notification_views import render_notifications_page
from webui import request_service
from webui.request_views import render_requests_page
_READ_ONLY_METHODS = frozenset({"GET", "HEAD", "OPTIONS"})
_AUDIT_MUTATION_PATHS = frozenset({"/audit", "/api/audit"})
@@ -203,6 +215,84 @@ async def system_health(request: Request) -> HTMLResponse:
)
async def api_recovery_diagnose(_request: Request) -> JSONResponse:
from webui.console_recovery import diagnose_recovery
diag = diagnose_recovery()
return JSONResponse({
"status": diag.status,
"clean": diag.clean,
"stale_runtime": diag.stale_runtime,
"master_parity": diag.master_parity,
"stale_binding": diag.stale_binding,
"contamination": diag.contamination,
"worktree_anomalies": list(diag.worktree_anomalies),
"reasons": list(diag.reasons),
"playbooks": [
{
"playbook_id": pb.playbook_id,
"label": pb.label,
"action_id": pb.action_id,
"description": pb.description,
"eligible": pb.eligible,
"requires_confirmation": pb.requires_confirmation,
"reason": pb.reason,
"params_schema": pb.params_schema,
}
for pb in diag.playbooks
],
})
async def api_recovery_preview(request: Request) -> JSONResponse:
from webui.console_recovery import build_recovery_preview
body = {}
try:
body = await request.json()
except Exception:
pass
playbook_id = body.get("playbook_id") or request.query_params.get("playbook_id") or ""
target = body.get("target") or request.query_params.get("target")
principal = resolve_principal(request.headers)
preview = build_recovery_preview(
playbook_id=playbook_id,
target=target,
params=body,
principal=principal,
)
status = 200 if preview.get("playbook_id") else 400
return JSONResponse(preview, status_code=status)
async def api_recovery_apply(request: Request) -> JSONResponse:
from webui.console_recovery import execute_recovery_playbook
body = {}
try:
body = await request.json()
except Exception:
pass
playbook_id = body.get("playbook_id", "")
confirmation = body.get("confirmation", "")
target = body.get("target")
principal = resolve_principal(request.headers)
request_id = getattr(request.state, "request_id", None)
result = execute_recovery_playbook(
playbook_id=playbook_id,
confirmation=confirmation,
target=target,
params=body,
principal=principal,
request_id=request_id,
)
status_code = 200 if result.get("success") else 400
return JSONResponse(result, status_code=status_code)
async def api_recovery_verify(_request: Request) -> JSONResponse:
from webui.console_recovery import verify_post_recovery
verification = verify_post_recovery()
return JSONResponse(verification, status_code=200)
async def queue(_request: Request) -> HTMLResponse:
snapshot = load_queue_snapshot()
return HTMLResponse(render_page(title="Queue", body_html=render_queue_page(snapshot)))
@@ -372,6 +462,41 @@ async def api_sessions(_request: Request) -> JSONResponse:
return JSONResponse(session_view_snapshot_to_dict(load_session_view_snapshot()))
def _linkage_snapshot(request: Request):
"""Load one linkage snapshot from the request's scope and focus parameters."""
return load_linkage_snapshot(
request.query_params.get("project") or None,
state=request.query_params.get("state"),
issue=_query_int(request, "issue"),
pr=_query_int(request, "pr"),
)
async def gitea_linkage(request: Request) -> HTMLResponse:
"""Gitea issue↔PR linkage console (#645) — read-only.
Always 200, including on a failed read: this is an operator view, and it
must render *why* linkage could not be loaded rather than withhold the page.
The snapshot itself carries ``ok=False`` and the page refuses to draw a
linkage table it cannot stand behind.
"""
return HTMLResponse(render_linkage_page(_linkage_snapshot(request)))
async def api_v1_gitea_linkage(request: Request) -> JSONResponse:
"""JSON export of the issue↔PR linkage model (#645).
Unlike the HTML view, the API answers with 502 when the snapshot could not
be loaded, so an automated consumer cannot read a fail-closed payload as a
successful "no links exist" result.
"""
snapshot = _linkage_snapshot(request)
return JSONResponse(
linkage_snapshot_to_dict(snapshot),
status_code=200 if snapshot.ok else 502,
)
async def _parse_audit_form(request: Request) -> tuple[str, str | None]:
if request.method == "GET":
return "", None
@@ -769,6 +894,125 @@ async def api_v1_analytics_ingest(request: Request) -> JSONResponse:
)
async def notifications_route(request: Request) -> HTMLResponse:
project_id = request.query_params.get("project_id")
attention_class = request.query_params.get("attention_class") or "inbox"
snap = load_notifications_snapshot(project_id)
html = render_notifications_page(
snap, filter_class=attention_class, filter_project=project_id
)
return HTMLResponse(html)
async def api_notifications(request: Request) -> JSONResponse:
project_id = request.query_params.get("project_id")
snap = load_notifications_snapshot(project_id)
data = notifications_snapshot_to_dict(snap)
return JSONResponse(data)
def _default_request_scope() -> dict[str, str]:
"""Resolve remote/org/repo from the project registry for request forms.
Returns an empty mapping when the registry cannot be read, which makes
``parse_request`` reject a request that did not name its own scope rather
than letting it default to some other repository.
"""
from webui.queue_loader import _host_from_url # host normalisation helper
registry, error = _load_project_registry()
if error is not None or not registry.projects:
return {}
project = registry.projects[0]
host = _host_from_url(project.remote_host)
return {
"remote": _derive_remote(host),
"org": project.gitea_owner or "",
"repo": project.repo_name or "",
}
async def _request_payload(request: Request) -> dict[str, object]:
"""Read a request body as JSON or form-encoded. Never raises."""
content_type = (request.headers.get("content-type") or "").lower()
if "application/json" in content_type:
try:
body = await request.json()
except Exception:
return {}
return dict(body) if isinstance(body, dict) else {}
try:
form = await request.form()
except Exception:
return {}
return {key: form[key] for key in form}
async def requests_page(request: Request) -> HTMLResponse:
"""Operator request form and intent preview (#643).
POST here only ever *previews*. Initiation is a separate confirmed call to
``/api/v1/requests/apply`` so that submitting this form cannot reserve
work as a side effect.
"""
submitted: dict[str, object] = {}
preview = None
error = None
if request.method == "POST":
submitted = await _request_payload(request)
work_request, error = request_service.parse_request(
submitted, default_scope=_default_request_scope()
)
if work_request is not None:
preview = request_service.preview_request(
work_request,
principal=resolve_principal(headers=dict(request.headers)),
)
return HTMLResponse(
render_requests_page(
preview=preview, error=error, submitted=submitted
)
)
async def api_v1_request_preview(request: Request) -> JSONResponse:
"""Dry-run authorization and intent preview for a work request (#643)."""
payload = await _request_payload(request)
work_request, error = request_service.parse_request(
payload, default_scope=_default_request_scope()
)
if work_request is None:
return JSONResponse(error.to_dict(), status_code=400)
preview = request_service.preview_request(
work_request,
principal=resolve_principal(headers=dict(request.headers)),
)
return JSONResponse(
preview.to_dict(), status_code=200 if preview.authorized else 403
)
async def api_v1_request_apply(request: Request) -> JSONResponse:
"""Initiate a previewed work request through the allocator (#643).
Fail-closed at every step: unauthorized, unconfirmed, not-next-safe, and
already-claimed all return without attempting an assignment.
"""
payload = await _request_payload(request)
work_request, error = request_service.parse_request(
payload, default_scope=_default_request_scope()
)
if work_request is None:
return JSONResponse(error.to_dict(), status_code=400)
confirm = _truthy_flag(str(payload.get("confirm") or ""))
result = request_service.apply_request(
work_request,
principal=resolve_principal(headers=dict(request.headers)),
confirm=confirm,
)
status = int(result.pop("status_code", 403))
return JSONResponse(result, status_code=status)
async def method_not_allowed(request: Request, _exc: Exception) -> Response:
path = request.url.path
if path in _AUDIT_MUTATION_PATHS and request.method == "POST":
@@ -797,6 +1041,9 @@ def create_app(*, bind_host: str | None = None) -> Starlette:
Route("/api/queue", api_queue, methods=["GET"]),
Route("/traffic", traffic, methods=["GET"]),
Route("/api/traffic", api_traffic, methods=["GET"]),
Route("/notifications", notifications_route, methods=["GET"]),
Route("/api/notifications", api_notifications, methods=["GET"]),
Route("/api/v1/notifications", api_notifications, methods=["GET"]),
Route("/projects", projects, methods=["GET"]),
Route("/projects/{project_id}", project_detail, methods=["GET"]),
Route("/api/projects", api_projects, methods=["GET"]),
@@ -822,6 +1069,8 @@ def create_app(*, bind_host: str | None = None) -> Starlette:
Route("/api/sessions", api_sessions, methods=["GET"]),
Route("/api/v1/sessions", api_sessions, methods=["GET"]),
Route("/api/v1/timeline", api_v1_timeline, methods=["GET"]),
Route("/gitea", gitea_linkage, methods=["GET"]),
Route("/api/v1/gitea/linkage", api_v1_gitea_linkage, methods=["GET"]),
Route("/analytics", analytics, methods=["GET"]),
Route("/api/analytics", api_v1_analytics, methods=["GET"]),
Route("/api/v1/analytics", api_v1_analytics, methods=["GET"]),
@@ -843,6 +1092,17 @@ def create_app(*, bind_host: str | None = None) -> Starlette:
api_action_attempt,
methods=["POST"],
),
Route("/requests", requests_page, methods=["GET", "POST"]),
Route(
"/api/v1/requests/preview",
api_v1_request_preview,
methods=["POST"],
),
Route(
"/api/v1/requests/apply",
api_v1_request_apply,
methods=["POST"],
),
Route("/api/leases", api_leases, methods=["GET"]),
Route("/api/v1/inventory", api_inventory, methods=["GET"]),
Route(
@@ -855,6 +1115,11 @@ def create_app(*, bind_host: str | None = None) -> Starlette:
api_console_security_model,
methods=["GET"],
),
# #644 Phase 2 Recovery API routes
Route("/api/v1/system/recovery/diagnose", api_recovery_diagnose, methods=["GET"]),
Route("/api/v1/system/recovery/preview", api_recovery_preview, methods=["POST", "GET"]),
Route("/api/v1/system/recovery/apply", api_recovery_apply, methods=["POST"]),
Route("/api/v1/system/recovery/verify", api_recovery_verify, methods=["POST", "GET"]),
*[
Route(path, phase_stub, methods=["GET"])
for path in STUB_PAGES
+105 -8
View File
@@ -115,6 +115,12 @@ class ConsoleAction:
break_glass: bool
phase: int
summary: str
# Opt-in switch for an action whose execution path is genuinely wired
# ahead of its phase becoming globally active (#643). Naming a variable
# here enables nothing on its own: the variable must also be set in the
# environment. An action that leaves this ``None`` can only execute once
# ACTIVE_PHASE reaches its phase, exactly as before.
execution_env_flag: str | None = None
@property
def mcp_permission(self) -> str:
@@ -277,6 +283,61 @@ _ACTION_SPECS: tuple[ConsoleAction, ...] = (
phase=2,
summary="Restart one MCP namespace via the host supervisor.",
),
# #644: Phase 2 recovery controls & playbooks.
ConsoleAction(
action_id="system.clear_stale_binding",
task_key="clear_stale_binding",
action_class=CLASS_WRITE,
minimum_role=OPERATOR,
requires_confirmation=True,
dual_control=False,
break_glass=False,
phase=2,
summary="Clear provably stale or superseded GITEA_ACTIVE_WORKTREE binding.",
),
ConsoleAction(
action_id="system.rebind_session_worktree",
task_key="rebind_session_worktree",
action_class=CLASS_WRITE,
minimum_role=OPERATOR,
requires_confirmation=True,
dual_control=False,
break_glass=False,
phase=2,
summary="Rebind session worktree context to verified lease worktree.",
),
ConsoleAction(
action_id="system.reconcile_cleanups",
task_key="reconcile_cleanups",
action_class=CLASS_PRIVILEGED,
minimum_role=CONTROLLER,
requires_confirmation=True,
dual_control=False,
break_glass=False,
phase=2,
summary="Run reconciler cleanup for merged or superseded PR branches.",
),
# #643: submit a work request — desired role, issue/PR, intent — and let
# the allocator reserve it. This is the one Phase 2 action whose execution
# path is actually implemented (``webui.request_service``), so it carries
# the opt-in flag; it stays denied until an operator sets that variable.
# Authority is operator-class because the outcome is a claim, not a Gitea
# verdict: initiating reviewer or merger *work* does not grant the right
# to approve or merge, which stays with the MCP role profile.
ConsoleAction(
action_id="initiate_workflow",
task_key="allocate_next_work",
action_class=CLASS_WRITE,
minimum_role=OPERATOR,
requires_confirmation=True,
dual_control=False,
break_glass=False,
phase=2,
summary=(
"Preview and initiate allocator-owned workflow work for a role."
),
execution_env_flag="WEBUI_REQUESTS_EXECUTION",
),
)
ACTIONS: dict[str, ConsoleAction] = {a.action_id: a for a in _ACTION_SPECS}
@@ -430,6 +491,33 @@ ALLOW_PREVIEW = "allowed_preview_only"
# gated on this model landing; nothing here enables it.
ACTIVE_PHASE = 1
_TRUTHY = frozenset({"1", "true", "yes", "on"})
def execution_wired(
action: ConsoleAction | None, env: dict[str, str] | None = None
) -> bool:
"""Whether *action* has a live execution path right now.
Two ways to be wired, and only two. The action's phase is active, or the
action declares an opt-in environment variable *and* that variable is set.
Everything else including every action that never declares a flag is
unwired, so the default across the registry stays deny.
Bumping ``ACTIVE_PHASE`` would enable execution for every action of that
phase at once. The per-action flag exists so a single implemented action
can go live without dragging its unimplemented phase-mates with it.
"""
if action is None:
return False
if action.phase <= ACTIVE_PHASE:
return True
flag = (action.execution_env_flag or "").strip()
if not flag:
return False
source = env if env is not None else os.environ
return (source.get(flag) or "").strip().lower() in _TRUTHY
@dataclass(frozen=True)
class AuthorizationDecision:
@@ -469,16 +557,19 @@ def authorize(
principal: Principal | None = None,
*,
for_execution: bool = False,
env: dict[str, str] | None = None,
) -> AuthorizationDecision:
"""Decide whether *principal* may invoke *action_id*. Deny by default.
``for_execution`` distinguishes a read-only preview from a real invocation.
Even an allowed decision reports ``execution_enabled=False`` while the
console is in Phase 1, so no caller can read an allow as permission to
mutate.
``execution_enabled`` reports whether the action has a live execution path
at all (:func:`execution_wired`) for every action without an explicit
opt-in flag that stays ``False`` while the console is in Phase 1, so no
caller can read an allow as permission to mutate.
"""
who = principal if principal is not None else ANONYMOUS
action = get_action(action_id)
wired = execution_wired(action, env)
if action is None:
return AuthorizationDecision(
@@ -497,7 +588,7 @@ def authorize(
"requires_confirmation": action.requires_confirmation,
"dual_control": action.dual_control,
"break_glass": action.break_glass,
"execution_enabled": False,
"execution_enabled": wired,
}
if not who.authenticated:
@@ -530,13 +621,19 @@ def authorize(
**base,
)
if for_execution and action.phase > ACTIVE_PHASE:
if for_execution and not wired:
return AuthorizationDecision(
allowed=False,
reason_code=DENY_PHASE_NOT_ACTIVE,
detail=(
f"Action {action_id!r} belongs to phase {action.phase}; the "
f"console is in phase {ACTIVE_PHASE}. Execution is not wired."
f"console is in phase {ACTIVE_PHASE}"
+ (
f" and {action.execution_env_flag} is not set"
if action.execution_env_flag
else ""
)
+ ". Execution is not wired."
),
**base,
)
@@ -545,8 +642,8 @@ def authorize(
allowed=True,
reason_code=ALLOW_PREVIEW,
detail=(
"Principal holds the required role. Preview only — execution "
"remains disabled until the Phase 2 action framework ships."
"Principal holds the required role. Execution proceeds only for an "
"action with a wired execution path; everything else is preview."
),
**base,
)
+721
View File
@@ -0,0 +1,721 @@
"""Web Console Phase 2 Recovery Controls & Playbooks (#644).
Provides canonical recovery controls for the web console:
1. Diagnosis: Surfaces stale runtimes, worktree binding errors, contamination markers,
and un-reconciled cleanups.
2. Gated Actions & Playbooks: Guided recovery (rebind session worktree, clear stale
binding, trigger reconciler cleanups, sanctioned restart).
3. RBAC, Contamination (#630), and Master Parity (#610) integration.
4. Audit trail via ``console_audit`` and mandatory post-recovery revalidation.
"""
from __future__ import annotations
import os
from dataclasses import asdict, dataclass, field
from pathlib import Path
from typing import Any
import master_parity_gate
import runtime_recovery_guard
import stale_binding_recovery
from webui import console_audit, console_authz, sanctioned_restart, system_health, worktree_scanner
# --- Recovery Statuses ------------------------------------------------------
STATUS_HEALTHY = "healthy"
STATUS_ACTION_REQUIRED = "action_required"
STATUS_BLOCKED_CONTAMINATION = "blocked_contamination"
STATUS_RECONNECT_REQUIRED = "reconnect_required"
# --- Playbook Identifiers ---------------------------------------------------
PLAYBOOK_CLEAR_STALE_BINDING = "clear_stale_binding"
PLAYBOOK_REBIND_SESSION = "rebind_session_worktree"
PLAYBOOK_RECONCILE_CLEANUPS = "reconcile_cleanups"
PLAYBOOK_SANCTIONED_RESTART = "sanctioned_restart"
KNOWN_PLAYBOOKS: tuple[str, ...] = (
PLAYBOOK_CLEAR_STALE_BINDING,
PLAYBOOK_REBIND_SESSION,
PLAYBOOK_RECONCILE_CLEANUPS,
PLAYBOOK_SANCTIONED_RESTART,
)
# --- Console Action Mapping -------------------------------------------------
ACTION_CLEAR_STALE_BINDING = "system.clear_stale_binding"
ACTION_REBIND_SESSION = "system.rebind_session_worktree"
ACTION_RECONCILE_CLEANUPS = "system.reconcile_cleanups"
PLAYBOOK_ACTIONS: dict[str, str] = {
PLAYBOOK_CLEAR_STALE_BINDING: ACTION_CLEAR_STALE_BINDING,
PLAYBOOK_REBIND_SESSION: ACTION_REBIND_SESSION,
PLAYBOOK_RECONCILE_CLEANUPS: ACTION_RECONCILE_CLEANUPS,
PLAYBOOK_SANCTIONED_RESTART: sanctioned_restart.ACTION_RESTART_NAMESPACE,
}
#: Task key handed to :func:`runtime_recovery_guard.assess_contamination_gate`.
#: A console *action id* is not a task name and is not a member of
#: ``CONTAMINATION_GATED_TASKS``, so passing one left the #630 gate inert. Every
#: writing recovery playbook shares this one gated task key; the reconciler
#: cleanup playbook is exempted separately because it is the designated remedy.
CONTAMINATION_GATED_TASK = "console_recovery_apply"
#: Remote whose contamination markers govern this console. Markers are written
#: per remote, so reading the wrong one reports a contaminated runtime clean.
REMOTE_ENV = "WEBUI_GITEA_REMOTE"
DEFAULT_REMOTE = "prgs"
def _console_remote(env: dict[str, str] | None = None) -> str:
env_map = env if env is not None else os.environ
return (env_map.get(REMOTE_ENV) or "").strip() or DEFAULT_REMOTE
def load_active_contamination_marker(
remote: str | None = None, env: dict[str, str] | None = None
) -> dict[str, Any] | None:
"""Return the live #630 contamination marker payload, or ``None``.
The gate is only meaningful when it is fed a real marker: with
``marker=None`` :func:`assess_contamination_gate` returns ``block: False``
on its first statement. The #641 session inventory already reads the durable
markers, so reuse that reader rather than adding a second source of truth.
Never raises into a diagnosis or execution path.
"""
try:
from webui import session_loader
except Exception: # noqa: BLE001 — never break recovery on an import problem
return None
try:
markers = session_loader._load_contamination_markers(
remote=remote or _console_remote(env)
)
except Exception: # noqa: BLE001 — fail soft; the caller degrades to no marker
return None
for marker in markers:
payload = marker.to_dict()
if payload.get("active"):
return payload
return None
def _active_binding(env_map: Any) -> str | None:
"""Read the live worktree binding so a no-op recovery cannot report success."""
value = env_map.get(stale_binding_recovery.ACTIVE_WORKTREE_ENV)
return value if value else None
@dataclass(frozen=True)
class RecoveryLedgerEntry:
"""One planned recovery step displayed before execution."""
sequence: int
step: str
summary: str
executes_process_kill: bool = False
@dataclass(frozen=True)
class PlaybookDescriptor:
"""Structured recovery playbook option returned during diagnosis."""
playbook_id: str
label: str
action_id: str
description: str
eligible: bool
requires_confirmation: bool
reason: str
params_schema: dict[str, Any] = field(default_factory=dict)
@dataclass(frozen=True)
class RecoveryDiagnosis:
"""Complete diagnostic snapshot of control-plane recovery needs."""
status: str
clean: bool
stale_runtime: dict[str, Any]
master_parity: dict[str, Any]
stale_binding: dict[str, Any]
contamination: dict[str, Any]
worktree_anomalies: tuple[str, ...]
playbooks: tuple[PlaybookDescriptor, ...]
reasons: tuple[str, ...]
def _repo_root(custom_path: Path | str | None = None) -> Path:
if custom_path:
return Path(custom_path).resolve()
override = (os.environ.get("WEBUI_REPO_ROOT") or "").strip()
if override:
return Path(override).resolve()
return Path(__file__).resolve().parent.parent
def confirmation_phrase(playbook_id: str, target: str | None = None) -> str:
"""Construct exact confirmation phrase required for a recovery playbook."""
clean_target = (target or "").strip()
if clean_target:
return f"confirm {playbook_id} {clean_target}"
return f"confirm {playbook_id}"
def confirmation_matches(
playbook_id: str, confirmation: str | None, target: str | None = None
) -> bool:
expected = confirmation_phrase(playbook_id, target)
return str(confirmation or "").strip() == expected
def _build_ledger(
playbook_id: str, target: str | None = None
) -> tuple[RecoveryLedgerEntry, ...]:
if playbook_id == PLAYBOOK_CLEAR_STALE_BINDING:
return (
RecoveryLedgerEntry(1, "quiesce", "Stop admitting new gated mutations."),
RecoveryLedgerEntry(
2,
"clear_env",
f"Remove stale env binding GITEA_ACTIVE_WORKTREE ({target or 'active'}).",
),
RecoveryLedgerEntry(
3, "audit", "Record clear_stale_binding event in console audit log."
),
RecoveryLedgerEntry(
4, "revalidate", "Re-run diagnosis to verify clean binding state."
),
)
if playbook_id == PLAYBOOK_REBIND_SESSION:
return (
RecoveryLedgerEntry(1, "quiesce", "Stop admitting new gated mutations."),
RecoveryLedgerEntry(
2,
"rebind_worktree",
f"Rebind session worktree context safely to {target or 'target worktree'}.",
),
RecoveryLedgerEntry(
3, "audit", "Record rebind_session_worktree event in console audit log."
),
RecoveryLedgerEntry(
4, "revalidate", "Re-run diagnosis to verify worktree binding state."
),
)
if playbook_id == PLAYBOOK_RECONCILE_CLEANUPS:
return (
RecoveryLedgerEntry(1, "quiesce", "Stop admitting new gated mutations."),
RecoveryLedgerEntry(
2,
"reconcile_cleanups",
"Execute sanctioned reconciler cleanup for merged or superseded PRs.",
),
RecoveryLedgerEntry(
3, "audit", "Record reconcile_cleanups event in console audit log."
),
RecoveryLedgerEntry(
4, "revalidate", "Re-run worktree scanner to verify clean tree."
),
)
if playbook_id == PLAYBOOK_SANCTIONED_RESTART:
restart_ledger = sanctioned_restart._mutation_ledger(
target or "gitea-author", sanctioned_restart.MODE_RESTART
)
return tuple(
RecoveryLedgerEntry(
sequence=e.sequence,
step=e.step,
summary=e.summary,
executes_process_kill=e.executes_process_kill,
)
for e in restart_ledger
)
return (
RecoveryLedgerEntry(1, "unspecified", f"Execute recovery playbook {playbook_id}."),
)
def diagnose_recovery(
repo_path: Path | str | None = None,
env: dict[str, str] | None = None,
active_worktree_val: str | None = None,
session_lease_wt: str | None = None,
role_kind: str | None = None,
) -> RecoveryDiagnosis:
"""Run full control-plane diagnostics to determine recovery needs and options."""
root = _repo_root(repo_path)
source_env = dict(env) if env is not None else dict(os.environ)
reasons: list[str] = []
# 1. Stale runtime assessment
stale_runtime_obj = system_health.assess_stale_runtime(root)
stale_runtime_dict = {
"daemon_head": stale_runtime_obj.daemon_head,
"checkout_head": stale_runtime_obj.checkout_head,
"remote_head": stale_runtime_obj.remote_head,
"stale": stale_runtime_obj.stale,
"determinable": stale_runtime_obj.determinable,
"mutation_safe": stale_runtime_obj.mutation_safe,
"reasons": list(stale_runtime_obj.reasons),
}
if stale_runtime_obj.stale:
reasons.append("Runtime HEAD disagrees with checkout/remote HEAD.")
# 2. Master parity assessment
#
# The baseline is the commit the *running process* started at, which is what
# the parity gate is about. Capturing it from ``checkout_head`` and then
# comparing it against that same value made ``in_parity`` structurally
# incapable of being false. ``live_remote_head`` restores the #610
# live-remote dimension, which was previously dropped.
checkout_head = stale_runtime_obj.checkout_head
startup_dict = master_parity_gate.capture_startup_parity(
str(root), head=stale_runtime_obj.daemon_head
)
parity_dict = master_parity_gate.assess_master_parity(
startup_dict, checkout_head, stale_runtime_obj.remote_head
)
if not parity_dict.get("in_parity", True):
reasons.append("Repository is not in master parity.")
# 3. Worktree binding classification
boot_bindings = stale_binding_recovery.snapshot_boot_bindings(source_env)
active_val = (
active_worktree_val
if active_worktree_val is not None
else source_env.get(stale_binding_recovery.ACTIVE_WORKTREE_ENV)
)
boot_inherited = bool(boot_bindings.get("active_worktree") and active_val == boot_bindings.get("active_worktree"))
path_exists = None
if active_val:
path_exists = os.path.exists(os.path.realpath(active_val))
binding_class = stale_binding_recovery.classify_active_worktree_binding(
active_value=active_val,
session_lease_worktree=session_lease_wt,
boot_inherited=boot_inherited,
path_exists=path_exists,
role_kind=role_kind,
)
if binding_class.get("clear_eligible"):
reasons.append(
f"Active worktree binding is stale ({binding_class.get('classification')})."
)
elif binding_class.get("classification") == stale_binding_recovery.CLASSIFICATION_UNVERIFIED_INHERITED:
reasons.append("Inherited worktree binding is unverified.")
# 4. Contamination assessment (#630)
#
# A real marker and a task key the gate actually gates on: with marker=None
# the gate short-circuits to ``block: False``, and with a console action id
# the task is outside CONTAMINATION_GATED_TASKS, so it could never block.
contamination_marker = load_active_contamination_marker(env=source_env)
contamination_dict = runtime_recovery_guard.assess_contamination_gate(
contamination_marker,
task=CONTAMINATION_GATED_TASK,
actual_role=role_kind,
)
contaminated = bool(contamination_dict.get("block"))
if contaminated:
reasons.append("Runtime is contaminated by manual process kill (#630).")
# 5. Worktree scanner hygiene & anomalies
hygiene = worktree_scanner.load_hygiene_snapshot(project_root=str(root))
worktree_anomalies = hygiene.anomalies
# Determine status & eligible playbooks
playbooks: list[PlaybookDescriptor] = []
# Playbook 1: Clear Stale Binding
clear_eligible = bool(binding_class.get("clear_eligible"))
playbooks.append(
PlaybookDescriptor(
playbook_id=PLAYBOOK_CLEAR_STALE_BINDING,
label="Clear Stale Worktree Binding",
action_id=ACTION_CLEAR_STALE_BINDING,
description="Clear provably stale or superseded GITEA_ACTIVE_WORKTREE environment binding.",
eligible=clear_eligible,
requires_confirmation=True,
reason=(
f"Binding classified as {binding_class.get('classification')}; clear is authorized."
if clear_eligible
else "Active worktree binding is clean, corroborated, or absent."
),
)
)
# Playbook 2: Rebind Session Worktree
rebind_eligible = bool(
active_val
or binding_class.get("classification") == stale_binding_recovery.CLASSIFICATION_UNVERIFIED_INHERITED
)
playbooks.append(
PlaybookDescriptor(
playbook_id=PLAYBOOK_REBIND_SESSION,
label="Rebind Session Worktree",
action_id=ACTION_REBIND_SESSION,
description="Rebind or synchronize session worktree binding safely with active lease.",
eligible=rebind_eligible,
requires_confirmation=True,
reason=(
"Session worktree binding can be rebound to verified lease worktree."
if rebind_eligible
else "Session worktree is properly bound."
),
params_schema={"target_worktree": "string"},
)
)
# Playbook 3: Reconcile Cleanups
reconcile_eligible = bool(hygiene.anomalies or any(e.classification in {"stale-clean", "detached-review"} for e in hygiene.entries))
playbooks.append(
PlaybookDescriptor(
playbook_id=PLAYBOOK_RECONCILE_CLEANUPS,
label="Trigger Reconciler Cleanups",
action_id=ACTION_RECONCILE_CLEANUPS,
description="Run sanctioned reconciler cleanup preview and apply for merged/superseded PR branches.",
eligible=reconcile_eligible,
requires_confirmation=True,
reason=(
f"Worktree hygiene scanner detected {len(hygiene.anomalies)} anomalies and cleanups needed."
if reconcile_eligible
else "No reconciler cleanups pending."
),
)
)
# Playbook 4: Sanctioned Restart
restart_eligible = bool(stale_runtime_obj.stale or contaminated)
playbooks.append(
PlaybookDescriptor(
playbook_id=PLAYBOOK_SANCTIONED_RESTART,
label="Sanctioned MCP Restart",
action_id=sanctioned_restart.ACTION_RESTART_NAMESPACE,
description="Restart MCP daemon via configured host supervisor without manual process kill.",
eligible=restart_eligible,
requires_confirmation=True,
reason=(
"Stale runtime or contamination detected; host supervisor restart available."
if restart_eligible
else "Runtime is healthy and clean."
),
params_schema={"namespace": "string", "mode": "restart|reload"},
)
)
clean = not reasons and not contaminated
if contaminated:
status = STATUS_BLOCKED_CONTAMINATION
elif reasons:
status = STATUS_ACTION_REQUIRED
else:
status = STATUS_HEALTHY
return RecoveryDiagnosis(
status=status,
clean=clean,
stale_runtime=stale_runtime_dict,
master_parity=parity_dict,
stale_binding=binding_class,
contamination=contamination_dict,
worktree_anomalies=tuple(worktree_anomalies),
playbooks=tuple(playbooks),
reasons=tuple(reasons),
)
def build_recovery_preview(
playbook_id: str,
target: str | None = None,
params: dict[str, Any] | None = None,
principal: console_authz.Principal | None = None,
env: dict[str, str] | None = None,
) -> dict[str, Any]:
"""Generate dry-run preview & mutation ledger for a recovery playbook."""
if playbook_id not in KNOWN_PLAYBOOKS:
return {
"allowed": False,
"error": "unknown_playbook",
"detail": f"Playbook {playbook_id!r} is not a registered recovery playbook.",
}
action_id = PLAYBOOK_ACTIONS[playbook_id]
action = console_authz.get_action(action_id)
decision = console_authz.authorize(action_id, principal)
# Preview and apply must answer the same question. ``execution_enabled`` was
# a hardcoded False beside an authorization decision taken without
# ``for_execution``, so the preview could not tell an operator *why*
# execution was disabled — and the apply path did not ask at all.
execution_decision = console_authz.authorize(
action_id, principal, for_execution=True
)
phrase = confirmation_phrase(playbook_id, target)
ledger = _build_ledger(playbook_id, target)
return {
"playbook_id": playbook_id,
"action_id": action_id,
"target": target,
"required_role": action.minimum_role if action else console_authz.OPERATOR,
"required_permission": action.mcp_permission if action else "gitea.read",
"requires_confirmation": True,
"confirmation_phrase": phrase,
"mutation_ledger": [asdict(entry) for entry in ledger],
"authorization": decision.to_dict(),
"execution_authorization": execution_decision.to_dict(),
"params": dict(params or {}),
"execution_enabled": bool(execution_decision.allowed),
"execution_blocked_reason": (
None if execution_decision.allowed else execution_decision.reason_code
),
}
def execute_recovery_playbook(
playbook_id: str,
confirmation: str | None = None,
target: str | None = None,
params: dict[str, Any] | None = None,
principal: console_authz.Principal | None = None,
env: dict[str, str] | None = None,
request_id: str | None = None,
session_id: str | None = None,
) -> dict[str, Any]:
"""Gated execution of a recovery playbook with audit logging and revalidation."""
if playbook_id not in KNOWN_PLAYBOOKS:
return {
"success": False,
"allowed": False,
"error": "unknown_playbook",
"detail": f"Playbook {playbook_id!r} is not known.",
}
action_id = PLAYBOOK_ACTIONS[playbook_id]
# The mapping the playbooks actually mutate. ``dict(os.environ)`` produced a
# throwaway copy: every env playbook wrote to it, verified against it, and
# left the running daemon bound to the value it claimed to have fixed.
mutation_env: Any = env if env is not None else os.environ
# 1. Authorization check — ``for_execution=True`` is what arms the phase
# gate (console_authz.authorize only applies it in that branch). Without it
# a phase-2 write executed while the console was in phase 1.
decision = console_authz.authorize(action_id, principal, for_execution=True)
if not decision.allowed:
console_audit.record_event(
action_id=action_id,
result=console_audit.RESULT_DENIED,
principal=principal,
target={"playbook_id": playbook_id, "target": target},
reason_code=decision.reason_code,
detail=decision.detail,
request_id=request_id,
session_id=session_id,
)
return {
"success": False,
"allowed": False,
"error": decision.reason_code,
"detail": decision.detail,
}
# 2. Confirmation phrase check
if not confirmation_matches(playbook_id, confirmation, target):
expected = confirmation_phrase(playbook_id, target)
detail = f"Confirmation phrase mismatch. Expected: {expected!r}"
console_audit.record_event(
action_id=action_id,
result=console_audit.RESULT_DENIED,
principal=principal,
target={"playbook_id": playbook_id, "target": target},
reason_code="confirmation_mismatch",
detail=detail,
request_id=request_id,
session_id=session_id,
)
return {
"success": False,
"allowed": False,
"error": "confirmation_mismatch",
"detail": detail,
"expected_confirmation_phrase": expected,
}
# 3. Contamination rule (#630) check
role_str = principal.role if principal else None
contamination_marker = load_active_contamination_marker(env=env)
contam = runtime_recovery_guard.assess_contamination_gate(
contamination_marker,
task=CONTAMINATION_GATED_TASK,
actual_role=role_str,
)
if contam.get("block"):
if playbook_id != PLAYBOOK_RECONCILE_CLEANUPS:
detail = "Runtime is contaminated by a manual process kill (#630). Run reconciler cleanup playbook first."
console_audit.record_event(
action_id=action_id,
result=console_audit.RESULT_DENIED,
principal=principal,
target={"playbook_id": playbook_id, "target": target},
reason_code="contaminated_runtime",
detail=detail,
request_id=request_id,
session_id=session_id,
)
return {
"success": False,
"allowed": False,
"error": "contaminated_runtime",
"detail": detail,
}
# 4. Execute playbook action
applied_result: dict[str, Any] = {"performed": False}
if playbook_id == PLAYBOOK_CLEAR_STALE_BINDING:
binding_before = _active_binding(mutation_env)
diagnosis = diagnose_recovery(env=env)
plan = stale_binding_recovery.plan_recovery(diagnosis.stale_binding)
applied_result = stale_binding_recovery.apply_recovery(plan, env=mutation_env)
binding_after = _active_binding(mutation_env)
applied_result = {
**applied_result,
"binding_before": binding_before,
"binding_after": binding_after,
"binding_changed": binding_before != binding_after,
}
# A clear that did not clear is not a success, whatever the plan said.
if not applied_result["binding_changed"]:
applied_result["performed"] = False
applied_result.setdefault("reasons", []).append(
"clear_stale_binding did not change the live worktree binding"
)
elif playbook_id == PLAYBOOK_REBIND_SESSION:
target_wt = target or (params or {}).get("target_worktree")
if target_wt:
binding_before = _active_binding(mutation_env)
mutation_env[stale_binding_recovery.ACTIVE_WORKTREE_ENV] = target_wt
binding_after = _active_binding(mutation_env)
applied_result = {
"performed": binding_after == target_wt,
"rebound_worktree": target_wt,
"binding_before": binding_before,
"binding_after": binding_after,
"binding_changed": binding_before != binding_after,
"cleared_stale": binding_before != binding_after,
}
if binding_after != target_wt:
applied_result["reasons"] = [
"rebind_session_worktree did not take effect on the live "
"environment"
]
else:
applied_result = {
"performed": False,
"reason": "No target_worktree specified for rebind.",
}
elif playbook_id == PLAYBOOK_RECONCILE_CLEANUPS:
# ``merged_cleanup_reconcile`` exposes the building blocks only; the
# orchestrator is the MCP tool. The previous call named a function that
# does not exist, and a bare ``except`` turned the AttributeError into a
# generic failure, so this playbook could never succeed. Imported lazily
# because the MCP server module is large and binds FastMCP at import.
try:
import gitea_mcp_server
snapshot = gitea_mcp_server.gitea_reconcile_merged_cleanups(
dry_run=False,
execute_confirmed=True,
remote=(params or {}).get("remote") or _console_remote(env),
org=(params or {}).get("org"),
repo=(params or {}).get("repo"),
)
performed_reconcile = bool(snapshot.get("success"))
applied_result = {
"performed": performed_reconcile,
"reconciled_count": len(snapshot.get("entries") or []),
"snapshot": snapshot,
}
if not performed_reconcile:
applied_result["reasons"] = list(snapshot.get("reasons") or [])
except Exception as exc: # noqa: BLE001 — surfaced with its type
applied_result = {
"performed": False,
"error": str(exc),
"error_type": type(exc).__name__,
}
elif playbook_id == PLAYBOOK_SANCTIONED_RESTART:
ns = target or (params or {}).get("namespace", "gitea-author")
md = (params or {}).get("mode", sanctioned_restart.MODE_RESTART)
restart_res = sanctioned_restart.execute_restart(
namespace=ns,
mode=md,
principal=principal,
confirmation=f"{md} {ns}",
# Without the marker the stricter guard at sanctioned_restart.py:375
# never fires and a restart can launder a contaminated runtime.
contamination_marker=contamination_marker,
env=mutation_env,
request_id=request_id,
session_id=session_id,
)
applied_result = restart_res
# ``allowed`` is not ``performed``: execute_restart documents that success is
# False in both directions because the host supervisor still has to act.
performed = bool(applied_result.get("performed"))
# 5. Record Audit Log
audit_record = console_audit.record_event(
action_id=action_id,
result=console_audit.RESULT_ALLOWED if performed else console_audit.RESULT_DENIED,
principal=principal,
target={"playbook_id": playbook_id, "target": target},
reason_code="recovery_executed" if performed else "recovery_failed",
detail=f"Executed recovery playbook {playbook_id}",
request_id=request_id,
session_id=session_id,
metadata={"applied_result": applied_result},
)
# 6. Post-recovery verification recheck.
#
# Re-read state rather than re-reading the mapping the mutation just wrote:
# verifying the mutated copy confirmed changes that never reached the
# process. Passing ``env`` through means a caller-supplied mapping is the
# live one for that caller, and ``None`` re-reads ``os.environ`` fresh.
post_verification = verify_post_recovery(env=env)
return {
"success": performed,
"allowed": True,
"playbook_id": playbook_id,
"action_id": action_id,
"applied_result": applied_result,
"audit": audit_record,
"post_recovery_verification": post_verification,
}
def verify_post_recovery(
repo_path: Path | str | None = None, env: dict[str, str] | None = None
) -> dict[str, Any]:
"""Revalidate control-plane state post-recovery before clean status."""
diag = diagnose_recovery(repo_path, env)
classification = diag.stale_binding.get("classification")
# ``not clear_eligible`` also reads clean for every binding recovery is not
# allowed to touch — an unverified inherited binding is unproven, not clean.
binding_clean = (
not diag.stale_binding.get("clear_eligible")
and classification != stale_binding_recovery.CLASSIFICATION_UNVERIFIED_INHERITED
)
return {
"clean": diag.clean,
"status": diag.status,
"stale_runtime_clean": not diag.stale_runtime.get("stale"),
"binding_clean": binding_clean,
"binding_classification": classification,
# The gate returns ``block``; it has never returned ``contaminated``, so
# reading that key reported every runtime clean unconditionally.
"contamination_clean": not diag.contamination.get("block"),
"anomalies_count": len(diag.worktree_anomalies),
"reasons": list(diag.reasons),
}
+7
View File
@@ -178,6 +178,13 @@ def build_action_registry() -> ActionRegistry:
("system.restart_namespace", "Restart MCP namespace",
"restart_namespace", "host.supervisor_restart",
"Restart one MCP namespace via the host supervisor."),
# #644: Phase 2 recovery playbooks & controls.
("system.clear_stale_binding", "Clear stale binding", "clear_stale_binding",
"console.clear_stale_binding", "Clear provably stale or superseded env binding."),
("system.rebind_session_worktree", "Rebind session worktree", "rebind_session_worktree",
"console.rebind_session_worktree", "Rebind session worktree to verified lease."),
("system.reconcile_cleanups", "Reconcile cleanups", "reconcile_cleanups",
"console.reconcile_cleanups", "Run reconciler cleanup for merged or superseded PRs."),
)
actions = tuple(
GatedAction(
+771
View File
@@ -0,0 +1,771 @@
"""Gitea issue↔PR linkage model for the console (#645, Phase 3).
Operators lose context between an issue and the PR that closes it: which PR
carries which issue, whether two PRs claim the same issue, and what the latest
canonical handoff on that thread said. The evidence exists in Gitea, but only
as free text scattered across PR titles, bodies, and branch names.
This module resolves that linkage into one read-only model:
* :func:`resolve_linkage` is a pure function from raw Gitea issue/PR payloads to
a :class:`LinkageIndex`. It records *how* each edge was found (a ``Closes #N``
keyword, the canonical ``feat/issue-N-`` branch marker, or a bare ``#N``
body reference) and never collapses several candidates into one silent guess.
* :func:`load_linkage_snapshot` scopes that index to a registry project and
optionally attaches the latest Canonical Thread Handoff (CTH) summary for one
focused issue or PR.
Design rules, matching the rest of the console:
- **Read-only.** Gitea is read through the shared authenticated helpers. No
endpoint here mutates anything, and no write action is registered.
- **Qualified absence.** Linkage is a claim about a *loaded* window of Gitea.
When pagination did not complete, when credentials were unavailable, or when
only open items were fetched, the snapshot says so and every "no linked PR"
is marked non-authoritative. An empty edge list from a partial read is not
evidence that no link exists.
- **Handoff is loaded, never assumed.** CTH comments are thread-scoped, so they
are fetched only for an explicitly focused issue or PR. Every other row
reports ``not_loaded`` rather than rendering as "no handoff".
- **Redaction at the boundary.** Titles, labels, handoff fields, and error
reasons are free text from Gitea and cross :mod:`webui.console_redaction`
before they leave this module.
- **Deep links are opt-in.** A link to the Gitea web UI is emitted only under
the ``GITEA_MCP_REVEAL_ENDPOINTS`` admin opt-in, exactly as the MCP tools
gate their own URL exposure.
Non-goals (from the issue): no issue/PR editor, no browser review or merge, no
reimplementation of Gitea search.
"""
from __future__ import annotations
import os
import re
from dataclasses import dataclass
from typing import Any, Callable, Iterable, Sequence
from gitea_auth import api_fetch_page, get_auth_header, gitea_url, repo_api_url
from webui import console_redaction
from webui.project_registry import ProjectRecord, load_registry
from webui.queue_loader import (
PaginationMeta,
_fetch_issues,
_fetch_prs,
_host_from_url,
)
#: Version of the serialized linkage contract. Bump on any breaking change.
LINKAGE_SCHEMA_VERSION = 1
# --- Linkage evidence -------------------------------------------------------
# Ordered strongest to weakest. The strength ordering is what makes an
# ambiguous PR detectable: two candidates at the same strength are a genuine
# ambiguity, while a weaker candidate alongside a stronger one is not.
EVIDENCE_CLOSES = "closes_keyword"
EVIDENCE_BRANCH = "branch_marker"
EVIDENCE_REFERENCE = "body_reference"
EVIDENCE_ORDER: tuple[str, ...] = (
EVIDENCE_CLOSES,
EVIDENCE_BRANCH,
EVIDENCE_REFERENCE,
)
_EVIDENCE_RANK = {name: rank for rank, name in enumerate(EVIDENCE_ORDER)}
EVIDENCE_DESCRIPTIONS: dict[str, str] = {
EVIDENCE_CLOSES: (
"the PR title or body declares 'closes/fixes/resolves #N' — Gitea itself "
"acts on this keyword, so it is the strongest available evidence"
),
EVIDENCE_BRANCH: (
"the PR head branch carries the canonical issue marker "
"'(fix|feat|docs|chore)/issue-N-…' minted by the issue lock"
),
EVIDENCE_REFERENCE: (
"the PR body mentions '#N' without a closing keyword; a mention is not "
"a claim that the PR closes that issue"
),
}
_CLOSES_RE = re.compile(r"(?:closes|fixes|resolves)\s+#(\d+)", re.IGNORECASE)
_REFERENCE_RE = re.compile(r"#(\d+)")
_BRANCH_MARKER_RE = re.compile(
r"^(?:fix|feat|docs|chore)/issue-(\d+)(?:[-/]|$)", re.IGNORECASE
)
# Handoff-source states. ``not_loaded`` is deliberately distinct from "none
# found": a row whose comments were never fetched proves nothing about whether
# a handoff exists on that thread.
HANDOFF_NOT_LOADED = "not_loaded"
HANDOFF_LOADED = "loaded"
HANDOFF_UNAVAILABLE = "unavailable"
# Which item states were fetched. Linkage claims are scoped to this window.
STATE_OPEN = "open"
STATE_ALL = "all"
_SUPPORTED_STATES = (STATE_OPEN, STATE_ALL)
def _redact(value: Any) -> Any:
"""Redact one free-text field, failing closed to the placeholder."""
if value is None:
return None
return console_redaction.redact_text(str(value))
def deep_links_enabled(env: dict[str, str] | None = None) -> bool:
"""Whether Gitea web-UI deep links may be emitted (admin/debug opt-in)."""
source = env if env is not None else os.environ
return (source.get("GITEA_MCP_REVEAL_ENDPOINTS") or "").strip().lower() in {
"1",
"true",
"yes",
"on",
}
def _deep_link(host: str, org: str, repo: str, kind: str, number: int) -> str | None:
"""Build a Gitea web link for one item, or None when reveal is not enabled."""
if not deep_links_enabled() or not (host and org and repo):
return None
segment = "pulls" if kind == "pr" else "issues"
try:
return gitea_url(host, f"/{org}/{repo}/{segment}/{int(number)}")
except Exception:
return None
# --- Pure linkage resolution -------------------------------------------------
@dataclass(frozen=True)
class IssueLink:
"""One resolved edge from a PR to an issue, with the evidence that found it."""
issue_number: int
evidence: tuple[str, ...]
@property
def strength(self) -> int:
"""Rank of the strongest evidence backing this edge (lower is stronger)."""
return min(
(_EVIDENCE_RANK.get(name, len(EVIDENCE_ORDER)) for name in self.evidence),
default=len(EVIDENCE_ORDER),
)
@property
def closes(self) -> bool:
"""True only when the PR *declares* it closes the issue."""
return EVIDENCE_CLOSES in self.evidence
def to_dict(self) -> dict[str, Any]:
return {
"issue_number": self.issue_number,
"evidence": list(self.evidence),
"closes": self.closes,
}
def _sorted_links(links: Iterable[IssueLink]) -> tuple[IssueLink, ...]:
return tuple(sorted(links, key=lambda link: (link.strength, link.issue_number)))
def resolve_pr_links(pr: dict[str, Any]) -> tuple[IssueLink, ...]:
"""Resolve every issue a PR points at, strongest evidence first.
Every candidate is kept. Collapsing to a single "linked issue" is what makes
a mislinked or double-claimed PR invisible, so the caller decides what to do
with several candidates rather than being handed one guess.
"""
found: dict[int, set[str]] = {}
def _add(number: Any, evidence: str) -> None:
try:
issue_number = int(number)
except (TypeError, ValueError):
return
if issue_number <= 0:
return
found.setdefault(issue_number, set()).add(evidence)
title = str(pr.get("title") or "")
body = str(pr.get("body") or "")
for text in (title, body):
for match in _CLOSES_RE.finditer(text):
_add(match.group(1), EVIDENCE_CLOSES)
head_ref = str((pr.get("head") or {}).get("ref") or "")
branch_match = _BRANCH_MARKER_RE.match(head_ref.strip())
if branch_match:
_add(branch_match.group(1), EVIDENCE_BRANCH)
# The ``#N`` inside "Closes #N" is the *same* textual occurrence as the
# closing keyword, not a second, independent mention. Blanking the closing
# phrases first keeps "mention" meaning what the legend says it means: a
# reference the PR made without claiming to close anything.
for match in _REFERENCE_RE.finditer(_CLOSES_RE.sub(" ", body)):
_add(match.group(1), EVIDENCE_REFERENCE)
# A PR's own number appearing in its body is self-reference, not linkage.
try:
found.pop(int(pr.get("number")), None)
except (TypeError, ValueError):
pass
return _sorted_links(
IssueLink(
issue_number=number,
evidence=tuple(name for name in EVIDENCE_ORDER if name in evidence),
)
for number, evidence in found.items()
)
@dataclass(frozen=True)
class LinkageIndex:
"""Resolved linkage over one loaded window of issues and PRs."""
pr_links: dict[int, tuple[IssueLink, ...]]
issue_prs: dict[int, tuple[int, ...]]
def primary_issue(self, pr_number: int) -> IssueLink | None:
"""The strongest edge for a PR, or None when it points at no issue."""
links = self.pr_links.get(int(pr_number)) or ()
return links[0] if links else None
def ambiguous(self, pr_number: int) -> bool:
"""True when two or more issues tie at the PR's strongest evidence."""
links = self.pr_links.get(int(pr_number)) or ()
if len(links) < 2:
return False
best = links[0].strength
return sum(1 for link in links if link.strength == best) > 1
def contested_issues(self) -> tuple[int, ...]:
"""Issues claimed by more than one PR in the loaded window."""
return tuple(
number for number, prs in sorted(self.issue_prs.items()) if len(prs) > 1
)
def resolve_linkage(prs: Sequence[dict[str, Any]]) -> LinkageIndex:
"""Build the bidirectional linkage index for a loaded window of PRs.
Only *closing* and *branch-marker* edges populate the issuePR direction: a
bare ``#N`` mention is a reference, and treating it as "this PR is the work
for issue N" would invent contested issues out of ordinary cross-links. The
weaker edge stays visible on the PRissue side, where it is labelled.
"""
pr_links: dict[int, tuple[IssueLink, ...]] = {}
issue_prs: dict[int, list[int]] = {}
for pr in prs or []:
try:
pr_number = int(pr["number"])
except (KeyError, TypeError, ValueError):
continue
links = resolve_pr_links(pr)
pr_links[pr_number] = links
for link in links:
if link.evidence == (EVIDENCE_REFERENCE,):
continue
bucket = issue_prs.setdefault(link.issue_number, [])
if pr_number not in bucket:
bucket.append(pr_number)
return LinkageIndex(
pr_links=pr_links,
issue_prs={number: tuple(sorted(items)) for number, items in issue_prs.items()},
)
# --- Canonical handoff summary ----------------------------------------------
@dataclass(frozen=True)
class HandoffSummary:
"""The latest CTH comment on one thread, redacted for display."""
comment_id: int | None
created_at: str | None
author: str | None
cth_type: str
cth_type_known: bool
status: str | None
next_owner: str | None
current_blocker: str | None
decision: str | None
next_action: str | None
def to_dict(self) -> dict[str, Any]:
return {
"comment_id": self.comment_id,
"created_at": self.created_at,
"author": self.author,
"cth_type": self.cth_type,
"cth_type_known": self.cth_type_known,
"status": self.status,
"next_owner": self.next_owner,
"current_blocker": self.current_blocker,
"decision": self.decision,
"next_action": self.next_action,
}
@dataclass(frozen=True)
class HandoffStatus:
"""Why a thread's handoff summary is present, absent, or unknown."""
state: str
reason: str | None = None
target: str | None = None
@property
def loaded(self) -> bool:
return self.state == HANDOFF_LOADED
def to_dict(self) -> dict[str, Any]:
return {"state": self.state, "reason": self.reason, "target": self.target}
def summarize_handoff(comments: Sequence[dict[str, Any]]) -> HandoffSummary | None:
"""Summarize the newest CTH comment in *comments*, or None when there is none.
Every field is redacted before it is returned: a handoff body is operator
free text that regularly quotes commands, and it is rendered verbatim on the
page this feeds.
"""
from canonical_thread_handoff import find_latest_cth, is_known_cth_type
try:
latest = find_latest_cth(list(comments or []))
except Exception:
return None
if not latest:
return None
fields = latest.get("fields") or {}
cth_type = str(latest.get("cth_type") or "").strip()
known = is_known_cth_type(cth_type)
try:
comment_id: int | None = int(latest.get("comment_id"))
except (TypeError, ValueError):
comment_id = None
return HandoffSummary(
comment_id=comment_id,
created_at=_redact(latest.get("created_at")),
author=_redact(latest.get("author")),
# An unrecognised heading is reported as such rather than republished:
# the heading is free text, and CTH_TYPES is the only authority for what
# a handoff type may be.
cth_type=cth_type if known else "unrecognized",
cth_type_known=known,
status=_redact(fields.get("status")),
next_owner=_redact(fields.get("next owner")),
current_blocker=_redact(fields.get("current blocker")),
decision=_redact(fields.get("decision")),
next_action=_redact(fields.get("next action")),
)
CommentSource = Callable[[str, int], list[dict[str, Any]]]
def build_comment_source(host: str, org: str, repo: str) -> CommentSource | None:
"""Build an authenticated ``(kind, number) -> comments`` fetcher, or None.
Returns None when the console is running in offline test mode or when no
credential is available for *host*, so the caller reports the handoff source
as unavailable instead of as an empty thread.
"""
if _offline_test_mode() or not (host and org and repo):
return None
auth = get_auth_header(host)
if not auth:
return None
def _fetch(kind: str, number: int) -> list[dict[str, Any]]:
segment = "pulls" if kind == "pr" else "issues"
url = f"{repo_api_url(host, org, repo)}/{segment}/{int(number)}/comments"
comments: list[dict[str, Any]] = []
page = 1
while page <= 20:
raw, meta = api_fetch_page(url, auth, page=page, limit=50)
comments.extend(raw)
if bool(meta["is_final_page"]):
break
page += 1
return comments
return _fetch
# --- Snapshot ----------------------------------------------------------------
@dataclass(frozen=True)
class LinkageNode:
"""One issue or PR row with its resolved links and display metadata."""
kind: str
number: int
title: str
state: str
labels: tuple[str, ...] = ()
links: tuple[IssueLink, ...] = ()
linked_prs: tuple[int, ...] = ()
ambiguous: bool = False
contested: bool = False
deep_link: str | None = None
handoff: HandoffSummary | None = None
handoff_status: HandoffStatus = HandoffStatus(HANDOFF_NOT_LOADED)
links_authoritative: bool = True
def to_dict(self) -> dict[str, Any]:
return {
"kind": self.kind,
"number": self.number,
"title": self.title,
"state": self.state,
"labels": list(self.labels),
"links": [link.to_dict() for link in self.links],
"linked_prs": list(self.linked_prs),
"ambiguous": self.ambiguous,
"contested": self.contested,
"deep_link": self.deep_link,
"links_authoritative": self.links_authoritative,
"handoff": self.handoff.to_dict() if self.handoff else None,
"handoff_status": self.handoff_status.to_dict(),
}
@dataclass(frozen=True)
class LinkageSnapshot:
"""One answered linkage query over a scoped window of a Gitea repo."""
ok: bool
project_id: str
repo_label: str
host: str
state_scope: str
issues: tuple[LinkageNode, ...] = ()
prs: tuple[LinkageNode, ...] = ()
contested_issues: tuple[int, ...] = ()
focus: tuple[str, int] | None = None
inventory_complete: bool = False
deep_links_enabled: bool = False
handoff_status: HandoffStatus = HandoffStatus(HANDOFF_NOT_LOADED)
fetch_error: str | None = None
@property
def orphan_prs(self) -> tuple[LinkageNode, ...]:
"""PRs in the loaded window that point at no issue at all."""
return tuple(node for node in self.prs if not node.links)
def to_dict(self) -> dict[str, Any]:
return {
"ok": self.ok,
"schema_version": LINKAGE_SCHEMA_VERSION,
"project_id": self.project_id,
"repo": self.repo_label,
"state_scope": self.state_scope,
"inventory_complete": self.inventory_complete,
"deep_links_enabled": self.deep_links_enabled,
"focus": (
None
if self.focus is None
else {"kind": self.focus[0], "number": self.focus[1]}
),
"handoff_source": self.handoff_status.to_dict(),
"fetch_error": self.fetch_error,
"contested_issues": list(self.contested_issues),
"issues": [node.to_dict() for node in self.issues],
"prs": [node.to_dict() for node in self.prs],
"evidence_kinds": [
{"name": name, "description": EVIDENCE_DESCRIPTIONS[name]}
for name in EVIDENCE_ORDER
],
}
def snapshot_to_dict(snapshot: LinkageSnapshot) -> dict[str, Any]:
"""JSON-serializable export for ``/api/v1/gitea/linkage``."""
return snapshot.to_dict()
def _offline_test_mode() -> bool:
return (os.environ.get("WEBUI_TEST_OFFLINE") or "").strip().lower() in {
"1",
"true",
"yes",
}
def _labels_of(item: dict[str, Any]) -> tuple[str, ...]:
return tuple(
str(_redact(label.get("name")))
for label in (item.get("labels") or [])
if label.get("name")
)
def _failed_snapshot(
*,
project_id: str,
repo_label: str,
host: str,
state_scope: str,
reason: str,
) -> LinkageSnapshot:
"""A read that could not be answered. Never an empty-and-healthy snapshot."""
return LinkageSnapshot(
ok=False,
project_id=project_id,
repo_label=repo_label,
host=host,
state_scope=state_scope,
inventory_complete=False,
deep_links_enabled=deep_links_enabled(),
handoff_status=HandoffStatus(
HANDOFF_UNAVAILABLE, reason="linkage inventory could not be loaded"
),
fetch_error=str(_redact(reason)),
)
def _resolve_project(project_id: str | None) -> ProjectRecord | None:
registry = load_registry()
if project_id:
for entry in registry.projects:
if entry.id == project_id:
return entry
return None
return registry.projects[0] if registry.projects else None
def _normalize_state(state: str | None) -> str:
text = (state or STATE_OPEN).strip().lower()
return text if text in _SUPPORTED_STATES else STATE_OPEN
def load_linkage_snapshot(
project_id: str | None = None,
*,
state: str | None = None,
issue: int | None = None,
pr: int | None = None,
fetch_prs: Callable[..., tuple[list[dict], PaginationMeta]] | None = None,
fetch_issues: Callable[..., tuple[list[dict], PaginationMeta]] | None = None,
comment_source: CommentSource | None = None,
) -> LinkageSnapshot:
"""Load issue↔PR linkage for a registry project.
``issue``/``pr`` focus one thread: the focused row is the only one whose
Canonical Thread Handoff comments are fetched, because handoff comments are
thread-scoped and loading them for a whole queue would be one request per
row. Every unfocused row reports its handoff as ``not_loaded``.
"""
state_scope = _normalize_state(state)
try:
project = _resolve_project(project_id)
except Exception as exc: # registry invalid — fail closed with the reason
return _failed_snapshot(
project_id=project_id or "",
repo_label="",
host="",
state_scope=state_scope,
reason=f"project registry unavailable: {exc}",
)
if project is None:
return _failed_snapshot(
project_id=project_id or "",
repo_label="",
host="",
state_scope=state_scope,
reason=(
f"project {project_id!r} not found in registry"
if project_id
else "no projects registered"
),
)
host = _host_from_url(project.remote_host)
repo_label = f"{project.gitea_owner}/{project.repo_name}"
offline_test = _offline_test_mode()
def _empty_fetch(*_args, **_kwargs):
return [], PaginationMeta(
page=1,
per_page=50,
returned_count=0,
has_more=False,
is_final_page=True,
# An offline stub loaded nothing; claiming a complete inventory here
# would let the page assert that no issue has a linked PR.
inventory_complete=False,
pages_fetched=0,
)
pr_fetch = fetch_prs or (_empty_fetch if offline_test else _fetch_prs)
issue_fetch = fetch_issues or (_empty_fetch if offline_test else _fetch_issues)
using_live_fetch = not offline_test and (fetch_prs is None or fetch_issues is None)
auth = get_auth_header(host) if using_live_fetch else "test-auth"
if using_live_fetch and not auth:
return _failed_snapshot(
project_id=project.id,
repo_label=repo_label,
host=host,
state_scope=state_scope,
reason=(
f"Gitea credentials unavailable for {host}; linkage cannot be "
"loaded (fail closed — not rendering an empty linkage table)"
),
)
try:
raw_prs, pr_pagination = pr_fetch(
host, project.gitea_owner, project.repo_name, auth, state=state_scope
)
raw_issues, issue_pagination = issue_fetch(
host, project.gitea_owner, project.repo_name, auth, state=state_scope
)
except Exception as exc: # noqa: BLE001 — operator-visible fetch failure
return _failed_snapshot(
project_id=project.id,
repo_label=repo_label,
host=host,
state_scope=state_scope,
reason=f"Gitea fetch failed: {exc}",
)
inventory_complete = bool(
getattr(pr_pagination, "inventory_complete", False)
and getattr(issue_pagination, "inventory_complete", False)
)
index = resolve_linkage(raw_prs)
contested = index.contested_issues()
focus: tuple[str, int] | None = None
if pr is not None:
focus = ("pr", int(pr))
elif issue is not None:
focus = ("issue", int(issue))
unfocused_reason = (
"canonical handoff comments are thread-scoped; focus one issue or PR "
"to load its latest handoff"
)
handoff_status = HandoffStatus(HANDOFF_NOT_LOADED, reason=unfocused_reason)
focus_handoff: HandoffSummary | None = None
if focus is not None:
source = comment_source
if source is None and not offline_test:
source = build_comment_source(host, project.gitea_owner, project.repo_name)
target = f"{focus[0]}#{focus[1]}"
if source is None:
handoff_status = HandoffStatus(
HANDOFF_UNAVAILABLE,
reason="no authenticated comment source available for this read",
target=target,
)
else:
try:
focus_handoff = summarize_handoff(source(focus[0], focus[1]) or [])
handoff_status = HandoffStatus(HANDOFF_LOADED, target=target)
except Exception as exc: # fail soft: degrade this source only
handoff_status = HandoffStatus(
HANDOFF_UNAVAILABLE,
reason=str(_redact(f"handoff fetch failed: {exc}")),
target=target,
)
def _node_handoff(
kind: str, number: int
) -> tuple[HandoffSummary | None, HandoffStatus]:
"""Attach the handoff only to the focused row; qualify every other row."""
if focus == (kind, number):
return (focus_handoff, handoff_status)
return (
None,
HandoffStatus(
HANDOFF_NOT_LOADED,
reason=unfocused_reason if focus is None else "not the focused thread",
),
)
def _number_of(raw: dict[str, Any]) -> int | None:
try:
return int(raw["number"])
except (KeyError, TypeError, ValueError):
return None
def _sort_key(raw: dict[str, Any]) -> int:
number = _number_of(raw)
return -1 if number is None else number
issue_nodes: list[LinkageNode] = []
for raw in sorted(raw_issues or [], key=_sort_key, reverse=True):
number = _number_of(raw)
if number is None:
continue
node_handoff, node_status = _node_handoff("issue", number)
linked_prs = index.issue_prs.get(number, ())
issue_nodes.append(
LinkageNode(
kind="issue",
number=number,
title=str(_redact(raw.get("title")) or ""),
state=str(raw.get("state") or ""),
labels=_labels_of(raw),
linked_prs=linked_prs,
contested=len(linked_prs) > 1,
deep_link=_deep_link(
host, project.gitea_owner, project.repo_name, "issue", number
),
handoff=node_handoff,
handoff_status=node_status,
links_authoritative=inventory_complete,
)
)
pr_nodes: list[LinkageNode] = []
for raw in sorted(raw_prs or [], key=_sort_key, reverse=True):
number = _number_of(raw)
if number is None:
continue
node_handoff, node_status = _node_handoff("pr", number)
links = index.pr_links.get(number, ())
pr_nodes.append(
LinkageNode(
kind="pr",
number=number,
title=str(_redact(raw.get("title")) or ""),
state=str(raw.get("state") or ""),
labels=_labels_of(raw),
links=links,
ambiguous=index.ambiguous(number),
contested=any(link.issue_number in contested for link in links),
deep_link=_deep_link(
host, project.gitea_owner, project.repo_name, "pr", number
),
handoff=node_handoff,
handoff_status=node_status,
links_authoritative=inventory_complete,
)
)
return LinkageSnapshot(
ok=True,
project_id=project.id,
repo_label=repo_label,
host=host,
state_scope=state_scope,
issues=tuple(issue_nodes),
prs=tuple(pr_nodes),
contested_issues=contested,
focus=focus,
inventory_complete=inventory_complete,
deep_links_enabled=deep_links_enabled(),
handoff_status=handoff_status,
)
+364
View File
@@ -0,0 +1,364 @@
"""HTML views for the Gitea issue↔PR linkage console (#645, Phase 3).
Read-only renderer over :mod:`webui.linkage_loader`. The page's job is to make
three things impossible to misread:
* **why** an edge exists every link carries its evidence badge, so a bare
``#N`` mention never looks like a closing claim;
* **what was not loaded** a partial inventory, an unfocused thread, or an
unavailable handoff source renders as an explicit qualifier, never as an
affirmative "none";
* **that nothing here mutates** there is no review, merge, or edit control,
and the deep link out to Gitea appears only under the admin reveal opt-in.
"""
from __future__ import annotations
from html import escape
from typing import Sequence
from webui.layout import render_page
from webui.linkage_loader import (
EVIDENCE_BRANCH,
EVIDENCE_CLOSES,
EVIDENCE_DESCRIPTIONS,
EVIDENCE_ORDER,
EVIDENCE_REFERENCE,
HANDOFF_LOADED,
HANDOFF_NOT_LOADED,
HandoffSummary,
LinkageNode,
LinkageSnapshot,
)
_EVIDENCE_CSS = {
EVIDENCE_CLOSES: "badge-health-ok",
EVIDENCE_BRANCH: "badge-health-skipped",
EVIDENCE_REFERENCE: "badge-health-unproven",
}
_EVIDENCE_LABEL = {
EVIDENCE_CLOSES: "closes",
EVIDENCE_BRANCH: "branch",
EVIDENCE_REFERENCE: "mention",
}
def _badge(text: str, css: str) -> str:
return f'<span class="badge {css}">{escape(text)}</span>'
def _labels(names: Sequence[str]) -> str:
if not names:
return '<span class="muted">—</span>'
return " ".join(_badge(name, "badge-health-skipped") for name in names)
def _ref(node: LinkageNode) -> str:
"""Render an item reference, hyperlinked only when deep links are revealed."""
label = f"#{node.number}"
if node.deep_link:
return f'<a href="{escape(node.deep_link)}"><code>{escape(label)}</code></a>'
return f"<code>{escape(label)}</code>"
def _scope_card(snapshot: LinkageSnapshot) -> str:
focus = (
"none"
if snapshot.focus is None
else f"{snapshot.focus[0]}#{snapshot.focus[1]}"
)
completeness = (
_badge("complete", "badge-health-ok")
if snapshot.inventory_complete
else _badge("partial", "badge-health-degraded")
)
links_note = (
"Every linkage edge below is a claim about this loaded window only."
if snapshot.inventory_complete
else (
"Pagination did not complete for this window, so an empty link list "
"means <em>none found in what was loaded</em> — not that no link exists."
)
)
deep_links = (
_badge("enabled", "badge-health-ok")
if snapshot.deep_links_enabled
else _badge("hidden", "badge-health-skipped")
)
return f"""<div class="health-card">
<h3>Scope</h3>
<table class="detail">
<tr><th>Project</th><td><code>{escape(snapshot.project_id or "")}</code></td></tr>
<tr><th>Repository</th><td><code>{escape(snapshot.repo_label or "")}</code></td></tr>
<tr><th>Item state</th><td><code>{escape(snapshot.state_scope)}</code></td></tr>
<tr><th>Focused thread</th><td><code>{escape(focus)}</code></td></tr>
<tr><th>Inventory</th><td>{completeness}</td></tr>
<tr><th>Gitea deep links</th><td>{deep_links}</td></tr>
</table>
<p class="muted">{links_note}</p>
</div>"""
def _error_card(snapshot: LinkageSnapshot) -> str:
if snapshot.ok and not snapshot.fetch_error:
return ""
return (
'<div class="health-card health-stale"><strong>Linkage unavailable:</strong> '
f"{escape(snapshot.fetch_error or 'the linkage read did not complete')}. "
"No linkage table is rendered: an empty table would read as "
"<em>no issue is linked to any PR</em>, which this read cannot claim."
"</div>"
)
def _contested_card(snapshot: LinkageSnapshot) -> str:
if not snapshot.contested_issues:
return ""
refs = ", ".join(f"<code>#{number}</code>" for number in snapshot.contested_issues)
return (
'<div class="health-card health-stale">'
f"<strong>Contested issues:</strong> {refs}. More than one PR in this "
"window claims each of these — duplicate work or a superseded PR. "
"Resolution stays in Gitea and the workflow; this console only reports it."
"</div>"
)
def _evidence_badges(evidence: Sequence[str]) -> str:
return " ".join(
_badge(
_EVIDENCE_LABEL.get(name, name),
_EVIDENCE_CSS.get(name, "badge-health-skipped"),
)
for name in EVIDENCE_ORDER
if name in evidence
)
def _handoff_inline(handoff: HandoffSummary) -> str:
type_css = "badge-health-ok" if handoff.cth_type_known else "badge-health-degraded"
return (
f'{_badge(handoff.cth_type or "", type_css)}'
f'<div class="muted" style="font-size:0.82rem;">'
f'{escape(handoff.status or "")}{escape(handoff.next_owner or "")}</div>'
)
def _handoff_cell(node: LinkageNode) -> str:
"""Render the handoff column, distinguishing 'none found' from 'not loaded'."""
status = node.handoff_status
if status.state == HANDOFF_LOADED:
if node.handoff is None:
return '<span class="muted">no canonical handoff on this thread</span>'
return _handoff_inline(node.handoff)
if status.state == HANDOFF_NOT_LOADED:
return (
f'{_badge("not loaded", "badge-health-skipped")}'
f'<div class="muted" style="font-size:0.82rem;">'
f'{escape(status.reason or "")}</div>'
)
return (
f'{_badge("unavailable", "badge-health-degraded")}'
f'<div class="muted" style="font-size:0.82rem;">'
f'{escape(status.reason or "")}</div>'
)
def _issue_rows(snapshot: LinkageSnapshot) -> str:
rows = []
for node in snapshot.issues:
if node.linked_prs:
linked = ", ".join(f"<code>#{number}</code>" for number in node.linked_prs)
if node.contested:
linked += " " + _badge("contested", "badge-blocked")
elif node.links_authoritative:
linked = '<span class="muted">none</span>'
else:
# The distinction an operator needs: nothing found in a window that
# was not fully loaded is not the same as nothing existing.
linked = '<span class="muted">none found (partial inventory)</span>'
rows.append(
"<tr>"
f"<td>{_ref(node)}</td>"
f"<td>{escape(node.title)}</td>"
f"<td>{escape(node.state or '')}</td>"
f"<td>{_labels(node.labels)}</td>"
f"<td>{linked}</td>"
f"<td>{_handoff_cell(node)}</td>"
"</tr>"
)
if not rows:
return '<tr><td colspan="6" class="muted">No issues in the loaded window.</td></tr>'
return "".join(rows)
def _pr_rows(snapshot: LinkageSnapshot) -> str:
rows = []
for node in snapshot.prs:
if node.links:
linked = "".join(
f"<div><code>#{link.issue_number}</code> "
f"{_evidence_badges(link.evidence)}</div>"
for link in node.links
)
if node.ambiguous:
linked += _badge("ambiguous", "badge-blocked")
if node.contested:
linked += " " + _badge("contested", "badge-blocked")
elif node.links_authoritative:
linked = '<span class="muted">no issue reference</span>'
else:
linked = '<span class="muted">none found (partial inventory)</span>'
rows.append(
"<tr>"
f"<td>{_ref(node)}</td>"
f"<td>{escape(node.title)}</td>"
f"<td>{escape(node.state or '')}</td>"
f"<td>{_labels(node.labels)}</td>"
f"<td>{linked}</td>"
f"<td>{_handoff_cell(node)}</td>"
"</tr>"
)
if not rows:
return (
'<tr><td colspan="6" class="muted">No pull requests in the loaded '
"window.</td></tr>"
)
return "".join(rows)
def _focus_card(snapshot: LinkageSnapshot) -> str:
"""Render the focused thread's latest canonical handoff, when one was loaded."""
if snapshot.focus is None:
return f"""<div class="prompt-card">
<h3>Canonical handoff</h3>
<p class="muted">{escape(snapshot.handoff_status.reason or "")}
Add <code>?issue=N</code> or <code>?pr=N</code> to load the latest
Canonical Thread Handoff for one thread.</p>
</div>"""
kind, number = snapshot.focus
target = f"{kind} #{number}"
if not snapshot.handoff_status.loaded:
return f"""<div class="prompt-card">
<h3>Canonical handoff {escape(target)}</h3>
<p class="muted">{_badge("unavailable", "badge-health-degraded")}
{escape(snapshot.handoff_status.reason or "handoff source did not run")}.
This is not evidence that the thread carries no handoff.</p>
</div>"""
handoff = next(
(
node.handoff
for node in (snapshot.issues + snapshot.prs)
if node.kind == kind and node.number == number and node.handoff
),
None,
)
if handoff is None:
return f"""<div class="prompt-card">
<h3>Canonical handoff {escape(target)}</h3>
<p class="muted">Comments loaded; no Canonical Thread Handoff comment found on
this thread.</p>
</div>"""
type_css = "badge-health-ok" if handoff.cth_type_known else "badge-health-degraded"
unknown_note = (
""
if handoff.cth_type_known
else (
'<p class="muted">The comment\'s heading is not a declared CTH type, '
"so it is reported as unrecognized rather than republished.</p>"
)
)
return f"""<div class="prompt-card">
<h3>Canonical handoff {escape(target)}</h3>
<p class="meta">{_badge(handoff.cth_type or "", type_css)}
by <code>{escape(handoff.author or "unknown")}</code>
at <code>{escape(handoff.created_at or "unknown")}</code></p>
{unknown_note}
<table class="detail">
<tr><th>Status</th><td>{escape(handoff.status or "")}</td></tr>
<tr><th>Next owner</th><td>{escape(handoff.next_owner or "")}</td></tr>
<tr><th>Current blocker</th><td>{escape(handoff.current_blocker or "")}</td></tr>
<tr><th>Decision</th><td>{escape(handoff.decision or "")}</td></tr>
<tr><th>Next action</th><td>{escape(handoff.next_action or "")}</td></tr>
</table>
<p class="muted">Full event history:
<a href="/api/v1/timeline?{escape(kind)}={number}"><code>/api/v1/timeline</code></a></p>
</div>"""
def _legend_card(snapshot: LinkageSnapshot) -> str:
items = "".join(
f"<li>{_badge(_EVIDENCE_LABEL[name], _EVIDENCE_CSS[name])}"
f"{escape(EVIDENCE_DESCRIPTIONS[name])}</li>"
for name in EVIDENCE_ORDER
)
reveal_note = (
"Gitea deep links are shown because the "
"<code>GITEA_MCP_REVEAL_ENDPOINTS</code> admin opt-in is set."
if snapshot.deep_links_enabled
else (
"Gitea deep links are withheld. Set "
"<code>GITEA_MCP_REVEAL_ENDPOINTS=1</code> server-side to reveal "
"them; item numbers stay usable without them."
)
)
return f"""<div class="prompt-card">
<h3>How an edge was found</h3>
<ul class="reasons">{items}</ul>
<p class="muted">{reveal_note}</p>
<p class="muted">Read-only surface: no issue or PR editing, no review, and no merge.
JSON export: <a href="/api/v1/gitea/linkage"><code>/api/v1/gitea/linkage</code></a></p>
</div>"""
def render_linkage_page(snapshot: LinkageSnapshot) -> str:
"""Render the full HTML page for the Gitea linkage console."""
if not snapshot.ok:
return render_page(
title="Gitea linkage",
body_html=f"""<h2>Gitea issue and PR linkage</h2>
<p class="meta">Phase 3 read-only linkage console (#645).</p>
{_error_card(snapshot)}
{_scope_card(snapshot)}""",
)
body = f"""<h2>Gitea issue and PR linkage</h2>
<p class="meta">Phase 3 read-only linkage console (#645). Gitea remains the source of
truth; this page reads it and never writes to it.</p>
{_scope_card(snapshot)}
{_contested_card(snapshot)}
<div class="prompt-card">
<h3>Issues pull requests</h3>
<table class="registry">
<thead>
<tr>
<th>Issue</th><th>Title</th><th>State</th><th>Labels</th>
<th>Linked PRs</th><th>Latest handoff</th>
</tr>
</thead>
<tbody>{_issue_rows(snapshot)}</tbody>
</table>
</div>
<div class="prompt-card">
<h3>Pull requests issues</h3>
<table class="registry">
<thead>
<tr>
<th>PR</th><th>Title</th><th>State</th><th>Labels</th>
<th>Linked issues</th><th>Latest handoff</th>
</tr>
</thead>
<tbody>{_pr_rows(snapshot)}</tbody>
</table>
</div>
{_focus_card(snapshot)}
{_legend_card(snapshot)}
"""
return render_page(title="Gitea linkage", body_html=body)
+7 -1
View File
@@ -6,7 +6,8 @@ destination is a GET view or a Phase 1 placeholder. No mutation links.
Nav groups follow the #631 Phase 1 information architecture: Health, Traffic,
Runtime/Sessions, Projects, Inventory, Timeline, Policy (placeholder), and
Insights (placeholder). Later-phase surfaces are declared as ``stub`` items and
Insights (placeholder), joined by the Phase 3 Gitea linkage group (#645).
Later-phase surfaces are declared as ``stub`` items and
backed by ``STUB_PAGES`` so their nav links resolve to a graceful placeholder
instead of a 404.
"""
@@ -45,6 +46,8 @@ NAV_GROUPS: tuple[NavGroup, ...] = (
NavItem("/queue", "Queue"),
NavItem("/leases", "Leases"),
NavItem("/actions", "Actions"),
NavItem("/notifications", "Notifications"),
NavItem("/requests", "Requests"),
)),
NavGroup("Runtime/Sessions", (
NavItem("/runtime", "Runtime health"),
@@ -61,6 +64,9 @@ NAV_GROUPS: tuple[NavGroup, ...] = (
NavGroup("Timeline", (
NavItem("/timeline", "Timeline", "stub"),
)),
NavGroup("Gitea", (
NavItem("/gitea", "Issue/PR linkage"),
)),
NavGroup("Policy", (
NavItem("/policy", "Policy", "stub"),
NavItem("/prompts", "Prompts"),
+158
View File
@@ -0,0 +1,158 @@
"""HTML rendering for Phase 3 Notifications and Human-Attention Console (#648)."""
from __future__ import annotations
from html import escape
from typing import Sequence
from webui.layout import render_page
from webui.notifications import (
ATTENTION_HUMAN_REQUIRED,
ATTENTION_OPERATOR,
ATTENTION_ROUTINE,
NotificationItem,
NotificationSnapshot,
)
def _render_attention_badge(attention_class: str) -> str:
cls = "badge"
if attention_class == ATTENTION_HUMAN_REQUIRED:
cls += " badge-blocked"
elif attention_class == ATTENTION_OPERATOR:
cls += " badge-claimed"
else:
cls += " muted"
return f'<span class="{cls}">{escape(attention_class)}</span>'
def _render_notification_row(item: NotificationItem) -> str:
category_label = escape(item.category.upper())
id_str = escape(item.id)
title_str = escape(item.title)
summary_str = escape(item.summary)
att_badge = _render_attention_badge(item.attention_class)
work_item_html = ""
if item.work_number and item.work_kind:
kind_label = escape(item.work_kind.upper())
num_str = f"#{item.work_number}"
link = item.deep_link or "#"
work_item_html = f'<a href="{escape(link)}"><code>{kind_label} {num_str}</code></a>'
requires_human_label = (
'<span class="badge badge-blocked" style="font-size:0.75rem;">HUMAN REQUIRED</span>'
if item.requires_human
else ""
)
return f"""<tr>
<td><code>{category_label}</code><br><span class="muted" style="font-size:0.75rem;">{id_str}</span></td>
<td>
<div><strong>{title_str}</strong> {att_badge} {requires_human_label}</div>
<div class="muted" style="font-size:0.85rem; margin-top:0.25rem;">{summary_str}</div>
</td>
<td>{work_item_html}</td>
<td><span class="muted" style="font-size:0.8rem;">{escape(item.created_at[:19])}</span></td>
</tr>"""
def _render_notifications_table(items: Sequence[NotificationItem], empty_message: str) -> str:
if not items:
return f'<p class="muted" style="padding:1rem 0;">{escape(empty_message)}</p>'
rows = "".join(_render_notification_row(item) for item in items)
return f"""<table class="registry">
<thead>
<tr>
<th style="width: 18%;">Category & ID</th>
<th style="width: 52%;">Title & Attention Summary</th>
<th style="width: 15%;">Work Item</th>
<th style="width: 15%;">Time</th>
</tr>
</thead>
<tbody>
{rows}
</tbody>
</table>"""
def render_notifications_page(
snapshot: NotificationSnapshot,
*,
filter_class: str = "inbox",
filter_project: str | None = None,
) -> str:
"""Render the notifications and attention inbox page."""
title = "Notifications & Attention Inbox"
err_html = ""
if snapshot.fetch_error:
err_html = f'<div class="stub" style="border-color:#e53e3e; background:#fff5f5; color:#c53030; margin-bottom:1rem;"><p><strong>Fetch Warning:</strong> {escape(snapshot.fetch_error)}</p></div>'
# Determine items to render based on filter_class
if filter_class == ATTENTION_HUMAN_REQUIRED:
display_items = snapshot.human_required_items
active_tab_title = "Human-Required Escalations"
elif filter_class == ATTENTION_OPERATOR:
display_items = snapshot.operator_items
active_tab_title = "Operator Inbox Items"
elif filter_class == ATTENTION_ROUTINE:
display_items = snapshot.routine_items
active_tab_title = "Routine Workflow Transitions"
elif filter_class == "all":
display_items = snapshot.items
active_tab_title = "All Events (including Routine)"
else: # "inbox" default
display_items = snapshot.inbox_items
active_tab_title = "Attention Inbox (Human + Operator)"
hr_cls = "badge-blocked" if snapshot.human_required_count > 0 else "muted"
op_cls = "badge-claimed" if snapshot.operator_count > 0 else "muted"
metrics_html = f"""<div style="display:flex; gap:1rem; margin-bottom:1.5rem;">
<div class="health-card" style="flex:1;">
<span class="muted" style="font-size:0.85rem;">Human Required</span>
<h2 style="margin:0.2rem 0;"><span class="badge {hr_cls}" style="font-size:1.4rem;">{snapshot.human_required_count}</span></h2>
<p class="muted" style="font-size:0.8rem; margin:0;">Critical escalation boundary</p>
</div>
<div class="health-card" style="flex:1;">
<span class="muted" style="font-size:0.85rem;">Operator Inbox</span>
<h2 style="margin:0.2rem 0;"><span class="badge {op_cls}" style="font-size:1.4rem;">{snapshot.operator_count}</span></h2>
<p class="muted" style="font-size:0.8rem; margin:0;">Operational items needing review</p>
</div>
<div class="health-card" style="flex:1;">
<span class="muted" style="font-size:0.85rem;">Routine Transitions</span>
<h2 style="margin:0.2rem 0;"><span class="badge muted" style="font-size:1.4rem;">{snapshot.routine_count}</span></h2>
<p class="muted" style="font-size:0.8rem; margin:0;">Background transitions (filtered)</p>
</div>
</div>"""
# Filter navigation links
def _tab_link(target_class: str, label: str) -> str:
is_active = (filter_class == target_class)
style = "font-weight:bold; border-bottom:2px solid currentColor;" if is_active else "color:#4a5568;"
return f'<a href="/notifications?attention_class={target_class}" style="margin-right:1.25rem; text-decoration:none; padding-bottom:0.25rem; {style}">{label}</a>'
tabs_html = f"""<div style="margin-bottom:1.25rem; border-bottom:1px solid #e2e8f0; padding-bottom:0.5rem;">
{_tab_link("inbox", f"Attention Inbox ({snapshot.human_required_count + snapshot.operator_count})")}
{_tab_link("human-required", f"Human Required ({snapshot.human_required_count})")}
{_tab_link("operator", f"Operator ({snapshot.operator_count})")}
{_tab_link("routine", f"Routine ({snapshot.routine_count})")}
{_tab_link("all", f"All Events ({snapshot.total_count})")}
</div>"""
table_html = _render_notifications_table(
display_items,
f"No items match attention filter '{filter_class}'.",
)
body = f"""<h2>{escape(title)}</h2>
<p class="muted">Phase 3 console surface for human-attention routing (#648). Routine workflow transitions are filtered by default to eliminate notification fatigue.</p>
{err_html}
{metrics_html}
{tabs_html}
<h3>{escape(active_tab_title)}</h3>
{table_html}"""
return render_page(title=title, body_html=body)
+486
View File
@@ -0,0 +1,486 @@
"""Notifications and human-attention routing module for Phase 3 web console (#648).
Defines attention classes, event classification rules, and inbox aggregation so
operators receive direct alerts only for human-required escalation boundaries
(#628) while routine workflow transitions remain available for pull-based review.
"""
from __future__ import annotations
from dataclasses import dataclass, field
from datetime import datetime, timezone
from typing import Any, Callable
from webui import console_redaction
from webui.project_registry import load_registry
from webui.queue_loader import QueueSnapshot, load_queue_snapshot
from webui.lease_loader import LeaseSnapshot, load_lease_snapshot
from webui.system_health import SystemHealthSnapshot, load_system_health
# Attention class definitions (#628, #648)
ATTENTION_ROUTINE = "routine"
ATTENTION_OPERATOR = "operator"
ATTENTION_HUMAN_REQUIRED = "human-required"
ATTENTION_CLASSES = (
ATTENTION_ROUTINE,
ATTENTION_OPERATOR,
ATTENTION_HUMAN_REQUIRED,
)
# Notification categories
CATEGORY_AUTH = "auth"
CATEGORY_BLOCKER = "blocker"
CATEGORY_LEASE = "lease"
CATEGORY_VALIDATION = "validation"
CATEGORY_WORKFLOW = "workflow"
CATEGORY_SYSTEM = "system"
CATEGORIES = (
CATEGORY_AUTH,
CATEGORY_BLOCKER,
CATEGORY_LEASE,
CATEGORY_VALIDATION,
CATEGORY_WORKFLOW,
CATEGORY_SYSTEM,
)
@dataclass(frozen=True)
class NotificationItem:
"""A single notification or inbox event."""
id: str
attention_class: str # "routine", "operator", "human-required"
category: str # "auth", "blocker", "lease", "validation", etc.
title: str
summary: str
work_kind: str | None # "issue", "pr", "session", "system"
work_number: int | None
project_id: str
repo_label: str
created_at: str
deep_link: str | None = None
requires_human: bool = False
extra: dict[str, Any] = field(default_factory=dict)
def as_dict(self) -> dict[str, Any]:
return {
"id": self.id,
"attention_class": self.attention_class,
"category": self.category,
"title": self.title,
"summary": console_redaction.redact_text(self.summary),
"work_kind": self.work_kind,
"work_number": self.work_number,
"project_id": self.project_id,
"repo_label": self.repo_label,
"created_at": self.created_at,
"deep_link": self.deep_link,
"requires_human": self.requires_human,
"extra": self.extra,
}
@dataclass(frozen=True)
class NotificationSnapshot:
"""Snapshot of notifications and attention inbox state."""
project_id: str
repo_label: str
items: tuple[NotificationItem, ...]
human_required_count: int
operator_count: int
routine_count: int
total_count: int
fetch_error: str | None = None
@property
def inbox_items(self) -> tuple[NotificationItem, ...]:
"""Items requiring operator or human attention (excluding routine)."""
return tuple(
item
for item in self.items
if item.attention_class in {ATTENTION_OPERATOR, ATTENTION_HUMAN_REQUIRED}
)
@property
def human_required_items(self) -> tuple[NotificationItem, ...]:
return tuple(
item for item in self.items if item.attention_class == ATTENTION_HUMAN_REQUIRED
)
@property
def operator_items(self) -> tuple[NotificationItem, ...]:
return tuple(
item for item in self.items if item.attention_class == ATTENTION_OPERATOR
)
@property
def routine_items(self) -> tuple[NotificationItem, ...]:
return tuple(
item for item in self.items if item.attention_class == ATTENTION_ROUTINE
)
def as_dict(self) -> dict[str, Any]:
return {
"project_id": self.project_id,
"repo_label": self.repo_label,
"human_required_count": self.human_required_count,
"operator_count": self.operator_count,
"routine_count": self.routine_count,
"total_count": self.total_count,
"fetch_error": self.fetch_error,
"inbox_items": [item.as_dict() for item in self.inbox_items],
"all_items": [item.as_dict() for item in self.items],
}
def classify_attention_event(
category: str,
title: str,
summary: str,
*,
is_hard_stop: bool = False,
is_auth_failure: bool = False,
is_irrecoverable: bool = False,
is_decision_lock: bool = False,
is_validation_failure: bool = False,
is_stale: bool = False,
is_blocker: bool = False,
) -> tuple[str, bool]:
"""Classify an event into an attention class and human requirement flag.
Rules (#628, #648):
1. Critical boundaries (hard stop, auth failure, irrecoverable state,
decision lock, validation failure) -> ATTENTION_HUMAN_REQUIRED (requires_human=True).
2. Operational queues (blocker, stale lease, unassigned ready work, queue collision)
-> ATTENTION_OPERATOR (requires_human=False).
3. Routine state transitions (clean progression, healthy heartbeats) -> ATTENTION_ROUTINE (requires_human=False).
Classification uses structured flags and category only. Human-authored
``title`` / ``summary`` text is never substring-matched for escalation
(PR #905 review B1) — callers that need text signals must set flags from
machine-generated status/detail fields before calling this function.
"""
del title, summary # kept for API stability; never used for classification
if (
is_hard_stop
or is_auth_failure
or is_irrecoverable
or is_decision_lock
or is_validation_failure
or category in {CATEGORY_AUTH, CATEGORY_VALIDATION}
):
return ATTENTION_HUMAN_REQUIRED, True
if is_stale or is_blocker or category in {CATEGORY_BLOCKER, CATEGORY_LEASE}:
return ATTENTION_OPERATOR, False
return ATTENTION_ROUTINE, False
def load_notifications_snapshot(
project_id: str | None = None,
*,
load_queue: Callable[..., QueueSnapshot] | None = None,
load_leases: Callable[..., LeaseSnapshot] | None = None,
load_health: Callable[..., SystemHealthSnapshot] | None = None,
) -> NotificationSnapshot:
"""Load and classify attention notifications across queue, leases, and system health."""
registry = load_registry()
project = None
if project_id:
for entry in registry.projects:
if entry.id == project_id:
project = entry
break
else:
project = registry.projects[0] if registry.projects else None
if project is None:
return NotificationSnapshot(
project_id=project_id or "",
repo_label="",
items=(),
human_required_count=0,
operator_count=0,
routine_count=0,
total_count=0,
fetch_error="project not found in registry",
)
queue_loader_fn = load_queue or load_queue_snapshot
lease_loader_fn = load_leases or load_lease_snapshot
health_loader_fn = load_health or load_system_health
try:
queue_snap = queue_loader_fn(project.id)
except TypeError:
queue_snap = queue_loader_fn(project_id=project.id)
try:
lease_snap = lease_loader_fn(project_id=project.id)
except TypeError:
lease_snap = lease_loader_fn(project.id)
try:
health_snap = health_loader_fn(project_id=project.id)
except TypeError:
try:
health_snap = health_loader_fn(project.id)
except TypeError:
health_snap = health_loader_fn()
items: list[NotificationItem] = []
now_iso = datetime.now(timezone.utc).isoformat()
# 1. System health alerts (highest priority)
for err_idx, probe_err in enumerate(getattr(health_snap, "probe_errors", ())):
att_cls, req_human = classify_attention_event(
CATEGORY_SYSTEM,
"System Health Probe Error",
probe_err,
is_blocker=True,
)
items.append(
NotificationItem(
id=f"notif-sys-err-{project.id}-{err_idx}",
attention_class=att_cls,
category=CATEGORY_SYSTEM,
title="System Health Error",
summary=f"System health error: {probe_err}",
work_kind="system",
work_number=None,
project_id=project.id,
repo_label=f"{project.gitea_owner}/{project.repo_name}",
created_at=now_iso,
deep_link="/system",
requires_human=req_human,
)
)
for probe in getattr(health_snap, "dependencies", ()):
if probe.status not in ("ok", "healthy"):
att_cls, req_human = classify_attention_event(
CATEGORY_SYSTEM,
f"Probe Failure: {probe.name}",
probe.detail or probe.status,
is_hard_stop=("stop" in probe.status or "fatal" in probe.status),
is_auth_failure=("auth" in probe.name.lower() or "unauthorized" in probe.status.lower()),
is_blocker=True,
)
items.append(
NotificationItem(
id=f"notif-probe-{probe.name}",
attention_class=att_cls,
category=CATEGORY_AUTH if "auth" in probe.name.lower() else CATEGORY_SYSTEM,
title=f"Health Probe Alert: {probe.name}",
summary=f"Probe '{probe.name}' reported status '{probe.status}': {probe.detail}",
work_kind="system",
work_number=None,
project_id=project.id,
repo_label=f"{project.gitea_owner}/{project.repo_name}",
created_at=now_iso,
deep_link="/system",
requires_human=req_human,
)
)
# 2. Queue items (PRs and Issues)
for pr in queue_snap.prs:
if "blocked" in pr.badges:
att_cls, req_human = classify_attention_event(
CATEGORY_BLOCKER,
f"PR #{pr.number} Blocked",
f"PR #{pr.number} '{pr.title}' is blocked or has merge conflicts.",
is_blocker=True,
)
items.append(
NotificationItem(
id=f"notif-pr-block-{pr.number}",
attention_class=att_cls,
category=CATEGORY_BLOCKER,
title=f"Blocked PR #{pr.number}",
summary=f"PR #{pr.number} ({pr.title}) requires merge conflict resolution.",
work_kind="pr",
work_number=pr.number,
project_id=project.id,
repo_label=f"{project.gitea_owner}/{project.repo_name}",
created_at=now_iso,
deep_link=f"/traffic",
requires_human=req_human,
)
)
elif "stale" in pr.badges:
att_cls, req_human = classify_attention_event(
CATEGORY_WORKFLOW,
f"PR #{pr.number} Stale",
f"PR #{pr.number} '{pr.title}' has had no activity for over 14 days.",
is_stale=True,
)
items.append(
NotificationItem(
id=f"notif-pr-stale-{pr.number}",
attention_class=att_cls,
category=CATEGORY_WORKFLOW,
title=f"Stale PR #{pr.number}",
summary=f"PR #{pr.number} ({pr.title}) is stale.",
work_kind="pr",
work_number=pr.number,
project_id=project.id,
repo_label=f"{project.gitea_owner}/{project.repo_name}",
created_at=now_iso,
deep_link=f"/queue",
requires_human=req_human,
)
)
else:
# Routine PR transition
att_cls, req_human = classify_attention_event(
CATEGORY_WORKFLOW,
f"PR #{pr.number} Active",
f"PR #{pr.number} '{pr.title}' is in routine state {', '.join(pr.badges)}.",
)
items.append(
NotificationItem(
id=f"notif-pr-routine-{pr.number}",
attention_class=att_cls,
category=CATEGORY_WORKFLOW,
title=f"Routine PR #{pr.number}",
summary=f"PR #{pr.number} ({pr.title}) state: {', '.join(pr.badges)}.",
work_kind="pr",
work_number=pr.number,
project_id=project.id,
repo_label=f"{project.gitea_owner}/{project.repo_name}",
created_at=now_iso,
deep_link=f"/queue",
requires_human=req_human,
)
)
for issue in queue_snap.issues:
if "duplicate" in issue.badges:
att_cls, req_human = classify_attention_event(
CATEGORY_BLOCKER,
f"Issue #{issue.number} Duplicate PRs",
f"Issue #{issue.number} has multiple linked PRs.",
is_blocker=True,
)
items.append(
NotificationItem(
id=f"notif-issue-dup-{issue.number}",
attention_class=att_cls,
category=CATEGORY_BLOCKER,
title=f"Duplicate PRs on Issue #{issue.number}",
summary=f"Issue #{issue.number} ({issue.title}) linked to multiple PRs.",
work_kind="issue",
work_number=issue.number,
project_id=project.id,
repo_label=f"{project.gitea_owner}/{project.repo_name}",
created_at=now_iso,
deep_link=f"/traffic",
requires_human=req_human,
)
)
elif "claimed" in issue.badges or "in-review" in issue.badges:
att_cls, req_human = classify_attention_event(
CATEGORY_WORKFLOW,
f"Issue #{issue.number} Active",
f"Issue #{issue.number} '{issue.title}' in state {', '.join(issue.badges)}.",
)
items.append(
NotificationItem(
id=f"notif-issue-routine-{issue.number}",
attention_class=att_cls,
category=CATEGORY_WORKFLOW,
title=f"Routine Issue #{issue.number}",
summary=f"Issue #{issue.number} ({issue.title}) state: {', '.join(issue.badges)}.",
work_kind="issue",
work_number=issue.number,
project_id=project.id,
repo_label=f"{project.gitea_owner}/{project.repo_name}",
created_at=now_iso,
deep_link=f"/queue",
requires_human=req_human,
)
)
# 3. Leases / Collisions
for lease in lease_snap.reviewer_leases:
if lease.get("is_expired") or lease.get("status") == "expired":
pr_num = lease.get("pr_number") or lease.get("work_item_number")
att_cls, req_human = classify_attention_event(
CATEGORY_LEASE,
f"Reviewer Lease Expired for PR #{pr_num}",
f"Reviewer lease for PR #{pr_num} has expired.",
is_stale=True,
)
items.append(
NotificationItem(
id=f"notif-lease-exp-pr-{pr_num}",
attention_class=att_cls,
category=CATEGORY_LEASE,
title=f"Expired Reviewer Lease (PR #{pr_num})",
summary=f"Reviewer lease for PR #{pr_num} expired.",
work_kind="pr",
work_number=pr_num,
project_id=project.id,
repo_label=f"{project.gitea_owner}/{project.repo_name}",
created_at=now_iso,
deep_link="/leases",
requires_human=req_human,
)
)
for col_idx, collision in enumerate(lease_snap.duplicate_prs):
att_cls, req_human = classify_attention_event(
CATEGORY_BLOCKER,
f"Duplicate PR Collision ({collision.kind})",
collision.message,
is_blocker=True,
)
issue_part = collision.issue_number if collision.issue_number is not None else "none"
kind_part = (collision.kind or "unknown").replace(" ", "-")
items.append(
NotificationItem(
id=f"notif-collision-{kind_part}-{issue_part}-{col_idx}",
attention_class=att_cls,
category=CATEGORY_BLOCKER,
title=f"Collision Alert ({collision.kind})",
summary=collision.message,
work_kind="issue" if collision.issue_number else "pr",
work_number=collision.issue_number,
project_id=project.id,
repo_label=f"{project.gitea_owner}/{project.repo_name}",
created_at=now_iso,
deep_link="/leases",
requires_human=req_human,
)
)
human_req_count = sum(1 for i in items if i.attention_class == ATTENTION_HUMAN_REQUIRED)
operator_count = sum(1 for i in items if i.attention_class == ATTENTION_OPERATOR)
routine_count = sum(1 for i in items if i.attention_class == ATTENTION_ROUTINE)
# Fetch errors are transport/load failures only — not probe results that
# already surface as first-class notification items (PR #905 review B3).
fetch_err = queue_snap.fetch_error or lease_snap.fetch_error
if isinstance(fetch_err, (tuple, list)):
fetch_err = "; ".join(fetch_err) if fetch_err else None
return NotificationSnapshot(
project_id=project.id,
repo_label=f"{project.gitea_owner}/{project.repo_name}",
items=tuple(items),
human_required_count=human_req_count,
operator_count=operator_count,
routine_count=routine_count,
total_count=len(items),
fetch_error=fetch_err,
)
def snapshot_to_dict(snapshot: NotificationSnapshot) -> dict[str, Any]:
"""JSON-serializable export for /api/v1/notifications."""
return snapshot.as_dict()
+27 -2
View File
@@ -211,6 +211,19 @@ def _pagination_from_pages(
)
_SUPPORTED_FETCH_STATES = ("open", "closed", "all")
def _safe_state(state: str | None) -> str:
"""Constrain a caller-supplied item state before it reaches a query string.
The value is interpolated into the Gitea URL, so an unrecognised state falls
back to ``open`` rather than being passed through.
"""
text = (state or "open").strip().lower()
return text if text in _SUPPORTED_FETCH_STATES else "open"
def _fetch_prs(
host: str,
org: str,
@@ -218,8 +231,15 @@ def _fetch_prs(
auth: str,
*,
per_page: int = 50,
state: str = "open",
) -> tuple[list[dict], PaginationMeta]:
url = f"{repo_api_url(host, org, repo)}/pulls?state=open"
"""Fetch PRs in *state* (``open``, ``closed``, or ``all``).
The queue dashboard only ever wants the open window, so ``open`` stays the
default. The linkage console (#645) widens it, because a landed issue↔PR
edge lives on a merged PR.
"""
url = f"{repo_api_url(host, org, repo)}/pulls?state={_safe_state(state)}"
all_raw: list[dict] = []
pages_fetched = 0
is_final = False
@@ -248,8 +268,13 @@ def _fetch_issues(
auth: str,
*,
per_page: int = 50,
state: str = "open",
) -> tuple[list[dict], PaginationMeta]:
url = f"{repo_api_url(host, org, repo)}/issues?state=open&type=issues"
"""Fetch issues in *state* (``open``, ``closed``, or ``all``); see :func:`_fetch_prs`."""
url = (
f"{repo_api_url(host, org, repo)}/issues"
f"?state={_safe_state(state)}&type=issues"
)
all_raw: list[dict] = []
page = 1
pages_fetched = 0
File diff suppressed because it is too large Load Diff
+164
View File
@@ -0,0 +1,164 @@
"""HTML views for the operator request surface (#643).
The form is deliberately a *preview* form. It has no initiate button, because
initiating requires a confirmed POST to ``/api/v1/requests/apply`` and a stray
form submission must not be able to produce one by accident.
Nothing rendered here is trusted input: every interpolated value is escaped,
and the page renders only values the service already produced rather than
echoing a raw request body back.
"""
from __future__ import annotations
import html
import json
from typing import Any
from webui.layout import render_page
from webui.request_service import (
REQUESTABLE_ROLES,
WORK_KINDS,
RequestError,
RequestPreview,
)
REQUESTS_PATH = "/requests"
PREVIEW_API_PATH = "/api/v1/requests/preview"
APPLY_API_PATH = "/api/v1/requests/apply"
def _escape(text: Any) -> str:
return html.escape(str(text if text is not None else ""), quote=True)
REQUEST_PAGE_STYLES = """
<style>
.request-form { display: grid; gap: 0.75rem; max-width: 44rem; }
.request-form label { display: grid; gap: 0.25rem; font-size: 0.9rem; }
.request-check { margin: 0.35rem 0; }
.request-check .verdict-ok { color: var(--accent); }
.request-check .verdict-fail { color: #d14; }
.request-prohibited code { margin-right: 0.4rem; }
</style>
"""
def _options(values: tuple[str, ...], selected: Any) -> str:
return "".join(
f"<option value='{_escape(value)}'"
+ (" selected" if selected == value else "")
+ f">{_escape(value)}</option>"
for value in values
)
def _form(values: dict[str, Any] | None = None) -> str:
current = dict(values or {})
number = current.get("work_number")
return (
f"<form class='request-form' method='post' action='{REQUESTS_PATH}'>"
"<label>Desired role<select name='desired_role'>"
f"{_options(REQUESTABLE_ROLES, current.get('desired_role'))}"
"</select></label>"
"<label>Work kind<select name='work_kind'>"
f"{_options(WORK_KINDS, current.get('work_kind'))}"
"</select></label>"
"<label>Issue or PR number"
"<input type='number' name='work_number' min='1' required "
f"value='{_escape(number) if number else ''}'></label>"
"<label>Intent summary"
"<input type='text' name='intent_summary' maxlength='500' required "
f"value='{_escape(current.get('intent_summary'))}'></label>"
"<label>Expected head SHA <span class='muted'>(PR work only)</span>"
"<input type='text' name='expected_head_sha' "
f"value='{_escape(current.get('expected_head_sha'))}'></label>"
"<button type='submit' class='copy-btn'>Preview request</button>"
"<p class='muted meta'>Preview is read-only and creates no assignment. "
f"Initiating requires a confirmed POST to <code>{APPLY_API_PATH}</code>."
"</p>"
"</form>"
)
def _checks_block(preview: RequestPreview) -> str:
rows = []
for check in preview.checks:
verdict = "PASS" if check.ok else "FAIL"
css = "verdict-ok" if check.ok else "verdict-fail"
rows.append(
"<li class='request-check'>"
f"<span class='{css}'><strong>{verdict}</strong></span> "
f"<code>{_escape(check.name)}</code> — {_escape(check.detail)} "
f"<span class='muted meta'>({_escape(check.reason_code)})</span>"
"</li>"
)
return "<ul>" + "".join(rows) + "</ul>"
def _preview_block(preview: RequestPreview) -> str:
verdict = "AUTHORIZED" if preview.authorized else "DENIED"
prohibited = "".join(
f"<code>{_escape(action)}</code>" for action in preview.prohibited_actions
)
request = preview.request
evidence = json.dumps(preview.allocator_evidence, indent=2, default=str)
return (
"<h3>Intent preview</h3>"
f"<p><strong>{verdict}</strong> — {_escape(preview.detail)}</p>"
"<p class='meta'>"
f"Role <code>{_escape(request.desired_role)}</code> · "
f"{_escape(request.work_kind)} <code>{_escape(request.display_ref)}</code>"
f" · profile <code>{_escape(preview.required_profile)}</code> · "
f"namespace <code>{_escape(preview.required_namespace)}</code> · "
f"permission <code>{_escape(preview.required_permission)}</code>"
"</p>"
f"<p>Intent: {_escape(request.intent_summary)}</p>"
f"{_checks_block(preview)}"
f"<p><strong>Next safe action:</strong> "
f"{_escape(preview.next_safe_action)}</p>"
"<p class='request-prohibited'><strong>Prohibited for this role:</strong> "
+ (prohibited or "<span class='muted'>none declared</span>")
+ "</p>"
"<p class='muted meta'>Correlation id "
f"<code>{_escape(preview.correlation_id)}</code></p>"
"<details><summary>Allocator evidence</summary>"
f"<pre class='prompt-text'>{_escape(evidence)}</pre>"
"</details>"
)
def _error_block(error: RequestError) -> str:
field = (
f"<p class='meta'>Field: <code>{_escape(error.field_name)}</code></p>"
if error.field_name
else ""
)
return (
"<h3>Request rejected</h3>"
f"<p><strong>{_escape(error.reason_code)}</strong> — "
f"{_escape(error.detail)}</p>{field}"
)
def render_requests_page(
*,
preview: RequestPreview | None = None,
error: RequestError | None = None,
submitted: dict[str, Any] | None = None,
) -> str:
"""Render the request form, plus a preview or rejection when one exists."""
body = (
"<h2>Requests</h2>"
"<p>Submit a work request — desired role, issue or PR, and intent — "
"and see whether it would be authorized before anything is reserved. "
"Initiation goes through the allocator (#600/#613); this console never "
"self-selects work, never approves, and never merges.</p>"
+ _form(submitted)
+ (_error_block(error) if error is not None else "")
+ (_preview_block(preview) if preview is not None else "")
+ f"<p class='meta'><a href='{PREVIEW_API_PATH}'>Preview API</a> · "
"<a href='/api/console/security-model'>RBAC model</a></p>"
+ REQUEST_PAGE_STYLES
)
return render_page(title="Requests", body_html=body)
+47 -20
View File
@@ -270,26 +270,53 @@ def _probe_error_card(snapshot: SystemHealthSnapshot) -> str:
def _recovery_card() -> str:
"""Sanctioned recovery pointers only — never a manual process kill (#630)."""
return (
"<section class='health-card'>"
"<h3>Recovery</h3>"
"<p class='muted'>This dashboard is read-only. Restart and reload "
"controls arrive in Phase 2 (#642); until then recovery runs through "
"the sanctioned client reconnect / operator restart path.</p>"
"<ul class='reasons'>"
"<li><a href='/runtime'>Runtime health</a> — active profile, workflow "
"hashes, and shell health.</li>"
"<li><a href='/sessions'>Runtime and sessions</a> — namespaces, session "
"rows, worktree bindings, and contamination markers (#641).</li>"
"<li>Reconnect the MCP client from the IDE, then re-run the blocked "
"cycle. Never kill the daemon process manually: unmanaged kills are "
"recorded as runtime contamination (#630).</li>"
"<li>See <code>docs/webui-local-dev.md</code> for the documented "
"recovery sequence.</li>"
"</ul>"
"</section>"
)
"""Sanctioned recovery controls & playbooks (#644, Phase 2)."""
try:
from webui import console_recovery
diag = console_recovery.diagnose_recovery()
# Every other card in this file escapes at the interpolation boundary.
# This one did not, and it is where a #630 marker's operator-supplied
# command_summary lands once the contamination gate is wired.
status_badge = (
f"<span class='status-pill {_esc(diag.status)}'>{_esc(diag.status)}</span>"
)
playbook_lis = ""
for pb in diag.playbooks:
elig = "eligible" if pb.eligible else "disabled"
playbook_lis += (
f"<li><strong>{_esc(pb.label)}</strong> "
f"(<code>{_esc(pb.playbook_id)}</code>) — "
f"<span class='badge {elig}'>{elig}</span>: {_esc(pb.description)} "
f"<em class='muted'>({_esc(pb.reason)})</em></li>"
)
reasons_html = ""
if diag.reasons:
items = "".join(f"<li>{_esc(r)}</li>" for r in diag.reasons)
reasons_html = f"<ul class='reasons'>{items}</ul>"
else:
reasons_html = "<p class='clean-note'>No recovery actions currently required. Control plane is healthy.</p>"
return (
"<section class='health-card recovery-card'>"
f"<h3>Sanctioned Recovery Controls (Phase 2 #644) {status_badge}</h3>"
"<p class='muted'>Guided recovery wizard: Diagnose &rarr; Preview &rarr; Confirm &rarr; Verify. "
"Reconnect the MCP client from the IDE, then re-run the blocked cycle. "
"Never kill the daemon process manually: unmanaged kills are recorded as runtime contamination (#630).</p>"
f"{reasons_html}"
"<h4>Available Recovery Playbooks</h4>"
f"<ul class='playbooks-list'>{playbook_lis}</ul>"
"<p class='meta'>APIs: <code>/api/v1/system/recovery/diagnose</code>, "
"<code>/api/v1/system/recovery/preview</code>, <code>/api/v1/system/recovery/apply</code>, "
"<code>/api/v1/system/recovery/verify</code>.</p>"
"</section>"
)
except Exception as exc:
return (
"<section class='health-card'>"
"<h3>Sanctioned Recovery Controls (Phase 2 #644)</h3>"
f"<p class='error'>Recovery diagnostics unavailable: {_esc(exc)}</p>"
"</section>"
)
def render_system_health_page(snapshot: SystemHealthSnapshot) -> str:
+10
View File
@@ -201,6 +201,16 @@ def _candidates_from_queue_snapshot(q_snap: QueueSnapshot) -> list[WorkCandidate
return candidates
def candidates_from_queue_snapshot(q_snap: QueueSnapshot) -> list[WorkCandidate]:
"""Public alias for :func:`_candidates_from_queue_snapshot` (#643).
The request-initiation service ranks the same candidate set this view
renders, so both must agree on how a queue row becomes a candidate. One
construction, two callers not two that can drift apart.
"""
return _candidates_from_queue_snapshot(q_snap)
def _claim_lease_records(inventory: dict[str, Any] | None) -> list[dict[str, Any]]:
"""Normalize ``build_claim_inventory`` entries into lease records.