fix(webui): arm the recovery gates and make the playbooks reach the process (#644)
Reviewer REQUEST_CHANGES on PR #903 at head1c88b87raised five blockers, all reproduced by executing that head. The shared shape: a write path that declared itself gated, audited, and verified, but never armed the gate, mutated a copy of the state it claimed to fix, and then verified against that same copy. B1 - the apply path never asked the execution gate. execute_recovery_playbook called console_authz.authorize with the default for_execution=False, and the phase branch only fires when it is True. ACTIVE_PHASE is 1 and every new action is phase 2, so an operator executed a phase-2 write through POST /api/v1/system/recovery/apply while build_recovery_preview reported execution_enabled false. The call now passes for_execution=True and surfaces the phase_not_active refusal. Preview reports the same decision under execution_authorization / execution_blocked_reason instead of a hardcoded False it could not explain. B2 - both env playbooks mutated a discarded copy and verified against it. source_env = dict(os.environ) meant clear_stale_binding and rebind_session_worktree never touched the running process, and verify_post_recovery(env=source_env) re-diagnosed the same copy, confirming a change that had not happened. Mutations now target the live mapping (apply_recovery's sanctioned env=None -> os.environ path, #702 AC2) and verification re-reads state rather than the mutated input. binding_before / binding_after / binding_changed are returned, and a playbook that changed nothing reports performed: false. verify_post_recovery no longer reads an unverified_inherited binding as clean, because unproven is not clean. B3 - the reconcile playbook called a function that does not exist. merged_cleanup_reconcile.reconcile_merged_cleanups is absent from that module and a bare except turned the AttributeError into a generic failure, so the playbook could never succeed. It now calls gitea_mcp_server.gitea_reconcile_merged_cleanups, the real orchestrator, imported lazily; failures carry error_type. task_capability_map declared gitea.pr.close for reconcile_cleanups while the entry point gates on gitea.read; the two authority statements are reconciled to the one that is enforced. B4 - the #630 contamination integration could not block. assess_contamination_gate was fed marker=None, which short-circuits to block: False on its first statement; the task passed was a console action id outside CONTAMINATION_GATED_TASKS; and the result was read through a "contaminated" key the gate never returns, making STATUS_BLOCKED_CONTAMINATION unreachable. The live marker now comes from the #641 session inventory reader, the gated task key console_recovery_apply is added to CONTAMINATION_GATED_TASKS, every read uses the "block" key the gate actually returns, and the marker is forwarded to sanctioned_restart.execute_restart so a restart cannot launder a contaminated runtime. The reconciler cleanup playbook stays exempt as the designated remedy. B5 - the parity baseline was captured from the head it was compared against. capture_startup_parity(root, head=checkout_head) stores the head verbatim, so in_parity was structurally incapable of being false, and live_remote_head was never passed. The baseline is now the daemon start head that assess_stale_runtime already returns, and the #610 live-remote dimension is restored. Also: _recovery_card was the one renderer in system_health_views.py interpolating without _esc(), and it is where a marker's operator-supplied command_summary lands once B4 is wired; it now escapes, including the except branch. Docs no longer claim apply enforces master parity or that verify asserts clean: true, and the absolute file:///Users/... links are relative. Tests: the two that asserted the defects as intended are inverted - test_api_recovery_apply_with_dev_auth asserted the phase-gate bypass, and the rebind test asserted the input echoed back. Added coverage per blocker, including a no-op detection test that fails when a playbook reports success without changing anything, the previously untested reconcile playbook, contamination block and remedy-exemption tests, and parity baseline/live-remote tests. Both new guards were mutation-verified: disarming for_execution fails 2 tests, restoring the env copy fails 2 tests. Validation: WEBUI_TEST_OFFLINE=1 ../../venv/bin/python -m pytest tests/ -q from branches/feat-issue-644 gives 27 failed / 5291 passed / 6 skipped / 953 subtests; the same command from branches/baseline-master-76f293e at76f293eb28gives 28 failed / 5262 passed / 6 skipped / 926 subtests. Suites run one at a time. comm of the sorted FAILED lines shows no new signature at the head. The single absent signature, test_workspace_guard_alignment.py:: TestRuntimeContextGuardAlignment::test_declared_branches_worktree_passes_when_mcp_root_differs, is suite-order dependent: that file passes 9/9 in isolation at both revisions. Closes #644 Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
This commit is contained in:
+177
-29
@@ -17,7 +17,6 @@ from pathlib import Path
|
||||
from typing import Any
|
||||
|
||||
import master_parity_gate
|
||||
import merged_cleanup_reconcile
|
||||
import runtime_recovery_guard
|
||||
import stale_binding_recovery
|
||||
from webui import console_audit, console_authz, sanctioned_restart, system_health, worktree_scanner
|
||||
@@ -53,6 +52,57 @@ PLAYBOOK_ACTIONS: dict[str, str] = {
|
||||
PLAYBOOK_SANCTIONED_RESTART: sanctioned_restart.ACTION_RESTART_NAMESPACE,
|
||||
}
|
||||
|
||||
#: Task key handed to :func:`runtime_recovery_guard.assess_contamination_gate`.
|
||||
#: A console *action id* is not a task name and is not a member of
|
||||
#: ``CONTAMINATION_GATED_TASKS``, so passing one left the #630 gate inert. Every
|
||||
#: writing recovery playbook shares this one gated task key; the reconciler
|
||||
#: cleanup playbook is exempted separately because it is the designated remedy.
|
||||
CONTAMINATION_GATED_TASK = "console_recovery_apply"
|
||||
|
||||
#: Remote whose contamination markers govern this console. Markers are written
|
||||
#: per remote, so reading the wrong one reports a contaminated runtime clean.
|
||||
REMOTE_ENV = "WEBUI_GITEA_REMOTE"
|
||||
DEFAULT_REMOTE = "prgs"
|
||||
|
||||
|
||||
def _console_remote(env: dict[str, str] | None = None) -> str:
|
||||
env_map = env if env is not None else os.environ
|
||||
return (env_map.get(REMOTE_ENV) or "").strip() or DEFAULT_REMOTE
|
||||
|
||||
|
||||
def load_active_contamination_marker(
|
||||
remote: str | None = None, env: dict[str, str] | None = None
|
||||
) -> dict[str, Any] | None:
|
||||
"""Return the live #630 contamination marker payload, or ``None``.
|
||||
|
||||
The gate is only meaningful when it is fed a real marker: with
|
||||
``marker=None`` :func:`assess_contamination_gate` returns ``block: False``
|
||||
on its first statement. The #641 session inventory already reads the durable
|
||||
markers, so reuse that reader rather than adding a second source of truth.
|
||||
Never raises into a diagnosis or execution path.
|
||||
"""
|
||||
try:
|
||||
from webui import session_loader
|
||||
except Exception: # noqa: BLE001 — never break recovery on an import problem
|
||||
return None
|
||||
try:
|
||||
markers = session_loader._load_contamination_markers(
|
||||
remote=remote or _console_remote(env)
|
||||
)
|
||||
except Exception: # noqa: BLE001 — fail soft; the caller degrades to no marker
|
||||
return None
|
||||
for marker in markers:
|
||||
payload = marker.to_dict()
|
||||
if payload.get("active"):
|
||||
return payload
|
||||
return None
|
||||
|
||||
|
||||
def _active_binding(env_map: Any) -> str | None:
|
||||
"""Read the live worktree binding so a no-op recovery cannot report success."""
|
||||
value = env_map.get(stale_binding_recovery.ACTIVE_WORKTREE_ENV)
|
||||
return value if value else None
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class RecoveryLedgerEntry:
|
||||
@@ -210,9 +260,19 @@ def diagnose_recovery(
|
||||
reasons.append("Runtime HEAD disagrees with checkout/remote HEAD.")
|
||||
|
||||
# 2. Master parity assessment
|
||||
#
|
||||
# The baseline is the commit the *running process* started at, which is what
|
||||
# the parity gate is about. Capturing it from ``checkout_head`` and then
|
||||
# comparing it against that same value made ``in_parity`` structurally
|
||||
# incapable of being false. ``live_remote_head`` restores the #610
|
||||
# live-remote dimension, which was previously dropped.
|
||||
checkout_head = stale_runtime_obj.checkout_head
|
||||
startup_dict = master_parity_gate.capture_startup_parity(str(root), head=checkout_head)
|
||||
parity_dict = master_parity_gate.assess_master_parity(startup_dict, checkout_head)
|
||||
startup_dict = master_parity_gate.capture_startup_parity(
|
||||
str(root), head=stale_runtime_obj.daemon_head
|
||||
)
|
||||
parity_dict = master_parity_gate.assess_master_parity(
|
||||
startup_dict, checkout_head, stale_runtime_obj.remote_head
|
||||
)
|
||||
if not parity_dict.get("in_parity", True):
|
||||
reasons.append("Repository is not in master parity.")
|
||||
|
||||
@@ -245,10 +305,18 @@ def diagnose_recovery(
|
||||
reasons.append("Inherited worktree binding is unverified.")
|
||||
|
||||
# 4. Contamination assessment (#630)
|
||||
#
|
||||
# A real marker and a task key the gate actually gates on: with marker=None
|
||||
# the gate short-circuits to ``block: False``, and with a console action id
|
||||
# the task is outside CONTAMINATION_GATED_TASKS, so it could never block.
|
||||
contamination_marker = load_active_contamination_marker(env=source_env)
|
||||
contamination_dict = runtime_recovery_guard.assess_contamination_gate(
|
||||
marker=None, task=None, actual_role=role_kind
|
||||
contamination_marker,
|
||||
task=CONTAMINATION_GATED_TASK,
|
||||
actual_role=role_kind,
|
||||
)
|
||||
if contamination_dict.get("block"):
|
||||
contaminated = bool(contamination_dict.get("block"))
|
||||
if contaminated:
|
||||
reasons.append("Runtime is contaminated by manual process kill (#630).")
|
||||
|
||||
# 5. Worktree scanner hygiene & anomalies
|
||||
@@ -317,7 +385,7 @@ def diagnose_recovery(
|
||||
)
|
||||
|
||||
# Playbook 4: Sanctioned Restart
|
||||
restart_eligible = bool(stale_runtime_obj.stale or contamination_dict.get("contaminated"))
|
||||
restart_eligible = bool(stale_runtime_obj.stale or contaminated)
|
||||
playbooks.append(
|
||||
PlaybookDescriptor(
|
||||
playbook_id=PLAYBOOK_SANCTIONED_RESTART,
|
||||
@@ -335,8 +403,8 @@ def diagnose_recovery(
|
||||
)
|
||||
)
|
||||
|
||||
clean = not reasons and not contamination_dict.get("contaminated")
|
||||
if contamination_dict.get("contaminated"):
|
||||
clean = not reasons and not contaminated
|
||||
if contaminated:
|
||||
status = STATUS_BLOCKED_CONTAMINATION
|
||||
elif reasons:
|
||||
status = STATUS_ACTION_REQUIRED
|
||||
@@ -374,6 +442,13 @@ def build_recovery_preview(
|
||||
action_id = PLAYBOOK_ACTIONS[playbook_id]
|
||||
action = console_authz.get_action(action_id)
|
||||
decision = console_authz.authorize(action_id, principal)
|
||||
# Preview and apply must answer the same question. ``execution_enabled`` was
|
||||
# a hardcoded False beside an authorization decision taken without
|
||||
# ``for_execution``, so the preview could not tell an operator *why*
|
||||
# execution was disabled — and the apply path did not ask at all.
|
||||
execution_decision = console_authz.authorize(
|
||||
action_id, principal, for_execution=True
|
||||
)
|
||||
phrase = confirmation_phrase(playbook_id, target)
|
||||
ledger = _build_ledger(playbook_id, target)
|
||||
|
||||
@@ -387,8 +462,12 @@ def build_recovery_preview(
|
||||
"confirmation_phrase": phrase,
|
||||
"mutation_ledger": [asdict(entry) for entry in ledger],
|
||||
"authorization": decision.to_dict(),
|
||||
"execution_authorization": execution_decision.to_dict(),
|
||||
"params": dict(params or {}),
|
||||
"execution_enabled": False,
|
||||
"execution_enabled": bool(execution_decision.allowed),
|
||||
"execution_blocked_reason": (
|
||||
None if execution_decision.allowed else execution_decision.reason_code
|
||||
),
|
||||
}
|
||||
|
||||
|
||||
@@ -412,10 +491,15 @@ def execute_recovery_playbook(
|
||||
}
|
||||
|
||||
action_id = PLAYBOOK_ACTIONS[playbook_id]
|
||||
source_env = dict(env) if env is not None else dict(os.environ)
|
||||
# The mapping the playbooks actually mutate. ``dict(os.environ)`` produced a
|
||||
# throwaway copy: every env playbook wrote to it, verified against it, and
|
||||
# left the running daemon bound to the value it claimed to have fixed.
|
||||
mutation_env: Any = env if env is not None else os.environ
|
||||
|
||||
# 1. Authorization check
|
||||
decision = console_authz.authorize(action_id, principal)
|
||||
# 1. Authorization check — ``for_execution=True`` is what arms the phase
|
||||
# gate (console_authz.authorize only applies it in that branch). Without it
|
||||
# a phase-2 write executed while the console was in phase 1.
|
||||
decision = console_authz.authorize(action_id, principal, for_execution=True)
|
||||
if not decision.allowed:
|
||||
console_audit.record_event(
|
||||
action_id=action_id,
|
||||
@@ -458,7 +542,12 @@ def execute_recovery_playbook(
|
||||
|
||||
# 3. Contamination rule (#630) check
|
||||
role_str = principal.role if principal else None
|
||||
contam = runtime_recovery_guard.assess_contamination_gate(marker=None, task=action_id, actual_role=role_str)
|
||||
contamination_marker = load_active_contamination_marker(env=env)
|
||||
contam = runtime_recovery_guard.assess_contamination_gate(
|
||||
contamination_marker,
|
||||
task=CONTAMINATION_GATED_TASK,
|
||||
actual_role=role_str,
|
||||
)
|
||||
if contam.get("block"):
|
||||
if playbook_id != PLAYBOOK_RECONCILE_CLEANUPS:
|
||||
detail = "Runtime is contaminated by a manual process kill (#630). Run reconciler cleanup playbook first."
|
||||
@@ -482,37 +571,76 @@ def execute_recovery_playbook(
|
||||
# 4. Execute playbook action
|
||||
applied_result: dict[str, Any] = {"performed": False}
|
||||
if playbook_id == PLAYBOOK_CLEAR_STALE_BINDING:
|
||||
diagnosis = diagnose_recovery(env=source_env)
|
||||
binding_before = _active_binding(mutation_env)
|
||||
diagnosis = diagnose_recovery(env=env)
|
||||
plan = stale_binding_recovery.plan_recovery(diagnosis.stale_binding)
|
||||
applied_result = stale_binding_recovery.apply_recovery(plan, env=source_env)
|
||||
applied_result = stale_binding_recovery.apply_recovery(plan, env=mutation_env)
|
||||
binding_after = _active_binding(mutation_env)
|
||||
applied_result = {
|
||||
**applied_result,
|
||||
"binding_before": binding_before,
|
||||
"binding_after": binding_after,
|
||||
"binding_changed": binding_before != binding_after,
|
||||
}
|
||||
# A clear that did not clear is not a success, whatever the plan said.
|
||||
if not applied_result["binding_changed"]:
|
||||
applied_result["performed"] = False
|
||||
applied_result.setdefault("reasons", []).append(
|
||||
"clear_stale_binding did not change the live worktree binding"
|
||||
)
|
||||
elif playbook_id == PLAYBOOK_REBIND_SESSION:
|
||||
target_wt = target or (params or {}).get("target_worktree")
|
||||
if target_wt:
|
||||
source_env[stale_binding_recovery.ACTIVE_WORKTREE_ENV] = target_wt
|
||||
binding_before = _active_binding(mutation_env)
|
||||
mutation_env[stale_binding_recovery.ACTIVE_WORKTREE_ENV] = target_wt
|
||||
binding_after = _active_binding(mutation_env)
|
||||
applied_result = {
|
||||
"performed": True,
|
||||
"performed": binding_after == target_wt,
|
||||
"rebound_worktree": target_wt,
|
||||
"cleared_stale": True,
|
||||
"binding_before": binding_before,
|
||||
"binding_after": binding_after,
|
||||
"binding_changed": binding_before != binding_after,
|
||||
"cleared_stale": binding_before != binding_after,
|
||||
}
|
||||
if binding_after != target_wt:
|
||||
applied_result["reasons"] = [
|
||||
"rebind_session_worktree did not take effect on the live "
|
||||
"environment"
|
||||
]
|
||||
else:
|
||||
applied_result = {
|
||||
"performed": False,
|
||||
"reason": "No target_worktree specified for rebind.",
|
||||
}
|
||||
elif playbook_id == PLAYBOOK_RECONCILE_CLEANUPS:
|
||||
# ``merged_cleanup_reconcile`` exposes the building blocks only; the
|
||||
# orchestrator is the MCP tool. The previous call named a function that
|
||||
# does not exist, and a bare ``except`` turned the AttributeError into a
|
||||
# generic failure, so this playbook could never succeed. Imported lazily
|
||||
# because the MCP server module is large and binds FastMCP at import.
|
||||
try:
|
||||
snapshot = merged_cleanup_reconcile.reconcile_merged_cleanups(
|
||||
apply=True, project_root=str(_repo_root())
|
||||
import gitea_mcp_server
|
||||
|
||||
snapshot = gitea_mcp_server.gitea_reconcile_merged_cleanups(
|
||||
dry_run=False,
|
||||
execute_confirmed=True,
|
||||
remote=(params or {}).get("remote") or _console_remote(env),
|
||||
org=(params or {}).get("org"),
|
||||
repo=(params or {}).get("repo"),
|
||||
)
|
||||
performed_reconcile = bool(snapshot.get("success"))
|
||||
applied_result = {
|
||||
"performed": True,
|
||||
"reconciled_count": len(snapshot.get("reconciled") or []),
|
||||
"performed": performed_reconcile,
|
||||
"reconciled_count": len(snapshot.get("entries") or []),
|
||||
"snapshot": snapshot,
|
||||
}
|
||||
except Exception as exc:
|
||||
if not performed_reconcile:
|
||||
applied_result["reasons"] = list(snapshot.get("reasons") or [])
|
||||
except Exception as exc: # noqa: BLE001 — surfaced with its type
|
||||
applied_result = {
|
||||
"performed": False,
|
||||
"error": str(exc),
|
||||
"error_type": type(exc).__name__,
|
||||
}
|
||||
elif playbook_id == PLAYBOOK_SANCTIONED_RESTART:
|
||||
ns = target or (params or {}).get("namespace", "gitea-author")
|
||||
@@ -522,13 +650,18 @@ def execute_recovery_playbook(
|
||||
mode=md,
|
||||
principal=principal,
|
||||
confirmation=f"{md} {ns}",
|
||||
env=source_env,
|
||||
# Without the marker the stricter guard at sanctioned_restart.py:375
|
||||
# never fires and a restart can launder a contaminated runtime.
|
||||
contamination_marker=contamination_marker,
|
||||
env=mutation_env,
|
||||
request_id=request_id,
|
||||
session_id=session_id,
|
||||
)
|
||||
applied_result = restart_res
|
||||
|
||||
performed = bool(applied_result.get("performed") or applied_result.get("allowed"))
|
||||
# ``allowed`` is not ``performed``: execute_restart documents that success is
|
||||
# False in both directions because the host supervisor still has to act.
|
||||
performed = bool(applied_result.get("performed"))
|
||||
|
||||
# 5. Record Audit Log
|
||||
audit_record = console_audit.record_event(
|
||||
@@ -543,8 +676,13 @@ def execute_recovery_playbook(
|
||||
metadata={"applied_result": applied_result},
|
||||
)
|
||||
|
||||
# 6. Post-recovery verification recheck
|
||||
post_verification = verify_post_recovery(env=source_env)
|
||||
# 6. Post-recovery verification recheck.
|
||||
#
|
||||
# Re-read state rather than re-reading the mapping the mutation just wrote:
|
||||
# verifying the mutated copy confirmed changes that never reached the
|
||||
# process. Passing ``env`` through means a caller-supplied mapping is the
|
||||
# live one for that caller, and ``None`` re-reads ``os.environ`` fresh.
|
||||
post_verification = verify_post_recovery(env=env)
|
||||
|
||||
return {
|
||||
"success": performed,
|
||||
@@ -562,12 +700,22 @@ def verify_post_recovery(
|
||||
) -> dict[str, Any]:
|
||||
"""Revalidate control-plane state post-recovery before clean status."""
|
||||
diag = diagnose_recovery(repo_path, env)
|
||||
classification = diag.stale_binding.get("classification")
|
||||
# ``not clear_eligible`` also reads clean for every binding recovery is not
|
||||
# allowed to touch — an unverified inherited binding is unproven, not clean.
|
||||
binding_clean = (
|
||||
not diag.stale_binding.get("clear_eligible")
|
||||
and classification != stale_binding_recovery.CLASSIFICATION_UNVERIFIED_INHERITED
|
||||
)
|
||||
return {
|
||||
"clean": diag.clean,
|
||||
"status": diag.status,
|
||||
"stale_runtime_clean": not diag.stale_runtime.get("stale"),
|
||||
"binding_clean": not diag.stale_binding.get("clear_eligible"),
|
||||
"contamination_clean": not diag.contamination.get("contaminated"),
|
||||
"binding_clean": binding_clean,
|
||||
"binding_classification": classification,
|
||||
# The gate returns ``block``; it has never returned ``contaminated``, so
|
||||
# reading that key reported every runtime clean unconditionally.
|
||||
"contamination_clean": not diag.contamination.get("block"),
|
||||
"anomalies_count": len(diag.worktree_anomalies),
|
||||
"reasons": list(diag.reasons),
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user