Author SHA1 Message Date
sysadmin 1cbbde0089 feat(restart): pre-restart drain proof and hard gate (#661)
Add `drain_proof.py`: a machine-verifiable DrainProof artifact plus a
fail-closed verifier and the hard gate the sanctioned restart-apply path
must consult, so a restart can never proceed on a stale or false "ready"
claim (#655 umbrella, child of #658 coordinator / #659 drain / #660
checkpoints).

- DrainProof: HMAC-SHA256 keyed proof-id over a canonical body using a
  per-process secret -> non-forgeable within the process; a proof minted in
  a prior daemon process will not verify after restart. Short TTL (120s).
- build_drain_proof(): mints the proof from the #658 impact report + the
  drain-mode outcomes. Checklist: no in-flight mutations, assignments
  stopped, checkpoints complete, handoffs ok, leases handled, acks-or-
  timeout. Every check fails closed on missing/ambiguous evidence; the
  no-in-flight-mutations and leases-handled checks are derived from the
  authoritative impact report, not self-reported.
- verify_drain_proof(): fail-closed — rejects missing, malformed, expired,
  signature-mismatched (forged/tampered/prior-process), unclean, or
  stale-fingerprint proofs; recomputes cleanliness from the checks rather
  than trusting the flag.
- gate_apply_restart(): allow only on a valid clean proof; deny -> durable
  incident descriptor; break-glass is the only bypass and is never silent.
- Checkpoint completeness is a supplied input, not a hard dependency on the
  (still-unmerged #660) checkpoint schema.

Wire the gate into gitea_request_mcp_restart: dry_run=False now enforces the
hard gate (drain_proof_json required; break-glass via request_break_glass +
GITEA_BREAKGLASS_RESTART_AUTHORIZATION env). The tool still performs no
actual restart — execution remains a further child.

Tests: tests/test_drain_proof.py — 25 cases covering AC#1-4 (apply without
proof denied, successful drain verifiable, open unsafe mutation fails,
pass/fail/expired), forgery/tamper/wrong-secret/stale-fingerprint rejection,
break-glass bypass, and secret hygiene. 25/25 pass (coordinator suite
unaffected: 40/40 together).

Links #652 #653 #655 #658 #659 #660.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01E7Fv9Bp2XWgvaWa4M1kdR7
(cherry picked from commit e7bcc952bb3e820fda95acbecefeaebfa5f8fcff)
2026-07-24 17:17:13 -04:00
sysadmin 870843f999 Merge pull request 'feat(control-plane): durable MCP session checkpoint schema (Closes #660)' (#881) from feat/issue-660-session-checkpoint-schema into master 2026-07-24 15:52:52 -05:00
jcwalker3 6862049ef6 Merge branch 'master' into feat/issue-660-session-checkpoint-schema 2026-07-24 15:25:39 -05:00
sysadmin 9b5289940c Merge pull request 'fix(author-lock): dead-session recovery cross-session PR base-sync (Closes #872)' (#880) from fix/issue-872-dead-session-two-base-syncs into master 2026-07-24 15:13:56 -05:00
jcwalker3 64e6d7b7df Merge branch 'master' into fix/issue-872-dead-session-two-base-syncs 2026-07-24 15:04:57 -05:00
sysadmin 06c476b37d Merge master into feat/issue-660-session-checkpoint-schema 2026-07-24 16:02:31 -04:00
jcwalker3andClaude Opus 4.8 42657b3b65 docs(control-plane): document versioned session_checkpoints schema (#660)
Adds the session_checkpoints section to the control-plane DB substrate
architecture doc: schema version and row identity, the column groups, the
public API surface, and the hard rules (redaction at write, reconcile
rather than blind restore, unknown live state is not a mismatch, drain
fails closed on an incomplete checkpoint, no transcript storage).

Satisfies acceptance criterion 1 of #660 ("Schema documented and
versioned") for the implementation added in 7e18dcc.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-07-24 14:57:44 -05:00
sysadmin 35714258f0 Merge pull request 'feat(webui): unified session/lease/lock/worktree inventory API (Closes #636)' (#838) from feat/issue-636-inventory-api into master 2026-07-24 13:54:23 -05:00
jcwalker3 4e269f3a7a Merge branch 'master' into feat/issue-636-inventory-api 2026-07-24 13:51:13 -05:00
sysadminandClaude Opus 4.8 87c30484aa Merge master 36fe4785 into feat/issue-660-session-checkpoint-schema
Base-sync of the #660 durable session checkpoint schema branch onto current
master. The only conflicting file was control_plane_db.py, where master added
the #651 usage_events table, migration, and indexes while this branch added the

Resolution keeps both sides in full. Verified against both merge parents: the
resolved file contains every line of master's control_plane_db.py with zero
removals, plus exactly the #660 additions (session_checkpoints table, its two
indexes, _session_checkpoint_row, write_session_checkpoint,
get_session_checkpoint, list_session_checkpoints, reconcile_session_checkpoint,
and checkpoint_completeness).

Conflicts:
	control_plane_db.py

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-07-24 14:19:36 -04:00
sysadmin 36fe4785ec Merge pull request 'feat(mcp): post-restart reconciliation and completion proof (Closes #662)' (#879) from fix/issue-662-post-restart-reconcile into master 2026-07-24 09:04:23 -05:00
jcwalker3 1232789b41 Merge branch 'master' into fix/issue-662-post-restart-reconcile 2026-07-24 09:01:48 -05:00
sysadmin 2976c21ee6 Merge pull request 'test(#878): #628 building-block regression coverage (child of #628)' (#795) from feat/issue-628-autonomous-handoffs-orchestration into master 2026-07-24 08:51:56 -05:00
jcwalker3 2602605c83 Merge branch 'master' into feat/issue-628-autonomous-handoffs-orchestration 2026-07-24 08:50:34 -05:00
jcwalker3 ae31e1e852 Merge pull request 'feat(webui): model usage, token cost, latency, and performance analytics (Closes #651)' (#876) from feat/issue-651-usage-cost-analytics into master
Reviewed-on: #876
Reviewed-by: sysadmin <[email protected]>
2026-07-24 08:46:02 -05:00
jcwalker3 572a3cf8b8 Merge branch 'master' into feat/issue-628-autonomous-handoffs-orchestration 2026-07-24 08:43:12 -05:00
jcwalker3 22f9547fd8 Merge branch 'master' into feat/issue-651-usage-cost-analytics 2026-07-24 08:36:55 -05:00
sysadmin 2e4ed38c51 Merge pull request 'feat(lease): make the author task heartbeat load-bearing (Closes #790, Slice A)' (#794) from fix/issue-790-slice-a-heartbeat-policy into master 2026-07-24 08:26:09 -05:00
sysadminandClaude Opus 4.8 a7a283f449 feat(mcp): post-restart reconciliation and completion proof (Closes #662)
Add pure post_restart_reconcile.reconcile_after_restart classifier with a
machine-readable completion proof covering service health, sessions, leases,
capabilities, worktrees, interrupted mutations (never auto-resumed),
duplicates, and queue state. Soft-depends on #660 checkpoints (skipped with
reason when the schema module is absent).

Wire read-only MCP tool gitea_reconcile_after_restart, boot-once hook via
gitea_assess_master_parity, log_only/enforce modes (mutation_hold), and docs.

Closes #662
Related: #655 #652 #653 #660 #661

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-07-24 08:42:20 -04:00
sysadmin e42756b27f fix(author-lock): dead-session recovery cross-session PR base-sync (Closes #872) 2026-07-24 08:40:19 -04:00
jcwalker3 6f74552617 Merge branch 'master' into feat/issue-628-autonomous-handoffs-orchestration 2026-07-24 07:37:06 -05:00
jcwalker3 f7ef719bd6 Merge branch 'master' into feat/issue-636-inventory-api 2026-07-24 07:36:54 -05:00
sysadmin 38506a753f Merge remote-tracking branch 'prgs/master' into feat/issue-651-usage-cost-analytics
# Conflicts:
#	docs/webui-authz-audit.md
#	webui/console_authz.py
2026-07-24 08:33:35 -04:00
sysadmin 0987f93c67 Merge branch 'master' into fix/issue-790-slice-a-heartbeat-policy
Keep PR #794 current after master landed #863 (sanctioned restart controls).
Preserves Slice A heartbeat conflict-gate remediation for review #544.
2026-07-24 08:22:51 -04:00
sysadmin 103d0df289 Merge pull request 'feat(webui): sanctioned restart and graceful reload controls (Closes #642)' (#863) from feat/issue-642-sanctioned-restart-controls into master 2026-07-24 07:21:17 -05:00
sysadminandClaude Opus 4.8 b179610e7f fix(webui): remediate PR #876 REQUEST_CHANGES for analytics (#651)
Address review blockers and medium/low findings:

- F1: HTML-escape all dynamic analytics fields (html.escape quote=True)
- F2: Gate POST /api/v1/analytics/usage through console_authz
  record_analytics_usage (fail closed for unauthenticated / Phase 1)
- F3: Enforce usage_events retention (max rows + max age)
- F4: Coerce None remote/org/repo to empty strings in load_analytics

Add tests for XSS escaping, unauthorized ingest deny, retention, and
None scope coercion. Document authz + retention in analytics guide.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-07-24 08:11:12 -04:00
jcwalker3 cad5e44703 Merge branch 'master' into feat/issue-636-inventory-api 2026-07-24 07:08:19 -05:00
jcwalker3 4b436ea7d0 Merge branch 'master' into feat/issue-628-autonomous-handoffs-orchestration 2026-07-24 07:08:09 -05:00
jcwalker3 8ad3641fd3 Merge branch 'master' into fix/issue-790-slice-a-heartbeat-policy 2026-07-24 07:07:53 -05:00
sysadminandClaude Opus 4.8 7e18dccbd9 feat(control-plane): durable MCP session checkpoint schema (Closes #660)
Add a versioned, redacted, reconcile-on-boot session checkpoint store to
control_plane_db so a restart can recover session identity, stage, lease
ownership, and next valid action instead of forcing human reconstruction
(umbrella #655; #628 autonomous-handoff goal).

- Bump SCHEMA_VERSION 4->5. New `session_checkpoints` table + two indexes,
  added via `CREATE TABLE IF NOT EXISTS` so table creation is itself the
  additive, idempotent v4->v5 migration (dependency_edges precedent).
- Writer `write_session_checkpoint` upserts the current recoverable state
  per (remote, org, repo, session_id, work_kind, work_number); stage
  transitions audit to `events`. Readers `get_session_checkpoint` /
  `list_session_checkpoints`.
- Every free-text and JSON field is passed through `gitea_audit.redact`
  before storage — no tokens/credential URLs can land in a checkpoint (AC4).
- `reconcile_session_checkpoint` is pure and never restores: it diagnoses a
  stored checkpoint against live head/lease state and flags staleness (AC3).
- Drain gate: `require_complete=True` fails closed (writes nothing) when a
  checkpoint is missing a drain-required field, so drain cannot complete on
  an unrecoverable record.
- `lease_id`/`assignment_id` are soft references (no enforced FK) so a
  checkpoint survives deletion of the lease it names; `work_number=0` is the
  NULL-safe session-level sentinel.

Tests: 13 new cases (schema/version, multi-role fixtures, upsert+audit,
JSON round-trip, secret redaction, reconcile stale-head/dead-lease/
reassigned-lease/clean/unknown, drain fail-closed, sentinel key). Existing
schema_version assertion updated 4->5. Full tests/test_control_plane_db.py
suite: 33/33 pass.

Links #652 #653 #655.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-07-24 08:06:45 -04:00
jcwalker3 205207abb0 Merge branch 'master' into feat/issue-642-sanctioned-restart-controls 2026-07-24 07:03:50 -05:00
jcwalker3 0041542fe2 Merge branch 'master' into feat/issue-651-usage-cost-analytics 2026-07-24 07:03:24 -05:00
sysadmin a87a7d1da2 Merge pull request 'Implement native author issue worktree bootstrap (#850)' (#853) from fix/issue-850-native-mcp-bootstrap into master 2026-07-24 07:02:48 -05:00
sysadmin bc54effbd0 Merge branch 'master' into fix/issue-790-slice-a-heartbeat-policy 2026-07-24 08:01:27 -04:00
jcwalker3 c4d089f931 Merge branch 'master' into feat/issue-636-inventory-api 2026-07-24 06:55:53 -05:00
jcwalker3 9504fa8bbd Merge branch 'master' into fix/issue-850-native-mcp-bootstrap 2026-07-24 06:55:44 -05:00
jcwalker3 18bca47977 Merge branch 'master' into feat/issue-642-sanctioned-restart-controls 2026-07-24 06:55:15 -05:00
jcwalker3 1f144705e2 Merge branch 'master' into feat/issue-651-usage-cost-analytics 2026-07-24 06:55:03 -05:00
sysadmin e1d844bfed fix(bootstrap): path-shaped branches ancestry without isdir (#850 review #551)
resolve_canonical_repo_root fallback now uses commonpath only — no
string split and no os.path.isdir gate — so MCP project_root =
branches/<wt> still resolves to the repo root when the path is not yet
on disk (#274 / review #551 regression).
2026-07-24 07:54:07 -04:00
sysadmin 06e95254f0 fix(bootstrap): address PR #853 REQUEST_CHANGES findings (#850)
- Remove /branches/ string-split fallback in resolve_canonical_repo_root;
  recover roots via commonpath ancestry only (review #531 F2).
- Refuse existing branches that do not contain live master; no weak
  merge-base acceptance (F3).
- Verify caller-supplied assignment_id/lease_id against the control plane
  or fail closed (F4).
- Compensating recovery releases bound workflow leases via lease_lifecycle (F5).
- Regression tests for each finding.
2026-07-24 07:46:34 -04:00
sysadminandClaude Opus 4.8 b6a8989c77 fix(lease): honor heartbeat-lifecycle non-live reclaim bands in same-issue conflict gate (#790 review #502/#516)
assess_same_issue_lease_conflict entered the reclaim branch only under is_lease_expired, so a heartbeat-lifecycle lease that is non-live but whose expires_at is still in the future (stale_absolute_cap under default policy; stale_missed_heartbeat under TTL>grace) fell through to the foreign active-lease block and never reached assess_expired_lock_reclaim, which already permits those bands. The gate now also enters reclaim/renewal when a heartbeat-lifecycle lease is not is_lease_live even with a future expires_at; legacy locks are excluded and keep their absolute expires_at clock (AC-N8). Adds tests/test_issue_794_conflict_gate_reclaim.py.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-07-24 07:38:57 -04:00
sysadmin 3acfb039a9 fix(pr-795): honest scope for #628 building-block regression tests
Address REQUEST_CHANGES on PR #795: stop claiming all 21 umbrella ACs
from a single unit-test module. Scope docs and cases to child issue #878
(CTH/format, ownership classify_skip, dependency edges). Does not close
umbrella #628.
2026-07-24 07:35:26 -04:00
sysadmin a58b5d6e69 Merge remote-tracking branch 'prgs/master' into feat/issue-628-autonomous-handoffs-orchestration 2026-07-24 07:32:47 -04:00
sysadmin 5e935dffb4 Merge pull request 'feat(webui): system-health dashboard (Closes #639)' (#862) from feat/issue-639-webui-system-health-dashboard into master 2026-07-24 06:28:46 -05:00
sysadmin 5c5c1fdf77 feat(webui): model usage, token cost, latency, and performance analytics (Closes #651) 2026-07-24 07:16:44 -04:00
jcwalker3 baf3a474df Merge branch 'master' into feat/issue-636-inventory-api 2026-07-24 06:16:42 -05:00
jcwalker3 0ae05cb9bc Merge branch 'master' into fix/issue-850-native-mcp-bootstrap 2026-07-24 06:16:32 -05:00
jcwalker3 82464f4054 Merge branch 'master' into feat/issue-639-webui-system-health-dashboard 2026-07-24 06:16:12 -05:00
jcwalker3 2a5d6571ec Merge branch 'master' into feat/issue-642-sanctioned-restart-controls 2026-07-24 06:16:03 -05:00
sysadmin c33c69b3f3 Merge pull request 'fix: conflict-fix lease lifecycle chain termination and TTL handling (#842)' (#846) from fix/issue-842-conflict-fix-lease-lifecycle into master 2026-07-24 06:11:09 -05:00
jcwalker3 37c3e5dc39 Merge branch 'master' into feat/issue-642-sanctioned-restart-controls 2026-07-24 06:11:00 -05:00
sysadminandClaude Opus 4.8 784369cc25 merge(master): sync PR #838 with master after #875
No content conflicts. Master advanced with the #658 restart coordinator;
inventory API surfaces auto-merge cleanly.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-07-24 07:07:32 -04:00
sysadminandClaude Opus 4.8 2baf726ee6 merge(master): sync PR #853 with master; keep bootstrap and #860 transitions
Resolve task_capability_map.py by unioning #850 bootstrap_author_issue_worktree
preflight transitions with master's #860 dirty-orphan recovery and work_issue
commit transitions. No feature logic discarded.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-07-24 07:04:20 -04:00
sysadmin 67b4889984 Merge pull request 'feat(mcp-health): MCP restart coordinator and impact analysis (Closes #658)' (#875) from feat/issue-658-mcp-restart-coordinator into master 2026-07-24 06:04:14 -05:00
sysadmin e0b87a0ae5 Merge remote-tracking branch 'prgs/master' into feat/issue-636-inventory-api 2026-07-24 07:01:45 -04:00
sysadmin 34173e079c Merge branch 'master' of https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools into feat/issue-636-inventory-api
# Conflicts:
#	docs/webui-local-dev.md
#	webui/app.py
2026-07-24 07:01:32 -04:00
jcwalker3 d2eaca4949 Merge branch 'master' into feat/issue-642-sanctioned-restart-controls 2026-07-24 05:54:59 -05:00
jcwalker3 fd558ce5d8 Merge branch 'master' into feat/issue-658-mcp-restart-coordinator 2026-07-24 05:54:50 -05:00
sysadmin ae1161524d Merge pull request 'fix(author): bootstrap recovery for dirty orphaned issue worktrees (#860)' (#861) from fix/issue-860-dirty-orphan-worktree-recovery into master 2026-07-24 05:52:24 -05:00
jcwalker3 d456a763fa Merge branch 'master' into feat/issue-658-mcp-restart-coordinator 2026-07-24 05:48:25 -05:00
jcwalker3 cc56aeeacd Merge branch 'master' into fix/issue-860-dirty-orphan-worktree-recovery 2026-07-24 05:22:20 -05:00
sysadmin 44fe8d2eed Merge pull request 'feat(mcp-health): inventory and guard MCP restart/reload/kill paths (Closes #657)' (#870) from feat/issue-657-mcp-restart-path-inventory-guard into master 2026-07-24 04:28:10 -05:00
jcwalker3 d03d982e3b Merge branch 'master' into fix/issue-860-dirty-orphan-worktree-recovery 2026-07-24 04:11:02 -05:00
jcwalker3 657b5bc1b3 Merge branch 'master' into feat/issue-657-mcp-restart-path-inventory-guard 2026-07-24 04:10:08 -05:00
sysadmin b6ca778cef Merge pull request 'fix(reconciler): PR-scoped post-merge cleanup executor + expired reviewer-lease reclaim (Closes #855)' (#866) from fix/issue-855-pr-scoped-merged-cleanup into master 2026-07-24 03:24:53 -05:00
jcwalker3 73f82a2305 Merge branch 'master' into fix/issue-855-pr-scoped-merged-cleanup 2026-07-24 03:15:43 -05:00
sysadmin cc7dc8ac14 Merge pull request 'fix(author): refresh durable issue-lock head after branch sync; recover merge-sync-drifted dead-session locks (Closes #871)' (#874) from fix/issue-871-durable-lock-head-refresh into master 2026-07-24 03:04:24 -05:00
jcwalker3 e151759212 Merge branch 'master' into fix/issue-871-durable-lock-head-refresh 2026-07-24 02:42:52 -05:00
sysadmin 60df5087e9 Merge pull request 'docs(governance): MCP restart governance and authorization policy (Closes #656)' (#857) from docs/issue-656-mcp-restart-governance into master 2026-07-24 02:26:02 -05:00
jcwalker3 78e3befbbb Merge branch 'master' into feat/issue-658-mcp-restart-coordinator 2026-07-24 02:15:32 -05:00
jcwalker3 d7e69fbe77 Merge branch 'master' into feat/issue-639-webui-system-health-dashboard 2026-07-24 02:07:53 -05:00
jcwalker3 4ff4d2acc9 Merge branch 'master' into feat/issue-642-sanctioned-restart-controls 2026-07-24 02:07:44 -05:00
jcwalker3 c9aa09f341 Merge branch 'master' into fix/issue-855-pr-scoped-merged-cleanup 2026-07-24 02:07:33 -05:00
jcwalker3 b4afc8cefd Merge branch 'master' into feat/issue-657-mcp-restart-path-inventory-guard 2026-07-24 02:07:25 -05:00
jcwalker3 0088ecaf00 Merge branch 'master' into docs/issue-656-mcp-restart-governance 2026-07-24 02:07:17 -05:00
jcwalker3 e0536d344f Merge branch 'master' into fix/issue-850-native-mcp-bootstrap 2026-07-24 02:07:03 -05:00
jcwalker3 467e35504c Merge branch 'master' into fix/issue-842-conflict-fix-lease-lifecycle 2026-07-24 02:06:45 -05:00
jcwalker3 bf5c72e3ba Merge branch 'master' into feat/issue-628-autonomous-handoffs-orchestration 2026-07-24 02:06:27 -05:00
sysadminandClaude Opus 4.8 5d59c57c98 fix(author): refresh durable issue-lock head after branch sync; recover merge-sync-drifted dead-session locks (#871)
`gitea_update_pr_branch_by_merge` advanced a PR's remote head but never
advanced the linked durable issue lock's recorded head. After the owning
session died the drifted lock became unrecoverable and no further
synchronization was possible (PR #866 / issue #855).

Write-side (prevents future drift):
- issue_lock_store.assess/apply_durable_lock_head_refresh: on a successful
  sync, CAS-refresh the durable lock's recorded head from the exact expected
  PR head to the resulting head, re-verifying repo/issue/branch/worktree/
  identity/profile/live-session ownership, with read-after-write verification.
- gitea_update_pr_branch_by_merge now refreshes the lock after the remote
  advance and reports a PARTIAL LIFECYCLE FAILURE (success=False) when the
  refresh fails, instead of falsely reporting a full synchronization.

Read-side (recovers already-drifted locks):
- issue_lock_worktree.read_merge_sync_provenance: server-side git observation
  proving a remote head is a sanctioned base-into-branch merge that preserved
  the branch mainline back to the recorded head.
- issue_lock_recovery: new HEAD_RELATION_REMOTE_MERGE_SYNCED accepts a
  dead-session lock whose recorded head is a strict merge-sync ancestor of the
  live PR head — and only that. Rewrites, rebases, force-pushes, non-ancestor
  heads, dirty worktrees, live/competing owners, and wrong repo/issue/branch/
  identity/profile all stay protected.

No existing exact-head, branch-protection, parity, workspace, identity, role,
or mutation-safety gate is weakened. All provenance is server-derived; nothing
is reachable from an MCP caller.

Tests: tests/test_issue_871_durable_lock_head_refresh.py (32 cases) covering
first/second sync, CAS, ownership, partial-failure, merge-sync recovery
happy-path and every fail-closed branch. Full suite: 4789 passed, 13 pre-
existing baseline failures unchanged.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01FZPyVh2DGczQrDxtqwGH5p
2026-07-24 02:56:57 -04:00
sysadminandClaude Opus 4.8 a54b16676a merge(master): sync PR #861 with master; preserve F1–F10
Resolve conflicts from master (#858/#864/#868) while keeping published F10
PID-less lock freshness and dirty-orphan recovery (Issue #860 / PR #861).

Conflict resolutions:
- issue_lock_provenance.py: keep both SOURCE_RECOVER_DIRTY_ORPHANED and SOURCE_DIRTY_SAME_CLAIMANT_REBIND
- task_capability_map.py: keep both recover and rebind task entries
- gitea_mcp_server.py: keep both MCP tools
- issue_lock_store.py: restore F1–F9 helpers + surgical F10 owner_pid cascade

Closes #860

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-07-24 02:31:39 -04:00
sysadminandClaude Opus 4.8 2d0d8a682b feat(mcp-health): add MCP restart coordinator and impact analysis (Closes #658)
Child of umbrella #655 (governed MCP restart coordination); builds on the
#657 restart-path inventory. Adds a central coordinator that evaluates live
control-plane state before a restart and returns a blast-radius impact
preview, so operators and the web console (#642/#652) can see what a restart
would disrupt before concurrent LLM work is destroyed.

Changes
- restart_coordinator.py (new) — pure classification: inventory -> impact
  report DTO (RestartImpactReport/SessionImpact/LeaseImpact). Verdicts:
  safe / unsafe / override. Never restarts anything; fails closed on an
  incomplete inventory.
- control_plane_db.py — additive ControlPlaneDB.list_sessions() read-only
  session inventory (the process-level unit a restart kills).
- gitea_mcp_server.py — new dry-run MCP tool gitea_request_mcp_restart:
  gathers sessions/leases/terminal-lock from the #613 DB, calls the
  coordinator, returns the report. Override authority is read from the
  environment, never self-asserted (#630/#710 F1 pattern). Apply is gated
  by a later drain proof (non-goal here).
- docs/mcp-restart-coordinator.md + docs/mcp-restart-impact-sample.json — doc
  and a real dry-run sample report.
- docs/mcp-tool-inventory.md — register the new tool (inventory sync).
- tests/test_restart_coordinator.py (new) — 15 tests: multi-session fixtures,
  deny-when-critical-section-open, fail-closed deny, override, terminal lock,
  stale heartbeat, JSON-serializable DTO, list_sessions.

Tests: pytest tests/test_restart_coordinator.py -> 15 passed. Full suite:
13 failed / 4753 passed; all 13 reproduce identically on clean master
@ef14622 (0 regressions). The residual test_issue_781 doc-registry failure
is a pre-existing baseline gap for gitea_rebind_dirty_same_claimant_author_session
(merged #864, undocumented on master) — out of scope for #658.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-07-24 02:14:27 -04:00
jcwalker3 eb35c75514 fix(author): remediate test suite regression F10 for Issue #860 (#861) 2026-07-24 01:09:39 -05:00
sysadmin 89657a06c4 Merge pull request 'fix(author): harden dirty-session rebind inventory and journal identity (#868)' (#869) from fix/issue-868-dirty-rebind-inventory-journal into master 2026-07-24 01:00:53 -05:00
jcwalker3 d542b08ced Merge branch 'master' into feat/issue-657-mcp-restart-path-inventory-guard 2026-07-24 00:26:15 -05:00
jcwalker3 499b87c482 Merge branch 'master' into fix/issue-855-pr-scoped-merged-cleanup 2026-07-24 00:25:40 -05:00
sysadmin ba3ea3012c fix(author): harden dirty-session rebind inventory and journal identity (#868)
Complete dirty-inventory revalidation immediately before and after
bind_session_lock so added, removed, or renamed paths fail closed.
Persist and validate full recovery-journal operation identity (remote,
org, repo, claimant identity, claimant profile) on execute, resume,
retry, and already_rebound. Add focused regression coverage and the
reconciler success-path integration test.

Closes #868.
2026-07-24 01:13:50 -04:00
sysadminandClaude Opus 4.8 3428fb4190 feat(mcp-health): inventory and guard MCP restart/reload/kill paths (#657)
Enumerate every code/script/host path that can restart, reload, reconnect,
kill, or force-recreate an MCP process, classify each, and link it to the
guard that constrains it.

- mcp_restart_paths.py: machine-readable registry (single source of truth)
  with classifications (sanctioned_narrow / guarded_fail_closed / forbidden /
  removed / host_residual) plus fail-closed guards:
  * assert_restart_attempt_registered() -- unknown restart attempts fail closed
  * assert_no_daemon_self_replacement() -- daemon never os.execv/os.kill/os._exit
    itself (source-tree scan; comment/docstring mentions ignored)
  * assert_auto_restart_helper_absent() -- keeps the #685-removed
    _trigger_mcp_auto_restart from returning
  * assert_registry_wellformed() -- every path classified, guarded, referenced
- docs/mcp-restart-path-inventory.md: complete inventory table linked from
  #655; documents residual host behaviors (/mcp reconnect) and rollout.
- tests/test_mcp_restart_paths.py: 17 tests -- registry well-formedness,
  unknown-attempt fail-closed, daemon-self-replacement scan (with injected
  violation + comment/docstring negative case), legacy-helper-removed
  regression, pkill-stays-contamination (#630), and doc/module lock-step.

No behavior change to existing modules; regression assertions codify invariants
that already hold (per #657 flag-free-before-hard-block rollout). Links
#652 #653 #655 #656.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-07-24 00:57:33 -04:00
sysadmin ef14622ba0 Merge pull request 'feat(author): dirty-preserving same-claimant author-session rebind (#864)' (#865) from fix/issue-864-dirty-same-claimant-session-rebind into master 2026-07-23 23:55:01 -05:00
sysadminandClaude Opus 4.8 24c52abf6b feat(reconciler): PR-scoped post-merge cleanup executor + expired reviewer-lease reclaim (Closes #855)
Adds a single-target path for post-merge cleanup so a reconciler can
complete one merged PR without a batch sweep across unrelated PRs.

## PR-scoped selector

gitea_reconcile_merged_cleanups gains an optional pr_number. When set,
only that merged PR is assessed and acted on: the PR is resolved live and
fails closed on an invalid/non-positive number, an unresolvable or
ambiguous PR, or an unmerged PR; reviewer scratch worktrees are filtered
to that PR; and the report's entry set is pinned to exactly [pr_number],
failing closed on any drift. The existing execute loop then operates on
the single pinned entry only -- worktree removal, ownership reassessment,
then remote-branch delete -- with no unrelated target. Batch behaviour is
unchanged when pr_number is omitted.

## Expired reviewer-lease reclaim (AC4)

An expired or stale reviewer lease no longer protects an already-merged
branch forever. branch_cleanup_guard.assess_expired_reviewer_lease_reclaim
makes the decision explicitly and fail-closed: reclaim only when the lease
is a reviewer lease, its status is expired/stale, the PR is proven merged,
the owner process is proven dead, and no competing active claimant uses
the branch. _collect_branch_ownership_records supplies that evidence from
authoritative state (live PR merged-state, lease owner liveness, and the
full ownership inventory for competing-claimant detection) and evaluates
it only after the complete inventory is built, so the post-worktree-removal
reassessment is what unblocks the branch delete. Any unknown fails closed.

Author/merger/controller/reconciler leases are untouched; active leases,
worktree bindings, issue locks, and live sessions still block.

## Tests

- tests/test_issue_855_expired_reviewer_reclaim.py: full fail-closed matrix
  for the reclaim decision plus collector wiring (merged+dead+uncontested
  reclaims; unmerged, live-owner, competing-worktree, and author-lease
  cases stay protective).
- tests/test_branch_cleanup_guard.py: exact-PR selector coverage (ignores
  newer PRs in the batch queue, execute mutates only the selected PR,
  unknown/not-merged/invalid fail closed, batch mode preserved).

Changed-surface suites pass; the 2 pre-existing test_branch_cleanup_guard
failures and test_reconciler_supersession_close reproduce identically on
master 6d0015ca and are unrelated to this change.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-07-24 00:49:52 -04:00
jcwalker3 a3f8f67c93 feat(author): add dirty-preserving same-claimant author-session rebind (Closes #864)
Introduce an explicit, fail-closed recovery operation that rebinds a live
author session to an already-registered dirty issue worktree when the durable
lock still belongs to the same identity/profile and the recorded owner PID is
provably dead. Ordinary locking remains clean-worktree-only; this path updates
only stale lock/session provenance, preserves every dirty byte under fingerprint
pins, and grants create-PR-sanctioned provenance without remote sync or recovery
worktrees.

Also stop treating dead-owner session pointers as live mutation workspace
bindings so a stale dead-owner pointer cannot poison unrelated author work.
2026-07-23 22:54:15 -05:00
jcwalker3 0b60fd6557 Remediate PR #861 findings F1-F9 and TestWorktreeStart failures (#860)
- Fix F1: Prepare recovery worktree detached at remote_head without git checkout -B to avoid exit 128 when branch is held by source worktree
- Fix F2: Pass recovery_sanctioned=True in bind_session_lock and assess_same_issue_lease_conflict
- Fix F3: Add SOURCE_RECOVER_DIRTY_ORPHANED to SANCTIONED_LOCK_SOURCES
- Fix F4: Stop after Phase 4 dirty apply when conflicts exist; do not finalize session binding
- Fix F5: Dynamically query competing live locks and workflow leases in MCP server
- Fix F6: Fail closed on recovery worktree resume when HEAD does not match expected remote_head
- Fix F7: Fail closed on remote HEAD observation failure rather than copying expected_remote_head pin
- Fix F8: Enforce foreign overwrite protection requiring same claimant or sanctioned reclaim
- Fix F9: Add real multi-worktree integration tests for prepare_recovery_worktree and lock rebind
- Fix TestWorktreeStart: Bypass session lock check for dry-run and review/pr-* branches in scripts/worktree-start
2026-07-23 22:34:35 -05:00
jcwalker3 fc8fe329d2 Merge branch 'master' into feat/issue-642-sanctioned-restart-controls 2026-07-23 22:22:36 -05:00
sysadminandClaude Opus 4.8 e593444eea feat(webui): sanctioned restart and graceful reload controls (Closes #642)
Adds a gated `system.restart_namespace` action so operators and workers can
restart or gracefully reload MCP namespaces through an authorized, audited
path instead of manual host process killing (#630).

- webui/sanctioned_restart.py: restart/reload operation model, dry-run
  intent preview, confirmation-string enforcement, audit emission, and
  post-restart health verification. Fails closed on unknown auth, missing
  capability, or ambiguous target namespace.
- webui/console_authz.py: RBAC entries for the restart capability with
  secret redaction preserved.
- webui/gated_actions.py: registers the restart action in the gated action
  framework so it cannot be invoked without capability + confirmation.
- task_capability_map.py: capability mapping for the restart operation.
- docs/sanctioned-restart-controls.md: operator documentation for the
  sanctioned path and the explicit prohibition on pkill recovery.
- docs/webui-authz-audit.md: audit model updated for restart events.
- tests/test_webui_sanctioned_restart.py: authorized preview, unauthorized
  deny, confirmation enforcement, contamination classification, and audit
  emission coverage.

No unrestricted kill path is exposed; manual pkill remains classified as
contamination and continues to block clean claims.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-07-23 23:19:21 -04:00
sysadmin 6d0015cabc Merge pull request 'fix(worktree-audit): make cleanup audit merged-PR aware for issue worktrees (Closes #858)' (#859) from fix/issue-858-audit-merged-pr-aware into master 2026-07-23 22:02:22 -05:00
jcwalker3 67cd2da561 fix: remediate PR #853 review #528 findings for native MCP bootstrap (#850)
- Fix module reloading bug in task capability router (F-1)
- Harden journal persistence and pending creations crash window (F-3)
- Implement dirty worktree and author commit recovery preservation (F-4)
- Fail closed on missing identity, profile, or session parameters (F-5)
- Fix branches root path traversal and symlink validation (F-6)
- Enforce O_NOFOLLOW and symlink checking on transition locks (F-7)
- Support common ancestor merge-base verification for base SHA (F-8)
- Release transition lock on compensating recovery (F-10)
- Thread journal_dir through recovery and fix guidance strings (F-11, F-12)
- Fix unittest mock import in bootstrap test suite (F-13)
2026-07-23 20:40:09 -05:00
jcwalker3andGrok 4.5 18d6583e83 fix(author): bootstrap recovery for dirty orphaned issue worktrees (#860)
Add an explicit recovery operation for same-claimant dirty registered
worktrees under malformed PID-less durable locks, with crash-safe journals,
dirty byte preservation, path-level conflict detection, and live session
binding. PID-less locks are never treated as live merely because expiry is
absent.

Closes #860

Co-Authored-By: Grok 4.5 (xAI) <[email protected]>
2026-07-23 20:38:47 -05:00
sysadminandClaude Opus 4.8 fe259e6d38 fix(worktree-audit): make cleanup audit merged-PR aware for issue worktrees (Closes #858)
gitea_audit_worktree_cleanup had no PR linkage. Issue worktrees therefore
reported pr_number=null and classified as active_issue_work with
removable=false permanently, even once their PR was merged and the head was
already contained in master. Observed on live master 9301739910: 42
issue_work worktrees, 0 removable, 0 with pr_number populated, while
gitea_reconcile_merged_cleanups reported the same worktree safe to remove.

Two independent gaps caused it:

* build_worktree_metadata was never given a pr_number, and only open PRs were
  fetched, so no owning-PR evidence existed at all.
* clean_stale_removable was unreachable for issue_work: it required
  ttl_expired, derived from a last_used_at that nothing populates, and
  is_ttl_expired fail-safes to False when the timestamp is unknown.

This adds deterministic merged-PR linkage and gates removal on the complete
cleanup policy:

* build_pr_index / resolve_owning_pr link a worktree branch to exactly one
  owning PR. Competing PRs on one branch, a still-open owner, a head-branch
  mismatch, or missing PR state all fail closed while still reporting the
  resolved pr_number.
* assess_merged_pr_worktree_cleanup requires all of: conclusive merged
  ownership, branch agreement, containment of the head in authoritative
  master, no open/competing PR, no active lease, no issue lock, no live
  session, a clean tree, and a non-protected checkout. Unknown state blocks.
* Containment reuses merged_cleanup_reconcile.is_head_ancestor_of_ref so the
  audit and the PR-scoped reconciler agree on what "already landed" means.

Lease evidence is now supplied. audit_branches_directory already accepted
leased_branches but the MCP tool never passed it, so has_active_lease was
false for every worktree in a live run. That was inert only while issue
worktrees could never become removable; it is wired to authoritative
control-plane leases here, scoped so a lease on issue N protects that issue's
work worktree and not a baseline or review tree merely named after it.

Issue work no longer becomes removable on TTL age alone, since age is not
proof that a branch landed and would otherwise reclaim a worktree holding
unmerged commits. conflict_fix keeps its existing TTL behaviour, and review,
baseline, merge-simulation, detached, dirty, open-PR, and protected
classifications are unchanged.

The assessor still performs no deletion and gains no cleanup mutation. This
is assessor-side only and does not implement the PR-scoped executor or the
expired-lease reclaim policy tracked separately by #855.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-07-23 21:19:32 -04:00
sysadmin f80e3b33b0 Merge remote-tracking branch 'prgs/master' into feat/issue-639-webui-system-health-dashboard 2026-07-23 21:14:50 -04:00
jcwalker3 dc99c15ffa Merge branch 'master' into docs/issue-656-mcp-restart-governance 2026-07-23 19:53:16 -05:00
jcwalker3 e9f6d68bd7 Merge branch 'master' into fix/issue-850-native-mcp-bootstrap 2026-07-23 19:53:00 -05:00
jcwalker3 3a0d9e24ea Merge branch 'master' into fix/issue-842-conflict-fix-lease-lifecycle 2026-07-23 19:52:51 -05:00
jcwalker3 5adc328a2b Merge branch 'master' into feat/issue-628-autonomous-handoffs-orchestration 2026-07-23 19:52:36 -05:00
jcwalker3 6a636c58e7 Merge branch 'master' into fix/issue-790-slice-a-heartbeat-policy 2026-07-23 19:52:27 -05:00
sysadmin 9301739910 Merge pull request 'feat(webui): workflow-event and conversation timeline model (Closes #637)' (#849) from feat/issue-637-timeline-model into master 2026-07-23 19:33:22 -05:00
jcwalker3 b3859f6dad Merge branch 'master' into fix/issue-850-native-mcp-bootstrap 2026-07-23 19:13:38 -05:00
jcwalker3 dc0a05e5e9 Merge branch 'master' into fix/issue-842-conflict-fix-lease-lifecycle 2026-07-23 19:13:22 -05:00
jcwalker3 9cca5f3dd2 Merge branch 'master' into feat/issue-628-autonomous-handoffs-orchestration 2026-07-23 19:13:09 -05:00
jcwalker3 d2fe0110a0 Merge branch 'master' into fix/issue-790-slice-a-heartbeat-policy 2026-07-23 19:13:01 -05:00
sysadminandClaude Opus 4.8 edd5f813b2 Merge master into feat/issue-639-webui-system-health-dashboard
Resolve the #638 shell landing against the #639 dashboard:

- webui/layout.py: drop the flat NAV_ITEMS tuple in favor of master's
  grouped NAV_GROUPS nav-config module.
- webui/nav.py: register /system-health as a live item in the Health
  group, satisfying issue #639 AC5 through the canonical nav source.
- docs/webui-local-dev.md: keep both additive sections (#638 shell and
  #639 dashboard).
- tests/test_webui_system_health_dashboard.py: assert the nav entry via
  iter_nav_items() instead of the removed NAV_ITEMS tuple.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-07-23 19:56:57 -04:00
sysadminandClaude Opus 4.8 9f759150b8 docs(governance): correct linked-issue descriptions in restart ADR (#656)
The Related section and cross-reference lines described #652, #653, #630, #642,
and #591 by roles they do not hold. Align each description with the linked
issue's actual title and scope:

- #652 is the Control Plane Web Console product vision (restart controls live in
  its capability area A), not a restart-specific vision.
- #653 is the console phased-delivery roadmap; restart controls are Phase 2.
- #630 is the manual process-kill contamination guard, not the coordinator; the
  coordinator remains an unimplemented later child of #655.
- #642 is the sanctioned restart / graceful reload console UX.
- #591 is auto-restart on master advance (closed); only #584 is transport-flap
  reconnect. They were previously collapsed into one transport-recovery label.

Documentation-only wording change. Policy IDs RG-01..RG-08, the policy version
restart-governance/v1, the authorization matrix, and every normative statement
are unchanged.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-07-23 19:18:47 -04:00
sysadmin 347464a057 Merge branch 'master' into docs/issue-656-mcp-restart-governance 2026-07-23 19:17:05 -04:00
sysadminandClaude Opus 4.8 1301a57de4 docs(governance): MCP restart governance and authorization policy (#656)
Adds docs/architecture/mcp-restart-governance.md, the restart-governance/v1 ADR
defining who may restart the MCP control plane and under what conditions.

- Recovery ladder (reconnect -> rebind -> scoped restart -> full restart -> host)
  with restart stated as the last resort.
- Authorization matrix across author/reviewer/merger/reconciler/controller/
  operator/admin; no LLM worker role may perform or authorize a full or host
  restart.
- v1 authority decision recorded: controller approval + automated safety gates;
  quorum deferred to a superseding ADR.
- Break-glass path with pre-declared incident and mandatory post-hoc audit.
- Ambiguous policy state denies restart.
- Stable policy IDs RG-01..RG-08 for later enforcement code to bind to.

Cross-links the ADR from docs/safety-model.md and docs/webui-deployment.md, and
adds tests/test_mcp_restart_governance_docs.py asserting acceptance criteria 1-5.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-07-23 19:16:28 -04:00
jcwalker3 3b2b4e1dca Remediate PR #853 in response to review #525 for Issue #850 2026-07-23 17:49:31 -05:00
jcwalker3 b9ba43a5bf Merge branch 'master' into feat/issue-637-timeline-model 2026-07-23 17:47:29 -05:00
sysadminandClaude Opus 4.8 a1e5a4af8c fix(webui): validate externally influenced event_type at the timeline boundary (#637)
Review #526 found that event_type reached the serialized timeline payload
without crossing the redaction/validation boundary every other free-text
field on the same event crosses. A synthetic secret-shaped 40-hex canary
was redacted through message but survived verbatim through event_type on
the same control-plane record.

CTH heading path: CTH_TYPES is now the single authority for what a CTH
type may be. canonical_thread_handoff.is_known_cth_type() is that
authority, used by format_cth_body (write), assess_cth_comment (assess),
and now the read path too; parse_cth_comment reports membership as
cth_type_known and stays total. adapt_cth_comments serializes
handoff:<type> only for a declared type and otherwise emits the constant
handoff:unrecognized, so arbitrary, malformed, secret-shaped, or
whitespace-manipulated heading content never becomes an event_type.

Control-plane path: a stored event_type is treated as source data.
_safe_cp_event_type accepts only an ordinary identifier that is not a
bare secret-shaped hex run and that a redaction pass leaves unchanged;
anything else fails closed to the constant unsafe:redacted and marks the
event sensitive. The value is never emitted verbatim, never partially
sanitized, and never rewritten into a different valid-looking type.

Remaining serialized-field audit: event_key ids must be plain numeric
identifiers, the CTH adapter refuses a scope it cannot express, per-source
failure reasons are redacted (they can quote an authenticated fetch error),
and echoed scope/filter values are guarded so reflection is not a bypass.

Legitimate values are preserved: every declared CTH type, the real
producer types (assigned, lease_released, lease_adopted,
dependency_edge_state_change, allocation, pr.opened, lease.renew), and
the existing issue, PR, evidence, and full-SHA references.

Tests seed the canary independently through both event_type paths with a
benign message, so message redaction cannot be why they pass; each
inspects event_type directly, asserts the canary is absent from the
complete serialized payload, and asserts scan_for_secrets finds nothing.

tests/test_webui_timeline.py 64 passed, 15 subtests
pytest -k webui 362 passed, 285 subtests
pytest -k redact 79 passed
tests/test_canonical_thread_handoff.py tests/test_control_plane_db.py 29 passed

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-07-23 18:45:54 -04:00
sysadmin 188e83c4d6 Merge pull request 'fix: remove safe worktrees before reassessing remote delete ownership (#851)' (#852) from fix/issue-851-cleanup-worktree-before-remote-delete into master 2026-07-23 17:06:25 -05:00
sysadmin db5ed6042b Merge pull request 'feat(webui): evolve application shell for Phase 1 console IA (Closes #638)' (#818) from feat/issue-638-webui-app-shell-phase1 into master 2026-07-23 17:04:32 -05:00
sysadmin a942afe6c4 Implement native author issue worktree bootstrap (#850) 2026-07-23 17:28:23 -04:00
jcwalker3 0254a99336 Merge branch 'master' into feat/issue-628-autonomous-handoffs-orchestration 2026-07-23 16:21:56 -05:00
jcwalker3 15c75d2225 Merge branch 'master' into feat/issue-638-webui-app-shell-phase1 2026-07-23 16:21:31 -05:00
jcwalker3 8b34f9da0a Merge branch 'master' into fix/issue-842-conflict-fix-lease-lifecycle 2026-07-23 16:21:17 -05:00
jcwalker3 2d95e0fcc6 Merge branch 'master' into feat/issue-637-timeline-model 2026-07-23 16:21:09 -05:00
sysadminandClaude Opus 4.8 8ba1c5b87c fix(webui): make timeline session filter truthful and redact evidence refs (#637)
Remediates the two blocking findings in review #522 on PR #849.

F1 — the session filter dimension was dead end to end. No source could
produce an event carrying a session identifier, so filter_events dropped
every event whenever session was supplied and the API answered with a
green, empty page. An empty result reads to an operator as "no such
session activity", which is a stronger and false claim.

Each source now declares which filter dimensions its records can actually
carry. The control-plane events table is (event_id, work_item_id,
event_type, message, created_at) and records no session, so that source
declares the session dimension unsupported rather than pretending to
answer it; the dead read of a non-existent session_id column is removed.
A CTH handoff comment declares its own Session field, so the handoff
adapter populates session_id from that declared field — authoritative
source data, never inferred from an actor, work item, or message text.

When no source that ran can carry a requested dimension, load_timeline
refuses with ok=false and a structured error naming the unsupported
filters and the per-source reason, and the route answers 422. A source
that can answer the dimension and simply matched nothing still returns
200 with an honest empty page. Ordering, pagination, and the issue/PR
filters are unchanged.

F2 — evidence_refs bypassed redaction and could emit a credential
verbatim. proof and decision were passed to _extract_evidence_refs before
redaction, and the SHA pattern matched any 7-40 character lowercase hex
run, which is exactly the shape of a Gitea access token.

Redaction now runs first and every derived value is taken from the
redacted text. A commit reference is recognised only where the source
text declares one (commit, head, base, sha, ...), so an undeclared hex
run is never lifted out of prose into a structured field; this also drops
the ordinary-word noise the reviewer noted. Every reference is then
independently revalidated against an allowed shape and a second redaction
pass immediately before serialization, failing closed by dropping
anything unproven and flagging the event sensitive. Actor is redacted for
the same reason, and a secret-shaped session value is dropped rather than
emitted. Legitimate issue, PR, short-SHA and full 40-character SHA
references stay usable.

Tests: 16 added. Session filtering is now driven through the CTH adapter
and the composed load_timeline/API path rather than a hand-built
WorkflowEvent, covering a match, an honest empty result, pagination and
ordering under the filter, the 422 refusal, and the per-source support
declaration. Redaction coverage asserts a synthetic 40-character hex
value (not a real credential) appears nowhere in the complete serialized
payload including evidence_refs, that the independent revalidation drops
unproven references, and that legitimate references still resolve.

pytest tests/test_webui_timeline.py: 42 passed (was 26).
pytest -k webui: 340 passed, 270 subtests passed (was 324).
Full suite: 4607 passed, 12 failed, 6 skipped, 684 subtests passed — the
same 12 failures as the master baseline, none under webui/.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-07-23 16:45:43 -04:00
sysadminandClaude Opus 4.8 df58b5fb90 fix: remove safe worktrees before reassessing remote delete ownership (#851)
Post-merge cleanup previously continued past ownership-blocked remote
deletes, which skipped independently safe local worktree removal when the
only block was worktree_binding. Remove the clean owned worktree first,
reassess ownership, then delete the remote branch only if still safe.
Preserve fail-closed protection for dirty/foreign ownership categories.

Closes #851

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-07-23 16:41:27 -04:00
sysadminandClaude Opus 4.8 ecda200180 feat(webui): system-health dashboard (Closes #639)
Phase 1 child of the Web Console epic #631. Adds the operator-facing
system-health dashboard on top of the read-only system-health API landed
by #634, so runtime problems are visible on a surface instead of being
discovered late through failed LLM sessions.

- webui/system_health_views.py (new): renders the SystemHealthSnapshot as
  readiness, stale-runtime parity, version/uptime, dependency, MCP
  namespace, probe-error, and recovery cards.
- webui/app.py: GET /system-health, sharing load_system_health() with the
  JSON API so page and API cannot disagree. ?deep=1 behaves as on the API.
- webui/layout.py: nav entry and health card/badge styles.
- tests/test_webui_system_health_dashboard.py (new, 26 cases).
- docs/webui-local-dev.md: route, field authority, and redaction split.

Readiness honesty is preserved from the API: ready and readiness_complete
render separately, a probe that did not run is listed under "Not probed"
rather than counted healthy, and mutation safety is never claimed when the
runtime is stale or parity is indeterminate.

Redaction is split by field kind. Free text (probe details, reasons, probe
errors) passes through system_health.redact. Structured fields (commit
SHAs, probe names, statuses, timestamps) are HTML-escaped only: redact's
opaque-token rule matches any run of 32 or more characters, so routing a
40-character git SHA through it rendered "[redacted]" and blanked the
parity evidence the page exists to show.

Non-goals honored: no restart or reload controls (Phase 2, #642), no
manual process-kill guidance (#630). Read-only throughout.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-07-23 16:08:36 -04:00
sysadmin b70d5f3efa Merge pull request 'fix: make cross-role allocations consumable by independent workers (Closes #843)' (#845) from fix/issue-843-cross-role-allocation-handoff into master 2026-07-23 14:57:45 -05:00
sysadmin fa6ba8a162 Merge pull request 'fix: exclude epic and child-only containers from allocator selection (Closes #844)' (#848) from fix/issue-844-exclude-epic-containers into master 2026-07-23 14:56:53 -05:00
sysadmin a20975688d Merge branch 'master' into feat/issue-637-timeline-model
# Conflicts:
#	webui/app.py
2026-07-23 14:53:12 -04:00
sysadminandClaude Opus 4.8 25bc2a3291 feat(webui): workflow-event and conversation timeline model (Closes #637)
Phase 1 child of the Web Console epic #631. Adds a durable, versioned
WorkflowEvent schema with per-source adapters and a read-only query API so
operators can browse a unified timeline of workflow events, decisions, tool
calls, and handoffs instead of scattered evidence.

- webui/timeline.py (new): versioned WorkflowEvent schema; control-plane
  event adapter and Gitea CTH handoff-comment adapter; read-only mode=ro
  control-plane reader; conjunctive filter by issue/PR/session; stable
  (timestamp, source_rank, event_key) ordering; bounded pagination;
  fail-soft per-source status; redaction at the boundary, fail closed.
- webui/app.py: GET /api/v1/timeline read-only route with thread-scoped,
  fail-soft handoff comment source.
- tests/test_webui_timeline.py (new): schema, adapters, redaction of
  secret-like payloads, filter/sort/pagination, scoped CP reader,
  fail-soft composition, and API integration.
- docs/webui-local-dev.md: timeline route and field-authority notes.

Read-only Phase 1; no mutation of historical events; no full chat replay;
no unredacted tool-argument storage.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-07-23 14:48:46 -04:00
jcwalker3 f1e4809930 Merge branch 'master' into fix/issue-790-slice-a-heartbeat-policy 2026-07-23 12:31:44 -05:00
jcwalker3 0752b5d242 Merge branch 'master' into feat/issue-628-autonomous-handoffs-orchestration 2026-07-23 12:31:35 -05:00
jcwalker3 c040bd4674 Merge branch 'master' into feat/issue-638-webui-app-shell-phase1 2026-07-23 12:31:27 -05:00
jcwalker3 f0c9ffb25e Merge branch 'master' into fix/issue-843-cross-role-allocation-handoff 2026-07-23 12:31:10 -05:00
jcwalker3 04d9df559e Merge branch 'master' into fix/issue-842-conflict-fix-lease-lifecycle 2026-07-23 12:30:59 -05:00
jcwalker3 79256f9093 Merge branch 'master' into fix/issue-844-exclude-epic-containers 2026-07-23 12:30:51 -05:00
sysadmin 1c455b6ec0 Merge pull request 'feat(webui): read-only system-health API (Closes #634)' (#813) from feat/issue-634-readonly-system-health-api into master 2026-07-23 04:14:33 -05:00
sysadminandClaude Opus 4.8 c3f282ba44 fix: exclude epic and child-only containers from allocator selection (Closes #844)
Epics and parent issues whose body delegates implementation to children are
skipped before ranking with structured reason epic_or_child_only_container.
Title-only "epic" mentions without body/label evidence remain eligible.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-07-23 04:56:58 -04:00
sysadmin 8a63476787 fix: conflict-fix lease lifecycle chain termination and TTL handling (Closes #847, Refs #842) 2026-07-23 04:46:30 -04:00
sysadmin 5eb89f8830 fix: bind cross-role handoff consume role to authenticated profile (Closes #843)
Review #515 F1: gitea_adopt_workflow_lease trusted a caller-supplied role
((role or active_role)), so any namespace holding gitea.read could consume
an author-only cross-role handoff by passing role="author".

- Derive the adopter role authoritatively from the active profile; reject
  any supplied role that does not exactly match (no silent accept).
- Pass the profile-derived role and authoritative profile/namespace context
  to lease_lifecycle.adopt_lease; validate handoff provenance
  required_profile/required_namespace against it (fail closed).
- Fail closed when the profile role cannot be derived (no author default).
- Add MCP-boundary regression tests: reviewer/merger profiles cannot
  consume an author handoff via role="author"; the legitimate author
  profile still consumes; foreign required_profile rejected.
2026-07-23 03:41:28 -04:00
sysadminandClaude Opus 4.8 a6c15afec1 fix: make cross-role allocations consumable by independent workers (Closes #843)
Controller-created role=author allocations were owned by the allocating
controller session with no authorized consume path for independent author
workers. When the controller exited, the lease became stale_dead_process
and required abandon/reassign instead of a usable handoff.

- Mark cross-role apply with durable handoff provenance (pending)
- Allow gitea_adopt_workflow_lease to consume pending handoffs by the
  required role without sharing controller session identity or requiring
  the controller process to remain alive
- Atomically transfer assignment+lease ownership and set
  adopted_by_session_id with read-after-write evidence
- Reject wrong-role, second, and terminal adoptions
- Surface consume_allocation identifiers in process_work_queue results
- Preserve same-role allocation and genuine abandon recovery behavior

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-07-23 02:38:19 -04:00
jcwalker3 dc1d0e045f Merge branch 'master' into fix/issue-790-slice-a-heartbeat-policy 2026-07-23 01:13:11 -05:00
jcwalker3 6d4a0d12ec Merge branch 'master' into feat/issue-628-autonomous-handoffs-orchestration 2026-07-23 01:12:59 -05:00
jcwalker3 6868b345ee Merge branch 'master' into feat/issue-634-readonly-system-health-api 2026-07-23 01:12:52 -05:00
jcwalker3 64b6eb5d54 Merge branch 'master' into feat/issue-638-webui-app-shell-phase1 2026-07-23 01:01:24 -05:00
sysadminandClaude Opus 4.8 f21f81f9b5 Merge branch 'master' into feat/issue-638-webui-app-shell-phase1
Resolve conflict remediation for PR #818 (Closes #638) against master
caaae9b6. Two conflicts in webui/app.py, both resolved as unions since
the Phase 1 shell work (#638) and the merged master changes touch
disjoint concerns:

- Imports: keep the new webui.nav (NAV_GROUPS, STUB_PAGES) import from
  #638 alongside master's expanded project_registry / project_views API
  (ProjectRegistry, RegistryError, known_project_ids,
  project_detail_to_dict, render_registry_error).
- Route table: keep master's read-only /api/console/security-model route
  alongside #638's read-only Phase 1 stub routes (STUB_PAGES).

No behavior change beyond union; console stays read-only. docs/webui-local-dev.md
auto-merged. Full webui suite green (272 passed, 310 subtests).

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-07-23 01:14:51 -04:00
jcwalker3 badc4e636b Merge branch 'master' into fix/issue-790-slice-a-heartbeat-policy 2026-07-23 00:06:18 -05:00
jcwalker3 da6a864463 Merge branch 'master' into feat/issue-634-readonly-system-health-api 2026-07-23 00:06:06 -05:00
sysadminandClaude Opus 4.8 b7a63a5579 feat(webui): unified session/lease/lock/worktree inventory API (Closes #636)
Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-07-22 21:13:54 -04:00
jcwalker3andClaude Opus 4.8 08061b7b8a feat(webui): evolve application shell for Phase 1 console IA (Closes #638)
Introduce a single nav-config module (webui/nav.py) driving grouped
navigation across the epic #631 Phase 1 information architecture:
Health, Traffic, Runtime/Sessions, Projects, Inventory, Timeline,
Policy (placeholder), and Insights (placeholder). The shell header now
carries a read-only environment badge (local/remote from WEBUI_HOST), a
mode: read-only badge, and a Docs link. Home summarizes the console
purpose, lists the Phase 1 surfaces by group, and links the MVP legacy
pages.

Not-yet-implemented surfaces (/sessions, /inventory, /timeline,
/policy, /insights) resolve to graceful read-only stub pages rather than
404s; their backing views land in later child issues of #631 (inventory
surfaces backed by #636). Mutating methods on stub routes still fail
closed with read-only-mvp. No privileged action controls are added.

Adds tests/test_webui_shell.py covering nav groups, badge presence,
environment classification, docs link, stub routes (200 + read-only),
and home content. Updates docs/webui-local-dev.md with the new routes
and a Phase 1 shell section.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-07-22 17:56:28 -05:00
jcwalker3andClaude Opus 4.8 5494696227 feat(webui): read-only system-health API (Closes #634)
Adds `GET /api/v1/system/health`, a structured read-only health surface for
automated readiness checks, and keeps `/health` as the cheap liveness probe.

webui/system_health.py composes a DTO from fail-soft dependency probes: the
control-plane database, the local checkout, and — opt-in via `?deep=1` — live
Gitea reachability, each carrying status, reason, and probe latency. Required
probes drive readiness; the optional Gitea probe can only degrade overall
status, because local inventory stays serveable when the remote is
unreachable. A probe that did not run leaves readiness incomplete rather than
silently passing.

Read-only throughout: the control-plane database is opened through a `mode=ro`
URI because `ControlPlaneDB.__init__` creates directories and runs migrations,
which a health check must never do. No restart or reload control is exposed;
those are Phase 2 and #630 forbids process-kill recovery.

No unproven claims: `stale_runtime.mutation_safe` is true only when the
runtime, checkout, and remote commits are all known and equal, and MCP
namespaces always report `unproven` because a web process cannot exercise the
IDE-managed client path (#543). Probe details are redacted at the browser
boundary — URLs lose userinfo and query strings, credential-shaped text is
masked.

`/health` is expanded additively: every MVP key is retained, plus `started_at`,
`uptime_seconds`, and a pointer to the versioned API. The versioned route
returns 503 when not ready so automation can branch on the status code alone.

Verified at master 9eb0f29: focused file 40 passed / 11 subtests; `-k "webui or
health"` 230 passed / 159 subtests; full suite 4358 passed with the 11
pre-existing master-drift failures unchanged from the clean-master baseline.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-07-22 15:59:30 -05:00
jcwalker3 c1d2bad901 feat(issue-628): Add unit and integration tests for autonomous canonical handoffs and task orchestration 2026-07-22 04:08:43 -05:00
sysadminandClaude Opus 4.8 243f52dc79 feat(lease): make the author task heartbeat load-bearing (#790 Slice A)
Slice A of Issue #790, per the controller reassessment in comment 13958. Does
not close the issue: terminal retirement (Slice B) and the read-side generation
check plus the #760 renewal re-scope (Slice C) are deliberately not implemented.

The defect. `issue_lock_store.assess_lock_freshness` parsed `last_heartbeat_at`
and then never consulted it. Liveness was decided by an absolute four-hour
`expires_at` and by PID liveness, and the recorded PID is the long-lived MCP
daemon rather than the authoring task, so an abandoned claim stayed live for the
full four hours. A tree-wide search found the field written in exactly one place
and advanced by nothing. Issue #787 / PR #789 hit this; Issue #760 / PR #791 hit
it again, blocking reconciliation for over five hours after its work had landed.

A1 — central policy. New `lease_policy` declares every duration for every task
class in one place: author initial/sliding TTL 10 minutes, heartbeat cadence 2,
stale warning 5, missed-heartbeat grace 10, absolute cap 8 hours, recovery grace
10, terminal race-drain 2. It ships first so the first heartbeat and TTL
behavior to run reads from it (AC-N7). The duplicated four-hour literal is gone
from both `issue_lock_store` and `gitea_mcp_server`. Reviewer, merger, and
conflict-fix classes are declared but not rewired — Slice C moves those call
sites — and a test asserts the declaration still equals the constants #747 and
`pr_work_lease` own, so the two cannot drift apart unnoticed.

A2 — load-bearing freshness, with two deliberate asymmetries. An alive PID never
establishes freshness anywhere (AC-N2); it is recorded as evidence and no branch
returns live because of it. A dead PID still marks a lease stale, and that band
still precedes every heartbeat evaluation, so #753 dead-session recovery keys on
exactly the classification it always did. New bands `stale_missed_heartbeat` and
`stale_absolute_cap` are classified in `branch_cleanup_guard` rather than
falling through to unknown-status, and still block unless the ownership record
proves `reclaim_allowed is True`. A heartbeat lease carrying no heartbeat is
contradictory and fails closed. `assess_expired_lock_reclaim` accepts a lapsed
heartbeat as reclaim grounds for heartbeat-lifecycle leases only: under this
lifecycle the heartbeat is the liveness proof, and also requiring a dead PID
would reinstate the original defect.

A3/A4 — task-session identity and the writer. `mint_task_session_id` produces an
ownership key containing no process identifier, since the daemon PID is reused
by every task it serves and identifies none of them. `heartbeat_session_lock`
writes inside the existing per-issue flock under the #772 generation
compare-and-swap, verifying exact issue, branch, realpath-normalized worktree,
claimant username, claimant profile, and recorded session identifier. It cannot
acquire, take over, or revive: a lease past its grace is refused and must use
the reclaim path, so a session that stopped proving liveness cannot restore
ownership retroactively. New `gitea_heartbeat_issue_lock` gates on the same
authority as `lock_issue`, being strictly narrower.

A5 — legacy compatibility (AC-N8). The explicit `lifecycle_version` marker, never
a timestamp comparison, discriminates legacy from heartbeat leases: a legacy lock
has `last_heartbeat_at == created_at` forever precisely because nothing advanced
it, and a freshly minted heartbeat lease has them equal too, so the equality
carries no information in either direction. Legacy locks keep their recorded
absolute expiry and are never evaluated against the short grace, so deployment
cannot make an existing claim instantly reclaimable. They leave that state only
by terminal retirement (Slice B) or by `rebind_legacy_lock`, which re-verifies
the exact owner and mints a genuine identifier and first heartbeat while
preserving the original claim under `legacy_origin`. Rebinding a lapsed legacy
lease is refused; that belongs to #760 renewal or #601 reclaim.

A6 — native coverage. Review #499 proved assessor-level tests miss discard
points, so `tests/test_issue_790_heartbeat_mcp_path.py` drives the real tools
against a real git repository and a real durable lock: lock creation and
read-back, policy window, freshness, survival of `verify_lock_for_mutation`,
invariance of the duplicate-work and linked-open-PR gates, CAS rejection,
foreign-session and foreign-claimant refusal, alive-PID-only refusal, missed
heartbeat, legacy protection on deployment, and legacy rebinding.

Tests. New suites 55 passed. Lock and lease regression set (issue_lock_store,
lease_lifecycle, #753, #755, #760 x2, #768, #772, lock registration, worktree,
adoption, duplicate gate, branch cleanup guard, capability invariants, claim
heartbeat, worktrees) 383 passed with 98 subtests. Full suite 4295 passed, 11
failed, 6 skipped, 499 subtests passed, against a clean master baseline worktree
at 620ed6e9 that reports 11 failed and 4240 passed — the same eleven node IDs.
The 55-test delta is exactly the new suites; no new failures.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_011u6GKSJwwrrYjguPjs1aK5
2026-07-22 04:19:48 -04:00
86 changed files with 29586 additions and 201 deletions
+176 -5
View File
@@ -53,6 +53,8 @@ OUTCOME_CANDIDATE_SET_DRIFT = "candidate_set_drift"
SKIP_CLAIMED_BY_OTHER_SESSION = "claimed_by_other_session"
# #776: controller-supplied pre-rank exclusion.
SKIP_EXCLUDED_BY_CONTROLLER = "excluded_by_controller"
# #844: epic / child-only implementation container (pre-rank).
SKIP_EPIC_OR_CHILD_ONLY_CONTAINER = "epic_or_child_only_container"
# Ownership verdicts for a live claim on a candidate (#765).
OWNERSHIP_OWN = "own"
@@ -130,6 +132,39 @@ ROLE_ACTIONS: dict[str, tuple[tuple[str, ...], tuple[str, ...]]] = {
}
# Body phrases that prove an issue is an implementation container, not a
# unit of direct author work (#844). Matched case-insensitively against the
# issue body. Title alone is never sufficient (ordinary issues may mention
# "epic" incidentally).
_CHILD_ONLY_BODY_MARKERS: tuple[str, ...] = (
"implementation is delivered via child issues only",
"implementation is delivered through child issues only",
"implementation is delivered via child issues",
"implementation is delivered through child issues",
"do not implement product features in this epic",
"do not implement product features in this epic issue itself",
"no product feature implementation is claimed complete solely on this epic",
"implementable child issues remain independently eligible",
"owns the product roadmap and linkage",
"this epic owns the product roadmap",
"coordination container",
"child-only container",
"implementation is delegated to child",
)
# Explicit epic / umbrella labels (structured evidence preferred over title).
_EPIC_LABELS: frozenset[str] = frozenset(
{
"type:epic",
"epic",
"kind:epic",
"scope:epic",
"type:umbrella",
"umbrella",
}
)
@dataclass
class WorkCandidate:
"""One assignable Gitea issue or PR presented to the allocator."""
@@ -139,6 +174,7 @@ class WorkCandidate:
state: str = "open"
labels: tuple[str, ...] = ()
title: str = ""
body: str = ""
priority: int = 0
head_sha: str | None = None
# Routing signals (callers derive from Gitea / review feedback).
@@ -158,6 +194,7 @@ class WorkCandidate:
self.labels = tuple(
str(x).strip().lower() for x in (self.labels or ()) if str(x).strip()
)
self.body = str(self.body or "")
if self.kind not in WORK_KINDS:
raise InvalidWorkKindError(
f"candidate kind '{self.kind}' is not assignable; only "
@@ -171,6 +208,7 @@ class WorkCandidate:
"state": self.state,
"labels": list(self.labels),
"title": self.title,
"body": self.body,
"priority": self.priority,
"head_sha": self.head_sha,
"request_changes_current_head": self.request_changes_current_head,
@@ -184,6 +222,51 @@ class WorkCandidate:
}
def classify_epic_or_child_only_container(
c: WorkCandidate,
) -> tuple[bool, str | None]:
"""Return whether *c* is an epic / child-only implementation container (#844).
Exclusion uses structured evidence first (labels, body scope language).
A bare title containing the word "epic" is **not** enough — ordinary
implementable issues may mention epics incidentally. A title that is
explicitly prefixed ``Epic:`` only counts when the body also proves
child-only / no-direct-implementation scope (or an epic label is present).
PRs are never classified as containers here (they already have a head).
"""
if c.kind != "issue":
return False, None
labels = set(c.labels)
epic_label = sorted(labels & _EPIC_LABELS)
body_l = (c.body or "").lower()
title = (c.title or "").strip()
title_l = title.lower()
body_hits = [m for m in _CHILD_ONLY_BODY_MARKERS if m in body_l]
title_epic_prefix = title_l.startswith("epic:") or title_l.startswith("epic ")
if epic_label:
detail = f"label={epic_label[0]}"
if body_hits:
detail = f"{detail}; body_marker={body_hits[0]!r}"
return True, detail
if body_hits:
# Body proves child-only / umbrella scope. Title "Epic:" is corroborating
# but not required — containers without the word still exclude.
detail = f"body_marker={body_hits[0]!r}"
if title_epic_prefix:
detail = f"title_epic_prefix; {detail}"
return True, detail
# Title-only "Epic:" without body scope evidence is insufficient (#844 AC:
# eligibility does not rely solely on the word "Epic" in a title).
# Similarly, incidental "epic" mid-title without markers stays eligible.
return False, None
@dataclass
class SkipRecord:
kind: str
@@ -850,7 +933,8 @@ def allocate_next_work(
ownership_defects: list[dict[str, Any]] = []
controller_excluded: list[dict[str, Any]] = []
# #776 AC2: remove excluded numbers *before* ranking / selection / lease.
# #776 AC2 + #844: remove excluded numbers *and* epic/child-only containers
# *before* ranking / selection / lease so they never receive assignments.
rankable: list[WorkCandidate] = []
for c in candidates:
if int(c.number) in exclude_set:
@@ -929,6 +1013,23 @@ def allocate_next_work(
},
}
continue
# #844: epics / child-only containers are never direct implement targets.
is_container, container_detail = classify_epic_or_child_only_container(c)
if is_container:
detail = container_detail or "epic or child-only container"
reason = (
f"{c.kind}#{c.number} {SKIP_EPIC_OR_CHILD_ONLY_CONTAINER}: "
f"{detail}; implementation is delegated to child issues"
)
skipped.append(
SkipRecord(
c.kind,
c.number,
reason,
SKIP_EPIC_OR_CHILD_ONLY_CONTAINER,
)
)
continue
rankable.append(c)
ordered = sort_candidates(rankable)
@@ -1115,6 +1216,12 @@ def allocate_next_work(
"reasons": [
"dry-run only (apply=false); no assignment/lease created — "
"call again with apply=true to reserve via control-plane DB"
+ (
"; after apply, the required-role worker consumes via "
"gitea_adopt_workflow_lease (#843)"
if mode == ALLOCATION_MODE_CROSS_ROLE and expected_role != role_norm
else ""
)
],
"skipped": [s.as_dict() for s in skipped],
"terminal_pr": terminal_pr,
@@ -1146,6 +1253,9 @@ def allocate_next_work(
# Atomic reserve via #613 substrate.
ttl = lease_ttl_seconds if lease_ttl_seconds is not None else None
try:
cross_role_handoff = (
mode == ALLOCATION_MODE_CROSS_ROLE and lease_role != role_norm
)
kwargs: dict[str, Any] = {
"session_id": session_id,
"role": lease_role,
@@ -1157,7 +1267,8 @@ def allocate_next_work(
"expected_head_sha": selected.head_sha,
"allowed_actions": allowed,
"forbidden_actions": forbidden,
"phase": "allocated",
# #843: mark cross-role allocations as awaiting independent consume
"phase": "awaiting_handoff" if cross_role_handoff else "allocated",
}
if ttl is not None:
kwargs["lease_ttl_seconds"] = int(ttl)
@@ -1237,7 +1348,52 @@ def allocate_next_work(
"lease_role": lease_role,
"source": "control_plane_db.assign_and_lease",
}
return {
consume_allocation = None
if cross_role_handoff and result.lease_id:
# Durable handoff marker so independent required-role workers can
# consume without sharing the controller session (#843).
handoff_prov = {
"cross_role_handoff": True,
"handoff_status": "pending",
"allocating_session_id": session_id,
"allocating_role": role_norm,
"required_role": expected_role,
"required_profile": selection["required_profile"],
"required_namespace": selection["required_namespace"],
"assignment_id": result.assignment_id,
"lease_id": result.lease_id,
"allocation_mode": mode,
"adopted_by_session_id": None,
}
try:
db.attach_lease_provenance(result.lease_id, handoff_prov)
except ControlPlaneError:
# Still return assignment evidence; consume path may be unavailable
handoff_prov["attach_failed"] = True
consume_allocation = {
"tool": "gitea_adopt_workflow_lease",
"lease_id": result.lease_id,
"assignment_id": result.assignment_id,
"required_role": expected_role,
"required_profile": selection["required_profile"],
"required_namespace": selection["required_namespace"],
"handoff_status": "pending",
"controller_session_required": False,
"instructions": (
f"From an independent {expected_role} session "
f"({selection['required_namespace']} / "
f"{selection['required_profile']}), call "
f"gitea_adopt_workflow_lease(lease_id={result.lease_id!r}) "
"to consume this controller allocation. The allocating "
"controller process does not need to remain alive. Wrong-role "
"and second-adoption attempts fail closed."
),
}
lease_proof["cross_role_handoff"] = True
lease_proof["handoff_status"] = "pending"
lease_proof["consume_tool"] = "gitea_adopt_workflow_lease"
out = {
"success": True,
"outcome": OUTCOME_ASSIGNED,
"apply": True,
@@ -1271,8 +1427,17 @@ def allocate_next_work(
"lease_role": lease_role,
"lease_proof": lease_proof,
"selection_policy": SELECTION_POLICY,
"cross_role_handoff": bool(cross_role_handoff),
},
"next_valid_command": _next_command(lease_role, selected),
"next_valid_command": (
(
f"consume lease {result.lease_id} via gitea_adopt_workflow_lease "
f"as {expected_role}, then "
)
+ _next_command(lease_role, selected)
if cross_role_handoff
else _next_command(lease_role, selected)
),
"substrate": "control_plane_db",
"file_lock_only": False,
"comment_lease_only": False,
@@ -1287,9 +1452,14 @@ def allocate_next_work(
"downstream_note": (
"#612 incident bridge remains downstream of #600; "
"allocator never assigns raw monitoring incidents; "
"controller routes only under cross_role (#840)"
"controller routes only under cross_role (#840); "
"cross-role assignments are consumable by independent "
"required-role workers via gitea_adopt_workflow_lease (#843)"
),
}
if consume_allocation is not None:
out["consume_allocation"] = consume_allocation
return out
def _next_command(role: str, c: WorkCandidate) -> str:
@@ -1341,6 +1511,7 @@ def candidate_from_dict(data: dict[str, Any]) -> WorkCandidate:
state=str(data.get("state") or "open"),
labels=tuple(data.get("labels") or ()),
title=str(data.get("title") or ""),
body=str(data.get("body") or ""),
priority=priority,
head_sha=data.get("head_sha"),
request_changes_current_head=bool(data.get("request_changes_current_head")),
File diff suppressed because it is too large Load Diff
+80 -37
View File
@@ -40,22 +40,88 @@ def _normalize_path(path: str) -> str:
return (path or "").replace("\\", "/").rstrip("/")
def get_canonical_branches_root(project_root: str | None = None) -> str:
"""Return the absolute path of the canonical branches directory for *project_root*."""
root = os.path.realpath(project_root) if project_root else os.path.realpath(os.getcwd())
canonical_repo_root = resolve_canonical_repo_root(root, root)
return os.path.realpath(os.path.join(canonical_repo_root, "branches"))
def is_path_under_branches(path: str, project_root: str | None = None) -> bool:
"""True when *path* resolves inside ``<project_root>/branches/``."""
normalized = _normalize_path(path)
if not normalized:
"""True when *path* resolves inside a canonical ``branches/`` directory."""
if not path or not str(path).strip():
return False
if "/branches/" in f"{normalized}/":
return True
if normalized.endswith("/branches"):
return True
if project_root:
root = _normalize_path(os.path.realpath(project_root))
real = _normalize_path(os.path.realpath(path))
if real.startswith(f"{root}/"):
rel = real[len(root) + 1 :]
return rel == "branches" or rel.startswith("branches/")
return False
try:
real_path = os.path.realpath(os.path.abspath(str(path).strip()))
except Exception:
return False
branches_root = get_canonical_branches_root(project_root or real_path)
try:
common = os.path.commonpath([branches_root, real_path])
except Exception:
return False
if common != branches_root:
return False
rel = os.path.relpath(real_path, branches_root)
return rel != "." and not rel.startswith("..")
def resolve_canonical_repo_root(workspace_path: str, fallback_project_root: str) -> str:
"""Return the stable repository root for *workspace_path* via git metadata (#460)."""
p = (workspace_path or "").strip()
if p:
try:
res = subprocess.run(
["git", "-C", p, "rev-parse", "--git-common-dir"],
capture_output=True,
text=True,
check=True,
)
common = _realpath_git_common_dir(p, res.stdout)
if common.endswith(f"{os.sep}.git") or os.path.basename(common) == ".git":
candidate_root = os.path.dirname(common)
real_p = os.path.realpath(p)
try:
if os.path.commonpath([candidate_root, real_p]) == candidate_root:
return candidate_root
except Exception:
pass
except Exception:
pass
# Fallback when git metadata is unavailable. Never string-split on
# "/branches/" (review #531 F2 / #551): recover the repo root only via
# resolved-path commonpath ancestry. Do **not** require on-disk isdir —
# MCP may launch with project_root = branches/<wt> before that path
# exists, and #274 path-shaped worktree-as-project-root must still resolve.
fallback = os.path.realpath(fallback_project_root or workspace_path or ".")
cur = fallback
for _ in range(64):
parent = os.path.dirname(cur)
if parent == cur:
break
branches_dir = os.path.realpath(os.path.join(parent, "branches"))
try:
# Path-shaped: fallback is under parent/branches/ (commonpath).
if os.path.commonpath([branches_dir, fallback]) == branches_dir:
return parent
except ValueError:
pass
# Fallback path itself is the branches directory.
if os.path.basename(os.path.realpath(cur)) == "branches":
try:
if os.path.commonpath([os.path.realpath(cur), fallback]) == os.path.realpath(
cur
):
return parent
except ValueError:
pass
cur = parent
return fallback
def resolve_mutation_workspace(
@@ -87,29 +153,6 @@ def _realpath_git_common_dir(workspace_path: str, common_dir: str) -> str:
return os.path.realpath(os.path.join(workspace_path, raw))
def resolve_canonical_repo_root(workspace_path: str, fallback_project_root: str) -> str:
"""Return the stable repository root for *workspace_path* via git metadata (#460)."""
path = (workspace_path or "").strip()
fallback = os.path.realpath(fallback_project_root)
if not path:
return fallback
try:
res = subprocess.run(
["git", "-C", path, "rev-parse", "--git-common-dir"],
capture_output=True,
text=True,
check=True,
)
common = _realpath_git_common_dir(path, res.stdout)
except Exception:
return fallback
if common.endswith(f"{os.sep}.git"):
return os.path.dirname(common)
if os.path.basename(common) == ".git":
return os.path.dirname(common)
return fallback
def resolve_author_mutation_context(
worktree_path: str | None,
process_project_root: str,
+97 -1
View File
@@ -163,7 +163,19 @@ _TERMINAL_OWNERSHIP_STATUSES = frozenset(
{"released", "abandoned", "done", "blocked", "terminal", "closed"}
)
_EXPIRED_STATUSES = frozenset({"expired"})
_STALE_STATUSES = frozenset({"stale", "stale_dead_process", "stale_missing_worktree"})
_STALE_STATUSES = frozenset(
{
"stale",
"stale_dead_process",
"stale_missing_worktree",
# #790 Slice A heartbeat-lifecycle bands. Listed here so they are
# *classified* rather than falling through to the unknown-status branch;
# they still block unless the ownership record proves
# ``reclaim_allowed is True``, so the O2 fail-closed rule is unchanged.
"stale_missed_heartbeat",
"stale_absolute_cap",
}
)
def _norm_str(value: Any) -> str:
@@ -525,6 +537,90 @@ def assess_ownership_record_activity(record: dict[str, Any]) -> dict[str, Any]:
}
# Reviewer-lease reclaim is only reachable from a non-live (expired/stale) lease.
_RECLAIMABLE_REVIEWER_STATUSES = _EXPIRED_STATUSES | _STALE_STATUSES
def is_active_ownership_status(status: str | None) -> bool:
"""True when *status* denotes live/active ownership of a branch (#855).
Used to decide whether a *competing* active claimant still uses a branch
when weighing an expired reviewer lease for reclaim. Expired, stale,
released, and terminal statuses are not active.
"""
return _norm_str(status).lower() in _ACTIVE_OWNERSHIP_STATUSES
def assess_expired_reviewer_lease_reclaim(
*,
role: str,
status: str,
pr_merged: bool | None,
owner_pid_alive: bool | None,
competing_active_claimant: bool | None,
) -> dict[str, Any]:
"""Decide, explicitly and fail-closed, whether an expired reviewer lease
may stop protecting an already-merged branch (#855 AC4).
An expired reviewer lease should not protect a merged branch forever once
its work is done and no live claimant remains. Reclaim is permitted only
when **every** condition below is provably satisfied; any unknown
(``None``) or contrary value keeps the lease protective:
- the lease is a ``reviewer`` lease (author/merger/controller/reconciler
leases are out of scope and always keep protecting);
- its status is expired or stale (never an active/live lease);
- the PR is proven merged (``pr_merged is True``);
- the lease owner process is proven dead (``owner_pid_alive is False``);
- no competing active claimant uses the branch
(``competing_active_claimant is False``).
Returns a decision dict with ``reclaim_allowed`` and, when refused, the
fail-closed ``reasons``. The reasons never contain secrets — only the
role, the status, and which condition was unproven.
"""
reasons: list[str] = []
normalized_role = _norm_str(role).lower()
normalized_status = _norm_str(status).lower()
if normalized_role != "reviewer":
reasons.append(
f"lease role '{normalized_role or 'unknown'}' is not a reviewer "
"lease; expired-reviewer reclaim does not apply"
)
if normalized_status not in _RECLAIMABLE_REVIEWER_STATUSES:
reasons.append(
f"lease status '{normalized_status or 'unknown'}' is not expired "
"or stale; only a non-live reviewer lease may be reclaimed"
)
if pr_merged is not True:
reasons.append(
"PR merged state is not proven true; reclaim requires an "
"already-merged PR (fail closed)"
)
if owner_pid_alive is not False:
reasons.append(
"lease owner process liveness is not proven dead; a live owner "
"still protects the branch (fail closed)"
)
if competing_active_claimant is not False:
reasons.append(
"a competing active claimant may still use the branch; reclaim "
"requires no other active ownership (fail closed)"
)
allowed = not reasons
return {
"reclaim_allowed": allowed,
"role": normalized_role,
"status": normalized_status,
"decision": (
"reclaim_expired_reviewer_lease" if allowed else "keep_protecting"
),
"reasons": [] if allowed else reasons,
}
def assess_active_branch_ownership(
*,
remote: str,
+19 -2
View File
@@ -46,6 +46,17 @@ _FIELD_RE = re.compile(
)
def is_known_cth_type(value: str | None) -> bool:
"""True when *value* is a declared member of the :data:`CTH_TYPES` contract.
``CTH_TYPES`` is the single authority for what a CTH type may be. The
heading a comment carries is free text, so a *read* path that turns a parsed
type into something durable — a serialized field, a routing decision — must
check membership here rather than trust the parse or keep a list of its own.
"""
return (value or "").strip() in CTH_TYPES
def format_cth_body(
*,
cth_type: str,
@@ -60,7 +71,7 @@ def format_cth_body(
) -> str:
"""Render a canonical CTH comment body."""
normalized_type = (cth_type or "").strip()
if normalized_type not in CTH_TYPES:
if not is_known_cth_type(normalized_type):
raise ValueError(
f"unknown CTH type '{cth_type}'; expected one of {sorted(CTH_TYPES)}"
)
@@ -101,6 +112,12 @@ def parse_cth_comment(body: str) -> dict[str, Any] | None:
fields[key] = match.group(2).strip()
return {
"cth_type": cth_type,
# The heading capture is unconstrained free text, so the parse states
# whether it satisfies the CTH_TYPES contract instead of leaving every
# reader to decide (or forget). Parsing stays total — an unknown type is
# still parsed and reported, never raised on — but a reader that turns
# the type into a durable value can now tell the two apart.
"cth_type_known": is_known_cth_type(cth_type),
"fields": fields,
"raw_body": text,
}
@@ -119,7 +136,7 @@ def assess_cth_comment(body: str) -> dict[str, Any]:
}
cth_type = parsed.get("cth_type") or ""
if cth_type not in CTH_TYPES:
if not is_known_cth_type(cth_type):
reasons.append(
f"unknown CTH type '{cth_type}'; expected one of {sorted(CTH_TYPES)}"
)
+856 -4
View File
@@ -30,8 +30,9 @@ from datetime import datetime, timedelta, timezone
from typing import Any, Iterator, Sequence
import dependency_graph
import gitea_audit
SCHEMA_VERSION = 4
SCHEMA_VERSION = 5
# Assignable work kinds only — raw monitoring incidents are never work items.
WORK_KINDS = frozenset({"issue", "pr"})
@@ -177,6 +178,51 @@ CREATE TABLE IF NOT EXISTS dependency_edges (
)
);
-- Durable MCP session checkpoints (#660, umbrella #655 child; #628 handoff
-- goal). A restart otherwise loses session identity, stage, lease ownership,
-- and next action, forcing human reconstruction. Each row is the *current*
-- recoverable state of one session's work on one work unit; stage transitions
-- upsert the row and audit the change to ``events``. Creating the table is the
-- v4->v5 migration: additive, idempotent, and it never touches prior tables.
--
-- ``lease_id`` / ``assignment_id`` are recorded as soft references (plain TEXT,
-- no enforced FK) exactly as dependency_edges references issues/PRs by number:
-- a checkpoint is a recovery artifact that must survive the deletion of the
-- lease it names, so a hard FK would fail the pre-drain write the issue
-- requires to succeed. ``work_number`` uses 0 (never a real issue/PR number)
-- as the "session-level, no specific work" sentinel so UNIQUE is NULL-safe.
-- Every free-text / JSON field is redaction-filtered before storage (AC4).
CREATE TABLE IF NOT EXISTS session_checkpoints (
checkpoint_id TEXT PRIMARY KEY,
remote TEXT NOT NULL,
org TEXT NOT NULL,
repo TEXT NOT NULL,
session_id TEXT NOT NULL,
provider_identity TEXT NOT NULL DEFAULT '',
role TEXT NOT NULL DEFAULT '',
work_kind TEXT NOT NULL DEFAULT '' CHECK (work_kind IN ('issue', 'pr', '')),
work_number INTEGER NOT NULL DEFAULT 0,
worktree_path TEXT NOT NULL DEFAULT '',
branch TEXT NOT NULL DEFAULT '',
head_sha TEXT NOT NULL DEFAULT '',
capabilities TEXT NOT NULL DEFAULT '[]',
lease_id TEXT NOT NULL DEFAULT '',
assignment_id TEXT NOT NULL DEFAULT '',
workflow_stage TEXT NOT NULL DEFAULT '',
last_completed_action TEXT NOT NULL DEFAULT '',
current_operation TEXT NOT NULL DEFAULT '',
pending_mutation TEXT NOT NULL DEFAULT '',
evidence TEXT NOT NULL DEFAULT '{}',
blocker TEXT NOT NULL DEFAULT '',
next_valid_action TEXT NOT NULL DEFAULT '',
recovery_instructions TEXT NOT NULL DEFAULT '',
checkpoint_schema_version INTEGER NOT NULL DEFAULT 5,
status TEXT NOT NULL DEFAULT 'active',
created_at TEXT NOT NULL,
updated_at TEXT NOT NULL,
UNIQUE (remote, org, repo, session_id, work_kind, work_number)
);
CREATE INDEX IF NOT EXISTS idx_leases_work_status ON leases(work_item_id, status);
-- Reverse lookup ("what waits on this target") is the query automatic
-- resumption needs, so it gets its own index alongside the forward one.
@@ -186,6 +232,40 @@ CREATE INDEX IF NOT EXISTS idx_dependency_edges_target
ON dependency_edges(remote, org, repo, target_kind, target_number);
CREATE INDEX IF NOT EXISTS idx_assignments_session ON assignments(session_id, status);
CREATE INDEX IF NOT EXISTS idx_incident_gitea ON incident_links(gitea_org, gitea_repo, gitea_issue_number);
-- Resume queries hit by session (what was this session doing) and by work unit
-- (who was checkpointed on this issue/PR), so both get an index.
CREATE INDEX IF NOT EXISTS idx_session_checkpoints_session
ON session_checkpoints(remote, org, repo, session_id);
CREATE INDEX IF NOT EXISTS idx_session_checkpoints_work
ON session_checkpoints(remote, org, repo, work_kind, work_number);
-- Model usage, token cost, latency, and performance events (#651)
CREATE TABLE IF NOT EXISTS usage_events (
usage_id INTEGER PRIMARY KEY AUTOINCREMENT,
session_id TEXT,
remote TEXT NOT NULL DEFAULT 'dadeschools',
org TEXT NOT NULL DEFAULT '',
repo TEXT NOT NULL DEFAULT '',
project_id TEXT,
role TEXT NOT NULL DEFAULT 'unknown',
model TEXT NOT NULL DEFAULT 'unknown',
issue_number INTEGER,
pr_number INTEGER,
stage TEXT NOT NULL DEFAULT 'unknown',
input_tokens INTEGER,
output_tokens INTEGER,
total_tokens INTEGER,
estimated_cost_usd REAL,
latency_ms INTEGER,
duration_ms INTEGER,
status TEXT NOT NULL DEFAULT 'success',
metadata TEXT,
created_at TEXT NOT NULL
);
CREATE INDEX IF NOT EXISTS idx_usage_events_scope ON usage_events(remote, org, repo);
CREATE INDEX IF NOT EXISTS idx_usage_events_role_model ON usage_events(role, model);
CREATE INDEX IF NOT EXISTS idx_usage_events_stage ON usage_events(stage);
"""
@@ -339,6 +419,7 @@ class ControlPlaneDB:
self._migrate_incident_links_null_scope(conn)
self._migrate_lease_lifecycle_columns(conn)
self._migrate_session_ownership_columns(conn)
self._migrate_usage_events_table(conn)
conn.execute(
"INSERT OR REPLACE INTO schema_meta(key, value) VALUES (?, ?)",
("schema_version", str(SCHEMA_VERSION)),
@@ -518,6 +599,207 @@ class ControlPlaneDB:
f"UPDATE incident_links SET {col} = '' WHERE {col} IS NULL"
)
def _migrate_usage_events_table(self, conn: sqlite3.Connection) -> None:
"""Create usage_events table and indexes if they do not exist (#651)."""
conn.execute("""
CREATE TABLE IF NOT EXISTS usage_events (
usage_id INTEGER PRIMARY KEY AUTOINCREMENT,
session_id TEXT,
remote TEXT NOT NULL DEFAULT 'dadeschools',
org TEXT NOT NULL DEFAULT '',
repo TEXT NOT NULL DEFAULT '',
project_id TEXT,
role TEXT NOT NULL DEFAULT 'unknown',
model TEXT NOT NULL DEFAULT 'unknown',
issue_number INTEGER,
pr_number INTEGER,
stage TEXT NOT NULL DEFAULT 'unknown',
input_tokens INTEGER,
output_tokens INTEGER,
total_tokens INTEGER,
estimated_cost_usd REAL,
latency_ms INTEGER,
duration_ms INTEGER,
status TEXT NOT NULL DEFAULT 'success',
metadata TEXT,
created_at TEXT NOT NULL
);
""")
conn.execute("CREATE INDEX IF NOT EXISTS idx_usage_events_scope ON usage_events(remote, org, repo);")
conn.execute("CREATE INDEX IF NOT EXISTS idx_usage_events_role_model ON usage_events(role, model);")
conn.execute("CREATE INDEX IF NOT EXISTS idx_usage_events_stage ON usage_events(stage);")
# #651 retention: cap growth so unauthenticated or high-volume ingest
# cannot DoS the control-plane DB (PR #876 F3). Applied after every write.
USAGE_EVENTS_MAX_ROWS = 10_000
USAGE_EVENTS_RETENTION_DAYS = 90
def record_usage_event(
self,
*,
session_id: str | None = None,
remote: str = "dadeschools",
org: str = "",
repo: str = "",
project_id: str | None = None,
role: str = "unknown",
model: str = "unknown",
issue_number: int | None = None,
pr_number: int | None = None,
stage: str = "unknown",
input_tokens: int | None = None,
output_tokens: int | None = None,
total_tokens: int | None = None,
estimated_cost_usd: float | None = None,
latency_ms: int | None = None,
duration_ms: int | None = None,
status: str = "success",
metadata: str | dict[str, Any] | None = None,
created_at: str | None = None,
) -> int:
"""Record a model usage, token cost, latency, or stage performance event (#651)."""
ts = created_at or _ts()
meta_str: str | None = None
if metadata is not None:
from webui import console_redaction
redacted_meta = console_redaction.redact_payload(metadata)
if isinstance(redacted_meta, str):
meta_str = redacted_meta
else:
try:
meta_str = json.dumps(redacted_meta, default=str)
except Exception:
meta_str = str(redacted_meta)
if total_tokens is None and (input_tokens is not None or output_tokens is not None):
total_tokens = (input_tokens or 0) + (output_tokens or 0)
with self._tx(immediate=True) as conn:
cursor = conn.execute(
"""
INSERT INTO usage_events (
session_id, remote, org, repo, project_id, role, model,
issue_number, pr_number, stage, input_tokens, output_tokens,
total_tokens, estimated_cost_usd, latency_ms, duration_ms,
status, metadata, created_at
) VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)
""",
(
session_id,
remote,
org,
repo,
project_id,
role,
model,
issue_number,
pr_number,
stage,
input_tokens,
output_tokens,
total_tokens,
estimated_cost_usd,
latency_ms,
duration_ms,
status,
meta_str,
ts,
),
)
usage_id = cursor.lastrowid
self._enforce_usage_events_retention(conn)
return usage_id
def _enforce_usage_events_retention(self, conn: sqlite3.Connection) -> None:
"""Drop aged and excess usage_events rows (PR #876 F3)."""
# Age-based: ISO-8601 UTC timestamps compare lexicographically.
cutoff = (
datetime.now(timezone.utc)
- timedelta(days=int(self.USAGE_EVENTS_RETENTION_DAYS))
).strftime("%Y-%m-%dT%H:%M:%SZ")
conn.execute(
"DELETE FROM usage_events WHERE created_at < ?",
(cutoff,),
)
# Count-based: keep the newest USAGE_EVENTS_MAX_ROWS by usage_id.
max_rows = int(self.USAGE_EVENTS_MAX_ROWS)
if max_rows > 0:
conn.execute(
"""
DELETE FROM usage_events
WHERE usage_id NOT IN (
SELECT usage_id FROM usage_events
ORDER BY usage_id DESC
LIMIT ?
)
""",
(max_rows,),
)
def query_usage_events(
self,
*,
remote: str | None = None,
org: str | None = None,
repo: str | None = None,
project_id: str | None = None,
role: str | None = None,
model: str | None = None,
issue_number: int | None = None,
pr_number: int | None = None,
stage: str | None = None,
session_id: str | None = None,
limit: int = 500,
offset: int = 0,
) -> list[dict[str, Any]]:
"""Query stored usage events matching filters (#651)."""
conditions = []
params = []
if remote:
conditions.append("remote = ?")
params.append(remote)
if org:
conditions.append("org = ?")
params.append(org)
if repo:
conditions.append("repo = ?")
params.append(repo)
if project_id:
conditions.append("project_id = ?")
params.append(project_id)
if role:
conditions.append("role = ?")
params.append(role)
if model:
conditions.append("model = ?")
params.append(model)
if issue_number is not None:
conditions.append("issue_number = ?")
params.append(issue_number)
if pr_number is not None:
conditions.append("pr_number = ?")
params.append(pr_number)
if stage:
conditions.append("stage = ?")
params.append(stage)
if session_id:
conditions.append("session_id = ?")
params.append(session_id)
where_clause = f"WHERE {' AND '.join(conditions)}" if conditions else ""
sql = f"""
SELECT * FROM usage_events
{where_clause}
ORDER BY usage_id ASC
LIMIT ? OFFSET ?
"""
params.extend([limit, offset])
with self._tx(immediate=False) as conn:
cursor = conn.execute(sql, params)
rows = cursor.fetchall()
return [dict(row) for row in rows]
# ── sessions ──────────────────────────────────────────────────────────
def upsert_session(
@@ -599,6 +881,35 @@ class ControlPlaneDB:
(_ts(), session_id),
)
def list_sessions(
self,
*,
statuses: Sequence[str] | None = None,
limit: int = 500,
) -> list[dict[str, Any]]:
"""List session rows for restart / impact analysis (#658).
Read-only. Sessions are the process-level unit an MCP restart
disrupts, so the restart coordinator inventories them to compute blast
radius. Optional ``statuses`` filter (e.g. ``('active',)``) narrows to
live rows. Never returns secrets — only operational metadata.
"""
clauses: list[str] = []
params: list[Any] = []
if statuses:
placeholders = ", ".join("?" for _ in statuses)
clauses.append(f"status IN ({placeholders})")
params.extend(statuses)
where = ("WHERE " + " AND ".join(clauses)) if clauses else ""
sql = (
f"SELECT * FROM sessions {where} "
"ORDER BY last_heartbeat_at DESC LIMIT ?"
)
params.append(max(1, int(limit)))
with self._tx(immediate=False) as conn:
rows = conn.execute(sql, params).fetchall()
return [dict(r) for r in rows]
# ── work items ────────────────────────────────────────────────────────
def upsert_work_item(
@@ -1637,11 +1948,13 @@ class ControlPlaneDB:
provenance: dict[str, Any] | None = None,
lease_ttl_seconds: int = DEFAULT_LEASE_TTL_SECONDS,
) -> dict[str, Any]:
"""Transfer or refresh a lease with provenance (#601).
"""Transfer or refresh a lease with provenance (#601 / #843).
* Same owner + active → refresh (owner-resume).
* Cross-role handoff pending + matching required role → atomic consume
(even while the allocating controller session still "owns" the lease).
* Expired/abandoned/released → create new assignment+lease with provenance.
* Active foreign → raise ForeignLeaseError (never silent steal).
* Active foreign (non-handoff) → raise ForeignLeaseError (never silent steal).
"""
now = _utc_now()
now_s = _ts(now)
@@ -1677,7 +1990,35 @@ class ControlPlaneDB:
status = "expired"
owner = lease["session_id"]
if status == "active" and owner != adopter_session_id:
# Parse durable provenance for cross-role handoff consume (#843).
lease_prov: dict[str, Any] = {}
if "provenance_json" in lease.keys() and lease["provenance_json"]:
try:
loaded = json.loads(lease["provenance_json"])
if isinstance(loaded, dict):
lease_prov = loaded
except (TypeError, json.JSONDecodeError):
lease_prov = {}
handoff_pending = bool(lease_prov.get("cross_role_handoff")) and (
str(lease_prov.get("handoff_status") or "pending").strip().lower()
== "pending"
)
already_adopted = bool(
(lease["adopted_by_session_id"] if "adopted_by_session_id" in lease.keys() else None)
or lease_prov.get("adopted_by_session_id")
)
required_role = str(
lease_prov.get("required_role") or lease["role"] or ""
).strip().lower()
adopter_role = (role or "").strip().lower()
cross_role_consume = (
handoff_pending
and not already_adopted
and status == "active"
and owner != adopter_session_id
)
if status == "active" and owner != adopter_session_id and not cross_role_consume:
raise ForeignLeaseError(
f"cannot adopt active foreign lease {lease_id} owned by {owner}"
)
@@ -1761,6 +2102,142 @@ class ControlPlaneDB:
"reasons": ["owner-resume: refreshed lease with provenance"],
}
# #843: controller→required-role handoff consume (atomic, same lease_id)
if cross_role_consume:
if not required_role:
raise ControlPlaneError(
f"cross-role handoff lease {lease_id} missing required_role"
)
if adopter_role != required_role:
raise ForeignLeaseError(
f"wrong role for cross-role handoff consume: "
f"required={required_role} adopter={adopter_role or 'none'} "
f"(fail closed)"
)
# CAS: only transfer if still owned by allocating session and unadopted
cols = self._lease_columns(conn)
adopted_col_null = (
"(adopted_by_session_id IS NULL OR adopted_by_session_id = '')"
if "adopted_by_session_id" in cols
else "1=1"
)
cas = conn.execute(
f"""
UPDATE leases
SET session_id = ?,
heartbeat_at = ?,
expires_at = ?,
phase = ?,
role = ?
WHERE lease_id = ?
AND status = 'active'
AND session_id = ?
AND {adopted_col_null}
""",
(
adopter_session_id,
now_s,
expires,
"adopted",
required_role,
lease_id,
owner,
),
)
if cas.rowcount != 1:
raise ForeignLeaseError(
f"cross-role handoff consume lost race for lease {lease_id} "
"(already adopted or no longer pending; fail closed)"
)
if "adopted_from_session_id" in cols:
conn.execute(
"""
UPDATE leases
SET adopted_from_session_id = ?, adopted_by_session_id = ?
WHERE lease_id = ?
""",
(owner, adopter_session_id, lease_id),
)
if "worktree_path" in cols and worktree_path:
conn.execute(
"UPDATE leases SET worktree_path = ? WHERE lease_id = ?",
(worktree_path, lease_id),
)
if "owner_pid" in cols and owner_pid is not None:
conn.execute(
"UPDATE leases SET owner_pid = ? WHERE lease_id = ?",
(owner_pid, lease_id),
)
if "expected_head_sha" in cols and expected_head_sha:
conn.execute(
"UPDATE leases SET expected_head_sha = ? WHERE lease_id = ?",
(expected_head_sha, lease_id),
)
# Merge handoff provenance + caller provenance
merged = dict(lease_prov)
merged.update(provenance or {})
merged["cross_role_handoff"] = True
merged["handoff_status"] = "adopted"
merged["adopted_from_session_id"] = owner
merged["adopted_by_session_id"] = adopter_session_id
merged["required_role"] = required_role
if "provenance_json" in cols:
conn.execute(
"UPDATE leases SET provenance_json = ? WHERE lease_id = ?",
(json.dumps(merged), lease_id),
)
# Transfer active assignment ownership atomically
asn_cas = conn.execute(
"""
UPDATE assignments
SET session_id = ?, role = ?
WHERE lease_id = ? AND status = 'active' AND session_id = ?
""",
(adopter_session_id, required_role, lease_id, owner),
)
if asn_cas.rowcount < 1:
# Fail closed: assignment must move with the lease
raise ControlPlaneError(
f"cross-role handoff: no active assignment for lease {lease_id} "
f"owned by {owner}"
)
lease2 = conn.execute(
"SELECT * FROM leases WHERE lease_id = ?", (lease_id,)
).fetchone()
asn = conn.execute(
"""
SELECT * FROM assignments
WHERE lease_id = ? AND status = 'active'
ORDER BY created_at DESC LIMIT 1
""",
(lease_id,),
).fetchone()
conn.execute(
"""
INSERT INTO events(work_item_id, event_type, message, created_at)
VALUES (?, 'lease_adopted', ?, ?)
""",
(
lease["work_item_id"],
f"cross-role handoff: {adopter_session_id} consumed "
f"{lease_id} from {owner} as {required_role}",
now_s,
),
)
return {
"outcome": "adopted_cross_role_handoff",
"lease": dict(lease2) if lease2 else dict(lease),
"assignment": dict(asn) if asn else None,
"reasons": [
"cross-role handoff: independent required-role worker consumed "
"controller allocation without abandonment"
],
"adopted_by_session_id": adopter_session_id,
"adopted_from_session_id": owner,
"required_role": required_role,
"handoff_status": "adopted",
}
# Non-active: create new lease + assignment (transfer)
new_lease_id = f"lease-{uuid.uuid4().hex[:16]}"
new_asn_id = f"asn-{uuid.uuid4().hex[:16]}"
@@ -2173,3 +2650,378 @@ class ControlPlaneDB:
edge["prior_state"] = prior_state
edge["state_changed"] = prior_state != state_norm
return edge
# ------------------------------------------------------------------
# Session checkpoints (#660) — durable, redacted, reconcile-on-boot.
# ------------------------------------------------------------------
# Free-text columns that carry human/agent-authored strings. Each is
# redaction-filtered before storage so an accidentally pasted token or
# credential URL can never land in a checkpoint (#660 AC4).
_CHECKPOINT_TEXT_FIELDS: tuple[str, ...] = (
"provider_identity",
"role",
"worktree_path",
"branch",
"head_sha",
"lease_id",
"assignment_id",
"workflow_stage",
"last_completed_action",
"current_operation",
"blocker",
"next_valid_action",
"recovery_instructions",
"status",
)
# JSON columns whose *decoded* structure is redacted recursively.
_CHECKPOINT_JSON_FIELDS: tuple[str, ...] = (
"capabilities",
"pending_mutation",
"evidence",
)
# Minimum a checkpoint must carry to be a usable recovery record. The drain
# gate writes with ``require_complete=True``; an incomplete write fails
# closed so drain cannot complete on a checkpoint that can't resume (#660
# "fail closed if checkpoint incomplete during drain").
REQUIRED_CHECKPOINT_FIELDS_FOR_DRAIN: tuple[str, ...] = (
"session_id",
"role",
"workflow_stage",
"next_valid_action",
"recovery_instructions",
)
@classmethod
def checkpoint_completeness(cls, record: dict[str, Any]) -> list[str]:
"""Return the drain-required fields that are missing/blank in *record*.
Pure (no DB). An empty list means the record is drain-complete.
"""
missing: list[str] = []
for field in cls.REQUIRED_CHECKPOINT_FIELDS_FOR_DRAIN:
if not str(record.get(field) or "").strip():
missing.append(field)
return missing
def write_session_checkpoint(
self,
*,
remote: str,
org: str,
repo: str,
session_id: str,
provider_identity: str = "",
role: str = "",
work_kind: str | None = None,
work_number: int | None = None,
worktree_path: str = "",
branch: str = "",
head_sha: str = "",
capabilities: Any = None,
lease_id: str = "",
assignment_id: str = "",
workflow_stage: str = "",
last_completed_action: str = "",
current_operation: str = "",
pending_mutation: Any = None,
evidence: Any = None,
blocker: str = "",
next_valid_action: str = "",
recovery_instructions: str = "",
status: str = "active",
require_complete: bool = False,
) -> dict[str, Any]:
"""Upsert the current checkpoint for one session's work on one unit.
Keyed by ``(remote, org, repo, session_id, work_kind, work_number)`` so
a stage transition refreshes the single current row (transitions are
audited to ``events``) rather than appending unbounded history — the
row is *resumable state*, matching the dependency-edge current-state
model.
Every string / JSON field is redacted via ``gitea_audit.redact`` before
it touches the DB. With ``require_complete=True`` (the drain gate path)
a record missing any drain-required field raises ``ControlPlaneError``
and writes nothing, so drain cannot proceed on an unrecoverable record.
"""
if not str(session_id or "").strip():
raise ControlPlaneError("session_id is required for a checkpoint (fail closed)")
# Normalize the work-unit key. Absent/blank work collapses to the
# ('', 0) session-level sentinel so UNIQUE stays NULL-safe.
if work_kind and str(work_kind).strip():
kind_norm = dependency_graph.normalize_work_kind(work_kind)
number_norm = int(work_number) if work_number is not None else 0
else:
kind_norm = ""
number_norm = 0
# Assemble, then redact the whole record in one pass.
raw_record: dict[str, Any] = {
"provider_identity": provider_identity or "",
"role": role or "",
"worktree_path": worktree_path or "",
"branch": branch or "",
"head_sha": head_sha or "",
"lease_id": lease_id or "",
"assignment_id": assignment_id or "",
"workflow_stage": workflow_stage or "",
"last_completed_action": last_completed_action or "",
"current_operation": current_operation or "",
"blocker": blocker or "",
"next_valid_action": next_valid_action or "",
"recovery_instructions": recovery_instructions or "",
"status": (status or "active"),
"session_id": session_id,
"capabilities": capabilities if capabilities is not None else [],
"pending_mutation": pending_mutation if pending_mutation is not None else {},
"evidence": evidence if evidence is not None else {},
}
clean = gitea_audit.redact(raw_record)
if require_complete:
missing = self.checkpoint_completeness(clean)
if missing:
raise ControlPlaneError(
"checkpoint incomplete for drain; missing "
f"{', '.join(missing)} (fail closed)"
)
capabilities_json = json.dumps(clean.get("capabilities") or [])
pending_json = json.dumps(clean.get("pending_mutation") or {})
evidence_json = json.dumps(clean.get("evidence") or {})
now_s = _ts()
with self._tx() as conn:
existing = conn.execute(
"""
SELECT * FROM session_checkpoints
WHERE remote = ? AND org = ? AND repo = ?
AND session_id = ? AND work_kind = ? AND work_number = ?
""",
(remote, org, repo, session_id, kind_norm, number_norm),
).fetchone()
if existing is None:
checkpoint_id = uuid.uuid4().hex
conn.execute(
"""
INSERT INTO session_checkpoints(
checkpoint_id, remote, org, repo, session_id,
provider_identity, role, work_kind, work_number,
worktree_path, branch, head_sha, capabilities,
lease_id, assignment_id, workflow_stage,
last_completed_action, current_operation,
pending_mutation, evidence, blocker, next_valid_action,
recovery_instructions, checkpoint_schema_version,
status, created_at, updated_at
) VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?,
?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)
""",
(
checkpoint_id, remote, org, repo, session_id,
clean["provider_identity"], clean["role"], kind_norm,
number_norm, clean["worktree_path"], clean["branch"],
clean["head_sha"], capabilities_json, clean["lease_id"],
clean["assignment_id"], clean["workflow_stage"],
clean["last_completed_action"], clean["current_operation"],
pending_json, evidence_json, clean["blocker"],
clean["next_valid_action"], clean["recovery_instructions"],
SCHEMA_VERSION, clean["status"], now_s, now_s,
),
)
else:
checkpoint_id = str(existing["checkpoint_id"])
conn.execute(
"""
UPDATE session_checkpoints
SET provider_identity = ?, role = ?, worktree_path = ?,
branch = ?, head_sha = ?, capabilities = ?,
lease_id = ?, assignment_id = ?, workflow_stage = ?,
last_completed_action = ?, current_operation = ?,
pending_mutation = ?, evidence = ?, blocker = ?,
next_valid_action = ?, recovery_instructions = ?,
checkpoint_schema_version = ?, status = ?,
updated_at = ?
WHERE checkpoint_id = ?
""",
(
clean["provider_identity"], clean["role"],
clean["worktree_path"], clean["branch"], clean["head_sha"],
capabilities_json, clean["lease_id"], clean["assignment_id"],
clean["workflow_stage"], clean["last_completed_action"],
clean["current_operation"], pending_json, evidence_json,
clean["blocker"], clean["next_valid_action"],
clean["recovery_instructions"], SCHEMA_VERSION,
clean["status"], now_s, checkpoint_id,
),
)
prior_stage = str(existing["workflow_stage"] or "")
new_stage = str(clean["workflow_stage"] or "")
if prior_stage != new_stage:
conn.execute(
"""
INSERT INTO events(work_item_id, event_type, message, created_at)
VALUES (NULL, 'session_checkpoint_stage_change', ?, ?)
""",
(
f"checkpoint {checkpoint_id} session {session_id} "
f"stage {prior_stage or '(none)'} -> "
f"{new_stage or '(none)'}",
now_s,
),
)
row = conn.execute(
"SELECT * FROM session_checkpoints WHERE checkpoint_id = ?",
(checkpoint_id,),
).fetchone()
return self._session_checkpoint_row(row) or {}
@staticmethod
def _session_checkpoint_row(row: sqlite3.Row | None) -> dict[str, Any] | None:
"""Return a stored checkpoint as a plain dict with JSON fields decoded."""
if row is None:
return None
record = dict(row)
for field, empty in (
("capabilities", []),
("pending_mutation", {}),
("evidence", {}),
):
raw = record.get(field)
try:
record[field] = json.loads(raw) if raw else empty
except (TypeError, ValueError):
# A row written by an older/foreign writer must not break reads.
record[field] = {"unparsed": str(raw)}
return record
def get_session_checkpoint(
self,
*,
remote: str,
org: str,
repo: str,
session_id: str,
work_kind: str | None = None,
work_number: int | None = None,
) -> dict[str, Any] | None:
"""Return the current checkpoint for one session's work unit, or None."""
if work_kind and str(work_kind).strip():
kind_norm = dependency_graph.normalize_work_kind(work_kind)
number_norm = int(work_number) if work_number is not None else 0
else:
kind_norm = ""
number_norm = 0
with self._tx(immediate=False) as conn:
row = conn.execute(
"""
SELECT * FROM session_checkpoints
WHERE remote = ? AND org = ? AND repo = ?
AND session_id = ? AND work_kind = ? AND work_number = ?
""",
(remote, org, repo, session_id, kind_norm, number_norm),
).fetchone()
return self._session_checkpoint_row(row)
def list_session_checkpoints(
self,
*,
remote: str | None = None,
org: str | None = None,
repo: str | None = None,
session_id: str | None = None,
work_kind: str | None = None,
work_number: int | None = None,
status: str | None = None,
limit: int = 500,
) -> list[dict[str, Any]]:
"""Return stored checkpoints, filtered. Newest updated first."""
clauses: list[str] = []
params: list[Any] = []
if remote:
clauses.append("remote = ?")
params.append(remote)
if org:
clauses.append("org = ?")
params.append(org)
if repo:
clauses.append("repo = ?")
params.append(repo)
if session_id:
clauses.append("session_id = ?")
params.append(session_id)
if work_kind:
clauses.append("work_kind = ?")
params.append(dependency_graph.normalize_work_kind(work_kind))
if work_number is not None:
clauses.append("work_number = ?")
params.append(int(work_number))
if status:
clauses.append("status = ?")
params.append(status)
sql = "SELECT * FROM session_checkpoints"
if clauses:
sql += " WHERE " + " AND ".join(clauses)
sql += " ORDER BY updated_at DESC LIMIT ?"
params.append(int(limit))
with self._tx(immediate=False) as conn:
rows = conn.execute(sql, params).fetchall()
return [
record
for record in (self._session_checkpoint_row(r) for r in rows)
if record
]
@staticmethod
def reconcile_session_checkpoint(
checkpoint: dict[str, Any],
*,
live_head_sha: str | None = None,
live_lease_active: bool | None = None,
live_lease_id: str | None = None,
) -> dict[str, Any]:
"""Diagnose a stored checkpoint against live Git/Gitea/lease state.
Pure (no DB) and **never restores** — it returns a diagnosis and the
caller decides. A head that no longer matches, or a lease that is gone
or reassigned, marks the checkpoint stale so resumption reconciles
instead of blindly restoring stale ownership (#660 AC3).
A ``None`` live input means "unknown, not checked" and never flags a
mismatch on its own.
"""
stored_head = str(checkpoint.get("head_sha") or "")
stored_lease = str(checkpoint.get("lease_id") or "")
head_mismatch = bool(
live_head_sha is not None
and stored_head
and stored_head != str(live_head_sha)
)
lease_mismatch = False
if stored_lease:
if live_lease_active is False:
lease_mismatch = True
elif live_lease_id is not None and str(live_lease_id) != stored_lease:
lease_mismatch = True
stale = head_mismatch or lease_mismatch
return {
"checkpoint_id": checkpoint.get("checkpoint_id"),
"session_id": checkpoint.get("session_id"),
"stale": stale,
"head_mismatch": head_mismatch,
"lease_mismatch": lease_mismatch,
"stored_head_sha": stored_head,
"live_head_sha": None if live_head_sha is None else str(live_head_sha),
"stored_lease_id": stored_lease,
"live_lease_id": None if live_lease_id is None else str(live_lease_id),
"reconcile_action": "reconcile_required" if stale else "safe_to_resume",
}
+2 -1
View File
@@ -247,7 +247,8 @@ def bootstrap_permits_control_checkout(
"""
if not isinstance(assessment, dict):
return False
if not is_create_issue_task(task):
import author_issue_bootstrap
if not is_create_issue_task(task) and not author_issue_bootstrap.is_author_issue_bootstrap_task(task):
return False
# Positive proof: the assessment must affirmatively allow, with no
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
@@ -14,12 +14,13 @@
| Capability | Notes |
|------------|--------|
| Schema | `sessions`, `work_items`, `leases`, `assignments`, `terminal_locks`, `events`, `incident_links` |
| Schema | `sessions`, `work_items`, `leases`, `assignments`, `terminal_locks`, `events`, `incident_links`, `session_checkpoints` |
| Atomic assign+lease | `ControlPlaneDB.assign_and_lease` — one `BEGIN IMMEDIATE` transaction |
| Mutation gate | `require_valid_assignment` — live lease + allowed action + non-terminal work + non-stale head |
| Heartbeat / release / expire | Lease lifecycle helpers |
| Terminal-lock index | Routing signal for #600 (terminal path first) |
| `incident_links` | Provider-neutral link model for #612**not** assignable work; scope keys NULL-safe |
| `session_checkpoints` | Durable session resume state for #660 — redacted at write, reconciled (never restored) on boot |
## Hard rules (enforced in code)
@@ -92,6 +93,44 @@ Module: `incident_bridge.py` · MCP tools: `gitea_observability_*`
- Provider tokens never appear in issue bodies, links, or tool results.
- Allocator sees bridge work only after a Gitea issue exists.
## Session checkpoints (#660)
Module: `control_plane_db.py` · Table: `session_checkpoints` · Parent **#655** · Soft-depends **#659** · Vision **#652** · Roadmap **#653**
Workflow state that lived only in MCP process memory or chat did not survive a restart, so session identity, stage, lease ownership, and the next valid action had to be reconstructed by hand. The `session_checkpoints` table makes that state durable.
### Schema version
Rows carry `checkpoint_schema_version` (currently **5**, tracking the module-level `SCHEMA_VERSION`), so a reader can tell which field set a record was written under. The row identity is
`UNIQUE (remote, org, repo, session_id, work_kind, work_number)` — one current-state row per session per work unit, upserted rather than appended. A session-level checkpoint that is not bound to an issue or PR uses the `work_number = 0` sentinel with an empty `work_kind`.
| Group | Columns |
|-------|---------|
| Identity | `session_id`, `provider_identity`, `role`, `remote`/`org`/`repo` |
| Work unit | `work_kind` (`issue`/`pr`/empty), `work_number`, `worktree_path`, `branch`, `head_sha` |
| Ownership | `capabilities` (JSON), `lease_id`, `assignment_id` |
| Progress | `workflow_stage`, `last_completed_action`, `current_operation`, `pending_mutation` (JSON) |
| Recovery | `evidence` (JSON), `blocker`, `next_valid_action`, `recovery_instructions` |
| Bookkeeping | `checkpoint_schema_version`, `status`, `created_at`, `updated_at` |
### API
| Method | Purpose |
|--------|---------|
| `write_session_checkpoint` | Upsert the current-state row; redacts, and optionally enforces drain completeness |
| `get_session_checkpoint` | Read one checkpoint by session + work unit |
| `list_session_checkpoints` | List checkpoints by session or by work unit |
| `reconcile_session_checkpoint` | Pure diagnosis of a stored checkpoint against live head/lease state |
| `checkpoint_completeness` | Pure — the drain-required fields missing from a record |
### Hard rules
1. **Redaction at write.** Free-text columns are redaction-filtered and JSON columns are redacted recursively before storage, so a pasted token or credential URL cannot land in a checkpoint. Secrets are never stored (#660 AC4).
2. **Reconcile, never blind-restore.** `reconcile_session_checkpoint` returns a diagnosis (`stale`, `head_mismatch`, `lease_mismatch`, `reconcile_action`) and the caller decides. A stored head that no longer matches live Git, or a lease that is gone or reassigned, marks the checkpoint stale (#660 AC3).
3. **Unknown live state is not a mismatch.** A `None` live input means "not checked" and never flags staleness on its own.
4. **Drain fails closed.** A write with `require_complete=True` refuses when any of `session_id`, `role`, `workflow_stage`, `next_valid_action`, `recovery_instructions` is missing or blank, so drain cannot complete on a checkpoint that could not resume.
5. **No transcript storage.** Checkpoints hold resume state, not conversation history.
## Non-goals (intentionally deferred)
- Full unsupervised watchdog auto-filing (prefer explicit reconcile first)
+223
View File
@@ -0,0 +1,223 @@
# ADR: MCP restart governance and authorization policy
- **Status:** Accepted (policy effective immediately for LLM and operator sessions; enforcement tooling may lag)
- **Date:** 2026-07-23
- **Tracking issue:** [#656](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/656)
- **Policy version:** `restart-governance/v1`
- **Related:**
- Umbrella: [#655](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/655) — governed MCP restart coordination and zero-disruption recovery
- Vision: [#652](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/652) — MCP Control Plane Web Console product vision (§A system health and process control)
- Roadmap: [#653](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/653) — Control Plane Web Console phased delivery (Phase 2 restart controls)
- Contamination guard: [#630](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/630) — blocks manual process-kill recovery
- Console restart UX: [#642](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/642) — sanctioned restart and graceful reload
- Existing restart / reconnect paths to inventory: [#591](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/591) — auto-restart on master advance (closed); [#584](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/584) — host auto-reconnect on transport flap
- Stable-control runtime split: `docs/architecture/mcp-stable-control-runtime-policy-adr.md` (#615)
- Client-namespace health: `docs/mcp-namespace-health.md` (#543)
- Reconnect-only EOF recovery: `docs/mcp-namespace-eof-recovery.md`
## 1. Context
The Gitea MCP server is the **control plane** for real issue and PR mutations
(create, comment, lock, review, merge, reconcile). The same process serves every
role namespace (`gitea-author`, `gitea-reviewer`, `gitea-merger`,
`gitea-reconciler`, `gitea-controller`) and holds the in-memory capability-gate
code loaded at startup.
Restarting that process is destructive to concurrent work:
- It resets every session's identity, preflight, and capability-lease binding.
- It can interrupt a mutation mid-critical-section (a lock acquire, a review
submit, a merge), leaving durable state half-written.
- Relaunching from the wrong checkout or worktree silently changes which code
the control plane runs, defeating master-parity gates (#420 / #615).
Today there is **no durable written policy** stating who may restart MCP, under
what conditions, that restart is a last resort, and how controller approval,
automated safety gates, and break-glass interact. Operators and LLM sessions
therefore invent restart behavior ad hoc, which makes concurrent multi-role work
unsafe. #630 and #642 need this policy as their backbone.
This ADR defines that policy. It does **not** implement coordinator code or HA
multi-instance restart (those are later children of #655).
## 2. Decision
### 2.1 v1 decision (recorded)
**Restart authority in v1 is `controller approval + automated safety gates`.**
A restart of the stable control runtime is authorized only when **both** hold:
1. A **controller** role explicitly approves the restart, recording an audit
entry (who, why, scope, affected sessions), **and**
2. The **automated safety gates** pass: a completed drain acknowledgement (no
affected session is mid-critical-section) or a declared break-glass incident
(§2.5).
Quorum among multiple controllers is **not** required day-one. It is deferred
unless a later investigation (tracked under #653) proves single-controller
approval is insufficient. This ADR records the v1 decision so enforcement code
(#630) has a fixed target; changing it requires a superseding ADR.
### 2.2 Restart is a last resort — the recovery ladder
Restart is the **last** rung. Before any restart, exhaust the narrower
recoveries, in order:
1. **Reconnect** the IDE/client MCP namespace (transport EOF, `client is
closing: EOF`, transient `#584` flap). No process change. See
`docs/mcp-namespace-eof-recovery.md`.
2. **Refresh / rebind** the session workspace: re-run `gitea_whoami`,
`gitea_resolve_task_capability`, and pass an explicit validated
`worktree_path`. Fixes stale session context without touching the process.
3. **Scoped restart** of a single misbehaving namespace/service (where the
deployment supports per-service restart) rather than the whole control plane.
4. **Full restart** of the stable control runtime process — operator-owned,
controller-approved, drained.
5. **Host / infrastructure restart** — the broadest action; same authorization
as a full restart plus infrastructure ownership.
A session **must** try rungs 12 and record why they were insufficient before
requesting a restart at rung 3 or above. Skipping straight to restart is a
policy violation.
### 2.3 Authorization matrix
| Role | Reconnect (1) | Refresh/rebind (2) | Scoped restart (3) | Full restart (4) | Host restart (5) |
|---|---|---|---|---|---|
| **author** | self | self | request only | **forbidden** | forbidden |
| **reviewer** | self | self | request only | **forbidden** | forbidden |
| **merger** | self | self | request only | **forbidden** | forbidden |
| **reconciler** | self | self | request only | **forbidden** | forbidden |
| **controller** | self | self | **approve** (+gates) | **approve** (+gates) | request to operator |
| **operator** | self | self | execute (controller-approved) | execute (controller-approved) | execute (controller-approved) |
| **admin** | self | self | execute | execute | execute (break-glass) |
Legend: *self* = may perform for its own client session; *request only* = may
raise a restart request but not authorize or execute it; *approve* = may
authorize under §2.1 gates; *execute* = may perform the process action after the
authorization is recorded.
Key invariants:
- **No LLM worker role (author/reviewer/merger/reconciler) may perform or
authorize a full or host restart.** They may only reconnect/rebind their own
client and file a restart request.
- **Controller approval authorizes; operator/admin executes.** The approving
controller and the executing operator may be the same human, but both the
approval and the execution are audited.
- Privileged process actions (full restart, host restart) are reserved to
**operator/admin**, never to an automated worker.
### 2.4 Approved conditions
A restart at rung 3+ is approved only under one of these recorded conditions:
- **No affected sessions:** the control plane has no live session that would be
interrupted (verified, not assumed).
- **Full drain acknowledged:** every affected session has drained
(no open critical section — no held mutation lease mid-write) and the drain is
acknowledged in the audit record.
- **Controller + gates:** controller approval plus passing automated safety
gates (§2.1), the standard v1 path.
- **Quorum:** not required in v1; reserved for a future superseding ADR.
- **Break-glass:** an incident-backed emergency exception (§2.5).
Restart **never** bypasses mutation gates mid-critical-section. Drain before
restart is mandatory except under break-glass with a declared incident.
### 2.5 Break-glass
Break-glass is a **separate, narrower** authorization path for emergencies where
the normal drain-and-approve path cannot complete (e.g. the control plane is
wedged and cannot drain).
Break-glass conditions:
- A declared incident record exists (id, timestamp, declarer) **before** the
action.
- The action is taken by **operator or admin** authority only — never by an LLM
worker role, and never unilaterally by an operator with active peers when a
controller is reachable.
- The scope is the minimum necessary rung of the ladder.
- A **mandatory post-hoc audit** entry is filed: what was restarted, why the
normal path was impossible, which sessions were affected, and the incident id.
Break-glass suspends the drain requirement, not the audit requirement.
### 2.6 Explicit prohibitions
- **A unilateral LLM or operator full restart while active peer sessions
exist is forbidden.** An LLM worker role must not kill, restart, or relaunch
the MCP process; a lone operator must not full-restart over live peer work
without controller approval or a break-glass incident.
- Process-kill recovery is forbidden as a routine tool (#630). This ADR does not
introduce a kill path.
- Ambiguous policy state **denies** restart (§4).
## 3. Security requirements
- Full restart and host restart are **privileged**; only operator/admin execute
them, only after a controller approval or break-glass incident is recorded.
- Break-glass is a distinct authorization path with its own audit mandate; it is
never the default and never silent.
- **Every approval and every restart action is audited** (who approved, who
executed, scope, affected sessions, condition, policy version). No restart is
authorized without a durable audit entry.
## 4. Failure behavior
**Ambiguous policy → deny restart.** If it cannot be established that a
restart is authorized under §2 — unknown affected-session state, missing
controller approval, absent break-glass incident, or an unclassifiable request —
the safe action is to **refuse** the restart and stop with a recovery report,
never to restart on assumption.
## 5. Policy IDs (for enforcement code)
Enforcement code — the restart coordinator (a later child of #655), the #630
contamination guard, and the #642 console restart UX — binds to these stable
policy identifiers rather than to prose:
| Policy ID | Statement |
|---|---|
| `RG-01` | Restart is last resort; rungs 12 must be tried and recorded first (§2.2). |
| `RG-02` | v1 authority = controller approval + automated safety gates (§2.1). |
| `RG-03` | No LLM worker role performs or authorizes full/host restart (§2.3). |
| `RG-04` | Full/host restart executed by operator/admin only, post approval (§2.3). |
| `RG-05` | Drain before restart is mandatory except break-glass with incident (§2.4). |
| `RG-06` | Break-glass requires a pre-declared incident and post-hoc audit (§2.5). |
| `RG-07` | Unilateral LLM/operator full restart with active peers is forbidden (§2.6). |
| `RG-08` | Ambiguous policy state denies restart (§4). |
The `restart-governance/v1` **policy version** field is emitted on future
restart audit events so approvals can be reconciled against the policy revision
in force.
## 6. Dogfooding
Gitea-Tools governs its own MCP control plane by this policy. Author, reviewer,
merger, and reconciler sessions operating on this repository use the recovery
ladder (§2.2) — reconnect and rebind, never self-restart — and any real restart
of the Gitea-Tools stable control runtime follows the controller-approval +
drain path defined here.
## 7. Acceptance and cross-links
This ADR is the authoritative restart-governance policy. It **must** stay
cross-linked from the safety model and the web-console deployment boundary:
- `docs/safety-model.md` § Process restart governance references this ADR.
- `docs/webui-deployment.md` references this ADR for restart/reload disposition.
It is linked to its issue lineage — umbrella **#655**, vision **#652**, roadmap
**#653**, contamination guard **#630**, and console restart UX **#642** — in
§ Related above.
## 8. Non-goals
- Implementing the restart coordinator or approval state machine (#630, later
children of #655).
- Implementing HA multi-instance restart or quorum machinery.
- Introducing any process-kill or auto-restart tool; existing auto-restart
behavior must be inventoried before any new restart tool is enabled.
+95
View File
@@ -0,0 +1,95 @@
# MCP restart coordinator and impact analysis (#658)
Before any sanctioned MCP restart, a central coordinator evaluates the live
control-plane state and produces an **impact preview** so operators and the web
console (#642 / #652) can see the blast radius *before* concurrent LLM work is
disrupted. Uncoordinated restarts destroy in-flight author/reviewer/merger work
and give operators no way to see what they are about to break.
This lands the coordinator + impact DTO + a dry-run MCP tool. It is the single
sanctioned entry point for restart evaluation post-#657 (which inventoried the
restart/reload/kill paths). The **mutative apply** path — actually performing a
restart — is a later child gated by a drain proof and is explicitly out of
scope here.
## Components
| Piece | Where | Responsibility |
|-------|-------|----------------|
| `restart_coordinator.evaluate_restart_impact` | `restart_coordinator.py` | Pure classification: inventory → impact report DTO. No I/O, no restart. |
| `RestartImpactReport` / `SessionImpact` / `LeaseImpact` | `restart_coordinator.py` | Console-facing DTO (`.as_dict()` is JSON-serializable). |
| `ControlPlaneDB.list_sessions` | `control_plane_db.py` | Read-only session inventory (the process-level unit a restart kills). |
| `gitea_request_mcp_restart` | `gitea_mcp_server.py` | MCP tool: gathers inventory from the #613 DB, calls the coordinator, returns the report. Dry-run only. |
## Dimensions evaluated
The coordinator classifies the inventory across the dimensions #658 requires:
- **Sessions** — every active MCP session; a restart terminates all of them.
Liveness = `status == active` **and** the owner pid is alive **and** the
heartbeat is fresh (default window 15 min). Dead/stale sessions do not count
toward blast radius.
- **Leases / locks** — control-plane leases joined with work items and their
freshness (`lease_lifecycle.classify_lease_freshness`). Only `active` (live
owner) leases are *disruptive*; expired / released / dead-process leases never
withhold a restart.
- **Issue / PR work** — the issues and PRs behind disruptive leases.
- **Mutations / critical sections** — a live lease carrying an author worktree
or a mutating phase (`implementing`, `publishing`, `merging`, …) is a
critical section a restart must not sever.
- **Terminal (merge) lock** — an active terminal lock always makes a restart
unsafe.
- **Prior recovery attempts** — narrower recovery already tried (e.g. sanctioned
client reconnects) is echoed so the operator sees the escalation history.
## Verdict
Exactly three verdicts, matching the acceptance criteria:
| Verdict | `allow_restart` | Meaning |
|---------|-----------------|---------|
| `safe` | `true` | No other live sessions, no live leases, no terminal lock. |
| `unsafe` | `false` | Live work would be disrupted and no operator override is present — **or** the inventory could not be completed (fail closed). |
| `override` | `true` | Live work present, but an operator override accepts the blast radius. |
`override_would_allow` tells the console whether an override path exists for the
current state. `blast_radius` is a `none` / `low` / `medium` / `high` severity
band derived from the affected session and work counts.
### Fail closed
If the control-plane inventory cannot be completed (DB unavailable, a listing
failed), `inventory_complete` is `false` and the verdict is `unsafe` / deny. An
incomplete evaluation must never green-light a restart.
### Operator override authority
Override authority is read from the environment variable
`GITEA_OPERATOR_RESTART_OVERRIDE_AUTHORIZATION` and **never** from a tool
argument. A worker session cannot set an environment variable on an
already-running daemon, so override cannot be self-asserted (same pattern as the
#630 daemon-maintenance authorization). The `request_override` tool argument only
expresses caller intent; it takes effect solely when the environment
authorization is present.
## The tool
```text
gitea_request_mcp_restart(remote, host, org, repo,
dry_run=True, request_override=False,
session_id=None, limit=200)
```
Read-only, dry-run, and it **never restarts anything**. `apply_supported` is
always `false`; passing `dry_run=False` performs no restart and reports that
apply is gated by a drain proof (a separate child).
## Audit
Every evaluation carries an `audit_record` (event, coordinator version, verdict,
allow decision, blast radius, counts, timestamp) so restart decisions are
auditable. No secrets flow through the coordinator — session ids, pids, and
profiles are operational metadata only.
A representative dry-run report is in
[`mcp-restart-impact-sample.json`](./mcp-restart-impact-sample.json).
+148
View File
@@ -0,0 +1,148 @@
{
"coordinator_version": "1.0.0-issue-658",
"evaluated_at": "2026-07-24T06:00:00+00:00",
"dry_run": true,
"restart_performed": false,
"inventory_complete": true,
"incomplete_reasons": [],
"verdict": "unsafe",
"allow_restart": false,
"override_would_allow": true,
"operator_override": false,
"blast_radius": "high",
"reasons": [
"live work would be disrupted; restart denied without operator override",
"1 critical section(s) in flight (active lease with a live owner)"
],
"affected_sessions": [
{
"session_id": "prgs-author-30988-d6f43c25",
"role": "author",
"profile": "prgs-author",
"pid": 1,
"status": "active",
"alive": true,
"heartbeat_stale": false,
"is_requester": false,
"live": true
},
{
"session_id": "prgs-reviewer-4157-0ce9",
"role": "reviewer",
"profile": "prgs-reviewer",
"pid": 1,
"status": "active",
"alive": true,
"heartbeat_stale": false,
"is_requester": true,
"live": true
}
],
"affected_leases": [
{
"lease_id": "lease-abc",
"session_id": "prgs-author-30988-d6f43c25",
"role": "author",
"phase": "implementing",
"freshness": "active",
"work_kind": "issue",
"work_number": 658,
"worktree_path": "/repo/branches/feat-issue-658",
"disruptive": true,
"is_mutation": true,
"is_critical_section": true
},
{
"lease_id": "lease-dead",
"session_id": "prgs-author-91485",
"role": "author",
"phase": "allocated",
"freshness": "stale_dead_process",
"work_kind": "issue",
"work_number": 651,
"worktree_path": null,
"disruptive": false,
"is_mutation": false,
"is_critical_section": false
}
],
"critical_sections": [
{
"lease_id": "lease-abc",
"session_id": "prgs-author-30988-d6f43c25",
"role": "author",
"phase": "implementing",
"freshness": "active",
"work_kind": "issue",
"work_number": 658,
"worktree_path": "/repo/branches/feat-issue-658",
"disruptive": true,
"is_mutation": true,
"is_critical_section": true
}
],
"affected_issues": [
658
],
"affected_prs": [],
"mutations": [
{
"lease_id": "lease-abc",
"session_id": "prgs-author-30988-d6f43c25",
"role": "author",
"phase": "implementing",
"freshness": "active",
"work_kind": "issue",
"work_number": 658,
"worktree_path": "/repo/branches/feat-issue-658",
"disruptive": true,
"is_mutation": true,
"is_critical_section": true
}
],
"terminal_lock": null,
"ack_state": {
"prgs-author-30988-d6f43c25": "pending"
},
"prior_recovery_attempts": [
{
"kind": "client_reconnect",
"at": "2026-07-24T06:00:00+00:00",
"outcome": "insufficient"
}
],
"counts": {
"sessions_total": 2,
"sessions_live_other": 1,
"leases_total": 2,
"leases_disruptive": 1,
"critical_sections": 1,
"mutations": 1,
"affected_issues": 1,
"affected_prs": 0,
"prior_recovery_attempts": 1
},
"audit_record": {
"event": "restart_impact_evaluated",
"coordinator_version": "1.0.0-issue-658",
"evaluated_at": "2026-07-24T06:00:00+00:00",
"dry_run": true,
"operator_override": false,
"requesting_session_id": "prgs-reviewer-4157-0ce9",
"inventory_complete": true,
"verdict": "unsafe",
"allow_restart": false,
"blast_radius": "high",
"counts": {
"sessions_total": 2,
"sessions_live_other": 1,
"leases_total": 2,
"leases_disruptive": 1,
"critical_sections": 1,
"mutations": 1,
"affected_issues": 1,
"affected_prs": 0,
"prior_recovery_attempts": 1
}
}
}
+90
View File
@@ -0,0 +1,90 @@
# MCP restart / reload / kill path inventory (#657)
Complete inventory of every code, script, and host path that can **restart,
reload, reconnect, kill, or force-recreate** an MCP process in this project,
with each path classified and linked to the guard that constrains it.
This document is the human-readable companion to the machine-readable registry
in [`mcp_restart_paths.py`](../mcp_restart_paths.py). The two are kept in
lock-step by [`tests/test_mcp_restart_paths.py`](../tests/test_mcp_restart_paths.py):
every `path_id` below must appear in this file, and the source guards are run
against the live tree.
Roadmap linkage: this inventory is the enumeration step of the restart
governance work — parent **#655**, restart-governance ADR **#656**, vision
**#652**, roadmap **#653**. Related detection/guard work: master-advance
staleness **#591**/**#420**, side-effect-free resolver **#685**, transport flap
**#584**, manual-kill contamination **#630**.
## Classifications
| Classification | Meaning |
|---|---|
| `sanctioned_narrow_recovery` | One-shot, safe-by-construction recovery that never targets the running daemon. |
| `guarded_fail_closed` | Detects a restart-requiring condition, then fails mutations closed and emits reconnect guidance. Never self-restarts. |
| `forbidden` | A workflow-safety violation; where an LLM tool could invoke it, it is marked contamination. |
| `removed` | A previously-existing unguarded restart primitive that has been deleted; a regression guard keeps it absent. |
| `host_residual` | Behavior owned by the host/IDE, outside this process's control. Documented, not code-guarded here. |
## The rule
**No component may perform an unguarded full restart of the MCP daemon.** The
in-process daemon (`gitea_mcp_server.py`, `mcp_server.py`,
`role_session_router.py`) must never replace or terminate its own process:
replacing the process after the host has wired up the stdio pipes desyncs the
JSON-RPC transport (observed with Antigravity/Cascade hosts). Recovery is owned
by the host/operator via a client reconnect — the daemon only ever *detects*
and *fails closed*.
## Inventory
| path_id | Classification | Mechanism | Guard | Refs |
|---|---|---|---|---|
| `cli_venv_bootstrap_execv` | sanctioned_narrow_recovery | CLI wrapper scripts re-exec into `venv/bin/python3` via `os.execv`, guarded by `sys.executable != venv_python`. | One-shot pre-import bootstrap; runs before any MCP transport exists and only when not already on the venv interpreter; idempotent guard prevents a re-exec loop. | #657 |
| `daemon_self_replacement` | forbidden | The daemon replacing/terminating its own process (`os.execv`/`os.kill`/`os._exit`) to reload code. | Forbidden by design; enforced against the source tree by `assert_no_daemon_self_replacement()`. | #657, #584 |
| `legacy_auto_restart_helper` | removed | A helper (`_trigger_mcp_auto_restart`) that actively restarted the server from the read-only resolver path. | Removed in #685; kept absent by `assert_auto_restart_helper_absent()`. | #685, #657 |
| `config_touch_reload` | removed | Touching (utime) the MCP client config to make the host reload the server. | Removed from the resolver in #685: stale detection is report-only, never mutating config, spawning threads, or calling `os._exit`. | #685, #657 |
| `master_advance_auto_restart` | guarded_fail_closed | On-disk master advancing past the running code. | `master_parity_gate` captures startup parity and blocks mutations while stale, emitting restart guidance; the process never self-restarts. | #420, #591, #657 |
| `stale_runtime_resolver_reconnect` | guarded_fail_closed | The capability resolver detecting a stale serving process. | Report-only (#685): returns `restart_required`/`stop_required` and an exact reconnect action; no restart, thread, config touch, or `os._exit`. | #685, #657 |
| `manual_daemon_kill` | forbidden | Shell kills of the daemon: `pkill -f mcp_server.py`, `killall`, broad `pkill -f python` sweeps, or `kill <pid>` of a daemon pid. | Forbidden (#630): `runtime_recovery_guard` classifies these as contamination and `gitea_record_daemon_process_kill_attempt` writes a durable marker that fails later mutations closed. Operator maintenance authorization is read only from the environment. | #630, #657 |
| `conflict_marker_infra_stop` | guarded_fail_closed | The daemon entrypoint scans for unresolved merge-conflict markers at startup and stops (`sys.exit(1)`). | Fail-closed startup stop, not a restart: the process exits and waits for the operator to resolve conflicts and relaunch; never loops. | #657 |
| `ide_client_reconnect` | host_residual | A manual `/mcp reconnect` (or equivalent host action) that recreates the MCP client connection. | Outside this process's control; the sanctioned recovery the gates point operators toward. No in-process code initiates it. | #584, #656, #657 |
| `profile_switch_runtime` | sanctioned_narrow_recovery | Switching the active execution profile at runtime (dynamic-profile mode). | In-process and restart-free: `runtime_switching_supported` is true, so a switch rebinds capability without recreating the process. | #656, #657 |
## Guards enforced in CI
`tests/test_mcp_restart_paths.py` asserts, against the live source tree:
1. **Registry well-formedness** — every path has a valid classification, a
non-empty guard description, references, and locations; ids are unique; all
five classifications are represented.
2. **Unknown restart attempts fail closed**
`assert_restart_attempt_registered()` raises `UnknownRestartPathError` for
any path id not in this inventory, so a novel/unnamed restart primitive
cannot slip through silently.
3. **Daemon never self-replaces**`assert_no_daemon_self_replacement()` scans
the daemon modules for `os.execv`/`os.kill`/`os._exit`/`os.abort` calls
(comment/docstring mentions are ignored) and finds none.
4. **Legacy helper stays removed**`assert_auto_restart_helper_absent()`
confirms `_trigger_mcp_auto_restart` has not returned.
5. **pkill stays forbidden** — a daemon `pkill` command still classifies as
contamination via `runtime_recovery_guard`.
## Residual host behaviors (outside process control)
* `/mcp reconnect` in the IDE/host — the sanctioned recovery for stale-runtime,
transport-flap (#584), and worktree-binding conditions. The daemon can only
emit guidance toward it.
* Host-level process management (the operator relaunching the daemon after a
fail-closed stop, or after resolving merge conflicts).
These are documented rather than code-guarded because the process cannot
observe or gate them from inside itself.
## Rollout
Per #657, guards are introduced flag-free as **regression assertions** (they
codify invariants that already hold) before any hard runtime block is layered
on. When the restart coordinator (#655/#656) lands, registered paths gain a
coordinator token/capability check; unregistered attempts already fail closed
today via `assert_restart_attempt_registered()`.
+3
View File
@@ -69,6 +69,7 @@ that gates each call, not which tools exist.
- `gitea_audit_worktree_cleanup`
- `gitea_authorize_reconciliation_cleanup_phase`
- `gitea_authorize_review_correction`
- `gitea_bootstrap_author_issue_worktree`
- `gitea_capability_stop_terminal_report`
- `gitea_capture_branches_worktree_snapshot`
- `gitea_check_pr_eligibility`
@@ -100,6 +101,7 @@ that gates each call, not which tools exist.
- `gitea_get_profile`
- `gitea_get_runtime_context`
- `gitea_get_shell_health`
- `gitea_heartbeat_issue_lock`
- `gitea_heartbeat_reviewer_pr_lease`
- `gitea_inspect_workflow_lease`
- `gitea_issue_irrecoverable_provenance_authorization`
@@ -135,6 +137,7 @@ that gates each call, not which tools exist.
- `gitea_release_merger_pr_lease`
- `gitea_release_reviewer_pr_lease`
- `gitea_release_workflow_lease`
- `gitea_request_mcp_restart`
- `gitea_resolve_task_capability`
- `gitea_resume_review_draft`
- `gitea_review_pr`
@@ -0,0 +1,143 @@
# Model Usage, Token Cost, Latency, and Workflow Analytics (Phase 4)
- **Tracking Issue:** [#651](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/651)
- **Parent Epic:** [#631](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/651)
- **Console Surface:** `/analytics`, `/api/v1/analytics`, `/api/v1/analytics/usage`
## 1. Overview
The Web Console Analytics module provides durable, aggregate visibility into **model usage, token cost, latency percentiles, and workflow-stage performance** across projects, worker roles, AI models, issues, and PRs.
### Non-Goals
- No mandatory client-side telemetry that leaks prompts or secret keys.
- No third-party payment provider or billing integration.
- No automatic model routing changes without controller policy (#647).
---
## 2. Event Schema (`usage_events`)
Usage metrics are stored in the control-plane database under table `usage_events`.
| Column | Type | Description |
|---|---|---|
| `usage_id` | `INTEGER` | Primary key (autoincrement) |
| `session_id` | `TEXT` | Optional active session identifier |
| `remote` | `TEXT` | Known Gitea instance (`dadeschools` or `prgs`) |
| `org` | `TEXT` | Repository owner / organization |
| `repo` | `TEXT` | Repository name |
| `project_id` | `TEXT` | Optional project identifier |
| `role` | `TEXT` | Active worker role (`author`, `reviewer`, `merger`, `reconciler`, `controller`) |
| `model` | `TEXT` | LLM model identifier (e.g. `gemini-3.6-flash`, `claude-3-5-sonnet`) |
| `issue_number` | `INTEGER` | Correlated Gitea issue number (optional) |
| `pr_number` | `INTEGER` | Correlated Gitea PR number (optional) |
| `stage` | `TEXT` | Workflow stage (`preflight`, `implementation`, `review`, `merge`, `reconciliation`) |
| `input_tokens` | `INTEGER` | Input token count (optional / nullable) |
| `output_tokens` | `INTEGER` | Output token count (optional / nullable) |
| `total_tokens` | `INTEGER` | Total token count (optional / nullable) |
| `estimated_cost_usd` | `REAL` | Estimated USD cost (optional / nullable) |
| `latency_ms` | `INTEGER` | Request latency in milliseconds (optional / nullable) |
| `duration_ms` | `INTEGER` | Stage execution duration in milliseconds (optional / nullable) |
| `status` | `TEXT` | Outcome status (`success`, `failure`, `timeout`) |
| `metadata` | `TEXT` | Redacted metadata or summary string |
| `created_at` | `TEXT` | ISO 8601 UTC timestamp |
---
## 3. Handling of Missing Data ("Unknown" vs. Zero Fabrication)
To ensure operational metrics accurately reflect evidence:
- **Untracked or missing metrics are displayed as `Unknown`**, never zero-fabricated.
- If an event omits `estimated_cost_usd`, `latency_ms`, or token counts, the aggregator marks those fields as missing (`None`) rather than defaulting to `0` or `$0.00`.
- Summary tables and KPI cards explicitly indicate when data is unmeasured or partially reported.
---
## 4. Redaction & Security Rules
Per `#633` security policy:
- Free-text fields (`metadata`, `prompt_summary`, `session_id`) are run through `console_redaction.redact_text` before persistence and output serialization.
- Secret tokens, keychain commands, authorization headers, passwords, and JWTs are stripped automatically.
---
## 5. Opt-in Instrumentation Guide
Applications, MCP servers, and background sessions can report usage metrics through either Python API or HTTP ingestion.
### Python Ingestion
```python
from webui.analytics_loader import record_usage
record_usage(
remote="dadeschools",
org="Scaled-Tech-Consulting",
repo="Gitea-Tools",
role="author",
model="gemini-3.6-flash",
issue_number=651,
stage="implementation",
input_tokens=1420,
output_tokens=380,
total_tokens=1800,
estimated_cost_usd=0.00045,
latency_ms=320,
duration_ms=4500,
status="success",
metadata={"note": "Implementation of analytics module"},
)
```
### HTTP Ingestion API (authorized write)
`POST /api/v1/analytics/usage` is a **gated write**. It runs through
`console_authz` action `record_analytics_usage` (operator+, Phase 2 execution).
Unauthenticated or phase-inactive requests receive **403** and do not write.
Prefer in-process `record_usage` for MCP / session instrumentation.
```http
POST /api/v1/analytics/usage HTTP/1.1
Content-Type: application/json
# Requires authenticated principal with record_analytics_usage execution enabled
{
"remote": "dadeschools",
"org": "Scaled-Tech-Consulting",
"repo": "Gitea-Tools",
"role": "author",
"model": "gemini-3.6-flash",
"issue_number": 651,
"stage": "implementation",
"input_tokens": 1420,
"output_tokens": 380,
"total_tokens": 1800,
"estimated_cost_usd": 0.00045,
"latency_ms": 320,
"duration_ms": 4500,
"status": "success",
"metadata": "Analytics schema landed"
}
```
### Retention
`usage_events` is retained with hard caps applied on every write:
| Limit | Default |
|---|---|
| Max rows | 10,000 (`ControlPlaneDB.USAGE_EVENTS_MAX_ROWS`) |
| Max age | 90 days (`ControlPlaneDB.USAGE_EVENTS_RETENTION_DAYS`) |
Older rows (by `created_at`) and excess oldest rows (by `usage_id`) are deleted
after each insert so unbounded growth / DoS-by-volume cannot fill the DB.
---
## 6. Querying Analytics API
```http
GET /api/v1/analytics?role=author&stage=implementation HTTP/1.1
```
Returns `AnalyticsSnapshot` JSON containing aggregations (`by_model`, `by_stage`, `by_role`, `by_work_item`, `by_project`) and latency percentiles (`p50`, `p90`, `p95`, `p99`).
+56
View File
@@ -0,0 +1,56 @@
# Post-restart MCP reconciliation (#662)
After an MCP process restart, sessions, leases, capabilities, worktrees, and
interrupted mutations must be reconciled before operators claim a clean runtime.
This document describes the #662 completion-proof path.
## Components
| Piece | Where | Responsibility |
|-------|-------|----------------|
| `post_restart_reconcile.reconcile_after_restart` | `post_restart_reconcile.py` | Pure classification: inventory → completion proof DTO. No I/O. |
| `RestartCompletionProof` | `post_restart_reconcile.py` | Machine-readable proof (`.as_dict()` is JSON-serializable). |
| `gitea_reconcile_after_restart` | `gitea_mcp_server.py` | MCP tool: gathers inventory from the #613 control-plane DB + master-parity, classifies, returns the proof. Read-only. |
| Boot hook | `gitea_assess_master_parity` | First post-restart parity probe also runs reconcile once (log-only by default). |
## Dimensions
The assessor classifies:
- **service_health** — process healthy / parity mutation-safe
- **clients** — connected client descriptors (optional inventory)
- **sessions** — active session rows with dead owner pids are unresolved
- **checkpoints** — soft-depends on #660; skipped with reason when schema absent
- **leases** — live control-plane leases after restart
- **capabilities** — master-parity / stale-runtime (#610)
- **worktrees** — lease-bound paths missing on disk
- **interrupted_mutations** — mutating lease phases or explicit pending inventory; **never auto-resumed**
- **duplicates** — multiple live claims on the same work item
- **queue** — allocator resume safety
## Modes
| Mode | Env / arg | Behavior |
|------|-----------|----------|
| `log_only` (default) | unset or `GITEA_POST_RESTART_RECONCILE_MODE=log_only` | Proof only; `mutation_hold=false` |
| `enforce` | `GITEA_POST_RESTART_RECONCILE_MODE=enforce` or `mode=enforce` | Sets `mutation_hold=true` when overall status is degraded/failed or interrupted mutations remain |
## Follow-up issues
Unresolved dimensions produce `proposed_follow_ups` entries suitable for durable
Gitea issues. The MCP tool **does not create** those issues in v1 (rollout is
log-only first). Controllers may file them from the proof payload.
## Links
- Umbrella: #655
- Vision: #652 · Roadmap: #653
- Checkpoint schema: #660 (soft dependency)
- Drain proof: #661 (soft)
- This issue: #662
## Non-goals
- HA multi-instance failover
- Automatic silent mutation replay
- Implementing the #660 checkpoint schema itself
+14
View File
@@ -46,3 +46,17 @@ If shell helpers are unavailable and MCP commit cannot run, stop with a recovery
report (restart session, clear hung terminals, use MCP-native commit). See
[`llm-workflow-runbooks.md`](llm-workflow-runbooks.md) § MCP-native commit path
(#260) and agent temp artifact cleanup (#261).
## 7. Process restart governance
Restarting the MCP control-plane process is destructive to concurrent multi-role
work and is governed by a dedicated policy. Restart is a **last resort** behind
narrower recoveries (reconnect, rebind), full/host restart is reserved to
operator/admin under **controller approval + automated safety gates**, a
unilateral LLM or operator full restart with active peers is **forbidden**, and
ambiguous policy state **denies** restart. Break-glass is a separate,
incident-backed path with a mandatory audit.
See [`architecture/mcp-restart-governance.md`](architecture/mcp-restart-governance.md)
(#656) for the authorization matrix, the recovery ladder, break-glass
conditions, and the `RG-01``RG-08` policy IDs.
+122
View File
@@ -0,0 +1,122 @@
# Sanctioned restart and graceful reload controls (#642)
Sessions used to recover MCP connectivity by killing the host daemon
(`pkill -f mcp_server.py`, #630). That is forbidden and stays forbidden: it
kills every namespace on the host, contaminates whichever session survives, and
leaves no audit trail. This document describes the sanctioned replacement,
implemented in `webui/sanctioned_restart.py`.
## What the console will and will not do
The console **never** restarts anything. It authorizes an intent, records it,
and hands off to a host supervisor. There is no code path in which the console
sends a signal, spawns a process, or renders a kill command — a regression test
asserts the module contains no `subprocess`, `signal`, `os.kill`, `os.system`,
or `popen` reference, and that no returned payload contains a kill command.
## Operations
| Mode | Action | Minimum role | Behaviour |
|------|--------|--------------|-----------|
| `reload` | `system.reload_namespace` | controller | Host supervisor reloads the namespace in place, draining in-flight requests. |
| `restart` | `system.restart_namespace` | admin | Host supervisor restarts the namespace. In-flight requests are lost. |
Scope is always exactly one namespace. A fleet-wide restart is an explicit
non-goal: `all`, `*`, `fleet`, and an empty scope are refused with
`fleet_scope_not_permitted`, because that is precisely the blast radius the
forbidden kill already had. An unrecognised namespace is refused rather than
passed through to the host.
## The gate sequence
`assess_restart_request()` applies every gate in order and reports the first
failure with a stable reason code:
| Order | Gate | Reason code on failure |
|-------|------|------------------------|
| 1 | Mode is `restart` or `reload` | `unknown_mode` |
| 2 | Scope is a single known namespace | `fleet_scope_not_permitted`, `unknown_namespace` |
| 3 | Principal holds the required console role | `unauthorized` |
| 4 | Confirmation phrase supplied | `confirmation_required` |
| 5 | Confirmation names this namespace and mode | `confirmation_mismatch` |
| 6 | Out-of-band operator authorization present | `operator_authorization_missing` |
| 7 | Runtime is not contaminated | `contaminated_runtime` |
| 8 | Host restart hook configured | `restart_hook_not_configured` |
Passing every gate yields `host_action_required`, never "restarted".
### Confirmation binds the namespace
The required phrase is `"<mode> <namespace>"` — for example
`restart gitea-author`. Binding the namespace into the phrase is the point: a
confirmation typed for one namespace cannot be replayed against another.
### Operator authorization is not self-assertable
Host daemon maintenance is authorized out of band through
`GITEA_OPERATOR_DAEMON_MAINTENANCE_AUTHORIZATION`, read from the process
environment and nowhere else (#630; #710 finding F1). A worker session cannot
set an environment variable for an already-running daemon, so this cannot be
faked the way a tool argument could.
### The host hook
`GITEA_SANCTIONED_RESTART_HOOK` holds an opaque reference the *host* resolves —
a supervisor label such as a launchd job name, never a command line. With no
hook configured the request is refused; the console does not fall back to a
process kill. The value is read server-side and never rendered to a client.
## Manual kill remains contamination
`classify_restart_command()` classifies an operator-proposed recovery command.
A manual `pkill`/`kill`/`killall` of the MCP daemon is contamination, not a
restart: it returns `clean_claim_allowed: false` and builds a durable
contamination marker (redacted command only, never secrets) naming
`system.restart_namespace` as the sanctioned alternative.
A live, uncleared contamination marker also blocks a restart. This is stricter
than #630's task-scoped gate, which deliberately lets a contaminated worker keep
commenting and handing off: restarting a contaminated runtime would launder the
contamination rather than resolve it. Clear the marker through the reconciler
path first.
## Post-restart health verification
After the host supervisor acts, `verify_post_restart_health()` decides whether
the session may claim to be clean:
| Status | Meaning | Clean claim |
|--------|---------|-------------|
| `clean` | Required tool callable, proven through the live client namespace | Allowed |
| `unproven` | Reported healthy without live client-namespace evidence | Refused |
| `unhealthy` | Probe failed | Refused |
Only `probe_source=client_namespace` evidence clears a session. Static tool
registration is not proof, and neither is an offline subprocess probe — an IDE
client can hold a registered tool list while live calls fail with
`client is closing: EOF` (see
[`mcp-namespace-health.md`](mcp-namespace-health.md)).
## Audit
Every attempt — allowed or denied — is recorded through
`webui.console_audit` with actor, target namespace, mode, result, and reason
code, and is redacted before it is persisted. `system.restart_namespace` is
break-glass, so its records are retained for 730 days. Records carry
`process_kill_executed: false`, which is a fact about the code path rather than
a claim: no such path exists.
## Environment variables
| Variable | Purpose |
|----------|---------|
| `GITEA_SANCTIONED_RESTART_HOOK` | Host supervisor reference; absent means restart is refused. |
| `GITEA_OPERATOR_DAEMON_MAINTENANCE_AUTHORIZATION` | Out-of-band operator authorization reference. |
| `WEBUI_AUDIT_LOG` | Console audit sink; absent means records are built but not persisted. |
## Non-goals
* No unrestricted `kill` from the UI, in any role, in any phase.
* No fleet-wide restart.
* No silent auto-restart loop: every attempt is confirmed and audited.
* This does not implement the Phase 1 health API (#634).
+12 -2
View File
@@ -91,6 +91,9 @@ already define, and a regression test asserts each mapping matches.
| `close_pr` | controller | privileged | `gitea.pr.close` | Yes | No | No | 3 |
| `merge_pr` | controller | privileged | `gitea.pr.merge` | Yes | **Yes** | **Yes** | 3 |
| `delete_branch` | admin | destructive | `gitea.branch.delete` | Yes | **Yes** | **Yes** | 3 |
| `record_analytics_usage` | operator | gated_write | `runtime.record_analytics_usage` | Yes | No | No | 2 |
| `system.reload_namespace` | controller | privileged | `runtime.reload_namespace` | Yes | No | No | 2 |
| `system.restart_namespace` | admin | destructive | `runtime.restart_namespace` | Yes | **Yes** | **Yes** | 2 |
**Dual control** means the acting principal may not be the sole authority: a
second distinct principal must confirm. **Break-glass** means the action is
@@ -102,6 +105,13 @@ honouring it.
`delete_branch` is admin-only rather than controller because it is the one
irreversible action in the set.
`system.restart_namespace` is admin-only for the same reason: restarting a
namespace drops every in-flight request on it. `system.reload_namespace` drains
first, so it is privileged but not destructive. Neither action is ever executed
by the console — both hand off to a host supervisor, and neither exposes a raw
process kill. See
[`sanctioned-restart-controls.md`](sanctioned-restart-controls.md) (#642).
### Authorization decision
`authorize(action_id, principal, for_execution=False)` returns a decision
@@ -199,8 +209,8 @@ breaking the request it describes.
| Class | Applies to | Default |
|-------|-----------|---------|
| `standard` | Routine gated writes | 90 days |
| `privileged` | `review_pr`, `close_pr`, and any unclassifiable action | 365 days |
| `break_glass` | `merge_pr`, `delete_branch` | 730 days |
| `privileged` | `review_pr`, `close_pr`, `system.reload_namespace`, and any unclassifiable action | 365 days |
| `break_glass` | `merge_pr`, `delete_branch`, `system.restart_namespace` | 730 days |
Each record carries its own class, day count, and computed `expires_at`, so
retention is auditable per record rather than inferred from file age. An
+9
View File
@@ -55,6 +55,15 @@ shipped to the browser.
assumption paths, and the client-secret policy. Use it to verify an instance is
configured for internal-only operation.
## Process restart / reload disposition
The console never exposes a restart or reload control; process restart of the
MCP control-plane runtime is governed separately. Restart is a last resort behind
reconnect/rebind, full restart is operator/admin-only under controller approval
plus safety gates, and break-glass is an incident-backed path. See
[`architecture/mcp-restart-governance.md`](architecture/mcp-restart-governance.md)
(#656).
## Non-goals (MVP)
- Full SSO or session login in the UI
+287 -1
View File
@@ -52,7 +52,9 @@ status, onboarding checklist state, and the fail-closed error payloads (#635).
| Path | Description |
|------|-------------|
| `/` | Home / operator overview |
| `/health` | JSON liveness (`status`, `service`, `mode`, `timestamp`) |
| `/health` | JSON liveness (`status`, `service`, `mode`, `timestamp`, `uptime_seconds`) |
| `/api/v1/system/health` | Structured read-only system health (#634) |
| `/system-health` | System-health dashboard — readiness, version/uptime, dependencies, MCP namespaces, stale-runtime parity (#639) |
| `/queue` | Live PR and issue queue dashboard (#429) |
| `/api/queue` | JSON queue export with pagination metadata |
| `/projects` | Project registry list with status and onboarding progress (#427, #635) |
@@ -73,11 +75,95 @@ status, onboarding checklist state, and the fail-closed error payloads (#635).
| `/api/actions/{id}/preview` | Mutation ledger preview (GET, read-only) |
| `/leases` | Lease and collision visibility (#433) |
| `/api/leases` | JSON lease/collision export |
| `/sessions` | Phase 1 shell stub — session inventory (backed by #636) |
| `/inventory` | Phase 1 shell stub — unified inventory (backed by #636) |
| `/timeline` | Phase 1 shell stub — workflow event timeline |
| `/policy` | Phase 1 shell stub — capability/role policy placeholder |
| `/insights` | Phase 1 shell stub — operational insights placeholder |
Most routes are GET-only. POST/PUT/PATCH/DELETE return `405` with
`read-only-mvp`, except `/audit` and `/api/audit` which accept POST for
local validator preview only (no Gitea mutations, no server-side storage).
## System health API (#634)
`GET /api/v1/system/health` is the structured, read-only health surface for
automated readiness checks. It is the first console API under the `/api/v1`
prefix; the unversioned MVP exports remain as compatibility aliases.
`/health` is unchanged for existing consumers — every MVP key is still present
— and now also carries `started_at`, `uptime_seconds`, and a
`system_health_api` pointer. It stays deliberately cheap and runs no dependency
probe, because answering readiness costs real work.
**Status codes.** `200` when ready, `503` when a required dependency failed or
was never probed. Automation can branch on the code without parsing the body.
**Query flags.** The Gitea check is a network call, so it is opt-in:
`GET /api/v1/system/health?deep=1` runs it and caches the result for
`WEBUI_HEALTH_PROBE_TTL_SECONDS` (default 15s) so dashboard polling does not
amplify into remote load. Without the flag that probe reports `skipped`.
**Dependencies.** `control_plane_db` and `repository` are required and drive
readiness. `gitea` is optional: when it fails the overall `status` degrades but
`readiness.ready` stays true, because local inventory is still serveable. Each
entry carries `status`, `detail`, `required`, and `latency_ms`.
Two honesty rules are worth knowing before reading the payload:
* `stale_runtime.mutation_safe` is true only when the runtime, checkout, and
remote-tracking commits are all known and equal. An unfetched remote is
reported as indeterminate, never as safe.
* `mcp_namespaces` entries are always `unproven`. A web process runs outside
the IDE-managed MCP client and cannot invoke a namespace tool, so per #543
only a `client_namespace` probe can prove that path.
Sample response (abridged, healthy):
```json
{
"status": "ok",
"service": "mcp-control-plane-webui",
"mode": "read-only",
"api": "/api/v1/system/health",
"timestamp": "2026-07-22T11:04:18.512034+00:00",
"readiness": { "ready": true, "complete": true, "reasons": [] },
"version": {
"git_sha": "620ed6e9a9550b8da2ceb82d9ab8744e8920490f",
"git_describe": "v1.1.0-898-g620ed6e",
"control_plane_schema_version": 4,
"python_version": "3.14.5",
"known": true
},
"process": { "started_at": "2026-07-22T10:58:02.114+00:00", "uptime_seconds": 376.4 },
"deep_probes_requested": false,
"dependencies": [
{
"name": "control_plane_db",
"kind": "sqlite",
"status": "ok",
"detail": "schema v4 readable",
"required": true,
"healthy": true,
"latency_ms": 1.482,
"metadata": { "schema_version": 4, "active_leases": 3 }
},
{ "name": "repository", "kind": "git", "status": "ok", "required": true, "healthy": true },
{ "name": "gitea", "kind": "http", "status": "skipped", "required": false, "healthy": false }
],
"mcp_namespaces": [
{ "namespace": "gitea-author", "required_tool": "gitea_whoami", "status": "unproven" }
],
"stale_runtime": { "stale": false, "determinable": true, "mutation_safe": true, "reasons": [] },
"probe_errors": []
}
```
No restart, reload, or process-kill control is exposed here: those are Phase 2
at the earliest, and #630 forbids process-kill recovery outright. Every probe
opens its subject read-only — the control-plane database is opened through a
`mode=ro` URI so a health check can never create or migrate a schema.
## Report audit (#431)
Paste an LLM final report at `/audit` or POST JSON to `/api/audit`. The UI
@@ -153,6 +239,57 @@ health, workflow/schema SHA-256 hashes, and stale-runtime warnings when the
checkout is behind merged safety-gate changes. Restart guidance links to #420;
no tokens or MCP restart actions are exposed.
## Application shell — Phase 1 (#638)
The console shell (`webui/layout.py`) renders a grouped navigation driven by a
single nav-config module, `webui/nav.py`. Nav groups follow the epic #631
Phase 1 information architecture: **Health, Traffic, Runtime/Sessions,
Projects, Inventory, Timeline, Policy** (placeholder), and **Insights**
(placeholder). Live views and Phase 1 placeholders (`stub`) are declared in one
place so the layout and the route table cannot drift.
The header carries two read-only status badges — an **environment** badge
(`local` for loopback binds, `remote` otherwise, derived from `WEBUI_HOST`) and
a **mode: read-only** badge — plus a **Docs** link to this document. No
privileged action controls are present in the Phase 1 shell.
Not-yet-implemented surfaces (`/sessions`, `/inventory`, `/timeline`,
`/policy`, `/insights`) resolve to graceful read-only stub pages instead of
404s; their backing views land in later child issues of #631 (the inventory
surfaces are backed by #636). Mutating methods on stub routes still fail closed
with `read-only-mvp`.
## System-health dashboard (#639)
`/system-health` renders the same snapshot the `/api/v1/system/health` API
returns, so the page and the API can never disagree. Cards: overall readiness,
stale-runtime parity, version and uptime, dependency probes, MCP namespaces,
probe errors (only when present), and recovery pointers. `?deep=1` opts into
the network probe exactly as the API does; the plain page load stays cheap.
Field authority and honesty rules:
* `ready` and `readiness_complete` are shown separately. A snapshot whose
required probes never ran is not the same as one that ran them and passed,
and the page never collapses the two into an unproven green.
* A probe that did not run appears under **Not probed**, never as healthy.
* `stale_runtime.mutation_safe` is displayed verbatim from the API. When the
runtime is stale, or when parity is indeterminate, the page warns and does
not claim mutation safety.
* MCP namespaces are reported `unproven`: the web process runs outside the
IDE-managed MCP client and cannot prove that path (#543).
Redaction is split by field kind. Free text — probe details, readiness and
parity reasons, probe errors — passes through `system_health.redact`.
Structured fields — commit SHAs, probe names, statuses, timestamps — are
HTML-escaped only, because `redact`'s opaque-token rule matches any run of 32
or more characters and would otherwise blank every 40-character git SHA, which
is precisely the evidence the parity view exists to show.
The dashboard is read-only: no restart, reload, or process-kill control. Those
arrive in Phase 2 (#642). Recovery guidance points at the sanctioned client
reconnect / operator restart path — never a manual daemon kill (#630).
## Deployment boundary (#435)
MVP serves on loopback by default. Binding `0.0.0.0` or `::` is **refused**
@@ -212,6 +349,155 @@ health, workflow/schema SHA-256 hashes, and stale-runtime warnings when the
checkout is behind merged safety-gate changes. Restart guidance links to #420;
no tokens or MCP restart actions are exposed.
## Inventory API (#636)
`GET /api/v1/inventory` returns one versioned, read-only snapshot that unifies
what the lease (#433), worktree (#432), and runtime (#430) MVP views each show
separately, so traffic-control and recovery consumers read the same source.
`GET /api/v1/inventory/{section}` returns a single section under the identical
schema (`sessions`, `leases`, `locks`, `worktrees`, `namespaces`); an unknown
section is a `404` with `error: unknown_section`. Both routes are `GET`-only.
Each section carries its own `status` (`ok` / `degraded` / `unavailable`), a
`reason` when not `ok`, and a `scan_ms`. A subsystem that cannot be read
degrades to a reasoned section; it never raises and never emits an empty list
that would read as "nothing is there".
### Field authority
Every section names where its rows came from; authorities are never blended.
| Section | Authority | Source |
|---|---|---|
| `sessions` | `control_plane_db` | #613 control-plane DB (`mode=ro`), authoritative for exclusive ownership (#600/#601) |
| `leases` | `control_plane_db` | #613 control-plane DB; degrades if the `work_items` table is absent |
| `locks` | `filesystem` | durable per-issue lock files (`issue_lock_store`) |
| `worktrees` | `filesystem` | registered git worktrees via the #432 hygiene scanner |
| `namespaces` | `filesystem` | the active profile serving this web process (others are not enumerable) |
The payload restates this map under `field_authority` for machine consumers.
### Ownership safety
`ownership_authority_complete` is true only when every ownership-bearing section
(`sessions`, `leases`, `locks`) read cleanly. While it is false, nothing is
reported as unowned and no collision is asserted from a degraded source —
absence of evidence is reported as absence of evidence, never as free work.
`collisions` surfaces detectable conflicts, each with a `kind` and `severity`:
`lock-without-worktree`, `duplicate-live-lock`, `live-lock-dead-owner` (unexpired
lease, dead pid — a #753 recovery candidate that would read as live to a naive
timestamp check), `stale-lock-dead-owner`, `expired-lock-live-owner` (the
#635/#760 daemon-pid deadlock), `concurrent-active-lease`, `active-lease-past-expiry`,
and `orphan-lease`. Collisions are emitted only from sections that read cleanly.
The control-plane DB is opened through a `mode=ro` URI so a read never creates
or migrates it; paths are collapsed against `$HOME`, URLs lose userinfo and
query strings, and credential-shaped values are redacted at the boundary. Lease
steal/release and worktree deletion are Phase 2+ and have no representation here.
## Workflow-event timeline (#637)
`GET /api/v1/timeline` is a read-only, versioned aggregation of workflow
events from every available source into one normalised, filterable stream. It
is the model layer for the Phase 1 timeline console view (a later child issue
of #631); this issue ships the schema, adapters, and read API only.
### Schema (versioned)
`webui/timeline.py` declares `TIMELINE_SCHEMA_VERSION` (currently `1`) and the
frozen `WorkflowEvent` record. Every response carries `schema_version` so a
consumer can branch on shape. One event:
```json
{
"source": "control_plane",
"event_type": "lease.renew",
"event_key": "cp:1421",
"timestamp": "2026-07-23T02:00:00Z",
"actor": null,
"role": null,
"issue_number": 637,
"pr_number": null,
"session_id": null,
"tool_name": null,
"decision": null,
"message": "lease renewed",
"correlation_id": "issue#637",
"evidence_refs": [],
"sensitive": true
}
```
`event_key` is stable and unique per source (`cp:<event_id>`,
`cth:<kind>:<number>:<comment_id>`), so pagination and dedup are deterministic.
### Sources and field authority
| Source | Adapter | Authority |
|---|---|---|
| Control-plane `events``work_items` | `adapt_cp_events` | `event_type`, `message`, `timestamp`, issue/PR scope come from the CP database, read through a `mode=ro` URI (never creates the DB or runs migrations) |
| Gitea Canonical Thread Handoff comments | `adapt_cth_comments` | `actor`, `role` (next owner), `decision`, `evidence_refs`, `timestamp` come from the parsed CTH comment body (`canonical_thread_handoff`) |
Handoff comments are thread-scoped: they are only read when the request filters
by a single `issue` or `pr`. Otherwise the handoff source reports `not run`
with a reason — it is never rendered as empty-and-healthy. Each source degrades
independently: an unavailable control-plane DB or a failed comment fetch is a
`sources[]` entry with `ok:false` and a `reason`, never a dropped timeline.
### Query parameters
`issue`, `pr`, `session` (conjunctive filters); `limit` (default 50, max 500)
and `offset` for pagination; `remote`, `org`, `repo` to override the default
registry-project scope. Events sort ascending by
`(timestamp, source_rank, event_key)`; missing timestamps sort last.
### Filter authority, and refusing what cannot be answered
A filter dimension is only meaningful for a source whose records carry it.
Each source declares its own support in `_SOURCE_FILTER_SUPPORT` and reports it
per response as `supported_filters` / `unsupported_filters`:
| Source | issue | pr | session |
|---|---|---|---|
| `control_plane` | yes | yes | **no** — the `events` table is `(event_id, work_item_id, event_type, message, created_at)` and records no session |
| `gitea_handoff` | yes | yes | yes — a CTH comment declares its own `Session:` field |
`session_id` is read only from that declared CTH field. It is never inferred
from a work item, an actor, or message text, and a value that is
redaction-altering or bare-secret-shaped is dropped rather than emitted.
When **no source that ran** can carry a requested dimension, the request is
refused rather than answered: the response is `422` with `ok:false` and a
structured `error` naming `unsupported_filters` and the per-source reason. A
`200` with zero events would tell an operator that no such activity exists,
which is a stronger — and false — claim than "this cannot be answered here".
A source that *can* answer the dimension and simply matched nothing still
returns `200` with `ok:true` and an empty page.
### Redaction
Every free-text field (event messages, decision/proof text, roles, actors) is
passed through the console redaction policy (`webui.console_redaction`, backed
by `gitea_audit.redact`) before it leaves the module, failing closed to the
placeholder. No unredacted tool arguments or secrets are ever emitted, and a
generation error never drops raw data to a caller or a log.
Redaction also runs *before* any structured value is derived from free text.
`evidence_refs` are extracted from already-redacted proof/decision text, and a
commit reference is recognised only where the text declares one (`commit`,
`head`, `base`, `sha`, …). An undeclared 40-character hex run has the exact
shape of a Gitea access token, so it is never lifted out of prose into a
structured field. Every reference is then independently revalidated against an
allowed shape and a second redaction pass immediately before serialization;
anything unproven is dropped and the event is flagged `sensitive`.
### Tests
```bash
pytest tests/test_webui_timeline.py -q
```
## Tests
```bash
+726
View File
@@ -0,0 +1,726 @@
"""Pre-restart drain proof and hard gate (#661).
A sanctioned MCP restart may only proceed after a machine-verifiable *drain
proof* attests that unsafe work is clear and checkpoints are complete. Drain
mode alone (#659) is not enough: without a proof, an apply path could still
restart on a stale or false "ready" claim, dropping mutations and orphaning
leases. This module defines the :class:`DrainProof` artifact, a fail-closed
verifier, and the hard gate the sanctioned restart-apply path must consult.
Design rules (mirror :mod:`restart_coordinator` / :mod:`lease_lifecycle`):
* **Pure classification.** Every function here operates on already-gathered
inputs and returns a structured result. Nothing touches the network, the
filesystem, or a live process, so multi-session fixtures drive every branch.
This module never restarts anything; the gate only *authorizes or denies*.
* **Fail closed.** A missing, expired, tampered, or unclean proof denies the
restart. Unknown checkpoint completeness is treated as *not complete*. An
incomplete impact report can never yield a clean proof.
* **Non-forgeable within the process.** The proof id is a keyed hash over the
canonical proof contents using a per-process secret. A worker session cannot
hand-craft a passing proof without that secret, and a proof minted in a prior
daemon process will not verify after a restart (the secret is regenerated).
* **No secrets leak.** The per-process secret never appears in a proof, an
``as_dict``, an audit record, or an incident descriptor.
Relationship to siblings (#655 umbrella):
* **#658** ``restart_coordinator.evaluate_restart_impact`` — produces the
blast-radius impact report this proof consumes ("what would a restart
disrupt?"). ``mutations`` / ``critical_sections`` being empty is what the
no-in-flight-mutations check verifies.
* **#659** graceful drain mode — performs the drain actions and calls
:func:`build_drain_proof` to mint the artifact once its checklist passes.
* **#660** durable session checkpoints — supplies checkpoint completeness.
Because that schema may not yet be present, completeness is an *input* here,
never a hard table dependency; unknown fails closed.
Non-goals (separate children): the emergency break-glass *workflow* (this gate
only leaves a sanctioned bypass hole authorized elsewhere), the console UI, and
the actual restart execution.
"""
from __future__ import annotations
import hashlib
import hmac
import json
import os
from dataclasses import dataclass
from datetime import datetime, timedelta, timezone
from typing import Any, Mapping, Sequence
DRAIN_PROOF_VERSION = "1.0.0-issue-661"
# Short default lifetime for a drain proof. A proof attests to a *point-in-time*
# drained state; live work can resume the moment drain mode relaxes, so the
# window in which a proof is honoured must be small (#661 security: short TTL).
DEFAULT_PROOF_TTL_SECONDS = 120
# Gate verdicts.
GATE_ALLOW = "allow"
GATE_DENY = "deny"
GATE_BREAK_GLASS = "break_glass"
# The mandatory drain checklist. A proof is *clean* only when every one of these
# checks passed. Names are stable identifiers surfaced in audit + incidents.
CHECK_NO_INFLIGHT_MUTATIONS = "no_inflight_mutations"
CHECK_ASSIGNMENTS_STOPPED = "assignments_stopped"
CHECK_CHECKPOINTS_COMPLETE = "checkpoints_complete"
CHECK_HANDOFFS_OK = "handoffs_ok"
CHECK_LEASES_HANDLED = "leases_handled"
CHECK_ACKS_OR_TIMEOUT = "acks_or_timeout"
REQUIRED_CHECKS: tuple[str, ...] = (
CHECK_NO_INFLIGHT_MUTATIONS,
CHECK_ASSIGNMENTS_STOPPED,
CHECK_CHECKPOINTS_COMPLETE,
CHECK_HANDOFFS_OK,
CHECK_LEASES_HANDLED,
CHECK_ACKS_OR_TIMEOUT,
)
def _utc_now() -> datetime:
return datetime.now(timezone.utc)
def _parse_ts(value: str | None) -> datetime | None:
if not value:
return None
text = str(value).strip()
if not text:
return None
if text.endswith("Z"):
text = text[:-1] + "+00:00"
try:
dt = datetime.fromisoformat(text)
except ValueError:
return None
if dt.tzinfo is None:
dt = dt.replace(tzinfo=timezone.utc)
return dt
def _canonical(payload: Any) -> str:
"""Deterministic JSON encoding for hashing (stable key order, no spaces)."""
return json.dumps(payload, sort_keys=True, separators=(",", ":"), default=str)
# ---------------------------------------------------------------------------
# Per-process secret. Generated once per daemon process; regenerated on restart.
# Injectable for tests so build + verify share a secret. Never serialized.
# ---------------------------------------------------------------------------
_PROCESS_SECRET = os.urandom(32)
def process_secret() -> bytes:
"""Return the per-process proof-signing secret (never serialized)."""
return _PROCESS_SECRET
def _resolve_secret(secret: bytes | None) -> bytes:
return secret if secret is not None else _PROCESS_SECRET
@dataclass(frozen=True)
class DrainCheck:
"""One mandatory drain checklist result."""
name: str
passed: bool
detail: str
def as_dict(self) -> dict[str, Any]:
return {"name": self.name, "passed": self.passed, "detail": self.detail}
@dataclass(frozen=True)
class DrainProof:
"""Machine-verifiable proof that a restart's blast radius has been drained.
The ``proof_id`` is a keyed hash over the canonical proof body; it is the
tamper-evident signature verified at the gate. ``clean`` is True only when
every required check passed. The proof is honoured only until ``expires_at``.
"""
version: str
proof_id: str
clean: bool
issued_at: str
expires_at: str
requesting_session_id: str | None
impact_fingerprint: str
checks: list[DrainCheck]
failed_checks: list[str]
def as_dict(self) -> dict[str, Any]:
return {
"version": self.version,
"proof_id": self.proof_id,
"clean": self.clean,
"issued_at": self.issued_at,
"expires_at": self.expires_at,
"requesting_session_id": self.requesting_session_id,
"impact_fingerprint": self.impact_fingerprint,
"checks": [c.as_dict() for c in self.checks],
"failed_checks": list(self.failed_checks),
}
@dataclass(frozen=True)
class VerifyResult:
"""Outcome of :func:`verify_drain_proof` (fail closed)."""
valid: bool
reasons: list[str]
proof_id: str | None
clean: bool
expired: bool
tampered: bool
def as_dict(self) -> dict[str, Any]:
return {
"valid": self.valid,
"reasons": list(self.reasons),
"proof_id": self.proof_id,
"clean": self.clean,
"expired": self.expired,
"tampered": self.tampered,
}
@dataclass(frozen=True)
class GateDecision:
"""Outcome of :func:`gate_apply_restart`."""
allow: bool
verdict: str
reasons: list[str]
proof_id: str | None
break_glass: bool
incident: dict[str, Any] | None
audit_record: dict[str, Any]
def as_dict(self) -> dict[str, Any]:
return {
"allow": self.allow,
"verdict": self.verdict,
"reasons": list(self.reasons),
"proof_id": self.proof_id,
"break_glass": self.break_glass,
"incident": self.incident,
"audit_record": dict(self.audit_record),
}
def impact_fingerprint(impact_report: Mapping[str, Any] | None) -> str:
"""Stable fingerprint of the blast-radius state a proof was minted against.
Binds a proof to the specific impact evaluation. If the live state changes
(a new mutation appears) between minting and gate, the caller can pass the
fresh fingerprint and the proof will be rejected as stale.
"""
report = impact_report or {}
counts = report.get("counts") or {}
material = {
"inventory_complete": bool(report.get("inventory_complete", False)),
"verdict": report.get("verdict"),
"affected_issues": sorted(report.get("affected_issues") or []),
"affected_prs": sorted(report.get("affected_prs") or []),
"mutations": sorted(
str(m.get("lease_id"))
for m in (report.get("mutations") or [])
if isinstance(m, Mapping)
),
"critical_sections": sorted(
str(c.get("lease_id"))
for c in (report.get("critical_sections") or [])
if isinstance(c, Mapping)
),
"terminal_lock": bool(report.get("terminal_lock")),
"counts": {
k: counts.get(k)
for k in (
"sessions_live_other",
"leases_disruptive",
"critical_sections",
"mutations",
)
},
}
return hashlib.sha256(_canonical(material).encode("utf-8")).hexdigest()
def _sign(body: Mapping[str, Any], secret: bytes) -> str:
"""Keyed (HMAC-SHA256) signature over the canonical proof body."""
return hmac.new(
secret, _canonical(body).encode("utf-8"), hashlib.sha256
).hexdigest()
def _proof_body(
*,
clean: bool,
issued_at: str,
expires_at: str,
requesting_session_id: str | None,
fingerprint: str,
checks: Sequence[DrainCheck],
) -> dict[str, Any]:
"""The exact fields covered by the signature. Order-independent (canonical)."""
return {
"version": DRAIN_PROOF_VERSION,
"clean": clean,
"issued_at": issued_at,
"expires_at": expires_at,
"requesting_session_id": requesting_session_id,
"impact_fingerprint": fingerprint,
"checks": [c.as_dict() for c in checks],
}
def _bool_input(value: Any) -> bool:
"""Strictly interpret a drain-state flag; anything not explicitly True fails."""
return value is True
def _evaluate_checks(
impact_report: Mapping[str, Any],
drain_state: Mapping[str, Any],
) -> list[DrainCheck]:
"""Compute the mandatory checklist from the impact report + drain outcomes.
The report answers "is unsafe work still in flight?"; ``drain_state`` reports
the drain-mode actions the coordinator/#659 performed. Every check fails
closed when its evidence is missing or ambiguous.
"""
checks: list[DrainCheck] = []
inventory_complete = bool(impact_report.get("inventory_complete", False))
mutations = list(impact_report.get("mutations") or [])
critical = list(impact_report.get("critical_sections") or [])
disruptive = [
l
for l in (impact_report.get("affected_leases") or [])
if isinstance(l, Mapping) and l.get("disruptive")
]
# 1. No in-flight mutations. Derived from the authoritative impact report,
# not self-reported: a proof cannot claim "no mutations" while the report
# still shows mutations or unsevered critical sections.
if not inventory_complete:
checks.append(
DrainCheck(
CHECK_NO_INFLIGHT_MUTATIONS,
False,
"impact report inventory incomplete; cannot confirm mutations "
"cleared (fail closed)",
)
)
elif mutations or critical:
checks.append(
DrainCheck(
CHECK_NO_INFLIGHT_MUTATIONS,
False,
f"{len(mutations)} mutation(s) and {len(critical)} critical "
"section(s) still in flight",
)
)
else:
checks.append(
DrainCheck(
CHECK_NO_INFLIGHT_MUTATIONS,
True,
"no in-flight mutations or critical sections in impact report",
)
)
# 2. New assignment halted (maintenance-drain entered).
checks.append(
DrainCheck(
CHECK_ASSIGNMENTS_STOPPED,
_bool_input(drain_state.get("assignments_stopped")),
"new work assignment halted"
if _bool_input(drain_state.get("assignments_stopped"))
else "assignments not confirmed stopped (fail closed)",
)
)
# 3. Durable checkpoints complete (#660). Unknown => not complete.
cp = drain_state.get("checkpoints_complete")
checks.append(
DrainCheck(
CHECK_CHECKPOINTS_COMPLETE,
_bool_input(cp),
"all live sessions checkpointed"
if _bool_input(cp)
else "checkpoint completeness unconfirmed (fail closed)",
)
)
# 4. Handoffs verified.
checks.append(
DrainCheck(
CHECK_HANDOFFS_OK,
_bool_input(drain_state.get("handoffs_verified")),
"pending handoffs verified"
if _bool_input(drain_state.get("handoffs_verified"))
else "handoffs not verified (fail closed)",
)
)
# 5. Leases resolved/transferred/preserved AND none left disruptive. Requires
# both the drain-mode assertion and the report showing no disruptive lease.
leases_asserted = _bool_input(drain_state.get("leases_handled"))
if not leases_asserted:
checks.append(
DrainCheck(
CHECK_LEASES_HANDLED,
False,
"lease disposition not asserted by drain (fail closed)",
)
)
elif disruptive:
checks.append(
DrainCheck(
CHECK_LEASES_HANDLED,
False,
f"{len(disruptive)} disruptive lease(s) still active in report",
)
)
else:
checks.append(
DrainCheck(
CHECK_LEASES_HANDLED,
True,
"leases resolved/transferred/preserved; none left disruptive",
)
)
# 6. Acknowledgements received, or an explicit timeout policy was applied.
acks = drain_state.get("acks") or {}
ack_values = list(acks.values()) if isinstance(acks, Mapping) else []
all_acked = bool(ack_values) and all(
str(v).strip().lower() in {"ack", "acked", "acknowledged"}
for v in ack_values
)
no_sessions_to_ack = isinstance(acks, Mapping) and len(ack_values) == 0
timeout_policy = _bool_input(drain_state.get("ack_timeout_policy_applied"))
acks_ok = all_acked or no_sessions_to_ack or timeout_policy
if acks_ok:
if timeout_policy and not all_acked:
detail = "explicit ack timeout policy applied"
elif no_sessions_to_ack:
detail = "no other live sessions required to acknowledge"
else:
detail = "all affected sessions acknowledged"
else:
detail = "outstanding acks with no timeout policy (fail closed)"
checks.append(DrainCheck(CHECK_ACKS_OR_TIMEOUT, acks_ok, detail))
return checks
def build_drain_proof(
*,
impact_report: Mapping[str, Any],
drain_state: Mapping[str, Any],
requesting_session_id: str | None = None,
now: datetime | None = None,
ttl_seconds: int = DEFAULT_PROOF_TTL_SECONDS,
secret: bytes | None = None,
) -> DrainProof:
"""Mint a drain proof from an impact report and the drain-mode outcomes.
A successful drain (every checklist item passes) yields a *clean* proof with
a valid signature (AC#2). An unclean drain still yields a signed proof, but
with ``clean=False`` and the failing checks named the gate will deny it
so the artifact is auditable rather than silently absent.
"""
moment = now or _utc_now()
ttl = max(1, int(ttl_seconds))
issued_at = moment.isoformat()
expires_at = (moment + timedelta(seconds=ttl)).isoformat()
fingerprint = impact_fingerprint(impact_report)
checks = _evaluate_checks(impact_report, drain_state)
failed = [c.name for c in checks if not c.passed]
clean = not failed
body = _proof_body(
clean=clean,
issued_at=issued_at,
expires_at=expires_at,
requesting_session_id=requesting_session_id,
fingerprint=fingerprint,
checks=checks,
)
proof_id = _sign(body, _resolve_secret(secret))
return DrainProof(
version=DRAIN_PROOF_VERSION,
proof_id=proof_id,
clean=clean,
issued_at=issued_at,
expires_at=expires_at,
requesting_session_id=requesting_session_id,
impact_fingerprint=fingerprint,
checks=checks,
failed_checks=failed,
)
def verify_drain_proof(
proof: Mapping[str, Any] | None,
*,
now: datetime | None = None,
secret: bytes | None = None,
expected_impact_fingerprint: str | None = None,
) -> VerifyResult:
"""Verify a drain proof, failing closed on any doubt.
A proof is valid only when: it is present and well-formed; its signature
recomputes with the per-process secret (not forged/tampered, not minted in a
prior process); it has not expired; every required check is present and
passed; and when ``expected_impact_fingerprint`` is supplied it was
minted against the current blast-radius state.
"""
moment = now or _utc_now()
reasons: list[str] = []
if not isinstance(proof, Mapping):
return VerifyResult(
valid=False,
reasons=["drain proof missing or not an object (fail closed)"],
proof_id=None,
clean=False,
expired=False,
tampered=False,
)
proof_id = proof.get("proof_id")
presented_clean = bool(proof.get("clean", False))
# Rebuild the signed body from the presented fields and re-sign. Any mutation
# of a covered field (including flipping ``clean`` to True) breaks the match.
raw_checks = proof.get("checks")
checks: list[DrainCheck] = []
checks_wellformed = isinstance(raw_checks, Sequence) and not isinstance(
raw_checks, (str, bytes)
)
if checks_wellformed:
for c in raw_checks:
if not isinstance(c, Mapping) or "name" not in c or "passed" not in c:
checks_wellformed = False
break
checks.append(
DrainCheck(
name=str(c.get("name")),
passed=bool(c.get("passed")),
detail=str(c.get("detail") or ""),
)
)
tampered = False
if not checks_wellformed:
reasons.append("drain proof checks malformed (fail closed)")
tampered = True
else:
body = _proof_body(
clean=presented_clean,
issued_at=str(proof.get("issued_at") or ""),
expires_at=str(proof.get("expires_at") or ""),
requesting_session_id=proof.get("requesting_session_id"),
fingerprint=str(proof.get("impact_fingerprint") or ""),
checks=checks,
)
expected_sig = _sign(body, _resolve_secret(secret))
if not (
isinstance(proof_id, str)
and hmac.compare_digest(expected_sig, proof_id)
):
tampered = True
reasons.append(
"drain proof signature mismatch: forged, tampered, or minted "
"in a prior process (fail closed)"
)
expires = _parse_ts(proof.get("expires_at"))
expired = expires is None or moment >= expires
if expires is None:
reasons.append("drain proof has no valid expiry (fail closed)")
elif expired:
reasons.append(f"drain proof expired at {proof.get('expires_at')}")
# Recompute cleanliness from the checks themselves — never trust the flag.
recomputed_failed = [c.name for c in checks if not c.passed]
present_names = {c.name for c in checks}
missing = [name for name in REQUIRED_CHECKS if name not in present_names]
recomputed_clean = checks_wellformed and not recomputed_failed and not missing
if missing:
reasons.append(f"drain proof missing required checks: {', '.join(missing)}")
if checks_wellformed and recomputed_failed:
reasons.append(
f"drain checks failed: {', '.join(sorted(set(recomputed_failed)))}"
)
if presented_clean and not recomputed_clean:
tampered = True
reasons.append("proof claims clean but its checks do not support it")
if expected_impact_fingerprint is not None:
if str(proof.get("impact_fingerprint") or "") != str(
expected_impact_fingerprint
):
reasons.append(
"drain proof was minted against a different blast-radius state "
"(stale; fail closed)"
)
valid = (not tampered) and (not expired) and recomputed_clean and not reasons
return VerifyResult(
valid=valid,
reasons=reasons,
proof_id=proof_id if isinstance(proof_id, str) else None,
clean=recomputed_clean,
expired=expired,
tampered=tampered,
)
def _incident_descriptor(
*,
reasons: Sequence[str],
requesting_session_id: str | None,
proof_id: str | None,
at: str,
) -> dict[str, Any]:
"""Durable incident descriptor for a denied restart (caller creates issue).
Kept as data (not a live Gitea call) so this module stays pure and testable;
the MCP tool layer turns it into a durable issue via the incident bridge.
"""
return {
"kind": "restart_drain_gate_denied",
"title": "Restart denied: drain proof failed the hard gate",
"labels": ["mcp-health", "safety", "stale-runtime", "workflow-hardening"],
"reasons": list(reasons),
"requesting_session_id": requesting_session_id,
"proof_id": proof_id,
"at": at,
"drain_proof_version": DRAIN_PROOF_VERSION,
}
def gate_apply_restart(
*,
proof: Mapping[str, Any] | None,
now: datetime | None = None,
secret: bytes | None = None,
break_glass: bool = False,
expected_impact_fingerprint: str | None = None,
requesting_session_id: str | None = None,
) -> GateDecision:
"""Hard gate for a sanctioned restart apply (#661 AC#1/#3).
Restart apply is authorized only with a valid, unexpired, clean drain proof.
A missing, expired, tampered, or unclean proof denies the restart and emits
a durable incident descriptor. ``break_glass`` is the *only* sanctioned
bypass authorization for it is the caller's responsibility (the emergency
workflow is a separate child); when set, the gate allows without a proof but
records the bypass so it is never silent.
"""
moment = now or _utc_now()
at = moment.isoformat()
if break_glass:
reasons = ["break-glass restart authorized; drain proof gate bypassed"]
audit = {
"event": "restart_gate_evaluated",
"drain_proof_version": DRAIN_PROOF_VERSION,
"evaluated_at": at,
"verdict": GATE_BREAK_GLASS,
"allow": True,
"break_glass": True,
"requesting_session_id": requesting_session_id,
"proof_id": None,
}
return GateDecision(
allow=True,
verdict=GATE_BREAK_GLASS,
reasons=reasons,
proof_id=None,
break_glass=True,
incident=None,
audit_record=audit,
)
result = verify_drain_proof(
proof,
now=moment,
secret=secret,
expected_impact_fingerprint=expected_impact_fingerprint,
)
if result.valid:
reasons = ["valid unexpired clean drain proof present; restart authorized"]
audit = {
"event": "restart_gate_evaluated",
"drain_proof_version": DRAIN_PROOF_VERSION,
"evaluated_at": at,
"verdict": GATE_ALLOW,
"allow": True,
"break_glass": False,
"requesting_session_id": requesting_session_id,
"proof_id": result.proof_id,
}
return GateDecision(
allow=True,
verdict=GATE_ALLOW,
reasons=reasons,
proof_id=result.proof_id,
break_glass=False,
incident=None,
audit_record=audit,
)
deny_reasons = ["restart denied: drain proof invalid (fail closed)"] + list(
result.reasons
)
incident = _incident_descriptor(
reasons=deny_reasons,
requesting_session_id=requesting_session_id,
proof_id=result.proof_id,
at=at,
)
audit = {
"event": "restart_gate_evaluated",
"drain_proof_version": DRAIN_PROOF_VERSION,
"evaluated_at": at,
"verdict": GATE_DENY,
"allow": False,
"break_glass": False,
"requesting_session_id": requesting_session_id,
"proof_id": result.proof_id,
"incident_kind": incident["kind"],
}
return GateDecision(
allow=False,
verdict=GATE_DENY,
reasons=deny_reasons,
proof_id=result.proof_id,
break_glass=False,
incident=incident,
audit_record=audit,
)
+1764 -42
View File
File diff suppressed because it is too large Load Diff
+7
View File
@@ -16,11 +16,18 @@ ISSUE_LOCK_FILE = os.environ.get("GITEA_ISSUE_LOCK_FILE", "/tmp/gitea_issue_lock
SOURCE_LOCK_ISSUE = "gitea_lock_issue"
SOURCE_LOCK_ADOPTION = "gitea_lock_issue_adoption"
SOURCE_OPERATOR_OVERRIDE = "operator_override"
SOURCE_RECOVER_DIRTY_ORPHANED = "gitea_recover_dirty_orphaned_issue_worktree"
# #864: dirty-preserving same-claimant author-session rebind (dead owner PID).
SOURCE_DIRTY_SAME_CLAIMANT_REBIND = (
"gitea_rebind_dirty_same_claimant_author_session"
)
SANCTIONED_LOCK_SOURCES = frozenset({
SOURCE_LOCK_ISSUE,
SOURCE_LOCK_ADOPTION,
SOURCE_OPERATOR_OVERRIDE,
SOURCE_RECOVER_DIRTY_ORPHANED,
SOURCE_DIRTY_SAME_CLAIMANT_REBIND,
})
_OPERATOR_OVERRIDE_ENV = "GITEA_ISSUE_LOCK_OPERATOR_OVERRIDE"
+130 -8
View File
@@ -85,6 +85,12 @@ HEAD_RELATION_STRICT_DESCENDANT = "strict_descendant"
# #772: an unpublished claim has no recorded head to compare against at all, so
# its head is measured against the base the branch was cut from instead.
HEAD_RELATION_DESCENDS_FROM_BASE = "descends_from_recorded_base"
# #871: the remote/PR head advanced *past* the recorded head via a sanctioned
# merge-based branch synchronization (``gitea_update_pr_branch_by_merge``) while
# the local worktree stayed at the recorded head. This is the inverse of the
# #768 descendant relation — here the *remote* strictly descends the local head,
# and only because a base was merged into the branch, proven server-side.
HEAD_RELATION_REMOTE_MERGE_SYNCED = "remote_merge_synced"
# Which body of evidence a recovery was decided on (#772 AC10). These are not
# interchangeable: a published claim proves ownership against a remote/PR head,
@@ -266,6 +272,70 @@ def _assess_base_descendancy(
]
def _assess_remote_merge_synced(
sync_provenance: Mapping[str, Any] | None,
*,
recorded_head: str,
remote_head: str,
) -> tuple[bool, list[str]]:
"""Did ``remote_head`` advance past ``recorded_head`` via a sanctioned
merge-based branch sync (#871)?
``sync_provenance`` is the server-side git observation from
``issue_lock_worktree.read_merge_sync_provenance``. Its own
``prior_head_sha`` / ``synced_head_sha`` are re-checked against the heads
this assessment is actually reasoning about, so an observation taken for some
other pair of commits stale, mismatched, or hand-built can never
authorize recovery. This is the inverse of ``_assess_strict_descendant``: the
recorded head is the ancestor and the *remote* head is the descendant, and it
is accepted only because the remote head is a base-into-branch merge that
preserved the branch mainline back to the recorded head.
Returns ``(proven, notes)``. Notes name the exact missing element so a
refused caller sees why, never a bare "unproven".
"""
if not isinstance(sync_provenance, Mapping):
return False, [
"no server-derived merge-sync provenance observation was available; a "
"remote head ahead of the recorded head cannot be accepted"
]
probe_prior = _text(sync_provenance.get("prior_head_sha"))
probe_synced = _text(sync_provenance.get("synced_head_sha"))
if probe_prior != recorded_head or probe_synced != remote_head:
return False, [
f"merge-sync observation covers {probe_prior or 'unknown'} -> "
f"{probe_synced or 'unknown'}, not the heads under assessment "
f"({recorded_head} -> {remote_head})"
]
if not sync_provenance.get("probe_ok"):
return False, (
list(sync_provenance.get("reasons") or [])
or ["merge-sync provenance probe did not complete; provenance unproven"]
)
if not sync_provenance.get("prior_is_ancestor"):
return False, [
f"recorded head {recorded_head} is not an ancestor of remote head "
f"{remote_head}; a rewritten or force-moved head cannot be recovered"
]
if not sync_provenance.get("is_merge_sync"):
return False, (
list(sync_provenance.get("reasons") or [])
or [
f"remote head {remote_head} is not a sanctioned merge-based sync "
f"of the base into the branch above {recorded_head}"
]
)
proof = _text(sync_provenance.get("proof")) or (
f"{remote_head} merged the base into the branch above {recorded_head}"
)
return True, [
f"remote head {remote_head} advanced past recorded head {recorded_head} "
f"via a sanctioned merge-based branch sync ({proof})"
]
def assess_dead_session_lock_recovery(
existing_lock: Mapping[str, Any] | None,
*,
@@ -290,6 +360,7 @@ def assess_dead_session_lock_recovery(
remote_branch_exists: bool | None = None,
recorded_base_sha: str | None = None,
base_ancestry: Mapping[str, Any] | None = None,
sync_provenance: Mapping[str, Any] | None = None,
) -> dict[str, Any]:
"""Decide whether a dead-session author lock may be natively recovered.
@@ -469,19 +540,39 @@ def assess_dead_session_lock_recovery(
head_relation = HEAD_RELATION_STRICT_DESCENDANT
ancestry_proof = notes[0] if notes else None
else:
reasons.append(
f"local head {local_head} does not match remote branch head "
f"{remote_head}"
# #871: the reverse relation — the remote head advanced past
# the recorded/local head via a sanctioned merge-based branch
# sync while the local worktree stayed put. Accepted only on
# server-proven merge-sync provenance, never a caller claim.
synced, sync_notes = _assess_remote_merge_synced(
sync_provenance,
recorded_head=local_head,
remote_head=remote_head,
)
reasons.extend(notes)
if synced:
head_relation = HEAD_RELATION_REMOTE_MERGE_SYNCED
ancestry_proof = sync_notes[0] if sync_notes else None
else:
reasons.append(
f"local head {local_head} does not match remote branch "
f"head {remote_head}"
)
reasons.extend(notes)
reasons.extend(sync_notes)
evidence["recorded_base"] = recorded_base or None
evidence["local_head"] = local_head or None
evidence["remote_head"] = remote_head or None
# ``recorded_head`` is the head recovery is being measured against;
# ``accepted_head`` is the head this recovery actually adopts. They differ
# only in the descendant case, and downstream gates need both (#768 AC2/AC7).
# #871: in the merge-sync case the branch/PR already carries the synced
# remote head, so that is the head recovery adopts; the local worktree stays
# at the ancestor recorded head.
evidence["recorded_head"] = remote_head or None
evidence["accepted_head"] = local_head or None
if head_relation == HEAD_RELATION_REMOTE_MERGE_SYNCED:
evidence["accepted_head"] = remote_head or None
else:
evidence["accepted_head"] = local_head or None
evidence["head_relation"] = head_relation
evidence["ancestry_proof"] = ancestry_proof
@@ -493,12 +584,17 @@ def assess_dead_session_lock_recovery(
# contradictory; re-stating it as a head mismatch would only obscure why.
if not unpublished and local_head and pr_head != local_head:
# A descendant recovery has not been published yet, so the open PR
# legitimately still points at the recorded head. Any other
# disagreement is a real mismatch.
# legitimately still points at the recorded head. A merge-sync
# recovery's PR legitimately sits at the advanced remote head. Any
# other disagreement is a real mismatch.
if not (
head_relation == HEAD_RELATION_STRICT_DESCENDANT
and remote_head
and pr_head == remote_head
) and not (
head_relation == HEAD_RELATION_REMOTE_MERGE_SYNCED
and remote_head
and pr_head == remote_head
):
reasons.append(
f"open PR #{pr_number} head {pr_head} does not match local head "
@@ -619,7 +715,11 @@ def assess_dead_session_lock_recovery(
)
if (
head_relation
in (HEAD_RELATION_STRICT_DESCENDANT, HEAD_RELATION_DESCENDS_FROM_BASE)
in (
HEAD_RELATION_STRICT_DESCENDANT,
HEAD_RELATION_DESCENDS_FROM_BASE,
HEAD_RELATION_REMOTE_MERGE_SYNCED,
)
and ancestry_proof
):
proof.append(ancestry_proof)
@@ -697,6 +797,16 @@ def owning_pr_recovery_evidence(
return None
if accepted_head != local_head:
return None
elif relation == HEAD_RELATION_REMOTE_MERGE_SYNCED:
# #871: the PR already sits at the advanced remote head; the local
# worktree is the ancestor the merge preserved. The head the open PR
# shows and the head recovery adopts are both the synced remote head.
if not remote_head or pr_head != remote_head:
return None
if accepted_head != remote_head:
return None
if not local_head or local_head == remote_head:
return None
else:
return None
try:
@@ -765,6 +875,18 @@ def recovered_owning_pr_from_lock(
return None
if not accepted_head or accepted_head == recorded_head:
return None
elif relation == HEAD_RELATION_REMOTE_MERGE_SYNCED:
# #871: PR sits at the advanced remote head, which is both the recorded
# measured-against head and the adopted head; the local worktree is the
# ancestor the merge preserved.
remote_head = _text(record.get("remote_head"))
local_head = _text(record.get("local_head"))
if not remote_head or pr_head != remote_head:
return None
if accepted_head and accepted_head != remote_head:
return None
if not local_head or local_head == remote_head:
return None
else:
return None
try:
+944 -36
View File
File diff suppressed because it is too large Load Diff
+178
View File
@@ -145,6 +145,184 @@ def read_head_ancestry(
return result
def read_merge_sync_provenance(
worktree_path: str,
*,
prior_head_sha: str | None,
synced_head_sha: str | None,
remote: str | None = None,
) -> dict:
"""Observe whether ``synced_head_sha`` is a sanctioned merge-sync of a base
into the branch above ``prior_head_sha`` (#871/#872).
Reports server-derived git facts only; the recovery disposition lives in
``issue_lock_recovery``. All comparisons are executed locally in the
declared worktree -- nothing is taken from caller parameters.
Provenance is proven only when ALL hold:
* both commits are present (a rewritten/force-moved prior head leaves the
object graph and fails closed);
* ``prior_head_sha`` is a strict ancestor of ``synced_head_sha`` (the branch
history is preserved, never replaced);
* ``synced_head_sha`` is a merge commit (two or more parents), i.e. a base
merged in a plain fast-forward of new direct commits is not a sync;
* ``prior_head_sha`` is an ancestor of the merge's **first** parent, so the
branch mainline (first-parent lineage) still reaches the prior head a
rebase/force-push that re-authored the branch side fails this.
"""
path = (worktree_path or "").strip()
prior = (prior_head_sha or "").strip()
synced = (synced_head_sha or "").strip()
target_remote = (remote or "").strip() or None
result: dict = {
"prior_head_sha": prior or None,
"synced_head_sha": synced or None,
"probe_ok": False,
"prior_present": False,
"synced_present": False,
"prior_is_ancestor": False,
"synced_is_merge": False,
"first_parent_reaches_prior": False,
"is_merge_sync": False,
"first_parent_sha": None,
"parent_count": None,
"proof": None,
"reasons": [],
}
if not path or not prior or not synced:
result["reasons"].append(
"merge-sync provenance probe requires a worktree path and both "
"commit SHAs"
)
return result
if prior == synced:
result["reasons"].append(
"prior and synced heads are identical; no branch sync occurred"
)
return result
def _present(sha: str) -> bool:
res = subprocess.run(
["git", "-C", path, "rev-parse", "--verify", "--quiet", f"{sha}^{{commit}}"],
capture_output=True,
text=True,
check=False,
)
return res.returncode == 0
def _is_ancestor(ancestor: str, descendant: str) -> bool | None:
res = subprocess.run(
["git", "-C", path, "merge-base", "--is-ancestor", ancestor, descendant],
capture_output=True,
text=True,
check=False,
)
if res.returncode == 0:
return True
if res.returncode == 1:
return False
return None # failed probe — never a silent "no"
try:
result["prior_present"] = _present(prior)
result["synced_present"] = _present(synced)
if not result["synced_present"] and path and os.path.isdir(path):
if target_remote:
subprocess.run(
["git", "-C", path, "fetch", target_remote, "--quiet"],
capture_output=True,
text=True,
check=False,
)
result["synced_present"] = _present(synced)
if not result["synced_present"]:
subprocess.run(
["git", "-C", path, "fetch", "--quiet"],
capture_output=True,
text=True,
check=False,
)
result["synced_present"] = _present(synced)
except OSError as exc: # git unavailable — fail closed, never assume
result["reasons"].append(f"merge-sync provenance probe could not run: {exc}")
return result
if not result["prior_present"]:
result["reasons"].append(
f"prior head {prior} is not reachable in '{path}'; history may have "
"been rewritten or force-moved"
)
if not result["synced_present"]:
result["reasons"].append(
f"synced head {synced} is not reachable in '{path}'"
)
if not (result["prior_present"] and result["synced_present"]):
return result
ancestor = _is_ancestor(prior, synced)
if ancestor is None:
result["reasons"].append(
"ancestry probe failed; merge-sync provenance unproven"
)
return result
result["prior_is_ancestor"] = bool(ancestor)
if not ancestor:
result["reasons"].append(
f"prior head {prior} is not an ancestor of synced head {synced}; "
"the branch history was not preserved (not a merge-based sync)"
)
return result
parents_res = subprocess.run(
["git", "-C", path, "rev-list", "--parents", "-n", "1", synced],
capture_output=True,
text=True,
check=False,
)
if parents_res.returncode != 0:
result["reasons"].append(
f"could not read parents of {synced}; merge-sync provenance unproven"
)
return result
tokens = (parents_res.stdout or "").split()
# tokens[0] is the commit itself; the rest are its parents.
parents = tokens[1:]
result["parent_count"] = len(parents)
result["synced_is_merge"] = len(parents) >= 2
if not result["synced_is_merge"]:
result["probe_ok"] = True
result["reasons"].append(
f"synced head {synced} has {len(parents)} parent(s); a merge-based "
"branch sync produces a merge commit (two or more parents)"
)
return result
first_parent = parents[0]
result["first_parent_sha"] = first_parent
fp_reaches = _is_ancestor(prior, first_parent) if prior != first_parent else True
if fp_reaches is None:
result["reasons"].append(
"first-parent ancestry probe failed; merge-sync provenance unproven"
)
return result
result["first_parent_reaches_prior"] = bool(fp_reaches)
result["probe_ok"] = True
if not fp_reaches:
result["reasons"].append(
f"merge first parent {first_parent} does not reach prior head "
f"{prior}; the branch mainline was re-authored (not a sanctioned sync)"
)
return result
result["is_merge_sync"] = True
result["proof"] = (
f"synced head {synced} is a merge commit (parents={len(parents)}) whose "
f"first-parent lineage reaches prior head {prior}; base merged into branch"
)
return result
def read_recorded_base(
worktree_path: str,
*,
+209 -13
View File
@@ -39,6 +39,7 @@ SAFE_RELEASE_OWNED = "release_owned"
SAFE_STALE_PROMPT = "stale_prompt_lease"
SAFE_UNKNOWN = "inspect_only"
SAFE_NO_AUTHORITY = "file_or_comment_not_authoritative"
SAFE_CONSUME_CROSS_ROLE = "consume_cross_role_handoff"
LEASE_STATUS_ACTIVE = "active"
LEASE_STATUS_RELEASED = "released"
@@ -250,6 +251,23 @@ def decide_safe_next_action(
"same_owner": True,
"also_allowed": [SAFE_ABANDON_ALLOWED, SAFE_RELEASE_OWNED],
}
handoff = is_pending_cross_role_handoff({"lease": lease})
if handoff:
return {
"safe_next_action": SAFE_CONSUME_CROSS_ROLE,
"reasons": [
f"controller allocation pending handoff (freshness={status}); "
"required-role worker may consume without abandon/reassign; "
f"required_role={handoff['required_role']}"
],
"block": False,
"same_owner": False,
"owner_session_id": owner,
"required_role": handoff["required_role"],
"cross_role_handoff": True,
"handoff_status": "pending",
"also_allowed": [SAFE_ABANDON_ALLOWED],
}
return {
"safe_next_action": SAFE_ABANDON_ALLOWED,
"reasons": [
@@ -272,6 +290,24 @@ def decide_safe_next_action(
}
if not same_owner and status == "active":
# #843: pending cross-role handoff is consumable by required role
handoff = is_pending_cross_role_handoff({"lease": lease})
if handoff:
return {
"safe_next_action": SAFE_CONSUME_CROSS_ROLE,
"reasons": [
"controller cross-role allocation pending handoff; "
f"required_role={handoff['required_role']}; "
"consume via gitea_adopt_workflow_lease without "
"abandonment or sharing the controller session"
],
"block": False,
"same_owner": False,
"owner_session_id": owner,
"required_role": handoff["required_role"],
"cross_role_handoff": True,
"handoff_status": "pending",
}
return {
"safe_next_action": SAFE_WAIT_FOREIGN,
"reasons": [
@@ -440,6 +476,84 @@ def list_active_leases(
}
def parse_lease_provenance(lease_or_state: Mapping[str, Any] | None) -> dict[str, Any]:
"""Return durable lease provenance dict (empty when absent/unparseable)."""
if not lease_or_state:
return {}
if "provenance" in lease_or_state and isinstance(lease_or_state.get("provenance"), dict):
return dict(lease_or_state["provenance"])
raw = None
if "provenance_json" in lease_or_state:
raw = lease_or_state.get("provenance_json")
elif "lease" in lease_or_state and isinstance(lease_or_state.get("lease"), Mapping):
raw = lease_or_state["lease"].get("provenance_json")
if not raw:
return {}
if isinstance(raw, dict):
return dict(raw)
try:
loaded = json.loads(raw)
except (TypeError, json.JSONDecodeError):
return {}
return dict(loaded) if isinstance(loaded, dict) else {}
def is_pending_cross_role_handoff(
state: Mapping[str, Any] | None,
) -> dict[str, Any] | None:
"""Return handoff evidence when a controller allocation awaits consume (#843).
A pending handoff is identified by durable provenance written at
cross-role apply time not by title heuristics or session-id guessing.
"""
if not state:
return None
lease = state.get("lease") if isinstance(state.get("lease"), Mapping) else state
if not isinstance(lease, Mapping):
return None
status = str(lease.get("status") or "").strip().lower()
if status in (LEASE_STATUS_ABANDONED, LEASE_STATUS_RELEASED, LEASE_STATUS_EXPIRED):
return None
prov = parse_lease_provenance(state)
if not prov and isinstance(lease, Mapping):
prov = parse_lease_provenance(lease)
if not prov.get("cross_role_handoff"):
return None
handoff_status = str(prov.get("handoff_status") or "pending").strip().lower()
if handoff_status != "pending":
return None
adopted_by = (
lease.get("adopted_by_session_id")
or prov.get("adopted_by_session_id")
or ""
)
if str(adopted_by).strip():
return None
required_role = str(
prov.get("required_role") or lease.get("role") or ""
).strip().lower()
if not required_role:
return None
return {
"cross_role_handoff": True,
"handoff_status": "pending",
"required_role": required_role,
"allocating_session_id": str(
prov.get("allocating_session_id") or lease.get("session_id") or ""
),
"allocating_role": str(prov.get("allocating_role") or "controller"),
"lease_id": str(lease.get("lease_id") or ""),
"assignment_id": (
str(state["assignment"]["assignment_id"])
if isinstance(state.get("assignment"), Mapping)
and state["assignment"].get("assignment_id")
else None
),
"provenance": prov,
}
def adopt_lease(
db: cpd.ControlPlaneDB,
*,
@@ -450,8 +564,17 @@ def adopt_lease(
expected_head_sha: str | None = None,
owner_pid: int | None = None,
operator_authorized: bool = False,
adopter_profile_name: str | None = None,
adopter_namespace: str | None = None,
) -> dict[str, Any]:
"""Sanctioned adopt path with provenance; never silent foreign steal."""
"""Sanctioned adopt path with provenance; never silent foreign steal.
#843 F1: for a pending cross-role handoff, ``role`` must be the
authoritative profile-derived role supplied by the MCP boundary never
caller-asserted authority. When the handoff provenance declares
``required_profile`` / ``required_namespace`` and the caller context is
provided, both are validated exactly; a mismatch fails closed.
"""
state = db.get_lease_workflow_state(lease_id)
if not state:
raise LeaseLifecycleError(
@@ -463,11 +586,8 @@ def adopt_lease(
owner = str(lease.get("session_id") or "")
same_owner = owner == str(adopter_session_id)
if freshness["freshness"] == "active" and not same_owner:
raise LeaseLifecycleError(
f"refusing to steal active foreign lease {lease_id} owned by "
f"{owner} (fail closed)"
)
handoff = is_pending_cross_role_handoff(state)
adopter_role = (role or "").strip().lower()
if freshness["freshness"] in ("abandoned", "released"):
raise LeaseLifecycleError(
@@ -475,13 +595,64 @@ def adopt_lease(
"(fail closed)"
)
# Expired or stale: require abandon-style safety before ownership transfer
# when not same owner; same owner may reclaim.
if not same_owner and freshness["freshness"] in (
if handoff and not same_owner:
# Terminal statuses already rejected above. Freshness may be
# active OR stale_dead_process (controller exited) — both are
# consumable without abandonment when handoff is still pending.
if freshness["freshness"] not in (
"active",
"stale_dead_process",
"stale_missing_worktree",
):
raise LeaseLifecycleError(
f"lease {lease_id} freshness={freshness['freshness']}; "
"terminal or non-active allocation cannot be handoff-consumed "
"(fail closed)"
)
required = handoff["required_role"]
if adopter_role != required:
raise LeaseLifecycleError(
f"wrong role for cross-role handoff consume of {lease_id}: "
f"required={required} adopter={adopter_role or 'none'} "
"(fail closed)"
)
# #843 F1: provenance profile/namespace restrictions are validated
# against the authoritative caller context when declared. Caller
# input can never widen authority; a mismatch fails closed.
handoff_prov = handoff.get("provenance") or {}
required_profile = str(
handoff_prov.get("required_profile") or ""
).strip()
if required_profile and adopter_profile_name is not None:
if str(adopter_profile_name).strip() != required_profile:
raise LeaseLifecycleError(
f"wrong profile for cross-role handoff consume of "
f"{lease_id}: required_profile={required_profile} "
f"adopter_profile={adopter_profile_name} (fail closed)"
)
required_namespace = str(
handoff_prov.get("required_namespace") or ""
).strip()
if required_namespace and adopter_namespace is not None:
if str(adopter_namespace).strip() != required_namespace:
raise LeaseLifecycleError(
f"wrong namespace for cross-role handoff consume of "
f"{lease_id}: required_namespace={required_namespace} "
f"adopter_namespace={adopter_namespace} (fail closed)"
)
reason = "cross-role-handoff-consume"
elif freshness["freshness"] == "active" and not same_owner:
raise LeaseLifecycleError(
f"refusing to steal active foreign lease {lease_id} owned by "
f"{owner} (fail closed)"
)
elif not same_owner and freshness["freshness"] in (
"expired",
"stale_dead_process",
"stale_missing_worktree",
):
# Expired or stale (non-handoff): require abandon-style safety before
# ownership transfer when not same owner; same owner may reclaim.
if not operator_authorized and freshness["freshness"] == "expired":
# Deterministic reclaim of expired foreign lease is allowed
# without operator flag (sanctioned expire reclaim).
@@ -492,6 +663,9 @@ def adopt_lease(
f"lease {lease_id} freshness={freshness['freshness']}; "
"use abandon with proof before foreign adopt (fail closed)"
)
reason = "sanctioned-reclaim-adopt"
else:
reason = "owner-resume-adopt" if same_owner else "sanctioned-reclaim-adopt"
provenance = build_adopt_provenance(
adopted_from_session_id=owner,
@@ -504,10 +678,14 @@ def adopt_lease(
worktree_path=worktree_path,
expected_head_sha=expected_head_sha or lease.get("expected_head_sha"),
prior_lease_id=lease_id,
reason=(
"owner-resume-adopt" if same_owner else "sanctioned-reclaim-adopt"
),
reason=reason,
)
if handoff and not same_owner:
provenance["cross_role_handoff"] = True
provenance["handoff_status"] = "adopted"
provenance["required_role"] = handoff["required_role"]
provenance["allocating_session_id"] = handoff["allocating_session_id"]
provenance["allocating_role"] = handoff["allocating_role"]
result = db.adopt_lease(
lease_id=lease_id,
@@ -518,7 +696,7 @@ def adopt_lease(
owner_pid=owner_pid if owner_pid is not None else os.getpid(),
provenance=provenance,
)
return {
out = {
"success": True,
"outcome": result.get("outcome"),
"same_owner": same_owner,
@@ -531,6 +709,24 @@ def adopt_lease(
"comment_lease_only": False,
"reasons": result.get("reasons") or [],
}
if handoff and not same_owner:
out["cross_role_handoff"] = True
out["handoff_status"] = "adopted"
out["required_role"] = handoff["required_role"]
out["adopted_by_session_id"] = adopter_session_id
out["adopted_from_session_id"] = owner
lease_row = result.get("lease") or {}
if isinstance(lease_row, Mapping):
out["read_after_write"] = {
"lease_id": lease_row.get("lease_id"),
"session_id": lease_row.get("session_id"),
"role": lease_row.get("role"),
"status": lease_row.get("status"),
"adopted_by_session_id": lease_row.get("adopted_by_session_id"),
"adopted_from_session_id": lease_row.get("adopted_from_session_id"),
"phase": lease_row.get("phase"),
}
return out
def release_lease(
+212
View File
@@ -0,0 +1,212 @@
"""Central lease policy configuration (#790 Slice A, AC-N7).
The single authoritative source for every lease duration in the project. Before
this module the numbers were scattered: a four-hour author TTL was declared
twice (``issue_lock_store`` and ``gitea_mcp_server``), the reviewer/merger
sliding window lived in ``reviewer_pr_lease``, the conflict-fix window in
``pr_work_lease``, and the control-plane default in ``control_plane_db``.
Nothing tied them together, so tuning one class silently diverged from the
others and no reader could answer "how long does a lease live?" without
grepping four files.
AC-N7 requires that this configuration exist *before* the first heartbeat and
TTL behavior that reads from it, so it ships in Slice A rather than trailing the
code it governs.
Deliberate boundaries:
* **Declaration is not rewiring.** Every task class is declared here, but only
those with ``heartbeat_lifecycle_active`` were migrated onto the shared
heartbeat lifecycle in Slice A currently ``author_issue_work`` alone.
Reviewer, merger, and conflict-fix leases keep their own existing behavior
until Slice C moves them; their numbers are recorded here so the two cannot
drift apart unnoticed, and ``tests/test_issue_790_lease_policy.py`` asserts
the recorded values still equal the constants those modules use.
* **No policy decision lives here.** This module answers "how long", never "may
this session proceed". Freshness, reclaim, and renewal dispositions stay in
``issue_lock_store``.
"""
from __future__ import annotations
import os
from dataclasses import dataclass
from typing import Any
# Task classes. Only the first is migrated onto the shared lifecycle in Slice A.
TASK_CLASS_AUTHOR_ISSUE_WORK = "author_issue_work"
TASK_CLASS_REVIEWER_PR = "reviewer_pr"
TASK_CLASS_MERGER_PR = "merger_pr"
TASK_CLASS_CONFLICT_FIX = "conflict_fix"
# Durable marker for a lease minted under the shared heartbeat lifecycle.
#
# #790 AC-N8: this explicit marker — never a timestamp comparison — is what
# distinguishes a heartbeat-lifecycle lease from a legacy one. A lock written
# before this lifecycle existed carries no marker and reads as
# ``LIFECYCLE_LEGACY``.
LIFECYCLE_HEARTBEAT_V1 = "heartbeat-v1"
LIFECYCLE_LEGACY = "legacy"
_ENV_PREFIX = "GITEA_LEASE_POLICY"
@dataclass(frozen=True)
class LeasePolicy:
"""Durations governing one task class.
All intervals are minutes except ``absolute_cap_hours``. ``None`` for the
cap means the class has no maximum continuous duration.
"""
task_class: str
initial_ttl_minutes: float
heartbeat_cadence_minutes: float
stale_warning_minutes: float
missed_heartbeat_grace_minutes: float
absolute_cap_hours: float | None
recovery_grace_minutes: float
terminal_race_drain_minutes: float
terminal_retirement_eligible: bool
heartbeat_lifecycle_active: bool
# Defaults. ``author_issue_work`` adopts the reviewer window proven by #747
# rather than inventing new numbers: a lease expires 10 minutes after its last
# valid heartbeat, warns at half that, and an actively heartbeating session is
# never evicted. The prior value was a fixed four hours (240 minutes) that no
# heartbeat could shorten — the defect this issue exists to correct.
_DEFAULTS: dict[str, LeasePolicy] = {
TASK_CLASS_AUTHOR_ISSUE_WORK: LeasePolicy(
task_class=TASK_CLASS_AUTHOR_ISSUE_WORK,
initial_ttl_minutes=10.0,
heartbeat_cadence_minutes=2.0,
stale_warning_minutes=5.0,
missed_heartbeat_grace_minutes=10.0,
absolute_cap_hours=8.0,
recovery_grace_minutes=10.0,
terminal_race_drain_minutes=2.0,
terminal_retirement_eligible=True,
heartbeat_lifecycle_active=True,
),
# Declared, not rewired. These mirror reviewer_pr_lease.LEASE_TTL_MINUTES
# and STALE_WARNING_MINUTES; Slice C migrates the call sites.
TASK_CLASS_REVIEWER_PR: LeasePolicy(
task_class=TASK_CLASS_REVIEWER_PR,
initial_ttl_minutes=10.0,
heartbeat_cadence_minutes=2.0,
stale_warning_minutes=5.0,
missed_heartbeat_grace_minutes=10.0,
absolute_cap_hours=None,
recovery_grace_minutes=10.0,
terminal_race_drain_minutes=2.0,
terminal_retirement_eligible=False,
heartbeat_lifecycle_active=False,
),
TASK_CLASS_MERGER_PR: LeasePolicy(
task_class=TASK_CLASS_MERGER_PR,
initial_ttl_minutes=10.0,
heartbeat_cadence_minutes=2.0,
stale_warning_minutes=5.0,
missed_heartbeat_grace_minutes=10.0,
absolute_cap_hours=None,
recovery_grace_minutes=10.0,
terminal_race_drain_minutes=2.0,
terminal_retirement_eligible=False,
heartbeat_lifecycle_active=False,
),
# Mirrors pr_work_lease.DEFAULT_CONFLICT_FIX_TTL_MINUTES. Deliberately left
# at its current window; shortening it is Slice C's call, not this slice's.
TASK_CLASS_CONFLICT_FIX: LeasePolicy(
task_class=TASK_CLASS_CONFLICT_FIX,
initial_ttl_minutes=120.0,
heartbeat_cadence_minutes=2.0,
stale_warning_minutes=5.0,
missed_heartbeat_grace_minutes=10.0,
absolute_cap_hours=None,
recovery_grace_minutes=10.0,
terminal_race_drain_minutes=2.0,
terminal_retirement_eligible=False,
heartbeat_lifecycle_active=False,
),
}
_NUMERIC_FIELDS = (
"initial_ttl_minutes",
"heartbeat_cadence_minutes",
"stale_warning_minutes",
"missed_heartbeat_grace_minutes",
"absolute_cap_hours",
"recovery_grace_minutes",
"terminal_race_drain_minutes",
)
def env_var_name(task_class: str, field: str) -> str:
"""Environment variable that overrides one field of one task class."""
return f"{_ENV_PREFIX}_{task_class.upper()}_{field.upper()}"
def _override(task_class: str, field: str, default: float | None) -> float | None:
"""Read one override, falling back to *default* on anything unusable.
A malformed or non-positive override is ignored rather than raised: a typo
in an environment variable must not be able to mint a zero-length lease that
makes every claim instantly reclaimable, nor crash the server at import.
"""
raw = (os.environ.get(env_var_name(task_class, field)) or "").strip()
if not raw:
return default
try:
value = float(raw)
except (TypeError, ValueError):
return default
if value <= 0:
return default
return value
def policy_for(task_class: str) -> LeasePolicy:
"""Return the effective policy for *task_class*.
Unknown task classes fall back to the author policy, which is the most
conservative migrated class, rather than raising a new caller must never
be able to crash a lock write by naming a class this table has not learned.
"""
key = str(task_class or "").strip() or TASK_CLASS_AUTHOR_ISSUE_WORK
base = _DEFAULTS.get(key) or _DEFAULTS[TASK_CLASS_AUTHOR_ISSUE_WORK]
resolved = {
field: _override(base.task_class, field, getattr(base, field))
for field in _NUMERIC_FIELDS
}
if all(resolved[field] == getattr(base, field) for field in _NUMERIC_FIELDS):
return base
return LeasePolicy(
task_class=base.task_class,
terminal_retirement_eligible=base.terminal_retirement_eligible,
heartbeat_lifecycle_active=base.heartbeat_lifecycle_active,
**resolved,
)
def known_task_classes() -> tuple[str, ...]:
"""Every declared task class, migrated or not."""
return tuple(_DEFAULTS)
def describe(task_class: str) -> dict[str, Any]:
"""Serializable view of a policy, for audit records and tool payloads."""
policy = policy_for(task_class)
return {
"task_class": policy.task_class,
"initial_ttl_minutes": policy.initial_ttl_minutes,
"heartbeat_cadence_minutes": policy.heartbeat_cadence_minutes,
"stale_warning_minutes": policy.stale_warning_minutes,
"missed_heartbeat_grace_minutes": policy.missed_heartbeat_grace_minutes,
"absolute_cap_hours": policy.absolute_cap_hours,
"recovery_grace_minutes": policy.recovery_grace_minutes,
"terminal_race_drain_minutes": policy.terminal_race_drain_minutes,
"terminal_retirement_eligible": policy.terminal_retirement_eligible,
"heartbeat_lifecycle_active": policy.heartbeat_lifecycle_active,
"lifecycle_version": LIFECYCLE_HEARTBEAT_V1,
}
+475
View File
@@ -0,0 +1,475 @@
"""Inventory and fail-closed guards for MCP restart/reload/kill paths (#657).
Single source of truth enumerating every code/script/doc path that can
restart, reload, reconnect, kill, or force-recreate an MCP process. Each path
is classified and linked to the guard that constrains it. The companion
human-readable inventory lives in ``docs/mcp-restart-path-inventory.md`` and is
kept in lock-step with this module by ``tests/test_mcp_restart_paths.py``.
Design intent (aligns with #655 restart-coordinator roadmap):
* **No unguarded full restart.** The in-process MCP daemon
(``gitea_mcp_server.py`` / ``mcp_server.py`` / ``role_session_router.py``)
must never replace or kill its own process replacing the process after the
host wired up the stdio pipes desyncs the JSON-RPC transport (observed with
Antigravity/Cascade hosts). ``assert_no_daemon_self_replacement`` enforces
this against the live source tree.
* **No legacy auto-restart helper.** ``_trigger_mcp_auto_restart`` was removed
when the stale-runtime resolver became side-effect free (#685);
``assert_auto_restart_helper_absent`` keeps it removed.
* **Unknown restart attempts fail closed.** LLM tools must route any restart
intent through a *registered* path. ``assert_restart_attempt_registered``
raises ``UnknownRestartPathError`` for anything not in this inventory.
* **pkill stays forbidden (#630).** Manual daemon kills are classified as
contamination by :mod:`runtime_recovery_guard`; this module records that path
and the test asserts the classification still holds.
This module performs no restarts, spawns no threads, and touches no config or
process state. It is pure inventory + read-only source assertions.
"""
from __future__ import annotations
import os
from dataclasses import dataclass
from pathlib import Path
from typing import Iterable
# --- Classifications -------------------------------------------------------
#: A narrow, one-shot recovery that is safe by construction (e.g. a CLI wrapper
#: re-execing into the venv interpreter before importing anything, or an
#: in-process profile switch). Never targets the running MCP daemon process.
CLASS_SANCTIONED_NARROW = "sanctioned_narrow_recovery"
#: The path detects a condition that would require a restart, then *fails
#: closed* on mutations and emits restart/reconnect guidance. It never restarts
#: the process itself (recovery is owned by the host/operator).
CLASS_GUARDED_FAIL_CLOSED = "guarded_fail_closed"
#: The path is forbidden. Attempting it is a workflow-safety violation and,
#: where an LLM tool could invoke it, is marked as contamination.
CLASS_FORBIDDEN = "forbidden"
#: A previously-existing unguarded restart primitive that has been deleted. A
#: regression guard keeps it absent.
CLASS_REMOVED = "removed"
#: Behavior that lives in the host/IDE and is outside this process's control
#: (e.g. a manual ``/mcp reconnect``). Documented, not code-guarded here.
CLASS_HOST_RESIDUAL = "host_residual"
VALID_CLASSIFICATIONS = frozenset(
{
CLASS_SANCTIONED_NARROW,
CLASS_GUARDED_FAIL_CLOSED,
CLASS_FORBIDDEN,
CLASS_REMOVED,
CLASS_HOST_RESIDUAL,
}
)
#: The in-process MCP daemon modules. These must never self-replace/self-kill.
DAEMON_MODULES = (
"gitea_mcp_server.py",
"mcp_server.py",
"role_session_router.py",
)
#: The legacy auto-restart helper removed in #685. Must stay removed.
LEGACY_AUTO_RESTART_HELPER = "_trigger_mcp_auto_restart"
#: Call patterns that would let the daemon replace or terminate its own
#: process. Matched as calls (trailing ``(``) so prose/docstring mentions such
#: as "we do NOT os.execv() here" or "never calls ``os._exit``" do not trip the
#: scanner (comment lines are stripped first regardless).
DAEMON_SELF_REPLACEMENT_PRIMITIVES = (
"os.execv(",
"os.execve(",
"os.execvp(",
"os.execvpe(",
"os.kill(",
"os.killpg(",
"os._exit(",
"os.abort(",
)
@dataclass(frozen=True)
class RestartPath:
"""One classified restart/reload/kill path in the inventory."""
path_id: str
title: str
mechanism: str
classification: str
guard: str
locations: tuple[str, ...]
references: tuple[str, ...]
residual_host: bool = False
notes: str = ""
class UnknownRestartPathError(RuntimeError):
"""Raised when a restart attempt is not a registered, classified path."""
# --- The inventory ---------------------------------------------------------
_RESTART_PATHS: tuple[RestartPath, ...] = (
RestartPath(
path_id="cli_venv_bootstrap_execv",
title="CLI wrapper venv re-exec",
mechanism=(
"Standalone CLI scripts re-exec into venv/bin/python3 via os.execv "
"at import top, guarded by `sys.executable != venv_python`."
),
classification=CLASS_SANCTIONED_NARROW,
guard=(
"One-shot, pre-import bootstrap; runs before any MCP transport "
"exists and only when not already on the venv interpreter, so it "
"cannot desync a live daemon. Idempotent guard condition prevents "
"a re-exec loop."
),
locations=(
"create_pr.py",
"create_issue.py",
"close_issue.py",
"merge_pr.py",
"review_pr.py",
"edit_pr.py",
"delete_branch.py",
"mark_issue.py",
"manage_labels.py",
"list_issues.py",
"list_prs.py",
),
references=("#657",),
),
RestartPath(
path_id="daemon_self_replacement",
title="MCP daemon self-replacement",
mechanism=(
"The in-process MCP daemon replacing/terminating its own process "
"(os.execv/os.kill/os._exit) to reload code."
),
classification=CLASS_FORBIDDEN,
guard=(
"Forbidden by design: replacing the process after the host wired "
"up stdio desyncs JSON-RPC (Antigravity/Cascade). Enforced against "
"the source tree by assert_no_daemon_self_replacement()."
),
locations=("gitea_mcp_server.py:~155 (decision comment)",) + DAEMON_MODULES,
references=("#657", "#584"),
),
RestartPath(
path_id="legacy_auto_restart_helper",
title="Legacy _trigger_mcp_auto_restart helper",
mechanism=(
"A helper that actively restarted the MCP server from the "
"read-only resolver path."
),
classification=CLASS_REMOVED,
guard=(
"Removed in #685 when the resolver became side-effect free. Kept "
"absent by assert_auto_restart_helper_absent()."
),
locations=("gitea_mcp_server.py", "mcp_server.py"),
references=("#685", "#657"),
),
RestartPath(
path_id="config_touch_reload",
title="MCP client config-touch reload",
mechanism=(
"Touching (utime) the MCP client config file to make the host "
"reload/recreate the server process."
),
classification=CLASS_REMOVED,
guard=(
"Removed from the resolver in #685: stale-runtime detection is "
"report-only and never mutates client config, spawns threads, or "
"calls os._exit."
),
locations=("gitea_mcp_server.py (resolve_task_capability)",),
references=("#685", "#657"),
),
RestartPath(
path_id="master_advance_auto_restart",
title="Master-advance staleness gate",
mechanism=(
"On-disk master advancing past the running code. The master-parity "
"gate detects it and fails mutations closed with restart guidance."
),
classification=CLASS_GUARDED_FAIL_CLOSED,
guard=(
"Detect + fail closed only; the process never self-restarts. "
"master_parity_gate captures startup parity and blocks mutations "
"while stale, emitting restart/reconnect guidance."
),
locations=(
"master_parity_gate.py",
"gitea_mcp_server.py (gitea_assess_master_parity)",
),
references=("#420", "#591", "#657"),
),
RestartPath(
path_id="stale_runtime_resolver_reconnect",
title="Stale-runtime resolver reconnect guidance",
mechanism=(
"The capability resolver detecting a stale serving process and "
"reporting restart_required/stop_required for a client reconnect."
),
classification=CLASS_GUARDED_FAIL_CLOSED,
guard=(
"Report-only (#685): returns restart_required/stop_required and an "
"exact_safe_next_action pointing at IDE/client reconnect; performs "
"no restart, thread spawn, config touch, or os._exit."
),
locations=("gitea_mcp_server.py (gitea_resolve_task_capability)",),
references=("#685", "#657"),
),
RestartPath(
path_id="manual_daemon_kill",
title="Manual daemon kill (pkill/killall/kill)",
mechanism=(
"Shell kills of the MCP daemon: `pkill -f mcp_server.py`, "
"`killall`, broad `pkill -f python` sweeps, or `kill <pid>` of a "
"daemon pid."
),
classification=CLASS_FORBIDDEN,
guard=(
"Forbidden (#630): runtime_recovery_guard classifies these as "
"contamination and gitea_record_daemon_process_kill_attempt writes "
"a durable marker that fails subsequent mutations closed. Operator "
"maintenance authorization is read only from the environment, not "
"from a tool argument."
),
locations=(
"runtime_recovery_guard.py",
"gitea_mcp_server.py (gitea_record_daemon_process_kill_attempt)",
),
references=("#630", "#657"),
),
RestartPath(
path_id="conflict_marker_infra_stop",
title="Startup conflict-marker infra stop",
mechanism=(
"The daemon entrypoint scans for unresolved merge-conflict markers "
"at startup and stops (sys.exit(1)) if found."
),
classification=CLASS_GUARDED_FAIL_CLOSED,
guard=(
"Fail-closed startup stop, not a restart: the process exits and "
"waits for the operator to resolve conflicts and relaunch. Never "
"self-restarts or loops."
),
locations=("mcp_server.py (check_conflict_markers)",),
references=("#657",),
),
RestartPath(
path_id="ide_client_reconnect",
title="Host/IDE MCP reconnect",
mechanism=(
"A manual `/mcp reconnect` (or equivalent host action) that the "
"IDE performs to recreate the MCP client connection."
),
classification=CLASS_HOST_RESIDUAL,
guard=(
"Outside this process's control. It is the sanctioned recovery the "
"gates point operators toward; documented as residual host "
"behavior. No in-process code initiates it."
),
locations=("host/IDE",),
references=("#584", "#656", "#657"),
residual_host=True,
),
RestartPath(
path_id="profile_switch_runtime",
title="Runtime profile switch",
mechanism=(
"Switching the active execution profile at runtime "
"(dynamic-profile mode)."
),
classification=CLASS_SANCTIONED_NARROW,
guard=(
"In-process and restart-free: runtime_switching_supported is true, "
"so a profile switch rebinds capability without recreating the "
"process. No restart primitive is invoked."
),
locations=("gitea_mcp_server.py (gitea_activate_profile)",),
references=("#656", "#657"),
),
)
_BY_ID: dict[str, RestartPath] = {p.path_id: p for p in _RESTART_PATHS}
# --- Read-only accessors ---------------------------------------------------
def iter_restart_paths() -> tuple[RestartPath, ...]:
"""Return the full inventory as an immutable tuple."""
return _RESTART_PATHS
def restart_path_ids() -> frozenset[str]:
"""Return the set of registered path ids."""
return frozenset(_BY_ID)
def get_restart_path(path_id: str) -> RestartPath:
"""Return the registered path, or raise :class:`UnknownRestartPathError`."""
try:
return _BY_ID[path_id]
except KeyError as exc:
raise UnknownRestartPathError(
f"unknown restart path id {path_id!r}; not in the #657 inventory"
) from exc
def paths_by_classification(classification: str) -> tuple[RestartPath, ...]:
"""Return all registered paths with the given classification."""
if classification not in VALID_CLASSIFICATIONS:
raise ValueError(f"unknown classification {classification!r}")
return tuple(p for p in _RESTART_PATHS if p.classification == classification)
def assert_restart_attempt_registered(path_id: str) -> RestartPath:
"""Fail closed unless ``path_id`` is a registered, classified restart path.
LLM tools that intend to trigger any restart/reload/reconnect must name a
registered path so an unknown/novel restart primitive cannot slip through
silently. Forbidden and removed paths are registered too this only
asserts the attempt is *known*, not that it is *permitted*; callers must
still honor the classification.
"""
return get_restart_path(path_id)
def assert_registry_wellformed() -> None:
"""Validate the inventory's own invariants (fail closed on drift)."""
seen: set[str] = set()
for path in _RESTART_PATHS:
if path.path_id in seen:
raise ValueError(f"duplicate restart path id {path.path_id!r}")
seen.add(path.path_id)
if path.classification not in VALID_CLASSIFICATIONS:
raise ValueError(
f"{path.path_id!r} has invalid classification "
f"{path.classification!r}"
)
if not path.guard.strip():
raise ValueError(f"{path.path_id!r} is missing a guard description")
if not path.references:
raise ValueError(f"{path.path_id!r} is missing references")
if not path.locations:
raise ValueError(f"{path.path_id!r} is missing locations")
if path.classification == CLASS_HOST_RESIDUAL and not path.residual_host:
raise ValueError(
f"{path.path_id!r} is host_residual but residual_host is False"
)
# --- Source-tree guards ----------------------------------------------------
def _repo_root(root: str | os.PathLike[str] | None = None) -> Path:
if root is not None:
return Path(root)
return Path(__file__).resolve().parent
def _iter_code_lines(text: str) -> Iterable[tuple[int, str]]:
"""Yield (1-based lineno, line) for lines that are not full-line comments."""
for lineno, line in enumerate(text.splitlines(), start=1):
if line.lstrip().startswith("#"):
continue
yield lineno, line
def scan_daemon_self_replacement(
root: str | os.PathLike[str] | None = None,
) -> list[dict[str, object]]:
"""Return violations where a daemon module could self-replace/self-kill.
Scans :data:`DAEMON_MODULES` for calls in
:data:`DAEMON_SELF_REPLACEMENT_PRIMITIVES`. Full-line comments are ignored,
and only call forms (with a trailing ``(``) match, so decision comments and
docstrings that merely mention the primitives do not produce false hits.
"""
repo = _repo_root(root)
violations: list[dict[str, object]] = []
for module in DAEMON_MODULES:
path = repo / module
if not path.exists():
continue
text = path.read_text(encoding="utf-8", errors="replace")
for lineno, line in _iter_code_lines(text):
for primitive in DAEMON_SELF_REPLACEMENT_PRIMITIVES:
if primitive in line:
violations.append(
{
"module": module,
"line": lineno,
"primitive": primitive,
"text": line.strip(),
}
)
return violations
def assert_no_daemon_self_replacement(
root: str | os.PathLike[str] | None = None,
) -> None:
"""Fail closed if any daemon module can restart/kill its own process."""
violations = scan_daemon_self_replacement(root)
if violations:
rendered = "; ".join(
f"{v['module']}:{v['line']} {v['primitive']}" for v in violations
)
raise AssertionError(
"MCP daemon must never self-replace/self-kill (#657); found: "
f"{rendered}"
)
def scan_auto_restart_helper(
root: str | os.PathLike[str] | None = None,
) -> list[dict[str, object]]:
"""Return occurrences of a *definition* of the legacy auto-restart helper."""
repo = _repo_root(root)
needle = f"def {LEGACY_AUTO_RESTART_HELPER}"
hits: list[dict[str, object]] = []
for module in DAEMON_MODULES:
path = repo / module
if not path.exists():
continue
text = path.read_text(encoding="utf-8", errors="replace")
for lineno, line in _iter_code_lines(text):
if needle in line:
hits.append({"module": module, "line": lineno})
return hits
def assert_auto_restart_helper_absent(
root: str | os.PathLike[str] | None = None,
) -> None:
"""Fail closed if the removed ``_trigger_mcp_auto_restart`` reappears."""
hits = scan_auto_restart_helper(root)
if hits:
rendered = "; ".join(f"{h['module']}:{h['line']}" for h in hits)
raise AssertionError(
f"{LEGACY_AUTO_RESTART_HELPER} was removed in #685 and must not "
f"return (#657); found definition at: {rendered}"
)
+58
View File
@@ -566,6 +566,10 @@ def build_pr_cleanup_entry(
worktree_state=worktree_state,
active_lock=active_lock,
)
planned = plan_cleanup_execution_order(
remote_assessment=remote,
local_assessment=local,
)
return {
"pr_number": pr_number,
"issue_number": issue_number,
@@ -576,9 +580,63 @@ def build_pr_cleanup_entry(
"merged": merged,
"remote_branch": remote,
"local_worktree": local,
# #851: dry-run and execute share the same lifecycle order description.
"planned_execution_order": planned,
}
def plan_cleanup_execution_order(
*,
remote_assessment: dict[str, Any] | None,
local_assessment: dict[str, Any] | None,
) -> list[dict[str, Any]]:
"""Describe independent worktree-then-reassess-then-remote cleanup order (#851).
Remote ownership protection remains fail-closed at execute time. A worktree
that is independently safe to remove is never skipped merely because remote
deletion may be blocked by that same ``worktree_binding``.
"""
remote = remote_assessment or {}
local = local_assessment or {}
steps: list[dict[str, Any]] = []
worktree_safe = bool(local.get("safe_to_remove_worktree"))
remote_safe = bool(remote.get("safe_to_delete_remote"))
if worktree_safe:
steps.append(
{
"action": "remove_local_worktree",
"reason": "independently_safe_to_remove",
"phase": 1,
}
)
if remote_safe:
if worktree_safe:
steps.append(
{
"action": "reassess_branch_ownership",
"reason": "after_worktree_removal_clear_worktree_binding",
"phase": 2,
}
)
steps.append(
{
"action": "delete_remote_branch",
"reason": "only_if_independently_safe_after_reassessment",
"phase": 3,
}
)
else:
steps.append(
{
"action": "delete_remote_branch",
"reason": "safe_to_delete_and_no_independent_worktree_removal",
"phase": 1,
}
)
return steps
def build_reconciliation_report(
*,
project_root: str,
+791
View File
@@ -0,0 +1,791 @@
"""Post-restart MCP reconciliation and completion proof (#662).
After an MCP process restart, sessions, leases, capabilities, worktrees, and
interrupted mutations are not systematically reconciled; operators rebuild
context from chat. This module is the pure classification core of the
post-restart reconcile path.
Design rules (mirrors ``restart_coordinator`` / ``workflow_dashboard``):
* **Pure classification.** :func:`reconcile_after_restart` takes an already
gathered inventory and returns a structured *completion proof*. It never
touches the network, the filesystem, or a live process, so multi-session
fixtures can drive every branch in unit tests.
* **Fail closed.** Incomplete inventory never reports overall ``complete``.
Ambiguous interrupted mutations are ``unresolved`` (never silently resumed).
* **No blind write resume.** The proof never authorizes replaying a mutation;
it only classifies evidence and names follow-up work.
* **#660 soft dependency.** When durable session checkpoints are not present
in the inventory, the checkpoint dimension is ``skipped`` with an explicit
reason rather than inventing a schema (#660 lands separately).
* **Log-only then enforce.** Default mode is ``log_only``. ``enforce`` sets
``mutation_hold`` when anything remains unresolved so callers can block
write ops until reconcile is complete or degraded mode is documented.
The single sanctioned gather+classify entry point is the MCP tool
``gitea_reconcile_after_restart`` (read-only inventory gather + pure classify).
Creating durable follow-up Gitea issues from unresolved items is an explicit
apply step outside this pure module.
"""
from __future__ import annotations
from dataclasses import dataclass, field
from datetime import datetime, timezone
from typing import Any, Mapping, Sequence
from uuid import uuid4
import lease_lifecycle
RECONCILE_VERSION = "1.0.0-issue-662"
SCHEMA_VERSION = 1
# Overall proof statuses.
STATUS_COMPLETE = "complete"
STATUS_DEGRADED = "degraded"
STATUS_FAILED = "failed"
# Per-dimension item statuses.
ITEM_RESOLVED = "resolved"
ITEM_UNRESOLVED = "unresolved"
ITEM_DEGRADED = "degraded"
ITEM_SKIPPED = "skipped"
# Modes.
MODE_LOG_ONLY = "log_only"
MODE_ENFORCE = "enforce"
# Lease / session phases that imply a write critical section was in flight.
MUTATING_PHASES = frozenset(
{
"implementing",
"publishing",
"merging",
"reviewing",
"committing",
"pushing",
"closing",
"mutating",
"critical_section",
}
)
# Dimensions the acceptance criteria require.
DIM_SERVICE_HEALTH = "service_health"
DIM_CLIENTS = "clients"
DIM_SESSIONS = "sessions"
DIM_CHECKPOINTS = "checkpoints"
DIM_LEASES = "leases"
DIM_CAPABILITIES = "capabilities"
DIM_WORKTREES = "worktrees"
DIM_MUTATIONS = "interrupted_mutations"
DIM_DUPLICATES = "duplicates"
DIM_QUEUE = "queue"
def _utc_now() -> datetime:
return datetime.now(timezone.utc)
def _ts(dt: datetime) -> str:
return dt.astimezone(timezone.utc).isoformat().replace("+00:00", "Z")
@dataclass(frozen=True)
class ReconcileItem:
"""One dimension of the post-restart reconcile report."""
dimension: str
status: str
summary: str
details: dict[str, Any] = field(default_factory=dict)
follow_up_required: bool = False
def as_dict(self) -> dict[str, Any]:
return {
"dimension": self.dimension,
"status": self.status,
"summary": self.summary,
"details": dict(self.details),
"follow_up_required": self.follow_up_required,
}
@dataclass(frozen=True)
class FollowUpIssue:
"""A durable follow-up issue the apply path may create for unresolved work."""
title: str
body: str
dimension: str
severity: str = "high"
def as_dict(self) -> dict[str, Any]:
return {
"title": self.title,
"body": self.body,
"dimension": self.dimension,
"severity": self.severity,
}
@dataclass(frozen=True)
class RestartCompletionProof:
"""Machine-readable post-restart completion proof (#662 AC2)."""
schema_version: int
reconcile_version: str
reconcile_id: str
started_at: str
finished_at: str
boot_head_sha: str | None
current_head_sha: str | None
inventory_complete: bool
incomplete_reasons: tuple[str, ...]
mode: str
mutation_hold: bool
overall_status: str
items: tuple[ReconcileItem, ...]
proposed_follow_ups: tuple[FollowUpIssue, ...]
resolved_count: int
unresolved_count: int
skipped_count: int
note: str
def as_dict(self) -> dict[str, Any]:
return {
"schema_version": self.schema_version,
"reconcile_version": self.reconcile_version,
"reconcile_id": self.reconcile_id,
"started_at": self.started_at,
"finished_at": self.finished_at,
"boot_head_sha": self.boot_head_sha,
"current_head_sha": self.current_head_sha,
"inventory_complete": self.inventory_complete,
"incomplete_reasons": list(self.incomplete_reasons),
"mode": self.mode,
"mutation_hold": self.mutation_hold,
"overall_status": self.overall_status,
"items": [i.as_dict() for i in self.items],
"proposed_follow_ups": [f.as_dict() for f in self.proposed_follow_ups],
"resolved_count": self.resolved_count,
"unresolved_count": self.unresolved_count,
"skipped_count": self.skipped_count,
"note": self.note,
"links": {
"umbrella": 655,
"vision": 652,
"roadmap": 653,
"issue": 662,
"checkpoint_schema": 660,
"drain_proof": 661,
},
}
def _item(
dimension: str,
status: str,
summary: str,
*,
details: dict[str, Any] | None = None,
follow_up: bool = False,
) -> ReconcileItem:
return ReconcileItem(
dimension=dimension,
status=status,
summary=summary,
details=dict(details or {}),
follow_up_required=follow_up,
)
def _lease_freshness(lease: Mapping[str, Any]) -> str:
fr = lease.get("freshness")
if isinstance(fr, Mapping):
return str(fr.get("freshness") or fr.get("status") or "unknown")
if isinstance(fr, str):
return fr
# Fall back to pure classifier when raw lease rows are supplied.
try:
return str(lease_lifecycle.classify_lease_freshness(dict(lease)).get("freshness") or "unknown")
except Exception: # noqa: BLE001 - pure path must not raise on bad rows
return "unknown"
def _is_live_freshness(freshness: str) -> bool:
return freshness in {"active", "live", "fresh"}
def _is_mutating_phase(phase: str | None) -> bool:
p = (phase or "").strip().lower()
if not p:
return False
if p in MUTATING_PHASES:
return True
# Soft match for compound phases like "author_implementing".
return any(token in p for token in MUTATING_PHASES)
def _detect_interrupted_mutations(
leases: Sequence[Mapping[str, Any]],
pending_mutations: Sequence[Mapping[str, Any]],
) -> list[dict[str, Any]]:
"""Return interrupted-mutation evidence (never auto-resumes writes)."""
found: list[dict[str, Any]] = []
for raw in pending_mutations or ():
if not isinstance(raw, Mapping):
continue
found.append(
{
"source": "pending_mutation_inventory",
"status": "unresolved",
"phase": raw.get("phase"),
"session_id": raw.get("session_id"),
"work_kind": raw.get("work_kind") or raw.get("kind"),
"work_number": raw.get("work_number") or raw.get("number"),
"reason": raw.get("reason")
or "pending mutation recorded across process restart",
"resume_allowed": False,
}
)
for lease in leases or ():
if not isinstance(lease, Mapping):
continue
phase = lease.get("phase")
freshness = _lease_freshness(lease)
if not _is_mutating_phase(str(phase) if phase is not None else None):
continue
# A mutating phase whose owner is not live is interrupted.
if _is_live_freshness(freshness):
# Still live after restart is itself surprising — flag for review.
found.append(
{
"source": "lease_mutating_phase",
"status": "unresolved",
"phase": phase,
"freshness": freshness,
"lease_id": lease.get("lease_id"),
"session_id": lease.get("session_id"),
"work_kind": lease.get("work_kind"),
"work_number": lease.get("work_number"),
"worktree_path": lease.get("worktree_path"),
"reason": (
"mutating lease phase still classified live after restart; "
"do not auto-resume writes"
),
"resume_allowed": False,
}
)
else:
found.append(
{
"source": "lease_mutating_phase",
"status": "unresolved",
"phase": phase,
"freshness": freshness,
"lease_id": lease.get("lease_id"),
"session_id": lease.get("session_id"),
"work_kind": lease.get("work_kind"),
"work_number": lease.get("work_number"),
"worktree_path": lease.get("worktree_path"),
"reason": (
f"mutating lease phase '{phase}' with non-live freshness "
f"'{freshness}' — interrupted by restart"
),
"resume_allowed": False,
}
)
return found
def _detect_duplicate_work(
leases: Sequence[Mapping[str, Any]],
) -> list[dict[str, Any]]:
"""Surface duplicate live claims on the same work item."""
by_work: dict[tuple[Any, Any], list[Mapping[str, Any]]] = {}
for lease in leases or ():
if not isinstance(lease, Mapping):
continue
if not _is_live_freshness(_lease_freshness(lease)):
continue
key = (lease.get("work_kind"), lease.get("work_number"))
if key[0] is None or key[1] is None:
continue
by_work.setdefault(key, []).append(lease)
dups: list[dict[str, Any]] = []
for (kind, number), rows in sorted(by_work.items(), key=lambda kv: str(kv[0])):
if len(rows) < 2:
continue
dups.append(
{
"work_kind": kind,
"work_number": number,
"claim_count": len(rows),
"session_ids": [r.get("session_id") for r in rows],
"lease_ids": [r.get("lease_id") for r in rows],
}
)
return dups
def _follow_up_for_item(item: ReconcileItem) -> FollowUpIssue | None:
if not item.follow_up_required:
return None
title = f"[post-restart] unresolved {item.dimension} after MCP restart"
body = (
f"## Post-restart reconcile follow-up (#662)\n\n"
f"**Dimension:** `{item.dimension}`\n"
f"**Status:** `{item.status}`\n"
f"**Summary:** {item.summary}\n\n"
f"```json\n{item.details!r}\n```\n\n"
f"Parent umbrella: #655 · Vision: #652 · Roadmap: #653 · Reconcile: #662\n"
f"Do **not** auto-resume write mutations; reconcile evidence first.\n"
)
return FollowUpIssue(
title=title,
body=body,
dimension=item.dimension,
severity="high" if item.dimension == DIM_MUTATIONS else "medium",
)
def reconcile_after_restart(
inventory: Mapping[str, Any],
*,
now: datetime | None = None,
mode: str = MODE_LOG_ONLY,
reconcile_id: str | None = None,
) -> RestartCompletionProof:
"""Classify a post-restart inventory into a completion proof (#662).
Parameters
----------
inventory:
Gathered facts. Expected keys (all optional except completeness):
* ``inventory_complete`` (bool) fail closed when false
* ``incomplete_reasons`` (list[str])
* ``service_health`` (dict with ``healthy`` bool)
* ``clients`` (list) connected client descriptors
* ``sessions`` (list)
* ``leases`` (list, optionally with ``freshness``)
* ``checkpoints`` (list | None) durable session checkpoints (#660)
* ``checkpoints_available`` (bool) False when #660 schema absent
* ``worktree_bindings`` (list)
* ``pending_mutations`` (list) explicit interrupted-mutation evidence
* ``capabilities`` (dict with optional ``stale`` / heads)
* ``boot_head_sha`` / ``current_head_sha``
* ``queue_state`` (dict)
mode:
``log_only`` (default) or ``enforce`` (sets mutation_hold on unresolved).
"""
started = now or _utc_now()
mode_norm = (mode or MODE_LOG_ONLY).strip().lower()
if mode_norm not in {MODE_LOG_ONLY, MODE_ENFORCE}:
mode_norm = MODE_LOG_ONLY
inventory_complete = bool(inventory.get("inventory_complete", False))
incomplete_reasons = tuple(
str(r) for r in (inventory.get("incomplete_reasons") or []) if str(r).strip()
)
items: list[ReconcileItem] = []
# --- service health -------------------------------------------------
health = inventory.get("service_health") or {}
if not isinstance(health, Mapping):
health = {}
if not inventory_complete and "service_health" not in inventory:
items.append(
_item(
DIM_SERVICE_HEALTH,
ITEM_UNRESOLVED,
"service health unknown because inventory is incomplete",
details={"inventory_complete": False},
follow_up=True,
)
)
elif health.get("healthy") is True:
items.append(
_item(
DIM_SERVICE_HEALTH,
ITEM_RESOLVED,
"service health verified",
details=dict(health),
)
)
elif health.get("healthy") is False:
items.append(
_item(
DIM_SERVICE_HEALTH,
ITEM_UNRESOLVED,
"service health check failed",
details=dict(health),
follow_up=True,
)
)
else:
items.append(
_item(
DIM_SERVICE_HEALTH,
ITEM_DEGRADED,
"service health not reported; treating as degraded",
details=dict(health),
follow_up=True,
)
)
# --- clients --------------------------------------------------------
clients = list(inventory.get("clients") or [])
disconnected = [
c
for c in clients
if isinstance(c, Mapping) and c.get("connected") is False
]
if "clients" not in inventory:
items.append(
_item(
DIM_CLIENTS,
ITEM_SKIPPED,
"client inventory not supplied",
details={},
)
)
elif disconnected:
items.append(
_item(
DIM_CLIENTS,
ITEM_UNRESOLVED,
f"{len(disconnected)} disconnected client(s) need reconnect",
details={"disconnected": disconnected, "total": len(clients)},
follow_up=True,
)
)
else:
items.append(
_item(
DIM_CLIENTS,
ITEM_RESOLVED,
f"{len(clients)} client(s) accounted for",
details={"total": len(clients)},
)
)
# --- sessions -------------------------------------------------------
sessions = [s for s in (inventory.get("sessions") or []) if isinstance(s, Mapping)]
orphan_sessions = [
s
for s in sessions
if str(s.get("status") or "").lower() == "active"
and s.get("pid") is not None
and not lease_lifecycle.is_process_alive(s.get("pid"))
]
if orphan_sessions:
items.append(
_item(
DIM_SESSIONS,
ITEM_UNRESOLVED,
f"{len(orphan_sessions)} active session row(s) with dead owner pid",
details={
"orphan_session_ids": [s.get("session_id") for s in orphan_sessions],
"total_sessions": len(sessions),
},
follow_up=True,
)
)
else:
items.append(
_item(
DIM_SESSIONS,
ITEM_RESOLVED,
f"{len(sessions)} session row(s) reconciled (no dead-pid orphans)",
details={"total_sessions": len(sessions)},
)
)
# --- checkpoints (#660 soft) ----------------------------------------
checkpoints_available = inventory.get("checkpoints_available")
checkpoints = inventory.get("checkpoints")
if checkpoints_available is False or (
checkpoints is None and "checkpoints" not in inventory
):
items.append(
_item(
DIM_CHECKPOINTS,
ITEM_SKIPPED,
"durable session checkpoint schema not available yet (#660)",
details={"depends_on": 660},
)
)
else:
cp_list = [c for c in (checkpoints or []) if isinstance(c, Mapping)]
stale_cp = [c for c in cp_list if c.get("stale") or c.get("invalid")]
if stale_cp:
items.append(
_item(
DIM_CHECKPOINTS,
ITEM_UNRESOLVED,
f"{len(stale_cp)} checkpoint(s) invalid or stale vs live state",
details={"stale_count": len(stale_cp), "total": len(cp_list)},
follow_up=True,
)
)
else:
items.append(
_item(
DIM_CHECKPOINTS,
ITEM_RESOLVED,
f"{len(cp_list)} checkpoint(s) consistent with live state",
details={"total": len(cp_list)},
)
)
# --- leases / locks -------------------------------------------------
leases = [L for L in (inventory.get("leases") or []) if isinstance(L, Mapping)]
live_leases = [L for L in leases if _is_live_freshness(_lease_freshness(L))]
items.append(
_item(
DIM_LEASES,
ITEM_RESOLVED if inventory_complete else ITEM_DEGRADED,
f"{len(live_leases)} live lease(s) of {len(leases)} inventoried",
details={
"live_count": len(live_leases),
"total": len(leases),
"live_lease_ids": [L.get("lease_id") for L in live_leases],
},
follow_up=not inventory_complete,
)
)
# --- capabilities / stale runtime -----------------------------------
caps = inventory.get("capabilities") or {}
if not isinstance(caps, Mapping):
caps = {}
if caps.get("stale") is True:
items.append(
_item(
DIM_CAPABILITIES,
ITEM_UNRESOLVED,
"runtime code is stale vs on-disk master; restart did not reach parity",
details=dict(caps),
follow_up=True,
)
)
else:
items.append(
_item(
DIM_CAPABILITIES,
ITEM_RESOLVED,
"capability/runtime parity acceptable",
details=dict(caps) if caps else {"stale": False},
)
)
# --- worktrees ------------------------------------------------------
bindings = [
b for b in (inventory.get("worktree_bindings") or []) if isinstance(b, Mapping)
]
missing_wt = [
b
for b in bindings
if b.get("missing") is True or b.get("exists") is False
]
if "worktree_bindings" not in inventory:
items.append(
_item(
DIM_WORKTREES,
ITEM_SKIPPED,
"worktree binding inventory not supplied",
)
)
elif missing_wt:
items.append(
_item(
DIM_WORKTREES,
ITEM_UNRESOLVED,
f"{len(missing_wt)} worktree binding(s) missing on disk",
details={"missing": missing_wt, "total": len(bindings)},
follow_up=True,
)
)
else:
items.append(
_item(
DIM_WORKTREES,
ITEM_RESOLVED,
f"{len(bindings)} worktree binding(s) present",
details={"total": len(bindings)},
)
)
# --- interrupted mutations (AC4) ------------------------------------
pending = [
m
for m in (inventory.get("pending_mutations") or [])
if isinstance(m, Mapping)
]
interrupted = _detect_interrupted_mutations(leases, pending)
if interrupted:
items.append(
_item(
DIM_MUTATIONS,
ITEM_UNRESOLVED,
f"{len(interrupted)} interrupted mutation(s); write resume forbidden",
details={"interrupted": interrupted},
follow_up=True,
)
)
else:
items.append(
_item(
DIM_MUTATIONS,
ITEM_RESOLVED,
"no interrupted mutations detected",
details={"interrupted": []},
)
)
# --- duplicates -----------------------------------------------------
dups = _detect_duplicate_work(leases)
if dups:
items.append(
_item(
DIM_DUPLICATES,
ITEM_UNRESOLVED,
f"{len(dups)} work item(s) have multiple live claims",
details={"duplicates": dups},
follow_up=True,
)
)
else:
items.append(
_item(
DIM_DUPLICATES,
ITEM_RESOLVED,
"no duplicate live claims detected",
details={"duplicates": []},
)
)
# --- queue ----------------------------------------------------------
queue = inventory.get("queue_state")
if queue is None:
items.append(
_item(
DIM_QUEUE,
ITEM_SKIPPED,
"allocator queue state not supplied",
)
)
elif isinstance(queue, Mapping) and queue.get("safe_to_resume") is False:
items.append(
_item(
DIM_QUEUE,
ITEM_UNRESOLVED,
"allocator queue not safe to resume",
details=dict(queue),
follow_up=True,
)
)
else:
items.append(
_item(
DIM_QUEUE,
ITEM_RESOLVED,
"allocator queue state acceptable",
details=dict(queue) if isinstance(queue, Mapping) else {},
)
)
# Incomplete inventory always degrades the whole proof.
if not inventory_complete:
# Ensure at least one follow-up names the incomplete inventory.
items.append(
_item(
"inventory",
ITEM_UNRESOLVED,
"control-plane inventory incomplete; reconcile cannot claim success",
details={"reasons": list(incomplete_reasons)},
follow_up=True,
)
)
resolved = sum(1 for i in items if i.status == ITEM_RESOLVED)
unresolved = sum(1 for i in items if i.status in {ITEM_UNRESOLVED, ITEM_DEGRADED})
skipped = sum(1 for i in items if i.status == ITEM_SKIPPED)
if not inventory_complete or any(i.status == ITEM_UNRESOLVED for i in items):
if any(i.status == ITEM_UNRESOLVED for i in items) and inventory_complete:
overall = STATUS_DEGRADED
elif not inventory_complete:
overall = STATUS_FAILED
else:
overall = STATUS_DEGRADED
elif any(i.status == ITEM_DEGRADED for i in items):
overall = STATUS_DEGRADED
else:
overall = STATUS_COMPLETE
# Enforce mode holds mutations whenever anything is unresolved/failed.
mutation_hold = False
if mode_norm == MODE_ENFORCE and overall in {STATUS_DEGRADED, STATUS_FAILED}:
mutation_hold = True
if mode_norm == MODE_ENFORCE and any(
i.dimension == DIM_MUTATIONS and i.status == ITEM_UNRESOLVED for i in items
):
mutation_hold = True
follow_ups = tuple(
fu for i in items if (fu := _follow_up_for_item(i)) is not None
)
finished = _utc_now() if now is None else now
note = (
"Read-only completion proof. Never auto-resumes write mutations. "
"Unresolved items require durable follow-up before claiming clean restart. "
f"Mode={mode_norm}."
)
return RestartCompletionProof(
schema_version=SCHEMA_VERSION,
reconcile_version=RECONCILE_VERSION,
reconcile_id=(reconcile_id or f"reconcile-{uuid4().hex[:12]}"),
started_at=_ts(started),
finished_at=_ts(finished),
boot_head_sha=(
str(inventory.get("boot_head_sha")).strip()
if inventory.get("boot_head_sha")
else None
),
current_head_sha=(
str(inventory.get("current_head_sha")).strip()
if inventory.get("current_head_sha")
else None
),
inventory_complete=inventory_complete,
incomplete_reasons=incomplete_reasons,
mode=mode_norm,
mutation_hold=mutation_hold,
overall_status=overall,
items=tuple(items),
proposed_follow_ups=follow_ups,
resolved_count=resolved,
unresolved_count=unresolved,
skipped_count=skipped,
note=note,
)
def mutations_allowed(proof: RestartCompletionProof | Mapping[str, Any] | None) -> bool:
"""Return whether write mutations may proceed under the given proof."""
if proof is None:
return True # no proof yet → caller decides; enforce path sets hold
if isinstance(proof, RestartCompletionProof):
return not proof.mutation_hold
if isinstance(proof, Mapping):
return not bool(proof.get("mutation_hold"))
return True
+51 -2
View File
@@ -228,25 +228,74 @@ def find_active_reviewer_lease(
return None
def _conflict_fix_chain_key(lease: dict) -> tuple | None:
"""Identity of the lease chain a conflict-fix marker belongs to (#842).
Keyed by PR number, profile, head_before, and branch. Returns None when any
required component (pr_number, profile, head_before) is missing or malformed.
"""
raw = lease.get("raw_fields") or {}
pr_number = lease.get("pr_number")
profile = (lease.get("profile") or "").strip().lower()
head_before = lease.get("head_before")
branch = (lease.get("branch") or raw.get("branch") or "").strip()
if not (pr_number and profile and head_before):
return None
return (pr_number, profile, head_before, branch)
def _conflict_fix_chain_matches(key1: tuple, key2: tuple) -> bool:
"""True when two conflict-fix chain keys refer to the same lease chain."""
pr1, profile1, head1, branch1 = key1
pr2, profile2, head2, branch2 = key2
if pr1 != pr2 or profile1 != profile2 or head1 != head2:
return False
if branch1 and branch2 and branch1 != branch2:
return False
return True
def _conflict_fix_chain_terminated_after(entries: list[dict], index: int) -> bool:
"""True when a later marker terminates the conflict-fix chain of ``entries[index]``.
Append-only newest-wins: a terminal marker (phase=released/blocked/done)
ends only its matching claim chain (#842).
"""
key = _conflict_fix_chain_key(entries[index])
if key is None:
return False
for later in entries[index + 1:]:
phase = (later.get("phase") or "").strip().lower()
if phase not in _TERMINAL_CONFLICT_FIX_PHASES:
continue
later_key = _conflict_fix_chain_key(later)
if later_key and _conflict_fix_chain_matches(key, later_key):
return True
return False
def find_active_conflict_fix_lease(
comments: list[dict],
*,
pr_number: int,
now: datetime | None = None,
) -> dict[str, Any] | None:
"""Return the newest unexpired conflict-fix lease for *pr_number*, if any."""
"""Return the newest unexpired, non-terminated conflict-fix lease for *pr_number*, if any."""
now = now or datetime.now(timezone.utc)
candidates = [
entry for entry in _comment_entries(comments, pr_number=pr_number)
if entry.get("lease_kind") == "conflict_fix"
]
for lease in reversed(candidates):
for index in range(len(candidates) - 1, -1, -1):
lease = candidates[index]
if _lease_expired(lease, now=now):
continue
phase = (lease.get("phase") or "").strip().lower()
if phase in _TERMINAL_CONFLICT_FIX_PHASES:
continue
if phase in _ACTIVE_CONFLICT_FIX_PHASES or phase:
if _conflict_fix_chain_terminated_after(candidates, index):
continue
return lease
return None
+451
View File
@@ -0,0 +1,451 @@
"""MCP restart coordinator and impact analysis (#658).
Before any sanctioned MCP restart, a central coordinator must evaluate the
live control-plane state active sessions, leases/locks, in-flight issue/PR
work, mutations, worktrees, and recovery history and produce an *impact
preview* so operators (and the web console, #642/#652) can see the blast
radius **before** concurrent LLM work is disrupted.
Design rules (mirrors the read-only posture of ``workflow_dashboard`` /
``lease_lifecycle``):
* **Pure classification.** :func:`evaluate_restart_impact` takes an already
gathered inventory and returns a structured report. It never touches the
network, the filesystem, or a live process, so multi-session fixtures can
drive every branch in unit tests. The coordinator *never restarts anything*;
a mutative apply path is a later child gated by a drain proof (non-goal here).
* **Fail closed.** If the inventory is not explicitly complete, the verdict is
``unsafe`` / deny an incomplete evaluation must never green-light a restart.
* **No secrets.** Session ids, pids, and profiles are operational metadata, not
credentials; nothing secret flows through this module.
The single sanctioned entry point post-#657 is the MCP tool
``gitea_request_mcp_restart`` (dry-run by default), which gathers the inventory
from the #613 control-plane DB and calls :func:`evaluate_restart_impact`.
"""
from __future__ import annotations
from dataclasses import dataclass, field
from datetime import datetime, timezone
from typing import Any, Mapping, Sequence
import lease_lifecycle
COORDINATOR_VERSION = "1.0.0-issue-658"
# Restart verdicts. Exactly the three the acceptance criteria name.
VERDICT_SAFE = "safe"
VERDICT_UNSAFE = "unsafe"
VERDICT_OVERRIDE = "override"
# Blast-radius severity bands.
BLAST_NONE = "none"
BLAST_LOW = "low"
BLAST_MEDIUM = "medium"
BLAST_HIGH = "high"
# A live lease with a live owner process is treated as active in-flight work.
LEASE_FRESHNESS_LIVE = "active"
# Default staleness window for a session heartbeat (seconds). A session whose
# last heartbeat is older than this is not counted as live even if its row is
# still marked ``active`` — it is assumed dead/detached.
DEFAULT_SESSION_HEARTBEAT_STALE_SECONDS = 900
def _utc_now() -> datetime:
return datetime.now(timezone.utc)
def _parse_ts(value: str | None) -> datetime | None:
return lease_lifecycle._parse_ts(value)
@dataclass(frozen=True)
class SessionImpact:
"""One MCP session a restart would terminate."""
session_id: str
role: str | None
profile: str | None
pid: int | None
status: str | None
alive: bool | None
heartbeat_stale: bool
is_requester: bool
live: bool
def as_dict(self) -> dict[str, Any]:
return {
"session_id": self.session_id,
"role": self.role,
"profile": self.profile,
"pid": self.pid,
"status": self.status,
"alive": self.alive,
"heartbeat_stale": self.heartbeat_stale,
"is_requester": self.is_requester,
"live": self.live,
}
@dataclass(frozen=True)
class LeaseImpact:
"""One control-plane lease a restart would disrupt."""
lease_id: str | None
session_id: str | None
role: str | None
phase: str | None
freshness: str | None
work_kind: str | None
work_number: int | None
worktree_path: str | None
disruptive: bool
is_mutation: bool
is_critical_section: bool
def as_dict(self) -> dict[str, Any]:
return {
"lease_id": self.lease_id,
"session_id": self.session_id,
"role": self.role,
"phase": self.phase,
"freshness": self.freshness,
"work_kind": self.work_kind,
"work_number": self.work_number,
"worktree_path": self.worktree_path,
"disruptive": self.disruptive,
"is_mutation": self.is_mutation,
"is_critical_section": self.is_critical_section,
}
@dataclass(frozen=True)
class RestartImpactReport:
"""Impact preview DTO returned to the console / operator (#642/#652)."""
coordinator_version: str
evaluated_at: str
dry_run: bool
restart_performed: bool
inventory_complete: bool
verdict: str
allow_restart: bool
override_would_allow: bool
operator_override: bool
blast_radius: str
reasons: list[str]
affected_sessions: list[SessionImpact]
affected_leases: list[LeaseImpact]
critical_sections: list[LeaseImpact]
affected_issues: list[int]
affected_prs: list[int]
mutations: list[LeaseImpact]
terminal_lock: dict[str, Any] | None
ack_state: dict[str, str]
prior_recovery_attempts: list[dict[str, Any]]
counts: dict[str, int]
audit_record: dict[str, Any]
incomplete_reasons: list[str] = field(default_factory=list)
def as_dict(self) -> dict[str, Any]:
return {
"coordinator_version": self.coordinator_version,
"evaluated_at": self.evaluated_at,
"dry_run": self.dry_run,
"restart_performed": self.restart_performed,
"inventory_complete": self.inventory_complete,
"incomplete_reasons": list(self.incomplete_reasons),
"verdict": self.verdict,
"allow_restart": self.allow_restart,
"override_would_allow": self.override_would_allow,
"operator_override": self.operator_override,
"blast_radius": self.blast_radius,
"reasons": list(self.reasons),
"affected_sessions": [s.as_dict() for s in self.affected_sessions],
"affected_leases": [l.as_dict() for l in self.affected_leases],
"critical_sections": [l.as_dict() for l in self.critical_sections],
"affected_issues": list(self.affected_issues),
"affected_prs": list(self.affected_prs),
"mutations": [l.as_dict() for l in self.mutations],
"terminal_lock": self.terminal_lock,
"ack_state": dict(self.ack_state),
"prior_recovery_attempts": list(self.prior_recovery_attempts),
"counts": dict(self.counts),
"audit_record": dict(self.audit_record),
}
def _classify_session(
row: Mapping[str, Any],
*,
now: datetime,
requesting_session_id: str | None,
heartbeat_stale_seconds: int,
) -> SessionImpact:
session_id = str(row.get("session_id") or "")
pid = row.get("pid")
status = (row.get("status") or "").strip().lower() or None
alive = lease_lifecycle.is_process_alive(pid) if pid is not None else None
hb = _parse_ts(row.get("last_heartbeat_at"))
heartbeat_stale = bool(
hb is not None and (now - hb).total_seconds() > heartbeat_stale_seconds
)
live = bool(status == "active" and alive is not False and not heartbeat_stale)
return SessionImpact(
session_id=session_id,
role=row.get("role"),
profile=row.get("profile"),
pid=pid,
status=status,
alive=alive,
heartbeat_stale=heartbeat_stale,
is_requester=bool(
requesting_session_id and session_id == requesting_session_id
),
live=live,
)
# Lease phases that represent an active mutation in flight (as opposed to a
# mere allocation/claim with no work committed yet). An active lease in any of
# these phases is a critical section a restart must not sever.
_MUTATING_PHASES = frozenset(
{
"implementing",
"publishing",
"pushing",
"committing",
"reviewing",
"merging",
"reconciling",
"conflict_fix",
}
)
def _classify_lease(row: Mapping[str, Any]) -> LeaseImpact:
freshness_obj = row.get("freshness")
if isinstance(freshness_obj, Mapping):
freshness = str(freshness_obj.get("freshness") or "").strip().lower() or None
else:
freshness = str(freshness_obj or "").strip().lower() or None
phase = (row.get("phase") or "").strip().lower() or None
worktree = row.get("worktree_path")
disruptive = freshness == LEASE_FRESHNESS_LIVE
# A live lease is a mutation-in-flight if it carries an author worktree or
# its phase names a mutating step. All disruptive leases are critical
# sections a restart would sever regardless.
is_mutation = bool(
disruptive and (bool(worktree) or (phase in _MUTATING_PHASES))
)
number = row.get("work_number")
try:
number = int(number) if number is not None else None
except (TypeError, ValueError):
number = None
return LeaseImpact(
lease_id=row.get("lease_id"),
session_id=row.get("session_id"),
role=row.get("role"),
phase=phase,
freshness=freshness,
work_kind=(str(row.get("work_kind") or "").strip().lower() or None),
work_number=number,
worktree_path=worktree,
disruptive=disruptive,
is_mutation=is_mutation,
is_critical_section=disruptive,
)
def _blast_radius(*, session_count: int, work_count: int, mutation_count: int) -> str:
if mutation_count > 0 or work_count >= 3 or session_count >= 3:
return BLAST_HIGH
if work_count > 0 or session_count == 2:
return BLAST_MEDIUM
if session_count == 1:
return BLAST_LOW
return BLAST_NONE
def evaluate_restart_impact(
inventory: Mapping[str, Any],
*,
now: datetime | None = None,
operator_override: bool = False,
requesting_session_id: str | None = None,
dry_run: bool = True,
session_heartbeat_stale_seconds: int = DEFAULT_SESSION_HEARTBEAT_STALE_SECONDS,
) -> RestartImpactReport:
"""Evaluate a proposed MCP restart and return an impact preview.
``inventory`` is a mapping with:
* ``sessions`` session rows (session_id, role, profile, pid, status,
last_heartbeat_at).
* ``leases`` control-plane lease rows, each ideally carrying an enriched
``freshness`` dict (as :func:`lease_lifecycle.list_active_leases` returns);
a bare string freshness is also accepted.
* ``terminal_lock`` the active terminal (merge) lock, if any.
* ``prior_recovery_attempts`` narrower recovery attempts already tried
(e.g. sanctioned reconnects) so the operator sees escalation history.
* ``inventory_complete`` bool. **Must** be explicitly True; a missing or
falsy value forces a deny (fail closed).
* ``incomplete_reasons`` optional reasons the inventory is incomplete.
The coordinator never restarts anything: ``restart_performed`` is always
False and the mutative apply path is a later drain-gated child.
"""
moment = now or _utc_now()
reasons: list[str] = []
inventory_complete = bool(inventory.get("inventory_complete", False))
incomplete_reasons = [str(r) for r in (inventory.get("incomplete_reasons") or [])]
sessions_raw: Sequence[Mapping[str, Any]] = inventory.get("sessions") or []
leases_raw: Sequence[Mapping[str, Any]] = inventory.get("leases") or []
terminal_lock = inventory.get("terminal_lock") or None
prior_recovery_attempts = [
dict(a) for a in (inventory.get("prior_recovery_attempts") or [])
]
session_impacts = [
_classify_session(
s,
now=moment,
requesting_session_id=requesting_session_id,
heartbeat_stale_seconds=session_heartbeat_stale_seconds,
)
for s in sessions_raw
]
lease_impacts = [_classify_lease(l) for l in leases_raw]
# Only *other* live sessions and live leases constitute blast radius: a
# restart that would kill only the requesting session with no other work in
# flight is safe.
other_live_sessions = [
s for s in session_impacts if s.live and not s.is_requester
]
disruptive_leases = [l for l in lease_impacts if l.disruptive]
critical_sections = [l for l in lease_impacts if l.is_critical_section]
mutations = [l for l in lease_impacts if l.is_mutation]
affected_issues = sorted(
{
l.work_number
for l in disruptive_leases
if l.work_kind == "issue" and l.work_number is not None
}
)
affected_prs = sorted(
{
l.work_number
for l in disruptive_leases
if l.work_kind == "pr" and l.work_number is not None
}
)
disruptive = bool(disruptive_leases or other_live_sessions or terminal_lock)
if not inventory_complete:
verdict = VERDICT_UNSAFE
allow_restart = False
reasons.append(
"inventory incomplete: restart evaluation cannot confirm blast "
"radius — deny (fail closed, #658)"
)
reasons.extend(incomplete_reasons)
elif not disruptive:
verdict = VERDICT_SAFE
allow_restart = True
reasons.append("no other live sessions, live leases, or terminal lock")
elif operator_override:
verdict = VERDICT_OVERRIDE
allow_restart = True
reasons.append(
"live work present; operator override accepts the blast radius"
)
else:
verdict = VERDICT_UNSAFE
allow_restart = False
reasons.append(
"live work would be disrupted; restart denied without operator "
"override"
)
if critical_sections and inventory_complete:
reasons.append(
f"{len(critical_sections)} critical section(s) in flight "
"(active lease with a live owner)"
)
if terminal_lock:
reasons.append("active terminal (merge) lock present")
override_would_allow = bool(inventory_complete and disruptive)
blast_radius = _blast_radius(
session_count=len(other_live_sessions),
work_count=len(affected_issues) + len(affected_prs),
mutation_count=len(mutations),
)
# Acknowledgement is a later child (drain protocol); expose per-session
# placeholders so the console can render the ack column now.
ack_state = {s.session_id: "pending" for s in other_live_sessions}
counts = {
"sessions_total": len(session_impacts),
"sessions_live_other": len(other_live_sessions),
"leases_total": len(lease_impacts),
"leases_disruptive": len(disruptive_leases),
"critical_sections": len(critical_sections),
"mutations": len(mutations),
"affected_issues": len(affected_issues),
"affected_prs": len(affected_prs),
"prior_recovery_attempts": len(prior_recovery_attempts),
}
audit_record = {
"event": "restart_impact_evaluated",
"coordinator_version": COORDINATOR_VERSION,
"evaluated_at": moment.isoformat(),
"dry_run": dry_run,
"operator_override": bool(operator_override),
"requesting_session_id": requesting_session_id,
"inventory_complete": inventory_complete,
"verdict": verdict,
"allow_restart": allow_restart,
"blast_radius": blast_radius,
"counts": counts,
}
return RestartImpactReport(
coordinator_version=COORDINATOR_VERSION,
evaluated_at=moment.isoformat(),
dry_run=dry_run,
restart_performed=False,
inventory_complete=inventory_complete,
verdict=verdict,
allow_restart=allow_restart,
override_would_allow=override_would_allow,
operator_override=bool(operator_override),
blast_radius=blast_radius,
reasons=reasons,
affected_sessions=session_impacts,
affected_leases=lease_impacts,
critical_sections=critical_sections,
affected_issues=affected_issues,
affected_prs=affected_prs,
mutations=mutations,
terminal_lock=dict(terminal_lock)
if isinstance(terminal_lock, Mapping)
else terminal_lock,
ack_state=ack_state,
prior_recovery_attempts=prior_recovery_attempts,
counts=counts,
audit_record=audit_record,
incomplete_reasons=incomplete_reasons,
)
+2
View File
@@ -73,6 +73,8 @@ AUTHOR_TASKS = frozenset({
"claim_issue",
"create_branch",
"push_branch",
"bootstrap_author_issue_worktree",
"gitea_bootstrap_author_issue_worktree",
"create_pr",
"comment_pr",
"address_pr_change_requests",
+10 -8
View File
@@ -43,19 +43,21 @@ repo_root="$(cd "$script_dir/.." && pwd)"
# Enforce issue-linked, traceable branch names (issue → branch → worktree → PR).
if [[ "$allow_unlinked" -eq 0 ]]; then
locked_branch=$(python3 -c "
if [[ "$dry_run" -eq 0 ]] && [[ ! "$branch" =~ ^review/pr-[0-9]+-.+ ]]; then
locked_branch=$(python3 -c "
import sys
sys.path.insert(0, '$repo_root')
import issue_lock_store
print(issue_lock_store.resolve_locked_branch_for_session('$branch'))
")
if [[ -z "$locked_branch" ]]; then
echo "Error: No session issue lock is bound. Call gitea_lock_issue before branch creation (fail closed)." >&2
exit 2
fi
if [[ "$branch" != "$locked_branch" ]]; then
echo "Error: Requested branch '$branch' does not match locked branch '$locked_branch' (fail closed)." >&2
exit 2
if [[ -z "$locked_branch" ]]; then
echo "Error: No session issue lock is bound. Call gitea_lock_issue before branch creation (fail closed)." >&2
exit 2
fi
if [[ "$branch" != "$locked_branch" ]]; then
echo "Error: Requested branch '$branch' does not match locked branch '$locked_branch' (fail closed)." >&2
exit 2
fi
fi
if [[ "$branch" =~ ^(fix|feat|docs|chore)/issue-[0-9]+-.+ ]] \
+83
View File
@@ -32,6 +32,35 @@ TASK_CAPABILITY_MAP: dict[str, dict[str, str]] = {
"permission": "gitea.issue.comment",
"role": "author",
},
# #790 Slice A: prove an owned author lease is still active. Strictly
# narrower than lock_issue — it can only slide a lease this exact session
# already owns, never acquire, take over, or revive one — so it gates on the
# same authority rather than introducing an operation name that every
# already-configured author profile would be missing.
"heartbeat_issue_lock": {
"permission": "gitea.issue.comment",
"role": "author",
},
# #860: dirty orphaned same-claimant worktree recovery (explicit operation).
"recover_dirty_orphaned_issue_worktree": {
"permission": "gitea.issue.comment",
"role": "author",
},
"gitea_recover_dirty_orphaned_issue_worktree": {
"permission": "gitea.issue.comment",
"role": "author",
},
# #864: dirty-preserving same-claimant author-session rebind (dead owner PID).
# Author MCP tool path. Reconciler execute is gated inside the tool via
# authorize_reconciler_execute + role_kind checks (not this map entry).
"rebind_dirty_same_claimant_author_session": {
"permission": "gitea.issue.comment",
"role": "author",
},
"gitea_rebind_dirty_same_claimant_author_session": {
"permission": "gitea.issue.comment",
"role": "author",
},
"set_issue_labels": {
"permission": "gitea.issue.comment",
"role": "author",
@@ -58,6 +87,14 @@ TASK_CAPABILITY_MAP: dict[str, dict[str, str]] = {
"permission": "gitea.branch.create",
"role": "author",
},
"bootstrap_author_issue_worktree": {
"permission": "gitea.branch.create",
"role": "author",
},
"gitea_bootstrap_author_issue_worktree": {
"permission": "gitea.branch.create",
"role": "author",
},
"push_branch": {
"permission": "gitea.branch.push",
"role": "author",
@@ -95,6 +132,16 @@ TASK_CAPABILITY_MAP: dict[str, dict[str, str]] = {
"permission": "gitea.branch.push",
"role": "author",
},
# #662: post-restart reconcile is read-only inventory + pure classification.
# Durable follow-up issue creation is a separate apply path (not this task).
"reconcile_after_restart": {
"permission": "gitea.read",
"role": "author",
},
"gitea_reconcile_after_restart": {
"permission": "gitea.read",
"role": "author",
},
# PR synchronization lifecycle: assess is read-only (any role with gitea.read);
# update-by-merge is author-only and mutates the PR head via Gitea API.
"assess_pr_sync_status": {
@@ -335,6 +382,21 @@ TASK_CAPABILITY_MAP: dict[str, dict[str, str]] = {
"role": "controller",
},
# #642: sanctioned host-daemon lifecycle controls. Deliberately *not* a
# ``gitea.*`` operation — restarting an MCP namespace is a host action, not
# a Gitea API call, and no configured Gitea profile should be able to
# satisfy it by accident. Authority comes from the console RBAC model plus
# out-of-band operator authorization (#630); these entries exist so the
# console cannot invent an authority the capability layer never declared.
"restart_namespace": {
"permission": "runtime.restart_namespace",
"role": "controller",
},
"reload_namespace": {
"permission": "runtime.reload_namespace",
"role": "controller",
},
# #601 first-class lease lifecycle — inspect/list need read; mutations gate on
# ownership in the control-plane DB (not a separate Gitea write permission).
"list_workflow_leases": {
@@ -467,6 +529,15 @@ TASK_CAPABILITY_MAP: dict[str, dict[str, str]] = {
"permission": "gitea.issue.comment",
"role": "author",
},
# #651 console analytics ingest — control-plane DB write, not a Gitea API
# call. Authority comes from console RBAC (operator+) plus phase gating;
# permission string is a non-Gitea runtime capability so no Gitea profile
# can satisfy it by accident.
"record_analytics_usage": {
"permission": "runtime.record_analytics_usage",
"role": "author",
},
}
@@ -477,6 +548,15 @@ TASK_CAPABILITY_MAP: dict[str, dict[str, str]] = {
# merger lease (#763).
_PREFLIGHT_TASK_TRANSITIONS = frozenset({
("review_pr", "acquire_reviewer_pr_lease"),
# #850: native author issue worktree bootstrap
("work_issue", "bootstrap_author_issue_worktree"),
("bootstrap_author_issue_worktree", "lock_issue"),
# #860: dirty-orphan recovery and related work_issue transitions (master)
("work_issue", "lock_issue"),
("work_issue", "recover_dirty_orphaned_issue_worktree"),
("work_issue", "gitea_recover_dirty_orphaned_issue_worktree"),
("work_issue", "commit_files"),
("work_issue", "gitea_commit_files"),
})
@@ -523,6 +603,8 @@ ROLE_EXCLUSIVE_TASKS: frozenset[str] = frozenset(
"gitea_release_merger_pr_lease",
"create_branch",
"push_branch",
"bootstrap_author_issue_worktree",
"gitea_bootstrap_author_issue_worktree",
"publish_unpublished_branch",
"create_pr",
"commit_files",
@@ -548,6 +630,7 @@ ISSUE_MUTATION_TOOL_TASKS: dict[str, str] = {
"gitea_set_issue_labels": "set_issue_labels",
"gitea_cleanup_terminal_pr_labels": "cleanup_terminal_pr_labels",
"gitea_create_label": "create_label",
"gitea_bootstrap_author_issue_worktree": "bootstrap_author_issue_worktree",
"gitea_commit_files": "commit_files",
}
@@ -0,0 +1,243 @@
"""Allocator epic / child-only container pre-rank exclusion (#844).
Covers:
* Issue #631-shaped child-only epic is excluded before ranking.
* Implementable child issues remain eligible and can be selected.
* Ordinary issues that merely mention "epic" in title/body are not excluded.
* Excluded containers never receive assignments or workflow leases.
* Structured skip reason ``epic_or_child_only_container`` is reported.
"""
from __future__ import annotations
import os
import tempfile
import unittest
from allocator_service import (
OUTCOME_ASSIGNED,
OUTCOME_PREVIEW,
SKIP_EPIC_OR_CHILD_ONLY_CONTAINER,
WorkCandidate,
allocate_next_work,
classify_epic_or_child_only_container,
)
from control_plane_db import ControlPlaneDB
REMOTE = "prgs"
ORG = "Scaled-Tech-Consulting"
REPO = "Gitea-Tools"
# Minimal body mirroring issue #631 authoritative scope language.
_EPIC_631_BODY = """
## Scope (umbrella)
This epic owns the **product roadmap and linkage** for the Web Console.
Implementation is delivered via child issues only.
## Explicit non-goals
* Do not implement product features in this epic issue itself.
* No product feature implementation is claimed complete solely on this epic.
"""
_CHILD_BODY = """
## Problem
Operators need a workflow-event timeline model for Phase 1.
## Acceptance criteria
- [ ] Timeline model API exists
"""
def _issue(
number: int,
*,
title: str = "",
body: str = "",
labels: tuple[str, ...] = ("status:ready", "type:feature"),
priority: int = 20,
) -> WorkCandidate:
return WorkCandidate(
kind="issue",
number=number,
state="open",
labels=labels,
title=title or f"issue {number}",
body=body,
priority=priority,
)
class ClassifyEpicContainerTest(unittest.TestCase):
def test_631_shaped_body_and_title_is_container(self) -> None:
c = _issue(
631,
title="Epic: MCP Control Plane Web Console",
body=_EPIC_631_BODY,
)
is_c, detail = classify_epic_or_child_only_container(c)
self.assertTrue(is_c)
self.assertIsNotNone(detail)
self.assertIn("body_marker", detail or "")
def test_body_markers_without_epic_title(self) -> None:
c = _issue(
900,
title="Control plane roadmap tracker",
body="Implementation is delivered via child issues only.",
)
is_c, _ = classify_epic_or_child_only_container(c)
self.assertTrue(is_c)
def test_epic_label_alone_is_container(self) -> None:
c = _issue(
901,
title="Roadmap linkage",
body="Track children.",
labels=("status:ready", "type:epic"),
)
is_c, detail = classify_epic_or_child_only_container(c)
self.assertTrue(is_c)
self.assertIn("type:epic", detail or "")
def test_title_epic_prefix_alone_not_container(self) -> None:
"""Title-only 'Epic:' without body scope evidence stays eligible (#844)."""
c = _issue(
902,
title="Epic: something mentioned only in title",
body="Implement a concrete fix for the allocator skip list.",
)
is_c, detail = classify_epic_or_child_only_container(c)
self.assertFalse(is_c)
self.assertIsNone(detail)
def test_incidental_epic_word_not_container(self) -> None:
c = _issue(
903,
title="Document epic handoff conventions",
body=(
"Update the docs so implementable issues that mention an epic "
"remain independently executable."
),
)
is_c, _ = classify_epic_or_child_only_container(c)
self.assertFalse(is_c)
def test_prs_never_classified(self) -> None:
pr = WorkCandidate(
kind="pr",
number=10,
state="open",
title="Epic: fake",
body="Implementation is delivered via child issues only.",
head_sha="a" * 40,
priority=5,
)
is_c, _ = classify_epic_or_child_only_container(pr)
self.assertFalse(is_c)
class AllocateEpicContainerExclusionTest(unittest.TestCase):
def setUp(self) -> None:
self._tmp = tempfile.TemporaryDirectory()
self.addCleanup(self._tmp.cleanup)
self.db = ControlPlaneDB(os.path.join(self._tmp.name, "cp.sqlite3"))
def _alloc(self, candidates, **kwargs):
defaults = dict(
session_id="sess-844",
role="author",
remote=REMOTE,
org=ORG,
repo=REPO,
profile_name="prgs-author",
username="jcwalker3",
claims={},
apply=False,
)
defaults.update(kwargs)
return allocate_next_work(self.db, candidates=candidates, **defaults)
def test_631_shaped_epic_excluded_child_selected(self) -> None:
epic = _issue(
631,
title="Epic: MCP Control Plane Web Console",
body=_EPIC_631_BODY,
)
child = _issue(
637,
title="Web Console: Workflow-event timeline model (Phase 1)",
body=_CHILD_BODY,
)
res = self._alloc([epic, child], apply=False)
self.assertTrue(res["success"], res)
self.assertEqual(res["outcome"], OUTCOME_PREVIEW)
self.assertEqual(res["selected"]["number"], 637)
skipped = {s["number"]: s for s in res["skipped"]}
self.assertIn(631, skipped)
self.assertEqual(
skipped[631]["reason_code"], SKIP_EPIC_OR_CHILD_ONLY_CONTAINER
)
self.assertIn(SKIP_EPIC_OR_CHILD_ONLY_CONTAINER, skipped[631]["reason"])
def test_container_cannot_receive_assignment_or_lease(self) -> None:
epic = _issue(
631,
title="Epic: MCP Control Plane Web Console",
body=_EPIC_631_BODY,
)
res = self._alloc([epic], apply=True)
self.assertTrue(res["success"], res)
# Only container present → no safe work; never assigned_work.
self.assertNotEqual(res["outcome"], OUTCOME_ASSIGNED)
self.assertIsNone(res.get("assignment"))
self.assertIsNone(res.get("selected"))
skipped = {s["number"]: s for s in res["skipped"]}
self.assertEqual(
skipped[631]["reason_code"], SKIP_EPIC_OR_CHILD_ONLY_CONTAINER
)
# No lease row for the epic.
leases = self.db.list_active_leases(
remote=REMOTE, org=ORG, repo=REPO
) if hasattr(self.db, "list_active_leases") else []
# Prefer generic inventory if available.
if not leases and hasattr(self.db, "list_leases"):
leases = self.db.list_leases(remote=REMOTE, org=ORG, repo=REPO)
for lease in leases or []:
work_number = lease.get("work_number") if isinstance(lease, dict) else None
self.assertNotEqual(work_number, 631)
def test_incidental_epic_title_remains_eligible(self) -> None:
ordinary = _issue(
700,
title="Document epic handoff conventions",
body="Write runbook text about epic vs child issues.",
)
res = self._alloc([ordinary], apply=False)
self.assertTrue(res["success"], res)
self.assertEqual(res["selected"]["number"], 700)
self.assertEqual(res["skipped"], [])
def test_apply_selects_child_not_epic(self) -> None:
epic = _issue(
631,
title="Epic: MCP Control Plane Web Console",
body=_EPIC_631_BODY,
)
child = _issue(
637,
title="Web Console: Workflow-event timeline model (Phase 1)",
body=_CHILD_BODY,
)
res = self._alloc([epic, child], apply=True)
self.assertTrue(res["success"], res)
self.assertEqual(res["outcome"], OUTCOME_ASSIGNED)
self.assertEqual(res["selected"]["number"], 637)
self.assertEqual(res["assignment"]["work_number"], 637)
if __name__ == "__main__":
unittest.main()
+601
View File
@@ -0,0 +1,601 @@
"""Regression test suite for native author issue worktree bootstrap (#850)."""
from __future__ import annotations
import json
import os
import shutil
import subprocess
import tempfile
import unittest
from unittest import mock
import author_issue_bootstrap
import task_capability_map
def _concurrent_bootstrap_worker(args: tuple[str, int, str, str, str, str]) -> dict:
repo_dir, issue_num, key, lock_dir, journal_dir, master_sha = args
os.environ["GITEA_BOOTSTRAP_JOURNAL_DIR"] = journal_dir
return author_issue_bootstrap.bootstrap_author_issue_worktree(
issue_number=issue_num,
canonical_repo_root=repo_dir,
expected_base_sha=master_sha,
idempotency_key=key,
lock_dir=lock_dir,
owner_session="session-concurrent-test",
active_identity="jcwalker3",
active_profile="prgs-author",
)
class TestAuthorIssueBootstrap(unittest.TestCase):
"""Test suite covering AC1-AC10 and comment #14959 specification."""
def setUp(self):
self.tmp_dir = tempfile.mkdtemp(prefix="test_bootstrap_")
self.repo_dir = os.path.join(self.tmp_dir, "repo")
os.makedirs(self.repo_dir)
# Initialize synthetic git repo
subprocess.run(["git", "init", "-b", "master"], cwd=self.repo_dir, check=True, capture_output=True)
subprocess.run(["git", "config", "user.name", "Test User"], cwd=self.repo_dir, check=True)
subprocess.run(["git", "config", "user.email", "[email protected]"], cwd=self.repo_dir, check=True)
readme = os.path.join(self.repo_dir, "README.md")
with open(readme, "w", encoding="utf-8") as f:
f.write("# Test Repo\n")
subprocess.run(["git", "add", "README.md"], cwd=self.repo_dir, check=True, capture_output=True)
subprocess.run(["git", "commit", "-m", "initial commit"], cwd=self.repo_dir, check=True, capture_output=True)
rev_res = subprocess.run(["git", "rev-parse", "HEAD"], cwd=self.repo_dir, capture_output=True, text=True, check=True)
self.master_sha = rev_res.stdout.strip()
self.branches_dir = os.path.join(self.repo_dir, "branches")
os.makedirs(self.branches_dir, exist_ok=True)
self.lock_dir = os.path.join(self.tmp_dir, "locks")
os.makedirs(self.lock_dir, exist_ok=True)
self.journal_dir = os.path.join(self.tmp_dir, "journals")
os.makedirs(self.journal_dir, exist_ok=True)
os.environ["GITEA_BOOTSTRAP_JOURNAL_DIR"] = self.journal_dir
def tearDown(self):
os.environ.pop("GITEA_BOOTSTRAP_JOURNAL_DIR", None)
shutil.rmtree(self.tmp_dir, ignore_errors=True)
def test_bootstrap_success_path(self):
"""AC1/AC3/AC8: Successful bootstrap creates branch, worktree, registration, and lock proof."""
key = "test_key_success_1"
res = author_issue_bootstrap.bootstrap_author_issue_worktree(
issue_number=850,
canonical_repo_root=self.repo_dir,
# assignment_id/lease_id omitted: optional unless verified live.
expected_base_sha=self.master_sha,
idempotency_key=key,
remote="prgs",
lock_dir=self.lock_dir,
owner_session="session-test-1234",
)
self.assertTrue(res.get("success"), f"Bootstrap failed: {res}")
self.assertFalse(res.get("replayed"))
self.assertEqual(res.get("issue_number"), 850)
self.assertEqual(res.get("base_sha"), self.master_sha)
self.assertIn("branches/fix-issue-850-native-mcp-bootstrap", res.get("worktree_path"))
# Verify worktree directory exists and is registered
worktree_path = res["worktree_path"]
self.assertTrue(os.path.isdir(worktree_path))
wt_list = subprocess.run(["git", "-C", self.repo_dir, "worktree", "list"], capture_output=True, text=True, check=True)
self.assertIn(worktree_path, wt_list.stdout)
# Verify phase journal written
journal = author_issue_bootstrap.load_phase_journal(key, journal_dir=self.lock_dir)
self.assertIsNotNone(journal)
self.assertTrue(journal.get("completed"))
self.assertEqual(journal.get("current_phase"), author_issue_bootstrap.PHASE_7_TRANSITION_COMPLETED)
def test_idempotent_replay(self):
"""Item 2: Replaying with identical key returns cached transition without duplicate creation."""
key = "test_key_idempotent_1"
res1 = author_issue_bootstrap.bootstrap_author_issue_worktree(
issue_number=850,
canonical_repo_root=self.repo_dir,
idempotency_key=key,
lock_dir=self.lock_dir,
owner_session="session-test-1234",
)
self.assertTrue(res1["success"], f"res1 failed: {res1}")
self.assertFalse(res1.get("replayed"))
# Second call
res2 = author_issue_bootstrap.bootstrap_author_issue_worktree(
issue_number=850,
canonical_repo_root=self.repo_dir,
idempotency_key=key,
lock_dir=self.lock_dir,
owner_session="session-test-1234",
)
self.assertTrue(res2["success"], f"res2 failed: {res2}")
self.assertTrue(res2.get("replayed"))
self.assertEqual(res1["worktree_path"], res2["worktree_path"])
def test_stale_concurrency_pin_refusal(self):
"""Item 3: Mismatched expected base SHA fails closed without silent rebasing."""
stale_sha = "0000000000000000000000000000000000000000"
res = author_issue_bootstrap.bootstrap_author_issue_worktree(
issue_number=850,
canonical_repo_root=self.repo_dir,
expected_base_sha=stale_sha,
lock_dir=self.lock_dir,
owner_session="session-test-1234",
)
self.assertFalse(res["success"])
self.assertEqual(res.get("reason_code"), "stale_concurrency_pin")
self.assertIn("exact_next_action", res)
def test_path_outside_branches_root_refusal(self):
"""Item 6: Worktree path outside branches/ root is refused."""
outside_path = os.path.join(self.tmp_dir, "outside_worktree")
res = author_issue_bootstrap.bootstrap_author_issue_worktree(
issue_number=850,
canonical_repo_root=self.repo_dir,
worktree_path=outside_path,
lock_dir=self.lock_dir,
owner_session="session-test-1234",
)
self.assertFalse(res["success"])
self.assertEqual(res.get("reason_code"), "path_outside_canonical_branches_root")
def test_preexisting_dirty_worktree_preservation(self):
"""Item 6: Preexisting dirty worktree fails closed and is NOT modified or cleaned."""
branch = "fix/issue-850-dirty-test"
wt_path = os.path.join(self.branches_dir, "fix-issue-850-dirty-test")
subprocess.run(["git", "-C", self.repo_dir, "worktree", "add", "-b", branch, wt_path], check=True, capture_output=True)
# Create dirty untracked file
dirty_file = os.path.join(wt_path, "dirty.txt")
with open(dirty_file, "w") as f:
f.write("dirty edits\n")
res = author_issue_bootstrap.bootstrap_author_issue_worktree(
issue_number=850,
canonical_repo_root=self.repo_dir,
branch_name=branch,
worktree_path=wt_path,
lock_dir=self.lock_dir,
owner_session="session-test-1234",
)
self.assertFalse(res["success"])
self.assertEqual(res.get("reason_code"), "preexisting_dirty_worktree")
# Prove dirty file is preserved byte-for-byte
self.assertTrue(os.path.exists(dirty_file))
with open(dirty_file, "r") as f:
self.assertEqual(f.read(), "dirty edits\n")
def test_compensating_recovery_on_failed_phase(self):
"""AC4/Item 4: Failure during transition rolls back ONLY newly created artifacts."""
key = "test_key_recovery_1"
# Simulate partial progress in journal
journal = {
"idempotency_key": key,
"issue_number": 850,
"branch_name": "fix/issue-850-recovery-test",
"worktree_path": os.path.join(self.branches_dir, "fix-issue-850-recovery-test"),
"artifacts_created": {
"branch_created": True,
"worktree_dir_created": True,
"worktree_registered": True,
"lock_created": False,
},
"failure_reason": "simulated lock failure",
"current_phase": author_issue_bootstrap.PHASE_5_REGISTRATION_VERIFIED,
"completed": False,
}
# Create the branch and worktree manually to simulate partial state
subprocess.run(["git", "-C", self.repo_dir, "branch", journal["branch_name"]], check=True, capture_output=True)
subprocess.run(["git", "-C", self.repo_dir, "worktree", "add", journal["worktree_path"], journal["branch_name"]], check=True, capture_output=True)
# Run compensating recovery
rec = author_issue_bootstrap.run_compensating_recovery(journal, self.repo_dir)
self.assertTrue(rec["executed"])
self.assertIn(f"worktree_path:{journal['worktree_path']}", rec["rolled_back"])
self.assertIn(f"branch:{journal['branch_name']}", rec["rolled_back"])
# Prove worktree directory and branch were rolled back
self.assertFalse(os.path.exists(journal["worktree_path"]))
branch_check = subprocess.run(["git", "-C", self.repo_dir, "rev-parse", "--verify", journal["branch_name"]], capture_output=True, text=True, check=False)
self.assertNotEqual(branch_check.returncode, 0)
def test_cross_process_concurrency(self):
"""Review #525 Finding 1: Genuine cross-process concurrency locking prevents corruption."""
import concurrent.futures
key = "test_concurrent_key_850"
args = (self.repo_dir, 850, key, self.lock_dir, self.journal_dir, self.master_sha)
with concurrent.futures.ProcessPoolExecutor(max_workers=2) as executor:
fut1 = executor.submit(_concurrent_bootstrap_worker, args)
fut2 = executor.submit(_concurrent_bootstrap_worker, args)
res1 = fut1.result(timeout=10)
res2 = fut2.result(timeout=10)
self.assertTrue(res1["success"], f"res1 failed: {res1}")
self.assertTrue(res2["success"], f"res2 failed: {res2}")
# One process performs creation, the other process receives idempotent replay
replayed_count = sum(1 for r in (res1, res2) if r.get("replayed"))
created_count = sum(1 for r in (res1, res2) if not r.get("replayed"))
self.assertEqual(replayed_count, 1)
self.assertEqual(created_count, 1)
self.assertEqual(res1["worktree_path"], res2["worktree_path"])
def test_interrupted_replay_preserves_artifacts_created_provenance(self):
"""Review #525 Finding 2: Replaying incomplete journal preserves creation provenance monotonically."""
key = "test_key_interrupted_replay_1"
branch = "fix/issue-850-interrupted-replay"
wt_path = os.path.join(self.branches_dir, "fix-issue-850-interrupted-replay")
# Simulate Phase 2/3 completion where branch and worktree directory were created by this transition
journal = {
"idempotency_key": key,
"issue_number": 850,
"branch_name": branch,
"worktree_path": wt_path,
"active_identity": "jcwalker3",
"active_profile": "prgs-author",
"remote": "prgs",
"org": "Scaled-Tech-Consulting",
"repo": "Gitea-Tools",
"phases": {
author_issue_bootstrap.PHASE_1_REQUEST_ACCEPTED: {"status": "completed"},
author_issue_bootstrap.PHASE_2_BRANCH_CONFIRMED: {"status": "completed", "created": True},
},
"artifacts_created": {
"branch_created": True,
"worktree_dir_created": True,
"worktree_registered": True,
"lock_created": False,
},
"current_phase": author_issue_bootstrap.PHASE_3_PATH_RESERVED,
"completed": False,
}
# Pre-create the branch and worktree on disk to simulate partial state after crash
subprocess.run(["git", "-C", self.repo_dir, "branch", branch, self.master_sha], check=True, capture_output=True)
subprocess.run(["git", "-C", self.repo_dir, "worktree", "add", wt_path, branch], check=True, capture_output=True)
author_issue_bootstrap.save_phase_journal(journal, journal_dir=self.lock_dir)
# Now resume/replay the transition but simulate lock binding failure during Phase 6
with mock.patch("issue_lock_store.bind_session_lock", side_effect=RuntimeError("Lock failure test")):
res = author_issue_bootstrap.bootstrap_author_issue_worktree(
issue_number=850,
canonical_repo_root=self.repo_dir,
branch_name=branch,
worktree_path=wt_path,
idempotency_key=key,
lock_dir=self.lock_dir,
owner_session="session-test-1234",
)
self.assertFalse(res["success"])
self.assertEqual(res.get("reason_code"), "issue_lock_acquisition_failed")
# Verify that compensating recovery correctly deleted transition-created branch & worktree
# because creation provenance was preserved across replay (NOT downgraded to False!)
self.assertFalse(os.path.exists(wt_path))
branch_check = subprocess.run(["git", "-C", self.repo_dir, "rev-parse", "--verify", branch], capture_output=True, text=True, check=False)
self.assertNotEqual(branch_check.returncode, 0)
def test_transition_created_only_compensation(self):
"""Review #525 Finding 4: Preexisting branch is NOT deleted by compensation when only worktree was transition-created."""
key = "test_key_preexisting_branch_compensation"
preexisting_branch = "fix/issue-850-preexisting"
wt_path = os.path.join(self.branches_dir, "fix-issue-850-preexisting")
# Create branch BEFORE bootstrap (preexisting branch)
subprocess.run(["git", "-C", self.repo_dir, "branch", preexisting_branch, self.master_sha], check=True, capture_output=True)
# Call bootstrap with simulated failure during Phase 6 (lock binding)
with mock.patch("issue_lock_store.bind_session_lock", side_effect=RuntimeError("Simulated lock failure")):
res = author_issue_bootstrap.bootstrap_author_issue_worktree(
issue_number=850,
canonical_repo_root=self.repo_dir,
branch_name=preexisting_branch,
worktree_path=wt_path,
idempotency_key=key,
lock_dir=self.lock_dir,
owner_session="session-test-1234",
)
self.assertFalse(res["success"])
# Worktree dir was created by transition -> removed by compensation
self.assertFalse(os.path.exists(wt_path))
# Preexisting branch was NOT created by transition -> MUST BE PRESERVED!
branch_check = subprocess.run(["git", "-C", self.repo_dir, "rev-parse", "--verify", preexisting_branch], capture_output=True, text=True, check=False)
self.assertEqual(branch_check.returncode, 0, "Preexisting branch was deleted by mistake!")
def test_incompatible_idempotency_replay_refusal(self):
"""Review #525 Finding 4: Replaying key with incompatible parameters returns refusal."""
key = "test_key_incompatible_replay"
res1 = author_issue_bootstrap.bootstrap_author_issue_worktree(
issue_number=850,
canonical_repo_root=self.repo_dir,
branch_name="fix/issue-850-param-a",
idempotency_key=key,
lock_dir=self.lock_dir,
owner_session="session-test-1234",
)
self.assertTrue(res1["success"])
# Second call with different branch_name
res2 = author_issue_bootstrap.bootstrap_author_issue_worktree(
issue_number=850,
canonical_repo_root=self.repo_dir,
branch_name="fix/issue-850-param-b",
idempotency_key=key,
lock_dir=self.lock_dir,
owner_session="session-test-1234",
)
self.assertFalse(res2["success"])
self.assertEqual(res2.get("reason_code"), "incompatible_idempotency_replay")
def test_exact_next_action_satisfiable_via_mcp(self):
"""Review #525 Finding 4: exact_next_action provides satisfiable MCP actions, not shell commands."""
key = "test_key_next_action_mcp"
stale_sha = "0000000000000000000000000000000000000000"
res = author_issue_bootstrap.bootstrap_author_issue_worktree(
issue_number=850,
canonical_repo_root=self.repo_dir,
expected_base_sha=stale_sha,
lock_dir=self.lock_dir,
owner_session="session-test-1234",
)
next_action = res.get("exact_next_action", "")
self.assertNotIn("scripts/worktree-start", next_action)
self.assertNotIn("git worktree add", next_action)
self.assertNotIn("bash", next_action.lower())
def test_missing_owner_session_refusal(self):
"""Finding D: Missing owner_session context fails closed with typed refusal and zero mutation."""
res = author_issue_bootstrap.bootstrap_author_issue_worktree(
issue_number=850,
canonical_repo_root=self.repo_dir,
owner_session=None,
lock_dir=self.lock_dir,
)
self.assertFalse(res["success"])
self.assertEqual(res.get("reason_code"), "missing_owner_session")
self.assertIn("exact_next_action", res)
def test_symlink_lock_file_refusal(self):
"""Finding C: BootstrapTransitionLock refuses to follow symlinks."""
key = "test_symlink_lock_key"
safe_key = "".join(c if c.isalnum() or c in ("-", "_", ".") else "_" for c in key)
lock_path = os.path.join(self.lock_dir, f"{safe_key}.lock")
target_file = os.path.join(self.tmp_dir, "fake_target")
with open(target_file, "w") as f:
f.write("target")
os.symlink(target_file, lock_path)
with self.assertRaises(RuntimeError) as ctx:
with author_issue_bootstrap.BootstrapTransitionLock(key, journal_dir=self.lock_dir):
pass
self.assertIn("symlink", str(ctx.exception).lower())
def test_lock_directory_escape_refusal(self):
"""Finding C: BootstrapTransitionLock refuses keys that escape lock directory."""
with mock.patch("os.path.abspath", return_value="/tmp/outside/evil_key.lock"):
with self.assertRaises(RuntimeError) as ctx:
author_issue_bootstrap.BootstrapTransitionLock("key", journal_dir=self.lock_dir)
self.assertIn("escapes", str(ctx.exception).lower())
def test_missing_active_identity_refusal(self):
"""F-5: Missing active_identity parameter fails closed."""
res = author_issue_bootstrap.bootstrap_author_issue_worktree(
issue_number=850,
canonical_repo_root=self.repo_dir,
owner_session="session-test-1234",
active_identity=None,
active_profile="prgs-author",
lock_dir=self.lock_dir,
)
self.assertFalse(res["success"])
self.assertEqual(res.get("reason_code"), "missing_active_identity")
def test_missing_active_profile_refusal(self):
"""F-5: Missing active_profile parameter fails closed."""
res = author_issue_bootstrap.bootstrap_author_issue_worktree(
issue_number=850,
canonical_repo_root=self.repo_dir,
owner_session="session-test-1234",
active_identity="jcwalker3",
active_profile=None,
lock_dir=self.lock_dir,
)
self.assertFalse(res["success"])
self.assertEqual(res.get("reason_code"), "missing_active_profile")
def test_dirty_worktree_preserved_during_recovery(self):
"""F-4: Compensating recovery does not delete dirty worktree."""
branch = "fix/issue-850-rec-dirty"
wt_path = os.path.join(self.branches_dir, "fix-issue-850-rec-dirty")
subprocess.run(["git", "-C", self.repo_dir, "worktree", "add", "-b", branch, wt_path], check=True, capture_output=True)
dirty_file = os.path.join(wt_path, "dirty.txt")
with open(dirty_file, "w") as f:
f.write("uncommitted work")
journal = {
"idempotency_key": "test_dirty_rec",
"issue_number": 850,
"branch_name": branch,
"worktree_path": wt_path,
"artifacts_created": {
"worktree_dir_created": True,
"worktree_registered": True,
},
"failure_reason": "test dirty recovery",
}
rec = author_issue_bootstrap.run_compensating_recovery(journal, self.repo_dir, journal_dir=self.lock_dir)
self.assertTrue(os.path.exists(wt_path))
self.assertIn(f"worktree_path_preserved_dirty:{wt_path}", rec["rolled_back"])
def test_branch_with_commits_preserved_during_recovery(self):
"""F-4: Compensating recovery does not delete branch with author commits."""
branch = "fix/issue-850-rec-commits"
subprocess.run(["git", "-C", self.repo_dir, "branch", branch, self.master_sha], check=True, capture_output=True)
# Add a commit on the branch
wt_path = os.path.join(self.branches_dir, "fix-issue-850-rec-commits")
subprocess.run(["git", "-C", self.repo_dir, "worktree", "add", wt_path, branch], check=True, capture_output=True)
cfile = os.path.join(wt_path, "commit.txt")
with open(cfile, "w") as f:
f.write("author commit")
subprocess.run(["git", "-C", wt_path, "add", "commit.txt"], check=True, capture_output=True)
subprocess.run(["git", "-C", wt_path, "commit", "-m", "author commit"], check=True, capture_output=True)
subprocess.run(["git", "-C", self.repo_dir, "worktree", "remove", "--force", wt_path], check=True, capture_output=True)
journal = {
"idempotency_key": "test_commits_rec",
"issue_number": 850,
"branch_name": branch,
"resolved_base_sha": self.master_sha,
"artifacts_created": {
"branch_created": True,
},
"failure_reason": "test commit branch recovery",
}
rec = author_issue_bootstrap.run_compensating_recovery(journal, self.repo_dir, journal_dir=self.lock_dir)
branch_check = subprocess.run(["git", "-C", self.repo_dir, "rev-parse", "--verify", branch], capture_output=True, text=True, check=False)
self.assertEqual(branch_check.returncode, 0, "Branch with commits was deleted!")
self.assertIn(f"branch_preserved_commits:{branch}", rec["rolled_back"])
def test_task_capability_map_integration(self):
"""Verify task_capability_map has bootstrap_author_issue_worktree configured correctly."""
self.assertEqual(task_capability_map.required_role("bootstrap_author_issue_worktree"), "author")
self.assertEqual(task_capability_map.required_permission("bootstrap_author_issue_worktree"), "gitea.branch.create")
self.assertTrue(task_capability_map.preflight_task_matches("work_issue", "bootstrap_author_issue_worktree"))
self.assertTrue(task_capability_map.preflight_task_matches("bootstrap_author_issue_worktree", "lock_issue"))
def test_unverified_assignment_lease_ids_fail_closed(self):
"""Review #531 Finding 4: fabricated assignment/lease IDs are refused."""
res = author_issue_bootstrap.bootstrap_author_issue_worktree(
issue_number=850,
canonical_repo_root=self.repo_dir,
assignment_id="asn-fabricated",
lease_id="lease-fabricated",
expected_base_sha=self.master_sha,
lock_dir=self.lock_dir,
owner_session="session-test-1234",
)
self.assertFalse(res["success"])
self.assertIn(
res.get("reason_code"),
{
"unknown_lease_id",
"assignment_lease_lookup_failed",
"incomplete_assignment_lease_ids",
},
)
def test_partial_assignment_lease_ids_fail_closed(self):
"""Review #531 Finding 4: one of assignment_id/lease_id alone is incomplete."""
res = author_issue_bootstrap.bootstrap_author_issue_worktree(
issue_number=850,
canonical_repo_root=self.repo_dir,
assignment_id="asn-only",
lock_dir=self.lock_dir,
owner_session="session-test-1234",
)
self.assertFalse(res["success"])
self.assertEqual(res.get("reason_code"), "incomplete_assignment_lease_ids")
def test_stale_diverged_branch_is_not_accepted_via_merge_base(self):
"""Review #531 Finding 3: any common ancestor is not enough; require master ⊆ branch."""
branch = "fix/issue-850-stale-divergent"
# Create branch from current master, then advance master so branch lacks tip.
subprocess.run(
["git", "-C", self.repo_dir, "branch", branch, self.master_sha],
check=True,
capture_output=True,
)
# Make a new commit on master (orphan path so branch does not contain it).
marker = os.path.join(self.repo_dir, "master-advance.txt")
with open(marker, "w") as f:
f.write("advance master\n")
subprocess.run(["git", "-C", self.repo_dir, "add", "master-advance.txt"], check=True, capture_output=True)
subprocess.run(
["git", "-C", self.repo_dir, "commit", "-m", "advance master past branch"],
check=True,
capture_output=True,
)
new_master = subprocess.run(
["git", "-C", self.repo_dir, "rev-parse", "HEAD"],
capture_output=True,
text=True,
check=True,
).stdout.strip()
res = author_issue_bootstrap.bootstrap_author_issue_worktree(
issue_number=850,
canonical_repo_root=self.repo_dir,
branch_name=branch,
expected_base_sha=new_master,
lock_dir=self.lock_dir,
owner_session="session-test-1234",
)
self.assertFalse(res["success"])
self.assertEqual(res.get("reason_code"), "incompatible_existing_branch")
def test_compensating_recovery_attempts_lease_release(self):
"""Review #531 Finding 5: recovery invokes lease release when lease_id is present."""
from unittest import mock
journal = {
"idempotency_key": "test_lease_rec",
"issue_number": 850,
"owner_session": "session-test-1234",
"lease_id": "lease-abc",
"branch_name": "fix/issue-850-lease-rec",
"artifacts_created": {"lock_created": True},
"failure_reason": "simulated",
"completed": False,
}
with mock.patch.object(
author_issue_bootstrap.lease_lifecycle,
"release_lease",
return_value={"success": True},
) as rel, mock.patch.object(
author_issue_bootstrap.control_plane_db,
"ControlPlaneDB",
return_value=mock.Mock(),
):
rec = author_issue_bootstrap.run_compensating_recovery(
journal, self.repo_dir, journal_dir=self.lock_dir
)
self.assertTrue(rec["executed"])
rel.assert_called_once()
self.assertIn("lease:lease-abc", rec["rolled_back"])
class TestCanonicalRootNoStringSplit(unittest.TestCase):
def test_fallback_uses_commonpath_not_substring_split(self):
"""Review #531 Finding 2: no norm.split('/branches/') fallback."""
import inspect
import author_mutation_worktree as amw
src = inspect.getsource(amw.resolve_canonical_repo_root)
self.assertNotIn('split("/branches/")', src)
self.assertNotIn("split('/branches/')", src)
# Fallback recovers repo root from a nested branches worktree path.
with tempfile.TemporaryDirectory() as tmp:
repo = os.path.join(tmp, "repo")
wt = os.path.join(repo, "branches", "fix-issue-850-x")
os.makedirs(wt)
# git unavailable path: pass missing workspace so fallback is used.
root = amw.resolve_canonical_repo_root("/missing/path", wt)
self.assertEqual(root, os.path.realpath(repo))
if __name__ == "__main__":
unittest.main()
+22
View File
@@ -28,6 +28,21 @@ class TestPathUnderBranches(unittest.TestCase):
amw.is_path_under_branches("/repo/other-checkout", self.ROOT)
)
def test_unrelated_branches_dir_fails(self):
self.assertFalse(
amw.is_path_under_branches("/tmp/branches/evil", self.ROOT)
)
def test_prefix_confusion_fails(self):
self.assertFalse(
amw.is_path_under_branches(f"{self.ROOT}/branches-other/foo", self.ROOT)
)
def test_traversal_fails(self):
self.assertFalse(
amw.is_path_under_branches(f"{self.ROOT}/branches/../evil", self.ROOT)
)
class TestAssessAuthorMutationWorktree(unittest.TestCase):
ROOT = "/repo/Gitea-Tools"
@@ -71,6 +86,13 @@ class TestAssessAuthorMutationWorktree(unittest.TestCase):
self.assertTrue(result["proven"])
self.assertFalse(result["block"])
def test_path_shaped_branches_ancestry_without_isdir(self):
"""Review #551: commonpath recovery must not require on-disk isdir."""
fake_wt = "/repo/Gitea-Tools/branches/issue-274"
root = amw.resolve_canonical_repo_root(fake_wt, fake_wt)
self.assertEqual(root, "/repo/Gitea-Tools")
self.assertNotIn('split("/branches/")', open(amw.__file__).read())
class TestPreflightIntegration(unittest.TestCase):
def test_verify_preflight_blocks_control_checkout_with_test_porcelain(self):
+724
View File
@@ -1266,6 +1266,730 @@ class TestSecondRemediationIntegration(unittest.TestCase):
self.assertIn("delete_acknowledged", delete_actions[0])
self.assertTrue(delete_actions[0].get("verified_absent"))
def test_issue_851_worktree_removed_when_remote_blocked_only_by_worktree_binding(self):
"""#851: remote blocked by worktree_binding must not skip safe worktree removal.
Lifecycle: remove clean owned worktree reassess ownership delete
remote only if independently safe. Unrelated entries stay untouched.
"""
from mcp_server import gitea_reconcile_merged_cleanups
target_branch = "fix/issue-844-exclude-epic-containers"
foreign_branch = "fix/issue-999-unrelated-active"
worktree_path = "/tmp/branches/fix-issue-844-exclude-epic-containers"
ownership_calls = []
remove_calls = []
delete_api_calls = []
def fake_collect(**kwargs):
ownership_calls.append(dict(kwargs))
# Ownership is reassessed *after* independent worktree removal (#851).
# Target worktree is already gone → no worktree_binding remains.
# Foreign branch keeps an active author lease → remote delete blocked.
if kwargs.get("branch") == foreign_branch:
# Match session-bound org/repo + host used by the tool resolve path.
return {
"records": [
{
"category": guard.OWNERSHIP_CATEGORY_AUTHOR_LEASE,
"status": "active",
"remote": kwargs.get("remote") or "prgs",
"host": kwargs.get("host") or "gitea.example.com",
"org": kwargs.get("org") or "Scaled-Tech-Consulting",
"repo": kwargs.get("repo") or "Gitea-Tools",
"branch": foreign_branch,
"reclaim_allowed": False,
}
],
"inventory_error": False,
}
return {"records": [], "inventory_error": False}
def fake_remove(project_root, branch, worktree_path=None):
remove_calls.append(
{"branch": branch, "worktree_path": worktree_path}
)
return {
"success": True,
"performed": True,
"message": f"removed worktree {worktree_path}",
"worktree_path": worktree_path,
}
def fake_probe(h, o, r, auth, br):
return guard.classify_branch_readback_http_status(
404, not_found_scope=guard.NOT_FOUND_SCOPE_BRANCH
)
def fake_api(method, url, auth, **kwargs):
if method == "DELETE":
delete_api_calls.append(url)
return {}
report = {
"entries": [
{
"pr_number": 848,
"head_branch": target_branch,
"remote_branch": {"safe_to_delete_remote": True},
"local_worktree": {
"safe_to_remove_worktree": True,
"worktree_path": worktree_path,
},
},
{
"pr_number": 999,
"head_branch": foreign_branch,
"remote_branch": {"safe_to_delete_remote": True},
"local_worktree": {
"safe_to_remove_worktree": False,
"worktree_path": None,
},
},
],
"reviewer_scratch_entries": [],
}
patch(
"mcp_server.get_profile",
return_value={
"profile_name": "prgs-reconciler",
"role": "reconciler",
"allowed_operations": [
"gitea.read",
"gitea.branch.delete",
"gitea.pr.close",
],
"forbidden_operations": [],
},
).start()
patch("mcp_server.api_get_all", return_value=[]).start()
patch(
"mcp_server.merged_cleanup_reconcile.build_reconciliation_report",
return_value=report,
).start()
patch(
"mcp_server.merged_cleanup_reconcile.discover_reviewer_scratch_worktrees",
return_value=[],
).start()
patch(
"mcp_server.audit_reconciliation_mode.check_cleanup_execution_allowed",
return_value=(True, []),
).start()
patch("mcp_server.verify_preflight_purity", return_value=None).start()
patch(
"mcp_server._collect_branch_ownership_records",
side_effect=fake_collect,
).start()
patch("mcp_server._probe_remote_branch", side_effect=fake_probe).start()
patch(
"mcp_server.merged_cleanup_reconcile.remove_local_worktree",
side_effect=fake_remove,
).start()
self.mock_api.side_effect = fake_api
res = gitea_reconcile_merged_cleanups(
dry_run=False,
execute_confirmed=True,
remote="prgs",
)
self.assertTrue(res.get("performed") or res.get("executed"))
actions = res.get("actions") or []
remove_actions = [
a for a in actions if a.get("action") == "remove_local_worktree"
]
self.assertEqual(len(remove_actions), 1, actions)
self.assertTrue(remove_actions[0].get("success"))
self.assertEqual(remove_calls[0]["branch"], target_branch)
self.assertEqual(remove_calls[0]["worktree_path"], worktree_path)
# Target remote delete succeeds after worktree removal + reassessment.
target_deletes = [
a
for a in actions
if a.get("action") == "delete_remote_branch"
and a.get("branch") == target_branch
]
self.assertEqual(len(target_deletes), 1, actions)
self.assertTrue(target_deletes[0].get("success"))
self.assertTrue(target_deletes[0].get("after_worktree_removal"))
self.assertTrue(target_deletes[0].get("ownership_reassessed"))
self.assertTrue(target_deletes[0].get("verified_absent"))
# Foreign branch remains protected (author lease) and is not deleted.
foreign_deletes = [
a
for a in actions
if a.get("action") == "delete_remote_branch"
and a.get("branch") == foreign_branch
]
self.assertEqual(len(foreign_deletes), 1, actions)
self.assertFalse(foreign_deletes[0].get("success"))
self.assertEqual(
foreign_deletes[0].get("blocker_kind"), "active_branch_ownership"
)
self.assertIn(
guard.OWNERSHIP_CATEGORY_AUTHOR_LEASE,
foreign_deletes[0].get("blocking_categories") or [],
)
# Only the target branch should hit the DELETE API.
self.assertEqual(len(delete_api_calls), 1)
# Ownership collected for target (post-removal) and foreign; worktree
# removal happened before target remote delete in the action log.
target_idx = next(
i
for i, a in enumerate(actions)
if a.get("action") == "remove_local_worktree"
)
delete_idx = next(
i
for i, a in enumerate(actions)
if a.get("action") == "delete_remote_branch"
and a.get("branch") == target_branch
and a.get("success")
)
self.assertLess(target_idx, delete_idx)
def test_issue_851_dirty_worktree_not_removed_and_remote_stays_protected(self):
"""#851: dirty/foreign worktrees remain protected; no unsafe cleanup."""
from mcp_server import gitea_reconcile_merged_cleanups
branch = "fix/issue-851-dirty"
remove_calls = []
def fake_collect(**kwargs):
return {
"records": [
{
"category": guard.OWNERSHIP_CATEGORY_WORKTREE_BINDING,
"status": "active",
"remote": kwargs.get("remote") or "prgs",
"host": kwargs.get("host") or "gitea.example.com",
"org": kwargs.get("org") or "Scaled-Tech-Consulting",
"repo": kwargs.get("repo") or "Gitea-Tools",
"branch": branch,
"reclaim_allowed": False,
}
],
"inventory_error": False,
}
report = {
"entries": [
{
"pr_number": 851,
"head_branch": branch,
"remote_branch": {"safe_to_delete_remote": True},
"local_worktree": {
"safe_to_remove_worktree": False,
"worktree_path": "/tmp/dirty-wt",
},
}
],
"reviewer_scratch_entries": [],
}
patch(
"mcp_server.get_profile",
return_value={
"profile_name": "prgs-reconciler",
"role": "reconciler",
"allowed_operations": [
"gitea.read",
"gitea.branch.delete",
],
"forbidden_operations": [],
},
).start()
patch("mcp_server.api_get_all", return_value=[]).start()
patch(
"mcp_server.merged_cleanup_reconcile.build_reconciliation_report",
return_value=report,
).start()
patch(
"mcp_server.merged_cleanup_reconcile.discover_reviewer_scratch_worktrees",
return_value=[],
).start()
patch(
"mcp_server.audit_reconciliation_mode.check_cleanup_execution_allowed",
return_value=(True, []),
).start()
patch("mcp_server.verify_preflight_purity", return_value=None).start()
patch(
"mcp_server._collect_branch_ownership_records",
side_effect=fake_collect,
).start()
patch(
"mcp_server.merged_cleanup_reconcile.remove_local_worktree",
side_effect=lambda *a, **k: remove_calls.append(k) or {
"success": True,
"performed": True,
},
).start()
self.mock_api.side_effect = lambda *a, **k: {}
res = gitea_reconcile_merged_cleanups(
dry_run=False,
execute_confirmed=True,
remote="prgs",
)
actions = res.get("actions") or []
self.assertEqual(remove_calls, [])
self.assertFalse(
any(a.get("action") == "remove_local_worktree" for a in actions)
)
deletes = [
a for a in actions if a.get("action") == "delete_remote_branch"
]
self.assertEqual(len(deletes), 1)
self.assertFalse(deletes[0].get("success"))
self.assertEqual(deletes[0].get("blocker_kind"), "active_branch_ownership")
self.assertIn(
guard.OWNERSHIP_CATEGORY_WORKTREE_BINDING,
deletes[0].get("blocking_categories") or [],
)
def test_issue_851_idempotent_resume_when_worktree_already_absent(self):
"""#851: partial failures remain resumable and idempotent."""
from mcp_server import gitea_reconcile_merged_cleanups
branch = "fix/issue-851-resume"
ownership_calls = []
def fake_collect(**kwargs):
ownership_calls.append(kwargs)
return {"records": [], "inventory_error": False}
def fake_remove(project_root, branch, worktree_path=None):
return {
"success": False,
"performed": False,
"message": f"worktree not found: {worktree_path}",
}
def fake_probe(h, o, r, auth, br):
return guard.classify_branch_readback_http_status(
404, not_found_scope=guard.NOT_FOUND_SCOPE_BRANCH
)
report = {
"entries": [
{
"pr_number": 851,
"head_branch": branch,
"remote_branch": {"safe_to_delete_remote": True},
"local_worktree": {
"safe_to_remove_worktree": True,
"worktree_path": "/tmp/already-gone",
},
}
],
"reviewer_scratch_entries": [],
}
patch(
"mcp_server.get_profile",
return_value={
"profile_name": "prgs-reconciler",
"role": "reconciler",
"allowed_operations": [
"gitea.read",
"gitea.branch.delete",
],
"forbidden_operations": [],
},
).start()
patch("mcp_server.api_get_all", return_value=[]).start()
patch(
"mcp_server.merged_cleanup_reconcile.build_reconciliation_report",
return_value=report,
).start()
patch(
"mcp_server.merged_cleanup_reconcile.discover_reviewer_scratch_worktrees",
return_value=[],
).start()
patch(
"mcp_server.audit_reconciliation_mode.check_cleanup_execution_allowed",
return_value=(True, []),
).start()
patch("mcp_server.verify_preflight_purity", return_value=None).start()
patch(
"mcp_server._collect_branch_ownership_records",
side_effect=fake_collect,
).start()
patch("mcp_server._probe_remote_branch", side_effect=fake_probe).start()
patch(
"mcp_server.merged_cleanup_reconcile.remove_local_worktree",
side_effect=fake_remove,
).start()
self.mock_api.side_effect = lambda *a, **k: {}
res = gitea_reconcile_merged_cleanups(
dry_run=False,
execute_confirmed=True,
remote="prgs",
)
actions = res.get("actions") or []
removes = [a for a in actions if a.get("action") == "remove_local_worktree"]
deletes = [a for a in actions if a.get("action") == "delete_remote_branch"]
self.assertEqual(len(removes), 1)
self.assertFalse(removes[0].get("success"))
self.assertEqual(len(deletes), 1)
self.assertTrue(deletes[0].get("success"))
self.assertTrue(deletes[0].get("after_worktree_removal"))
self.assertTrue(ownership_calls)
class TestIssue855ExactPrSelector(unittest.TestCase):
"""#855: exact pr_number pin for reconcile_merged_cleanups (#851 lifecycle)."""
def setUp(self):
self._remotes = patch.dict(
mcp_server.REMOTES,
{
"prgs": {
"host": "gitea.example.com",
"org": "Scaled-Tech-Consulting",
"repo": "Gitea-Tools",
}
},
)
self._remotes.start()
patch("gitea_audit.audit_enabled", return_value=False).start()
self.mock_api = patch("mcp_server.api_request").start()
self.mock_all = patch("mcp_server.api_get_all", return_value=[]).start()
patch("mcp_server.get_auth_header", return_value=FAKE_AUTH).start()
patch(
"mcp_server.merged_cleanup_reconcile.is_head_ancestor_of_ref",
return_value=True,
).start()
patch(
"mcp_server.get_profile",
return_value=dict(RECONCILER_WITH_DELETE),
).start()
patch(
"mcp_server._profile_operation_gate",
return_value=[],
).start()
patch(
"mcp_server._collect_branch_ownership_records",
return_value={"records": [], "inventory_error": False},
).start()
patch(
"mcp_server.merged_cleanup_reconcile.discover_reviewer_scratch_worktrees",
return_value=[],
).start()
patch("mcp_server.verify_preflight_purity", return_value=None).start()
patch(
"mcp_server.audit_reconciliation_mode.check_cleanup_execution_allowed",
return_value=(True, []),
).start()
def tearDown(self):
patch.stopall()
def _merged_pr(self, number, branch, sha="c" * 40):
return {
"number": number,
"title": f"PR {number}",
"body": f"Closes #{number - 4}",
"merged": True,
"merged_at": "2026-07-23T12:00:00Z",
"merge_commit_sha": "f" * 40,
"state": "closed",
"head": {"ref": branch, "sha": sha},
"base": {"ref": "master"},
}
def test_exact_pr_848_ignores_newer_852_in_batch_queue(self):
"""pr_number=848 selects only #848 even when #852 is newer/first."""
from mcp_server import gitea_reconcile_merged_cleanups
pr_848 = self._merged_pr(
848, "fix/issue-844-exclude-epic-containers", sha="c3f282ba" + "0" * 32
)
# Closed list would rank #852 first in batch mode; exact pin must ignore it.
closed_batch = [
self._merged_pr(852, "fix/issue-851-cleanup-worktree-before-remote-delete"),
pr_848,
self._merged_pr(849, "fix/issue-849-other"),
self._merged_pr(846, "fix/issue-846-other"),
self._merged_pr(845, "fix/issue-845-other"),
]
batch_fetch_calls = []
def fake_api(method, url, *args, **kwargs):
if method == "GET" and url.rstrip("/").endswith("/pulls/848"):
return dict(pr_848)
if method == "GET" and "/pulls/" in url:
raise AssertionError(f"unexpected PR fetch: {url}")
if method == "GET" and "/branches/" in url:
return {"name": "present"}
return {}
def fake_all(url, auth, limit=None):
batch_fetch_calls.append((url, limit))
if "state=open" in url:
return []
if "state=closed" in url:
# Exact mode must not use the closed batch list.
raise AssertionError(
"exact pr_number mode must not page closed PRs: " + url
)
return []
self.mock_api.side_effect = fake_api
self.mock_all.side_effect = fake_all
patch(
"mcp_server._remote_branch_exists",
return_value=True,
).start()
patch(
"mcp_server.merged_cleanup_reconcile.build_reconciliation_report",
side_effect=lambda **kwargs: {
"entries": [
{
"pr_number": int(pr["number"]),
"head_branch": (pr.get("head") or {}).get("ref"),
"issue_number": 844,
"remote_branch": {
"safe_to_delete_remote": True,
"head_branch": (pr.get("head") or {}).get("ref"),
},
"local_worktree": {
"safe_to_remove_worktree": True,
"worktree_path": (
"/tmp/branches/fix-issue-844-exclude-epic-containers"
),
},
"planned_execution_order": (
mcp_server.merged_cleanup_reconcile.plan_cleanup_execution_order(
remote_assessment={"safe_to_delete_remote": True},
local_assessment={"safe_to_remove_worktree": True},
)
),
}
for pr in kwargs.get("closed_prs") or []
if pr.get("merged_at") or pr.get("merged")
],
"reviewer_scratch_entries": [],
"merged_pr_count": len(kwargs.get("closed_prs") or []),
},
).start()
res = gitea_reconcile_merged_cleanups(
dry_run=True,
pr_number=848,
remote="prgs",
org="Scaled-Tech-Consulting",
repo="Gitea-Tools",
)
self.assertTrue(res.get("success"))
self.assertFalse(res.get("performed"))
self.assertEqual(res.get("selection_mode"), "exact_pr")
self.assertEqual(res.get("selected_pr_number"), 848)
entries = res.get("entries") or []
self.assertEqual(len(entries), 1, entries)
self.assertEqual(entries[0].get("pr_number"), 848)
self.assertEqual(
entries[0].get("head_branch"),
"fix/issue-844-exclude-epic-containers",
)
# No other PR appears in plan.
self.assertEqual(list((res.get("planned_execution_orders") or {}).keys()), ["848"])
plan = (res.get("planned_execution_orders") or {}).get("848") or []
actions = [s.get("action") for s in plan]
self.assertEqual(
actions,
[
"remove_local_worktree",
"reassess_branch_ownership",
"delete_remote_branch",
],
)
# Prove we never scanned the multi-PR closed batch.
self.assertFalse(any("state=closed" in (u or "") for u, _ in batch_fetch_calls))
# closed_batch fixture must remain unused (sanity).
self.assertEqual(closed_batch[0]["number"], 852)
def test_exact_pr_execute_only_mutates_selected_pr(self):
"""Execute with pr_number must never touch #845/#846/#849/#852."""
from mcp_server import gitea_reconcile_merged_cleanups
pr_848 = self._merged_pr(848, "fix/issue-844-exclude-epic-containers")
worktree_path = "/tmp/branches/fix-issue-844-exclude-epic-containers"
remove_calls = []
delete_api_calls = []
ownership_branches = []
def fake_api(method, url, *args, **kwargs):
if method == "GET" and url.rstrip("/").endswith("/pulls/848"):
return dict(pr_848)
if method == "DELETE":
delete_api_calls.append(url)
# Forbid foreign PR branch deletion by URL content.
for forbidden in ("845", "846", "849", "852"):
self.assertNotIn(forbidden, url)
return {}
def fake_remove(project_root, branch, worktree_path=None):
remove_calls.append({"branch": branch, "worktree_path": worktree_path})
return {
"success": True,
"performed": True,
"message": f"removed {worktree_path}",
"worktree_path": worktree_path,
}
def fake_collect(**kwargs):
ownership_branches.append(kwargs.get("branch"))
return {"records": [], "inventory_error": False}
def fake_probe(h, o, r, auth, br):
return guard.classify_branch_readback_http_status(
404, not_found_scope=guard.NOT_FOUND_SCOPE_BRANCH
)
self.mock_api.side_effect = fake_api
self.mock_all.side_effect = lambda url, auth, limit=None: []
patch("mcp_server._remote_branch_exists", return_value=True).start()
patch(
"mcp_server.merged_cleanup_reconcile.build_reconciliation_report",
return_value={
"entries": [
{
"pr_number": 848,
"head_branch": "fix/issue-844-exclude-epic-containers",
"remote_branch": {"safe_to_delete_remote": True},
"local_worktree": {
"safe_to_remove_worktree": True,
"worktree_path": worktree_path,
},
"planned_execution_order": [
{"action": "remove_local_worktree", "phase": 1},
{"action": "reassess_branch_ownership", "phase": 2},
{"action": "delete_remote_branch", "phase": 3},
],
}
],
"reviewer_scratch_entries": [
# Foreign scratch must be filtered before report execute loop;
# if present here it would still be a test failure if acted on.
],
"merged_pr_count": 1,
},
).start()
patch(
"mcp_server.merged_cleanup_reconcile.remove_local_worktree",
side_effect=fake_remove,
).start()
patch(
"mcp_server._collect_branch_ownership_records",
side_effect=fake_collect,
).start()
patch("mcp_server._probe_remote_branch", side_effect=fake_probe).start()
res = gitea_reconcile_merged_cleanups(
dry_run=False,
execute_confirmed=True,
pr_number=848,
remote="prgs",
org="Scaled-Tech-Consulting",
repo="Gitea-Tools",
)
self.assertTrue(res.get("performed") or res.get("executed"))
self.assertEqual(res.get("selection_mode"), "exact_pr")
self.assertEqual(res.get("selected_pr_number"), 848)
actions = res.get("actions") or []
pr_numbers_touched = {
a.get("pr_number") for a in actions if a.get("pr_number") is not None
}
self.assertTrue(pr_numbers_touched.issubset({None, 848}) or not pr_numbers_touched)
removes = [a for a in actions if a.get("action") == "remove_local_worktree"]
deletes = [a for a in actions if a.get("action") == "delete_remote_branch"]
self.assertEqual(len(removes), 1)
self.assertEqual(remove_calls[0]["branch"], "fix/issue-844-exclude-epic-containers")
self.assertEqual(len(deletes), 1)
self.assertTrue(deletes[0].get("success"))
self.assertTrue(deletes[0].get("after_worktree_removal"))
self.assertEqual(len(delete_api_calls), 1)
self.assertEqual(
ownership_branches, ["fix/issue-844-exclude-epic-containers"]
)
def test_exact_pr_unknown_fails_closed_without_mutation(self):
from mcp_server import gitea_reconcile_merged_cleanups
def fake_api(method, url, *args, **kwargs):
if method == "GET" and "/pulls/99999" in url:
raise RuntimeError("HTTP 404 Not Found")
raise AssertionError(f"unexpected API call {method} {url}")
self.mock_api.side_effect = fake_api
res = gitea_reconcile_merged_cleanups(
dry_run=True,
pr_number=99999,
remote="prgs",
)
self.assertFalse(res.get("success"))
self.assertFalse(res.get("performed"))
self.assertEqual(res.get("blocker_kind"), "pr_unresolvable")
self.assertIn("99999", " ".join(res.get("reasons") or []))
def test_exact_pr_not_merged_fails_closed(self):
from mcp_server import gitea_reconcile_merged_cleanups
def fake_api(method, url, *args, **kwargs):
if method == "GET" and url.rstrip("/").endswith("/pulls/900"):
return {
"number": 900,
"merged": False,
"merged_at": None,
"state": "open",
"head": {"ref": "feat/x", "sha": "a" * 40},
}
raise AssertionError(f"unexpected {method} {url}")
self.mock_api.side_effect = fake_api
res = gitea_reconcile_merged_cleanups(
dry_run=False,
execute_confirmed=True,
pr_number=900,
remote="prgs",
)
self.assertFalse(res.get("success"))
self.assertFalse(res.get("performed"))
self.assertEqual(res.get("blocker_kind"), "pr_not_merged")
def test_exact_pr_invalid_number_fails_closed(self):
from mcp_server import gitea_reconcile_merged_cleanups
res = gitea_reconcile_merged_cleanups(
dry_run=True,
pr_number=0,
remote="prgs",
)
self.assertFalse(res.get("success"))
self.assertEqual(res.get("blocker_kind"), "invalid_pr_number")
self.mock_api.assert_not_called()
def test_batch_mode_still_works_without_pr_number(self):
"""Unfiltered batch path remains backward compatible."""
from mcp_server import gitea_reconcile_merged_cleanups
self.mock_all.side_effect = lambda url, auth, limit=None: []
self.mock_api.side_effect = lambda *a, **k: {}
patch(
"mcp_server.merged_cleanup_reconcile.build_reconciliation_report",
return_value={
"entries": [],
"reviewer_scratch_entries": [],
"merged_pr_count": 0,
},
).start()
res = gitea_reconcile_merged_cleanups(dry_run=True, remote="prgs", limit=10)
self.assertTrue(res.get("success"))
self.assertEqual(res.get("selection_mode"), "batch")
self.assertIsNone(res.get("selected_pr_number"))
if __name__ == "__main__":
+226 -1
View File
@@ -11,6 +11,7 @@ from datetime import timedelta
from control_plane_db import (
ControlPlaneDB,
ControlPlaneError,
InvalidWorkKindError,
LeaseRequiredError,
WORK_KINDS,
@@ -36,7 +37,7 @@ class ControlPlaneDBTest(unittest.TestCase):
rows = dict(conn.execute("SELECT key, value FROM schema_meta").fetchall())
finally:
conn.close()
self.assertEqual(rows["schema_version"], "4")
self.assertEqual(rows["schema_version"], "5")
self.assertIn("DB coordinates", rows["architecture"])
self.assertIn("bridge", rows["architecture"].lower())
@@ -820,5 +821,229 @@ class ControlPlaneDBTest(unittest.TestCase):
self.assertEqual(n, 1)
class SessionCheckpointTest(unittest.TestCase):
"""Durable MCP session checkpoint schema, redaction, and reconcile (#660)."""
def setUp(self) -> None:
self._tmp = tempfile.TemporaryDirectory()
self.db_path = os.path.join(self._tmp.name, "cp.sqlite3")
self.db = ControlPlaneDB(self.db_path)
def tearDown(self) -> None:
self._tmp.cleanup()
def _write(self, **overrides):
base = dict(
remote="prgs",
org="Scaled-Tech-Consulting",
repo="Gitea-Tools",
session_id="prgs-author-1-abc",
role="author",
work_kind="issue",
work_number=660,
branch="feat/issue-660-session-checkpoint-schema",
head_sha="deadbeef",
lease_id="lease-1",
workflow_stage="implementing",
last_completed_action="wrote schema",
next_valid_action="write tests",
recovery_instructions="re-lock #660 then continue tests",
)
base.update(overrides)
return self.db.write_session_checkpoint(**base)
# AC1 — schema documented and versioned.
def test_table_exists_and_row_carries_schema_version(self) -> None:
import sqlite3
conn = sqlite3.connect(self.db_path)
try:
names = {
r[0]
for r in conn.execute(
"SELECT name FROM sqlite_master WHERE type='table'"
).fetchall()
}
finally:
conn.close()
self.assertIn("session_checkpoints", names)
record = self._write()
self.assertEqual(record["checkpoint_schema_version"], 5)
# AC2 — checkpoints written for multi-role session fixtures.
def test_multi_role_fixtures_each_get_a_row(self) -> None:
roles = [
("prgs-author-1", "author", "issue", 660),
("prgs-reviewer-2", "reviewer", "pr", 795),
("prgs-merger-3", "merger", "pr", 862),
("prgs-controller-4", "controller", "issue", 653),
]
for session_id, role, kind, number in roles:
self._write(
session_id=session_id,
role=role,
work_kind=kind,
work_number=number,
lease_id=f"lease-{session_id}",
)
rows = self.db.list_session_checkpoints(remote="prgs")
self.assertEqual(len(rows), 4)
self.assertEqual(
{r["role"] for r in rows},
{"author", "reviewer", "merger", "controller"},
)
def test_upsert_is_current_state_and_audits_stage_change(self) -> None:
first = self._write(workflow_stage="implementing")
second = self._write(workflow_stage="testing")
self.assertEqual(first["checkpoint_id"], second["checkpoint_id"])
rows = self.db.list_session_checkpoints(
remote="prgs", session_id="prgs-author-1-abc"
)
self.assertEqual(len(rows), 1)
self.assertEqual(rows[0]["workflow_stage"], "testing")
import sqlite3
conn = sqlite3.connect(self.db_path)
try:
n = conn.execute(
"SELECT COUNT(*) FROM events "
"WHERE event_type = 'session_checkpoint_stage_change'"
).fetchone()[0]
finally:
conn.close()
self.assertEqual(n, 1)
def test_get_and_roundtrip_json_fields(self) -> None:
self._write(
capabilities=["gitea.repo.commit", "gitea.pr.create"],
evidence={"tests": "4 passing"},
pending_mutation={"op": "commit_files", "files": ["control_plane_db.py"]},
)
got = self.db.get_session_checkpoint(
remote="prgs",
org="Scaled-Tech-Consulting",
repo="Gitea-Tools",
session_id="prgs-author-1-abc",
work_kind="issue",
work_number=660,
)
self.assertIsNotNone(got)
self.assertEqual(got["capabilities"], ["gitea.repo.commit", "gitea.pr.create"])
self.assertEqual(got["evidence"], {"tests": "4 passing"})
self.assertEqual(got["pending_mutation"]["op"], "commit_files")
# AC4 — no secrets in stored records.
def test_secrets_are_redacted_before_storage(self) -> None:
self._write(
recovery_instructions=(
"resume with Authorization: Bearer sk-supersecrettoken then retry"
),
evidence={"authorization": "Bearer sk-anothersecret"},
pending_mutation={"url": "https://user:[email protected]/repo.git"},
)
import sqlite3
conn = sqlite3.connect(self.db_path)
try:
row = conn.execute(
"SELECT recovery_instructions, evidence, pending_mutation "
"FROM session_checkpoints"
).fetchone()
finally:
conn.close()
blob = " ".join(str(v) for v in row)
self.assertNotIn("sk-supersecrettoken", blob)
self.assertNotIn("sk-anothersecret", blob)
self.assertNotIn("password", blob)
self.assertIn("REDACTED", blob)
# AC3 — reconcile detects stale head / lease mismatch.
def test_reconcile_flags_stale_head(self) -> None:
record = self._write(head_sha="aaaa1111")
result = self.db.reconcile_session_checkpoint(
record, live_head_sha="bbbb2222", live_lease_active=True,
live_lease_id="lease-1",
)
self.assertTrue(result["stale"])
self.assertTrue(result["head_mismatch"])
self.assertFalse(result["lease_mismatch"])
self.assertEqual(result["reconcile_action"], "reconcile_required")
def test_reconcile_flags_dead_lease(self) -> None:
record = self._write(lease_id="lease-1", head_sha="aaaa1111")
result = self.db.reconcile_session_checkpoint(
record, live_head_sha="aaaa1111", live_lease_active=False,
)
self.assertTrue(result["stale"])
self.assertFalse(result["head_mismatch"])
self.assertTrue(result["lease_mismatch"])
def test_reconcile_reassigned_lease_is_stale(self) -> None:
record = self._write(lease_id="lease-1")
result = self.db.reconcile_session_checkpoint(
record, live_lease_active=True, live_lease_id="lease-999",
)
self.assertTrue(result["lease_mismatch"])
def test_reconcile_clean_state_is_safe_to_resume(self) -> None:
record = self._write(head_sha="aaaa1111", lease_id="lease-1")
result = self.db.reconcile_session_checkpoint(
record, live_head_sha="aaaa1111", live_lease_active=True,
live_lease_id="lease-1",
)
self.assertFalse(result["stale"])
self.assertEqual(result["reconcile_action"], "safe_to_resume")
def test_unknown_live_state_never_flags_mismatch(self) -> None:
record = self._write(head_sha="aaaa1111", lease_id="lease-1")
result = self.db.reconcile_session_checkpoint(record)
self.assertFalse(result["stale"])
# Drain gate — fail closed when a checkpoint is incomplete.
def test_drain_requires_complete_checkpoint(self) -> None:
with self.assertRaises(ControlPlaneError):
self.db.write_session_checkpoint(
remote="prgs",
org="Scaled-Tech-Consulting",
repo="Gitea-Tools",
session_id="prgs-author-1-abc",
role="author",
workflow_stage="draining",
# next_valid_action + recovery_instructions intentionally absent
require_complete=True,
)
# Nothing was written.
rows = self.db.list_session_checkpoints(remote="prgs")
self.assertEqual(rows, [])
def test_drain_write_succeeds_when_complete(self) -> None:
record = self._write(require_complete=True)
self.assertEqual(record["status"], "active")
self.assertEqual(self.db.checkpoint_completeness(record), [])
def test_missing_session_id_fails_closed(self) -> None:
with self.assertRaises(ControlPlaneError):
self.db.write_session_checkpoint(
remote="prgs", org="o", repo="r", session_id="",
)
def test_session_level_checkpoint_uses_sentinel_key(self) -> None:
# No work unit -> ('', 0) sentinel; a second session-level write upserts.
self.db.write_session_checkpoint(
remote="prgs", org="o", repo="r", session_id="s-sess",
workflow_stage="idle",
)
self.db.write_session_checkpoint(
remote="prgs", org="o", repo="r", session_id="s-sess",
workflow_stage="booting",
)
rows = self.db.list_session_checkpoints(remote="prgs", session_id="s-sess")
self.assertEqual(len(rows), 1)
self.assertEqual(rows[0]["work_kind"], "")
self.assertEqual(rows[0]["work_number"], 0)
if __name__ == "__main__":
unittest.main()
@@ -0,0 +1,483 @@
"""Synthetic regression coverage for dirty orphaned worktree recovery (#860).
Modeled on the #850 / #855 shape without mutating their real state.
"""
from __future__ import annotations
import json
import os
import shutil
import tempfile
import unittest
from unittest import mock
import dirty_orphan_worktree_recovery as dorec
import issue_lock_store
DEAD_PID = 999_999_999
LIVE_PID = os.getpid()
BRANCH = "fix/issue-901-dirty-orphan"
SOURCE_WT = "/repo/branches/issue-901-dirty-orphan"
RECOVERY_WT_NAME = "recovery-issue-901-dirty-orphan"
LOCAL_HEAD = "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa"
REMOTE_HEAD = "bbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbb"
OTHER_HEAD = "cccccccccccccccccccccccccccccccccccccccc"
FP_A = dorec.sha256_bytes(b"dirty-a")
FP_B = dorec.sha256_bytes(b"dirty-b")
FP_C = dorec.sha256_bytes(b"dirty-c-conflict")
def durable_lock(**overrides):
"""#850-shaped PID-less malformed same-claimant lock."""
lock = {
"issue_number": 901,
"branch_name": BRANCH,
"worktree_path": SOURCE_WT,
"remote": "prgs",
"org": "Example-Org",
"repo": "Example-Repo",
# intentionally no pid / session_pid / work_lease expiry
"claimant": {"username": "author-user", "profile": "prgs-author"},
}
lock.update(overrides)
return lock
def base_kwargs(**overrides):
kwargs = {
"issue_number": 901,
"branch_name": BRANCH,
"source_worktree_path": SOURCE_WT,
"remote": "prgs",
"org": "Example-Org",
"repo": "Example-Repo",
"identity": "author-user",
"profile": "prgs-author",
"expected_local_head": LOCAL_HEAD,
"expected_remote_head": REMOTE_HEAD,
"expected_dirty_fingerprints": {"a.py": FP_A, "b.py": FP_B},
"current_branch": BRANCH,
"porcelain_status": " M a.py\n M b.py\n",
"observed_local_head": LOCAL_HEAD,
"observed_remote_head": REMOTE_HEAD,
"observed_dirty_fingerprints": {"a.py": FP_A, "b.py": FP_B},
"competing_live_locks": [],
"competing_live_sessions": [],
"workflow_lease_active": False,
"workflow_lease_expired": True,
"canonical_repo_root": "/repo",
"worktree_registered": True,
"current_pid": LIVE_PID,
}
kwargs.update(overrides)
return kwargs
def assess(lock=None, **overrides):
return dorec.assess_dirty_orphan_recovery(
durable_lock() if lock is None else lock, **base_kwargs(**overrides)
)
class FreshnessPidLess(unittest.TestCase):
def test_pid_less_lock_is_not_live(self):
freshness = issue_lock_store.assess_lock_freshness(durable_lock())
self.assertFalse(freshness["live"])
self.assertTrue(freshness.get("pid_missing"))
self.assertEqual(freshness["status"], "malformed")
def test_pid_less_with_far_future_expiry_still_not_live(self):
lock = durable_lock(
work_lease={
"operation_type": "author_issue_work",
"expires_at": "2999-01-01T00:00:00Z",
"last_heartbeat_at": "2999-01-01T00:00:00Z",
}
)
freshness = issue_lock_store.assess_lock_freshness(lock)
self.assertFalse(freshness["live"])
self.assertTrue(freshness.get("pid_missing"))
class EligibilityGranted(unittest.TestCase):
def test_dead_same_claimant_pid_less_dirty(self):
result = assess()
self.assertEqual(result["outcome"], dorec.ELIGIBLE)
self.assertTrue(result["eligible"])
def test_expired_workflow_lease_corroboration(self):
result = assess(workflow_lease_active=False, workflow_lease_expired=True)
self.assertTrue(result["eligible"])
def test_older_local_newer_remote_heads(self):
result = assess()
self.assertTrue(result["evidence"].get("heads_diverged"))
self.assertTrue(result["eligible"])
class EligibilityRefused(unittest.TestCase):
def test_active_owner_with_pid(self):
lock = durable_lock(pid=LIVE_PID, session_pid=LIVE_PID)
result = assess(lock=lock, owner_process_alive_override=True)
self.assertEqual(result["outcome"], dorec.REFUSED)
self.assertFalse(result["eligible"])
self.assertTrue(any("alive" in r for r in result["reasons"]))
def test_foreign_claimant(self):
result = assess(identity="other-user")
self.assertEqual(result["outcome"], dorec.REFUSED)
self.assertTrue(any("foreign claimant identity" in r for r in result["reasons"]))
def test_foreign_profile(self):
result = assess(profile="prgs-reviewer")
self.assertEqual(result["outcome"], dorec.REFUSED)
def test_fingerprint_mismatch(self):
result = assess(observed_dirty_fingerprints={"a.py": "0" * 64, "b.py": FP_B})
self.assertEqual(result["outcome"], dorec.REFUSED)
self.assertTrue(any("fingerprint mismatch" in r for r in result["reasons"]))
def test_head_mismatch(self):
result = assess(observed_local_head=OTHER_HEAD)
self.assertEqual(result["outcome"], dorec.REFUSED)
def test_remote_head_mismatch(self):
result = assess(observed_remote_head=OTHER_HEAD)
self.assertEqual(result["outcome"], dorec.REFUSED)
def test_path_not_under_branches(self):
result = assess(
source_worktree_path="/tmp/branches/evil",
# lock path also changed so worktree agreement holds
lock=durable_lock(worktree_path="/tmp/branches/evil"),
)
self.assertEqual(result["outcome"], dorec.REFUSED)
self.assertTrue(any("canonical branches" in r for r in result["reasons"]))
def test_unregistered_worktree(self):
result = assess(worktree_registered=False)
self.assertEqual(result["outcome"], dorec.REFUSED)
def test_active_workflow_lease(self):
result = assess(workflow_lease_active=True, workflow_lease_expired=False)
self.assertEqual(result["outcome"], dorec.REFUSED)
def test_unsafe_dirty_path_pin(self):
result = assess(
expected_dirty_fingerprints={"../etc/passwd": FP_A},
observed_dirty_fingerprints={"../etc/passwd": FP_A},
)
self.assertEqual(result["outcome"], dorec.REFUSED)
def test_symlink_escape_rejected_by_ancestry(self):
ok, reasons = dorec.is_path_under_canonical_branches(
"/tmp/branches/evil", canonical_repo_root="/repo"
)
self.assertFalse(ok)
self.assertTrue(reasons)
class ConflictDetection(unittest.TestCase):
def test_overlapping_upstream_change(self):
conflicts = dorec.detect_path_conflicts(
dirty_paths=["c.py"],
local_head_contents={"c.py": b"local-base"},
remote_head_contents={"c.py": b"remote-changed"},
dirty_contents={"c.py": b"dirty-c-conflict"},
)
self.assertEqual(len(conflicts), 1)
self.assertEqual(conflicts[0]["path"], "c.py")
def test_unchanged_upstream_no_conflict(self):
conflicts = dorec.detect_path_conflicts(
dirty_paths=["a.py"],
local_head_contents={"a.py": b"same"},
remote_head_contents={"a.py": b"same"},
dirty_contents={"a.py": b"dirty-a"},
)
self.assertEqual(conflicts, [])
class CrashSafeRecovery(unittest.TestCase):
def setUp(self):
self.tmp = tempfile.mkdtemp(prefix="dirty-orphan-")
self.repo = os.path.join(self.tmp, "repo")
self.branches = os.path.join(self.repo, "branches")
self.source = os.path.join(self.branches, "issue-901-dirty-orphan")
self.recovery = os.path.join(self.branches, RECOVERY_WT_NAME)
os.makedirs(self.source, exist_ok=True)
os.makedirs(self.branches, exist_ok=True)
# seed dirty files in source
with open(os.path.join(self.source, "a.py"), "wb") as fh:
fh.write(b"dirty-a")
with open(os.path.join(self.source, "b.py"), "wb") as fh:
fh.write(b"dirty-b")
self.journal_dir = os.path.join(self.tmp, "journals")
self.lock = durable_lock(worktree_path=self.source)
self.assessment = dorec.assess_dirty_orphan_recovery(
self.lock,
**base_kwargs(
source_worktree_path=self.source,
canonical_repo_root=self.repo,
),
)
class FakeGit(dorec.GitOps):
def __init__(self, recovery_path, head):
self.recovery_path = recovery_path
self.head = head
self.calls = []
def run(self, args, *, cwd):
self.calls.append((args, cwd))
if args[:3] == ["git", "worktree", "add"]:
os.makedirs(self.recovery_path, exist_ok=True)
return mock.Mock(returncode=0, stdout="", stderr="")
if args[:2] == ["git", "checkout"]:
return mock.Mock(returncode=0, stdout="", stderr="")
if args[:2] == ["git", "rev-parse"]:
return mock.Mock(returncode=0, stdout=self.head + "\n", stderr="")
return mock.Mock(returncode=0, stdout="", stderr="")
self.git = FakeGit(self.recovery, REMOTE_HEAD)
self.written_locks = []
def lock_writer(record):
self.written_locks.append(record)
self.lock_writer = lock_writer
def tearDown(self):
shutil.rmtree(self.tmp, ignore_errors=True)
def _run(self, **overrides):
kwargs = {
"assessment": self.assessment,
"existing_lock": self.lock,
"issue_number": 901,
"branch_name": BRANCH,
"source_worktree_path": self.source,
"recovery_worktree_path": self.recovery,
"remote": "prgs",
"org": "Example-Org",
"repo": "Example-Repo",
"identity": "author-user",
"profile": "prgs-author",
"expected_local_head": LOCAL_HEAD,
"expected_remote_head": REMOTE_HEAD,
"expected_dirty_fingerprints": {"a.py": FP_A, "b.py": FP_B},
"dirty_contents": {"a.py": b"dirty-a", "b.py": b"dirty-b"},
"local_head_contents": {"a.py": b"base-a", "b.py": b"base-b"},
"remote_head_contents": {"a.py": b"base-a", "b.py": b"base-b"},
"canonical_repo_root": self.repo,
"bind_lock": True,
"lock_writer": self.lock_writer,
"git_ops": self.git,
"journal_dir": self.journal_dir,
"session_pid": LIVE_PID,
}
kwargs.update(overrides)
return dorec.run_dirty_orphan_recovery(**kwargs)
def test_success_preserves_dirty_bytes_and_source(self):
result = self._run()
self.assertTrue(result["success"])
self.assertEqual(result["outcome"], dorec.RECOVERY_COMPLETED)
self.assertTrue(os.path.isdir(self.source))
with open(os.path.join(self.source, "a.py"), "rb") as fh:
self.assertEqual(fh.read(), b"dirty-a")
with open(os.path.join(self.recovery, "a.py"), "rb") as fh:
self.assertEqual(fh.read(), b"dirty-a")
with open(os.path.join(self.recovery, "b.py"), "rb") as fh:
self.assertEqual(fh.read(), b"dirty-b")
self.assertEqual(len(self.written_locks), 1)
rec = self.written_locks[0]
self.assertEqual(rec["session_pid"], LIVE_PID)
self.assertTrue(rec["dirty_orphan_recovery"]["recovered"])
self.assertTrue(rec["dirty_orphan_recovery"]["source_frozen"])
def test_conflict_leaves_governed_state(self):
result = self._run(
expected_dirty_fingerprints={"c.py": FP_C},
dirty_contents={"c.py": b"dirty-c-conflict"},
local_head_contents={"c.py": b"local-base"},
remote_head_contents={"c.py": b"remote-changed"},
)
# #860 F4: session binding is NOT finalized while conflicts remain
self.assertFalse(result["success"])
self.assertEqual(result["outcome"], dorec.CONFLICTS_PRESENT)
sidecar = os.path.join(self.recovery, "c.py.recovered-dirty")
self.assertTrue(os.path.isfile(sidecar))
state = os.path.join(
self.recovery, dorec.CONFLICT_STATE_DIR, dorec.CONFLICT_STATE_FILE
)
self.assertTrue(os.path.isfile(state))
with open(state, "r", encoding="utf-8") as fh:
payload = json.load(fh)
self.assertEqual(payload["resolution"], "author_edit_required")
def test_interrupt_before_journal_no_artifacts(self):
result = self._run(interrupt_after_phase=dorec.PHASE_ELIGIBILITY)
self.assertFalse(result["success"])
self.assertEqual(result["outcome"], "INTERRUPTED")
self.assertFalse(os.path.isdir(self.recovery))
def test_interrupt_after_journal_then_retry_idempotent(self):
first = self._run(interrupt_after_phase=dorec.PHASE_JOURNAL_PERSISTED)
self.assertEqual(first["outcome"], "INTERRUPTED")
self.assertTrue(first["journal"]["artifacts_created"]["journal"])
second = self._run()
self.assertTrue(second["success"])
# source still recoverable
with open(os.path.join(self.source, "a.py"), "rb") as fh:
self.assertEqual(fh.read(), b"dirty-a")
def test_interrupt_after_worktree_then_retry(self):
first = self._run(interrupt_after_phase=dorec.PHASE_RECOVERY_WORKTREE)
self.assertEqual(first["outcome"], "INTERRUPTED")
self.assertTrue(os.path.isdir(self.recovery))
second = self._run()
self.assertTrue(second["success"])
def test_interrupt_after_binding_then_retry_complete(self):
first = self._run(interrupt_after_phase=dorec.PHASE_BINDING)
self.assertEqual(first["outcome"], "INTERRUPTED")
second = self._run()
self.assertTrue(second["success"])
# completed journal makes further retries no-ops
third = self._run()
self.assertEqual(third["outcome"], dorec.RECOVERY_RESUMED)
def test_source_worktree_never_deleted(self):
self._run()
self.assertTrue(os.path.isdir(self.source))
self.assertTrue(os.path.isfile(os.path.join(self.source, "a.py")))
def test_fingerprint_drift_refuses_without_mutation(self):
result = self._run(dirty_contents={"a.py": b"CHANGED", "b.py": b"dirty-b"})
self.assertFalse(result["success"])
self.assertFalse(os.path.isdir(self.recovery))
class SessionBindingPreflight(unittest.TestCase):
def test_canonical_session_binding_recognized(self):
lock = {
"worktree_path": "/repo/branches/recovery",
"session_pid": LIVE_PID,
"dirty_orphan_recovery": {
"recovered": True,
"conflicts": [],
"recovery_worktree_path": "/repo/branches/recovery",
"source_worktree_path": SOURCE_WT,
"accepted_head": REMOTE_HEAD,
},
}
result = dorec.preflight_recognizes_recovered_provenance(lock)
self.assertTrue(result["recognized"])
def test_conflicts_block_commit_preflight(self):
lock = {
"worktree_path": "/repo/branches/recovery",
"session_pid": LIVE_PID,
"dirty_orphan_recovery": {
"recovered": True,
"conflicts": [{"path": "c.py"}],
},
}
result = dorec.preflight_recognizes_recovered_provenance(lock)
self.assertFalse(result["recognized"])
def test_active_foreign_does_not_mutate(self):
# assess-only path: foreign refused before run
result = assess(identity="intruder")
self.assertFalse(result["eligible"])
class JournalSymlinkRefusal(unittest.TestCase):
def test_symlink_journal_path_refused_on_load(self):
tmp = tempfile.mkdtemp()
try:
real = os.path.join(tmp, "real.json")
with open(real, "w", encoding="utf-8") as fh:
fh.write("{}")
link = os.path.join(tmp, "link.json")
os.symlink(real, link)
key = "symlink-test"
jdir = tmp
path = dorec._journal_path(key, journal_dir=jdir)
with open(path, "w", encoding="utf-8") as fh:
json.dump({"idempotency_key": key}, fh)
os.remove(path)
os.symlink(real, path)
with self.assertRaises(ValueError):
dorec.load_journal(key, journal_dir=jdir)
finally:
shutil.rmtree(tmp, ignore_errors=True)
class RealGitMultiWorktreeIntegration(unittest.TestCase):
def setUp(self):
import subprocess
self.tmp = tempfile.mkdtemp(prefix="git-integration-")
self.repo = os.path.join(self.tmp, "repo")
os.makedirs(self.repo, exist_ok=True)
subprocess.run(["git", "init"], cwd=self.repo, check=True, capture_output=True)
subprocess.run(["git", "config", "user.name", "Test User"], cwd=self.repo, check=True)
subprocess.run(["git", "config", "user.email", "[email protected]"], cwd=self.repo, check=True)
with open(os.path.join(self.repo, "init.txt"), "w") as fh:
fh.write("init")
subprocess.run(["git", "add", "."], cwd=self.repo, check=True)
subprocess.run(["git", "commit", "-m", "init"], cwd=self.repo, check=True)
branch = "fix/issue-999-test"
subprocess.run(["git", "branch", branch], cwd=self.repo, check=True)
self.branches = os.path.join(self.repo, "branches")
self.source = os.path.join(self.branches, "issue-999-test")
subprocess.run(["git", "worktree", "add", self.source, branch], cwd=self.repo, check=True)
self.dirty_path = os.path.join(self.source, "dirty.txt")
with open(self.dirty_path, "w") as fh:
fh.write("dirty-data")
def tearDown(self):
shutil.rmtree(self.tmp, ignore_errors=True)
def test_prepare_recovery_worktree_detached_no_exit_128(self):
import subprocess
head_sha = subprocess.check_output(["git", "rev-parse", "HEAD"], cwd=self.repo, text=True).strip()
rec_wt = os.path.join(self.branches, "recovery-issue-999-test")
res = dorec.prepare_recovery_worktree(
canonical_repo_root=self.repo,
recovery_worktree_path=rec_wt,
branch_name="fix/issue-999-test",
remote_head=head_sha,
)
self.assertTrue(res["success"], res.get("reasons"))
self.assertTrue(os.path.isdir(rec_wt))
def test_real_lock_rebind_recovery_sanctioned(self):
lock_dir = os.path.join(self.tmp, "locks")
rec_wt = os.path.join(self.branches, "recovery-issue-999-test")
os.makedirs(rec_wt, exist_ok=True)
record = {
"remote": "prgs",
"org": "Example-Org",
"repo": "Example-Repo",
"issue_number": 999,
"branch_name": "fix/issue-999-test",
"worktree_path": rec_wt,
"claimant": {"username": "author-user", "profile": "prgs-author"},
}
record_src = dict(record)
record_src["worktree_path"] = self.source
issue_lock_store.bind_session_lock(record_src, lock_dir=lock_dir)
path = issue_lock_store.bind_session_lock(
record,
lock_dir=lock_dir,
recovery_sanctioned=True,
)
self.assertTrue(os.path.isfile(path))
if __name__ == "__main__":
unittest.main()
File diff suppressed because it is too large Load Diff
+383
View File
@@ -0,0 +1,383 @@
"""Tests for the pre-restart drain proof and hard gate (#661).
Covers the acceptance criteria:
1. Restart apply without a proof fails closed.
2. A successful drain produces a verifiable proof.
3. An open unsafe mutation makes the proof fail (multi-session fixture).
4. Pass / fail / expired verification paths.
Plus the security posture: forged/tampered proofs are rejected, break-glass is
the only bypass and is never silent, a stale blast-radius fingerprint rejects a
proof, and no per-process secret ever leaks into a serialized artifact.
"""
from __future__ import annotations
import os
import unittest
from datetime import datetime, timedelta, timezone
import drain_proof as dp
import restart_coordinator as rc
NOW = datetime(2026, 7, 24, 6, 0, 0, tzinfo=timezone.utc)
SECRET = b"unit-test-drain-proof-secret-0123456789abcdef"
def _live_pid() -> int:
return os.getpid()
def _clean_drain_state() -> dict:
"""Every drain action succeeded, no sessions outstanding."""
return {
"assignments_stopped": True,
"checkpoints_complete": True,
"handoffs_verified": True,
"leases_handled": True,
"acks": {}, # no other live sessions to acknowledge
"ack_timeout_policy_applied": False,
}
def _safe_report() -> dict:
"""Impact report with no other live work: a restart here is safe."""
report = rc.evaluate_restart_impact(
{"sessions": [], "leases": [], "inventory_complete": True},
now=NOW,
requesting_session_id="prgs-controller-1-req",
)
return report.as_dict()
def _unsafe_mutation_report() -> dict:
"""Multi-session report: a second session holds a live author mutation."""
sessions = [
{
"session_id": "prgs-controller-1-req",
"role": "controller",
"profile": "prgs-controller",
"pid": _live_pid(),
"status": "active",
"last_heartbeat_at": NOW.isoformat(),
},
{
"session_id": "prgs-author-99",
"role": "author",
"profile": "prgs-author",
"pid": _live_pid(),
"status": "active",
"last_heartbeat_at": NOW.isoformat(),
},
]
leases = [
{
"lease_id": "lease-mut",
"session_id": "prgs-author-99",
"role": "author",
"phase": "implementing",
"work_kind": "issue",
"work_number": 661,
"worktree_path": "branches/issue-661",
"freshness": {"freshness": "active"},
}
]
report = rc.evaluate_restart_impact(
{"sessions": sessions, "leases": leases, "inventory_complete": True},
now=NOW,
requesting_session_id="prgs-controller-1-req",
)
return report.as_dict()
class BuildDrainProofTests(unittest.TestCase):
def test_clean_drain_produces_verifiable_clean_proof(self):
"""AC#2: a successful drain produces a verifiable proof."""
proof = dp.build_drain_proof(
impact_report=_safe_report(),
drain_state=_clean_drain_state(),
requesting_session_id="prgs-controller-1-req",
now=NOW,
secret=SECRET,
)
self.assertTrue(proof.clean)
self.assertEqual(proof.failed_checks, [])
self.assertEqual(
{c.name for c in proof.checks}, set(dp.REQUIRED_CHECKS)
)
result = dp.verify_drain_proof(
proof.as_dict(), now=NOW, secret=SECRET
)
self.assertTrue(result.valid, result.reasons)
self.assertFalse(result.expired)
self.assertFalse(result.tampered)
def test_open_mutation_makes_proof_unclean(self):
"""AC#3: an unsafe mutation still in flight fails the proof."""
proof = dp.build_drain_proof(
impact_report=_unsafe_mutation_report(),
drain_state=_clean_drain_state(),
now=NOW,
secret=SECRET,
)
self.assertFalse(proof.clean)
self.assertIn(dp.CHECK_NO_INFLIGHT_MUTATIONS, proof.failed_checks)
# Leases-handled also fails: the report still shows a disruptive lease.
self.assertIn(dp.CHECK_LEASES_HANDLED, proof.failed_checks)
result = dp.verify_drain_proof(proof.as_dict(), now=NOW, secret=SECRET)
self.assertFalse(result.valid)
def test_incomplete_inventory_fails_no_mutations_check(self):
proof = dp.build_drain_proof(
impact_report={"inventory_complete": False},
drain_state=_clean_drain_state(),
now=NOW,
secret=SECRET,
)
self.assertFalse(proof.clean)
self.assertIn(dp.CHECK_NO_INFLIGHT_MUTATIONS, proof.failed_checks)
def test_missing_checkpoint_flag_fails_closed(self):
state = _clean_drain_state()
del state["checkpoints_complete"]
proof = dp.build_drain_proof(
impact_report=_safe_report(), drain_state=state, now=NOW, secret=SECRET
)
self.assertFalse(proof.clean)
self.assertIn(dp.CHECK_CHECKPOINTS_COMPLETE, proof.failed_checks)
def test_non_true_flags_fail_closed(self):
"""A truthy-but-not-True value (e.g. the string 'yes') must not pass."""
state = _clean_drain_state()
state["assignments_stopped"] = "yes"
proof = dp.build_drain_proof(
impact_report=_safe_report(), drain_state=state, now=NOW, secret=SECRET
)
self.assertIn(dp.CHECK_ASSIGNMENTS_STOPPED, proof.failed_checks)
def test_ack_timeout_policy_satisfies_ack_check(self):
state = _clean_drain_state()
state["acks"] = {"prgs-author-99": "pending"}
state["ack_timeout_policy_applied"] = True
proof = dp.build_drain_proof(
impact_report=_safe_report(), drain_state=state, now=NOW, secret=SECRET
)
names = {c.name: c.passed for c in proof.checks}
self.assertTrue(names[dp.CHECK_ACKS_OR_TIMEOUT])
def test_outstanding_acks_without_timeout_fail(self):
state = _clean_drain_state()
state["acks"] = {"prgs-author-99": "pending"}
state["ack_timeout_policy_applied"] = False
proof = dp.build_drain_proof(
impact_report=_safe_report(), drain_state=state, now=NOW, secret=SECRET
)
self.assertIn(dp.CHECK_ACKS_OR_TIMEOUT, proof.failed_checks)
def test_all_acked_satisfies_ack_check(self):
state = _clean_drain_state()
state["acks"] = {"prgs-author-99": "acked", "prgs-author-2": "acknowledged"}
proof = dp.build_drain_proof(
impact_report=_safe_report(), drain_state=state, now=NOW, secret=SECRET
)
names = {c.name: c.passed for c in proof.checks}
self.assertTrue(names[dp.CHECK_ACKS_OR_TIMEOUT])
class VerifyDrainProofTests(unittest.TestCase):
def _clean_proof_dict(self) -> dict:
return dp.build_drain_proof(
impact_report=_safe_report(),
drain_state=_clean_drain_state(),
now=NOW,
secret=SECRET,
).as_dict()
def test_missing_proof_is_invalid(self):
result = dp.verify_drain_proof(None, now=NOW, secret=SECRET)
self.assertFalse(result.valid)
self.assertIsNone(result.proof_id)
def test_expired_proof_is_invalid(self):
"""AC#4: an expired proof fails verification."""
proof = self._clean_proof_dict()
later = NOW + timedelta(seconds=dp.DEFAULT_PROOF_TTL_SECONDS + 1)
result = dp.verify_drain_proof(proof, now=later, secret=SECRET)
self.assertFalse(result.valid)
self.assertTrue(result.expired)
def test_proof_valid_just_before_expiry(self):
proof = self._clean_proof_dict()
almost = NOW + timedelta(seconds=dp.DEFAULT_PROOF_TTL_SECONDS - 1)
result = dp.verify_drain_proof(proof, now=almost, secret=SECRET)
self.assertTrue(result.valid, result.reasons)
def test_wrong_secret_rejected(self):
"""A proof minted in a prior process (different secret) will not verify."""
proof = self._clean_proof_dict()
result = dp.verify_drain_proof(proof, now=NOW, secret=b"other-secret")
self.assertFalse(result.valid)
self.assertTrue(result.tampered)
def test_flipping_clean_flag_is_detected(self):
"""Forging clean=True on an unclean proof breaks the signature."""
unclean = dp.build_drain_proof(
impact_report=_unsafe_mutation_report(),
drain_state=_clean_drain_state(),
now=NOW,
secret=SECRET,
).as_dict()
self.assertFalse(unclean["clean"])
unclean["clean"] = True # forge
result = dp.verify_drain_proof(unclean, now=NOW, secret=SECRET)
self.assertFalse(result.valid)
self.assertTrue(result.tampered)
def test_tampering_a_check_is_detected(self):
unclean = dp.build_drain_proof(
impact_report=_unsafe_mutation_report(),
drain_state=_clean_drain_state(),
now=NOW,
secret=SECRET,
).as_dict()
for c in unclean["checks"]:
if c["name"] == dp.CHECK_NO_INFLIGHT_MUTATIONS:
c["passed"] = True # forge the failing check to pass
result = dp.verify_drain_proof(unclean, now=NOW, secret=SECRET)
self.assertFalse(result.valid)
self.assertTrue(result.tampered)
def test_missing_required_check_rejected(self):
proof = self._clean_proof_dict()
proof["checks"] = [
c for c in proof["checks"] if c["name"] != dp.CHECK_HANDOFFS_OK
]
result = dp.verify_drain_proof(proof, now=NOW, secret=SECRET)
self.assertFalse(result.valid)
def test_stale_fingerprint_rejected(self):
proof = self._clean_proof_dict()
result = dp.verify_drain_proof(
proof,
now=NOW,
secret=SECRET,
expected_impact_fingerprint="deadbeef",
)
self.assertFalse(result.valid)
def test_matching_fingerprint_accepted(self):
report = _safe_report()
proof = dp.build_drain_proof(
impact_report=report,
drain_state=_clean_drain_state(),
now=NOW,
secret=SECRET,
).as_dict()
fp = dp.impact_fingerprint(report)
result = dp.verify_drain_proof(
proof, now=NOW, secret=SECRET, expected_impact_fingerprint=fp
)
self.assertTrue(result.valid, result.reasons)
class GateApplyRestartTests(unittest.TestCase):
def _clean_proof_dict(self) -> dict:
return dp.build_drain_proof(
impact_report=_safe_report(),
drain_state=_clean_drain_state(),
now=NOW,
secret=SECRET,
).as_dict()
def test_apply_without_proof_denied(self):
"""AC#1: restart apply without a proof fails closed + raises incident."""
decision = dp.gate_apply_restart(proof=None, now=NOW, secret=SECRET)
self.assertFalse(decision.allow)
self.assertEqual(decision.verdict, dp.GATE_DENY)
self.assertIsNotNone(decision.incident)
self.assertEqual(
decision.incident["kind"], "restart_drain_gate_denied"
)
def test_apply_with_valid_proof_allowed(self):
decision = dp.gate_apply_restart(
proof=self._clean_proof_dict(), now=NOW, secret=SECRET
)
self.assertTrue(decision.allow)
self.assertEqual(decision.verdict, dp.GATE_ALLOW)
self.assertIsNone(decision.incident)
def test_apply_with_expired_proof_denied_with_incident(self):
later = NOW + timedelta(seconds=dp.DEFAULT_PROOF_TTL_SECONDS + 5)
decision = dp.gate_apply_restart(
proof=self._clean_proof_dict(), now=later, secret=SECRET
)
self.assertFalse(decision.allow)
self.assertIsNotNone(decision.incident)
def test_apply_with_unclean_proof_denied(self):
"""AC#3 at the gate: an unsafe-mutation proof is denied."""
unclean = dp.build_drain_proof(
impact_report=_unsafe_mutation_report(),
drain_state=_clean_drain_state(),
now=NOW,
secret=SECRET,
).as_dict()
decision = dp.gate_apply_restart(proof=unclean, now=NOW, secret=SECRET)
self.assertFalse(decision.allow)
self.assertIsNotNone(decision.incident)
def test_break_glass_allows_without_proof_but_records_bypass(self):
decision = dp.gate_apply_restart(
proof=None, now=NOW, secret=SECRET, break_glass=True
)
self.assertTrue(decision.allow)
self.assertEqual(decision.verdict, dp.GATE_BREAK_GLASS)
self.assertTrue(decision.break_glass)
self.assertIsNone(decision.incident)
self.assertTrue(decision.audit_record["break_glass"])
def test_denied_gate_carries_stale_fingerprint_reason(self):
decision = dp.gate_apply_restart(
proof=self._clean_proof_dict(),
now=NOW,
secret=SECRET,
expected_impact_fingerprint="not-the-fingerprint",
)
self.assertFalse(decision.allow)
class SecretHygieneTests(unittest.TestCase):
def test_secret_never_serialized(self):
proof = dp.build_drain_proof(
impact_report=_safe_report(),
drain_state=_clean_drain_state(),
now=NOW,
secret=SECRET,
)
blob = dp._canonical(proof.as_dict())
self.assertNotIn(SECRET.decode(), blob)
# The signature is a hex digest, not the raw secret.
self.assertNotIn(SECRET.hex(), blob)
def test_incident_descriptor_has_no_secret(self):
decision = dp.gate_apply_restart(proof=None, now=NOW, secret=SECRET)
blob = dp._canonical(decision.incident)
self.assertNotIn(SECRET.decode(), blob)
if __name__ == "__main__":
unittest.main()
+199
View File
@@ -0,0 +1,199 @@
"""Regression tests for #628 building blocks (child scope only).
Honest scope: unit coverage of pre-existing APIs used by umbrella #628.
This module does **not** implement or verify all 21 umbrella acceptance
criteria, automatic handoff store/retrieve, multi-worker product wiring,
or end-to-end orchestration.
Covered building blocks:
- CTH format / parse / assess (`format_cth_body`, `parse_cth_comment`,
`assess_cth_comment`)
- Exclusive-ownership skip classification (`classify_skip` with
OWNERSHIP_FOREIGN vs OWNERSHIP_OWN)
- Durable dependency edges (`upsert_dependency_edge` / list) and skip
when dependency_unmet
- Edge state transition UNMET -> MET
Parent umbrella remains #628; this slice is a scoped child issue only.
"""
import unittest
from unittest.mock import MagicMock, patch
import os
import json
import tempfile
from canonical_thread_handoff import (
format_cth_body,
parse_cth_comment,
assess_cth_comment,
)
import dependency_graph
from control_plane_db import ControlPlaneDB
from allocator_service import (
WorkCandidate,
classify_skip,
ROLE_AUTHOR,
ROLE_REVIEWER,
ROLE_MERGER,
ROLE_RECONCILER,
OWNERSHIP_OWN,
OWNERSHIP_FOREIGN,
)
class TestIssue628Orchestration(unittest.TestCase):
def setUp(self):
self._tmp = tempfile.TemporaryDirectory()
self.db_path = os.path.join(self._tmp.name, "cp.sqlite3")
self.db = ControlPlaneDB(self.db_path)
def tearDown(self):
self._tmp.cleanup()
def test_canonical_handoff_serialization_and_retrieval(self):
"""Building block: format/parse/assess a CTH body (not full AC1/AC2 product path)."""
handoff = format_cth_body(
cth_type="Author Handoff",
status="completed",
next_owner="reviewer",
current_blocker="none",
decision="Implementation complete, tests passing",
proof="pytest tests/test_issue_628_orchestration.py passed",
next_action="Review PR and run reviewer pre-flight",
ready_to_paste_prompt="Review PR for child issue #878 (parent #628)",
)
self.assertIn("CTH: Author Handoff", handoff)
parsed = parse_cth_comment(handoff)
self.assertIsNotNone(parsed)
self.assertEqual(parsed["cth_type"], "Author Handoff")
assessment = assess_cth_comment(handoff)
self.assertFalse(assessment["block"])
def test_exclusive_task_unit_single_owner(self):
"""Building block: classify_skip foreign vs own ownership."""
candidate = WorkCandidate(
kind="issue",
number=878,
title="Child #878 ownership classify candidate",
state="open",
labels=["status:in-progress"],
blocked=False,
dependency_unmet=False,
)
# Foreign ownership MUST be skipped
skip_foreign = classify_skip(
c=candidate,
role=ROLE_AUTHOR,
terminal_pr=None,
claim_ownership=OWNERSHIP_FOREIGN,
)
self.assertIsNotNone(skip_foreign)
self.assertIn("active lease", skip_foreign)
# Own/Self claim remains selectable for session resumption
skip_self = classify_skip(
c=candidate,
role=ROLE_AUTHOR,
terminal_pr=None,
claim_ownership=OWNERSHIP_OWN,
)
self.assertIsNone(skip_self)
def test_durable_dependency_graph_blocking(self):
"""Building block: unmet dependency edge + classify_skip on dependency_unmet."""
# Upsert a blocking dependency edge between issue 878 and blocker 601
self.db.upsert_dependency_edge(
remote="prgs",
org="Scaled-Tech-Consulting",
repo="Gitea-Tools",
source_kind="issue",
source_number=878,
target_kind="issue",
target_number=601,
edge_type=dependency_graph.EDGE_ISSUE_BLOCKED_BY_ISSUE,
state=dependency_graph.STATE_UNMET,
blocking_condition="Target issue #601 is not closed",
completion_condition="Target issue #601 is closed",
evidence={"source": "unit_test"},
)
edges = self.db.list_dependency_edges(
remote="prgs",
org="Scaled-Tech-Consulting",
repo="Gitea-Tools",
source_kind="issue",
source_number=878,
)
self.assertEqual(len(edges), 1)
self.assertEqual(edges[0]["state"], "unmet")
self.assertEqual(edges[0]["target_number"], 601)
# When dependency is unmet, candidate is blocked from selection
candidate = WorkCandidate(
kind="issue",
number=878,
title="Blocked candidate",
state="open",
labels=[],
blocked=False,
dependency_unmet=True,
dependency_reason="issue#878 is blocked by unmet dependency issue#601",
)
skip_reason = classify_skip(
c=candidate,
role=ROLE_AUTHOR,
terminal_pr=None,
claim_ownership=OWNERSHIP_OWN,
)
self.assertIsNotNone(skip_reason)
self.assertIn("issue#601", skip_reason)
def test_dependency_completion_reevaluation(self):
"""Building block: dependency edge state can transition UNMET -> MET."""
self.db.upsert_dependency_edge(
remote="prgs",
org="Scaled-Tech-Consulting",
repo="Gitea-Tools",
source_kind="issue",
source_number=878,
target_kind="issue",
target_number=601,
edge_type=dependency_graph.EDGE_ISSUE_BLOCKED_BY_ISSUE,
state=dependency_graph.STATE_UNMET,
blocking_condition="Target issue #601 is open",
completion_condition="Target issue #601 is closed",
evidence={"source": "unit_test"},
)
# Mark edge as met upon target issue closure
self.db.upsert_dependency_edge(
remote="prgs",
org="Scaled-Tech-Consulting",
repo="Gitea-Tools",
source_kind="issue",
source_number=878,
target_kind="issue",
target_number=601,
edge_type=dependency_graph.EDGE_ISSUE_BLOCKED_BY_ISSUE,
state=dependency_graph.STATE_MET,
blocking_condition="Target issue #601 is open",
completion_condition="Target issue #601 is closed",
evidence={"source": "target_closed_event"},
)
edges = self.db.list_dependency_edges(
remote="prgs",
org="Scaled-Tech-Consulting",
repo="Gitea-Tools",
source_kind="issue",
source_number=878,
)
self.assertEqual(len(edges), 1)
self.assertEqual(edges[0]["state"], "met")
if __name__ == "__main__":
unittest.main()
@@ -0,0 +1,249 @@
"""Tests for post-restart MCP reconciliation and completion proof (#662)."""
from __future__ import annotations
import json
import os
import unittest
from datetime import datetime, timezone
import post_restart_reconcile as prr
NOW = datetime(2026, 7, 24, 12, 0, 0, tzinfo=timezone.utc)
def _base_inventory(**overrides):
inv = {
"inventory_complete": True,
"incomplete_reasons": [],
"service_health": {"healthy": True},
"clients": [{"session_id": "c1", "connected": True}],
"sessions": [
{
"session_id": "s-live",
"status": "active",
"pid": os.getpid(),
"role": "author",
}
],
"leases": [],
"checkpoints_available": False,
"worktree_bindings": [{"path": "/tmp/wt", "exists": True}],
"pending_mutations": [],
"capabilities": {"stale": False},
"boot_head_sha": "a" * 40,
"current_head_sha": "a" * 40,
"queue_state": {"safe_to_resume": True},
}
inv.update(overrides)
return inv
class IncompleteInventoryTest(unittest.TestCase):
def test_incomplete_inventory_fails_closed(self) -> None:
proof = prr.reconcile_after_restart(
{
"inventory_complete": False,
"incomplete_reasons": ["control-plane DB unavailable"],
},
now=NOW,
mode=prr.MODE_ENFORCE,
reconcile_id="test-incomplete",
)
self.assertEqual(proof.overall_status, prr.STATUS_FAILED)
self.assertTrue(proof.mutation_hold)
self.assertFalse(proof.inventory_complete)
self.assertTrue(proof.proposed_follow_ups)
self.assertIn("control-plane DB unavailable", proof.incomplete_reasons)
class HappyPathTest(unittest.TestCase):
def test_clean_restart_is_complete_without_mutation_hold(self) -> None:
proof = prr.reconcile_after_restart(
_base_inventory(),
now=NOW,
mode=prr.MODE_ENFORCE,
reconcile_id="test-clean",
)
self.assertEqual(proof.overall_status, prr.STATUS_COMPLETE)
self.assertFalse(proof.mutation_hold)
self.assertEqual(proof.unresolved_count, 0)
cp = next(i for i in proof.items if i.dimension == prr.DIM_CHECKPOINTS)
self.assertEqual(cp.status, prr.ITEM_SKIPPED)
links = proof.as_dict()["links"]
self.assertEqual(links["umbrella"], 655)
self.assertEqual(links["issue"], 662)
self.assertEqual(links["vision"], 652)
self.assertEqual(links["roadmap"], 653)
class InterruptedMutationTest(unittest.TestCase):
def test_mutating_lease_with_dead_owner_is_unresolved(self) -> None:
proof = prr.reconcile_after_restart(
_base_inventory(
leases=[
{
"lease_id": "lease-mut",
"session_id": "s-dead",
"phase": "implementing",
"work_kind": "issue",
"work_number": 662,
"worktree_path": "/tmp/wt-662",
"freshness": {"freshness": "stale_dead_process"},
}
]
),
now=NOW,
mode=prr.MODE_ENFORCE,
)
mut = next(i for i in proof.items if i.dimension == prr.DIM_MUTATIONS)
self.assertEqual(mut.status, prr.ITEM_UNRESOLVED)
self.assertTrue(mut.follow_up_required)
interrupted = mut.details["interrupted"]
self.assertEqual(len(interrupted), 1)
self.assertFalse(interrupted[0]["resume_allowed"])
self.assertTrue(proof.mutation_hold)
self.assertTrue(
any(f.dimension == prr.DIM_MUTATIONS for f in proof.proposed_follow_ups)
)
def test_explicit_pending_mutation_inventory(self) -> None:
proof = prr.reconcile_after_restart(
_base_inventory(
pending_mutations=[
{
"session_id": "s1",
"phase": "publishing",
"work_kind": "pr",
"work_number": 856,
"reason": "push interrupted mid-flight",
}
]
),
now=NOW,
mode=prr.MODE_LOG_ONLY,
)
mut = next(i for i in proof.items if i.dimension == prr.DIM_MUTATIONS)
self.assertEqual(mut.status, prr.ITEM_UNRESOLVED)
# log_only never holds mutations even when unresolved
self.assertFalse(proof.mutation_hold)
self.assertEqual(proof.overall_status, prr.STATUS_DEGRADED)
class DuplicateClaimsTest(unittest.TestCase):
def test_duplicate_live_claims_flagged(self) -> None:
proof = prr.reconcile_after_restart(
_base_inventory(
leases=[
{
"lease_id": "l1",
"session_id": "s1",
"phase": "allocated",
"work_kind": "issue",
"work_number": 100,
"freshness": {"freshness": "active"},
},
{
"lease_id": "l2",
"session_id": "s2",
"phase": "allocated",
"work_kind": "issue",
"work_number": 100,
"freshness": {"freshness": "active"},
},
]
),
now=NOW,
mode=prr.MODE_ENFORCE,
)
dups = next(i for i in proof.items if i.dimension == prr.DIM_DUPLICATES)
self.assertEqual(dups.status, prr.ITEM_UNRESOLVED)
self.assertEqual(dups.details["duplicates"][0]["claim_count"], 2)
self.assertTrue(proof.mutation_hold)
class OrphanSessionTest(unittest.TestCase):
def test_active_session_dead_pid_is_unresolved(self) -> None:
proof = prr.reconcile_after_restart(
_base_inventory(
sessions=[
{
"session_id": "ghost",
"status": "active",
"pid": 2_000_000_000,
"role": "author",
}
]
),
now=NOW,
mode=prr.MODE_ENFORCE,
)
sess = next(i for i in proof.items if i.dimension == prr.DIM_SESSIONS)
self.assertEqual(sess.status, prr.ITEM_UNRESOLVED)
self.assertIn("ghost", sess.details["orphan_session_ids"])
class CapabilityStaleTest(unittest.TestCase):
def test_stale_runtime_unresolved(self) -> None:
proof = prr.reconcile_after_restart(
_base_inventory(capabilities={"stale": True, "startup_head": "aaa"}),
now=NOW,
mode=prr.MODE_ENFORCE,
)
caps = next(i for i in proof.items if i.dimension == prr.DIM_CAPABILITIES)
self.assertEqual(caps.status, prr.ITEM_UNRESOLVED)
self.assertTrue(proof.mutation_hold)
class MutationsAllowedHelperTest(unittest.TestCase):
def test_mutations_allowed_respects_hold(self) -> None:
held = prr.reconcile_after_restart(
_base_inventory(
pending_mutations=[{"phase": "merging", "session_id": "x"}]
),
now=NOW,
mode=prr.MODE_ENFORCE,
)
self.assertFalse(prr.mutations_allowed(held))
self.assertFalse(prr.mutations_allowed(held.as_dict()))
clean = prr.reconcile_after_restart(
_base_inventory(), now=NOW, mode=prr.MODE_ENFORCE
)
self.assertTrue(prr.mutations_allowed(clean))
class CheckpointSoftDependencyTest(unittest.TestCase):
def test_checkpoints_when_schema_present(self) -> None:
proof = prr.reconcile_after_restart(
_base_inventory(
checkpoints_available=True,
checkpoints=[{"session_id": "s1", "stale": False}],
),
now=NOW,
)
cp = next(i for i in proof.items if i.dimension == prr.DIM_CHECKPOINTS)
self.assertEqual(cp.status, prr.ITEM_RESOLVED)
def test_stale_checkpoints_unresolved(self) -> None:
proof = prr.reconcile_after_restart(
_base_inventory(
checkpoints_available=True,
checkpoints=[{"session_id": "s1", "stale": True}],
),
now=NOW,
mode=prr.MODE_ENFORCE,
)
cp = next(i for i in proof.items if i.dimension == prr.DIM_CHECKPOINTS)
self.assertEqual(cp.status, prr.ITEM_UNRESOLVED)
class ProofSerializationTest(unittest.TestCase):
def test_as_dict_is_json_friendly(self) -> None:
proof = prr.reconcile_after_restart(_base_inventory(), now=NOW)
blob = json.dumps(proof.as_dict())
self.assertIn("reconcile_id", blob)
self.assertIn("proposed_follow_ups", blob)
if __name__ == "__main__":
unittest.main()
+444
View File
@@ -0,0 +1,444 @@
"""Task heartbeat through the native MCP author path (#790 Slice A, AC-N6).
Assessor-level coverage is not sufficient here, and this project has already
paid for learning that: in review #499 on PR #791 the #760 renewal waiver was
computed correctly and then *discarded* at two later gates, so every real
renewal still failed while the unit suite stayed green. AC-N6 exists because of
that, and requires driving the real tools against a real git repository and a
real durable lock file, composing the gates in production order.
These tests therefore call ``gitea_lock_issue`` and
``gitea_heartbeat_issue_lock`` themselves and assert on what lands on disk,
never on an assessor's return value alone.
"""
from __future__ import annotations
import os
import subprocess
import sys
import tempfile
import unittest
from datetime import datetime, timedelta, timezone
from unittest.mock import patch
sys.path.insert(0, os.path.dirname(os.path.dirname(os.path.abspath(__file__))))
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
from mutation_profile_fixture import shared_mutation_env # noqa: E402
import issue_lock_provenance # noqa: E402
import issue_lock_store # noqa: E402
import lease_policy # noqa: E402
import mcp_server # noqa: E402
ISSUE = 9791
BRANCH = f"fix/issue-{ISSUE}-heartbeat-mcp"
IDENTITY = "example-user"
PROFILE = "test-author-prgs"
ORG = "Scaled-Tech-Consulting"
REPO = "Gitea-Tools"
def _ts(moment: datetime) -> str:
return (
moment.astimezone(timezone.utc)
.replace(microsecond=0)
.isoformat()
.replace("+00:00", "Z")
)
class _HeartbeatMcpBase(unittest.TestCase):
"""Real git repo plus a real durable lock, driven through the real tools."""
def setUp(self):
self.lock_dir = tempfile.TemporaryDirectory()
self.addCleanup(self.lock_dir.cleanup)
self.repo = tempfile.mkdtemp(prefix="issue790-mcp-")
self.addCleanup(lambda: subprocess.run(["rm", "-rf", self.repo], check=False))
self._init_worktree()
self.remotes = patch.dict(
mcp_server.REMOTES,
{"prgs": {"host": "gitea.prgs.cc", "org": ORG, "repo": REPO}},
)
self.remotes.start()
self.addCleanup(patch.stopall)
mcp_server._IDENTITY_CACHE.clear()
def _git(self, *args):
return subprocess.run(
["git", "-C", self.repo, *args], capture_output=True, text=True, check=True
)
def _init_worktree(self):
self._git("init", "-q", "-b", "master")
self._git("config", "user.email", "[email protected]")
self._git("config", "user.name", "Test")
with open(os.path.join(self.repo, "seed.txt"), "w") as fh:
fh.write("seed\n")
self._git("add", "seed.txt")
self._git("commit", "-q", "-m", "seed")
self.base_sha = self._git("rev-parse", "HEAD").stdout.strip()
# A fresh claim starts base-equivalent, which is the ordinary first-lock
# shape and exercises assess_issue_lock_worktree on its normal path.
self._git("checkout", "-q", "-b", BRANCH)
self.head_sha = self.base_sha
self.worktree = os.path.realpath(self.repo)
def _lock_path(self):
return issue_lock_store.lock_file_path(
remote="prgs",
org=ORG,
repo=REPO,
issue_number=ISSUE,
lock_dir=self.lock_dir.name,
)
def _tool_env(self):
env = shared_mutation_env(
PROFILE, include_example_repo=True, GITEA_ISSUE_LOCK_DIR=self.lock_dir.name
)
env["GITEA_ISSUE_LOCK_DIR"] = self.lock_dir.name
return env
def _git_state(self, *, porcelain="", base_equivalent=True):
return {
"current_branch": BRANCH,
"porcelain_status": porcelain,
"base_equivalent": base_equivalent,
"head_sha": self.head_sha,
"inspected_git_root": self.worktree,
"base_branch": "master",
}
def run_lock_issue(
self,
*,
branch_entries=None,
open_prs=None,
git_state=None,
identity=IDENTITY,
profile=PROFILE,
):
branch_entries = branch_entries if branch_entries is not None else []
open_prs = open_prs if open_prs is not None else []
git_state = git_state or self._git_state()
env = self._tool_env()
with patch(
"mcp_server.api_get_all", return_value=list(branch_entries)
), patch(
"mcp_server._list_open_pulls", return_value=list(open_prs)
), patch(
"mcp_server.get_auth_header", return_value="token x"
), patch(
"mcp_server._work_lease_claimant",
return_value={"username": identity, "profile": profile},
), patch(
"mcp_server.issue_lock_worktree.read_worktree_git_state",
return_value=git_state,
), patch(
"mcp_server.issue_duplicate_context_fetcher",
side_effect=lambda h, o, r, auth, issue_number: (
list(open_prs),
[b.get("name") for b in branch_entries if isinstance(b, dict)],
{"status": "not_claimed"},
),
), patch.dict(os.environ, env, clear=True):
os.environ["GITEA_ISSUE_LOCK_DIR"] = self.lock_dir.name
return mcp_server.gitea_lock_issue(
issue_number=ISSUE,
branch_name=BRANCH,
remote="prgs",
worktree_path=self.worktree,
)
def run_heartbeat(
self, *, task_session_id, identity=IDENTITY, profile=PROFILE, **kwargs
):
env = self._tool_env()
with patch(
"mcp_server._work_lease_claimant",
return_value={"username": identity, "profile": profile},
), patch("mcp_server.get_auth_header", return_value="token x"), patch.dict(
os.environ, env, clear=True
):
os.environ["GITEA_ISSUE_LOCK_DIR"] = self.lock_dir.name
return mcp_server.gitea_heartbeat_issue_lock(
issue_number=ISSUE,
branch_name=kwargs.pop("branch_name", BRANCH),
task_session_id=task_session_id,
remote="prgs",
worktree_path=kwargs.pop("worktree_path", self.worktree),
**kwargs,
)
def write_legacy_lock(self, *, hours_old: float = 3.0, ttl_hours: float = 4.0):
"""A durable lock in the shape the store wrote before this slice."""
now = datetime.now(timezone.utc)
claimant = {"username": IDENTITY, "profile": PROFILE}
created = now - timedelta(hours=hours_old)
record = {
"issue_number": ISSUE,
"branch_name": BRANCH,
"remote": "prgs",
"org": ORG,
"repo": REPO,
"worktree_path": self.worktree,
"session_pid": os.getpid(),
"pid": os.getpid(),
"lock_generation": 1,
"work_lease": {
"operation_type": issue_lock_store.AUTHOR_ISSUE_WORK_LEASE,
"issue_number": ISSUE,
"pr_number": None,
"branch": BRANCH,
"worktree_path": self.worktree,
"claimant": claimant,
"created_at": _ts(created),
# The legacy signature: never advanced past creation.
"last_heartbeat_at": _ts(created),
"expires_at": _ts(created + timedelta(hours=ttl_hours)),
},
"lock_provenance": issue_lock_provenance.build_sanctioned_lock_provenance(
tool="gitea_lock_issue", claimant=claimant
),
}
path = self._lock_path()
record["lock_file_path"] = path
issue_lock_store.save_lock_file(path, record)
return record
class TestLockIssueMintsTheLifecycle(_HeartbeatMcpBase):
"""Durable lock creation and read-back through the real tool."""
def test_native_lock_writes_the_marker_and_a_task_session_id(self):
result = self.run_lock_issue()
self.assertTrue(result["success"], result)
written = issue_lock_store.read_lock_file(result["lock_file_path"])
lease = written["work_lease"]
self.assertEqual(
lease["lifecycle_version"], lease_policy.LIFECYCLE_HEARTBEAT_V1
)
self.assertTrue(lease["task_session_id"])
self.assertFalse(issue_lock_store.is_legacy_lease(written))
# AC-N1: the ownership key is not the daemon pid, which is recorded
# separately as evidence.
self.assertNotIn(str(written["session_pid"]), lease["task_session_id"])
self.assertEqual(written["session_pid"], os.getpid())
def test_native_lease_uses_the_policy_window_not_four_hours(self):
result = self.run_lock_issue()
lease = result["work_lease"]
created = datetime.fromisoformat(lease["created_at"].replace("Z", "+00:00"))
expires = datetime.fromisoformat(lease["expires_at"].replace("Z", "+00:00"))
policy = lease_policy.policy_for(lease_policy.TASK_CLASS_AUTHOR_ISSUE_WORK)
self.assertEqual(
(expires - created).total_seconds() / 60.0, policy.initial_ttl_minutes
)
def test_freshness_of_a_new_native_lock_is_live(self):
result = self.run_lock_issue()
self.assertEqual(
result["lock_freshness"]["status"], issue_lock_store.STATUS_LIVE
)
self.assertTrue(result["lock_freshness"]["live"])
class TestHeartbeatThroughTheTool(_HeartbeatMcpBase):
def _lock_and_session(self):
result = self.run_lock_issue()
self.assertTrue(result["success"], result)
return result, result["work_lease"]["task_session_id"]
def test_heartbeat_slides_the_lease_and_advances_the_generation(self):
locked, session = self._lock_and_session()
before = issue_lock_store.read_lock_file(locked["lock_file_path"])
beat = self.run_heartbeat(task_session_id=session)
self.assertTrue(beat["success"], beat)
self.assertEqual(beat["operation"], "heartbeat")
after = issue_lock_store.read_lock_file(locked["lock_file_path"])
self.assertGreater(
issue_lock_store.lock_generation(after),
issue_lock_store.lock_generation(before),
)
self.assertGreaterEqual(
after["work_lease"]["expires_at"], before["work_lease"]["expires_at"]
)
self.assertEqual(after["work_lease"]["heartbeat_count"], 2)
def test_heartbeat_evidence_survives_the_downstream_mutation_gate(self):
"""The #499 F2 lesson, applied.
A sanction that is computed and then discarded downstream is worthless.
After a heartbeat the lock must still satisfy the gate every author
mutation runs through.
"""
locked, session = self._lock_and_session()
self.run_heartbeat(task_session_id=session)
written = issue_lock_store.read_lock_file(locked["lock_file_path"])
verdict = issue_lock_store.verify_lock_for_mutation(
written,
issue_number=ISSUE,
branch_name=BRANCH,
worktree_path=self.worktree,
)
self.assertTrue(verdict["proven"], verdict)
self.assertFalse(verdict["block"])
def _duplicate_gate(self, *, open_prs, branches):
env = self._tool_env()
with patch("mcp_server.get_auth_header", return_value="token x"), patch(
"mcp_server.issue_duplicate_context_fetcher",
side_effect=lambda h, o, r, auth, issue_number: (
list(open_prs),
list(branches),
{"status": "not_claimed"},
),
), patch.dict(os.environ, env, clear=True):
os.environ["GITEA_ISSUE_LOCK_DIR"] = self.lock_dir.name
return mcp_server.gitea_assess_work_issue_duplicate(
issue_number=ISSUE, branch_name=BRANCH, remote="prgs"
)
def test_heartbeat_does_not_change_the_duplicate_gate_verdict(self):
"""The gate must be invariant under heartbeating.
The point is not that the gate passes with a linked open PR at the
lock phase it correctly blocks (#400), heartbeat or not. The property
that matters is that sliding a lease neither loosens the gate nor
corrupts the lock state it reads: the verdict before and after a
heartbeat must be identical, for both the clear and the blocking shape.
"""
_, session = self._lock_and_session()
linked = [{"number": 4242, "head": {"ref": BRANCH, "sha": self.head_sha}}]
clear_before = self._duplicate_gate(open_prs=[], branches=[])
blocked_before = self._duplicate_gate(open_prs=linked, branches=[BRANCH])
self.assertTrue(self.run_heartbeat(task_session_id=session)["success"])
clear_after = self._duplicate_gate(open_prs=[], branches=[])
blocked_after = self._duplicate_gate(open_prs=linked, branches=[BRANCH])
self.assertEqual(clear_before["outcome"], clear_after["outcome"])
self.assertFalse(clear_after["block"])
self.assertEqual(blocked_before["outcome"], blocked_after["outcome"])
self.assertTrue(blocked_after["block"])
self.assertEqual(blocked_after["linked_open_pr"], 4242)
def test_foreign_session_id_is_refused_through_the_tool(self):
self._lock_and_session()
beat = self.run_heartbeat(task_session_id="author_issue_work-ffffffffffffffff")
self.assertFalse(beat["success"])
self.assertIn("task_session_id does not match", " ".join(beat["reasons"]))
def test_stale_generation_is_refused_through_the_tool(self):
locked, session = self._lock_and_session()
current = issue_lock_store.lock_generation(
issue_lock_store.read_lock_file(locked["lock_file_path"])
)
beat = self.run_heartbeat(
task_session_id=session, expected_generation=current + 5
)
self.assertFalse(beat["success"])
self.assertIn("generation changed", beat["reasons"][0])
def test_foreign_claimant_is_refused_through_the_tool(self):
_, session = self._lock_and_session()
beat = self.run_heartbeat(task_session_id=session, identity="someone-else")
self.assertFalse(beat["success"])
def test_heartbeat_cannot_acquire_a_missing_lock(self):
beat = self.run_heartbeat(task_session_id="author_issue_work-000000000000")
self.assertFalse(beat["success"])
self.assertIn("no durable lock", beat["reasons"][0])
def test_alive_pid_alone_does_not_keep_a_lease_live_through_the_tool(self):
"""PID-only refusal, end to end.
The recorded pid is this live process. The lock is aged past its grace
with no heartbeat, so the tool must refuse to slide it and the durable
record must classify as a missed heartbeat rather than as live.
"""
locked, session = self._lock_and_session()
record = issue_lock_store.read_lock_file(locked["lock_file_path"])
record["work_lease"]["last_heartbeat_at"] = _ts(
datetime.now(timezone.utc) - timedelta(minutes=30)
)
record["work_lease"]["expires_at"] = _ts(
datetime.now(timezone.utc) + timedelta(hours=2)
)
issue_lock_store.save_lock_file(locked["lock_file_path"], record)
self.assertTrue(issue_lock_store.is_process_alive(record["session_pid"]))
fresh = issue_lock_store.assess_lock_freshness(record)
self.assertEqual(
fresh["status"], issue_lock_store.STATUS_STALE_MISSED_HEARTBEAT
)
self.assertTrue(fresh["pid_alive"])
beat = self.run_heartbeat(task_session_id=session)
self.assertFalse(beat["success"])
self.assertIn("reclaimed", " ".join(beat["reasons"]))
class TestLegacyLocksThroughTheTool(_HeartbeatMcpBase):
"""AC-N8 end to end: protected on deployment, and rebindable."""
def test_legacy_lock_stays_protected_after_deployment(self):
record = self.write_legacy_lock(hours_old=3.0, ttl_hours=4.0)
fresh = issue_lock_store.assess_lock_freshness(record)
self.assertEqual(fresh["status"], issue_lock_store.STATUS_LIVE)
self.assertTrue(fresh["legacy_lease"])
self.assertTrue(fresh["legacy_expiry_preserved"])
# It had never heartbeated, so under the new grace alone it would be
# long gone; the preserved absolute expiry is what protects it.
self.assertEqual(
record["work_lease"]["created_at"],
record["work_lease"]["last_heartbeat_at"],
)
def test_tool_rebinds_a_legacy_lock_and_mints_a_first_heartbeat(self):
self.write_legacy_lock(hours_old=3.0, ttl_hours=4.0)
result = self.run_heartbeat(task_session_id=None)
self.assertTrue(result["success"], result)
self.assertEqual(result["operation"], "legacy_rebind")
self.assertTrue(result["task_session_id"])
written = issue_lock_store.read_lock_file(self._lock_path())
lease = written["work_lease"]
self.assertEqual(
lease["lifecycle_version"], lease_policy.LIFECYCLE_HEARTBEAT_V1
)
self.assertEqual(lease["heartbeat_count"], 1)
self.assertNotEqual(
lease["created_at"],
written["legacy_rebind"]["legacy_origin"]["created_at"],
)
self.assertFalse(issue_lock_store.is_legacy_lease(written))
def test_rebound_lock_then_heartbeats_through_the_tool(self):
self.write_legacy_lock(hours_old=3.0, ttl_hours=4.0)
rebound = self.run_heartbeat(task_session_id=None)
beat = self.run_heartbeat(task_session_id=rebound["task_session_id"])
self.assertTrue(beat["success"], beat)
self.assertEqual(beat["operation"], "heartbeat")
self.assertEqual(beat["heartbeat_count"], 2)
def test_rebind_refuses_a_foreign_owner_through_the_tool(self):
self.write_legacy_lock(hours_old=3.0, ttl_hours=4.0)
result = self.run_heartbeat(task_session_id=None, identity="someone-else")
self.assertFalse(result["success"])
self.assertEqual(result["operation"], "legacy_rebind")
if __name__ == "__main__":
unittest.main()
+594
View File
@@ -0,0 +1,594 @@
"""Central lease policy and load-bearing heartbeat freshness (#790 Slice A).
Before this slice, ``issue_lock_store.assess_lock_freshness`` parsed
``last_heartbeat_at`` and then never consulted it: liveness was decided by an
absolute four-hour ``expires_at`` and by PID liveness. Because the recorded PID
is the long-lived MCP daemon rather than the authoring task, an abandoned claim
stayed "live" for the full four hours, and a claim whose work had already landed
blocked reconciliation for just as long (Issue #787 / PR #789, and again Issue
#760 / PR #791).
These tests pin the corrected semantics, including the two asymmetries that are
easy to lose in a refactor:
* an **alive** PID must never make anything live (AC-N2), while
* a **dead** PID must still mark a lease stale, because #753 dead-session
recovery keys on exactly that classification.
Durable-state helpers here write real lock files through the real flock path;
they are not mocks of the store.
"""
from __future__ import annotations
import os
import sys
import tempfile
import unittest
from datetime import datetime, timedelta, timezone
from unittest.mock import patch
sys.path.insert(0, os.path.dirname(os.path.dirname(os.path.abspath(__file__))))
import issue_lock_store # noqa: E402
import lease_policy # noqa: E402
import pr_work_lease # noqa: E402
import reviewer_pr_lease # noqa: E402
ISSUE = 9790
BRANCH = f"fix/issue-{ISSUE}-heartbeat"
IDENTITY = "example-user"
PROFILE = "test-author-prgs"
ORG = "Example-Org"
REPO = "Example-Repo"
REMOTE = "prgs"
DEAD_PID = 2**22 # far above any live pid on a test host
def _ts(moment: datetime) -> str:
return (
moment.astimezone(timezone.utc)
.replace(microsecond=0)
.isoformat()
.replace("+00:00", "Z")
)
class _LockFixture(unittest.TestCase):
def setUp(self):
self.lock_dir = tempfile.TemporaryDirectory()
self.addCleanup(self.lock_dir.cleanup)
self.now = datetime.now(timezone.utc)
self.worktree = os.path.realpath(tempfile.mkdtemp(prefix="issue790-"))
self.addCleanup(patch.stopall)
def _path(self):
return issue_lock_store.lock_file_path(
remote=REMOTE,
org=ORG,
repo=REPO,
issue_number=ISSUE,
lock_dir=self.lock_dir.name,
)
def write_lock(
self,
*,
lifecycle: str | None = lease_policy.LIFECYCLE_HEARTBEAT_V1,
created_delta: timedelta = timedelta(minutes=1),
heartbeat_delta: timedelta = timedelta(minutes=1),
expires_delta: timedelta = timedelta(minutes=9),
pid: int | None = None,
task_session_id: str | None = "author_issue_work-aaaabbbbccccdddd",
generation: int = 1,
identity: str = IDENTITY,
profile: str = PROFILE,
branch: str = BRANCH,
worktree: str | None = None,
) -> dict:
"""Write a real durable lock and return the record.
Deltas are relative to ``self.now``; ``expires_delta`` is added, the
others subtracted, so "in the past" reads naturally at each call site.
"""
lease: dict = {
"operation_type": issue_lock_store.AUTHOR_ISSUE_WORK_LEASE,
"issue_number": ISSUE,
"pr_number": None,
"branch": branch,
"worktree_path": worktree or self.worktree,
"claimant": {"username": identity, "profile": profile},
"created_at": _ts(self.now - created_delta),
"last_heartbeat_at": _ts(self.now - heartbeat_delta),
"expires_at": _ts(self.now + expires_delta),
}
if lifecycle is not None:
lease["lifecycle_version"] = lifecycle
if task_session_id is not None:
lease["task_session_id"] = task_session_id
pid_value = os.getpid() if pid is None else pid
record = {
"issue_number": ISSUE,
"branch_name": branch,
"remote": REMOTE,
"org": ORG,
"repo": REPO,
"worktree_path": worktree or self.worktree,
"session_pid": pid_value,
"pid": pid_value,
"lock_generation": generation,
"work_lease": lease,
}
path = self._path()
record["lock_file_path"] = path
issue_lock_store.save_lock_file(path, record)
return record
class TestPolicyIsTheSingleSource(unittest.TestCase):
"""AC-N7: one authoritative configuration source for every duration."""
def test_author_policy_carries_the_agreed_values(self):
policy = lease_policy.policy_for(lease_policy.TASK_CLASS_AUTHOR_ISSUE_WORK)
self.assertEqual(policy.initial_ttl_minutes, 10.0)
self.assertEqual(policy.heartbeat_cadence_minutes, 2.0)
self.assertEqual(policy.stale_warning_minutes, 5.0)
self.assertEqual(policy.missed_heartbeat_grace_minutes, 10.0)
self.assertEqual(policy.absolute_cap_hours, 8.0)
self.assertEqual(policy.recovery_grace_minutes, 10.0)
self.assertEqual(policy.terminal_race_drain_minutes, 2.0)
self.assertTrue(policy.terminal_retirement_eligible)
self.assertTrue(policy.heartbeat_lifecycle_active)
def test_the_four_hour_author_ttl_literal_is_gone(self):
"""The duplicated literal AC-N7 exists to remove."""
self.assertFalse(hasattr(issue_lock_store, "WORK_LEASE_TTL_HOURS"))
import gitea_mcp_server
self.assertFalse(hasattr(gitea_mcp_server, "WORK_LEASE_TTL_HOURS"))
def test_declared_reviewer_values_match_the_module_still_using_them(self):
"""Slice A declares reviewer/merger numbers without rewiring them.
Recording a value in two places is only safe if drift is detectable, so
this asserts the declaration still equals the constants #747 owns. When
Slice C migrates those call sites, this test becomes the proof the
migration changed nothing.
"""
policy = lease_policy.policy_for(lease_policy.TASK_CLASS_REVIEWER_PR)
self.assertEqual(
policy.initial_ttl_minutes, float(reviewer_pr_lease.LEASE_TTL_MINUTES)
)
self.assertEqual(
policy.stale_warning_minutes,
float(reviewer_pr_lease.STALE_WARNING_MINUTES),
)
self.assertFalse(policy.heartbeat_lifecycle_active)
def test_declared_conflict_fix_value_matches_its_module(self):
policy = lease_policy.policy_for(lease_policy.TASK_CLASS_CONFLICT_FIX)
self.assertEqual(
policy.initial_ttl_minutes,
float(pr_work_lease.DEFAULT_CONFLICT_FIX_TTL_MINUTES),
)
self.assertFalse(policy.heartbeat_lifecycle_active)
def test_environment_override_applies(self):
var = lease_policy.env_var_name(
lease_policy.TASK_CLASS_AUTHOR_ISSUE_WORK, "initial_ttl_minutes"
)
with patch.dict(os.environ, {var: "7"}):
self.assertEqual(
lease_policy.policy_for(
lease_policy.TASK_CLASS_AUTHOR_ISSUE_WORK
).initial_ttl_minutes,
7.0,
)
def test_unusable_override_falls_back_instead_of_minting_a_zero_lease(self):
"""A typo must not make every claim instantly reclaimable."""
var = lease_policy.env_var_name(
lease_policy.TASK_CLASS_AUTHOR_ISSUE_WORK, "initial_ttl_minutes"
)
for bad in ("0", "-5", "not-a-number", " "):
with self.subTest(value=bad), patch.dict(os.environ, {var: bad}):
self.assertEqual(
lease_policy.policy_for(
lease_policy.TASK_CLASS_AUTHOR_ISSUE_WORK
).initial_ttl_minutes,
10.0,
)
def test_unknown_task_class_does_not_raise(self):
policy = lease_policy.policy_for("something-new")
self.assertEqual(policy.task_class, lease_policy.TASK_CLASS_AUTHOR_ISSUE_WORK)
class TestLifecycleDiscrimination(_LockFixture):
"""AC-N8: the marker, never a timestamp, decides legacy vs heartbeat."""
def test_missing_marker_reads_as_legacy(self):
record = self.write_lock(lifecycle=None)
self.assertTrue(issue_lock_store.is_legacy_lease(record))
self.assertEqual(
issue_lock_store.lease_lifecycle_version(record),
lease_policy.LIFECYCLE_LEGACY,
)
def test_marker_present_reads_as_heartbeat_lifecycle(self):
record = self.write_lock()
self.assertFalse(issue_lock_store.is_legacy_lease(record))
def test_equal_created_and_heartbeat_never_implies_a_fresh_heartbeat(self):
"""The exact inversion AC-N8 forbids.
A legacy lock has ``last_heartbeat_at == created_at`` forever because
nothing ever advanced it. Reading that equality as "recently
heartbeated" would classify every never-heartbeated lock as fresh.
"""
legacy = self.write_lock(
lifecycle=None,
created_delta=timedelta(hours=3),
heartbeat_delta=timedelta(hours=3),
)
lease = legacy["work_lease"]
self.assertEqual(lease["created_at"], lease["last_heartbeat_at"])
self.assertTrue(issue_lock_store.is_legacy_lease(legacy))
# A brand-new heartbeat lease has them equal too, so the equality
# carries no information in either direction.
fresh = self.write_lock(
created_delta=timedelta(seconds=0), heartbeat_delta=timedelta(seconds=0)
)
self.assertEqual(
fresh["work_lease"]["created_at"],
fresh["work_lease"]["last_heartbeat_at"],
)
self.assertFalse(issue_lock_store.is_legacy_lease(fresh))
def test_minted_session_id_contains_no_pid(self):
"""AC-N1: the ownership key must not be derived from the daemon pid."""
minted = issue_lock_store.mint_task_session_id()
self.assertNotIn(str(os.getpid()), minted)
self.assertNotEqual(minted, issue_lock_store.mint_task_session_id())
class TestFreshnessIsHeartbeatDriven(_LockFixture):
"""AC-N2 and the new bands."""
def test_fresh_heartbeat_is_live(self):
record = self.write_lock(heartbeat_delta=timedelta(minutes=1))
fresh = issue_lock_store.assess_lock_freshness(record, now=self.now)
self.assertEqual(fresh["status"], issue_lock_store.STATUS_LIVE)
self.assertTrue(fresh["live"])
self.assertFalse(fresh["heartbeat_warning"])
def test_heartbeat_past_warning_is_still_live_but_flagged(self):
record = self.write_lock(heartbeat_delta=timedelta(minutes=6))
fresh = issue_lock_store.assess_lock_freshness(record, now=self.now)
self.assertEqual(fresh["status"], issue_lock_store.STATUS_LIVE)
self.assertTrue(fresh["heartbeat_warning"])
def test_missed_heartbeat_past_grace_is_classified_explicitly(self):
record = self.write_lock(
heartbeat_delta=timedelta(minutes=11),
expires_delta=timedelta(minutes=30),
)
fresh = issue_lock_store.assess_lock_freshness(record, now=self.now)
self.assertEqual(
fresh["status"], issue_lock_store.STATUS_STALE_MISSED_HEARTBEAT
)
self.assertFalse(fresh["live"])
self.assertTrue(fresh["stale"])
def test_alive_pid_never_establishes_freshness(self):
"""The defect in one assertion.
The recorded PID is this very process, so it is unambiguously alive
and the lease is still not live, because the task stopped heartbeating.
"""
record = self.write_lock(
pid=os.getpid(),
heartbeat_delta=timedelta(hours=4),
expires_delta=timedelta(hours=4),
)
fresh = issue_lock_store.assess_lock_freshness(record, now=self.now)
self.assertTrue(fresh["pid_alive"])
self.assertFalse(fresh["live"])
self.assertEqual(
fresh["status"], issue_lock_store.STATUS_STALE_MISSED_HEARTBEAT
)
def test_dead_pid_still_marks_stale_for_issue_753(self):
"""The opposite asymmetry: dead-PID corroboration is preserved."""
record = self.write_lock(pid=DEAD_PID, heartbeat_delta=timedelta(minutes=1))
fresh = issue_lock_store.assess_lock_freshness(record, now=self.now)
self.assertEqual(fresh["status"], issue_lock_store.STATUS_STALE)
self.assertFalse(fresh["live"])
self.assertIn("not alive", fresh["reason"])
def test_absolute_cap_requires_readoption(self):
record = self.write_lock(
created_delta=timedelta(hours=9), heartbeat_delta=timedelta(minutes=1)
)
fresh = issue_lock_store.assess_lock_freshness(record, now=self.now)
self.assertEqual(fresh["status"], issue_lock_store.STATUS_STALE_ABSOLUTE_CAP)
self.assertIn("re-adoption", fresh["reason"])
def test_heartbeat_lifecycle_without_a_heartbeat_fails_closed(self):
record = self.write_lock()
del record["work_lease"]["last_heartbeat_at"]
issue_lock_store.save_lock_file(self._path(), record)
fresh = issue_lock_store.assess_lock_freshness(record, now=self.now)
self.assertEqual(
fresh["status"], issue_lock_store.STATUS_STALE_MISSED_HEARTBEAT
)
self.assertIn("fail closed", fresh["reason"])
def test_absent_lock(self):
fresh = issue_lock_store.assess_lock_freshness(None)
self.assertEqual(fresh["status"], issue_lock_store.STATUS_ABSENT)
self.assertFalse(fresh["stale"])
class TestLegacyLocksStayProtected(_LockFixture):
"""AC-N8: deployment must not retroactively shorten an existing claim."""
def test_legacy_lock_with_a_stale_heartbeat_remains_live(self):
"""The deployment-safety case.
A four-hour legacy lease minted three hours ago has not heartbeated
once. Under the new grace it would be long gone; under its preserved
absolute expiry it is still live, and must stay that way.
"""
record = self.write_lock(
lifecycle=None,
created_delta=timedelta(hours=3),
heartbeat_delta=timedelta(hours=3),
expires_delta=timedelta(hours=1),
)
fresh = issue_lock_store.assess_lock_freshness(record, now=self.now)
self.assertEqual(fresh["status"], issue_lock_store.STATUS_LIVE)
self.assertTrue(fresh["live"])
self.assertTrue(fresh["legacy_lease"])
self.assertTrue(fresh["legacy_expiry_preserved"])
def test_legacy_lock_past_its_absolute_expiry_is_expired_as_before(self):
record = self.write_lock(
lifecycle=None,
created_delta=timedelta(hours=5),
heartbeat_delta=timedelta(hours=5),
expires_delta=timedelta(hours=-1),
)
fresh = issue_lock_store.assess_lock_freshness(record, now=self.now)
self.assertEqual(fresh["status"], issue_lock_store.STATUS_EXPIRED)
def test_legacy_lock_is_never_reclaimed_by_the_heartbeat_band(self):
record = self.write_lock(
lifecycle=None,
created_delta=timedelta(hours=3),
heartbeat_delta=timedelta(hours=3),
expires_delta=timedelta(hours=1),
)
reclaim = issue_lock_store.assess_expired_lock_reclaim(record, now=self.now)
self.assertFalse(reclaim["reclaim_allowed"])
class TestReclaimAfterMissedHeartbeat(_LockFixture):
def test_missed_heartbeat_makes_ownership_reclaimable(self):
record = self.write_lock(
pid=os.getpid(),
heartbeat_delta=timedelta(minutes=15),
expires_delta=timedelta(hours=3),
)
reclaim = issue_lock_store.assess_expired_lock_reclaim(record, now=self.now)
self.assertTrue(reclaim["reclaim_allowed"])
self.assertIn("stale_missed_heartbeat", reclaim["reasons"][0])
def test_live_lease_is_never_reclaimable(self):
record = self.write_lock(heartbeat_delta=timedelta(minutes=1))
reclaim = issue_lock_store.assess_expired_lock_reclaim(record, now=self.now)
self.assertFalse(reclaim["reclaim_allowed"])
def test_dead_pid_reclaim_path_is_unchanged(self):
"""#753 must keep working through its original conditions."""
record = self.write_lock(pid=DEAD_PID, heartbeat_delta=timedelta(minutes=1))
reclaim = issue_lock_store.assess_expired_lock_reclaim(record, now=self.now)
self.assertTrue(reclaim["reclaim_allowed"])
self.assertTrue(reclaim["owner_pid_dead"])
class TestHeartbeatWriter(_LockFixture):
"""A4: flock + CAS + exact verification, and no revival path."""
def _heartbeat(self, **kwargs):
params = {
"remote": REMOTE,
"org": ORG,
"repo": REPO,
"issue_number": ISSUE,
"branch_name": BRANCH,
"worktree_path": self.worktree,
"identity": IDENTITY,
"profile": PROFILE,
"task_session_id": "author_issue_work-aaaabbbbccccdddd",
"lock_dir": self.lock_dir.name,
"now": self.now,
}
params.update(kwargs)
return issue_lock_store.heartbeat_session_lock(**params)
def test_heartbeat_slides_expiry_and_advances_generation(self):
self.write_lock(heartbeat_delta=timedelta(minutes=4), generation=5)
result = self._heartbeat()
self.assertTrue(result["success"], result)
self.assertEqual(result["prior_generation"], 5)
self.assertEqual(result["lock_generation"], 6)
self.assertEqual(result["heartbeat_count"], 1)
self.assertEqual(result["last_heartbeat_at"], _ts(self.now))
self.assertEqual(result["expires_at"], _ts(self.now + timedelta(minutes=10)))
self.assertTrue(result["freshness"]["live"])
def test_heartbeat_is_durable_and_repeatable(self):
self.write_lock(heartbeat_delta=timedelta(minutes=4))
self._heartbeat()
second = self._heartbeat(now=self.now + timedelta(minutes=1))
self.assertTrue(second["success"], second)
self.assertEqual(second["heartbeat_count"], 2)
written = issue_lock_store.read_lock_file(self._path())
self.assertEqual(written["work_lease"]["heartbeat_count"], 2)
def test_stale_generation_is_refused(self):
self.write_lock(generation=5)
result = self._heartbeat(expected_generation=4)
self.assertFalse(result["success"])
self.assertIn("generation changed", result["reasons"][0])
def test_foreign_session_is_refused(self):
self.write_lock()
result = self._heartbeat(task_session_id="author_issue_work-ffffffffffffffff")
self.assertFalse(result["success"])
self.assertIn("task_session_id does not match", " ".join(result["reasons"]))
def test_missing_session_id_is_refused(self):
self.write_lock()
result = self._heartbeat(task_session_id="")
self.assertFalse(result["success"])
def test_foreign_claimant_is_refused(self):
self.write_lock()
for field, value in (
("identity", "someone-else"),
("profile", "other-profile"),
):
with self.subTest(field=field):
result = self._heartbeat(**{field: value})
self.assertFalse(result["success"])
def test_branch_and_worktree_mismatch_are_refused(self):
self.write_lock()
wrong_branch = self._heartbeat(branch_name=f"fix/issue-{ISSUE}-other")
self.assertFalse(wrong_branch["success"])
wrong_worktree = self._heartbeat(worktree_path="/tmp/not-the-worktree")
self.assertFalse(wrong_worktree["success"])
def test_lapsed_lease_cannot_be_heartbeated_back_to_life(self):
"""No revival path (A4).
A session that stopped proving liveness must reclaim under a fresh
generation, not restore ownership retroactively.
"""
self.write_lock(
heartbeat_delta=timedelta(minutes=30), expires_delta=timedelta(hours=1)
)
result = self._heartbeat()
self.assertFalse(result["success"])
self.assertIn("reclaimed", " ".join(result["reasons"]))
def test_absent_lock_cannot_be_created_by_heartbeat(self):
result = self._heartbeat()
self.assertFalse(result["success"])
self.assertIn("no durable lock", result["reasons"][0])
def test_legacy_lock_is_refused_until_rebound(self):
self.write_lock(lifecycle=None)
result = self._heartbeat()
self.assertFalse(result["success"])
self.assertTrue(result["legacy_lease"])
self.assertIn("rebound", " ".join(result["reasons"]))
class TestLegacyRebind(_LockFixture):
"""AC-N8 exit route: canonical exact-owner rebinding."""
def _rebind(self, **kwargs):
params = {
"remote": REMOTE,
"org": ORG,
"repo": REPO,
"issue_number": ISSUE,
"branch_name": BRANCH,
"worktree_path": self.worktree,
"identity": IDENTITY,
"profile": PROFILE,
"lock_dir": self.lock_dir.name,
"now": self.now,
}
params.update(kwargs)
return issue_lock_store.rebind_legacy_lock(**params)
def test_rebind_mints_a_session_and_a_genuine_first_heartbeat(self):
self.write_lock(
lifecycle=None,
created_delta=timedelta(hours=3),
heartbeat_delta=timedelta(hours=3),
expires_delta=timedelta(hours=1),
generation=2,
)
result = self._rebind()
self.assertTrue(result["success"], result)
self.assertTrue(result["task_session_id"])
self.assertEqual(result["lock_generation"], 3)
written = issue_lock_store.read_lock_file(self._path())
lease = written["work_lease"]
self.assertEqual(
lease["lifecycle_version"], lease_policy.LIFECYCLE_HEARTBEAT_V1
)
self.assertEqual(lease["last_heartbeat_at"], _ts(self.now))
self.assertEqual(lease["expires_at"], _ts(self.now + timedelta(minutes=10)))
self.assertFalse(issue_lock_store.is_legacy_lease(written))
# The original claim is preserved for audit rather than overwritten.
origin = written["legacy_rebind"]["legacy_origin"]
self.assertTrue(origin["created_at"])
self.assertEqual(origin["lifecycle"], lease_policy.LIFECYCLE_LEGACY)
def test_rebound_lock_can_then_heartbeat(self):
self.write_lock(
lifecycle=None,
created_delta=timedelta(hours=3),
heartbeat_delta=timedelta(hours=3),
expires_delta=timedelta(hours=1),
)
rebound = self._rebind()
beat = issue_lock_store.heartbeat_session_lock(
remote=REMOTE,
org=ORG,
repo=REPO,
issue_number=ISSUE,
branch_name=BRANCH,
worktree_path=self.worktree,
identity=IDENTITY,
profile=PROFILE,
task_session_id=rebound["task_session_id"],
lock_dir=self.lock_dir.name,
now=self.now + timedelta(minutes=1),
)
self.assertTrue(beat["success"], beat)
def test_rebind_refuses_a_foreign_owner(self):
self.write_lock(lifecycle=None, expires_delta=timedelta(hours=1))
result = self._rebind(identity="someone-else")
self.assertFalse(result["success"])
def test_rebind_refuses_a_lock_already_on_the_lifecycle(self):
self.write_lock()
result = self._rebind()
self.assertFalse(result["success"])
self.assertFalse(result["legacy_lease"])
def test_rebind_is_not_a_recovery_path_for_a_lapsed_legacy_lease(self):
"""An expired legacy lease belongs to #760 renewal or #601 reclaim."""
self.write_lock(
lifecycle=None,
created_delta=timedelta(hours=5),
heartbeat_delta=timedelta(hours=5),
expires_delta=timedelta(hours=-1),
)
result = self._rebind()
self.assertFalse(result["success"])
self.assertIn("not a recovery path", " ".join(result["reasons"]))
if __name__ == "__main__":
unittest.main()
@@ -0,0 +1,193 @@
"""Conflict gate honors heartbeat-lifecycle non-live reclaim bands (#790 review #502/#516).
``assess_expired_lock_reclaim`` already permits reclaim for the heartbeat-lifecycle
stale bands (``stale_missed_heartbeat``, ``stale_absolute_cap``) without a dead PID.
But ``assess_same_issue_lease_conflict`` used to enter that reclaim branch only under
``is_lease_expired`` (``expires_at <= now``). For a heartbeat-lifecycle lease that is
non-live yet whose ``expires_at`` is still in the future, the acquisition gate fell
through to the foreign "already has an active lease" block and never consulted the
reclaim assessor so the load-bearing heartbeat was not load-bearing for foreign
reclaim, the exact abandonment scenario #790 exists to fix.
Two future-``expires_at`` non-live shapes are reachable:
* ``stale_absolute_cap`` a session that keeps heartbeating past the 8h absolute cap
has ``expires_at = last_heartbeat + TTL`` in the future (default policy).
* ``stale_missed_heartbeat`` under an independent TTL>grace policy the heartbeat
grace lapses while ``expires_at`` is still ahead.
These tests pin: both reclaim from a foreign acquirer; a live lease still blocks a
foreign acquirer; the same owner may reclaim its own abandoned heartbeat lease; and a
legacy (pre-lifecycle) lock keeps its absolute-``expires_at`` clock a non-expired
legacy lock with a dead PID is *not* reclaimable through this path.
"""
from __future__ import annotations
import os
import sys
import unittest
from datetime import datetime, timedelta, timezone
from unittest import mock
sys.path.insert(0, os.path.dirname(os.path.dirname(os.path.abspath(__file__))))
import issue_lock_store as ils # noqa: E402
import lease_policy # noqa: E402
ISSUE = 790
OWNER_BRANCH = f"fix/issue-{ISSUE}-slice-a-heartbeat-policy"
OWNER_WORKTREE = "/tmp/wt-790-owner"
FOREIGN_BRANCH = f"fix/issue-{ISSUE}-foreign-attempt"
FOREIGN_WORKTREE = "/tmp/wt-790-foreign"
def _ts(moment: datetime) -> str:
return (
moment.astimezone(timezone.utc)
.replace(microsecond=0)
.isoformat()
.replace("+00:00", "Z")
)
class _ConflictGateBase(unittest.TestCase):
def setUp(self):
self.now = datetime(2026, 7, 23, 12, 0, 0, tzinfo=timezone.utc)
def _lock(
self,
*,
lifecycle: str | None,
created_ago: timedelta,
heartbeat_ago: timedelta,
expires_in: timedelta,
pid: int,
) -> dict:
lease: dict = {
"operation_type": ils.AUTHOR_ISSUE_WORK_LEASE,
"issue_number": ISSUE,
"branch": OWNER_BRANCH,
"worktree_path": OWNER_WORKTREE,
"created_at": _ts(self.now - created_ago),
"last_heartbeat_at": _ts(self.now - heartbeat_ago),
"expires_at": _ts(self.now + expires_in),
}
if lifecycle is not None:
lease["lifecycle_version"] = lifecycle
lease["task_session_id"] = "author_issue_work-deadbeefdeadbeef"
return {
"issue_number": ISSUE,
"branch_name": OWNER_BRANCH,
"remote": "prgs",
"org": "Scaled-Tech-Consulting",
"repo": "Gitea-Tools",
"worktree_path": OWNER_WORKTREE,
"session_pid": pid,
"pid": pid,
"work_lease": lease,
}
def _foreign_conflict(self, existing: dict) -> str | None:
return ils.assess_same_issue_lease_conflict(
existing,
issue_number=ISSUE,
branch_name=FOREIGN_BRANCH,
worktree_path=FOREIGN_WORKTREE,
now=self.now,
)
def _same_owner_conflict(self, existing: dict) -> str | None:
return ils.assess_same_issue_lease_conflict(
existing,
issue_number=ISSUE,
branch_name=OWNER_BRANCH,
worktree_path=OWNER_WORKTREE,
now=self.now,
)
class TestHeartbeatNonLiveFutureExpiresReclaim(_ConflictGateBase):
def test_stale_absolute_cap_future_expires_allows_foreign_reclaim(self):
# Created >8h ago, heartbeated one minute ago, expires 9 min in the FUTURE,
# owner PID alive: freshness = stale_absolute_cap, live=False, not expired.
existing = self._lock(
lifecycle=lease_policy.LIFECYCLE_HEARTBEAT_V1,
created_ago=timedelta(hours=9),
heartbeat_ago=timedelta(minutes=1),
expires_in=timedelta(minutes=9),
pid=os.getpid(),
)
freshness = ils.assess_lock_freshness(existing, now=self.now)
self.assertEqual(freshness["status"], ils.STATUS_STALE_ABSOLUTE_CAP)
self.assertFalse(freshness["live"])
self.assertFalse(ils.is_lease_expired(existing, now=self.now))
with mock.patch.object(ils, "is_process_alive", return_value=True):
self.assertIsNone(self._foreign_conflict(existing))
def test_missed_heartbeat_ttl_gt_grace_future_expires_allows_foreign_reclaim(self):
# Heartbeat grace (default 10 min) lapsed 5 min ago, but a TTL>grace policy
# leaves expires_at 20 min in the FUTURE: stale_missed_heartbeat, not expired.
existing = self._lock(
lifecycle=lease_policy.LIFECYCLE_HEARTBEAT_V1,
created_ago=timedelta(minutes=30),
heartbeat_ago=timedelta(minutes=15),
expires_in=timedelta(minutes=20),
pid=os.getpid(),
)
freshness = ils.assess_lock_freshness(existing, now=self.now)
self.assertEqual(freshness["status"], ils.STATUS_STALE_MISSED_HEARTBEAT)
self.assertFalse(freshness["live"])
self.assertFalse(ils.is_lease_expired(existing, now=self.now))
with mock.patch.object(ils, "is_process_alive", return_value=True):
self.assertIsNone(self._foreign_conflict(existing))
def test_same_owner_may_reclaim_its_own_abandoned_heartbeat_lease(self):
existing = self._lock(
lifecycle=lease_policy.LIFECYCLE_HEARTBEAT_V1,
created_ago=timedelta(hours=9),
heartbeat_ago=timedelta(minutes=1),
expires_in=timedelta(minutes=9),
pid=os.getpid(),
)
with mock.patch.object(ils, "is_process_alive", return_value=True):
self.assertIsNone(self._same_owner_conflict(existing))
class TestLiveAndLegacyStillBlockForeign(_ConflictGateBase):
def test_live_heartbeat_lease_still_blocks_foreign(self):
existing = self._lock(
lifecycle=lease_policy.LIFECYCLE_HEARTBEAT_V1,
created_ago=timedelta(minutes=5),
heartbeat_ago=timedelta(minutes=1),
expires_in=timedelta(minutes=9),
pid=os.getpid(),
)
freshness = ils.assess_lock_freshness(existing, now=self.now)
self.assertTrue(freshness["live"])
with mock.patch.object(ils, "is_process_alive", return_value=True):
block = self._foreign_conflict(existing)
self.assertIn("already has an active", block or "")
def test_legacy_lease_future_expires_dead_pid_is_not_reclaimed_here(self):
# AC-N8: a legacy lock keeps its absolute expires_at clock. Non-expired +
# dead PID is non-live, but the widened band excludes legacy, so the
# foreign acquirer is still blocked rather than silently reclaiming.
existing = self._lock(
lifecycle=None,
created_ago=timedelta(hours=9),
heartbeat_ago=timedelta(hours=9),
expires_in=timedelta(hours=2),
pid=4_194_304, # far above any live pid on a test host
)
with mock.patch.object(ils, "is_process_alive", return_value=False):
freshness = ils.assess_lock_freshness(existing, now=self.now)
self.assertTrue(freshness["legacy_lease"])
self.assertFalse(freshness["live"])
self.assertFalse(ils.is_lease_expired(existing, now=self.now))
block = self._foreign_conflict(existing)
self.assertIn("already has an active", block or "")
if __name__ == "__main__":
unittest.main()
+682
View File
@@ -0,0 +1,682 @@
"""Cross-role allocation handoff consumable by independent workers (#843).
Regression coverage for the controllerrequired-role consume path:
* controller allocates author work; independent author adopts successfully
* author adoption succeeds after allocating controller process exits
* author adoption without sharing controller session identity
* wrong-role adoption rejected
* concurrent/second adoption rejected without state corruption
* terminal allocation adoption rejected
* successful adoption produces authoritative ownership evidence
* genuine abandoned-lease recovery remains valid
* process_work_queue / allocate results include consume identifiers
* same-role allocation behavior remains compatible
"""
from __future__ import annotations
import os
import tempfile
import unittest
from datetime import timedelta
from unittest.mock import patch
from allocator_service import (
ALLOCATION_MODE_CROSS_ROLE,
ALLOCATION_MODE_ROLE_SCOPED,
OUTCOME_ASSIGNED,
ROLE_AUTHOR,
ROLE_CONTROLLER,
ROLE_REVIEWER,
WorkCandidate,
allocate_next_work,
)
from control_plane_db import ControlPlaneDB, ForeignLeaseError, _ts, _utc_now
import lease_lifecycle as ll
class CrossRoleHandoffTest(unittest.TestCase):
def setUp(self) -> None:
self._tmp = tempfile.TemporaryDirectory()
self.db_path = os.path.join(self._tmp.name, "cp.sqlite3")
self.db = ControlPlaneDB(self.db_path)
self.db.upsert_session(
session_id="ctrl-session",
role="controller",
profile="prgs-controller",
pid=99999999, # dead-looking pid
)
self.db.upsert_session(
session_id="author-worker",
role="author",
profile="prgs-author",
pid=os.getpid(),
)
self.db.upsert_session(
session_id="author-worker-2",
role="author",
profile="prgs-author",
pid=os.getpid(),
)
self.db.upsert_session(
session_id="reviewer-worker",
role="reviewer",
profile="prgs-reviewer",
pid=os.getpid(),
)
self.wt = self._tmp.name
def tearDown(self) -> None:
self._tmp.cleanup()
def _ready_issue(self, number: int = 843, title: str = "handoff target") -> WorkCandidate:
return WorkCandidate(
kind="issue",
number=number,
labels=("status:ready", "type:bug"),
title=title,
priority=20,
)
def _controller_allocate(self, number: int = 843, **kwargs):
defaults = dict(
db=self.db,
session_id="ctrl-session",
role=ROLE_CONTROLLER,
remote="prgs",
org="Scaled-Tech-Consulting",
repo="Gitea-Tools",
candidates=[self._ready_issue(number)],
apply=True,
profile_name="prgs-controller",
username="controller-user",
allocation_mode=ALLOCATION_MODE_CROSS_ROLE,
)
defaults.update(kwargs)
return allocate_next_work(**defaults)
def test_controller_allocates_author_independent_author_adopts(self) -> None:
res = self._controller_allocate()
self.assertEqual(res["outcome"], OUTCOME_ASSIGNED)
self.assertEqual(res["required_role"], ROLE_AUTHOR)
self.assertIn("consume_allocation", res)
consume = res["consume_allocation"]
self.assertEqual(consume["tool"], "gitea_adopt_workflow_lease")
self.assertEqual(consume["required_role"], ROLE_AUTHOR)
self.assertFalse(consume["controller_session_required"])
lid = res["assignment"]["lease_id"]
self.assertEqual(consume["lease_id"], lid)
self.assertIn(lid, res["next_valid_command"])
adopted = ll.adopt_lease(
self.db,
lease_id=lid,
adopter_session_id="author-worker",
role=ROLE_AUTHOR,
worktree_path=self.wt,
)
self.assertTrue(adopted["success"])
self.assertEqual(adopted["outcome"], "adopted_cross_role_handoff")
self.assertEqual(adopted["adopted_by_session_id"], "author-worker")
self.assertEqual(adopted["adopted_from_session_id"], "ctrl-session")
raw = adopted["read_after_write"]
self.assertEqual(raw["session_id"], "author-worker")
self.assertEqual(raw["adopted_by_session_id"], "author-worker")
self.assertEqual(raw["status"], "active")
self.assertEqual(raw["phase"], "adopted")
# Authoritative re-read
state = self.db.get_lease_workflow_state(lid)
self.assertEqual(state["lease"]["session_id"], "author-worker")
self.assertEqual(state["lease"]["adopted_by_session_id"], "author-worker")
self.assertEqual(state["assignment"]["session_id"], "author-worker")
self.assertEqual(state["provenance"]["handoff_status"], "adopted")
def test_author_adoption_after_controller_process_exits(self) -> None:
res = self._controller_allocate(number=900)
lid = res["assignment"]["lease_id"]
# Force owner_pid dead + freshness stale_dead_process
import sqlite3
conn = sqlite3.connect(self.db_path)
try:
conn.execute(
"UPDATE leases SET owner_pid = 99999999 WHERE lease_id = ?",
(lid,),
)
conn.commit()
finally:
conn.close()
state = self.db.get_lease_workflow_state(lid)
fr = ll.classify_lease_freshness(
state["lease"], pid_checker=lambda _p: False
)
self.assertEqual(fr["freshness"], "stale_dead_process")
adopted = ll.adopt_lease(
self.db,
lease_id=lid,
adopter_session_id="author-worker",
role=ROLE_AUTHOR,
worktree_path=self.wt,
)
self.assertEqual(adopted["outcome"], "adopted_cross_role_handoff")
self.assertEqual(adopted["adopted_by_session_id"], "author-worker")
# No abandon required
state2 = self.db.get_lease_workflow_state(lid)
self.assertEqual(state2["lease"]["status"], "active")
self.assertNotEqual(state2["lease"]["status"], "abandoned")
def test_adoption_without_sharing_controller_session_identity(self) -> None:
res = self._controller_allocate(number=901)
lid = res["assignment"]["lease_id"]
adopted = ll.adopt_lease(
self.db,
lease_id=lid,
adopter_session_id="author-worker",
role=ROLE_AUTHOR,
worktree_path=self.wt,
)
self.assertNotEqual(adopted["adopted_by_session_id"], "ctrl-session")
self.assertFalse(adopted["same_owner"])
self.assertEqual(adopted["adopted_from_session_id"], "ctrl-session")
def test_wrong_role_adoption_rejected(self) -> None:
res = self._controller_allocate(number=902)
lid = res["assignment"]["lease_id"]
with self.assertRaises(ll.LeaseLifecycleError) as ctx:
ll.adopt_lease(
self.db,
lease_id=lid,
adopter_session_id="reviewer-worker",
role=ROLE_REVIEWER,
)
self.assertIn("wrong role", str(ctx.exception).lower())
# State unchanged
state = self.db.get_lease_workflow_state(lid)
self.assertEqual(state["lease"]["session_id"], "ctrl-session")
self.assertIsNone(state["lease"].get("adopted_by_session_id") or None)
self.assertEqual(state["provenance"]["handoff_status"], "pending")
def test_second_adoption_rejected_without_corruption(self) -> None:
res = self._controller_allocate(number=903)
lid = res["assignment"]["lease_id"]
first = ll.adopt_lease(
self.db,
lease_id=lid,
adopter_session_id="author-worker",
role=ROLE_AUTHOR,
worktree_path=self.wt,
)
self.assertEqual(first["outcome"], "adopted_cross_role_handoff")
with self.assertRaises(ll.LeaseLifecycleError):
ll.adopt_lease(
self.db,
lease_id=lid,
adopter_session_id="author-worker-2",
role=ROLE_AUTHOR,
worktree_path=self.wt,
)
state = self.db.get_lease_workflow_state(lid)
self.assertEqual(state["lease"]["session_id"], "author-worker")
self.assertEqual(state["lease"]["adopted_by_session_id"], "author-worker")
self.assertEqual(state["assignment"]["session_id"], "author-worker")
self.assertEqual(state["lease"]["status"], "active")
def test_terminal_allocation_adoption_rejected(self) -> None:
res = self._controller_allocate(number=904)
lid = res["assignment"]["lease_id"]
# Abandon as terminal
proof = ll.AbandonProof(
dead_process=True,
missing_worktree=True,
no_open_pr=True,
no_live_mutation_risk=True,
owner_pid=99999999,
worktree_path="/nonexistent/for-843",
)
# Attach dead pid / missing wt for abandon eligibility
import sqlite3
conn = sqlite3.connect(self.db_path)
try:
conn.execute(
"UPDATE leases SET owner_pid = 99999999, worktree_path = ? WHERE lease_id = ?",
("/nonexistent/for-843", lid),
)
conn.commit()
finally:
conn.close()
abandoned = ll.abandon_lease(
self.db,
lease_id=lid,
requester_session_id="author-worker",
proof=proof,
)
self.assertEqual(abandoned["outcome"], "abandoned")
with self.assertRaises(ll.LeaseLifecycleError) as ctx:
ll.adopt_lease(
self.db,
lease_id=lid,
adopter_session_id="author-worker",
role=ROLE_AUTHOR,
)
self.assertIn("abandoned", str(ctx.exception).lower())
def test_successful_adoption_read_after_write_ownership(self) -> None:
res = self._controller_allocate(number=905)
lid = res["assignment"]["lease_id"]
adopted = ll.adopt_lease(
self.db,
lease_id=lid,
adopter_session_id="author-worker",
role=ROLE_AUTHOR,
worktree_path=self.wt,
)
raw = adopted["read_after_write"]
self.assertEqual(raw["lease_id"], lid)
self.assertEqual(raw["session_id"], "author-worker")
self.assertEqual(raw["adopted_by_session_id"], "author-worker")
self.assertEqual(raw["adopted_from_session_id"], "ctrl-session")
# Re-fetch proves durable write
state = self.db.get_lease_workflow_state(lid)
self.assertEqual(state["lease"]["session_id"], raw["session_id"])
self.assertEqual(
state["lease"]["adopted_by_session_id"], raw["adopted_by_session_id"]
)
def test_genuine_abandoned_recovery_still_valid(self) -> None:
"""Same-role author lease abandoned remains reclaimable via abandon path."""
same = allocate_next_work(
self.db,
session_id="author-worker",
role=ROLE_AUTHOR,
remote="prgs",
org="Scaled-Tech-Consulting",
repo="Gitea-Tools",
candidates=[self._ready_issue(906, "same-role")],
apply=True,
profile_name="prgs-author",
username="author-user",
allocation_mode=ALLOCATION_MODE_ROLE_SCOPED,
)
self.assertEqual(same["outcome"], OUTCOME_ASSIGNED)
lid = same["assignment"]["lease_id"]
import sqlite3
conn = sqlite3.connect(self.db_path)
try:
conn.execute(
"UPDATE leases SET owner_pid = 99999999, worktree_path = ? WHERE lease_id = ?",
("/nonexistent/same-role", lid),
)
conn.commit()
finally:
conn.close()
proof = ll.AbandonProof(
dead_process=True,
missing_worktree=True,
no_open_pr=True,
no_live_mutation_risk=True,
owner_pid=99999999,
worktree_path="/nonexistent/same-role",
)
abandoned = ll.abandon_lease(
self.db,
lease_id=lid,
requester_session_id="author-worker-2",
proof=proof,
)
self.assertEqual(abandoned["outcome"], "abandoned")
# Foreign author cannot handoff-consume an abandoned non-handoff lease
with self.assertRaises(ll.LeaseLifecycleError):
ll.adopt_lease(
self.db,
lease_id=lid,
adopter_session_id="author-worker-2",
role=ROLE_AUTHOR,
)
# Reclaim path still works for expired/abandoned after force-expire
reclaimed = ll.reclaim_expired_lease(
self.db,
lease_id=lid,
session_id="author-worker-2",
role=ROLE_AUTHOR,
worktree_path=self.wt,
)
self.assertEqual(reclaimed["outcome"], "reclaimed")
self.assertEqual(reclaimed["assignment"]["session_id"], "author-worker-2")
def test_allocate_payload_includes_consume_identifiers(self) -> None:
res = self._controller_allocate(number=907)
self.assertIn("consume_allocation", res)
c = res["consume_allocation"]
for key in (
"tool",
"lease_id",
"assignment_id",
"required_role",
"required_profile",
"required_namespace",
"instructions",
"handoff_status",
):
self.assertIn(key, c)
self.assertEqual(c["required_namespace"], "gitea-author")
self.assertEqual(c["required_profile"], "prgs-author")
self.assertIn("gitea_adopt_workflow_lease", c["instructions"])
self.assertTrue(res["lease_proof"]["cross_role_handoff"])
self.assertEqual(res["lease_proof"]["handoff_status"], "pending")
def test_same_role_allocation_remains_compatible(self) -> None:
res = allocate_next_work(
self.db,
session_id="author-worker",
role=ROLE_AUTHOR,
remote="prgs",
org="Scaled-Tech-Consulting",
repo="Gitea-Tools",
candidates=[self._ready_issue(908)],
apply=True,
profile_name="prgs-author",
username="author-user",
)
self.assertEqual(res["outcome"], OUTCOME_ASSIGNED)
self.assertNotIn("consume_allocation", res)
lid = res["assignment"]["lease_id"]
state = self.db.get_lease_workflow_state(lid)
# No cross-role handoff provenance
prov = state.get("provenance") or {}
self.assertFalse(prov.get("cross_role_handoff"))
# Owner resume still works
resume = ll.adopt_lease(
self.db,
lease_id=lid,
adopter_session_id="author-worker",
role=ROLE_AUTHOR,
worktree_path=self.wt,
)
self.assertTrue(resume["same_owner"])
self.assertEqual(resume["outcome"], "adopted_owner_resume")
def test_inspect_points_required_role_at_consume(self) -> None:
res = self._controller_allocate(number=909)
lid = res["assignment"]["lease_id"]
decision = ll.inspect_lease(
self.db, lid, caller_session_id="author-worker"
)
self.assertEqual(
decision["safe_next_action"], ll.SAFE_CONSUME_CROSS_ROLE
)
self.assertFalse(decision["block"])
self.assertEqual(decision["required_role"], ROLE_AUTHOR)
def test_db_cas_rejects_concurrent_second_consume(self) -> None:
res = self._controller_allocate(number=910)
lid = res["assignment"]["lease_id"]
# First consume via DB layer directly
first = self.db.adopt_lease(
lease_id=lid,
adopter_session_id="author-worker",
role=ROLE_AUTHOR,
worktree_path=self.wt,
provenance={
"cross_role_handoff": True,
"handoff_status": "adopted",
"required_role": "author",
},
)
self.assertEqual(first["outcome"], "adopted_cross_role_handoff")
# Second CAS must fail
with self.assertRaises(ForeignLeaseError):
self.db.adopt_lease(
lease_id=lid,
adopter_session_id="author-worker-2",
role=ROLE_AUTHOR,
worktree_path=self.wt,
provenance={
"cross_role_handoff": True,
"handoff_status": "pending",
"required_role": "author",
},
)
state = self.db.get_lease_workflow_state(lid)
self.assertEqual(state["lease"]["session_id"], "author-worker")
class MCPBoundaryAdoptRoleBindingTest(unittest.TestCase):
"""#843 F1: MCP-boundary role binding for ``gitea_adopt_workflow_lease``.
The library-level wrong-role test calls ``lease_lifecycle.adopt_lease``
directly. These tests prove the MCP entry point derives the adopter role
authoritatively from the active authenticated profile and rejects any
caller-supplied role that disagrees, so a reviewer/merger profile cannot
consume an author handoff by passing ``role="author"``.
"""
AUTHOR_PROFILE = {
"profile_name": "prgs-author",
"role": "author",
"allowed_operations": [
"gitea.read",
"gitea.pr.create",
"gitea.branch.push",
],
"forbidden_operations": [],
}
REVIEWER_PROFILE = {
"profile_name": "prgs-reviewer",
"role": "reviewer",
"allowed_operations": [
"gitea.read",
"gitea.pr.review",
"gitea.pr.approve",
"gitea.pr.request_changes",
],
"forbidden_operations": ["gitea.pr.create", "gitea.branch.push"],
}
MERGER_PROFILE = {
"profile_name": "prgs-merger",
"role": "merger",
"allowed_operations": ["gitea.read", "gitea.pr.merge"],
"forbidden_operations": ["gitea.pr.create", "gitea.branch.push"],
}
FOREIGN_AUTHOR_PROFILE = {
"profile_name": "dadeschools-author",
"role": "author",
"allowed_operations": [
"gitea.read",
"gitea.pr.create",
"gitea.branch.push",
],
"forbidden_operations": [],
}
def setUp(self) -> None:
self._tmp = tempfile.TemporaryDirectory()
self.db_path = os.path.join(self._tmp.name, "cp.sqlite3")
self.db = ControlPlaneDB(self.db_path)
self.db.upsert_session(
session_id="ctrl-session",
role="controller",
profile="prgs-controller",
pid=99999999,
)
self.db.upsert_session(
session_id="author-worker",
role="author",
profile="prgs-author",
pid=os.getpid(),
)
self.wt = self._tmp.name
def tearDown(self) -> None:
self._tmp.cleanup()
def _ready_issue(self, number: int) -> WorkCandidate:
return WorkCandidate(
kind="issue",
number=number,
labels=("status:ready", "type:bug"),
title="handoff target",
priority=20,
)
def _handoff_lease(self, number: int = 843) -> str:
res = allocate_next_work(
db=self.db,
session_id="ctrl-session",
role=ROLE_CONTROLLER,
remote="prgs",
org="Scaled-Tech-Consulting",
repo="Gitea-Tools",
candidates=[self._ready_issue(number)],
apply=True,
profile_name="prgs-controller",
username="controller-user",
allocation_mode=ALLOCATION_MODE_CROSS_ROLE,
)
self.assertEqual(res["outcome"], OUTCOME_ASSIGNED)
self.assertEqual(res["required_role"], ROLE_AUTHOR)
return res["assignment"]["lease_id"]
def _call_adopt_tool(self, profile: dict, **kwargs):
import gitea_mcp_server as mcp_server
with (
patch.object(mcp_server, "get_profile", return_value=profile),
patch.object(
mcp_server,
"_control_plane_db_or_error",
return_value=(self.db, []),
),
):
return mcp_server.gitea_adopt_workflow_lease(
remote="prgs", **kwargs
)
def _assert_handoff_untouched(self, lease_id: str) -> None:
state = self.db.get_lease_workflow_state(lease_id)
self.assertEqual(state["lease"]["session_id"], "ctrl-session")
self.assertIsNone(state["lease"].get("adopted_by_session_id") or None)
self.assertEqual(state["lease"]["status"], "active")
self.assertEqual(state["provenance"]["handoff_status"], "pending")
def test_reviewer_profile_cannot_consume_author_handoff_via_role_author(
self,
) -> None:
lid = self._handoff_lease(920)
result = self._call_adopt_tool(
self.REVIEWER_PROFILE,
lease_id=lid,
session_id="reviewer-worker",
role="author",
worktree_path=self.wt,
)
self.assertFalse(result["success"])
self.assertEqual(result["outcome"], "blocked")
self.assertEqual(result["profile_role_kind"], "reviewer")
self.assertEqual(result["supplied_role"], "author")
self.assertIn("does not match", result["reasons"][0])
self._assert_handoff_untouched(lid)
def test_merger_profile_cannot_consume_author_handoff_via_role_author(
self,
) -> None:
lid = self._handoff_lease(921)
result = self._call_adopt_tool(
self.MERGER_PROFILE,
lease_id=lid,
session_id="merger-worker",
role="author",
worktree_path=self.wt,
)
self.assertFalse(result["success"])
self.assertEqual(result["outcome"], "blocked")
self.assertEqual(result["profile_role_kind"], "merger")
self._assert_handoff_untouched(lid)
def test_reviewer_profile_rejected_without_role_argument(self) -> None:
"""Even without a spoofed role, the profile-derived role binds."""
lid = self._handoff_lease(922)
result = self._call_adopt_tool(
self.REVIEWER_PROFILE,
lease_id=lid,
session_id="reviewer-worker",
worktree_path=self.wt,
)
self.assertFalse(result["success"])
self.assertEqual(result["outcome"], "blocked")
self.assertIn("wrong role", result["reasons"][0].lower())
self._assert_handoff_untouched(lid)
def test_author_profile_mismatching_supplied_role_rejected(self) -> None:
lid = self._handoff_lease(923)
result = self._call_adopt_tool(
self.AUTHOR_PROFILE,
lease_id=lid,
session_id="author-worker",
role="reviewer",
worktree_path=self.wt,
)
self.assertFalse(result["success"])
self.assertEqual(result["outcome"], "blocked")
self.assertEqual(result["profile_role_kind"], "author")
self.assertEqual(result["supplied_role"], "reviewer")
self._assert_handoff_untouched(lid)
def test_foreign_profile_name_rejected_for_author_handoff(self) -> None:
"""Provenance required_profile binds even when the role matches."""
lid = self._handoff_lease(924)
result = self._call_adopt_tool(
self.FOREIGN_AUTHOR_PROFILE,
lease_id=lid,
session_id="foreign-author-worker",
worktree_path=self.wt,
)
self.assertFalse(result["success"])
self.assertEqual(result["outcome"], "blocked")
self.assertIn("wrong profile", result["reasons"][0].lower())
self._assert_handoff_untouched(lid)
def test_author_profile_consumes_author_handoff(self) -> None:
lid = self._handoff_lease(925)
result = self._call_adopt_tool(
self.AUTHOR_PROFILE,
lease_id=lid,
session_id="author-worker",
role="author",
worktree_path=self.wt,
)
self.assertTrue(result["success"])
self.assertEqual(result["outcome"], "adopted_cross_role_handoff")
self.assertEqual(result["adopted_by_session_id"], "author-worker")
self.assertEqual(result["adopted_from_session_id"], "ctrl-session")
state = self.db.get_lease_workflow_state(lid)
self.assertEqual(state["lease"]["session_id"], "author-worker")
self.assertEqual(
state["lease"]["adopted_by_session_id"], "author-worker"
)
self.assertEqual(state["assignment"]["session_id"], "author-worker")
self.assertEqual(state["provenance"]["handoff_status"], "adopted")
def test_author_profile_consumes_author_handoff_without_role_argument(
self,
) -> None:
lid = self._handoff_lease(926)
result = self._call_adopt_tool(
self.AUTHOR_PROFILE,
lease_id=lid,
session_id="author-worker",
worktree_path=self.wt,
)
self.assertTrue(result["success"])
self.assertEqual(result["outcome"], "adopted_cross_role_handoff")
state = self.db.get_lease_workflow_state(lid)
self.assertEqual(state["lease"]["session_id"], "author-worker")
self.assertEqual(state["lease"]["role"], "author")
if __name__ == "__main__":
unittest.main()
@@ -0,0 +1,244 @@
"""#855 AC4: an expired reviewer lease must not indefinitely protect an
already-merged branch when no live claimant exists.
Two layers are covered:
* ``branch_cleanup_guard.assess_expired_reviewer_lease_reclaim`` the pure,
fail-closed reclaim decision. Every condition must be provably satisfied or
the lease keeps protecting the branch.
* ``gitea_mcp_server._collect_branch_ownership_records`` the wiring that
supplies authoritative evidence (PR merged state, owner-process liveness,
competing ownership) to that decision, and flips an expired reviewer lease
to reclaimable only under the full policy.
All inputs are fabricated; no real repository, lease, or credential is used.
"""
import importlib
import unittest
from unittest.mock import patch
import branch_cleanup_guard
mcp_server = importlib.import_module("gitea_mcp_server")
FAKE_AUTH = "token fake"
REMOTE = "prgs"
ORG = "Scaled-Tech-Consulting"
REPO = "Gitea-Tools"
HOST = "gitea.prgs.cc"
BRANCH = "feat/issue-638-webui-app-shell-phase1"
PR_NUMBER = 818
class TestAssessExpiredReviewerLeaseReclaim(unittest.TestCase):
"""Pure fail-closed reclaim decision (#855 AC4)."""
def _call(self, **overrides):
base = dict(
role="reviewer",
status="expired",
pr_merged=True,
owner_pid_alive=False,
competing_active_claimant=False,
)
base.update(overrides)
return branch_cleanup_guard.assess_expired_reviewer_lease_reclaim(**base)
def test_full_policy_satisfied_allows_reclaim(self):
out = self._call()
self.assertTrue(out["reclaim_allowed"])
self.assertEqual(out["reasons"], [])
self.assertEqual(out["decision"], "reclaim_expired_reviewer_lease")
def test_stale_dead_process_reviewer_also_reclaimable(self):
out = self._call(status="stale_dead_process")
self.assertTrue(out["reclaim_allowed"])
def test_non_reviewer_role_never_reclaims(self):
for role in ("author", "merger", "controller", "reconciler", "unknown"):
with self.subTest(role=role):
out = self._call(role=role)
self.assertFalse(out["reclaim_allowed"])
self.assertTrue(out["reasons"])
self.assertEqual(out["decision"], "keep_protecting")
def test_active_status_never_reclaims(self):
out = self._call(status="active")
self.assertFalse(out["reclaim_allowed"])
def test_pr_not_merged_blocks_reclaim(self):
out = self._call(pr_merged=False)
self.assertFalse(out["reclaim_allowed"])
def test_pr_merged_unknown_fails_closed(self):
out = self._call(pr_merged=None)
self.assertFalse(out["reclaim_allowed"])
def test_owner_process_alive_blocks_reclaim(self):
out = self._call(owner_pid_alive=True)
self.assertFalse(out["reclaim_allowed"])
def test_owner_liveness_unknown_fails_closed(self):
out = self._call(owner_pid_alive=None)
self.assertFalse(out["reclaim_allowed"])
def test_competing_active_claimant_blocks_reclaim(self):
out = self._call(competing_active_claimant=True)
self.assertFalse(out["reclaim_allowed"])
def test_competing_claimant_unknown_fails_closed(self):
out = self._call(competing_active_claimant=None)
self.assertFalse(out["reclaim_allowed"])
def test_reasons_never_leak_secrets(self):
out = self._call(role="author")
blob = " ".join(out["reasons"]).lower()
self.assertNotIn("token", blob)
self.assertNotIn("password", blob)
class _FakeLease(dict):
pass
class TestCollectorExpiredReviewerReclaimWiring(unittest.TestCase):
"""`_collect_branch_ownership_records` supplies authoritative evidence and
flips an expired reviewer lease to reclaimable only under the full policy."""
def _run(
self,
*,
lease_role="reviewer",
lease_freshness="stale_dead_process",
owner_pid_alive=False,
pr_merged=True,
extra_leases=None,
worktree_on_branch=False,
):
lease = _FakeLease(
role=lease_role,
work_kind="pr",
work_number=PR_NUMBER,
branch=BRANCH,
status="active",
owner_pid=999999,
remote=REMOTE,
org=ORG,
repo=REPO,
host=HOST,
freshness={
"freshness": lease_freshness,
"owner_pid": 999999,
"owner_pid_alive": owner_pid_alive,
"expired_by_time": lease_freshness == "expired",
},
)
leases = [lease] + list(extra_leases or [])
pr_payload = {
"number": PR_NUMBER,
"merged": pr_merged,
"merged_at": "2026-07-23T00:00:00Z" if pr_merged else None,
"head": {"ref": BRANCH},
}
def fake_api_request(method, url, *a, **k):
if method == "GET" and f"/pulls/{PR_NUMBER}" in url:
return pr_payload
raise AssertionError(f"unexpected api_request {method} {url}")
wt_entries = []
if worktree_on_branch:
wt_entries = [{"branch": BRANCH, "path": f"/x/branches/{BRANCH}"}]
with patch.object(
mcp_server.lease_lifecycle,
"list_active_leases",
return_value={"leases": leases},
), patch.object(
mcp_server.control_plane_db, "get_db", return_value=object(), create=True
), patch.object(
mcp_server.issue_lock_store, "iter_lock_files", return_value=[]
), patch.object(
mcp_server.worktree_cleanup_audit,
"list_worktrees",
return_value=wt_entries,
), patch.object(
mcp_server, "api_get_all", return_value=[]
), patch.object(
mcp_server, "api_request", side_effect=fake_api_request
):
return mcp_server._collect_branch_ownership_records(
remote=REMOTE,
host=HOST,
org=ORG,
repo=REPO,
branch=BRANCH,
pr_number=PR_NUMBER,
project_root="/x",
auth=FAKE_AUTH,
base_api="https://gitea.prgs.cc/api/v1/repos/x/y",
)
def _reviewer_records(self, bundle):
return [
rec
for rec in bundle["records"]
if rec.get("category")
== branch_cleanup_guard.OWNERSHIP_CATEGORY_REVIEWER_LEASE
]
def test_merged_dead_uncontested_reviewer_lease_is_reclaimable(self):
bundle = self._run()
self.assertFalse(bundle["inventory_error"])
recs = self._reviewer_records(bundle)
self.assertEqual(len(recs), 1)
self.assertTrue(recs[0]["reclaim_allowed"])
# And the guard consequently does not block deletion on it.
ownership = branch_cleanup_guard.assess_active_branch_ownership(
remote=REMOTE, org=ORG, repo=REPO, branch=BRANCH, host=HOST,
records=bundle["records"],
)
self.assertFalse(ownership["block"])
def test_unmerged_pr_keeps_reviewer_lease_protective(self):
bundle = self._run(pr_merged=False)
recs = self._reviewer_records(bundle)
self.assertEqual(len(recs), 1)
self.assertFalse(recs[0]["reclaim_allowed"])
ownership = branch_cleanup_guard.assess_active_branch_ownership(
remote=REMOTE, org=ORG, repo=REPO, branch=BRANCH, host=HOST,
records=bundle["records"],
)
self.assertTrue(ownership["block"])
def test_owner_process_alive_keeps_reviewer_lease_protective(self):
bundle = self._run(owner_pid_alive=True, lease_freshness="expired")
recs = self._reviewer_records(bundle)
self.assertFalse(recs[0]["reclaim_allowed"])
def test_competing_worktree_binding_keeps_reviewer_lease_protective(self):
bundle = self._run(worktree_on_branch=True)
recs = self._reviewer_records(bundle)
self.assertFalse(recs[0]["reclaim_allowed"])
ownership = branch_cleanup_guard.assess_active_branch_ownership(
remote=REMOTE, org=ORG, repo=REPO, branch=BRANCH, host=HOST,
records=bundle["records"],
)
self.assertTrue(ownership["block"])
def test_expired_author_lease_never_reclaimed_by_reviewer_policy(self):
bundle = self._run(lease_role="author")
author_recs = [
rec
for rec in bundle["records"]
if rec.get("category")
== branch_cleanup_guard.OWNERSHIP_CATEGORY_AUTHOR_LEASE
]
self.assertEqual(len(author_recs), 1)
self.assertFalse(author_recs[0]["reclaim_allowed"])
if __name__ == "__main__":
unittest.main()
@@ -0,0 +1,551 @@
"""Merged-PR awareness for the worktree cleanup audit (#858).
Before #858 an ``issue_work`` worktree could never leave ``active_issue_work``:
the audit had no PR linkage at all (``pr_number`` was structurally ``None``)
and its only route to ``clean_stale_removable`` was a TTL derived from a
``last_used_at`` that nothing ever populated. A merged, clean, unprotected
worktree was therefore reported as active work forever, disagreeing with the
PR-scoped reconciler.
These tests use fabricated temporary repositories and synthetic PR records
only. Nothing here removes a worktree or deletes a branch.
"""
import os
import subprocess
import sys
import tempfile
import unittest
from unittest.mock import patch
sys.path.insert(0, str(__import__("pathlib").Path(__file__).resolve().parent.parent))
import merged_cleanup_reconcile as mcr # noqa: E402
import worktree_cleanup_audit as wca # noqa: E402
MERGED_BRANCH = "feat/issue-777-timeline"
MERGED_PATH = "/repo/branches/issue-777-timeline"
HEAD_SHA = "a" * 40
def _pr(number, branch, *, merged=True, sha=HEAD_SHA, state=None):
"""Synthetic Gitea PR payload."""
return {
"number": number,
"head": {"ref": branch, "sha": sha},
"merged_at": "2026-07-24T01:00:00Z" if merged else None,
"state": state or ("closed" if merged else "open"),
}
def _porcelain(*entries):
out = []
for path, branch, sha in entries:
out.append(f"worktree {path}")
out.append(f"HEAD {sha}")
if branch is None:
out.append("detached")
else:
out.append(f"branch refs/heads/{branch}")
out.append("")
return "\n".join(out)
class _AuditHarness(unittest.TestCase):
"""Runs audit_branches_directory over a fabricated worktree listing."""
PORCELAIN = _porcelain(
("/repo", "master", "f" * 40),
(MERGED_PATH, MERGED_BRANCH, HEAD_SHA),
)
def run_audit(self, *, dirty_paths=(), contained=True, **kwargs):
def fake_dirty(path):
if path in dirty_paths:
return {"exists": True, "dirty": True, "dirty_files": [" M x.py"]}
return {"exists": True, "dirty": False, "dirty_files": []}
with patch.object(
wca, "list_worktrees",
return_value=wca.parse_worktree_porcelain(self.PORCELAIN),
), patch.object(
wca, "read_worktree_dirty", side_effect=fake_dirty
), patch.object(
wca, "git_worktree_list", return_value="(mocked)"
), patch.object(
wca, "is_head_ancestor_of_ref", return_value=contained
):
report = wca.audit_branches_directory("/repo", **kwargs)
return {wt["path"]: wt for wt in report["worktrees"]}, report
def merged_audit(self, **kwargs):
kwargs.setdefault("pr_index", wca.build_pr_index([_pr(849, MERGED_BRANCH)]))
kwargs.setdefault("master_ref", "prgs/master")
return self.run_audit(**kwargs)
class TestMergedWorktreeBecomesRemovable(_AuditHarness):
def test_clean_merged_issue_worktree_is_linked_and_removable(self):
by_path, report = self.merged_audit()
entry = by_path[MERGED_PATH]
self.assertEqual(entry["classification"], wca.CLASS_CLEAN_STALE_REMOVABLE)
self.assertTrue(entry["removable"])
self.assertEqual(entry["merged_pr_linkage"]["status"], wca.LINKAGE_MERGED)
self.assertEqual(entry["merged_pr_cleanup"]["block_reasons"], [])
self.assertIn(MERGED_PATH, [c["path"] for c in report["removable_candidates"]])
def test_pr_number_populated_from_authoritative_linkage(self):
by_path, _ = self.merged_audit()
self.assertEqual(by_path[MERGED_PATH]["pr_number"], 849)
def test_regression_without_pr_evidence_stays_active_issue_work(self):
"""The pre-#858 behaviour, still correct when no PR state is supplied."""
by_path, _ = self.run_audit()
entry = by_path[MERGED_PATH]
self.assertEqual(entry["classification"], wca.CLASS_ACTIVE_ISSUE_WORK)
self.assertFalse(entry["removable"])
self.assertIsNone(entry["pr_number"])
class TestProtectiveSignalsSurvive(_AuditHarness):
def test_open_pr_worktree_is_not_removable(self):
index = wca.build_pr_index([_pr(900, MERGED_BRANCH, merged=False)])
by_path, _ = self.run_audit(
pr_index=index,
master_ref="prgs/master",
open_pr_branches={MERGED_BRANCH},
)
entry = by_path[MERGED_PATH]
self.assertEqual(entry["classification"], wca.CLASS_ACTIVE_OPEN_PR)
self.assertFalse(entry["removable"])
# linkage still reports the owning PR, it just is not merge proof
self.assertEqual(entry["merged_pr_linkage"]["status"], wca.LINKAGE_OPEN)
self.assertEqual(entry["pr_number"], 900)
def test_dirty_tracked_worktree_is_not_removable(self):
by_path, _ = self.merged_audit(dirty_paths=(MERGED_PATH,))
entry = by_path[MERGED_PATH]
self.assertEqual(entry["classification"], wca.CLASS_DIRTY_LOCAL)
self.assertFalse(entry["removable"])
self.assertIn(
"worktree has uncommitted changes",
entry["merged_pr_cleanup"]["block_reasons"],
)
def test_untracked_only_worktree_is_not_removable(self):
"""``git status --porcelain`` reports untracked files as dirty too."""
def untracked(path):
if path == MERGED_PATH:
return {"exists": True, "dirty": True, "dirty_files": ["?? scratch.txt"]}
return {"exists": True, "dirty": False, "dirty_files": []}
with patch.object(
wca, "list_worktrees",
return_value=wca.parse_worktree_porcelain(self.PORCELAIN),
), patch.object(
wca, "read_worktree_dirty", side_effect=untracked
), patch.object(
wca, "git_worktree_list", return_value="(mocked)"
), patch.object(
wca, "is_head_ancestor_of_ref", return_value=True
):
report = wca.audit_branches_directory(
"/repo",
pr_index=wca.build_pr_index([_pr(849, MERGED_BRANCH)]),
master_ref="prgs/master",
)
entry = {wt["path"]: wt for wt in report["worktrees"]}[MERGED_PATH]
self.assertEqual(entry["classification"], wca.CLASS_DIRTY_LOCAL)
self.assertFalse(entry["removable"])
def test_active_lease_by_issue_number_is_protective(self):
by_path, _ = self.merged_audit(leased_issue_numbers={777})
entry = by_path[MERGED_PATH]
self.assertTrue(entry["has_active_lease"])
self.assertEqual(entry["classification"], wca.CLASS_ACTIVE_ISSUE_WORK)
self.assertFalse(entry["removable"])
def test_active_lease_by_branch_is_protective(self):
by_path, _ = self.merged_audit(leased_branches={MERGED_BRANCH})
entry = by_path[MERGED_PATH]
self.assertTrue(entry["has_active_lease"])
self.assertFalse(entry["removable"])
def test_active_issue_lock_is_protective(self):
by_path, _ = self.merged_audit(active_issue_branches={MERGED_BRANCH})
entry = by_path[MERGED_PATH]
self.assertTrue(entry["has_active_issue_lock"])
self.assertEqual(entry["classification"], wca.CLASS_ACTIVE_ISSUE_WORK)
self.assertFalse(entry["removable"])
def test_live_session_worktree_is_protective(self):
by_path, _ = self.merged_audit(live_session_paths={MERGED_PATH})
entry = by_path[MERGED_PATH]
self.assertTrue(entry["has_live_session"])
self.assertEqual(entry["classification"], wca.CLASS_ACTIVE_ISSUE_WORK)
self.assertFalse(entry["removable"])
def test_head_not_contained_in_master_is_not_removable(self):
by_path, _ = self.merged_audit(contained=False)
entry = by_path[MERGED_PATH]
self.assertEqual(entry["classification"], wca.CLASS_ACTIVE_ISSUE_WORK)
self.assertFalse(entry["removable"])
self.assertIn(
"worktree head is not contained in authoritative master "
"(unmerged commits remain)",
entry["merged_pr_cleanup"]["block_reasons"],
)
def test_unknown_containment_fails_closed(self):
by_path, _ = self.merged_audit(contained=None)
entry = by_path[MERGED_PATH]
self.assertFalse(entry["removable"])
self.assertIn(
"containment of the worktree head in master is unknown",
entry["merged_pr_cleanup"]["block_reasons"],
)
def test_missing_master_ref_fails_closed(self):
by_path, _ = self.run_audit(
pr_index=wca.build_pr_index([_pr(849, MERGED_BRANCH)])
)
self.assertFalse(by_path[MERGED_PATH]["removable"])
def test_unmerged_owning_pr_is_not_removable(self):
index = wca.build_pr_index([_pr(901, MERGED_BRANCH, merged=False)])
by_path, _ = self.run_audit(pr_index=index, master_ref="prgs/master")
entry = by_path[MERGED_PATH]
self.assertFalse(entry["removable"])
self.assertIn(
"owning PR #901 is not merged",
entry["merged_pr_cleanup"]["block_reasons"],
)
def test_control_checkout_is_never_removable(self):
by_path, _ = self.merged_audit()
control = by_path["/repo"]
self.assertTrue(control["is_protected"])
self.assertEqual(control["classification"], wca.CLASS_UNSAFE_UNKNOWN)
self.assertFalse(control["removable"])
def test_control_checkout_not_removable_even_if_linked_and_merged(self):
"""A merged PR on the control checkout must not unlock removal."""
porcelain = _porcelain(("/repo", MERGED_BRANCH, HEAD_SHA))
with patch.object(
wca, "list_worktrees", return_value=wca.parse_worktree_porcelain(porcelain)
), patch.object(
wca, "read_worktree_dirty",
return_value={"exists": True, "dirty": False, "dirty_files": []},
), patch.object(
wca, "git_worktree_list", return_value="(mocked)"
), patch.object(
wca, "is_head_ancestor_of_ref", return_value=True
):
report = wca.audit_branches_directory(
"/repo",
pr_index=wca.build_pr_index([_pr(849, MERGED_BRANCH)]),
master_ref="prgs/master",
)
entry = report["worktrees"][0]
self.assertEqual(entry["classification"], wca.CLASS_UNSAFE_UNKNOWN)
self.assertFalse(entry["removable"])
class TestAmbiguousLinkageFailsClosed(_AuditHarness):
def test_competing_prs_on_one_branch_fail_closed(self):
index = wca.build_pr_index(
[_pr(849, MERGED_BRANCH), _pr(860, MERGED_BRANCH)]
)
by_path, _ = self.run_audit(pr_index=index, master_ref="prgs/master")
entry = by_path[MERGED_PATH]
self.assertEqual(entry["merged_pr_linkage"]["status"], wca.LINKAGE_AMBIGUOUS)
self.assertIsNone(entry["pr_number"])
self.assertEqual(entry["classification"], wca.CLASS_ACTIVE_ISSUE_WORK)
self.assertFalse(entry["removable"])
def test_merged_plus_open_pr_on_one_branch_fails_closed(self):
index = wca.build_pr_index(
[_pr(849, MERGED_BRANCH), _pr(861, MERGED_BRANCH, merged=False)]
)
by_path, _ = self.run_audit(pr_index=index, master_ref="prgs/master")
entry = by_path[MERGED_PATH]
self.assertEqual(entry["merged_pr_linkage"]["status"], wca.LINKAGE_AMBIGUOUS)
self.assertFalse(entry["removable"])
def test_no_owning_pr_fails_closed(self):
by_path, _ = self.run_audit(
pr_index=wca.build_pr_index([_pr(849, "feat/other-branch")]),
master_ref="prgs/master",
)
entry = by_path[MERGED_PATH]
self.assertEqual(entry["merged_pr_linkage"]["status"], wca.LINKAGE_NONE)
self.assertFalse(entry["removable"])
def test_malformed_pr_records_are_dropped_not_guessed(self):
index = wca.build_pr_index(
[
{"number": None, "head": {"ref": MERGED_BRANCH}},
{"number": 5, "head": {}},
{"number": "not-an-int", "head": {"ref": MERGED_BRANCH}},
]
)
self.assertEqual(index, {})
self.assertEqual(
wca.resolve_owning_pr(branch=MERGED_BRANCH, pr_index=index)["status"],
wca.LINKAGE_NONE,
)
def test_detached_worktree_has_no_branch_linkage(self):
self.assertEqual(
wca.resolve_owning_pr(branch=None, pr_index={})["status"],
wca.LINKAGE_UNKNOWN,
)
class TestUnrelatedClassificationsUnchanged(unittest.TestCase):
"""Non-issue_work worktrees keep their pre-#858 classifications."""
PORCELAIN = _porcelain(
("/repo", "master", "f" * 40),
("/repo/branches/review-pr42", "review-pr42", "2" * 40),
("/repo/branches/baseline-master-x", "baseline-master-x", "3" * 40),
("/repo/branches/conflict-fix-pr50", "conflict-fix-pr50", "4" * 40),
("/repo/branches/review-pr99", None, "5" * 40),
)
def _audit(self, **kwargs):
with patch.object(
wca, "list_worktrees",
return_value=wca.parse_worktree_porcelain(self.PORCELAIN),
), patch.object(
wca, "read_worktree_dirty",
return_value={"exists": True, "dirty": False, "dirty_files": []},
), patch.object(
wca, "git_worktree_list", return_value="(mocked)"
), patch.object(
wca, "is_head_ancestor_of_ref", return_value=True
):
report = wca.audit_branches_directory("/repo", **kwargs)
return {wt["path"]: wt for wt in report["worktrees"]}
def test_classifications_identical_with_and_without_pr_evidence(self):
without = self._audit()
with_evidence = self._audit(
pr_index=wca.build_pr_index([_pr(849, MERGED_BRANCH)]),
master_ref="prgs/master",
)
self.assertEqual(
{p: e["classification"] for p, e in without.items()},
{p: e["classification"] for p, e in with_evidence.items()},
)
def test_lease_on_issue_does_not_capture_similarly_named_scratch_trees(self):
"""A lease on issue 777 protects issue work, not baseline/review trees."""
porcelain = _porcelain(
("/repo/branches/baseline-master-issue-777", "baseline-issue-777", "7" * 40),
("/repo/branches/issue-777-timeline", MERGED_BRANCH, HEAD_SHA),
)
with patch.object(
wca, "list_worktrees", return_value=wca.parse_worktree_porcelain(porcelain)
), patch.object(
wca, "read_worktree_dirty",
return_value={"exists": True, "dirty": False, "dirty_files": []},
), patch.object(
wca, "git_worktree_list", return_value="(mocked)"
), patch.object(
wca, "is_head_ancestor_of_ref", return_value=True
):
report = wca.audit_branches_directory(
"/repo",
pr_index=wca.build_pr_index([_pr(849, MERGED_BRANCH)]),
master_ref="prgs/master",
leased_issue_numbers={777},
)
by_path = {wt["path"]: wt for wt in report["worktrees"]}
baseline = by_path["/repo/branches/baseline-master-issue-777"]
self.assertFalse(baseline["has_active_lease"])
self.assertEqual(baseline["classification"], wca.CLASS_CLEAN_STALE_REMOVABLE)
issue_work = by_path["/repo/branches/issue-777-timeline"]
self.assertTrue(issue_work["has_active_lease"])
self.assertFalse(issue_work["removable"])
def test_review_and_baseline_still_removable(self):
by_path = self._audit(
pr_index=wca.build_pr_index([]), master_ref="prgs/master"
)
self.assertEqual(
by_path["/repo/branches/review-pr42"]["classification"],
wca.CLASS_CLEAN_STALE_REMOVABLE,
)
self.assertEqual(
by_path["/repo/branches/baseline-master-x"]["classification"],
wca.CLASS_CLEAN_STALE_REMOVABLE,
)
self.assertEqual(
by_path["/repo/branches/review-pr99"]["classification"],
wca.CLASS_DETACHED_REVIEW_LEFTOVER,
)
def test_conflict_fix_ttl_behaviour_unchanged(self):
"""conflict_fix still needs only TTL expiry; #858 did not touch it."""
self.assertEqual(
wca.classify_worktree(
workflow_type=wca.WORKFLOW_CONFLICT_FIX,
is_dirty=False,
ttl_expired=True,
),
wca.CLASS_CLEAN_STALE_REMOVABLE,
)
self.assertEqual(
wca.classify_worktree(
workflow_type=wca.WORKFLOW_CONFLICT_FIX,
is_dirty=False,
ttl_expired=False,
),
wca.CLASS_ACTIVE_ISSUE_WORK,
)
def test_issue_work_ttl_alone_no_longer_grants_removal(self):
"""Age is not landing proof: TTL alone must not reclaim issue work."""
self.assertEqual(
wca.classify_worktree(
workflow_type=wca.WORKFLOW_ISSUE_WORK,
is_dirty=False,
ttl_expired=True,
),
wca.CLASS_ACTIVE_ISSUE_WORK,
)
class TestAssessorPerformsNoDeletion(_AuditHarness):
def test_audit_never_removes_a_worktree(self):
with patch.object(wca, "remove_worktree") as removal:
self.merged_audit()
removal.assert_not_called()
def test_audit_shells_out_to_no_destructive_git_command(self):
seen = []
real_run = subprocess.run
def recording_run(cmd, *args, **kwargs):
seen.append(cmd)
return real_run(["true"], *args, **kwargs)
with patch.object(subprocess, "run", side_effect=recording_run):
wca.audit_branches_directory("/nonexistent-repo-for-audit")
joined = [" ".join(c) if isinstance(c, list) else str(c) for c in seen]
for cmd in joined:
self.assertNotIn("worktree remove", cmd)
self.assertNotIn("branch -D", cmd)
self.assertNotIn("push", cmd)
class TestAgreementWithPrScopedReconciler(unittest.TestCase):
"""The audit and merged_cleanup_reconcile must agree on identical input.
Uses a real throwaway git repository so containment is computed by git
rather than asserted. Nothing outside the temporary directory is touched.
"""
def _git(self, *args):
subprocess.run(
["git", "-C", self.root, *args],
check=True,
capture_output=True,
text=True,
)
def setUp(self):
self._tmp = tempfile.TemporaryDirectory()
self.root = os.path.realpath(self._tmp.name)
self._git("init", "-b", "master", ".")
self._git("config", "user.email", "[email protected]")
self._git("config", "user.name", "Test")
with open(os.path.join(self.root, "seed.txt"), "w") as fh:
fh.write("seed\n")
self._git("add", "seed.txt")
self._git("commit", "-m", "seed")
self.branch = "feat/issue-777-timeline"
self._git("checkout", "-b", self.branch)
with open(os.path.join(self.root, "feature.txt"), "w") as fh:
fh.write("feature\n")
self._git("add", "feature.txt")
self._git("commit", "-m", "feature")
self.head_sha = subprocess.run(
["git", "-C", self.root, "rev-parse", "HEAD"],
capture_output=True, text=True, check=True,
).stdout.strip()
self._git("checkout", "master")
self._git("merge", "--no-ff", "-m", "merge feature", self.branch)
self.worktree = os.path.join(self.root, "branches", "issue-777-timeline")
self._git("worktree", "add", self.worktree, self.branch)
def tearDown(self):
self._tmp.cleanup()
def _pr_index(self):
return wca.build_pr_index(
[
{
"number": 849,
"head": {"ref": self.branch, "sha": self.head_sha},
"merged_at": "2026-07-24T01:00:00Z",
}
]
)
def _audit_entry(self):
report = wca.audit_branches_directory(
self.root, pr_index=self._pr_index(), master_ref="master"
)
return next(wt for wt in report["worktrees"] if wt["path"] == self.worktree)
def _reconciler_entry(self):
return mcr.assess_local_worktree_cleanup(
pr_number=849,
head_branch=self.branch,
merged=True,
worktree_state=mcr.resolve_cleanup_worktree_state(
project_root=self.root,
head_branch=self.branch,
issue_number=777,
pr_head_sha=self.head_sha,
target_ref="master",
),
active_lock=False,
)
def test_both_assessors_agree_the_worktree_is_safe(self):
audit_entry = self._audit_entry()
reconciler = self._reconciler_entry()
self.assertTrue(reconciler["safe_to_remove_worktree"], reconciler)
self.assertTrue(audit_entry["removable"], audit_entry)
self.assertEqual(audit_entry["pr_number"], reconciler["pr_number"])
self.assertEqual(audit_entry["merged_pr_cleanup"]["block_reasons"], [])
self.assertEqual(reconciler["block_reasons"], [])
def test_both_assessors_agree_a_dirty_worktree_is_unsafe(self):
with open(os.path.join(self.worktree, "feature.txt"), "a") as fh:
fh.write("local edit\n")
audit_entry = self._audit_entry()
reconciler = self._reconciler_entry()
self.assertFalse(audit_entry["removable"])
self.assertFalse(reconciler["safe_to_remove_worktree"])
def test_worktree_still_present_after_audit(self):
self._audit_entry()
self.assertTrue(os.path.isdir(self.worktree))
if __name__ == "__main__":
unittest.main()
@@ -0,0 +1,630 @@
"""Durable linked-issue lock head refresh + merge-sync dead-session recovery (#871).
``gitea_update_pr_branch_by_merge`` advances a PR's *remote* head but historically
never advanced the linked durable issue lock's recorded head. After the owning
session died the drifted lock became unrecoverable and no further synchronization
was possible (PR #866 / issue #855).
Two halves are covered:
* the write-side refresh (``issue_lock_store.assess/apply_durable_lock_head_refresh``)
that records the new synced head under compare-and-swap with read-after-write; and
* the read-side recovery relation (``issue_lock_recovery`` +
``issue_lock_worktree.read_merge_sync_provenance``) that lets a dead-session lock
whose recorded head is a merge-sync *ancestor* of the live PR head be recovered
and nothing else.
"""
from __future__ import annotations
import os
import subprocess
import sys
import tempfile
import unittest
from datetime import datetime, timedelta, timezone
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
import issue_lock_recovery # noqa: E402
import issue_lock_store # noqa: E402
import issue_lock_worktree # noqa: E402
ISSUE = 8710
PR_NUMBER = 8711
BRANCH = f"fix/issue-{ISSUE}-durable-lock-head-refresh"
IDENTITY = "example-user"
PROFILE = "example-author"
OLD = "a" * 40
NEW1 = "b" * 40
NEW2 = "c" * 40
BASE = "d" * 40
REMOTE = "prgs"
ORG = "ExampleOrg"
REPO = "ExampleRepo"
def dead_pid() -> int:
proc = subprocess.Popen([sys.executable, "-c", "pass"])
proc.wait()
return proc.pid
def future_ts(hours: int = 4) -> str:
return (
(datetime.now(timezone.utc) + timedelta(hours=hours))
.isoformat()
.replace("+00:00", "Z")
)
def _git(cwd, *args):
return subprocess.run(
["git", "-C", cwd, *args],
capture_output=True,
text=True,
check=True,
)
def _rev(cwd, ref="HEAD") -> str:
return _git(cwd, "rev-parse", ref).stdout.strip()
def build_merge_sync_repo(tmp: str) -> dict:
"""Build a repo where a feature branch was synced by merging master in.
Returns a dict with the prior (branch) head, the synced merge-commit head,
the master tip, plus a rebase-style linear descendant and an unrelated head.
"""
_git(tmp, "init", "-q", "-b", "master")
_git(tmp, "config", "user.email", "[email protected]")
_git(tmp, "config", "user.name", "T")
Path(tmp, "base.txt").write_text("base\n")
_git(tmp, "add", "-A")
_git(tmp, "commit", "-q", "-m", "root")
# Feature branch cut from root, one commit — this is the PRIOR/recorded head.
_git(tmp, "checkout", "-q", "-b", BRANCH)
Path(tmp, "feature.txt").write_text("feature\n")
_git(tmp, "add", "-A")
_git(tmp, "commit", "-q", "-m", "feature work")
prior = _rev(tmp)
# Master advances (the base the sync will merge in).
_git(tmp, "checkout", "-q", "master")
Path(tmp, "base.txt").write_text("base\nmore\n")
_git(tmp, "add", "-A")
_git(tmp, "commit", "-q", "-m", "master advance")
master_tip = _rev(tmp)
# Sync: merge master INTO the feature branch → merge commit, first parent = prior.
_git(tmp, "checkout", "-q", BRANCH)
_git(tmp, "merge", "-q", "--no-ff", "-m", "Merge master into feature", "master")
synced = _rev(tmp)
# A plain linear descendant of prior (NOT a merge) — a rebase/extra-commit shape.
_git(tmp, "checkout", "-q", "-b", "linear-branch", prior)
Path(tmp, "extra.txt").write_text("extra\n")
_git(tmp, "add", "-A")
_git(tmp, "commit", "-q", "-m", "extra linear commit")
linear = _rev(tmp)
# An unrelated root (force-push / rewritten history shape).
unrelated_dir = tempfile.mkdtemp()
_git(unrelated_dir, "init", "-q", "-b", "x")
_git(unrelated_dir, "config", "user.email", "[email protected]")
_git(unrelated_dir, "config", "user.name", "T")
Path(unrelated_dir, "z.txt").write_text("z\n")
_git(unrelated_dir, "add", "-A")
_git(unrelated_dir, "commit", "-q", "-m", "unrelated")
unrelated = _rev(unrelated_dir)
# Leave the worktree checked out on the feature branch at the PRIOR head, as
# a dead author session that never advanced would have left it.
_git(tmp, "checkout", "-q", BRANCH)
_git(tmp, "reset", "-q", "--hard", prior)
return {
"prior": prior,
"master_tip": master_tip,
"synced": synced,
"linear": linear,
"unrelated": unrelated,
}
# ─────────────────────────── write-side refresh ───────────────────────────
class TestDurableLockHeadRefresh(unittest.TestCase):
def setUp(self):
self.lock_dir = tempfile.mkdtemp()
self.wt = tempfile.mkdtemp()
lock_data = {
"issue_number": ISSUE,
"branch_name": BRANCH,
"worktree_path": self.wt,
"remote": REMOTE,
"org": ORG,
"repo": REPO,
"claimant": {"username": IDENTITY, "profile": PROFILE},
"work_lease": {
"operation_type": issue_lock_store.AUTHOR_ISSUE_WORK_LEASE,
"issue_number": ISSUE,
"branch": BRANCH,
"worktree_path": self.wt,
"claimant": {"username": IDENTITY, "profile": PROFILE},
"expires_at": future_ts(),
},
}
issue_lock_store.bind_session_lock(lock_data, lock_dir=self.lock_dir)
def _apply(self, **over):
kw = dict(
remote=REMOTE, org=ORG, repo=REPO, issue_number=ISSUE,
branch_name=BRANCH, worktree_path=self.wt, pr_number=PR_NUMBER,
identity=IDENTITY, profile=PROFILE, current_pid=os.getpid(),
expected_old_head=OLD, new_head=NEW1, synced_at=future_ts(0),
base_head=BASE, lock_dir=self.lock_dir,
)
kw.update(over)
return issue_lock_store.apply_durable_lock_head_refresh(**kw)
def _load(self):
return issue_lock_store.load_issue_lock(
remote=REMOTE, org=ORG, repo=REPO, issue_number=ISSUE,
lock_dir=self.lock_dir,
)
def test_first_sync_updates_recorded_head(self):
"""AC1: first base sync writes the resulting head to the durable lock."""
res = self._apply()
self.assertTrue(res["refreshed"], res["reasons"])
self.assertTrue(res["read_after_write_ok"])
self.assertEqual(self._load().get("synced_pr_head"), NEW1)
def test_second_sync_after_master_advance(self):
"""AC2: a later master advance permits a second sanctioned sync."""
self.assertTrue(self._apply()["refreshed"])
res2 = self._apply(expected_old_head=NEW1, new_head=NEW2)
self.assertTrue(res2["refreshed"], res2["reasons"])
self.assertEqual(self._load().get("synced_pr_head"), NEW2)
history = self._load().get("branch_sync_history")
self.assertEqual(len(history), 2)
self.assertEqual(history[0]["last_synced_pr_head"], NEW1)
self.assertEqual(history[1]["prior_pr_head"], NEW1)
def test_cas_detects_concurrent_head_change(self):
"""AC6: CAS refuses when the recorded synced head is not the old head."""
self.assertTrue(self._apply()["refreshed"]) # recorded head now NEW1
# A second sync claiming the old head is still OLD must fail closed.
res = self._apply(expected_old_head=OLD, new_head=NEW2)
self.assertFalse(res["refreshed"])
self.assertTrue(any("CAS" in r or "concurrent" in r for r in res["reasons"]))
self.assertEqual(self._load().get("synced_pr_head"), NEW1)
def test_wrong_issue_fails_closed(self):
res = self._apply(issue_number=999999)
self.assertFalse(res["refreshed"])
def test_wrong_branch_fails_closed(self):
res = self._apply(branch_name="fix/issue-8710-wrong")
self.assertFalse(res["refreshed"])
def test_wrong_repo_fails_closed(self):
res = self._apply(repo="OtherRepo")
self.assertFalse(res["refreshed"])
def test_wrong_identity_fails_closed(self):
res = self._apply(identity="intruder")
self.assertFalse(res["refreshed"])
def test_wrong_profile_fails_closed(self):
res = self._apply(profile="prgs-reviewer")
self.assertFalse(res["refreshed"])
def test_foreign_session_fails_closed(self):
"""A refresh is not a recovery: the current process must own the lock."""
path = issue_lock_store.lock_file_path(
remote=REMOTE, org=ORG, repo=REPO, issue_number=ISSUE,
lock_dir=self.lock_dir,
)
rec = issue_lock_store.read_lock_file(path)
rec["session_pid"] = dead_pid()
rec["pid"] = rec["session_pid"]
issue_lock_store.save_lock_file(path, rec)
res = self._apply()
self.assertFalse(res["refreshed"])
self.assertTrue(any("current session" in r or "live owner" in r for r in res["reasons"]))
def test_new_equals_old_fails_closed(self):
res = self._apply(expected_old_head=OLD, new_head=OLD)
self.assertFalse(res["refreshed"])
def test_non_full_sha_fails_closed(self):
self.assertFalse(self._apply(new_head="deadbeef")["refreshed"])
self.assertFalse(self._apply(expected_old_head="xyz")["refreshed"])
def test_no_lock_fails_closed(self):
assessment = issue_lock_store.assess_durable_lock_head_refresh(
None, remote=REMOTE, org=ORG, repo=REPO, issue_number=ISSUE,
branch_name=BRANCH, worktree_path=self.wt, pr_number=PR_NUMBER,
identity=IDENTITY, profile=PROFILE, current_pid=os.getpid(),
expected_old_head=OLD, new_head=NEW1,
)
self.assertFalse(assessment["allowed"])
# ─────────────────────── merge-sync provenance (real git) ───────────────────
class TestMergeSyncProvenanceObservation(unittest.TestCase):
def setUp(self):
self.tmp = tempfile.mkdtemp()
self.shas = build_merge_sync_repo(self.tmp)
def test_merge_sync_is_recognized(self):
obs = issue_lock_worktree.read_merge_sync_provenance(
self.tmp, prior_head_sha=self.shas["prior"],
synced_head_sha=self.shas["synced"],
)
self.assertTrue(obs["is_merge_sync"], obs["reasons"])
self.assertTrue(obs["prior_is_ancestor"])
self.assertTrue(obs["synced_is_merge"])
self.assertTrue(obs["first_parent_reaches_prior"])
def test_linear_descendant_is_not_a_merge_sync(self):
"""A plain non-merge descendant (rebase/extra commit) is not a sync."""
obs = issue_lock_worktree.read_merge_sync_provenance(
self.tmp, prior_head_sha=self.shas["prior"],
synced_head_sha=self.shas["linear"],
)
self.assertTrue(obs["probe_ok"])
self.assertFalse(obs["is_merge_sync"])
self.assertFalse(obs["synced_is_merge"])
def test_unrelated_history_fails_closed(self):
"""A rewritten/force-pushed head where prior is unreachable fails closed."""
obs = issue_lock_worktree.read_merge_sync_provenance(
self.tmp, prior_head_sha=self.shas["prior"],
synced_head_sha=self.shas["unrelated"],
)
self.assertFalse(obs["is_merge_sync"])
def test_missing_args_fail_closed(self):
obs = issue_lock_worktree.read_merge_sync_provenance(
self.tmp, prior_head_sha=None, synced_head_sha=self.shas["synced"],
)
self.assertFalse(obs["is_merge_sync"])
# ──────────────────── merge-sync dead-session recovery ──────────────────────
def make_dead_lock(worktree, **over):
pid = dead_pid()
lock = {
"issue_number": ISSUE,
"branch_name": BRANCH,
"worktree_path": worktree,
"remote": REMOTE,
"org": ORG,
"repo": REPO,
"session_pid": pid,
"pid": pid,
"claimant": {"username": IDENTITY, "profile": PROFILE},
"work_lease": {
"operation_type": issue_lock_store.AUTHOR_ISSUE_WORK_LEASE,
"issue_number": ISSUE,
"branch": BRANCH,
"worktree_path": worktree,
"claimant": {"username": IDENTITY, "profile": PROFILE},
"expires_at": future_ts(),
},
}
lock.update(over)
return lock
def sync_prov(prior, synced, **over):
d = {
"prior_head_sha": prior,
"synced_head_sha": synced,
"probe_ok": True,
"prior_present": True,
"synced_present": True,
"prior_is_ancestor": True,
"synced_is_merge": True,
"first_parent_reaches_prior": True,
"is_merge_sync": True,
"first_parent_sha": prior,
"parent_count": 2,
"proof": f"{synced} merged base into branch above {prior}",
"reasons": [],
}
d.update(over)
return d
class TestMergeSyncRecovery(unittest.TestCase):
def setUp(self):
self.tmp = tempfile.mkdtemp()
self.shas = build_merge_sync_repo(self.tmp)
self.prior = self.shas["prior"]
self.synced = self.shas["synced"]
def _assess(self, **over):
lock = over.pop("_lock", None) or make_dead_lock(self.tmp)
kw = dict(
issue_number=ISSUE, branch_name=BRANCH, worktree_path=self.tmp,
remote=REMOTE, org=ORG, repo=REPO, identity=IDENTITY, profile=PROFILE,
current_branch=BRANCH, porcelain_status="",
head_sha=self.prior, remote_head_sha=self.synced,
pr_head_sha=self.synced, pr_number=PR_NUMBER,
competing_live_locks=[], candidate_branches=[BRANCH],
current_pid=os.getpid(),
remote_branch_exists=True,
sync_provenance=sync_prov(self.prior, self.synced),
)
kw.update(over)
return issue_lock_recovery.assess_dead_session_lock_recovery(lock, **kw)
def test_merge_sync_drift_is_recoverable(self):
"""AC3/AC4: dead session, recorded head is a merge-sync ancestor of PR head."""
res = self._assess()
self.assertEqual(res["outcome"], issue_lock_recovery.RECOVERY_SANCTIONED, res["reasons"])
self.assertEqual(
res["evidence"]["head_relation"],
issue_lock_recovery.HEAD_RELATION_REMOTE_MERGE_SYNCED,
)
self.assertEqual(res["evidence"]["accepted_head"], self.synced)
def test_missing_provenance_fails_closed(self):
"""No server-derived provenance → cannot accept a remote ahead of local."""
res = self._assess(sync_provenance=None)
self.assertEqual(res["outcome"], issue_lock_recovery.REFUSED)
def test_non_ancestor_recorded_head_fails_closed(self):
"""AC7: provenance that does not prove ancestry is rejected."""
res = self._assess(
sync_provenance=sync_prov(
self.prior, self.synced, prior_is_ancestor=False, is_merge_sync=False,
reasons=["prior head is not an ancestor"],
)
)
self.assertEqual(res["outcome"], issue_lock_recovery.REFUSED)
def test_force_pushed_history_fails_closed(self):
"""AC8: a rewritten head (not a merge sync) stays protected."""
res = self._assess(
sync_provenance=sync_prov(
self.prior, self.synced, is_merge_sync=False, synced_is_merge=False,
reasons=["not a merge-based sync"],
)
)
self.assertEqual(res["outcome"], issue_lock_recovery.REFUSED)
def test_provenance_for_other_commits_fails_closed(self):
"""Provenance whose endpoints differ from the heads under assessment is rejected."""
res = self._assess(
sync_provenance=sync_prov("f" * 40, self.synced),
)
self.assertEqual(res["outcome"], issue_lock_recovery.REFUSED)
def test_dirty_worktree_fails_closed(self):
"""AC11: dirty worktrees remain protected."""
res = self._assess(porcelain_status=" M feature.txt\n")
self.assertEqual(res["outcome"], issue_lock_recovery.REFUSED)
def test_live_owner_fails_closed(self):
"""AC10: a live recorded owner is not a dead-session recovery."""
lock = make_dead_lock(self.tmp, session_pid=os.getpid(), pid=os.getpid())
res = self._assess(_lock=lock)
self.assertEqual(res["outcome"], issue_lock_recovery.REFUSED)
def test_competing_claimant_fails_closed(self):
"""AC13: a competing live lock blocks recovery."""
res = self._assess(
competing_live_locks=[{
"issue_number": ISSUE, "branch_name": BRANCH,
"worktree_path": "/some/other/wt", "pid": os.getpid(),
}]
)
self.assertEqual(res["outcome"], issue_lock_recovery.REFUSED)
def test_wrong_branch_fails_closed(self):
"""AC9: worktree on a different branch fails closed."""
res = self._assess(current_branch="fix/issue-8710-other")
self.assertEqual(res["outcome"], issue_lock_recovery.REFUSED)
def test_wrong_identity_fails_closed(self):
res = self._assess(identity="intruder")
self.assertEqual(res["outcome"], issue_lock_recovery.REFUSED)
def test_pr_head_mismatch_fails_closed(self):
"""The open PR must sit at the synced remote head."""
res = self._assess(pr_head_sha="e" * 40)
self.assertEqual(res["outcome"], issue_lock_recovery.REFUSED)
def test_owning_pr_evidence_for_merge_sync(self):
res = self._assess()
ev = issue_lock_recovery.owning_pr_recovery_evidence(res)
self.assertIsNotNone(ev)
self.assertEqual(ev["pr_number"], PR_NUMBER)
self.assertEqual(ev["head_sha"], self.synced)
self.assertEqual(
ev["head_relation"],
issue_lock_recovery.HEAD_RELATION_REMOTE_MERGE_SYNCED,
)
def test_recovered_owning_pr_from_persisted_record(self):
res = self._assess()
record = issue_lock_recovery.build_recovery_record(res, recovered_at=future_ts(0))
lock = {"issue_number": ISSUE, "branch_name": BRANCH,
"dead_session_recovery": record}
rebuilt = issue_lock_recovery.recovered_owning_pr_from_lock(lock)
self.assertIsNotNone(rebuilt)
self.assertEqual(rebuilt["head_sha"], self.synced)
self.assertEqual(
rebuilt["head_relation"],
issue_lock_recovery.HEAD_RELATION_REMOTE_MERGE_SYNCED,
)
class TestExistingRelationsUnchanged(unittest.TestCase):
"""AC14/AC15: equal-head recovery still works; merge-sync did not weaken it."""
def setUp(self):
self.tmp = tempfile.mkdtemp()
self.shas = build_merge_sync_repo(self.tmp)
def test_equal_head_recovery_still_sanctioned(self):
# Worktree at prior head; remote also at prior head → the #753 equal case.
prior = self.shas["prior"]
lock = make_dead_lock(self.tmp)
res = issue_lock_recovery.assess_dead_session_lock_recovery(
lock, issue_number=ISSUE, branch_name=BRANCH, worktree_path=self.tmp,
remote=REMOTE, org=ORG, repo=REPO, identity=IDENTITY, profile=PROFILE,
current_branch=BRANCH, porcelain_status="",
head_sha=prior, remote_head_sha=prior,
pr_head_sha=prior, pr_number=PR_NUMBER,
competing_live_locks=[], candidate_branches=[BRANCH],
current_pid=os.getpid(), remote_branch_exists=True,
)
self.assertEqual(res["outcome"], issue_lock_recovery.RECOVERY_SANCTIONED, res["reasons"])
self.assertEqual(
res["evidence"]["head_relation"], issue_lock_recovery.HEAD_RELATION_EQUAL,
)
class TestUpdatePrWrapperPartialFailure(unittest.TestCase):
"""AC5/AC16: the tool advances the remote head then refreshes the durable lock.
When the durable refresh fails after the remote advance, the tool must report a
partial lifecycle failure and NOT a fully successful synchronization. Exact PR-
head / base-head pinning is preserved (delegated to the real preflight, stubbed
here only to isolate the post-update lifecycle branch).
"""
def setUp(self):
import gitea_mcp_server as gms # noqa: E402
self.gms = gms
self._orig = {}
def _patch(name, value):
self._orig[name] = getattr(gms, name)
setattr(gms, name, value)
_patch("get_profile", lambda *a, **k: {
"allowed_operations": ["gitea.branch.push"],
"forbidden_operations": [],
"profile_name": "prgs-author",
})
_patch("_role_kind", lambda *a, **k: "author")
_patch("_profile_operation_gate", lambda *a, **k: None)
_patch("_permission_block_report", lambda *a, **k: {})
_patch("_resolve", lambda *a, **k: ("gitea.prgs.cc", ORG, REPO))
_patch("_verify_role_mutation_workspace", lambda *a, **k: None)
_patch("_get_workspace_porcelain", lambda *a, **k: "")
_patch("_canonical_local_git_root", lambda *a, **k: "/x")
_patch("_master_parity_block", lambda *a, **k: None)
_patch("_auth", lambda *a, **k: {"token": "x"})
_patch("repo_api_url", lambda *a, **k: "http://api")
_patch("_redact", lambda s: s)
_patch("_work_lease_claimant", lambda *a, **k: {
"username": IDENTITY, "profile": PROFILE,
})
_patch("_prove_author_ownership_for_pr", lambda *a, **k: {
"has_author_lock": True, "matched_issue": ISSUE,
"matched_via": "branch", "linked_issues": [ISSUE],
"recovered_owning_pr": None, "reasons": [],
})
# Real preflight is unit-tested elsewhere; stub it to isolate the
# post-update durable-lock lifecycle branch under test.
orig_pf = gms.pr_sync_status.assess_update_pr_branch_preflight
self._orig_pf = orig_pf
gms.pr_sync_status.assess_update_pr_branch_preflight = (
lambda *a, **k: {"mutation_allowed": True, "reasons": [], "performed": False}
)
# Sequence the two GET /pulls calls: OLD before update, NEW after.
self._pull_calls = {"n": 0}
def fake_api_request(method, url, auth, *a, **k):
m = method.upper()
if m == "GET" and url.endswith(f"/pulls/{PR_NUMBER}"):
self._pull_calls["n"] += 1
head = OLD if self._pull_calls["n"] == 1 else NEW1
return {
"state": "open",
"head": {"sha": head, "ref": BRANCH},
"base": {"sha": BASE, "ref": "master"},
"mergeable": True, "title": "t", "body": "b",
}
if m == "GET" and "/branches/" in url:
return {"commit": {"id": BASE}}
if m == "POST" and "/update" in url:
return {}
return {}
_patch("api_request", fake_api_request)
def tearDown(self):
for name, value in self._orig.items():
setattr(self.gms, name, value)
self.gms.pr_sync_status.assess_update_pr_branch_preflight = self._orig_pf
def _run(self):
return self.gms.gitea_update_pr_branch_by_merge(
pr_number=PR_NUMBER,
expected_pr_head_sha=OLD,
expected_base_head_sha=BASE,
remote=REMOTE,
worktree_path="/tmp/branches/wt-871",
)
def test_partial_failure_when_refresh_fails(self):
self._orig["apply_durable_lock_head_refresh"] = (
self.gms.issue_lock_store.apply_durable_lock_head_refresh
)
self.gms.issue_lock_store.apply_durable_lock_head_refresh = (
lambda **k: {"refreshed": False, "reasons": ["forced refresh failure"]}
)
try:
res = self._run()
finally:
self.gms.issue_lock_store.apply_durable_lock_head_refresh = (
self._orig["apply_durable_lock_head_refresh"]
)
self.assertTrue(res["performed"])
self.assertEqual(res["new_pr_head_sha"], NEW1)
self.assertFalse(res["success"])
self.assertTrue(res["partial_lifecycle_failure"])
self.assertFalse(res["durable_lock_refreshed"])
def test_full_success_when_refresh_succeeds(self):
self._orig["apply_durable_lock_head_refresh"] = (
self.gms.issue_lock_store.apply_durable_lock_head_refresh
)
self.gms.issue_lock_store.apply_durable_lock_head_refresh = (
lambda **k: {"refreshed": True, "read_after_write_ok": True,
"new_head": NEW1, "reasons": ["ok"]}
)
try:
res = self._run()
finally:
self.gms.issue_lock_store.apply_durable_lock_head_refresh = (
self._orig["apply_durable_lock_head_refresh"]
)
self.assertTrue(res["success"])
self.assertTrue(res["performed"])
self.assertTrue(res["durable_lock_refreshed"])
self.assertTrue(res["fully_synchronized"])
self.assertEqual(res["new_pr_head_sha"], NEW1)
if __name__ == "__main__":
unittest.main()
@@ -0,0 +1,332 @@
"""Dead-session recovery for cross-session PR base-syncs (#872).
Validates that when a sanctioned server-side base-sync (``gitea_update_pr_branch_by_merge``)
advances a PR's remote head from A to B (a merge commit whose first parent is A),
a subsequent author session whose local worktree is at A can recover the dead-session
lock via HEAD_RELATION_REMOTE_MERGE_SYNCED and execute a second base-sync (B to C)
without deadlock or RuntimeError.
"""
from __future__ import annotations
import os
import subprocess
import sys
import tempfile
import unittest
from datetime import datetime, timedelta, timezone
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
import issue_lock_recovery # noqa: E402
import issue_lock_store # noqa: E402
import issue_lock_worktree # noqa: E402
ISSUE = 8720
PR_NUMBER = 8721
BRANCH = f"fix/issue-{ISSUE}-dead-session-two-base-syncs"
IDENTITY = "example-author"
PROFILE = "prgs-author"
REMOTE = "prgs"
ORG = "ExampleOrg"
REPO = "ExampleRepo"
def dead_pid() -> int:
proc = subprocess.Popen([sys.executable, "-c", "pass"])
proc.wait()
return proc.pid
def future_ts(hours: int = 4) -> str:
return (
(datetime.now(timezone.utc) + timedelta(hours=hours))
.isoformat()
.replace("+00:00", "Z")
)
def _git(cwd, *args):
return subprocess.run(
["git", "-C", cwd, *args],
capture_output=True,
text=True,
check=True,
)
def _rev(cwd, ref="HEAD") -> str:
return _git(cwd, "rev-parse", ref).stdout.strip()
def build_two_base_sync_repo(tmp: str) -> dict:
"""Build a repo simulating two sequential base-sync merges."""
_git(tmp, "init", "-q", "-b", "master")
_git(tmp, "config", "user.email", "[email protected]")
_git(tmp, "config", "user.name", "Author")
Path(tmp, "base.txt").write_text("base 1\n")
_git(tmp, "add", "-A")
_git(tmp, "commit", "-q", "-m", "initial master")
# Feature branch cut from initial master -> Head A (prior/recorded head)
_git(tmp, "checkout", "-q", "-b", BRANCH)
Path(tmp, "feature.txt").write_text("feature work\n")
_git(tmp, "add", "-A")
_git(tmp, "commit", "-q", "-m", "feature commit A")
head_a = _rev(tmp)
# Master advances -> Master 1
_git(tmp, "checkout", "-q", "master")
Path(tmp, "base.txt").write_text("base 1\nbase 2\n")
_git(tmp, "add", "-A")
_git(tmp, "commit", "-q", "-m", "master advance 1")
master_1 = _rev(tmp)
# First sync: merge master into feature -> Head B (merge commit, first parent = A)
_git(tmp, "checkout", "-q", BRANCH)
_git(tmp, "merge", "-q", "--no-ff", "-m", "First base-sync (merge master)", "master")
head_b = _rev(tmp)
# Master advances again -> Master 2
_git(tmp, "checkout", "-q", "master")
Path(tmp, "base.txt").write_text("base 1\nbase 2\nbase 3\n")
_git(tmp, "add", "-A")
_git(tmp, "commit", "-q", "-m", "master advance 2")
master_2 = _rev(tmp)
# Second sync: merge master into feature -> Head C (merge commit, first parent = B)
_git(tmp, "checkout", "-q", BRANCH)
_git(tmp, "merge", "-q", "--no-ff", "-m", "Second base-sync (merge master)", "master")
head_c = _rev(tmp)
# Non-merge rebase/force-pushed branch shape
_git(tmp, "checkout", "-q", "-b", "rebased-branch", head_a)
Path(tmp, "rebase.txt").write_text("rebased\n")
_git(tmp, "add", "-A")
_git(tmp, "commit", "-q", "-m", "rebased commit")
rebased_head = _rev(tmp)
# Reset worktree back to head A, as a dead session leaving local worktree at A
_git(tmp, "checkout", "-q", BRANCH)
_git(tmp, "reset", "-q", "--hard", head_a)
return {
"head_a": head_a,
"master_1": master_1,
"head_b": head_b,
"master_2": master_2,
"head_c": head_c,
"rebased_head": rebased_head,
}
def make_dead_lock(worktree, **over):
pid = dead_pid()
lock = {
"issue_number": ISSUE,
"branch_name": BRANCH,
"worktree_path": worktree,
"remote": REMOTE,
"org": ORG,
"repo": REPO,
"session_pid": pid,
"pid": pid,
"claimant": {"username": IDENTITY, "profile": PROFILE},
"work_lease": {
"operation_type": issue_lock_store.AUTHOR_ISSUE_WORK_LEASE,
"issue_number": ISSUE,
"branch": BRANCH,
"worktree_path": worktree,
"claimant": {"username": IDENTITY, "profile": PROFILE},
"expires_at": future_ts(),
},
}
lock.update(over)
return lock
class TestDeadSessionTwoBaseSyncs(unittest.TestCase):
def setUp(self):
self.tmp = tempfile.mkdtemp()
self.shas = build_two_base_sync_repo(self.tmp)
self.head_a = self.shas["head_a"]
self.head_b = self.shas["head_b"]
self.head_c = self.shas["head_c"]
def test_first_base_sync_dead_session_recovery(self):
"""AC1: dead session after first base sync (remote=B, local=A) recovers via merge-sync."""
lock = make_dead_lock(self.tmp, synced_pr_head=self.head_b)
obs = issue_lock_worktree.read_merge_sync_provenance(
self.tmp,
prior_head_sha=self.head_a,
synced_head_sha=self.head_b,
remote=REMOTE,
)
self.assertTrue(obs["is_merge_sync"], obs["reasons"])
self.assertTrue(obs["prior_is_ancestor"])
self.assertTrue(obs["synced_is_merge"])
self.assertTrue(obs["first_parent_reaches_prior"])
res = issue_lock_recovery.assess_dead_session_lock_recovery(
lock,
issue_number=ISSUE,
branch_name=BRANCH,
worktree_path=self.tmp,
remote=REMOTE,
org=ORG,
repo=REPO,
identity=IDENTITY,
profile=PROFILE,
current_branch=BRANCH,
porcelain_status="",
head_sha=self.head_a,
remote_head_sha=self.head_b,
pr_head_sha=self.head_b,
pr_number=PR_NUMBER,
competing_live_locks=[],
candidate_branches=[BRANCH],
current_pid=os.getpid(),
remote_branch_exists=True,
sync_provenance=obs,
)
self.assertEqual(res["outcome"], issue_lock_recovery.RECOVERY_SANCTIONED, res["reasons"])
self.assertEqual(
res["evidence"]["head_relation"],
issue_lock_recovery.HEAD_RELATION_REMOTE_MERGE_SYNCED,
)
self.assertEqual(res["evidence"]["accepted_head"], self.head_b)
def test_second_base_sync_head_refresh_after_recovery(self):
"""AC2: after recovery, second base sync (remote B -> C) updates durable lock head."""
lock_dir = tempfile.mkdtemp()
lock_data = make_dead_lock(self.tmp, synced_pr_head=self.head_b)
# Rebind to current PID as gitea_lock_issue does upon sanctioned recovery
lock_data["session_pid"] = os.getpid()
lock_data["pid"] = os.getpid()
issue_lock_store.save_lock_file(
issue_lock_store.lock_file_path(
remote=REMOTE, org=ORG, repo=REPO, issue_number=ISSUE, lock_dir=lock_dir
),
lock_data,
)
refresh_res = issue_lock_store.apply_durable_lock_head_refresh(
remote=REMOTE,
org=ORG,
repo=REPO,
issue_number=ISSUE,
branch_name=BRANCH,
worktree_path=self.tmp,
pr_number=PR_NUMBER,
identity=IDENTITY,
profile=PROFILE,
current_pid=os.getpid(),
expected_old_head=self.head_b,
new_head=self.head_c,
synced_at=future_ts(0),
base_head=self.shas["master_2"],
lock_dir=lock_dir,
)
self.assertTrue(refresh_res["refreshed"], refresh_res["reasons"])
self.assertTrue(refresh_res["read_after_write_ok"])
loaded = issue_lock_store.load_issue_lock(
remote=REMOTE, org=ORG, repo=REPO, issue_number=ISSUE, lock_dir=lock_dir
)
self.assertEqual(loaded.get("synced_pr_head"), self.head_c)
def test_dirty_worktree_blocks_recovery(self):
"""AC3: tracked dirty edits block dead-session recovery."""
lock = make_dead_lock(self.tmp)
obs = issue_lock_worktree.read_merge_sync_provenance(
self.tmp, prior_head_sha=self.head_a, synced_head_sha=self.head_b
)
res = issue_lock_recovery.assess_dead_session_lock_recovery(
lock,
issue_number=ISSUE,
branch_name=BRANCH,
worktree_path=self.tmp,
remote=REMOTE,
org=ORG,
repo=REPO,
identity=IDENTITY,
profile=PROFILE,
current_branch=BRANCH,
porcelain_status=" M feature.txt\n",
head_sha=self.head_a,
remote_head_sha=self.head_b,
pr_head_sha=self.head_b,
pr_number=PR_NUMBER,
competing_live_locks=[],
candidate_branches=[BRANCH],
current_pid=os.getpid(),
remote_branch_exists=True,
sync_provenance=obs,
)
self.assertEqual(res["outcome"], issue_lock_recovery.REFUSED)
self.assertTrue(any("dirty" in r or "tracked" in r for r in res["reasons"]))
def test_foreign_reclaimer_blocks_recovery(self):
"""AC4: a foreign claimant cannot recover a dead-session lock."""
lock = make_dead_lock(self.tmp)
obs = issue_lock_worktree.read_merge_sync_provenance(
self.tmp, prior_head_sha=self.head_a, synced_head_sha=self.head_b
)
res = issue_lock_recovery.assess_dead_session_lock_recovery(
lock,
issue_number=ISSUE,
branch_name=BRANCH,
worktree_path=self.tmp,
remote=REMOTE,
org=ORG,
repo=REPO,
identity="intruder-user",
profile=PROFILE,
current_branch=BRANCH,
porcelain_status="",
head_sha=self.head_a,
remote_head_sha=self.head_b,
pr_head_sha=self.head_b,
pr_number=PR_NUMBER,
competing_live_locks=[],
candidate_branches=[BRANCH],
current_pid=os.getpid(),
remote_branch_exists=True,
sync_provenance=obs,
)
self.assertEqual(res["outcome"], issue_lock_recovery.REFUSED)
def test_non_merge_rebased_remote_head_blocks_recovery(self):
"""AC5: a rebased/force-pushed remote head (not a merge commit) is refused."""
lock = make_dead_lock(self.tmp)
obs = issue_lock_worktree.read_merge_sync_provenance(
self.tmp, prior_head_sha=self.head_a, synced_head_sha=self.shas["rebased_head"]
)
self.assertFalse(obs["is_merge_sync"])
res = issue_lock_recovery.assess_dead_session_lock_recovery(
lock,
issue_number=ISSUE,
branch_name=BRANCH,
worktree_path=self.tmp,
remote=REMOTE,
org=ORG,
repo=REPO,
identity=IDENTITY,
profile=PROFILE,
current_branch=BRANCH,
porcelain_status="",
head_sha=self.head_a,
remote_head_sha=self.shas["rebased_head"],
pr_head_sha=self.shas["rebased_head"],
pr_number=PR_NUMBER,
competing_live_locks=[],
candidate_branches=[BRANCH],
current_pid=os.getpid(),
remote_branch_exists=True,
sync_provenance=obs,
)
self.assertEqual(res["outcome"], issue_lock_recovery.REFUSED)
if __name__ == "__main__":
unittest.main()
+6
View File
@@ -24,6 +24,8 @@ def _lease(expires_at: str) -> dict:
def _lock_record(**overrides) -> dict:
# #860: live locks require a usable session pid; PID-less records are never
# classified live merely because expiry/heartbeat fields are present.
record = {
"issue_number": 420,
"branch_name": "feat/issue-420-server-code-parity",
@@ -31,6 +33,8 @@ def _lock_record(**overrides) -> dict:
"org": "Scaled-Tech-Consulting",
"repo": "Gitea-Tools",
"worktree_path": "/tmp/wt-420",
"session_pid": os.getpid(),
"pid": os.getpid(),
"work_lease": _lease("2999-01-01T00:00:00Z"),
}
record.update(overrides)
@@ -88,6 +92,8 @@ class TestIssueLockStore(unittest.TestCase):
existing = _lock_record(
branch_name="feat/issue-420-other",
worktree_path="/tmp/other",
session_pid=os.getpid(),
pid=os.getpid(),
work_lease=_lease("2999-01-01T00:00:00Z"),
)
path = ils.lock_file_path(
+107
View File
@@ -0,0 +1,107 @@
"""Documentation acceptance for the MCP restart governance ADR (#656).
Enforces issue #656 acceptance criteria:
* AC1 policy document exists with an authorization matrix and the recorded
v1 decision (controller approval + automated safety gates).
* AC2 restart is stated as a last resort with enumerated narrower recoveries.
* AC3 a unilateral LLM full restart with affected sessions is forbidden.
* AC4 break-glass conditions are listed.
* AC5 the ADR is linked to #655, #652, #653, #630, #642, and is cross-linked
from the safety model and the web-console deployment boundary docs.
"""
from pathlib import Path
REPO_ROOT = Path(__file__).resolve().parent.parent
ADR = REPO_ROOT / "docs" / "architecture" / "mcp-restart-governance.md"
ADR_BASENAME = "mcp-restart-governance.md"
CROSS_LINK_DOCS = (
REPO_ROOT / "docs" / "safety-model.md",
REPO_ROOT / "docs" / "webui-deployment.md",
)
LINKED_ISSUES = ("#655", "#652", "#653", "#630", "#642")
POLICY_IDS = ("RG-01", "RG-02", "RG-03", "RG-04", "RG-05", "RG-06", "RG-07", "RG-08")
def _read(path: Path) -> str:
assert path.is_file(), f"missing {path.relative_to(REPO_ROOT)}"
return path.read_text(encoding="utf-8")
def test_ac1_adr_exists_with_matrix_and_v1_decision():
text = _read(ADR)
lower = text.lower()
assert text.lstrip().startswith("#"), "ADR lacks a title"
assert "#656" in text
assert "authorization matrix" in lower
# The matrix is a real table with the worker and privileged roles.
for role in ("author", "reviewer", "merger", "reconciler", "controller",
"operator", "admin"):
assert role in lower, f"authorization matrix missing role {role!r}"
# Recorded v1 decision.
assert "restart-governance/v1" in text
assert "controller approval" in lower and "automated safety gates" in lower
def test_ac2_restart_is_last_resort_with_narrower_recoveries():
text = _read(ADR)
lower = text.lower()
assert "last resort" in lower
# Enumerated narrower recoveries precede full restart on the ladder.
for rung in ("reconnect", "rebind", "scoped restart", "full restart",
"host"):
assert rung in lower, f"recovery ladder missing rung {rung!r}"
def test_ac3_forbids_unilateral_llm_full_restart_with_affected_sessions():
text = _read(ADR)
lower = text.lower()
assert "forbidden" in lower
assert "llm" in lower and "restart" in lower
assert "unilateral" in lower
# A worker role must not perform or authorize full/host restart.
assert "must not" in lower
def test_ac4_break_glass_conditions_listed():
text = _read(ADR)
lower = text.lower()
assert "break-glass" in lower
assert "incident" in lower
assert "audit" in lower
def test_ac5_adr_links_issue_lineage():
text = _read(ADR)
for issue in LINKED_ISSUES:
assert issue in text, f"ADR must link issue {issue}"
def test_ac5_safety_model_and_deployment_cross_link_adr():
for path in CROSS_LINK_DOCS:
text = _read(path)
assert ADR_BASENAME in text, (
f"{path.relative_to(REPO_ROOT)} must cross-link {ADR_BASENAME} "
f"(issue #656 acceptance criterion 5)"
)
def test_policy_ids_present_for_enforcement_code():
text = _read(ADR)
for pid in POLICY_IDS:
assert pid in text, f"policy id {pid} missing from ADR"
def test_failure_behavior_denies_on_ambiguity():
text = _read(ADR)
lower = text.lower()
assert "ambiguous" in lower and "deny" in lower
def test_cross_links_do_not_embed_secrets():
for path in (ADR,) + CROSS_LINK_DOCS:
text = _read(path)
for marker in ("ghp_", "BEGIN PRIVATE KEY", "Authorization: Bearer"):
assert marker not in text, f"{path} contains {marker!r}"
+146
View File
@@ -0,0 +1,146 @@
"""Tests for the MCP restart-path inventory and guards (#657).
Covers:
* the registry is well-formed and every path is classified;
* unknown restart attempts fail closed (AC "fail closed on unknown restart");
* the previously-unguarded full-restart primitives stay guarded/absent
against the real source tree (AC "tests for at least one previously
unguarded path");
* pkill of the daemon is still classified as contamination (#630, AC3);
* the inventory doc and module stay in lock-step.
"""
import os
import tempfile
import unittest
from pathlib import Path
import mcp_restart_paths as rp
import runtime_recovery_guard
REPO_ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
DOC_PATH = os.path.join(REPO_ROOT, "docs", "mcp-restart-path-inventory.md")
class TestRegistryWellformed(unittest.TestCase):
def test_registry_is_wellformed(self):
# Must not raise.
rp.assert_registry_wellformed()
def test_every_path_has_valid_classification(self):
for path in rp.iter_restart_paths():
self.assertIn(path.classification, rp.VALID_CLASSIFICATIONS)
self.assertTrue(path.guard.strip(), path.path_id)
self.assertTrue(path.references, path.path_id)
self.assertTrue(path.locations, path.path_id)
def test_ids_are_unique(self):
ids = [p.path_id for p in rp.iter_restart_paths()]
self.assertEqual(len(ids), len(set(ids)))
def test_covers_every_classification(self):
present = {p.classification for p in rp.iter_restart_paths()}
self.assertEqual(present, set(rp.VALID_CLASSIFICATIONS))
class TestUnknownAttemptFailsClosed(unittest.TestCase):
def test_unknown_path_raises(self):
with self.assertRaises(rp.UnknownRestartPathError):
rp.assert_restart_attempt_registered("totally_novel_restart_hack")
def test_get_unknown_raises(self):
with self.assertRaises(rp.UnknownRestartPathError):
rp.get_restart_path("nope")
def test_registered_attempt_returns_path(self):
path = rp.assert_restart_attempt_registered("manual_daemon_kill")
self.assertEqual(path.classification, rp.CLASS_FORBIDDEN)
class TestDaemonNeverSelfReplaces(unittest.TestCase):
"""Previously-unguarded full-restart primitive: daemon self-replacement."""
def test_no_self_replacement_in_source(self):
# The live daemon modules must contain no os.execv/os.kill/os._exit
# self-restart call. Must not raise.
rp.assert_no_daemon_self_replacement(REPO_ROOT)
def test_scanner_flags_injected_violation(self):
# Guard the guard: prove the scanner catches a real self-replace call.
with tempfile.TemporaryDirectory() as tmp:
bad = Path(tmp) / "gitea_mcp_server.py"
bad.write_text(
"import os\n"
"def restart():\n"
" os.execv('/usr/bin/python', ['python'])\n",
encoding="utf-8",
)
found = rp.scan_daemon_self_replacement(tmp)
self.assertTrue(found)
with self.assertRaises(AssertionError):
rp.assert_no_daemon_self_replacement(tmp)
def test_scanner_ignores_comment_and_docstring_mentions(self):
with tempfile.TemporaryDirectory() as tmp:
ok = Path(tmp) / "gitea_mcp_server.py"
ok.write_text(
"import os\n"
"# NOT os.execv() to re-point the interpreter here.\n"
'"""Never calls os._exit to restart."""\n'
"value = 1\n",
encoding="utf-8",
)
self.assertEqual(rp.scan_daemon_self_replacement(tmp), [])
class TestLegacyAutoRestartHelperRemoved(unittest.TestCase):
"""Previously-unguarded full-restart path: _trigger_mcp_auto_restart."""
def test_helper_absent_in_source(self):
# Must not raise: helper was removed in #685.
rp.assert_auto_restart_helper_absent(REPO_ROOT)
def test_scanner_flags_reintroduced_helper(self):
with tempfile.TemporaryDirectory() as tmp:
bad = Path(tmp) / "mcp_server.py"
bad.write_text(
"def _trigger_mcp_auto_restart():\n return True\n",
encoding="utf-8",
)
with self.assertRaises(AssertionError):
rp.assert_auto_restart_helper_absent(tmp)
class TestPkillStaysForbidden(unittest.TestCase):
"""AC3: pkill of the daemon remains forbidden/contaminating (#630)."""
def test_manual_daemon_kill_registered_as_forbidden(self):
path = rp.get_restart_path("manual_daemon_kill")
self.assertEqual(path.classification, rp.CLASS_FORBIDDEN)
def test_pkill_classified_as_contamination(self):
assessment = runtime_recovery_guard.assess_recovery_command(
"pkill -f mcp_server.py"
)
self.assertTrue(assessment["contaminated"])
def test_read_only_probe_not_contamination(self):
assessment = runtime_recovery_guard.assess_recovery_command(
"ps aux | grep mcp_server"
)
self.assertFalse(assessment["contaminated"])
class TestInventoryDocInSync(unittest.TestCase):
def test_doc_exists(self):
self.assertTrue(os.path.exists(DOC_PATH), DOC_PATH)
def test_doc_mentions_every_path_id(self):
with open(DOC_PATH, encoding="utf-8") as handle:
doc = handle.read()
for path in rp.iter_restart_paths():
self.assertIn(path.path_id, doc, f"doc missing {path.path_id}")
if __name__ == "__main__":
unittest.main()
+53
View File
@@ -12,6 +12,59 @@ import merged_cleanup_reconcile as mcr # noqa: E402
class TestMergedCleanupAssessment(unittest.TestCase):
def test_issue_851_plan_order_worktree_then_reassess_then_remote(self):
"""#851 dry-run plan: remove worktree, reassess ownership, then remote."""
plan = mcr.plan_cleanup_execution_order(
remote_assessment={"safe_to_delete_remote": True},
local_assessment={"safe_to_remove_worktree": True},
)
actions = [s["action"] for s in plan]
self.assertEqual(
actions,
[
"remove_local_worktree",
"reassess_branch_ownership",
"delete_remote_branch",
],
)
self.assertEqual(plan[0]["phase"], 1)
self.assertEqual(plan[-1]["phase"], 3)
self.assertIn("independently_safe", plan[0]["reason"])
self.assertIn("reassessment", plan[-1]["reason"])
def test_issue_851_plan_remote_only_when_worktree_not_safe(self):
plan = mcr.plan_cleanup_execution_order(
remote_assessment={"safe_to_delete_remote": True},
local_assessment={"safe_to_remove_worktree": False},
)
self.assertEqual([s["action"] for s in plan], ["delete_remote_branch"])
self.assertNotIn("reassess_branch_ownership", [s["action"] for s in plan])
def test_issue_851_plan_worktree_only_when_remote_not_safe(self):
plan = mcr.plan_cleanup_execution_order(
remote_assessment={"safe_to_delete_remote": False},
local_assessment={"safe_to_remove_worktree": True},
)
self.assertEqual([s["action"] for s in plan], ["remove_local_worktree"])
def test_issue_851_entry_includes_planned_execution_order(self):
entry = mcr.build_pr_cleanup_entry(
pr={
"number": 848,
"title": "Closes #844",
"body": "",
"merged_at": "2026-07-23T00:00:00Z",
"head": {"ref": "fix/issue-844-x", "sha": "a" * 40},
},
project_root="/tmp/not-a-real-root",
open_pr_heads=set(),
remote_branch_exists=True,
head_on_master=True,
delete_capability_allowed=True,
)
self.assertIn("planned_execution_order", entry)
self.assertIsInstance(entry["planned_execution_order"], list)
def test_extract_linked_issue_from_closes(self):
issue = mcr.extract_linked_issue(
"feat: cleanup (Closes #269)",
+16 -2
View File
@@ -37,6 +37,7 @@ def _live_lock(
"operation_type": issue_lock_store.AUTHOR_ISSUE_WORK_LEASE,
"acquired_at": now.isoformat(),
"expires_at": (now + timedelta(hours=2)).isoformat(),
"session_pid": os.getpid(),
"owner_pid": os.getpid(),
"status": "active",
}
@@ -177,11 +178,24 @@ class TestAuthorOwnershipIssuePrMismatch(unittest.TestCase):
self.assertFalse(result["proven"], result)
self.assertTrue(any("branch" in r for r in result["reasons"]))
def test_no_lock_fail_closed(self):
def test_pidless_durable_lock_rejected(self):
"""A lock without any PID identity must be classified as malformed/non-live and fail closed."""
lock = _live_lock(issue_number=727)
lock.pop("session_pid", None)
lock.pop("owner_pid", None)
lock.pop("pid", None)
path = issue_lock_store.lock_file_path(
remote="prgs",
org="Scaled-Tech-Consulting",
repo="Gitea-Tools",
issue_number=727,
lock_dir=self.lock_dir,
)
issue_lock_store.save_lock_file(path, lock)
result = mcp._prove_author_ownership_for_pr(
pr_number=728,
pr_title="feat: pr sync",
pr_body="Closes #727",
pr_body="Fixes #727",
source_branch="feat/issue-727-pr-sync-status",
remote="prgs",
host=None,
+153
View File
@@ -19,6 +19,7 @@ from pr_work_lease import ( # noqa: E402
assess_reviewer_mutation_blocked,
assess_reviewer_stale_head_final_report,
format_conflict_fix_lease_body,
find_active_conflict_fix_lease,
parse_conflict_fix_lease_comment,
parse_reviewer_lease_comment,
)
@@ -203,5 +204,157 @@ class TestFormatLease(unittest.TestCase):
self.assertEqual(parsed["pr_number"], 376)
class TestConflictFixLeaseLifecycle(unittest.TestCase):
def test_claim_followed_by_matching_release(self):
claim_body = _conflict_fix_body(phase="claimed", worktree="branches/fix-376")
expires = (NOW + timedelta(minutes=60)).isoformat().replace("+00:00", "Z")
release_body = "\n".join([
CONFLICT_FIX_LEASE_MARKER,
"pr: #376",
"branch: feat/fix-376",
"worktree: branches/fix-376",
"profile: prgs-author",
"phase: released",
f"head_before: {HEAD_A}",
f"head_after: {HEAD_B}",
f"expires_at: {expires}",
])
comments = [{"body": claim_body}, {"body": release_body}]
lease = find_active_conflict_fix_lease(comments, pr_number=376, now=NOW)
self.assertIsNone(lease)
def test_expired_claim_without_release(self):
past_expires = (NOW - timedelta(minutes=10)).isoformat().replace("+00:00", "Z")
claim_body = "\n".join([
CONFLICT_FIX_LEASE_MARKER,
"pr: #376",
"phase: claimed",
f"head_before: {HEAD_A}",
f"expires_at: {past_expires}",
"profile: prgs-author",
])
comments = [{"body": claim_body}]
lease = find_active_conflict_fix_lease(comments, pr_number=376, now=NOW)
self.assertIsNone(lease)
def test_mismatched_release_different_head(self):
claim_body = _conflict_fix_body(phase="claimed", worktree="branches/fix-376")
expires = (NOW + timedelta(minutes=60)).isoformat().replace("+00:00", "Z")
release_body = "\n".join([
CONFLICT_FIX_LEASE_MARKER,
"pr: #376",
"profile: prgs-author",
"phase: released",
f"head_before: {HEAD_B}",
f"expires_at: {expires}",
])
comments = [{"body": claim_body}, {"body": release_body}]
lease = find_active_conflict_fix_lease(comments, pr_number=376, now=NOW)
self.assertIsNotNone(lease)
self.assertEqual(lease["phase"], "claimed")
def test_mismatched_release_different_branch(self):
claim_body = "\n".join([
CONFLICT_FIX_LEASE_MARKER,
"pr: #376",
"branch: feat/branch-A",
"phase: claimed",
f"head_before: {HEAD_A}",
f"expires_at: {(NOW + timedelta(minutes=60)).isoformat().replace('+00:00', 'Z')}",
"profile: prgs-author",
])
release_body = "\n".join([
CONFLICT_FIX_LEASE_MARKER,
"pr: #376",
"branch: feat/branch-B",
"phase: released",
f"head_before: {HEAD_A}",
f"expires_at: {(NOW + timedelta(minutes=60)).isoformat().replace('+00:00', 'Z')}",
"profile: prgs-author",
])
comments = [{"body": claim_body}, {"body": release_body}]
lease = find_active_conflict_fix_lease(comments, pr_number=376, now=NOW)
self.assertIsNotNone(lease)
self.assertEqual(lease["phase"], "claimed")
def test_release_followed_by_newer_claim(self):
claim_1 = _conflict_fix_body(phase="claimed", worktree="branches/fix-376")
expires = (NOW + timedelta(minutes=60)).isoformat().replace("+00:00", "Z")
release_1 = "\n".join([
CONFLICT_FIX_LEASE_MARKER,
"pr: #376",
"profile: prgs-author",
"phase: released",
f"head_before: {HEAD_A}",
f"head_after: {HEAD_B}",
f"expires_at: {expires}",
])
claim_2 = "\n".join([
CONFLICT_FIX_LEASE_MARKER,
"pr: #376",
"profile: prgs-author",
"phase: claimed",
f"head_before: {HEAD_B}",
f"expires_at: {expires}",
])
comments = [{"body": claim_1}, {"body": release_1}, {"body": claim_2}]
lease = find_active_conflict_fix_lease(comments, pr_number=376, now=NOW)
self.assertIsNotNone(lease)
self.assertEqual(lease["head_before"], HEAD_B)
def test_malformed_or_ambiguous_markers(self):
malformed_release = "\n".join([
CONFLICT_FIX_LEASE_MARKER,
"pr: #376",
"phase: released",
# missing head_before and profile
])
claim_body = _conflict_fix_body(phase="claimed")
comments = [{"body": claim_body}, {"body": malformed_release}]
lease = find_active_conflict_fix_lease(comments, pr_number=376, now=NOW)
self.assertIsNotNone(lease)
def test_pr818_historical_sequence(self):
comment_14696 = "\n".join([
"<!-- mcp-conflict-fix-lease:v1 -->",
"pr: #818",
"branch: feat/issue-638-webui-app-shell-phase1",
"worktree: /Users/jasonwalker/Development/Gitea-Tools/branches/issue-638-webui-app-shell-phase1",
"profile: prgs-author",
"session_id: unknown",
"phase: claimed",
"head_before: 08061b7b8aebdd099a37d1abf5dafcf38e4fd3fb",
"expires_at: 2026-07-23T07:12:13Z",
"reviewer_active: no",
])
comment_14730 = "\n".join([
"<!-- mcp-conflict-fix-lease:v1 -->",
"pr: #818",
"branch: feat/issue-638-webui-app-shell-phase1",
"worktree: /Users/jasonwalker/Development/Gitea-Tools/branches/issue-638-webui-app-shell-phase1",
"profile: prgs-author",
"session_id: prgs-author-61241-e5129c60",
"phase: released",
"head_before: 08061b7b8aebdd099a37d1abf5dafcf38e4fd3fb",
"head_after: 64b6eb5d5402663098de5ded3b0617cc3b3df98f",
"expires_at: 2026-07-23T06:05:00Z",
"reviewer_active: no",
])
comments = [{"body": comment_14696}, {"body": comment_14730}]
check_now = datetime(2026, 7, 23, 6, 30, tzinfo=timezone.utc)
lease = find_active_conflict_fix_lease(comments, pr_number=818, now=check_now)
self.assertIsNone(lease)
reviewer_gate = assess_reviewer_mutation_blocked(
pr_number=818,
comments=comments,
reviewed_head_sha="64b6eb5d5402663098de5ded3b0617cc3b3df98f",
live_head_sha="64b6eb5d5402663098de5ded3b0617cc3b3df98f",
mutation="approve",
now=check_now,
)
self.assertTrue(reviewer_gate["mutation_allowed"])
if __name__ == "__main__":
unittest.main()
+340
View File
@@ -0,0 +1,340 @@
"""Tests for the MCP restart coordinator and impact analysis (#658).
Multi-session fixtures exercise every verdict branch: safe, unsafe (live work),
override, and the fail-closed deny on incomplete inventory. Also covers the
critical-section deny path and the new ``ControlPlaneDB.list_sessions``.
"""
from __future__ import annotations
import os
import tempfile
import unittest
from datetime import datetime, timedelta, timezone
import restart_coordinator as rc
from control_plane_db import ControlPlaneDB
NOW = datetime(2026, 7, 24, 6, 0, 0, tzinfo=timezone.utc)
def _ts(dt: datetime) -> str:
return dt.isoformat()
def _live_pid() -> int:
return os.getpid()
def _dead_pid() -> int:
# A pid that is essentially never alive. os.kill(0) on it raises
# ProcessLookupError → is_process_alive False.
return 2_000_000_000
def _session(session_id, *, pid, status="active", heartbeat=None, role="author"):
return {
"session_id": session_id,
"role": role,
"profile": "prgs-author",
"pid": pid,
"status": status,
"last_heartbeat_at": _ts(heartbeat or NOW),
}
def _lease(
lease_id,
*,
session_id,
freshness,
kind="issue",
number=658,
phase="allocated",
worktree=None,
role="author",
):
return {
"lease_id": lease_id,
"session_id": session_id,
"role": role,
"phase": phase,
"work_kind": kind,
"work_number": number,
"worktree_path": worktree,
"freshness": {"freshness": freshness},
}
class EvaluateRestartImpactTest(unittest.TestCase):
def test_incomplete_inventory_denies_fail_closed(self) -> None:
report = rc.evaluate_restart_impact(
{"inventory_complete": False, "incomplete_reasons": ["db down"]},
now=NOW,
)
self.assertEqual(report.verdict, rc.VERDICT_UNSAFE)
self.assertFalse(report.allow_restart)
self.assertFalse(report.restart_performed)
self.assertIn("db down", report.incomplete_reasons)
self.assertTrue(
any("fail closed" in reasoning for reasoning in report.reasons)
)
def test_missing_completeness_flag_denies(self) -> None:
# No inventory_complete key at all → treated as incomplete.
report = rc.evaluate_restart_impact({}, now=NOW)
self.assertEqual(report.verdict, rc.VERDICT_UNSAFE)
self.assertFalse(report.allow_restart)
def test_no_other_work_is_safe(self) -> None:
report = rc.evaluate_restart_impact(
{
"inventory_complete": True,
"sessions": [_session("requester", pid=_live_pid())],
"leases": [],
},
now=NOW,
requesting_session_id="requester",
)
self.assertEqual(report.verdict, rc.VERDICT_SAFE)
self.assertTrue(report.allow_restart)
self.assertEqual(report.blast_radius, rc.BLAST_NONE)
self.assertEqual(report.affected_issues, [])
def test_dead_foreign_session_and_lease_are_not_disruptive(self) -> None:
report = rc.evaluate_restart_impact(
{
"inventory_complete": True,
"sessions": [
_session("requester", pid=_live_pid()),
_session("dead", pid=_dead_pid()),
],
"leases": [
_lease("l-dead", session_id="dead", freshness="stale_dead_process")
],
},
now=NOW,
requesting_session_id="requester",
)
self.assertEqual(report.verdict, rc.VERDICT_SAFE)
self.assertTrue(report.allow_restart)
self.assertEqual(report.counts["leases_disruptive"], 0)
self.assertEqual(report.counts["sessions_live_other"], 0)
def test_live_foreign_lease_denies_without_override(self) -> None:
report = rc.evaluate_restart_impact(
{
"inventory_complete": True,
"sessions": [
_session("requester", pid=_live_pid()),
_session("worker", pid=_live_pid()),
],
"leases": [
_lease(
"l1",
session_id="worker",
freshness="active",
worktree="/tmp/wt-658",
phase="implementing",
)
],
},
now=NOW,
requesting_session_id="requester",
)
self.assertEqual(report.verdict, rc.VERDICT_UNSAFE)
self.assertFalse(report.allow_restart)
# Critical section detected: active lease with a live owner.
self.assertEqual(len(report.critical_sections), 1)
self.assertEqual(report.affected_issues, [658])
self.assertEqual(report.counts["mutations"], 1)
self.assertTrue(report.override_would_allow)
self.assertEqual(report.blast_radius, rc.BLAST_HIGH)
# Placeholder ack state for the affected session.
self.assertEqual(report.ack_state.get("worker"), "pending")
def test_operator_override_allows_despite_live_work(self) -> None:
inv = {
"inventory_complete": True,
"sessions": [
_session("requester", pid=_live_pid()),
_session("worker", pid=_live_pid()),
],
"leases": [_lease("l1", session_id="worker", freshness="active")],
}
report = rc.evaluate_restart_impact(
inv,
now=NOW,
requesting_session_id="requester",
operator_override=True,
)
self.assertEqual(report.verdict, rc.VERDICT_OVERRIDE)
self.assertTrue(report.allow_restart)
self.assertFalse(report.restart_performed)
def test_deny_when_critical_section_open(self) -> None:
# A single live author lease in a mutating phase is a critical section
# that must deny an un-overridden restart.
report = rc.evaluate_restart_impact(
{
"inventory_complete": True,
"sessions": [_session("worker", pid=_live_pid())],
"leases": [
_lease(
"l1",
session_id="worker",
freshness="active",
phase="merging",
kind="pr",
number=900,
)
],
},
now=NOW,
requesting_session_id="requester",
)
self.assertEqual(report.verdict, rc.VERDICT_UNSAFE)
self.assertFalse(report.allow_restart)
self.assertEqual(report.affected_prs, [900])
self.assertEqual(len(report.critical_sections), 1)
def test_terminal_lock_makes_restart_unsafe(self) -> None:
report = rc.evaluate_restart_impact(
{
"inventory_complete": True,
"sessions": [_session("requester", pid=_live_pid())],
"leases": [],
"terminal_lock": {"terminal_pr": 812},
},
now=NOW,
requesting_session_id="requester",
)
self.assertEqual(report.verdict, rc.VERDICT_UNSAFE)
self.assertFalse(report.allow_restart)
self.assertIsNotNone(report.terminal_lock)
self.assertTrue(
any("terminal" in reasoning for reasoning in report.reasons)
)
def test_other_live_session_without_lease_is_disruptive(self) -> None:
report = rc.evaluate_restart_impact(
{
"inventory_complete": True,
"sessions": [
_session("requester", pid=_live_pid()),
_session("idle-but-live", pid=_live_pid()),
],
"leases": [],
},
now=NOW,
requesting_session_id="requester",
)
self.assertEqual(report.verdict, rc.VERDICT_UNSAFE)
self.assertEqual(report.counts["sessions_live_other"], 1)
def test_stale_heartbeat_session_not_counted_live(self) -> None:
stale = NOW - timedelta(hours=2)
report = rc.evaluate_restart_impact(
{
"inventory_complete": True,
"sessions": [
_session("requester", pid=_live_pid()),
_session("stale", pid=_live_pid(), heartbeat=stale),
],
"leases": [],
},
now=NOW,
requesting_session_id="requester",
)
self.assertEqual(report.verdict, rc.VERDICT_SAFE)
self.assertEqual(report.counts["sessions_live_other"], 0)
def test_prior_recovery_attempts_echoed(self) -> None:
report = rc.evaluate_restart_impact(
{
"inventory_complete": True,
"sessions": [_session("requester", pid=_live_pid())],
"leases": [],
"prior_recovery_attempts": [
{"kind": "client_reconnect", "at": _ts(NOW)}
],
},
now=NOW,
requesting_session_id="requester",
)
self.assertEqual(len(report.prior_recovery_attempts), 1)
self.assertEqual(report.counts["prior_recovery_attempts"], 1)
def test_bare_string_freshness_accepted(self) -> None:
lease = _lease("l1", session_id="worker", freshness="active")
lease["freshness"] = "active" # bare string, not a dict
report = rc.evaluate_restart_impact(
{
"inventory_complete": True,
"sessions": [_session("worker", pid=_live_pid())],
"leases": [lease],
},
now=NOW,
requesting_session_id="requester",
)
self.assertEqual(report.counts["leases_disruptive"], 1)
def test_as_dict_is_serializable_dto(self) -> None:
import json
report = rc.evaluate_restart_impact(
{
"inventory_complete": True,
"sessions": [_session("requester", pid=_live_pid())],
"leases": [],
},
now=NOW,
requesting_session_id="requester",
)
payload = report.as_dict()
# Round-trips through JSON — safe for the console DTO.
encoded = json.dumps(payload)
decoded = json.loads(encoded)
self.assertEqual(decoded["verdict"], rc.VERDICT_SAFE)
self.assertIn("audit_record", decoded)
self.assertEqual(decoded["audit_record"]["event"], "restart_impact_evaluated")
self.assertFalse(decoded["restart_performed"])
self.assertIn("coordinator_version", decoded)
class ListSessionsTest(unittest.TestCase):
def setUp(self) -> None:
self._tmp = tempfile.TemporaryDirectory()
self.db = ControlPlaneDB(os.path.join(self._tmp.name, "cp.sqlite3"))
def tearDown(self) -> None:
self._tmp.cleanup()
def test_list_sessions_filters_by_status(self) -> None:
self.db.upsert_session(session_id="a", role="author", pid=1, status="active")
self.db.upsert_session(session_id="b", role="author", pid=2, status="ended")
active = self.db.list_sessions(statuses=("active",))
ids = {row["session_id"] for row in active}
self.assertEqual(ids, {"a"})
every = self.db.list_sessions()
self.assertEqual({row["session_id"] for row in every}, {"a", "b"})
def test_list_sessions_feeds_coordinator(self) -> None:
self.db.upsert_session(
session_id="requester", role="author", pid=os.getpid(), status="active"
)
report = rc.evaluate_restart_impact(
{
"inventory_complete": True,
"sessions": self.db.list_sessions(statuses=("active",)),
"leases": [],
},
now=NOW,
requesting_session_id="requester",
)
self.assertEqual(report.counts["sessions_total"], 1)
if __name__ == "__main__": # pragma: no cover
unittest.main()
@@ -139,6 +139,8 @@ EXPECTED_ROLE_EXCLUSIVE_TASKS = frozenset(
"gitea_release_merger_pr_lease",
"create_branch",
"push_branch",
"bootstrap_author_issue_worktree",
"gitea_bootstrap_author_issue_worktree",
# #812 AC20: publishing an unpublished local head is author-only for the
# same reason every other push is — it writes a branch to the remote.
"publish_unpublished_branch",
+252
View File
@@ -0,0 +1,252 @@
"""Unit and integration tests for Model Usage & Performance Analytics (#651)."""
from __future__ import annotations
import os
import tempfile
import unittest
from starlette.testclient import TestClient
import control_plane_db
from webui.analytics_loader import (
ANALYTICS_SCHEMA_VERSION,
compute_percentile,
load_analytics,
record_usage,
)
from webui.app import create_app
from webui import console_redaction
class AnalyticsLoaderTest(unittest.TestCase):
def setUp(self) -> None:
self.temp_dir = tempfile.TemporaryDirectory()
self.db_path = os.path.join(self.temp_dir.name, "test_control_plane.sqlite3")
os.environ["GITEA_CONTROL_PLANE_DB"] = self.db_path
self.db = control_plane_db.ControlPlaneDB(db_path=self.db_path)
def tearDown(self) -> None:
self.temp_dir.cleanup()
def test_compute_percentile(self) -> None:
self.assertIsNone(compute_percentile([], 50.0))
self.assertEqual(compute_percentile([100], 50.0), 100.0)
# 2 elements: [100, 200]
self.assertEqual(compute_percentile([100, 200], 50.0), 150.0)
# 100 elements: 1..100
vals = list(range(1, 101))
self.assertEqual(compute_percentile(vals, 50.0), 50.5)
self.assertAlmostEqual(compute_percentile(vals, 90.0), 90.1)
def test_record_and_aggregate_usage(self) -> None:
# Record event 1 (complete data)
u1 = record_usage(
db_path=self.db_path,
remote="dadeschools",
org="Scaled-Tech-Consulting",
repo="Gitea-Tools",
role="author",
model="gemini-3.6-flash",
issue_number=651,
stage="implementation",
input_tokens=1000,
output_tokens=500,
estimated_cost_usd=0.0015,
latency_ms=200,
duration_ms=3000,
metadata={"secret_key": "secret123", "note": "token=secret123"},
)
self.assertGreater(u1, 0)
# Record event 2 (missing tokens and cost -> unknown)
u2 = record_usage(
db_path=self.db_path,
remote="dadeschools",
org="Scaled-Tech-Consulting",
repo="Gitea-Tools",
role="reviewer",
model="claude-3-5-sonnet",
pr_number=846,
stage="review",
latency_ms=500,
duration_ms=6000,
)
self.assertGreater(u2, u1)
snapshot = load_analytics(
db_path=self.db_path,
remote="dadeschools",
org="Scaled-Tech-Consulting",
repo="Gitea-Tools",
)
self.assertTrue(snapshot.ok)
self.assertEqual(snapshot.schema_version, ANALYTICS_SCHEMA_VERSION)
self.assertEqual(snapshot.total_events, 2)
# Verify overall summary
summary = snapshot.overall_summary
self.assertEqual(summary.total_events, 2)
self.assertEqual(summary.events_with_tokens, 1)
self.assertEqual(summary.total_tokens, 1500)
self.assertEqual(summary.events_with_cost, 1)
self.assertEqual(summary.estimated_cost_usd, 0.0015)
self.assertEqual(summary.events_with_latency, 2)
self.assertEqual(summary.latency_p50_ms, 350.0)
# Verify missing data handling (AC 3: not zero-fabricated)
reviewer_model = snapshot.by_model.get("claude-3-5-sonnet")
self.assertIsNotNone(reviewer_model)
self.assertEqual(reviewer_model.total_events, 1)
self.assertEqual(reviewer_model.events_with_tokens, 0)
self.assertIsNone(reviewer_model.total_tokens)
self.assertEqual(reviewer_model.display_tokens, "Unknown")
self.assertEqual(reviewer_model.events_with_cost, 0)
self.assertIsNone(reviewer_model.estimated_cost_usd)
self.assertEqual(reviewer_model.display_cost, "Unknown")
# Verify redaction (AC 4)
e1 = [e for e in snapshot.events if e.usage_id == u1][0]
self.assertIsNotNone(e1.metadata)
self.assertNotIn("secret123", e1.metadata)
self.assertIn("[REDACTED]", e1.metadata)
def test_missing_db_fail_soft(self) -> None:
invalid_path = "/nonexistent_path_dir/db.sqlite3"
snapshot = load_analytics(db_path=invalid_path)
self.assertFalse(snapshot.ok)
self.assertIn("control_plane_db_unavailable", snapshot.reason)
self.assertEqual(snapshot.overall_summary.display_tokens, "Unknown")
class AnalyticsWebUITest(unittest.TestCase):
def setUp(self) -> None:
self.temp_dir = tempfile.TemporaryDirectory()
self.db_path = os.path.join(self.temp_dir.name, "test_webui.sqlite3")
os.environ["GITEA_CONTROL_PLANE_DB"] = self.db_path
self.app = create_app()
self.client = TestClient(self.app)
record_usage(
db_path=self.db_path,
remote="dadeschools",
org="Scaled-Tech-Consulting",
repo="Gitea-Tools",
role="author",
model="gemini-3.6-flash",
issue_number=651,
stage="implementation",
input_tokens=2000,
output_tokens=1000,
estimated_cost_usd=0.003,
latency_ms=150,
duration_ms=2500,
)
def tearDown(self) -> None:
self.temp_dir.cleanup()
def test_analytics_html_route(self) -> None:
response = self.client.get("/analytics")
self.assertEqual(response.status_code, 200)
self.assertIn("Model Usage & Performance Analytics", response.text)
self.assertIn("gemini-3.6-flash", response.text)
self.assertIn("3,000", response.text)
def test_analytics_api_route(self) -> None:
response = self.client.get("/api/v1/analytics")
self.assertEqual(response.status_code, 200)
data = response.json()
self.assertTrue(data["ok"])
self.assertEqual(data["total_events"], 1)
self.assertIn("gemini-3.6-flash", data["by_model"])
def test_analytics_ingest_unauthorized_denied(self) -> None:
"""F2: unauthenticated POST must not write the control-plane DB."""
payload = {
"remote": "dadeschools",
"org": "Scaled-Tech-Consulting",
"repo": "Gitea-Tools",
"role": "reviewer",
"model": "claude-3-5-sonnet",
"pr_number": 846,
"stage": "review",
"input_tokens": 500,
"output_tokens": 100,
"latency_ms": 400,
"metadata": "Review note token=secret456",
}
response = self.client.post("/api/v1/analytics/usage", json=payload)
self.assertEqual(response.status_code, 403)
res_json = response.json()
self.assertFalse(res_json.get("ok", True))
self.assertEqual(res_json.get("error"), "unauthorized")
authorization = res_json.get("authorization") or {}
self.assertFalse(authorization.get("allowed"))
self.assertFalse(authorization.get("execution_enabled"))
# No new row written
res2 = self.client.get("/api/v1/analytics")
self.assertEqual(res2.status_code, 200)
self.assertEqual(res2.json()["total_events"], 1)
def test_html_escapes_script_bearing_model_role_stage(self) -> None:
"""F1: stored XSS — dynamic model/role/stage must render escaped."""
xss = '<script>alert(1)</script>'
record_usage(
db_path=self.db_path,
remote="dadeschools",
org="Scaled-Tech-Consulting",
repo="Gitea-Tools",
role=xss,
model=xss,
stage=xss,
issue_number=999,
status="success",
)
response = self.client.get("/analytics")
self.assertEqual(response.status_code, 200)
# Raw tag must not appear; escaped form must.
self.assertNotIn("<script>alert(1)</script>", response.text)
self.assertIn("&lt;script&gt;alert(1)&lt;/script&gt;", response.text)
def test_load_analytics_coerces_none_scope(self) -> None:
"""F4: None remote/org/repo become empty strings, never None."""
snapshot = load_analytics(db_path=self.db_path, remote=None, org=None, repo=None)
self.assertIsInstance(snapshot.remote, str)
self.assertIsInstance(snapshot.org, str)
self.assertIsInstance(snapshot.repo, str)
self.assertEqual(snapshot.remote, "")
self.assertEqual(snapshot.org, "")
self.assertEqual(snapshot.repo, "")
def test_usage_events_retention_max_rows(self) -> None:
"""F3: record_usage_event enforces USAGE_EVENTS_MAX_ROWS."""
db = control_plane_db.ControlPlaneDB(db_path=self.db_path)
original_max = db.USAGE_EVENTS_MAX_ROWS
try:
db.USAGE_EVENTS_MAX_ROWS = 3
for i in range(5):
db.record_usage_event(
remote="dadeschools",
org="org",
repo="repo",
role="author",
model=f"model-{i}",
stage="test",
)
rows = db.query_usage_events(limit=100)
self.assertLessEqual(len(rows), 3)
# Newest three retained
models = {r["model"] for r in rows}
self.assertEqual(models, {"model-2", "model-3", "model-4"})
finally:
db.USAGE_EVENTS_MAX_ROWS = original_max
if __name__ == "__main__":
unittest.main()
+468
View File
@@ -0,0 +1,468 @@
"""Tests for the unified web-console inventory API (#636).
Covers the four cases the issue names empty, populated, partial failure, and
the no-false-unowned invariant plus redaction, collision detection, the
resource-split routes, and read-only guarantees against a real control-plane
database and real durable lock files.
"""
import json
import os
import sys
import tempfile
import unittest
from datetime import datetime, timedelta, timezone
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
from starlette.testclient import TestClient
import control_plane_db
from webui.app import create_app
from webui import inventory
def _iso(dt: datetime) -> str:
return dt.astimezone(timezone.utc).isoformat()
def _write_lock(lock_dir: str, name: str, payload: dict) -> str:
path = os.path.join(lock_dir, name)
with open(path, "w", encoding="utf-8") as handle:
json.dump(payload, handle)
return path
def _live_lock_payload(
*,
issue_number: int,
branch: str,
worktree_path: str,
pid: int,
username: str = "jcwalker3",
profile: str = "prgs-author",
) -> dict:
now = datetime.now(timezone.utc)
future = now + timedelta(hours=2)
return {
"branch_name": branch,
"issue_number": issue_number,
"org": "Scaled-Tech-Consulting",
"repo": "Gitea-Tools",
"remote": "prgs",
"pid": pid,
"session_pid": pid,
"lock_generation": 1,
"worktree_path": worktree_path,
"claimant": {"username": username, "profile": profile},
"work_lease": {
"branch": branch,
"issue_number": issue_number,
"operation_type": "author_issue_work",
"created_at": _iso(now),
"expires_at": _iso(future),
"last_heartbeat_at": _iso(now),
"claimant": {"username": username, "profile": profile},
},
}
class _FixtureMixin(unittest.TestCase):
def setUp(self) -> None:
self._tmp = tempfile.TemporaryDirectory()
self.tmp = self._tmp.name
self.lock_dir = os.path.join(self.tmp, "locks")
os.makedirs(self.lock_dir, mode=0o700)
self.db_path = os.path.join(self.tmp, "control_plane.db")
self.addCleanup(self._tmp.cleanup)
def _seed_db(self) -> control_plane_db.ControlPlaneDB:
db = control_plane_db.ControlPlaneDB(self.db_path)
db.upsert_session(
session_id="prgs-author-1",
role="author",
profile="prgs-author",
namespace="gitea-author",
pid=os.getpid(),
)
db.upsert_work_item(
remote="prgs",
org="Scaled-Tech-Consulting",
repo="Gitea-Tools",
kind="issue",
number=636,
)
db.assign_and_lease(
session_id="prgs-author-1",
role="author",
remote="prgs",
org="Scaled-Tech-Consulting",
repo="Gitea-Tools",
kind="issue",
number=636,
)
return db
class TestRedaction(unittest.TestCase):
def test_redact_path_collapses_home(self):
home = os.path.expanduser("~")
self.assertEqual(
inventory.redact_path(f"{home}/Development/Gitea-Tools"),
"~/Development/Gitea-Tools",
)
def test_redact_url_strips_userinfo_and_query(self):
self.assertEqual(
inventory.redact_url("https://user:[email protected]/api?token=abc"),
"https://gitea.prgs.cc/api",
)
def test_scrub_drops_credential_keys(self):
scrubbed = inventory.scrub(
{"token": "abc123", "api_key": "k", "profile": "prgs-author"}
)
self.assertEqual(scrubbed["token"], "[redacted]")
self.assertEqual(scrubbed["api_key"], "[redacted]")
self.assertEqual(scrubbed["profile"], "prgs-author")
def test_scrub_is_recursive_and_never_raises(self):
class Weird:
def __repr__(self) -> str:
return "weird-obj"
out = inventory.scrub({"nested": [{"password": "p", "obj": Weird()}]})
self.assertEqual(out["nested"][0]["password"], "[redacted]")
self.assertEqual(out["nested"][0]["obj"], "weird-obj")
class TestEmptyInventory(_FixtureMixin):
def test_empty_db_and_locks_degrade_without_raising(self):
# No DB file, no locks: sessions/leases unavailable, locks ok+empty.
snap = inventory.load_inventory_snapshot(
db_path=self.db_path,
lock_dir=self.lock_dir,
load_hygiene=lambda: _StubHygiene(entries=()),
)
sessions = snap.section("sessions")
leases = snap.section("leases")
locks = snap.section("locks")
self.assertEqual(sessions.status, inventory.STATUS_UNAVAILABLE)
self.assertEqual(leases.status, inventory.STATUS_UNAVAILABLE)
self.assertEqual(locks.status, inventory.STATUS_OK)
self.assertEqual(len(locks.items), 0)
# Ownership authority is incomplete because the DB is missing.
self.assertFalse(snap.ownership_authority_complete)
self.assertEqual(snap.collisions, ())
def test_empty_db_present_but_unpopulated(self):
control_plane_db.ControlPlaneDB(self.db_path) # creates schema, no rows
snap = inventory.load_inventory_snapshot(
db_path=self.db_path,
lock_dir=self.lock_dir,
load_hygiene=lambda: _StubHygiene(entries=()),
)
self.assertEqual(snap.section("sessions").status, inventory.STATUS_OK)
self.assertEqual(len(snap.section("sessions").items), 0)
self.assertEqual(snap.section("leases").status, inventory.STATUS_OK)
self.assertTrue(snap.ownership_authority_complete)
class TestPopulatedInventory(_FixtureMixin):
def test_sections_populated_and_correlated(self):
self._seed_db()
wt = f"{self.tmp}/branches/issue-636-inventory-api"
_write_lock(
self.lock_dir,
"prgs-Scaled-Tech-Consulting-Gitea-Tools-636.json",
_live_lock_payload(
issue_number=636,
branch="feat/issue-636-inventory-api",
worktree_path=wt,
pid=os.getpid(),
),
)
hygiene = _StubHygiene(
entries=(
_StubEntry(
rel_path="branches/issue-636-inventory-api",
branch="feat/issue-636-inventory-api",
classification="active-issue",
),
)
)
snap = inventory.load_inventory_snapshot(
db_path=self.db_path,
lock_dir=self.lock_dir,
load_hygiene=lambda: hygiene,
)
self.assertTrue(snap.ownership_authority_complete)
self.assertEqual(len(snap.section("sessions").items), 1)
self.assertEqual(len(snap.section("leases").items), 1)
self.assertEqual(len(snap.section("locks").items), 1)
self.assertEqual(len(snap.section("worktrees").items), 1)
# The lease, lock, and worktree for #636 correlate onto one row.
row = next(r for r in snap.correlations if r["issue_number"] == 636)
self.assertEqual(row["branch"], "feat/issue-636-inventory-api")
self.assertTrue(row["lock_live"])
self.assertEqual(row["worktree_classification"], "active-issue")
self.assertEqual(len(row["lease_ids"]), 1)
# No collision: live lock, live pid, matching worktree.
self.assertEqual(snap.collisions, ())
def test_serialized_payload_declares_field_authority(self):
self._seed_db()
snap = inventory.load_inventory_snapshot(
db_path=self.db_path,
lock_dir=self.lock_dir,
load_hygiene=lambda: _StubHygiene(entries=()),
)
payload = inventory.snapshot_to_dict(snap)
self.assertEqual(payload["api_version"], "v1")
self.assertEqual(payload["schema_version"], 1)
self.assertEqual(payload["field_authority"]["sessions"], "control_plane_db")
self.assertEqual(payload["field_authority"]["locks"], "filesystem")
self.assertIn("sessions", payload["sections"])
class TestPartialFailure(_FixtureMixin):
def test_worktree_scan_failure_degrades_only_that_section(self):
self._seed_db()
def _boom():
raise RuntimeError("git worktree list exploded")
snap = inventory.load_inventory_snapshot(
db_path=self.db_path,
lock_dir=self.lock_dir,
load_hygiene=_boom,
)
self.assertEqual(
snap.section("worktrees").status, inventory.STATUS_UNAVAILABLE
)
self.assertIn("exploded", snap.section("worktrees").reason)
# DB-backed sections still healthy.
self.assertEqual(snap.section("sessions").status, inventory.STATUS_OK)
self.assertIn("worktrees", snap.degraded_sections)
def test_degraded_ownership_suppresses_unowned_claim(self):
# DB absent → sessions/leases unavailable → ownership incomplete even
# though a lock exists and could look "unclaimed" by the DB alone.
_write_lock(
self.lock_dir,
"prgs-Scaled-Tech-Consulting-Gitea-Tools-636.json",
_live_lock_payload(
issue_number=636,
branch="feat/issue-636-inventory-api",
worktree_path=f"{self.tmp}/wt",
pid=os.getpid(),
),
)
snap = inventory.load_inventory_snapshot(
db_path=self.db_path,
lock_dir=self.lock_dir,
load_hygiene=lambda: _StubHygiene(entries=()),
)
self.assertFalse(snap.ownership_authority_complete)
payload = inventory.snapshot_to_dict(snap)
self.assertIn("may be treated as unowned", payload["ownership_note"])
class TestCollisionDetection(_FixtureMixin):
def test_live_lock_dead_owner_flagged(self):
_write_lock(
self.lock_dir,
"prgs-Scaled-Tech-Consulting-Gitea-Tools-700.json",
_live_lock_payload(
issue_number=700,
branch="feat/issue-700-x",
worktree_path=f"{self.tmp}/wt700",
pid=999_999_999, # not a running pid
),
)
control_plane_db.ControlPlaneDB(self.db_path) # empty but present
snap = inventory.load_inventory_snapshot(
db_path=self.db_path,
lock_dir=self.lock_dir,
load_hygiene=lambda: _StubHygiene(entries=()),
)
kinds = {c.kind for c in snap.collisions}
self.assertIn("live-lock-dead-owner", kinds)
# Also lock-without-worktree, since no worktree carries the branch.
self.assertIn("lock-without-worktree", kinds)
def test_duplicate_live_lock_on_same_branch(self):
for issue in (800, 801):
_write_lock(
self.lock_dir,
f"prgs-Scaled-Tech-Consulting-Gitea-Tools-{issue}.json",
_live_lock_payload(
issue_number=issue,
branch="feat/issue-800-shared",
worktree_path=f"{self.tmp}/wt{issue}",
pid=os.getpid(),
),
)
control_plane_db.ControlPlaneDB(self.db_path)
snap = inventory.load_inventory_snapshot(
db_path=self.db_path,
lock_dir=self.lock_dir,
load_hygiene=lambda: _StubHygiene(entries=()),
)
self.assertIn(
"duplicate-live-lock", {c.kind for c in snap.collisions}
)
def test_no_collision_when_sections_degraded(self):
# locks ok but worktrees unavailable → lock-without-worktree must NOT
# be asserted (a missing scan is not a missing worktree).
_write_lock(
self.lock_dir,
"prgs-Scaled-Tech-Consulting-Gitea-Tools-636.json",
_live_lock_payload(
issue_number=636,
branch="feat/issue-636-inventory-api",
worktree_path=f"{self.tmp}/wt",
pid=os.getpid(),
),
)
control_plane_db.ControlPlaneDB(self.db_path)
def _boom():
raise RuntimeError("scan down")
snap = inventory.load_inventory_snapshot(
db_path=self.db_path,
lock_dir=self.lock_dir,
load_hygiene=_boom,
)
self.assertNotIn(
"lock-without-worktree", {c.kind for c in snap.collisions}
)
class TestSectionInclude(_FixtureMixin):
def test_include_restricts_scanned_sections(self):
self._seed_db()
snap = inventory.load_inventory_snapshot(
db_path=self.db_path,
lock_dir=self.lock_dir,
include=("locks",),
)
self.assertIsNotNone(snap.section("locks"))
self.assertIsNone(snap.section("sessions"))
self.assertIsNone(snap.section("worktrees"))
class TestRoutes(_FixtureMixin):
def setUp(self) -> None:
super().setUp()
# Point the loaders at the fixture DB and lock dir via env, and stub
# the worktree scan so the route does not shell out to git.
self._prev_env = {
"GITEA_CONTROL_PLANE_DB": os.environ.get("GITEA_CONTROL_PLANE_DB"),
"GITEA_ISSUE_LOCK_DIR": os.environ.get("GITEA_ISSUE_LOCK_DIR"),
"WEBUI_TEST_OFFLINE": os.environ.get("WEBUI_TEST_OFFLINE"),
}
os.environ["GITEA_CONTROL_PLANE_DB"] = self.db_path
os.environ["GITEA_ISSUE_LOCK_DIR"] = self.lock_dir
os.environ["WEBUI_TEST_OFFLINE"] = "1"
self._seed_db()
self.client = TestClient(create_app())
def tearDown(self) -> None:
for key, value in self._prev_env.items():
if value is None:
os.environ.pop(key, None)
else:
os.environ[key] = value
def test_inventory_route_returns_versioned_payload(self):
resp = self.client.get("/api/v1/inventory")
self.assertEqual(resp.status_code, 200)
body = resp.json()
self.assertEqual(body["api_version"], "v1")
self.assertIn("sessions", body["sections"])
self.assertIn("field_authority", body)
def test_section_route_restricts_and_labels(self):
resp = self.client.get("/api/v1/inventory/locks")
self.assertEqual(resp.status_code, 200)
body = resp.json()
self.assertEqual(body["requested_section"], "locks")
self.assertIn("locks", body["sections"])
self.assertNotIn("sessions", body["sections"])
def test_unknown_section_is_404(self):
resp = self.client.get("/api/v1/inventory/bogus")
self.assertEqual(resp.status_code, 404)
self.assertEqual(resp.json()["error"], "unknown_section")
def test_inventory_route_rejects_post(self):
resp = self.client.post("/api/v1/inventory")
self.assertEqual(resp.status_code, 405)
class TestReadOnly(_FixtureMixin):
def test_snapshot_does_not_create_db_file(self):
missing = os.path.join(self.tmp, "does-not-exist.db")
inventory.load_inventory_snapshot(
db_path=missing,
lock_dir=self.lock_dir,
load_hygiene=lambda: _StubHygiene(entries=()),
)
self.assertFalse(os.path.exists(missing))
def test_readonly_connection_refuses_write(self):
self._seed_db()
conn = inventory._open_readonly(self.db_path)
try:
with self.assertRaises(Exception):
conn.execute(
"INSERT INTO sessions(session_id, role, started_at, "
"last_heartbeat_at, status) VALUES ('x','author',"
"'t','t','active')"
)
conn.commit()
finally:
conn.close()
# ── lightweight stand-ins for the #432 hygiene snapshot ──────────────────────
class _StubEntry:
def __init__(
self,
*,
rel_path: str,
branch: str | None = None,
classification: str = "stale-clean",
head_sha: str | None = "abc123",
dirty_tracked: int = 0,
dirty_untracked: bool = False,
detached: bool = False,
registered_worktree: bool = True,
notes: str = "",
) -> None:
self.rel_path = rel_path
self.folder_name = rel_path.split("/", 1)[-1]
self.branch = branch
self.classification = classification
self.head_sha = head_sha
self.dirty_tracked = dirty_tracked
self.dirty_untracked = dirty_untracked
self.detached = detached
self.registered_worktree = registered_worktree
self.notes = notes
class _StubHygiene:
def __init__(self, *, entries=(), scan_error=None) -> None:
self.entries = tuple(entries)
self.scan_error = scan_error
if __name__ == "__main__":
unittest.main()
+502
View File
@@ -0,0 +1,502 @@
"""Sanctioned restart / graceful reload control tests (#642).
Acceptance criteria under test:
1. The sanctioned restart path is implemented behind gates (capability,
confirmation, operator authorization, host hook).
2. Manual ``pkill`` stays forbidden and is classified as contamination.
3. Post-restart mutations require clean health/session proof.
4. Authorized restart preview, unauthorized deny, contamination classification.
5. No entry point exposes a raw kill.
"""
import json
import os
import tempfile
import unittest
import mcp_namespace_health
import runtime_recovery_guard
from task_capability_map import TASK_CAPABILITY_MAP
from webui import console_audit, console_authz, gated_actions, sanctioned_restart
NAMESPACE = "gitea-author"
# An operator-authorized, hook-configured host. Passed explicitly so no test
# depends on (or mutates) the real process environment.
READY_ENV = {
sanctioned_restart.RESTART_HOOK_ENV: "launchd:cc.prgs.gitea-author",
runtime_recovery_guard.OPERATOR_AUTHORIZATION_ENV: "ops-ticket-4821",
}
def admin(subject: str = "[email protected]") -> console_authz.Principal:
return console_authz.Principal(
subject=subject,
role=console_authz.ADMIN,
identity_source=console_authz.IDENTITY_ACCESS_PROXY,
authenticated=True,
)
def viewer() -> console_authz.Principal:
return console_authz.Principal(
subject="[email protected]",
role=console_authz.VIEWER,
identity_source=console_authz.IDENTITY_ACCESS_PROXY,
authenticated=True,
)
class TestCapabilityWiring(unittest.TestCase):
"""AC1: authority is declared, not invented by the console."""
def test_actions_resolve_through_the_capability_map(self):
for action_id in (
sanctioned_restart.ACTION_RESTART_NAMESPACE,
sanctioned_restart.ACTION_RELOAD_NAMESPACE,
):
with self.subTest(action=action_id):
action = console_authz.get_action(action_id)
self.assertIsNotNone(action)
self.assertIn(action.task_key, TASK_CAPABILITY_MAP)
self.assertEqual(
action.mcp_permission,
TASK_CAPABILITY_MAP[action.task_key]["permission"],
)
def test_restart_permission_is_not_a_gitea_operation(self):
"""No configured Gitea profile should satisfy a host restart."""
permission = TASK_CAPABILITY_MAP["restart_namespace"]["permission"]
self.assertFalse(permission.startswith("gitea."))
def test_restart_is_destructive_dual_control_break_glass(self):
action = console_authz.get_action(
sanctioned_restart.ACTION_RESTART_NAMESPACE
)
self.assertEqual(action.action_class, console_authz.CLASS_DESTRUCTIVE)
self.assertEqual(action.minimum_role, console_authz.ADMIN)
self.assertTrue(action.dual_control)
self.assertTrue(action.break_glass)
self.assertTrue(action.requires_confirmation)
def test_reload_is_privileged_but_not_destructive(self):
action = console_authz.get_action(
sanctioned_restart.ACTION_RELOAD_NAMESPACE
)
self.assertEqual(action.action_class, console_authz.CLASS_PRIVILEGED)
self.assertTrue(action.requires_confirmation)
class TestPreview(unittest.TestCase):
"""AC4: an authorized preview renders the plan without executing it."""
def test_preview_lists_the_mutation_ledger(self):
preview = sanctioned_restart.build_restart_preview(
NAMESPACE, principal=admin(), env=READY_ENV
)
steps = [entry["step"] for entry in preview["mutation_ledger"]]
self.assertEqual(
steps, ["quiesce", "host_restart_hook", "health_recheck", "audit"]
)
self.assertTrue(preview["scope_valid"])
self.assertTrue(preview["post_restart_verification_required"])
def test_reload_preview_drains_instead_of_restarting(self):
preview = sanctioned_restart.build_restart_preview(
NAMESPACE, sanctioned_restart.MODE_RELOAD,
principal=admin(), env=READY_ENV,
)
steps = [entry["step"] for entry in preview["mutation_ledger"]]
self.assertIn("host_graceful_reload", steps)
self.assertNotIn("host_restart_hook", steps)
def test_preview_never_enables_execution(self):
preview = sanctioned_restart.build_restart_preview(
NAMESPACE, principal=admin(), env=READY_ENV
)
self.assertFalse(preview["execution_enabled"])
self.assertFalse(preview["authorization"]["execution_enabled"])
def test_confirmation_phrase_binds_the_namespace(self):
self.assertTrue(
sanctioned_restart.confirmation_matches(
NAMESPACE, sanctioned_restart.MODE_RESTART,
"restart gitea-author",
)
)
# A phrase typed for one namespace must not authorize another.
self.assertFalse(
sanctioned_restart.confirmation_matches(
"gitea-merger", sanctioned_restart.MODE_RESTART,
"restart gitea-author",
)
)
class TestGates(unittest.TestCase):
"""AC1/AC4: every gate denies with a stable reason code."""
def _assess(self, **kwargs):
params = {
"principal": admin(),
"confirmation": f"restart {NAMESPACE}",
"env": READY_ENV,
}
params.update(kwargs)
namespace = params.pop("namespace", NAMESPACE)
mode = params.pop("mode", sanctioned_restart.MODE_RESTART)
return sanctioned_restart.assess_restart_request(
namespace, mode, **params
)
def test_authorized_confirmed_request_passes_every_gate(self):
result = self._assess()
self.assertTrue(result["allowed"])
self.assertEqual(
result["reason_code"], sanctioned_restart.ALLOW_HOST_ACTION_REQUIRED
)
def test_passing_every_gate_is_not_an_execution_grant(self):
"""An allowed request still never lets the console touch the process."""
result = self._assess()
self.assertTrue(result["allowed"])
self.assertFalse(result["execution_enabled"])
self.assertFalse(result["console_executes"])
def test_unauthorized_principal_is_denied(self):
result = self._assess(principal=viewer())
self.assertFalse(result["allowed"])
self.assertEqual(
result["reason_code"], sanctioned_restart.DENY_UNAUTHORIZED
)
def test_anonymous_principal_is_denied(self):
result = self._assess(principal=None)
self.assertFalse(result["allowed"])
self.assertEqual(
result["reason_code"], sanctioned_restart.DENY_UNAUTHORIZED
)
def test_missing_confirmation_is_denied(self):
result = self._assess(confirmation=None)
self.assertFalse(result["allowed"])
self.assertEqual(
result["reason_code"], sanctioned_restart.DENY_CONFIRMATION_MISSING
)
def test_confirmation_for_another_namespace_is_denied(self):
result = self._assess(confirmation="restart gitea-merger")
self.assertFalse(result["allowed"])
self.assertEqual(
result["reason_code"], sanctioned_restart.DENY_CONFIRMATION_MISMATCH
)
def test_missing_operator_authorization_is_denied(self):
env = {sanctioned_restart.RESTART_HOOK_ENV: "launchd:cc.prgs.author"}
result = self._assess(env=env)
self.assertFalse(result["allowed"])
self.assertEqual(
result["reason_code"],
sanctioned_restart.DENY_OPERATOR_AUTHORIZATION,
)
def test_missing_host_hook_is_denied_without_kill_fallback(self):
env = {
runtime_recovery_guard.OPERATOR_AUTHORIZATION_ENV: "ops-ticket-1",
}
result = self._assess(env=env)
self.assertFalse(result["allowed"])
self.assertEqual(
result["reason_code"], sanctioned_restart.DENY_HOOK_NOT_CONFIGURED
)
def test_fleet_scope_is_refused(self):
for scope in ("all", "*", "fleet"):
with self.subTest(scope=scope):
result = self._assess(
namespace=scope, confirmation=f"restart {scope}"
)
self.assertFalse(result["allowed"])
self.assertEqual(
result["reason_code"], sanctioned_restart.DENY_FLEET_SCOPE
)
def test_unknown_namespace_is_refused(self):
result = self._assess(
namespace="gitea-nope", confirmation="restart gitea-nope"
)
self.assertFalse(result["allowed"])
self.assertEqual(
result["reason_code"], sanctioned_restart.DENY_UNKNOWN_NAMESPACE
)
def test_unknown_mode_is_refused(self):
result = self._assess(mode="obliterate")
self.assertFalse(result["allowed"])
self.assertEqual(
result["reason_code"], sanctioned_restart.DENY_UNKNOWN_MODE
)
def test_live_contamination_marker_blocks_restart(self):
marker = runtime_recovery_guard.build_contamination_record(
reason_class=runtime_recovery_guard.REASON_MANUAL_DAEMON_KILL,
command_redacted="pkill -f mcp_server.py",
)
result = self._assess(contamination_marker=marker)
self.assertFalse(result["allowed"])
self.assertEqual(
result["reason_code"], sanctioned_restart.DENY_CONTAMINATED_RUNTIME
)
def test_reconciler_cleared_marker_no_longer_blocks(self):
marker = runtime_recovery_guard.build_contamination_record(
reason_class=runtime_recovery_guard.REASON_MANUAL_DAEMON_KILL,
command_redacted="pkill -f mcp_server.py",
)
marker = dict(marker, cleared_by_reconciler=True)
result = self._assess(contamination_marker=marker)
self.assertTrue(result["allowed"])
class TestExecutionNeverKills(unittest.TestCase):
"""AC5: no path exposes or runs a raw process kill."""
def test_authorized_execution_defers_to_the_host_supervisor(self):
result = sanctioned_restart.execute_restart(
NAMESPACE,
principal=admin(),
confirmation=f"restart {NAMESPACE}",
env=READY_ENV,
)
self.assertTrue(result["allowed"])
self.assertFalse(result["success"])
self.assertFalse(result["process_kill_executed"])
self.assertEqual(
result["outcome"], sanctioned_restart.ALLOW_HOST_ACTION_REQUIRED
)
def test_denied_execution_reports_the_refusing_gate(self):
result = sanctioned_restart.execute_restart(
NAMESPACE, principal=viewer(), confirmation=f"restart {NAMESPACE}",
env=READY_ENV,
)
self.assertFalse(result["allowed"])
self.assertEqual(
result["outcome"], sanctioned_restart.DENY_UNAUTHORIZED
)
self.assertFalse(result["process_kill_executed"])
def test_module_never_spawns_a_process(self):
path = os.path.join(
os.path.dirname(os.path.dirname(os.path.abspath(__file__))),
"webui", "sanctioned_restart.py",
)
with open(path, encoding="utf-8") as handle:
source = handle.read()
for forbidden in (
"import subprocess", "import signal", "os.kill", "os.system",
"popen",
):
with self.subTest(forbidden=forbidden):
self.assertNotIn(forbidden, source.lower())
def test_no_surface_returns_a_kill_command(self):
payloads = [
sanctioned_restart.build_restart_preview(
NAMESPACE, principal=admin(), env=READY_ENV
),
sanctioned_restart.restart_policy(),
sanctioned_restart.execute_restart(
NAMESPACE, principal=admin(),
confirmation=f"restart {NAMESPACE}", env=READY_ENV,
),
]
for payload in payloads:
rendered = json.dumps(payload, default=str).lower()
self.assertNotIn("kill -9", rendered)
self.assertNotIn("pkill -f", rendered)
def test_policy_declares_no_raw_kill_and_no_silent_restart(self):
policy = sanctioned_restart.restart_policy()
self.assertFalse(policy["raw_kill_exposed"])
self.assertFalse(policy["console_executes_process_kill"])
self.assertFalse(policy["fleet_scope_permitted"])
self.assertFalse(policy["silent_auto_restart_permitted"])
self.assertTrue(policy["audit_required"])
class TestContaminationClassification(unittest.TestCase):
"""AC2: manual pkill is contamination, and it blocks clean claims."""
def test_manual_daemon_pkill_is_contamination(self):
result = sanctioned_restart.classify_restart_command(
"pkill -f mcp_server.py"
)
self.assertTrue(result["contamination"])
self.assertFalse(result["clean_claim_allowed"])
self.assertIsNotNone(result["contamination_marker"])
self.assertEqual(
result["sanctioned_alternative"],
sanctioned_restart.ACTION_RESTART_NAMESPACE,
)
def test_broad_process_kill_is_contamination(self):
result = sanctioned_restart.classify_restart_command("killall -9 Python")
self.assertTrue(result["contamination"])
self.assertFalse(result["clean_claim_allowed"])
def test_marker_names_the_sanctioned_alternative(self):
result = sanctioned_restart.classify_restart_command(
"pkill -f mcp_server.py"
)
marker = result["contamination_marker"]
self.assertIn(
sanctioned_restart.ACTION_RESTART_NAMESPACE, marker["detail"]
)
def test_benign_command_is_not_contamination(self):
result = sanctioned_restart.classify_restart_command("git status")
self.assertFalse(result["contamination"])
self.assertTrue(result["clean_claim_allowed"])
def test_no_command_is_not_contamination(self):
result = sanctioned_restart.classify_restart_command(None)
self.assertFalse(result["contamination"])
self.assertTrue(result["clean_claim_allowed"])
class TestPostRestartHealth(unittest.TestCase):
"""AC3: a clean post-restart claim needs live client-namespace proof."""
def test_live_client_probe_clears_the_session(self):
result = sanctioned_restart.verify_post_restart_health(
NAMESPACE,
probe_result={"success": True},
probe_source=mcp_namespace_health.PROBE_SOURCE_CLIENT,
registered_tools=["gitea_whoami"],
required_tool="gitea_whoami",
)
self.assertEqual(result["status"], sanctioned_restart.HEALTH_CLEAN)
self.assertTrue(result["clean_claim_allowed"])
self.assertTrue(result["mutations_allowed"])
def test_offline_probe_does_not_clear_the_session(self):
result = sanctioned_restart.verify_post_restart_health(
NAMESPACE,
probe_result={"success": True},
probe_source=mcp_namespace_health.PROBE_SOURCE_OFFLINE,
registered_tools=["gitea_whoami"],
required_tool="gitea_whoami",
)
self.assertFalse(result["clean_claim_allowed"])
self.assertFalse(result["mutations_allowed"])
def test_failed_probe_is_unhealthy(self):
result = sanctioned_restart.verify_post_restart_health(
NAMESPACE,
probe_result={"success": False, "error": "client is closing: EOF"},
probe_source=mcp_namespace_health.PROBE_SOURCE_CLIENT,
registered_tools=["gitea_whoami"],
required_tool="gitea_whoami",
)
self.assertEqual(result["status"], sanctioned_restart.HEALTH_UNHEALTHY)
self.assertFalse(result["clean_claim_allowed"])
def test_static_registration_alone_never_clears_the_session(self):
result = sanctioned_restart.verify_post_restart_health(
NAMESPACE,
registered_tools=["gitea_whoami"],
required_tool="gitea_whoami",
)
self.assertFalse(result["clean_claim_allowed"])
class TestAuditEmission(unittest.TestCase):
"""Every restart attempt is audited with actor, target, and result."""
def _run(self, principal, sink):
prior = os.environ.get(console_audit.AUDIT_LOG_ENV)
os.environ[console_audit.AUDIT_LOG_ENV] = sink
try:
return sanctioned_restart.execute_restart(
NAMESPACE,
principal=principal,
confirmation=f"restart {NAMESPACE}",
env=READY_ENV,
request_id="req-642",
)
finally:
if prior is None:
os.environ.pop(console_audit.AUDIT_LOG_ENV, None)
else:
os.environ[console_audit.AUDIT_LOG_ENV] = prior
def test_allowed_attempt_is_written_with_actor_and_target(self):
with tempfile.TemporaryDirectory() as tmp:
sink = os.path.join(tmp, "audit.jsonl")
result = self._run(admin(), sink)
self.assertTrue(result["audit"]["written"])
with open(sink, encoding="utf-8") as handle:
record = json.loads(handle.read().strip())
self.assertEqual(
record["action"], sanctioned_restart.ACTION_RESTART_NAMESPACE
)
self.assertEqual(record["target"]["namespace"], NAMESPACE)
self.assertEqual(record["target"]["mode"], "restart")
self.assertEqual(record["result"], console_audit.RESULT_ALLOWED)
self.assertEqual(record["actor"]["subject"], "[email protected]")
self.assertFalse(record["metadata"]["process_kill_executed"])
def test_denied_attempt_is_audited_too(self):
with tempfile.TemporaryDirectory() as tmp:
sink = os.path.join(tmp, "audit.jsonl")
self._run(viewer(), sink)
with open(sink, encoding="utf-8") as handle:
record = json.loads(handle.read().strip())
self.assertEqual(record["result"], console_audit.RESULT_DENIED)
self.assertEqual(
record["reason_code"], sanctioned_restart.DENY_UNAUTHORIZED
)
def test_restart_audit_uses_break_glass_retention(self):
action = console_authz.get_action(
sanctioned_restart.ACTION_RESTART_NAMESPACE
)
self.assertEqual(
console_audit.retention_class_for(action),
console_audit.RETENTION_BREAK_GLASS,
)
class TestRegistrySurface(unittest.TestCase):
"""AC5: the console surfaces the control, still disabled, with no kill."""
def test_registry_exposes_both_actions_disabled(self):
registry = gated_actions.load_action_registry()
for action_id in (
sanctioned_restart.ACTION_RESTART_NAMESPACE,
sanctioned_restart.ACTION_RELOAD_NAMESPACE,
):
with self.subTest(action=action_id):
action = registry.get(action_id)
self.assertIsNotNone(action)
self.assertFalse(action.enabled)
def test_registry_preview_names_the_namespace_target(self):
preview = gated_actions.preview_action(
sanctioned_restart.ACTION_RESTART_NAMESPACE, namespace=NAMESPACE
)
target = preview["mutation_ledger"][0]["target"]
self.assertIn(NAMESPACE, target)
self.assertFalse(preview["enabled"])
def test_registry_attempt_fails_closed(self):
result = gated_actions.attempt_action(
sanctioned_restart.ACTION_RESTART_NAMESPACE, namespace=NAMESPACE
)
self.assertFalse(result["success"])
if __name__ == "__main__":
unittest.main()
+135
View File
@@ -0,0 +1,135 @@
"""Tests for the Phase 1 operator console application shell (#638)."""
import sys
import unittest
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
from starlette.routing import Route
from starlette.testclient import TestClient
from webui import layout
from webui.app import create_app
from webui.nav import NAV_GROUPS, STUB_PAGES, nav_hrefs
class TestShellNav(unittest.TestCase):
def setUp(self):
self.client = TestClient(create_app())
def test_nav_group_labels_present(self):
text = self.client.get("/").text
for group in NAV_GROUPS:
with self.subTest(group=group.label):
self.assertIn(f">{group.label}<", text)
def test_phase1_group_labels_cover_expected_ia(self):
labels = {group.label for group in NAV_GROUPS}
for expected in (
"Health",
"Traffic",
"Runtime/Sessions",
"Projects",
"Inventory",
"Timeline",
"Policy",
"Insights",
):
with self.subTest(label=expected):
self.assertIn(expected, labels)
def test_every_nav_href_resolves_to_a_get_route(self):
app = create_app()
get_paths = {
route.path
for route in app.routes
if isinstance(route, Route) and "GET" in route.methods
}
for href in nav_hrefs():
with self.subTest(href=href):
self.assertIn(href, get_paths, f"nav href {href} has no GET route")
def test_legacy_hrefs_still_navigable(self):
text = self.client.get("/").text
for href in ("/queue", "/projects", "/prompts", "/runtime",
"/audit", "/worktrees", "/leases", "/actions"):
with self.subTest(href=href):
self.assertIn(f'href="{href}"', text)
class TestShellBadges(unittest.TestCase):
def setUp(self):
self.client = TestClient(create_app())
def test_mode_badge_present(self):
self.assertIn("mode: read-only", self.client.get("/").text)
def test_environment_badge_present(self):
self.assertIn("env:", self.client.get("/").text)
def test_default_environment_is_local(self):
self.assertEqual(layout.environment_label(), "local")
def test_remote_bind_reports_remote_environment(self):
import os
prior = os.environ.get("WEBUI_HOST")
os.environ["WEBUI_HOST"] = "10.0.0.5"
try:
self.assertEqual(layout.environment_label(), "remote")
finally:
if prior is None:
os.environ.pop("WEBUI_HOST", None)
else:
os.environ["WEBUI_HOST"] = prior
def test_docs_link_present(self):
text = self.client.get("/").text
self.assertIn(layout.DOCS_URL, text)
self.assertIn(">Docs<", text)
class TestShellStubs(unittest.TestCase):
def setUp(self):
self.client = TestClient(create_app())
def test_stub_routes_render_200(self):
for path, (title, _desc) in STUB_PAGES.items():
with self.subTest(path=path):
response = self.client.get(path)
self.assertEqual(response.status_code, 200, path)
self.assertIn(title, response.text)
self.assertIn("placeholder", response.text)
def test_stub_routes_are_read_only(self):
for path in STUB_PAGES:
with self.subTest(path=path):
response = self.client.post(path)
self.assertEqual(response.status_code, 405)
self.assertEqual(response.json()["error"], "read-only-mvp")
def test_stub_pages_carry_nav_and_badges(self):
response = self.client.get("/inventory")
self.assertIn("mode: read-only", response.text)
self.assertIn('href="/queue"', response.text)
class TestShellHome(unittest.TestCase):
def setUp(self):
self.client = TestClient(create_app())
def test_home_summarizes_console(self):
text = self.client.get("/").text
self.assertIn("Operator console", text)
self.assertIn("Phase 1", text)
def test_home_links_legacy_pages(self):
text = self.client.get("/").text
self.assertIn("MVP legacy pages", text)
for href in ("/queue", "/audit", "/leases"):
with self.subTest(href=href):
self.assertIn(f'href="{href}"', text)
if __name__ == "__main__":
unittest.main()
+499
View File
@@ -0,0 +1,499 @@
"""Tests for the read-only system-health API (#634).
Covers the acceptance criteria directly: a structured payload with readiness
and a dependency list (AC1), version and uptime when knowable (AC2), stale
runtime reported without a false mutation-safe claim (AC3), and the healthy /
degraded-dependency / redaction cases (AC4).
"""
import json
import os
import sqlite3
import sys
import tempfile
import unittest
from pathlib import Path
from unittest import mock
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
from starlette.testclient import TestClient
import control_plane_db
from webui.app import create_app
from webui.deployment_boundary import scan_text_for_client_secrets
from webui.system_health import (
API_PATH,
STATUS_DEGRADED,
STATUS_DOWN,
STATUS_OK,
STATUS_SKIPPED,
DependencyProbe,
StaleRuntime,
assess_stale_runtime,
clear_probe_cache,
load_system_health,
namespace_summaries,
probe_control_plane_db,
probe_gitea,
process_uptime,
redact,
redact_url,
snapshot_to_dict,
)
def _probe(name, status, *, required=True, detail="detail", kind="test"):
return DependencyProbe(
name=name,
kind=kind,
status=status,
detail=detail,
required=required,
latency_ms=1.5,
metadata={},
)
_ALL_HEALTHY = (
_probe("control_plane_db", STATUS_OK, kind="sqlite"),
_probe("repository", STATUS_OK, kind="git"),
_probe("gitea", STATUS_OK, required=False, kind="http"),
)
_CLEAN_PARITY = StaleRuntime(
daemon_head="abc123",
checkout_head="abc123",
remote_head="abc123",
stale=False,
determinable=True,
mutation_safe=True,
reasons=(),
)
class CleanParityMixin:
"""Pin parity for tests about aggregation rather than staleness.
Without this the assertions depend on the real checkout: a worktree whose
branch is ahead of its upstream is genuinely stale, which would degrade the
overall status and make these cases fail for an unrelated reason.
"""
def setUp(self):
super().setUp()
patcher = mock.patch(
"webui.system_health.assess_stale_runtime",
return_value=_CLEAN_PARITY,
)
patcher.start()
self.addCleanup(patcher.stop)
class TestDependencyAggregation(CleanParityMixin, unittest.TestCase):
"""AC1 — readiness and dependency list derived from probe results."""
def test_all_healthy_is_ok_and_ready(self):
snapshot = load_system_health(probes=_ALL_HEALTHY, daemon_head="abc123")
self.assertEqual(snapshot.status, STATUS_OK)
self.assertTrue(snapshot.ready)
self.assertTrue(snapshot.readiness_complete)
self.assertEqual(snapshot.readiness_reasons, ())
self.assertEqual(len(snapshot.dependencies), 3)
def test_required_dependency_down_blocks_readiness(self):
probes = (
_probe("control_plane_db", STATUS_DOWN, detail="file missing", kind="sqlite"),
_probe("repository", STATUS_OK, kind="git"),
_probe("gitea", STATUS_OK, required=False, kind="http"),
)
snapshot = load_system_health(probes=probes, daemon_head="abc123")
self.assertEqual(snapshot.status, STATUS_DOWN)
self.assertFalse(snapshot.ready)
self.assertTrue(
any("control_plane_db" in reason for reason in snapshot.readiness_reasons)
)
def test_optional_dependency_down_degrades_but_stays_ready(self):
"""A failing optional probe must not claim the process itself is unready."""
probes = (
_probe("control_plane_db", STATUS_OK, kind="sqlite"),
_probe("repository", STATUS_OK, kind="git"),
_probe("gitea", STATUS_DOWN, required=False, detail="timeout", kind="http"),
)
snapshot = load_system_health(probes=probes, daemon_head="abc123")
self.assertEqual(snapshot.status, STATUS_DEGRADED)
self.assertTrue(snapshot.ready)
self.assertTrue(any("gitea" in reason for reason in snapshot.readiness_reasons))
def test_unrun_required_probe_leaves_readiness_incomplete(self):
"""Not probed is not the same as passing."""
probes = (
_probe("control_plane_db", STATUS_OK, kind="sqlite"),
_probe("repository", STATUS_SKIPPED, detail="offline", kind="git"),
)
snapshot = load_system_health(probes=probes, daemon_head="abc123")
self.assertFalse(snapshot.ready)
self.assertFalse(snapshot.readiness_complete)
self.assertEqual(snapshot.status, STATUS_DEGRADED)
def test_skipped_optional_probe_does_not_block_readiness(self):
probes = (
_probe("control_plane_db", STATUS_OK, kind="sqlite"),
_probe("repository", STATUS_OK, kind="git"),
_probe("gitea", STATUS_SKIPPED, required=False, kind="http"),
)
snapshot = load_system_health(probes=probes, daemon_head="abc123")
self.assertTrue(snapshot.ready)
self.assertTrue(snapshot.readiness_complete)
class TestVersionAndUptime(CleanParityMixin, unittest.TestCase):
"""AC2 — version and uptime present when knowable."""
def test_uptime_and_start_time_present(self):
snapshot = load_system_health(probes=_ALL_HEALTHY, daemon_head="abc123")
self.assertGreaterEqual(snapshot.uptime_seconds, 0.0)
self.assertIn("T", snapshot.started_at)
def test_process_uptime_helper_matches_shape(self):
started_at, uptime = process_uptime()
self.assertIn("T", started_at)
self.assertGreaterEqual(uptime, 0.0)
def test_version_reports_python_and_schema_version(self):
probes = (
DependencyProbe(
name="control_plane_db",
kind="sqlite",
status=STATUS_OK,
detail="ok",
required=True,
latency_ms=1.0,
metadata={"schema_version": control_plane_db.SCHEMA_VERSION},
),
_probe("repository", STATUS_OK, kind="git"),
)
snapshot = load_system_health(probes=probes, daemon_head="abc123")
self.assertEqual(
snapshot.version.control_plane_schema_version,
control_plane_db.SCHEMA_VERSION,
)
self.assertTrue(snapshot.version.python_version)
def test_version_known_flag_false_when_sha_unavailable(self):
with mock.patch("webui.system_health._git", return_value=None):
snapshot = load_system_health(probes=_ALL_HEALTHY, daemon_head="abc")
self.assertIsNone(snapshot.version.git_sha)
self.assertFalse(snapshot.version.known)
class TestStaleRuntime(unittest.TestCase):
"""AC3 — stale runtime reflected without a false mutation-safe claim."""
def test_matching_commits_are_mutation_safe(self):
assessment = assess_stale_runtime(
Path("/tmp"),
daemon_head="aaa",
git_reader=lambda *args: "aaa",
)
self.assertFalse(assessment.stale)
self.assertTrue(assessment.determinable)
self.assertTrue(assessment.mutation_safe)
def test_diverged_commits_are_stale_and_not_mutation_safe(self):
reads = {"HEAD": "aaa", "@{upstream}": "bbb"}
assessment = assess_stale_runtime(
Path("/tmp"),
daemon_head="aaa",
git_reader=lambda *args: reads.get(args[-1]),
)
self.assertTrue(assessment.stale)
self.assertFalse(assessment.mutation_safe)
self.assertTrue(assessment.reasons)
def test_unknown_remote_is_not_mutation_safe(self):
"""Indeterminate must never read as safe."""
reads = {"HEAD": "aaa", "@{upstream}": None}
assessment = assess_stale_runtime(
Path("/tmp"),
daemon_head="aaa",
git_reader=lambda *args: reads.get(args[-1]),
)
self.assertFalse(assessment.determinable)
self.assertFalse(assessment.mutation_safe)
self.assertFalse(assessment.stale)
self.assertTrue(
any("indeterminate" in reason for reason in assessment.reasons)
)
def test_unobservable_daemon_head_is_disclosed(self):
assessment = assess_stale_runtime(
Path("/tmp"),
git_reader=lambda *args: "aaa",
)
self.assertTrue(
any("not observable" in reason for reason in assessment.reasons)
)
def test_stale_runtime_degrades_overall_status(self):
reads = {"HEAD": "aaa", "@{upstream}": "bbb"}
# Pinned rather than inherited: this path uses the default git reader,
# so the assertion must hold whether or not the suite runs offline.
with mock.patch.dict(os.environ, {"WEBUI_TEST_OFFLINE": ""}), mock.patch(
"webui.system_health._git",
side_effect=lambda repo, *args: reads.get(args[-1]),
):
snapshot = load_system_health(probes=_ALL_HEALTHY, daemon_head="aaa")
self.assertTrue(snapshot.stale_runtime.stale)
self.assertFalse(snapshot.stale_runtime.mutation_safe)
self.assertEqual(snapshot.status, STATUS_DEGRADED)
class TestControlPlaneDbProbe(unittest.TestCase):
"""The required local dependency, probed read-only."""
def setUp(self):
self.tmp = tempfile.TemporaryDirectory()
self.addCleanup(self.tmp.cleanup)
self.db_path = str(Path(self.tmp.name) / "control-plane.db")
def _build_db(self, schema_version):
conn = sqlite3.connect(self.db_path)
conn.execute("CREATE TABLE schema_meta (key TEXT PRIMARY KEY, value TEXT)")
conn.execute("CREATE TABLE leases (lease_id TEXT PRIMARY KEY, status TEXT)")
conn.execute(
"INSERT INTO schema_meta(key, value) VALUES ('schema_version', ?)",
(str(schema_version),),
)
conn.execute("INSERT INTO leases(lease_id, status) VALUES ('l1', 'active')")
conn.commit()
conn.close()
def test_missing_database_is_down(self):
probe = probe_control_plane_db(str(Path(self.tmp.name) / "absent.db"))
self.assertEqual(probe.status, STATUS_DOWN)
self.assertTrue(probe.required)
self.assertIsNotNone(probe.latency_ms)
def test_matching_schema_is_ok(self):
self._build_db(control_plane_db.SCHEMA_VERSION)
probe = probe_control_plane_db(self.db_path)
self.assertEqual(probe.status, STATUS_OK)
self.assertEqual(
probe.metadata["schema_version"], control_plane_db.SCHEMA_VERSION
)
self.assertEqual(probe.metadata["active_leases"], 1)
def test_mismatched_schema_is_degraded(self):
self._build_db(control_plane_db.SCHEMA_VERSION + 99)
probe = probe_control_plane_db(self.db_path)
self.assertEqual(probe.status, STATUS_DEGRADED)
def test_probe_does_not_create_a_database(self):
"""A health check must never initialise the substrate it inspects."""
absent = str(Path(self.tmp.name) / "never-created.db")
probe_control_plane_db(absent)
self.assertFalse(Path(absent).exists())
def test_unreadable_database_is_down_not_raised(self):
Path(self.db_path).write_text("this is not a sqlite database")
probe = probe_control_plane_db(self.db_path)
self.assertEqual(probe.status, STATUS_DOWN)
class TestRedaction(unittest.TestCase):
"""AC4 — redaction. No credential-shaped text crosses the boundary."""
def test_redacts_token_assignment(self):
cleaned = redact("failed with token=ghp_ABCDEFGHIJKLMNOPQRSTUVWXYZ012345")
self.assertNotIn("ghp_ABCDEFGHIJKLMNOPQRSTUVWXYZ012345", cleaned)
self.assertIn("[redacted]", cleaned)
def test_redacts_authorization_header_text(self):
cleaned = redact("Authorization: Bearer abcdefghijklmnopqrstuvwxyz123456")
self.assertNotIn("abcdefghijklmnopqrstuvwxyz123456", cleaned)
def test_redacts_long_opaque_strings(self):
cleaned = redact("value 0123456789abcdef0123456789abcdef here")
self.assertNotIn("0123456789abcdef0123456789abcdef", cleaned)
def test_url_userinfo_and_query_are_stripped(self):
cleaned = redact_url("https://user:[email protected]/api/v1?token=xyz")
self.assertNotIn("secretpass", cleaned)
self.assertNotIn("token=xyz", cleaned)
self.assertEqual(cleaned, "https://gitea.example.com/api/v1")
def test_url_inside_free_text_is_redacted(self):
cleaned = redact("GET https://u:[email protected]/x?token=abc failed")
self.assertNotIn("u:p@", cleaned)
self.assertNotIn("token=abc", cleaned)
def test_gitea_probe_failure_detail_is_redacted(self):
boom = RuntimeError(
"connection refused for https://user:[email protected]/api/v1/version"
)
with mock.patch("webui.system_health.get_auth_header", return_value="token x"), \
mock.patch("webui.system_health.api_request", side_effect=boom):
probe = probe_gitea("gitea.example.com")
self.assertEqual(probe.status, STATUS_DOWN)
self.assertNotIn("hunter2", probe.detail)
self.assertEqual(scan_text_for_client_secrets(probe.detail), [])
def test_credential_guard_refusal_is_a_status_not_a_crash(self):
with mock.patch(
"webui.system_health.get_auth_header",
side_effect=RuntimeError("daemon guard refused"),
):
probe = probe_gitea("gitea.example.com")
self.assertEqual(probe.status, STATUS_DEGRADED)
self.assertFalse(probe.required)
class TestNamespaceSummaries(unittest.TestCase):
"""A web process cannot prove IDE namespace health, and must not claim to."""
def test_every_namespace_reports_unproven(self):
rows = namespace_summaries()
self.assertTrue(rows)
for row in rows:
with self.subTest(namespace=row["namespace"]):
self.assertEqual(row["status"], "unproven")
self.assertFalse(row["ide_namespace_proven"])
self.assertIn("client_namespace", row["reason"])
class TestSystemHealthRoutes(CleanParityMixin, unittest.TestCase):
"""The HTTP surface: versioned path, status codes, read-only guard."""
def setUp(self):
super().setUp()
clear_probe_cache()
self.addCleanup(clear_probe_cache)
self.client = TestClient(create_app())
def _patch_snapshot(self, probes, daemon_head="abc123"):
snapshot = load_system_health(probes=probes, daemon_head=daemon_head)
patcher = mock.patch(
"webui.app.load_system_health",
return_value=snapshot,
)
patcher.start()
self.addCleanup(patcher.stop)
return snapshot
def test_versioned_route_is_registered(self):
self.assertEqual(API_PATH, "/api/v1/system/health")
self._patch_snapshot(_ALL_HEALTHY)
response = self.client.get(API_PATH)
self.assertEqual(response.status_code, 200)
def test_healthy_payload_shape(self):
self._patch_snapshot(_ALL_HEALTHY)
data = self.client.get(API_PATH).json()
self.assertEqual(data["status"], STATUS_OK)
self.assertTrue(data["readiness"]["ready"])
self.assertTrue(data["readiness"]["complete"])
self.assertEqual(data["api"], API_PATH)
self.assertEqual(len(data["dependencies"]), 3)
for key in ("version", "process", "stale_runtime", "mcp_namespaces"):
self.assertIn(key, data)
self.assertIn("uptime_seconds", data["process"])
self.assertIn("mutation_safe", data["stale_runtime"])
def test_degraded_dependency_returns_503(self):
probes = (
_probe("control_plane_db", STATUS_DOWN, detail="missing", kind="sqlite"),
_probe("repository", STATUS_OK, kind="git"),
)
self._patch_snapshot(probes)
response = self.client.get(API_PATH)
self.assertEqual(response.status_code, 503)
data = response.json()
self.assertFalse(data["readiness"]["ready"])
self.assertTrue(data["readiness"]["reasons"])
def test_dependency_entries_expose_status_and_latency(self):
self._patch_snapshot(_ALL_HEALTHY)
data = self.client.get(API_PATH).json()
names = {entry["name"] for entry in data["dependencies"]}
self.assertEqual(names, {"control_plane_db", "repository", "gitea"})
for entry in data["dependencies"]:
with self.subTest(dependency=entry["name"]):
self.assertIn("status", entry)
self.assertIn("required", entry)
self.assertIn("latency_ms", entry)
def test_response_body_carries_no_client_secrets(self):
self._patch_snapshot(_ALL_HEALTHY)
body = self.client.get(API_PATH).text
self.assertEqual(scan_text_for_client_secrets(body), [])
def test_deep_flag_is_forwarded(self):
snapshot = load_system_health(probes=_ALL_HEALTHY, daemon_head="abc")
with mock.patch(
"webui.app.load_system_health", return_value=snapshot
) as loader:
self.client.get(f"{API_PATH}?deep=1")
loader.assert_called_once_with(deep=True)
def test_shallow_is_the_default(self):
snapshot = load_system_health(probes=_ALL_HEALTHY, daemon_head="abc")
with mock.patch(
"webui.app.load_system_health", return_value=snapshot
) as loader:
self.client.get(API_PATH)
loader.assert_called_once_with(deep=False)
def test_route_rejects_mutation_methods(self):
for method in ("POST", "PUT", "PATCH", "DELETE"):
with self.subTest(method=method):
response = self.client.request(method, API_PATH)
self.assertEqual(response.status_code, 405)
self.assertEqual(response.json()["error"], "read-only-mvp")
def test_default_shallow_call_skips_the_network_probe(self):
"""The expensive probe must not run unless it was asked for."""
with mock.patch("webui.system_health.probe_gitea") as probe:
snapshot = load_system_health(deep=False)
probe.assert_not_called()
gitea = next(p for p in snapshot.dependencies if p.name == "gitea")
self.assertEqual(gitea.status, STATUS_SKIPPED)
class TestHealthRouteBackwardCompatibility(unittest.TestCase):
"""`/health` is expanded additively; MVP consumers must keep working."""
def setUp(self):
self.client = TestClient(create_app())
def test_mvp_keys_are_unchanged(self):
data = self.client.get("/health").json()
self.assertEqual(data["status"], "ok")
self.assertEqual(data["service"], "mcp-control-plane-webui")
self.assertEqual(data["mode"], "read-only-mvp")
self.assertIn("timestamp", data)
self.assertEqual(data["deployment"]["mode"], "internal-operator-console")
def test_health_points_at_the_versioned_api(self):
data = self.client.get("/health").json()
self.assertEqual(data["system_health_api"], API_PATH)
self.assertIn("uptime_seconds", data)
self.assertIn("started_at", data)
def test_health_runs_no_dependency_probe(self):
"""Liveness must stay cheap: no probe, no snapshot assembly."""
with mock.patch("webui.app.load_system_health") as loader:
response = self.client.get("/health")
self.assertEqual(response.status_code, 200)
loader.assert_not_called()
class TestSnapshotSerialisation(CleanParityMixin, unittest.TestCase):
def test_snapshot_dict_is_json_serialisable(self):
snapshot = load_system_health(probes=_ALL_HEALTHY, daemon_head="abc123")
encoded = json.dumps(snapshot_to_dict(snapshot))
self.assertIn("readiness", encoded)
if __name__ == "__main__":
unittest.main()
+345
View File
@@ -0,0 +1,345 @@
"""Tests for the system-health dashboard view (#639).
Covers the acceptance criteria directly: the page renders the health DTO
fields (AC1), degraded dependencies are visible (AC2), stale runtime is warned
prominently and never rendered as mutation-safe (AC3), healthy and degraded
fixtures both render (AC4), and the shell carries a nav entry (AC5).
"""
import sys
import unittest
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
from starlette.testclient import TestClient
from webui.app import create_app
from webui.deployment_boundary import scan_text_for_client_secrets
from webui.layout import render_page
from webui.nav import iter_nav_items
from webui.system_health import (
STATUS_DEGRADED,
STATUS_DOWN,
STATUS_OK,
STATUS_SKIPPED,
STATUS_UNPROVEN,
DependencyProbe,
StaleRuntime,
SystemHealthSnapshot,
VersionInfo,
)
from webui.system_health_views import render_system_health_page
DASHBOARD_PATH = "/system-health"
def _version(*, known: bool = True) -> VersionInfo:
return VersionInfo(
git_sha="1c455b6ec0f9cb761fe6248de68c17e061fb5ecd" if known else None,
git_describe="v0.4.1-12-g1c455b6" if known else None,
control_plane_schema_version=4 if known else None,
python_version="3.13.1",
known=known,
)
def _parity(*, stale: bool = False, determinable: bool = True) -> StaleRuntime:
if stale:
return StaleRuntime(
daemon_head="aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
checkout_head="bbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbb",
remote_head="cccccccccccccccccccccccccccccccccccccccc",
stale=True,
determinable=True,
mutation_safe=False,
reasons=("runtime, checkout, and remote commits disagree",),
)
if not determinable:
return StaleRuntime(
daemon_head=None,
checkout_head=None,
remote_head=None,
stale=False,
determinable=False,
mutation_safe=False,
reasons=("local checkout HEAD could not be read",),
)
return StaleRuntime(
daemon_head="1c455b6ec0f9cb761fe6248de68c17e061fb5ecd",
checkout_head="1c455b6ec0f9cb761fe6248de68c17e061fb5ecd",
remote_head="1c455b6ec0f9cb761fe6248de68c17e061fb5ecd",
stale=False,
determinable=True,
mutation_safe=True,
reasons=(),
)
def _snapshot(
*,
status: str = STATUS_OK,
ready: bool = True,
readiness_complete: bool = True,
readiness_reasons: tuple[str, ...] = (),
dependencies: tuple[DependencyProbe, ...] | None = None,
parity: StaleRuntime | None = None,
namespaces: tuple[dict, ...] = (),
probe_errors: tuple[str, ...] = (),
version_known: bool = True,
) -> SystemHealthSnapshot:
if dependencies is None:
dependencies = (
DependencyProbe(
name="control_plane_db",
kind="sqlite",
status=STATUS_OK,
detail="schema version 4",
required=True,
latency_ms=1.25,
metadata={"schema_version": 4},
),
)
return SystemHealthSnapshot(
status=status,
ready=ready,
readiness_complete=readiness_complete,
readiness_reasons=readiness_reasons,
service="mcp-control-plane-webui",
mode="read-only",
version=_version(known=version_known),
started_at="2026-07-23T19:50:47+00:00",
uptime_seconds=3661.5,
timestamp="2026-07-23T20:51:48+00:00",
deep_probes_requested=False,
dependencies=dependencies,
mcp_namespaces=namespaces,
stale_runtime=parity if parity is not None else _parity(),
probe_errors=probe_errors,
)
class TestHealthyRender(unittest.TestCase):
"""AC1 / AC4 — every health DTO field reaches the page."""
def setUp(self):
self.html = render_system_health_page(_snapshot())
def test_readiness_fields_render(self):
self.assertIn("System health", self.html)
self.assertIn("Ready", self.html)
self.assertIn("mcp-control-plane-webui", self.html)
self.assertIn("read-only", self.html)
self.assertIn("2026-07-23T20:51:48+00:00", self.html)
def test_version_and_uptime_render(self):
self.assertIn("1c455b6ec0f9cb761fe6248de68c17e061fb5ecd", self.html)
self.assertIn("v0.4.1-12-g1c455b6", self.html)
self.assertIn("3.13.1", self.html)
self.assertIn("3661.500s", self.html)
self.assertIn("1.02h", self.html)
def test_dependency_row_renders_with_latency(self):
self.assertIn("control_plane_db", self.html)
self.assertIn("sqlite", self.html)
self.assertIn("schema version 4", self.html)
self.assertIn("1.2 ms", self.html)
def test_healthy_page_shows_no_stale_warning(self):
self.assertNotIn("Stale runtime:", self.html)
self.assertNotIn("Staleness", self.html)
def test_unknown_version_is_labelled_not_faked(self):
html = render_system_health_page(_snapshot(version_known=False))
self.assertIn("unknown", html)
self.assertIn("unresolved", html)
class TestDegradedRender(unittest.TestCase):
"""AC2 — a degraded or unrun dependency is visible, not swallowed."""
def setUp(self):
self.deps = (
DependencyProbe(
name="control_plane_db",
kind="sqlite",
status=STATUS_OK,
detail="schema version 4",
required=True,
latency_ms=0.9,
),
DependencyProbe(
name="repository",
kind="git",
status=STATUS_DOWN,
detail="repository root is not a git checkout",
required=True,
latency_ms=4.0,
),
DependencyProbe(
name="gitea",
kind="http",
status=STATUS_SKIPPED,
detail="deep probe not requested",
required=False,
),
)
self.html = render_system_health_page(
_snapshot(
status=STATUS_DEGRADED,
ready=False,
readiness_complete=False,
readiness_reasons=("required dependency 'repository' is down",),
dependencies=self.deps,
)
)
def test_degraded_banner_names_the_dependency(self):
self.assertIn("Degraded dependencies:", self.html)
self.assertIn("repository", self.html)
def test_not_run_probe_is_reported_separately(self):
self.assertIn("Not probed:", self.html)
self.assertIn("gitea", self.html)
self.assertIn("not counted", self.html)
def test_not_ready_headline_and_reason(self):
self.assertIn("Not ready", self.html)
self.assertIn("required dependency &#x27;repository&#x27; is down", self.html)
def test_degraded_status_badge_present(self):
self.assertIn("badge-health-degraded", self.html)
self.assertIn("badge-health-down", self.html)
def test_ready_but_incomplete_is_not_shown_as_plain_ready(self):
html = render_system_health_page(
_snapshot(ready=True, readiness_complete=False)
)
self.assertIn("Ready (incomplete evidence)", html)
class TestStaleRuntimeWarning(unittest.TestCase):
"""AC3 — staleness is prominent and never claims mutation safety."""
def test_stale_runtime_warns_and_denies_mutation_safety(self):
html = render_system_health_page(_snapshot(parity=_parity(stale=True)))
self.assertIn("Stale runtime:", html)
self.assertIn("do not treat this runtime as mutation-safe", html)
self.assertIn("<tr><th>Mutation safe</th><td>False</td></tr>", html)
def test_indeterminate_parity_is_not_reported_safe(self):
html = render_system_health_page(
_snapshot(parity=_parity(determinable=False))
)
self.assertIn("Staleness", html)
self.assertIn("<tr><th>Mutation safe</th><td>False</td></tr>", html)
self.assertIn("<tr><th>Determinable</th><td>False</td></tr>", html)
def test_healthy_parity_reports_mutation_safe_true(self):
html = render_system_health_page(_snapshot())
self.assertIn("<tr><th>Mutation safe</th><td>True</td></tr>", html)
class TestNamespacesAndErrors(unittest.TestCase):
def test_unproven_namespace_rows_render(self):
html = render_system_health_page(
_snapshot(
namespaces=(
{
"namespace": "gitea-author",
"required_tool": "gitea_lock_issue",
"status": STATUS_UNPROVEN,
"ide_namespace_proven": False,
"reason": "the web console cannot invoke the IDE-managed MCP client",
},
)
)
)
self.assertIn("gitea-author", html)
self.assertIn("gitea_lock_issue", html)
self.assertIn("badge-health-unproven", html)
def test_no_namespaces_degrades_gracefully(self):
html = render_system_health_page(_snapshot(namespaces=()))
self.assertIn("No MCP namespaces are declared.", html)
def test_probe_errors_render_when_present(self):
html = render_system_health_page(
_snapshot(probe_errors=("probe raised: disk offline",))
)
self.assertIn("Probe errors", html)
self.assertIn("disk offline", html)
def test_probe_error_card_absent_when_clean(self):
self.assertNotIn("Probe errors", render_system_health_page(_snapshot()))
class TestReadOnlyAndRedaction(unittest.TestCase):
def test_no_restart_or_kill_controls(self):
html = render_system_health_page(_snapshot())
self.assertNotIn("<button", html)
self.assertNotIn("<form", html)
self.assertNotIn("pkill", html)
self.assertIn("read-only", html)
def test_recovery_points_at_sanctioned_path(self):
html = render_system_health_page(_snapshot())
self.assertIn("Reconnect the MCP client", html)
self.assertIn("Never kill the daemon process manually", html)
def test_secret_shaped_detail_is_redacted(self):
leaky = DependencyProbe(
name="gitea",
kind="http",
status=STATUS_DOWN,
detail="auth failed for token=ghp_ABCDEFGHIJKLMNOPQRSTUVWXYZ0123456789",
required=False,
latency_ms=12.0,
)
html = render_system_health_page(_snapshot(dependencies=(leaky,)))
self.assertNotIn("ghp_ABCDEFGHIJKLMNOPQRSTUVWXYZ0123456789", html)
def test_html_in_detail_is_escaped(self):
hostile = DependencyProbe(
name="repository",
kind="git",
status=STATUS_DOWN,
detail="<script>alert(1)</script>",
required=True,
)
html = render_system_health_page(_snapshot(dependencies=(hostile,)))
self.assertNotIn("<script>", html)
self.assertIn("&lt;script&gt;", html)
class TestNavAndRoute(unittest.TestCase):
"""AC5 — the shell links the dashboard, and the route serves it."""
def setUp(self):
self.client = TestClient(create_app())
def test_nav_contains_system_health(self):
self.assertIn(
(DASHBOARD_PATH, "System health"),
[(item.href, item.label) for item in iter_nav_items()],
)
def test_rendered_shell_links_dashboard(self):
page = render_page(title="Home", body_html="<p>x</p>")
self.assertIn(f'href="{DASHBOARD_PATH}"', page)
def test_route_renders_dashboard(self):
response = self.client.get(DASHBOARD_PATH)
self.assertEqual(response.status_code, 200)
self.assertIn("System health", response.text)
self.assertIn("Stale-runtime parity", response.text)
def test_route_is_read_only(self):
self.assertEqual(self.client.post(DASHBOARD_PATH).status_code, 405)
def test_live_page_leaks_no_client_secret(self):
findings = scan_text_for_client_secrets(self.client.get(DASHBOARD_PATH).text)
self.assertEqual(findings, [])
if __name__ == "__main__": # pragma: no cover
unittest.main()
File diff suppressed because it is too large Load Diff
+3 -1
View File
@@ -35,8 +35,10 @@ class TestCanonicalRepoRoot(unittest.TestCase):
self.assertEqual(root, CONTROL_ROOT)
def test_falls_back_when_git_unavailable(self):
# When fallback is a path under <repo>/branches/<worktree>, recover
# <repo> via commonpath ancestry (never string-split on "/branches/").
root = amw.resolve_canonical_repo_root("/missing/path", MCP_PROCESS_ROOT)
self.assertEqual(root, os.path.realpath(MCP_PROCESS_ROOT))
self.assertEqual(root, os.path.realpath(CONTROL_ROOT))
class TestWorkspaceRepoMembership(unittest.TestCase):
+24 -2
View File
@@ -134,13 +134,35 @@ class TestClassification(unittest.TestCase):
self.assertEqual(cls, wca.CLASS_ACTIVE_OPEN_PR)
self.assertFalse(wca.is_removable(cls))
def test_stale_clean_issue_worktree_removable(self):
# Scenario 5: clean issue worktree, TTL expired, no lock -> removable.
def test_stale_clean_issue_worktree_needs_merged_pr_proof(self):
# Scenario 5 (#858): age is not proof that the branch landed, so a
# TTL-expired issue worktree stays active work. Only authoritative
# merged-PR evidence makes it removable, which is what keeps a
# worktree holding unmerged commits from being reclaimed by age.
cls = wca.classify_worktree(
workflow_type=wca.WORKFLOW_ISSUE_WORK,
is_dirty=False,
ttl_expired=True,
)
self.assertEqual(cls, wca.CLASS_ACTIVE_ISSUE_WORK)
self.assertFalse(wca.is_removable(cls))
cls = wca.classify_worktree(
workflow_type=wca.WORKFLOW_ISSUE_WORK,
is_dirty=False,
ttl_expired=True,
merged_pr_cleanup={"proven": True},
)
self.assertEqual(cls, wca.CLASS_CLEAN_STALE_REMOVABLE)
self.assertTrue(wca.is_removable(cls))
def test_stale_clean_conflict_fix_worktree_removable(self):
# conflict_fix keeps the original TTL rule; #858 changed issue work only.
cls = wca.classify_worktree(
workflow_type=wca.WORKFLOW_CONFLICT_FIX,
is_dirty=False,
ttl_expired=True,
)
self.assertEqual(cls, wca.CLASS_CLEAN_STALE_REMOVABLE)
self.assertTrue(wca.is_removable(cls))
+433
View File
@@ -0,0 +1,433 @@
"""Model usage, token cost, latency, and workflow-performance analytics (#651, Phase 4).
Ingests session instrumentation metrics, aggregates usage/cost/latency percentiles
by project, role, model, issue/PR, and stage, enforcing secret redaction and
explicitly rendering missing metrics as "Unknown" without zero-fabrication.
"""
from __future__ import annotations
import math
from dataclasses import asdict, dataclass
from typing import Any, Sequence
import control_plane_db
from webui import console_redaction
ANALYTICS_SCHEMA_VERSION = 1
@dataclass(frozen=True)
class UsageEvent:
usage_id: int
session_id: str | None
remote: str
org: str
repo: str
project_id: str | None
role: str
model: str
issue_number: int | None
pr_number: int | None
stage: str
input_tokens: int | None
output_tokens: int | None
total_tokens: int | None
estimated_cost_usd: float | None
latency_ms: int | None
duration_ms: int | None
status: str
metadata: str | None
created_at: str
def to_dict(self) -> dict[str, Any]:
d = asdict(self)
if d["metadata"]:
d["metadata"] = console_redaction.redact_text(d["metadata"])
return d
@dataclass(frozen=True)
class GroupMetrics:
name: str
total_events: int
events_with_tokens: int
input_tokens: int | None
output_tokens: int | None
total_tokens: int | None
events_with_cost: int
estimated_cost_usd: float | None
events_with_latency: int
latency_p50_ms: float | None
latency_p90_ms: float | None
latency_p95_ms: float | None
latency_p99_ms: float | None
latency_avg_ms: float | None
events_with_duration: int
duration_avg_ms: float | None
display_tokens: str
display_cost: str
display_latency_p50: str
display_latency_p90: str
display_duration_avg: str
def to_dict(self) -> dict[str, Any]:
return asdict(self)
@dataclass(frozen=True)
class AnalyticsSnapshot:
ok: bool
reason: str
schema_version: int
remote: str
org: str
repo: str
total_events: int
overall_summary: GroupMetrics
by_project: dict[str, GroupMetrics]
by_role: dict[str, GroupMetrics]
by_model: dict[str, GroupMetrics]
by_work_item: dict[str, GroupMetrics]
by_stage: dict[str, GroupMetrics]
events: tuple[UsageEvent, ...]
def to_dict(self) -> dict[str, Any]:
return {
"ok": self.ok,
"reason": self.reason,
"schema_version": self.schema_version,
"remote": self.remote,
"org": self.org,
"repo": self.repo,
"total_events": self.total_events,
"overall_summary": self.overall_summary.to_dict(),
"by_project": {k: v.to_dict() for k, v in self.by_project.items()},
"by_role": {k: v.to_dict() for k, v in self.by_role.items()},
"by_model": {k: v.to_dict() for k, v in self.by_model.items()},
"by_work_item": {k: v.to_dict() for k, v in self.by_work_item.items()},
"by_stage": {k: v.to_dict() for k, v in self.by_stage.items()},
"events": [e.to_dict() for e in self.events],
}
def compute_percentile(values: Sequence[float | int], percentile: float) -> float | None:
if not values:
return None
sorted_vals = sorted(values)
n = len(sorted_vals)
if n == 1:
return float(sorted_vals[0])
k = (n - 1) * (percentile / 100.0)
f = math.floor(k)
c = math.ceil(k)
if f == c:
return float(sorted_vals[int(f)])
d0 = sorted_vals[int(f)] * (c - k)
d1 = sorted_vals[int(c)] * (k - f)
return float(d0 + d1)
def aggregate_events(group_name: str, events: Sequence[UsageEvent]) -> GroupMetrics:
total_events = len(events)
if total_events == 0:
return GroupMetrics(
name=group_name,
total_events=0,
events_with_tokens=0,
input_tokens=None,
output_tokens=None,
total_tokens=None,
events_with_cost=0,
estimated_cost_usd=None,
events_with_latency=0,
latency_p50_ms=None,
latency_p90_ms=None,
latency_p95_ms=None,
latency_p99_ms=None,
latency_avg_ms=None,
events_with_duration=0,
duration_avg_ms=None,
display_tokens="Unknown",
display_cost="Unknown",
display_latency_p50="Unknown",
display_latency_p90="Unknown",
display_duration_avg="Unknown",
)
token_events = [
e for e in events
if e.total_tokens is not None or e.input_tokens is not None or e.output_tokens is not None
]
events_with_tokens = len(token_events)
if events_with_tokens > 0:
input_tokens = sum(e.input_tokens or 0 for e in token_events)
output_tokens = sum(e.output_tokens or 0 for e in token_events)
total_tokens = sum(
e.total_tokens if e.total_tokens is not None else ((e.input_tokens or 0) + (e.output_tokens or 0))
for e in token_events
)
display_tokens = f"{total_tokens:,}"
else:
input_tokens = None
output_tokens = None
total_tokens = None
display_tokens = "Unknown"
cost_events = [e for e in events if e.estimated_cost_usd is not None]
events_with_cost = len(cost_events)
if events_with_cost > 0:
estimated_cost_usd = round(sum(e.estimated_cost_usd for e in cost_events), 6)
display_cost = f"${estimated_cost_usd:.4f}"
else:
estimated_cost_usd = None
display_cost = "Unknown"
latency_vals = [e.latency_ms for e in events if e.latency_ms is not None]
events_with_latency = len(latency_vals)
if events_with_latency > 0:
latency_p50_ms = compute_percentile(latency_vals, 50.0)
latency_p90_ms = compute_percentile(latency_vals, 90.0)
latency_p95_ms = compute_percentile(latency_vals, 95.0)
latency_p99_ms = compute_percentile(latency_vals, 99.0)
latency_avg_ms = round(sum(latency_vals) / events_with_latency, 2)
display_latency_p50 = f"{round(latency_p50_ms, 1)} ms" if latency_p50_ms is not None else "Unknown"
display_latency_p90 = f"{round(latency_p90_ms, 1)} ms" if latency_p90_ms is not None else "Unknown"
else:
latency_p50_ms = None
latency_p90_ms = None
latency_p95_ms = None
latency_p99_ms = None
latency_avg_ms = None
display_latency_p50 = "Unknown"
display_latency_p90 = "Unknown"
duration_vals = [e.duration_ms for e in events if e.duration_ms is not None]
events_with_duration = len(duration_vals)
if events_with_duration > 0:
duration_avg_ms = round(sum(duration_vals) / events_with_duration, 2)
display_duration_avg = f"{round(duration_avg_ms / 1000.0, 2)} s" if duration_avg_ms >= 1000 else f"{round(duration_avg_ms, 1)} ms"
else:
duration_avg_ms = None
display_duration_avg = "Unknown"
return GroupMetrics(
name=group_name,
total_events=total_events,
events_with_tokens=events_with_tokens,
input_tokens=input_tokens,
output_tokens=output_tokens,
total_tokens=total_tokens,
events_with_cost=events_with_cost,
estimated_cost_usd=estimated_cost_usd,
events_with_latency=events_with_latency,
latency_p50_ms=latency_p50_ms,
latency_p90_ms=latency_p90_ms,
latency_p95_ms=latency_p95_ms,
latency_p99_ms=latency_p99_ms,
latency_avg_ms=latency_avg_ms,
events_with_duration=events_with_duration,
duration_avg_ms=duration_avg_ms,
display_tokens=display_tokens,
display_cost=display_cost,
display_latency_p50=display_latency_p50,
display_latency_p90=display_latency_p90,
display_duration_avg=display_duration_avg,
)
def record_usage(
*,
db_path: str | None = None,
session_id: str | None = None,
remote: str = "dadeschools",
org: str = "",
repo: str = "",
project_id: str | None = None,
role: str = "unknown",
model: str = "unknown",
issue_number: int | None = None,
pr_number: int | None = None,
stage: str = "unknown",
input_tokens: int | None = None,
output_tokens: int | None = None,
total_tokens: int | None = None,
estimated_cost_usd: float | None = None,
latency_ms: int | None = None,
duration_ms: int | None = None,
status: str = "success",
metadata: str | dict[str, Any] | None = None,
created_at: str | None = None,
) -> int:
"""Ingest/record a single usage event with optional metrics."""
db = control_plane_db.ControlPlaneDB(db_path=db_path)
return db.record_usage_event(
session_id=session_id,
remote=remote,
org=org,
repo=repo,
project_id=project_id,
role=role,
model=model,
issue_number=issue_number,
pr_number=pr_number,
stage=stage,
input_tokens=input_tokens,
output_tokens=output_tokens,
total_tokens=total_tokens,
estimated_cost_usd=estimated_cost_usd,
latency_ms=latency_ms,
duration_ms=duration_ms,
status=status,
metadata=metadata,
created_at=created_at,
)
def load_analytics(
*,
db_path: str | None = None,
remote: str | None = None,
org: str | None = None,
repo: str | None = None,
project_id: str | None = None,
role: str | None = None,
model: str | None = None,
stage: str | None = None,
issue_number: int | None = None,
pr_number: int | None = None,
limit: int = 500,
) -> AnalyticsSnapshot:
"""Load analytics snapshot aggregated by project, role, model, issue/PR, and stage."""
remote_filter = (remote or "").strip() or None
org_filter = (org or "").strip() or None
repo_filter = (repo or "").strip() or None
role_filter = (role or "").strip() or None
model_filter = (model or "").strip() or None
stage_filter = (stage or "").strip() or None
# F4: coerce optional scope filters to str so AnalyticsSnapshot never holds None.
scope_remote = (remote or "").strip()
scope_org = (org or "").strip()
scope_repo = (repo or "").strip()
try:
db = control_plane_db.ControlPlaneDB(db_path=db_path)
rows = db.query_usage_events(
remote=remote_filter,
org=org_filter,
repo=repo_filter,
project_id=project_id,
role=role_filter,
model=model_filter,
stage=stage_filter,
issue_number=issue_number,
pr_number=pr_number,
limit=limit,
)
except Exception as exc:
empty_summary = aggregate_events("Overall", [])
return AnalyticsSnapshot(
ok=False,
reason=f"control_plane_db_unavailable: {exc}",
schema_version=ANALYTICS_SCHEMA_VERSION,
remote=scope_remote,
org=scope_org,
repo=scope_repo,
total_events=0,
overall_summary=empty_summary,
by_project={},
by_role={},
by_model={},
by_work_item={},
by_stage={},
events=(),
)
parsed_events: list[UsageEvent] = []
for r in rows:
meta = console_redaction.redact_text(r.get("metadata")) if r.get("metadata") else None
parsed_events.append(
UsageEvent(
usage_id=r["usage_id"],
session_id=r.get("session_id"),
remote=r.get("remote") or scope_remote,
org=r.get("org") or scope_org,
repo=r.get("repo") or scope_repo,
project_id=r.get("project_id"),
role=r.get("role") or "unknown",
model=r.get("model") or "unknown",
issue_number=r.get("issue_number"),
pr_number=r.get("pr_number"),
stage=r.get("stage") or "unknown",
input_tokens=r.get("input_tokens"),
output_tokens=r.get("output_tokens"),
total_tokens=r.get("total_tokens"),
estimated_cost_usd=r.get("estimated_cost_usd"),
latency_ms=r.get("latency_ms"),
duration_ms=r.get("duration_ms"),
status=r.get("status") or "success",
metadata=meta,
created_at=r.get("created_at") or "",
)
)
overall_summary = aggregate_events("Overall", parsed_events)
# Group by project
groups_by_project: dict[str, list[UsageEvent]] = {}
for e in parsed_events:
key = e.project_id or (f"{e.org}/{e.repo}" if e.org and e.repo else "default")
groups_by_project.setdefault(key, []).append(e)
by_project = {k: aggregate_events(k, v) for k, v in groups_by_project.items()}
# Group by role
groups_by_role: dict[str, list[UsageEvent]] = {}
for e in parsed_events:
groups_by_role.setdefault(e.role, []).append(e)
by_role = {k: aggregate_events(k, v) for k, v in groups_by_role.items()}
# Group by model
groups_by_model: dict[str, list[UsageEvent]] = {}
for e in parsed_events:
groups_by_model.setdefault(e.model, []).append(e)
by_model = {k: aggregate_events(k, v) for k, v in groups_by_model.items()}
# Group by work item
groups_by_work_item: dict[str, list[UsageEvent]] = {}
for e in parsed_events:
if e.issue_number:
key = f"issue #{e.issue_number}"
elif e.pr_number:
key = f"pr #{e.pr_number}"
else:
key = "unlinked"
groups_by_work_item.setdefault(key, []).append(e)
by_work_item = {k: aggregate_events(k, v) for k, v in groups_by_work_item.items()}
# Group by stage
groups_by_stage: dict[str, list[UsageEvent]] = {}
for e in parsed_events:
groups_by_stage.setdefault(e.stage, []).append(e)
by_stage = {k: aggregate_events(k, v) for k, v in groups_by_stage.items()}
return AnalyticsSnapshot(
ok=True,
reason="ok",
schema_version=ANALYTICS_SCHEMA_VERSION,
remote=scope_remote,
org=scope_org,
repo=scope_repo,
total_events=len(parsed_events),
overall_summary=overall_summary,
by_project=by_project,
by_role=by_role,
by_model=by_model,
by_work_item=by_work_item,
by_stage=by_stage,
events=tuple(parsed_events),
)
def snapshot_to_dict(snapshot: AnalyticsSnapshot) -> dict[str, Any]:
return snapshot.to_dict()
+248
View File
@@ -0,0 +1,248 @@
"""HTML views for the Model Usage & Performance Analytics console (#651)."""
from __future__ import annotations
import html
from webui.analytics_loader import AnalyticsSnapshot, GroupMetrics, UsageEvent
from webui.layout import render_page
def _escape(text: object) -> str:
"""HTML-escape dynamic analytics fields (mirrors audit_views / project_views)."""
return html.escape(str(text), quote=True)
def _render_badge(text: str, badge_type: str = "muted") -> str:
return f'<span class="badge badge-{_escape(badge_type)}">{_escape(text)}</span>'
def _render_group_table(title: str, groups: dict[str, GroupMetrics], key_header: str = "Group") -> str:
if not groups:
return (
f"<h3>{_escape(title)}</h3>"
'<div class="card"><p class="muted">No telemetry events recorded for this dimension.</p></div>'
)
rows = []
for key, g in sorted(groups.items(), key=lambda x: x[1].total_events, reverse=True):
cost_cell = (
f'<span class="accent">{_escape(g.display_cost)}</span>'
if g.events_with_cost > 0
else _render_badge("Unknown")
)
tokens_cell = (
_escape(g.display_tokens)
if g.events_with_tokens > 0
else _render_badge("Unknown")
)
lat_p50 = (
_escape(g.display_latency_p50)
if g.events_with_latency > 0
else _render_badge("Unknown")
)
lat_p90 = (
_escape(g.display_latency_p90)
if g.events_with_latency > 0
else _render_badge("Unknown")
)
dur_avg = (
_escape(g.display_duration_avg)
if g.events_with_duration > 0
else _render_badge("Unknown")
)
rows.append(
"<tr>"
f"<td><strong>{_escape(key)}</strong></td>"
f"<td>{g.total_events}</td>"
f"<td>{tokens_cell}</td>"
f"<td>{cost_cell}</td>"
f"<td>{lat_p50}</td>"
f"<td>{lat_p90}</td>"
f"<td>{dur_avg}</td>"
"</tr>"
)
rows_html = "".join(rows)
return f"""
<h3>{_escape(title)}</h3>
<div class="card" style="overflow-x: auto;">
<table class="data-table">
<thead>
<tr>
<th>{_escape(key_header)}</th>
<th>Events</th>
<th>Total Tokens</th>
<th>Est. Cost</th>
<th>Latency (p50)</th>
<th>Latency (p90)</th>
<th>Avg Stage Duration</th>
</tr>
</thead>
<tbody>
{rows_html}
</tbody>
</table>
</div>
"""
def _render_events_table(events: tuple[UsageEvent, ...]) -> str:
if not events:
return (
"<h3>Recent Usage & Instrumentation Events</h3>"
'<div class="card"><p class="muted">No individual telemetry events recorded yet. Opt-in instrumentation via session logging or authorized POST /api/v1/analytics/usage.</p></div>'
)
rows = []
for e in list(events)[-50:]: # Display latest 50
if e.issue_number is not None:
work_item = f"issue #{e.issue_number}"
elif e.pr_number is not None:
work_item = f"pr #{e.pr_number}"
else:
work_item = "unlinked"
tokens = (
_escape(f"{e.total_tokens:,}")
if e.total_tokens is not None
else _render_badge("Unknown")
)
cost = (
_escape(f"${e.estimated_cost_usd:.4f}")
if e.estimated_cost_usd is not None
else _render_badge("Unknown")
)
latency = (
_escape(f"{e.latency_ms} ms")
if e.latency_ms is not None
else _render_badge("Unknown")
)
duration = (
_escape(f"{e.duration_ms} ms")
if e.duration_ms is not None
else _render_badge("Unknown")
)
status_badge = _render_badge(
e.status, "success" if e.status == "success" else "danger"
)
rows.append(
"<tr>"
f"<td>#{e.usage_id}</td>"
f"<td><small>{_escape(e.created_at)}</small></td>"
f"<td><span class=\"badge\">{_escape(e.role)}</span></td>"
f"<td><strong>{_escape(e.model)}</strong></td>"
f"<td>{_escape(e.stage)}</td>"
f"<td>{_escape(work_item)}</td>"
f"<td>{tokens}</td>"
f"<td>{cost}</td>"
f"<td>{latency}</td>"
f"<td>{duration}</td>"
f"<td>{status_badge}</td>"
"</tr>"
)
rows_html = "".join(rows)
return f"""
<h3>Recent Telemetry Events</h3>
<div class="card" style="overflow-x: auto;">
<table class="data-table">
<thead>
<tr>
<th>ID</th>
<th>Timestamp</th>
<th>Role</th>
<th>Model</th>
<th>Stage</th>
<th>Work Item</th>
<th>Tokens</th>
<th>Cost</th>
<th>Latency</th>
<th>Duration</th>
<th>Status</th>
</tr>
</thead>
<tbody>
{rows_html}
</tbody>
</table>
</div>
"""
def render_analytics_page(snapshot: AnalyticsSnapshot) -> str:
"""Render the main Model Usage & Performance Analytics console page."""
summary = snapshot.overall_summary
kpi_tokens = (
_escape(summary.display_tokens)
if summary.events_with_tokens > 0
else _render_badge("Unknown")
)
kpi_cost = (
_escape(summary.display_cost)
if summary.events_with_cost > 0
else _render_badge("Unknown")
)
kpi_lat_p50 = (
_escape(summary.display_latency_p50)
if summary.events_with_latency > 0
else _render_badge("Unknown")
)
kpi_dur_avg = (
_escape(summary.display_duration_avg)
if summary.events_with_duration > 0
else _render_badge("Unknown")
)
status_notice = ""
if not snapshot.ok:
status_notice = (
f'<div class="card warning-card"><strong>Degraded Data Source:</strong> '
f'{_escape(snapshot.reason)}</div>'
)
body_html = f"""
<h2>Model Usage & Performance Analytics (Phase 4)</h2>
<p class="muted">
Durable console analytics for model usage, token cost, latency percentiles, and workflow-stage performance correlated to issues, PRs, and worker roles.
</p>
{status_notice}
<div class="notice-card" style="background: rgba(91, 159, 212, 0.1); border: 1px solid var(--border); padding: 0.75rem 1rem; border-radius: 6px; margin-bottom: 1.5rem;">
<small><strong>Note on telemetry fidelity:</strong> Missing data or untracked metrics are explicitly labeled as <em>Unknown</em>. No token costs or latency metrics are zero-fabricated.</small>
</div>
<div class="card-grid" style="display: grid; grid-template-columns: repeat(auto-fit, minmax(180px, 1fr)); gap: 1rem; margin-bottom: 1.5rem;">
<div class="card">
<span class="muted" style="font-size: 0.85rem;">Total Events</span>
<h3 style="margin: 0.25rem 0 0 0;">{summary.total_events}</h3>
</div>
<div class="card">
<span class="muted" style="font-size: 0.85rem;">Total Tokens</span>
<h3 style="margin: 0.25rem 0 0 0;">{kpi_tokens}</h3>
</div>
<div class="card">
<span class="muted" style="font-size: 0.85rem;">Est. Token Cost</span>
<h3 style="margin: 0.25rem 0 0 0;">{kpi_cost}</h3>
</div>
<div class="card">
<span class="muted" style="font-size: 0.85rem;">Latency (p50)</span>
<h3 style="margin: 0.25rem 0 0 0;">{kpi_lat_p50}</h3>
</div>
<div class="card">
<span class="muted" style="font-size: 0.85rem;">Avg Stage Duration</span>
<h3 style="margin: 0.25rem 0 0 0;">{kpi_dur_avg}</h3>
</div>
</div>
{_render_group_table("Usage & Cost by Model", snapshot.by_model, "Model")}
{_render_group_table("Performance by Workflow Stage", snapshot.by_stage, "Stage")}
{_render_group_table("Usage & Cost by Role", snapshot.by_role, "Role")}
{_render_group_table("Work Item Analytics", snapshot.by_work_item, "Work Item")}
{_render_events_table(snapshot.events)}
"""
return render_page(title="Model Usage & Performance Analytics", body_html=body_html)
+361 -11
View File
@@ -12,6 +12,7 @@ from starlette.routing import Route
from webui.deployment_boundary import deployment_snapshot
from webui.layout import render_page
from webui.nav import NAV_GROUPS, STUB_PAGES
from webui.project_registry import (
ProjectRegistry,
RegistryError,
@@ -45,6 +46,25 @@ from webui.worktree_scanner import load_hygiene_snapshot, snapshot_to_dict as wo
from webui.worktree_views import render_worktrees_page
from webui.runtime_health import load_runtime_snapshot, snapshot_to_dict as runtime_snapshot_to_dict
from webui.runtime_views import render_runtime_page
from webui.inventory import (
SECTION_NAMES as _INVENTORY_SECTIONS,
load_inventory_snapshot,
snapshot_to_dict as inventory_snapshot_to_dict,
)
from webui.timeline import load_timeline, snapshot_to_dict as timeline_snapshot_to_dict
from webui.analytics_loader import (
load_analytics,
record_usage,
snapshot_to_dict as analytics_snapshot_to_dict,
)
from webui.analytics_views import render_analytics_page
from webui.system_health import (
API_PATH as SYSTEM_HEALTH_API_PATH,
load_system_health,
process_uptime,
snapshot_to_dict as system_health_to_dict,
)
from webui.system_health_views import render_system_health_page
_READ_ONLY_METHODS = frozenset({"GET", "HEAD", "OPTIONS"})
_AUDIT_MUTATION_PATHS = frozenset({"/audit", "/api/audit"})
@@ -59,35 +79,118 @@ def _stub_page(title: str, description: str) -> HTMLResponse:
return HTMLResponse(render_page(title=title, body_html=body))
_LEGACY_PAGES = (
("/queue", "Queue", "live PR and issue dashboard (#429)"),
("/projects", "Projects", "registry and onboarding (#427)"),
("/prompts", "Prompts", "canonical workflow prompt library (#428)"),
("/runtime", "Runtime", "MCP health and stale-runtime detection (#430)"),
("/audit", "Audit", "final-report paste and validator preview (#431)"),
("/worktrees", "Worktrees", "branch hygiene dashboard (#432)"),
("/leases", "Leases", "collision and lease visibility (#433)"),
("/actions", "Actions", "gated write-action framework (#434)"),
)
def _render_home_nav_groups() -> str:
groups = []
for group in NAV_GROUPS:
items = "".join(
f'<li><a href="{item.href}">{item.label}</a>'
+ ("" if item.status == "live" else " <span class=\"muted\">(stub)</span>")
+ "</li>"
for item in group.items
)
groups.append(f"<h3>{group.label}</h3><ul>{items}</ul>")
return "".join(groups)
async def home(_request: Request) -> HTMLResponse:
legacy = "".join(
f"<li><strong>{label}</strong> — {desc} "
f'(<a href="{href}">{href}</a>)</li>'
for href, label, desc in _LEGACY_PAGES
)
body = (
"<h2>Operator console</h2>"
"<p>Local entry point for MCP Control Plane operational views.</p>"
"<ul>"
"<li><strong>Queue</strong> — live PR and issue dashboard (#429)</li>"
"<li><strong>Projects</strong> — registry and onboarding (#427)</li>"
"<li><strong>Prompts</strong> — canonical workflow prompt library (#428)</li>"
"<li><strong>Runtime</strong> — MCP health and stale-runtime detection (#430)</li>"
"<li><strong>Audit</strong> — final-report paste and validator preview (#431)</li>"
"<li><strong>Worktrees</strong> — branch hygiene dashboard (#432)</li>"
"<li><strong>Leases</strong> — collision and lease visibility (#433)</li>"
"<li><strong>Actions</strong> — gated write-action framework (#434)</li>"
"</ul>"
"<p>Read-only home for the MCP Control Plane Phase 1 operator console. "
"Gitea, MCP capability gates, and canonical workflows remain the source "
"of truth; this console never mutates them.</p>"
"<h2>Phase 1 surfaces</h2>"
+ _render_home_nav_groups()
+ "<h2>MVP legacy pages</h2>"
"<ul>" + legacy + "</ul>"
)
return HTMLResponse(render_page(title="Home", body_html=body))
async def phase_stub(request: Request) -> HTMLResponse:
"""Graceful read-only placeholder for a not-yet-implemented Phase 1 surface."""
title, description = STUB_PAGES[request.url.path]
body = (
f"<h2>{title}</h2>"
f'<div class="stub"><p>{description}</p>'
"<p>Phase 1 shell placeholder — no write actions. Tracked under "
"epic #631.</p></div>"
)
return HTMLResponse(render_page(title=title, body_html=body))
async def health(_request: Request) -> JSONResponse:
"""Liveness only — deliberately cheap, runs no dependency probe (#634).
Every MVP key is retained so existing pollers keep working; the additions
are a pointer to the structured API and the in-memory process uptime.
Readiness lives at that API because answering it costs real probes.
"""
bind_host = _request.app.state.webui_bind_host
started_at, uptime_seconds = process_uptime()
return JSONResponse({
"status": "ok",
"service": "mcp-control-plane-webui",
"mode": "read-only-mvp",
"timestamp": datetime.now(timezone.utc).isoformat(),
"deployment": deployment_snapshot(bind_host=bind_host),
"started_at": started_at,
"uptime_seconds": uptime_seconds,
"system_health_api": SYSTEM_HEALTH_API_PATH,
})
def _truthy_flag(value: str | None) -> bool:
return (value or "").strip().lower() in {"1", "true", "yes", "on"}
async def api_system_health(request: Request) -> JSONResponse:
"""Structured read-only system health (#634).
`?deep=1` opts into the expensive network probe. The response status code
reflects readiness so automated checks can branch on it without parsing the
body: 200 when ready, 503 when a required dependency failed or never ran.
"""
deep = _truthy_flag(request.query_params.get("deep"))
snapshot = load_system_health(deep=deep)
payload = system_health_to_dict(snapshot)
return JSONResponse(payload, status_code=200 if snapshot.ready else 503)
async def system_health(request: Request) -> HTMLResponse:
"""Read-only system-health dashboard (#639).
Shares the #634 snapshot loader with the JSON API so the page can never
disagree with it. `?deep=1` opts into the network probe exactly as the API
does; the default page load stays cheap. The response is always 200: this
is an operator view that must render the degraded state, not withhold it.
"""
deep = _truthy_flag(request.query_params.get("deep"))
snapshot = load_system_health(deep=deep)
return HTMLResponse(
render_page(
title="System health",
body_html=render_system_health_page(snapshot),
)
)
async def queue(_request: Request) -> HTMLResponse:
snapshot = load_queue_snapshot()
return HTMLResponse(render_page(title="Queue", body_html=render_queue_page(snapshot)))
@@ -368,6 +471,30 @@ async def api_action_attempt(request: Request) -> JSONResponse:
return JSONResponse(result, status_code=status)
async def api_inventory(_request: Request) -> JSONResponse:
"""Unified read-only session/lease/lock/worktree inventory (#636)."""
snapshot = load_inventory_snapshot()
return JSONResponse(inventory_snapshot_to_dict(snapshot))
async def api_inventory_section(request: Request) -> JSONResponse:
"""Resource-split view: one inventory section under the shared schema."""
section = request.path_params["section"]
if section not in _INVENTORY_SECTIONS:
return JSONResponse(
{
"error": "unknown_section",
"detail": f"no inventory section named {section!r}",
"available": sorted(_INVENTORY_SECTIONS),
},
status_code=404,
)
snapshot = load_inventory_snapshot(include=(section,))
payload = inventory_snapshot_to_dict(snapshot)
payload["requested_section"] = section
return JSONResponse(payload)
async def api_console_security_model(_request: Request) -> JSONResponse:
"""Read-only publication of the #633 authorization/redaction/audit model."""
return JSONResponse({
@@ -377,6 +504,212 @@ async def api_console_security_model(_request: Request) -> JSONResponse:
})
def _query_int(request: Request, key: str) -> int | None:
"""Parse an optional integer query parameter; None when absent/invalid."""
raw = request.query_params.get(key)
if raw is None or not str(raw).strip():
return None
try:
return int(str(raw).strip())
except (TypeError, ValueError):
return None
def _derive_remote(host: str) -> str:
"""Map a Gitea host to its known short remote name (control-plane scope key)."""
text = (host or "").lower()
if "prgs" in text:
return "prgs"
if "dadeschools" in text:
return "dadeschools"
return text.split(".")[0] if text else ""
def _timeline_comment_source(host: str, org: str, repo: str):
"""Build a fail-soft CTH-comment fetcher for one repo, or None when offline.
Returns a callable ``(kind, number) -> list[comment]``. Credentials or
network failures raise inside the callable so ``load_timeline`` degrades the
handoff source rather than the whole timeline. Offline test mode yields no
live source so the handoff section reports ``not run``.
"""
import os
from gitea_auth import api_fetch_page, get_auth_header, repo_api_url
offline = (os.environ.get("WEBUI_TEST_OFFLINE") or "").strip().lower() in {"1", "true", "yes"}
if offline:
return None
auth = get_auth_header(host)
if not auth:
return None
def _fetch(kind: str, number: int) -> list:
segment = "pulls" if kind == "pr" else "issues"
url = f"{repo_api_url(host, org, repo)}/{segment}/{int(number)}/comments"
comments: list = []
page = 1
while page <= 20:
raw, meta = api_fetch_page(url, auth, page=page, limit=50)
comments.extend(raw)
if bool(meta["is_final_page"]):
break
page += 1
return comments
return _fetch
async def api_v1_timeline(request: Request) -> JSONResponse:
"""Read-only workflow-event timeline (#637). Filter by issue/PR/session."""
from webui.queue_loader import _host_from_url # host normalisation helper
registry, error = _load_project_registry()
if error is not None:
return JSONResponse(error.to_dict(), status_code=500)
project = registry.projects[0] if registry.projects else None
org = request.query_params.get("org") or (project.gitea_owner if project else "")
repo = request.query_params.get("repo") or (project.repo_name if project else "")
host = _host_from_url(project.remote_host) if project else ""
remote = request.query_params.get("remote") or _derive_remote(host)
if not (remote and org and repo):
return JSONResponse(
{
"error": "timeline_scope_unresolved",
"detail": "no project in registry and no remote/org/repo query params provided",
},
status_code=400,
)
comment_source = _timeline_comment_source(host, org, repo) if (host and org and repo) else None
snapshot = load_timeline(
remote=remote,
org=org,
repo=repo,
issue=_query_int(request, "issue"),
pr=_query_int(request, "pr"),
session=(request.query_params.get("session") or None),
limit=_query_int(request, "limit"),
offset=_query_int(request, "offset"),
comment_source=comment_source,
)
# A filter no surviving source can carry is refused, not answered empty:
# a 200 with zero events would tell the operator no such activity exists.
status_code = 200 if snapshot.ok else 422
return JSONResponse(timeline_snapshot_to_dict(snapshot), status_code=status_code)
async def analytics(request: Request) -> HTMLResponse:
"""Read-only model usage, token cost, latency, and performance analytics HTML view (#651)."""
snapshot = load_analytics(
remote=request.query_params.get("remote"),
org=request.query_params.get("org"),
repo=request.query_params.get("repo"),
role=request.query_params.get("role"),
model=request.query_params.get("model"),
stage=request.query_params.get("stage"),
issue_number=_query_int(request, "issue"),
pr_number=_query_int(request, "pr"),
limit=_query_int(request, "limit") or 200,
)
return HTMLResponse(render_analytics_page(snapshot))
async def api_v1_analytics(request: Request) -> JSONResponse:
"""Read-only model usage, token cost, latency, and performance analytics API (#651)."""
snapshot = load_analytics(
remote=request.query_params.get("remote"),
org=request.query_params.get("org"),
repo=request.query_params.get("repo"),
role=request.query_params.get("role"),
model=request.query_params.get("model"),
stage=request.query_params.get("stage"),
issue_number=_query_int(request, "issue"),
pr_number=_query_int(request, "pr"),
limit=_query_int(request, "limit") or 500,
)
status_code = 200 if snapshot.ok else 500
return JSONResponse(analytics_snapshot_to_dict(snapshot), status_code=status_code)
async def api_v1_analytics_ingest(request: Request) -> JSONResponse:
"""Optional session instrumentation ingestion endpoint (#651).
Fail-closed write: every request is authorized through console_authz
(``record_analytics_usage``) before any control-plane DB mutation. Phase 1
keeps ``execution_enabled=False`` and denies unauthenticated callers, so
this route cannot be used as an unauthenticated write or XSS injection
vector (PR #876 F2).
"""
try:
body = await request.json()
except Exception:
body = {}
if not isinstance(body, dict):
body = {}
principal = resolve_principal(headers=dict(request.headers))
decision = authorize(
"record_analytics_usage", principal, for_execution=True
)
allowed = bool(decision.allowed and decision.execution_enabled)
console_audit.record_event(
action_id="record_analytics_usage",
result=(
console_audit.RESULT_ALLOWED
if allowed
else console_audit.RESULT_DENIED
),
decision=decision,
principal=principal,
target=_audit_target("record_analytics_usage", body),
request_id=_request_id(),
detail=decision.detail,
)
authorization = decision.to_dict()
if not allowed:
return JSONResponse(
{
"ok": False,
"error": "unauthorized",
"detail": (
"POST /api/v1/analytics/usage requires an authenticated "
"principal with record_analytics_usage execution enabled"
),
"authorization": authorization,
},
status_code=403,
)
usage_id = record_usage(
session_id=body.get("session_id"),
remote=body.get("remote", "dadeschools"),
org=body.get("org", ""),
repo=body.get("repo", ""),
project_id=body.get("project_id"),
role=body.get("role", "unknown"),
model=body.get("model", "unknown"),
issue_number=body.get("issue_number") or body.get("issue"),
pr_number=body.get("pr_number") or body.get("pr"),
stage=body.get("stage", "unknown"),
input_tokens=body.get("input_tokens"),
output_tokens=body.get("output_tokens"),
total_tokens=body.get("total_tokens"),
estimated_cost_usd=body.get("estimated_cost_usd"),
latency_ms=body.get("latency_ms"),
duration_ms=body.get("duration_ms"),
status=body.get("status", "success"),
metadata=body.get("metadata"),
)
return JSONResponse(
{"ok": True, "usage_id": usage_id, "authorization": authorization},
status_code=201,
)
async def method_not_allowed(request: Request, _exc: Exception) -> Response:
path = request.url.path
if path in _AUDIT_MUTATION_PATHS and request.method == "POST":
@@ -399,6 +732,8 @@ def create_app(*, bind_host: str | None = None) -> Starlette:
routes=[
Route("/", home, methods=["GET"]),
Route("/health", health, methods=["GET"]),
Route(SYSTEM_HEALTH_API_PATH, api_system_health, methods=["GET"]),
Route("/system-health", system_health, methods=["GET"]),
Route("/queue", queue, methods=["GET"]),
Route("/api/queue", api_queue, methods=["GET"]),
Route("/projects", projects, methods=["GET"]),
@@ -415,6 +750,11 @@ def create_app(*, bind_host: str | None = None) -> Starlette:
Route("/api/prompts", api_prompts, methods=["GET"]),
Route("/runtime", runtime, methods=["GET"]),
Route("/api/runtime", api_runtime, methods=["GET"]),
Route("/api/v1/timeline", api_v1_timeline, methods=["GET"]),
Route("/analytics", analytics, methods=["GET"]),
Route("/api/analytics", api_v1_analytics, methods=["GET"]),
Route("/api/v1/analytics", api_v1_analytics, methods=["GET"]),
Route("/api/v1/analytics/usage", api_v1_analytics_ingest, methods=["POST"]),
Route("/audit", audit, methods=["GET", "POST"]),
Route("/api/audit", api_audit, methods=["GET", "POST"]),
Route("/worktrees", worktrees, methods=["GET"]),
@@ -433,11 +773,21 @@ def create_app(*, bind_host: str | None = None) -> Starlette:
methods=["POST"],
),
Route("/api/leases", api_leases, methods=["GET"]),
Route("/api/v1/inventory", api_inventory, methods=["GET"]),
Route(
"/api/v1/inventory/{section}",
api_inventory_section,
methods=["GET"],
),
Route(
"/api/console/security-model",
api_console_security_model,
methods=["GET"],
),
*[
Route(path, phase_stub, methods=["GET"])
for path in STUB_PAGES
],
],
exception_handlers={405: method_not_allowed},
)
+41
View File
@@ -236,6 +236,47 @@ _ACTION_SPECS: tuple[ConsoleAction, ...] = (
phase=3,
summary="Remove a remote feature branch.",
),
# #651 analytics ingest: local control-plane write, not a Gitea mutation.
# Phase 2 gated write so Phase 1 (ACTIVE_PHASE=1) fails closed on execution.
ConsoleAction(
action_id="record_analytics_usage",
task_key="record_analytics_usage",
action_class=CLASS_WRITE,
minimum_role=OPERATOR,
requires_confirmation=True,
dual_control=False,
break_glass=False,
phase=2,
summary="Ingest a model-usage / latency analytics event into the control-plane DB.",
),
# #642: sanctioned daemon lifecycle. These exist so operators have an
# audited path off `pkill -f mcp_server.py` (#630). Restart drops every
# in-flight request on a namespace, so it carries the same dual-control and
# break-glass weight as a merge; reload drains first and is privileged but
# not destructive. Neither ever exposes a raw kill: execution is handed to
# a host supervisor by ``webui.sanctioned_restart``.
ConsoleAction(
action_id="system.reload_namespace",
task_key="reload_namespace",
action_class=CLASS_PRIVILEGED,
minimum_role=CONTROLLER,
requires_confirmation=True,
dual_control=False,
break_glass=False,
phase=2,
summary="Gracefully reload one MCP namespace via the host supervisor.",
),
ConsoleAction(
action_id="system.restart_namespace",
task_key="restart_namespace",
action_class=CLASS_DESTRUCTIVE,
minimum_role=ADMIN,
requires_confirmation=True,
dual_control=True,
break_glass=True,
phase=2,
summary="Restart one MCP namespace via the host supervisor.",
),
)
ACTIONS: dict[str, ConsoleAction] = {a.action_id: a for a in _ACTION_SPECS}
+13
View File
@@ -110,6 +110,8 @@ def _format_target(action_id: str, params: dict[str, Any]) -> str:
)
if action_id == "create_issue":
return f"issue {params.get('title', '?')!r}"
if action_id in {"system.restart_namespace", "system.reload_namespace"}:
return f"MCP namespace {params.get('namespace', '?')!r}"
return "unspecified"
@@ -165,6 +167,17 @@ def build_action_registry() -> ActionRegistry:
"gitea_create_issue_comment", "Post a PR review thread comment."),
("close_pr", "Close PR", "close_pr", "gitea_edit_pr",
"Close a pull request without merge."),
# #642: the sanctioned replacement for the forbidden manual daemon-kill
# recovery path (#630). The "tool" is a host supervisor hook, not an MCP
# call — the console never signals a process. Preview and gating live in
# ``webui.sanctioned_restart``; these stay disabled like every other
# registry entry.
("system.reload_namespace", "Reload MCP namespace", "reload_namespace",
"host.supervisor_reload",
"Gracefully reload one MCP namespace via the host supervisor."),
("system.restart_namespace", "Restart MCP namespace",
"restart_namespace", "host.supervisor_restart",
"Restart one MCP namespace via the host supervisor."),
)
actions = tuple(
GatedAction(
+952
View File
@@ -0,0 +1,952 @@
"""Unified session/lease/lock/worktree inventory for the web console (#636).
Leases (#433), worktrees (#432), and runtime (#430) each ship their own MVP
view, each with its own shape and its own idea of what "owned" means. A
traffic-control or recovery operator has to read all three and correlate them
by hand, which is exactly the step that goes wrong under collision pressure.
This module aggregates them into one versioned, read-only snapshot so the
console, and any worker asking "what is safe to do next", read the same
inventory from the same authority.
Field authority is explicit and never blended. Every section declares where its
rows came from:
* ``control_plane_db`` the #613 substrate: sessions, leases, assignments.
Authoritative for *exclusive ownership* (#600/#601).
* ``filesystem`` durable per-issue lock files (:mod:`issue_lock_store`) and
registered git worktrees. Authoritative for *what exists on this machine*.
* ``gitea`` remote issue/PR state, reached only through existing loaders.
Safety invariants:
* **Read-only.** The control-plane database is opened through a ``mode=ro``
URI. :class:`control_plane_db.ControlPlaneDB` creates directories and runs
migrations in its constructor, which an inventory read must never do, so this
module talks to sqlite directly rather than through that class.
* **Fail-soft, never fail-silent.** A subsystem that cannot be read degrades to
a section carrying ``status`` and ``reason``. It never raises, and it never
produces an empty list that reads like "nothing is there".
* **Never invent active ownership.** This is the invariant that matters most.
A degraded ownership source sets ``ownership_authority_complete`` false, and
while that flag is false no work item is reported unowned and no collision is
asserted. Absence of evidence is reported as absence of evidence.
* **Redaction at the boundary.** Absolute paths are collapsed against the home
directory, URLs lose userinfo and query strings, and no credential-shaped
value is emitted. No session token exists in these sources and none is read.
Phase 1 is read-only. Lease steal/release and worktree deletion are Phase 2+
and deliberately have no representation here, not even a disabled one.
"""
from __future__ import annotations
import os
import re
import sqlite3
import time
from dataclasses import dataclass, field
from datetime import datetime, timezone
from typing import Any, Callable
from urllib.parse import urlparse
import control_plane_db
import issue_lock_store
SCHEMA_VERSION = 1
API_VERSION = "v1"
#: Sections whose absence would make an ownership claim unprovable. If any of
#: these is not ``ok``, the snapshot refuses to describe anything as unowned.
OWNERSHIP_SECTIONS = ("sessions", "leases", "locks")
SECTION_NAMES = ("sessions", "leases", "locks", "worktrees", "namespaces")
STATUS_OK = "ok"
STATUS_DEGRADED = "degraded"
STATUS_UNAVAILABLE = "unavailable"
AUTHORITY_CONTROL_PLANE_DB = "control_plane_db"
AUTHORITY_FILESYSTEM = "filesystem"
AUTHORITY_GITEA = "gitea"
_CREDENTIAL_KEY_RE = re.compile(
r"(token|secret|password|passwd|api[_-]?key|authorization|bearer|credential)",
re.IGNORECASE,
)
_REDACTED = "[redacted]"
@dataclass(frozen=True)
class InventorySection:
"""One subsystem's contribution, with its authority and health."""
name: str
authority: str
status: str
items: tuple[dict[str, Any], ...] = ()
reason: str | None = None
scan_ms: float | None = None
@property
def ok(self) -> bool:
return self.status == STATUS_OK
def to_dict(self) -> dict[str, Any]:
return {
"name": self.name,
"authority": self.authority,
"status": self.status,
"count": len(self.items),
"reason": self.reason,
"scan_ms": self.scan_ms,
"items": [dict(item) for item in self.items],
}
@dataclass(frozen=True)
class CollisionSignal:
"""A detected conflict between two ownership records."""
kind: str
message: str
severity: str = "warning"
issue_number: int | None = None
branch: str | None = None
worktree_path: str | None = None
session_ids: tuple[str, ...] = ()
def to_dict(self) -> dict[str, Any]:
return {
"kind": self.kind,
"severity": self.severity,
"message": self.message,
"issue_number": self.issue_number,
"branch": self.branch,
"worktree_path": self.worktree_path,
"session_ids": list(self.session_ids),
}
@dataclass(frozen=True)
class InventorySnapshot:
"""Versioned aggregate of every inventory section."""
generated_at: str
sections: tuple[InventorySection, ...]
collisions: tuple[CollisionSignal, ...] = ()
correlations: tuple[dict[str, Any], ...] = ()
schema_version: int = SCHEMA_VERSION
api_version: str = API_VERSION
scan_ms: float | None = None
_section_index: dict[str, InventorySection] = field(
default_factory=dict, repr=False, compare=False
)
def section(self, name: str) -> InventorySection | None:
return self._section_index.get(name)
@property
def degraded_sections(self) -> tuple[str, ...]:
return tuple(s.name for s in self.sections if not s.ok)
@property
def ownership_authority_complete(self) -> bool:
"""True only when every ownership-bearing section read cleanly.
While this is false the snapshot must not describe any work item as
unowned: a lease the reader could not load is not an absent lease.
"""
for name in OWNERSHIP_SECTIONS:
section = self._section_index.get(name)
if section is None or not section.ok:
return False
return True
@property
def status(self) -> str:
if all(s.ok for s in self.sections):
return STATUS_OK
return STATUS_DEGRADED
# ── redaction ────────────────────────────────────────────────────────────────
def redact_path(path: str | None) -> str | None:
"""Collapse an absolute path against ``$HOME`` for browser display."""
if not path:
return path
text = str(path)
home = os.path.expanduser("~")
if home and home != "/" and text.startswith(home):
return "~" + text[len(home) :]
return text
def redact_url(value: str | None) -> str | None:
"""Strip userinfo and query string from a URL."""
if not value:
return value
text = str(value)
try:
parsed = urlparse(text)
except ValueError:
return _REDACTED
if not parsed.scheme or not parsed.netloc:
return text
netloc = parsed.hostname or ""
if parsed.port:
netloc = f"{netloc}:{parsed.port}"
rebuilt = f"{parsed.scheme}://{netloc}{parsed.path}"
return rebuilt.rstrip("/") or rebuilt
def scrub(value: Any, *, key: str | None = None) -> Any:
"""Recursively drop credential-shaped values and redact paths/URLs.
Never raises: an unexpected object degrades to its ``repr`` rather than
propagating out of a read-only view.
"""
if key and _CREDENTIAL_KEY_RE.search(key):
return _REDACTED
if isinstance(value, dict):
return {str(k): scrub(v, key=str(k)) for k, v in value.items()}
if isinstance(value, (list, tuple)):
return [scrub(v, key=key) for v in value]
if isinstance(value, str):
if value.startswith(("http://", "https://")):
return redact_url(value)
if value.startswith("/") or value.startswith("~"):
return redact_path(value)
return value
if isinstance(value, (int, float, bool)) or value is None:
return value
return repr(value)
# ── control-plane database (read-only) ───────────────────────────────────────
def _open_readonly(db_path: str) -> sqlite3.Connection:
"""Open the control-plane DB without creating or migrating anything."""
conn = sqlite3.connect(f"file:{db_path}?mode=ro", uri=True, timeout=5)
conn.row_factory = sqlite3.Row
return conn
def _table_names(conn: sqlite3.Connection) -> set[str]:
rows = conn.execute(
"SELECT name FROM sqlite_master WHERE type = 'table'"
).fetchall()
return {str(row[0]) for row in rows}
def _load_cp_db_sections(
*,
db_path: str | None = None,
limit: int = 200,
) -> tuple[InventorySection, InventorySection]:
"""Return the ``sessions`` and ``leases`` sections from the #613 DB."""
path = (db_path or control_plane_db.default_db_path()).strip()
def _both_unavailable(reason: str) -> tuple[InventorySection, InventorySection]:
return (
InventorySection(
name="sessions",
authority=AUTHORITY_CONTROL_PLANE_DB,
status=STATUS_UNAVAILABLE,
reason=reason,
),
InventorySection(
name="leases",
authority=AUTHORITY_CONTROL_PLANE_DB,
status=STATUS_UNAVAILABLE,
reason=reason,
),
)
if not path:
return _both_unavailable("control-plane database path is not configured")
if not os.path.exists(path):
return _both_unavailable(
f"control-plane database not present at {redact_path(path)}; "
"no session or lease authority available"
)
started = time.perf_counter()
try:
conn = _open_readonly(path)
except sqlite3.Error as exc:
return _both_unavailable(f"control-plane database could not be opened: {exc}")
try:
tables = _table_names(conn)
if "sessions" not in tables or "leases" not in tables:
missing = sorted({"sessions", "leases"} - tables)
return _both_unavailable(
"control-plane database is missing required tables: "
+ ", ".join(missing)
)
session_rows = [
dict(row)
for row in conn.execute(
"SELECT session_id, role, profile, namespace, pid, started_at,"
" last_heartbeat_at, status FROM sessions"
" ORDER BY last_heartbeat_at DESC LIMIT ?",
(max(1, int(limit)),),
).fetchall()
]
has_work_items = "work_items" in tables
if has_work_items:
lease_sql = (
"SELECT l.lease_id, l.session_id, l.role, l.phase, l.status,"
" l.expires_at, w.remote, w.org, w.repo, w.kind AS work_kind,"
" w.number AS work_number, w.state AS work_state,"
" s.pid AS session_pid, s.profile AS session_profile,"
" s.namespace AS session_namespace, s.status AS session_status"
" FROM leases l"
" JOIN work_items w ON w.work_item_id = l.work_item_id"
" LEFT JOIN sessions s ON s.session_id = l.session_id"
" ORDER BY l.expires_at DESC LIMIT ?"
)
else:
lease_sql = (
"SELECT l.lease_id, l.session_id, l.role, l.phase, l.status,"
" l.expires_at FROM leases l"
" ORDER BY l.expires_at DESC LIMIT ?"
)
lease_rows = [
dict(row)
for row in conn.execute(lease_sql, (max(1, int(limit)),)).fetchall()
]
except sqlite3.Error as exc:
return _both_unavailable(f"control-plane database read failed: {exc}")
finally:
conn.close()
elapsed = (time.perf_counter() - started) * 1000.0
now = datetime.now(timezone.utc)
sessions = tuple(
scrub(
{
"session_id": row.get("session_id"),
"role": row.get("role"),
"profile": row.get("profile"),
"namespace": row.get("namespace"),
"pid": row.get("pid"),
"pid_alive": issue_lock_store.is_process_alive(row.get("pid")),
"started_at": row.get("started_at"),
"last_heartbeat_at": row.get("last_heartbeat_at"),
"status": row.get("status"),
}
)
for row in session_rows
)
leases = tuple(
scrub(
{
"lease_id": row.get("lease_id"),
"session_id": row.get("session_id"),
"role": row.get("role"),
"phase": row.get("phase"),
"status": row.get("status"),
"expires_at": row.get("expires_at"),
"expired": _is_expired(row.get("expires_at"), now=now),
"remote": row.get("remote"),
"org": row.get("org"),
"repo": row.get("repo"),
"work_kind": row.get("work_kind"),
"work_number": row.get("work_number"),
"work_state": row.get("work_state"),
"session_pid": row.get("session_pid"),
"session_profile": row.get("session_profile"),
"session_namespace": row.get("session_namespace"),
"session_status": row.get("session_status"),
}
)
for row in lease_rows
)
degraded_reason = (
None
if has_work_items
else "work_items table absent; lease rows carry no work linkage"
)
lease_status = STATUS_OK if has_work_items else STATUS_DEGRADED
return (
InventorySection(
name="sessions",
authority=AUTHORITY_CONTROL_PLANE_DB,
status=STATUS_OK,
items=sessions,
scan_ms=round(elapsed, 3),
),
InventorySection(
name="leases",
authority=AUTHORITY_CONTROL_PLANE_DB,
status=lease_status,
items=leases,
reason=degraded_reason,
scan_ms=round(elapsed, 3),
),
)
def _is_expired(expires_at: str | None, *, now: datetime) -> bool | None:
if not expires_at:
return None
text = str(expires_at).strip().replace("Z", "+00:00")
try:
parsed = datetime.fromisoformat(text)
except ValueError:
return None
if parsed.tzinfo is None:
parsed = parsed.replace(tzinfo=timezone.utc)
return parsed <= now
# ── durable issue locks (filesystem) ─────────────────────────────────────────
def _load_locks_section(*, lock_dir: str | None = None) -> InventorySection:
started = time.perf_counter()
try:
paths = issue_lock_store.iter_lock_files(lock_dir)
except OSError as exc:
return InventorySection(
name="locks",
authority=AUTHORITY_FILESYSTEM,
status=STATUS_UNAVAILABLE,
reason=f"issue lock directory could not be listed: {exc}",
)
items: list[dict[str, Any]] = []
unreadable = 0
for path in paths:
try:
record = issue_lock_store.read_lock_file(path)
except (OSError, ValueError):
unreadable += 1
continue
if not record:
unreadable += 1
continue
try:
freshness = issue_lock_store.assess_lock_freshness(record)
except Exception: # noqa: BLE001 — a read-only view never raises
freshness = {"status": "unknown", "live": False, "stale": False}
claimant = record.get("claimant") or (
(record.get("work_lease") or {}).get("claimant") or {}
)
items.append(
scrub(
{
"issue_number": record.get("issue_number"),
"branch_name": record.get("branch_name"),
"remote": record.get("remote"),
"org": record.get("org"),
"repo": record.get("repo"),
"worktree_path": record.get("worktree_path"),
"pid": record.get("session_pid") or record.get("pid"),
"pid_alive": issue_lock_store.is_process_alive(
record.get("session_pid") or record.get("pid")
),
"claimant_username": (claimant or {}).get("username"),
"claimant_profile": (claimant or {}).get("profile"),
"lock_generation": record.get("lock_generation"),
"freshness_status": freshness.get("status"),
"live": bool(freshness.get("live")),
"stale": bool(freshness.get("stale")),
"freshness_reason": freshness.get("reason"),
"lock_path": record.get("lock_file_path") or path,
}
)
)
elapsed = (time.perf_counter() - started) * 1000.0
reason = (
f"{unreadable} lock file(s) were unreadable and are not represented"
if unreadable
else None
)
return InventorySection(
name="locks",
authority=AUTHORITY_FILESYSTEM,
status=STATUS_DEGRADED if unreadable else STATUS_OK,
items=tuple(items),
reason=reason,
scan_ms=round(elapsed, 3),
)
# ── worktrees (filesystem, via the #432 scanner) ─────────────────────────────
def _load_worktrees_section(
*, load_hygiene: Callable[[], Any] | None = None
) -> InventorySection:
started = time.perf_counter()
try:
loader = load_hygiene
if loader is None:
from webui.worktree_scanner import load_hygiene_snapshot
loader = load_hygiene_snapshot
snapshot = loader()
except Exception as exc: # noqa: BLE001 — fail soft, never fail the request
return InventorySection(
name="worktrees",
authority=AUTHORITY_FILESYSTEM,
status=STATUS_UNAVAILABLE,
reason=f"worktree scan failed: {exc}",
)
items = tuple(
scrub(
{
"rel_path": entry.rel_path,
"folder_name": entry.folder_name,
"classification": entry.classification,
"branch": entry.branch,
"head_sha": entry.head_sha,
"dirty_tracked": entry.dirty_tracked,
"dirty_untracked": entry.dirty_untracked,
"detached": entry.detached,
"registered_worktree": entry.registered_worktree,
"notes": entry.notes,
}
)
for entry in snapshot.entries
)
scan_error = getattr(snapshot, "scan_error", None)
elapsed = (time.perf_counter() - started) * 1000.0
return InventorySection(
name="worktrees",
authority=AUTHORITY_FILESYSTEM,
status=STATUS_DEGRADED if scan_error else STATUS_OK,
items=items,
reason=scan_error,
scan_ms=round(elapsed, 3),
)
# ── namespaces / capability summary ──────────────────────────────────────────
def _load_namespaces_section() -> InventorySection:
started = time.perf_counter()
try:
from gitea_auth import get_profile
profile = get_profile() or {}
except Exception as exc: # noqa: BLE001
return InventorySection(
name="namespaces",
authority=AUTHORITY_FILESYSTEM,
status=STATUS_UNAVAILABLE,
reason=f"active profile could not be resolved: {exc}",
)
allowed = list(profile.get("allowed_operations") or [])
forbidden = list(profile.get("forbidden_operations") or [])
profile_name = str(profile.get("profile_name") or "")
namespace = None
try:
import role_namespace_gate
namespace = role_namespace_gate.infer_mcp_namespace(profile_name)
except Exception: # noqa: BLE001 — namespace inference is advisory
namespace = None
item = scrub(
{
"profile_name": profile_name,
"role": profile.get("role"),
"mcp_namespace": namespace,
"allowed_operations": sorted(allowed),
"forbidden_operations": sorted(forbidden),
"capability_summary": {
"can_author": "gitea.pr.create" in allowed,
"can_review": "gitea.pr.approve" in allowed,
"can_merge": "gitea.pr.merge" in allowed,
"can_close_pr": "gitea.pr.close" in allowed,
},
"active": True,
}
)
elapsed = (time.perf_counter() - started) * 1000.0
return InventorySection(
name="namespaces",
authority=AUTHORITY_FILESYSTEM,
status=STATUS_OK,
items=(item,),
reason=(
"only the profile serving this web process is observable; other "
"namespaces are not enumerable from here"
),
scan_ms=round(elapsed, 3),
)
# ── correlation and collision detection ──────────────────────────────────────
def _issue_from_branch(branch: str | None) -> int | None:
match = re.search(r"issue-(\d+)", str(branch or ""), re.IGNORECASE)
return int(match.group(1)) if match else None
def correlate(
*,
leases: InventorySection,
locks: InventorySection,
worktrees: InventorySection,
sessions: InventorySection,
) -> tuple[tuple[dict[str, Any], ...], tuple[CollisionSignal, ...]]:
"""Join lease owner ↔ lock ↔ worktree ↔ namespace where evidence allows.
Correlation rows are emitted from whatever sections did load. Collision
signals are only emitted from sections that are ``ok``: a conflict inferred
from a partially-read source would be a false accusation.
"""
correlations: list[dict[str, Any]] = []
collisions: list[CollisionSignal] = []
worktree_by_branch: dict[str, dict[str, Any]] = {}
for entry in worktrees.items:
branch = (entry.get("branch") or "").strip()
if branch:
worktree_by_branch.setdefault(branch, entry)
session_by_id = {
str(s.get("session_id")): s for s in sessions.items if s.get("session_id")
}
# Lock-centred rows: a durable lock names an issue, a branch, and a worktree.
for lock in locks.items:
branch = (lock.get("branch_name") or "").strip()
worktree = worktree_by_branch.get(branch)
matching_leases = [
lease
for lease in leases.items
if lease.get("work_kind") == "issue"
and lease.get("work_number") == lock.get("issue_number")
]
correlations.append(
{
"issue_number": lock.get("issue_number"),
"branch": branch or None,
"lock_live": bool(lock.get("live")),
"lock_claimant": lock.get("claimant_profile"),
"lock_pid": lock.get("pid"),
"lock_pid_alive": lock.get("pid_alive"),
"worktree_rel_path": (worktree or {}).get("rel_path"),
"worktree_classification": (worktree or {}).get("classification"),
"worktree_registered": (worktree or {}).get("registered_worktree"),
"lease_ids": [
lease.get("lease_id")
for lease in matching_leases
if lease.get("lease_id")
],
"lease_sessions": [
lease.get("session_id")
for lease in matching_leases
if lease.get("session_id")
],
}
)
if locks.ok and worktrees.ok:
# A claim whose lease window is still open but has no registered
# worktree is an anomaly regardless of whether its pid is alive; a
# fully time-expired lease is on its way out and is not flagged.
if (
lock.get("freshness_status") != "expired"
and branch
and worktree is None
):
collisions.append(
CollisionSignal(
kind="lock-without-worktree",
severity="warning",
issue_number=lock.get("issue_number"),
branch=branch,
worktree_path=lock.get("worktree_path"),
message=(
f"Live lock on issue #{lock.get('issue_number')} names "
f"branch {branch!r} but no registered worktree carries "
"that branch (#404)"
),
)
)
if locks.ok:
# A lock whose recorded pid is gone is held by nobody: a clean #753
# dead-session recovery candidate. Subclassify by the lease window,
# because the two cases need different operator urgency. When the
# window is still open the lock would read as live to a naive
# timestamp check even though the owner is dead — the more dangerous
# case — so it is flagged distinctly from a fully time-expired lease.
if lock.get("pid_alive") is False and lock.get("stale"):
if lock.get("freshness_status") == "expired":
collisions.append(
CollisionSignal(
kind="stale-lock-dead-owner",
severity="warning",
issue_number=lock.get("issue_number"),
branch=branch or None,
message=(
f"Lock on issue #{lock.get('issue_number')} is stale "
f"and its recorded pid {lock.get('pid')} is not running "
"(#753 dead-session recovery candidate)"
),
)
)
else:
collisions.append(
CollisionSignal(
kind="live-lock-dead-owner",
severity="warning",
issue_number=lock.get("issue_number"),
branch=branch or None,
message=(
f"Lock on issue #{lock.get('issue_number')} has an "
"unexpired lease but its recorded pid "
f"{lock.get('pid')} is not running; it would read as "
"live to a timestamp check (#753 dead-session recovery "
"candidate)"
),
)
)
# The #635 trap: the lease has expired but the recorded pid is a
# still-running daemon, so neither dead-pid reclaim nor exact-owner
# renewal applies. This is the collision an operator must see.
elif (
lock.get("freshness_status") == "expired"
and lock.get("pid_alive") is True
):
collisions.append(
CollisionSignal(
kind="expired-lock-live-owner",
severity="error",
issue_number=lock.get("issue_number"),
branch=branch or None,
message=(
f"Lock on issue #{lock.get('issue_number')} has an expired "
f"lease but its recorded pid {lock.get('pid')} is still "
"running (daemon-pid deadlock; needs an operator decision, "
"#635/#760)"
),
)
)
# Two live locks on one branch, or two active leases on one work item.
if locks.ok:
by_branch: dict[str, list[dict[str, Any]]] = {}
for lock in locks.items:
if not lock.get("live"):
continue
branch = (lock.get("branch_name") or "").strip()
if branch:
by_branch.setdefault(branch, []).append(lock)
for branch, entries in sorted(by_branch.items()):
if len(entries) > 1:
collisions.append(
CollisionSignal(
kind="duplicate-live-lock",
severity="error",
branch=branch,
message=(
f"{len(entries)} live locks name branch {branch!r}: "
"issues "
+ ", ".join(
f"#{e.get('issue_number')}" for e in entries
)
),
)
)
if leases.ok:
by_work: dict[tuple[str, int], list[dict[str, Any]]] = {}
for lease in leases.items:
if str(lease.get("status") or "").lower() != "active":
continue
kind = str(lease.get("work_kind") or "").strip().lower()
number = lease.get("work_number")
if not kind or number is None:
continue
by_work.setdefault((kind, int(number)), []).append(lease)
for (kind, number), entries in sorted(by_work.items()):
sessions_held = {
str(e.get("session_id")) for e in entries if e.get("session_id")
}
if len(sessions_held) > 1:
collisions.append(
CollisionSignal(
kind="concurrent-active-lease",
severity="error",
issue_number=number if kind == "issue" else None,
session_ids=tuple(sorted(sessions_held)),
message=(
f"{len(sessions_held)} sessions hold an active lease on "
f"{kind} #{number}"
),
)
)
for entry in entries:
if entry.get("expired") is True:
collisions.append(
CollisionSignal(
kind="active-lease-past-expiry",
severity="warning",
issue_number=number if kind == "issue" else None,
session_ids=(
(str(entry.get("session_id")),)
if entry.get("session_id")
else ()
),
message=(
f"Lease {entry.get('lease_id')} on {kind} #{number} "
"is still marked active past its expiry"
),
)
)
# A lease whose owning session is gone is an orphan, not free work.
if leases.ok and sessions.ok:
for lease in leases.items:
if str(lease.get("status") or "").lower() != "active":
continue
session_id = str(lease.get("session_id") or "")
if session_id and session_id not in session_by_id:
collisions.append(
CollisionSignal(
kind="orphan-lease",
severity="error",
session_ids=(session_id,),
message=(
f"Active lease {lease.get('lease_id')} names session "
f"{session_id}, which has no session record"
),
)
)
return tuple(correlations), tuple(collisions)
# ── snapshot assembly ────────────────────────────────────────────────────────
def load_inventory_snapshot(
*,
db_path: str | None = None,
lock_dir: str | None = None,
load_hygiene: Callable[[], Any] | None = None,
include: tuple[str, ...] | None = None,
) -> InventorySnapshot:
"""Build the unified inventory snapshot.
Every section is loaded independently and fails soft. *include* restricts
the sections that are scanned; omitted sections are simply absent rather
than reported as empty, so a resource-split request cannot be mistaken for
a whole-inventory answer.
"""
started = time.perf_counter()
wanted = tuple(include) if include else SECTION_NAMES
sections: list[InventorySection] = []
sessions_section: InventorySection | None = None
leases_section: InventorySection | None = None
if "sessions" in wanted or "leases" in wanted:
sessions_section, leases_section = _load_cp_db_sections(db_path=db_path)
if "sessions" in wanted:
sections.append(sessions_section)
if "leases" in wanted:
sections.append(leases_section)
locks_section = (
_load_locks_section(lock_dir=lock_dir)
if "locks" in wanted
else _empty_section("locks", AUTHORITY_FILESYSTEM)
)
if "locks" in wanted:
sections.append(locks_section)
worktrees_section = (
_load_worktrees_section(load_hygiene=load_hygiene)
if "worktrees" in wanted
else _empty_section("worktrees", AUTHORITY_FILESYSTEM)
)
if "worktrees" in wanted:
sections.append(worktrees_section)
if "namespaces" in wanted:
sections.append(_load_namespaces_section())
correlations, collisions = correlate(
leases=leases_section or _empty_section("leases", AUTHORITY_CONTROL_PLANE_DB),
locks=locks_section,
worktrees=worktrees_section,
sessions=sessions_section
or _empty_section("sessions", AUTHORITY_CONTROL_PLANE_DB),
)
elapsed = (time.perf_counter() - started) * 1000.0
index = {section.name: section for section in sections}
return InventorySnapshot(
generated_at=datetime.now(timezone.utc).isoformat(),
sections=tuple(sections),
collisions=collisions,
correlations=correlations,
scan_ms=round(elapsed, 3),
_section_index=index,
)
def _empty_section(name: str, authority: str) -> InventorySection:
"""A section that was not requested — never a claim that it is empty."""
return InventorySection(
name=name,
authority=authority,
status=STATUS_UNAVAILABLE,
reason="section not requested in this scan",
)
def snapshot_to_dict(snapshot: InventorySnapshot) -> dict[str, Any]:
"""Serialize the snapshot for the versioned API."""
return {
"api_version": snapshot.api_version,
"schema_version": snapshot.schema_version,
"generated_at": snapshot.generated_at,
"status": snapshot.status,
"scan_ms": snapshot.scan_ms,
"ownership_authority_complete": snapshot.ownership_authority_complete,
"ownership_note": (
"Every ownership source read cleanly; an item absent from leases "
"and locks is genuinely unclaimed."
if snapshot.ownership_authority_complete
else "One or more ownership sources are degraded; nothing in this "
"snapshot may be treated as unowned. Collisions are reported only "
"from sections that read cleanly."
),
"degraded_sections": list(snapshot.degraded_sections),
"field_authority": {
"sessions": AUTHORITY_CONTROL_PLANE_DB,
"leases": AUTHORITY_CONTROL_PLANE_DB,
"locks": AUTHORITY_FILESYSTEM,
"worktrees": AUTHORITY_FILESYSTEM,
"namespaces": AUTHORITY_FILESYSTEM,
},
"sections": {section.name: section.to_dict() for section in snapshot.sections},
"correlations": [dict(row) for row in snapshot.correlations],
"collisions": [signal.to_dict() for signal in snapshot.collisions],
}
+114 -17
View File
@@ -2,28 +2,66 @@
from __future__ import annotations
NAV_ITEMS = (
("/", "Home"),
("/queue", "Queue"),
("/projects", "Projects"),
("/prompts", "Prompts"),
("/runtime", "Runtime"),
("/audit", "Audit"),
("/worktrees", "Worktrees"),
("/leases", "Leases"),
("/actions", "Actions"),
)
import os
from webui.nav import NAV_GROUPS
MVP_NOTICE = (
"Read-only MVP — Gitea, MCP tools, and canonical workflows remain the "
"source of truth. No mutation endpoints."
)
# Canonical docs entry point surfaced from the shell header (#638).
DOCS_URL = (
"https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/src/branch/"
"master/docs/webui-local-dev.md"
)
_LOCAL_HOSTS = frozenset({"", "127.0.0.1", "localhost", "::1"})
def environment_label() -> str:
"""Classify the serving environment as ``local`` or ``remote`` (#638).
Derived from the same ``WEBUI_HOST`` default the app binds to; loopback
hosts are ``local``, anything else is ``remote``. Read-only signal only.
"""
host = (os.environ.get("WEBUI_HOST", "127.0.0.1") or "").strip().lower()
return "local" if host in _LOCAL_HOSTS else "remote"
def _render_nav() -> str:
groups_html = []
for group in NAV_GROUPS:
links = "".join(
f'<a href="{item.href}"'
+ (' class="nav-stub"' if item.status == "stub" else "")
+ f'>{item.label}</a>'
for item in group.items
)
groups_html.append(
'<div class="nav-group">'
f'<span class="nav-group-label">{group.label}</span>'
f'<span class="nav-group-links">{links}</span>'
"</div>"
)
return "".join(groups_html)
def _render_badges() -> str:
env = environment_label()
return (
'<div class="header-badges">'
f'<span class="badge env-badge env-{env}">env: {env}</span>'
'<span class="badge mode-badge">mode: read-only</span>'
f'<a class="badge docs-link" href="{DOCS_URL}">Docs</a>'
"</div>"
)
def render_page(*, title: str, body_html: str, extra_head: str = "") -> str:
nav_links = "".join(
f'<a href="{href}">{label}</a>' for href, label in NAV_ITEMS
)
nav_links = _render_nav()
header_badges = _render_badges()
return f"""<!DOCTYPE html>
<html lang="en">
<head>
@@ -53,21 +91,58 @@ def render_page(*, title: str, body_html: str, extra_head: str = "") -> str:
padding: 0.75rem 1.25rem;
}}
header h1 {{
margin: 0 0 0.5rem;
margin: 0;
font-size: 1.1rem;
font-weight: 600;
}}
.header-top {{
display: flex;
flex-wrap: wrap;
align-items: center;
justify-content: space-between;
gap: 0.5rem 1rem;
margin-bottom: 0.6rem;
}}
.header-badges {{ display: inline-flex; flex-wrap: wrap; gap: 0.4rem; }}
.env-badge.env-local {{ color: #8fd19e; border-color: #3d6b4a; }}
.env-badge.env-remote {{ color: #e0c27a; border-color: #6b5730; }}
.mode-badge {{ color: #9ec8f0; border-color: #3d5f7a; }}
a.docs-link {{
color: var(--accent);
border-color: var(--accent);
text-decoration: none;
text-transform: none;
}}
a.docs-link:hover {{ filter: brightness(1.12); }}
nav {{
display: flex;
flex-wrap: wrap;
gap: 0.75rem 1rem;
gap: 0.5rem 1.25rem;
}}
.nav-group {{
display: flex;
flex-direction: column;
gap: 0.15rem;
}}
.nav-group-label {{
font-size: 0.68rem;
text-transform: uppercase;
letter-spacing: 0.04em;
color: var(--muted);
}}
.nav-group-links {{ display: inline-flex; flex-wrap: wrap; gap: 0.6rem; }}
nav a {{
color: var(--accent);
text-decoration: none;
font-size: 0.9rem;
}}
nav a:hover {{ text-decoration: underline; }}
nav a.nav-stub {{ color: var(--muted); }}
nav a.nav-stub::after {{
content: " ·stub";
font-size: 0.7rem;
color: var(--muted);
}}
main {{
max-width: 52rem;
margin: 0 auto;
@@ -161,12 +236,34 @@ def render_page(*, title: str, body_html: str, extra_head: str = "") -> str:
.badge-in-review {{ color: #9ec8f0; border-color: #3d5f7a; }}
.badge-duplicate {{ color: #e0c27a; border-color: #6b5730; }}
.badge-stale {{ color: #c9b8e8; border-color: #5a4a78; }}
.badge-health-ok {{ color: #8fd19e; border-color: #3d6b4a; }}
.badge-health-degraded {{ color: #e0c27a; border-color: #6b5730; }}
.badge-health-down {{ color: #f0a8a8; border-color: #7a3b3b; }}
.badge-health-skipped {{ color: var(--muted); }}
.badge-health-unproven {{ color: #c9b8e8; border-color: #5a4a78; }}
.health-card {{
margin: 1.25rem 0;
padding: 0.85rem 1rem 1rem;
border: 1px solid var(--border);
border-radius: 8px;
background: var(--surface);
}}
.health-card h3 {{ margin: 0 0 0.5rem; font-size: 1.05rem; }}
.health-card h4 {{ margin: 1rem 0 0.35rem; font-size: 0.92rem; color: var(--muted); }}
.health-headline {{ color: var(--text); font-size: 1rem; margin: 0 0 0.5rem; }}
.health-degraded {{ border-left-color: #e0c27a; }}
.health-stale {{ border-left-color: #f0a8a8; }}
ul.reasons {{ margin: 0.35rem 0; padding-left: 1.15rem; color: var(--muted); font-size: 0.9rem; }}
ul.reasons li {{ margin-bottom: 0.3rem; }}
</style>
{extra_head}
</head>
<body>
<header>
<h1>MCP Control Plane</h1>
<div class="header-top">
<h1>MCP Control Plane</h1>
{header_badges}
</div>
<nav>{nav_links}</nav>
</header>
<main>
+113
View File
@@ -0,0 +1,113 @@
"""Navigation IA for the Phase 1 operator console shell (#638).
Single source of truth for the console navigation so ``webui/layout.py`` and
the ``webui/app.py`` route table stay aligned with epic #631. Read-only: every
destination is a GET view or a Phase 1 placeholder. No mutation links.
Nav groups follow the #631 Phase 1 information architecture: Health, Traffic,
Runtime/Sessions, Projects, Inventory, Timeline, Policy (placeholder), and
Insights (placeholder). Later-phase surfaces are declared as ``stub`` items and
backed by ``STUB_PAGES`` so their nav links resolve to a graceful placeholder
instead of a 404.
"""
from __future__ import annotations
from dataclasses import dataclass
@dataclass(frozen=True)
class NavItem:
"""A single navigation destination.
``status`` is ``"live"`` for implemented views and ``"stub"`` for Phase 1
placeholders whose backing view lands in a later child issue.
"""
href: str
label: str
status: str = "live"
@dataclass(frozen=True)
class NavGroup:
label: str
items: tuple[NavItem, ...]
NAV_GROUPS: tuple[NavGroup, ...] = (
NavGroup("Health", (
NavItem("/health", "Liveness"),
NavItem("/system-health", "System health"),
)),
NavGroup("Traffic", (
NavItem("/queue", "Queue"),
NavItem("/leases", "Leases"),
NavItem("/actions", "Actions"),
)),
NavGroup("Runtime/Sessions", (
NavItem("/runtime", "Runtime health"),
NavItem("/sessions", "Sessions", "stub"),
)),
NavGroup("Projects", (
NavItem("/projects", "Projects"),
)),
NavGroup("Inventory", (
NavItem("/inventory", "Inventory", "stub"),
NavItem("/worktrees", "Worktrees"),
)),
NavGroup("Timeline", (
NavItem("/timeline", "Timeline", "stub"),
)),
NavGroup("Policy", (
NavItem("/policy", "Policy", "stub"),
NavItem("/prompts", "Prompts"),
)),
NavGroup("Insights", (
NavItem("/insights", "Insights", "stub"),
NavItem("/analytics", "Analytics"),
NavItem("/audit", "Audit"),
)),
)
# Phase 1 placeholder destinations whose backing views land in later child
# issues of epic #631. Each maps a path to (title, description). Routes are
# registered so nav links resolve to a graceful, read-only stub page.
STUB_PAGES: dict[str, tuple[str, str]] = {
"/sessions": (
"Sessions",
"Active session, capability, and role inventory. Backed by the unified "
"inventory API (#636) once it lands.",
),
"/inventory": (
"Inventory",
"Unified sessions, leases, locks, namespaces, and worktree inventory. "
"Backed by the Phase 1 inventory API (#636).",
),
"/timeline": (
"Timeline",
"Workflow event timeline across issues and PRs. A later Phase 1 surface.",
),
"/policy": (
"Policy",
"Capability and role policy surface. Placeholder until a later phase.",
),
"/insights": (
"Insights",
"Aggregate operational insights and trends. Placeholder until a later "
"phase.",
),
}
def iter_nav_items():
"""Yield every ``NavItem`` across all groups in declared order."""
for group in NAV_GROUPS:
for item in group.items:
yield item
def nav_hrefs() -> tuple[str, ...]:
"""Return every navigation href in declared order."""
return tuple(item.href for item in iter_nav_items())
+613
View File
@@ -0,0 +1,613 @@
"""Sanctioned MCP restart and graceful reload controls (#642, Phase 2).
Sessions have historically recovered MCP connectivity by killing the host
daemon (``pkill -f mcp_server.py``, #630). That path stays forbidden: it kills
every namespace on the host, contaminates the surviving session, and leaves no
audit trail. This module is the sanctioned replacement.
A restart is modelled as a *gated action*, never as a command:
1. **Capability** the console action resolves through ``console_authz``
against ``task_capability_map``, so the console cannot invent an authority
the MCP layer does not already define.
2. **Preview** :func:`build_restart_preview` renders a mutation ledger and
the exact confirmation phrase. It never returns a shell command.
3. **Confirmation** the operator echoes a phrase naming the exact namespace
and mode. A phrase for one namespace never authorizes another.
4. **Operator authorization** host daemon maintenance is authorized out of
band through the environment (#630, and #710 finding F1: a worker session
cannot set an env var for an already-running daemon, so this cannot be
self-asserted the way a tool argument could).
5. **Execution** :func:`execute_restart` never spawns a process. Once every
gate passes it hands the request to the configured host-managed restart
hook; with no hook configured it fails closed.
6. **Health recheck** :func:`verify_post_restart_health` requires live
client-namespace probe evidence before any post-restart clean claim.
Manual ``pkill`` remains forbidden and is classified as contamination by
:func:`classify_restart_command`, which blocks clean claims (#630 AC3).
This module performs no I/O beyond reading its own environment configuration,
imports no MCP client, and holds no credential.
"""
from __future__ import annotations
import os
from dataclasses import asdict, dataclass
from typing import Any
import mcp_namespace_health
import runtime_recovery_guard
from webui import console_audit, console_authz
# --- Operations -------------------------------------------------------------
MODE_RESTART = "restart"
MODE_RELOAD = "reload"
MODES: tuple[str, ...] = (MODE_RESTART, MODE_RELOAD)
ACTION_RESTART_NAMESPACE = "system.restart_namespace"
ACTION_RELOAD_NAMESPACE = "system.reload_namespace"
ACTION_FOR_MODE: dict[str, str] = {
MODE_RESTART: ACTION_RESTART_NAMESPACE,
MODE_RELOAD: ACTION_RELOAD_NAMESPACE,
}
# Namespaces the console may target. An unlisted name fails closed rather than
# being passed through to a host hook.
KNOWN_NAMESPACES: tuple[str, ...] = tuple(
sorted(
set(mcp_namespace_health.DEFAULT_NAMESPACES)
| {"gitea-author", "gitea-reviewer", "gitea-merger",
"gitea-reconciler", "gitea-controller"}
)
)
# Scope tokens that would mean "everything at once". Explicit non-goal: the
# console never offers a fleet-wide restart, because that is the blast radius
# `pkill -f mcp_server.py` already had.
_FLEET_TOKENS = frozenset({"*", "all", "fleet", "any", ""})
# --- Environment configuration ----------------------------------------------
# Read server-side only; the value is an opaque host hook reference (e.g. a
# launchd label), never a command line, and is never rendered to a client.
RESTART_HOOK_ENV = "GITEA_SANCTIONED_RESTART_HOOK"
# --- Reason codes -----------------------------------------------------------
DENY_UNKNOWN_MODE = "unknown_mode"
DENY_UNKNOWN_NAMESPACE = "unknown_namespace"
DENY_FLEET_SCOPE = "fleet_scope_not_permitted"
DENY_UNAUTHORIZED = "unauthorized"
DENY_CONFIRMATION_MISSING = "confirmation_required"
DENY_CONFIRMATION_MISMATCH = "confirmation_mismatch"
DENY_OPERATOR_AUTHORIZATION = "operator_authorization_missing"
DENY_HOOK_NOT_CONFIGURED = "restart_hook_not_configured"
DENY_CONTAMINATED_RUNTIME = "contaminated_runtime"
ALLOW_HOST_ACTION_REQUIRED = "host_action_required"
# Post-restart verification outcomes.
HEALTH_CLEAN = "clean"
HEALTH_UNPROVEN = "unproven"
HEALTH_UNHEALTHY = "unhealthy"
def _clean(value: Any) -> str:
return str(value or "").strip()
# --- Mutation ledger --------------------------------------------------------
@dataclass(frozen=True)
class RestartLedgerEntry:
"""One planned step, shown before anything is asked of the host."""
sequence: int
step: str
summary: str
executes_process_kill: bool = False
def _mutation_ledger(namespace: str, mode: str) -> tuple[RestartLedgerEntry, ...]:
if mode == MODE_RELOAD:
middle = RestartLedgerEntry(
sequence=2,
step="host_graceful_reload",
summary=(
f"Ask the configured host supervisor to reload {namespace} "
"in place, draining in-flight requests. The console does not "
"signal the process itself."
),
)
else:
middle = RestartLedgerEntry(
sequence=2,
step="host_restart_hook",
summary=(
f"Ask the configured host supervisor to restart {namespace}. "
"The console never sends a signal and never runs a kill."
),
)
return (
RestartLedgerEntry(
sequence=1,
step="quiesce",
summary=(
f"Stop admitting new gated mutations for {namespace} and "
"record the intent before anything restarts."
),
),
middle,
RestartLedgerEntry(
sequence=3,
step="health_recheck",
summary=(
f"Re-probe {namespace} through the live client namespace and "
"prove the required tool is callable again."
),
),
RestartLedgerEntry(
sequence=4,
step="audit",
summary=(
"Append actor, target namespace, mode, and result to the "
"console audit log."
),
),
)
# --- Confirmation -----------------------------------------------------------
def confirmation_phrase(namespace: str, mode: str) -> str:
"""Exact phrase an operator must echo, naming the namespace and mode.
Binding the namespace into the phrase is the point: a confirmation typed
for ``gitea-author`` cannot be replayed against ``gitea-merger``.
"""
return f"{_clean(mode)} {_clean(namespace)}"
def confirmation_matches(
namespace: str, mode: str, confirmation: str | None
) -> bool:
"""Compare *confirmation* to the required phrase (exact, whitespace-trimmed)."""
return _clean(confirmation) == confirmation_phrase(namespace, mode)
# --- Scope validation -------------------------------------------------------
def _validate_scope(namespace: str, mode: str) -> tuple[str, str] | None:
"""Return ``(reason_code, detail)`` when the scope is refused."""
ns = _clean(namespace)
md = _clean(mode)
if md not in MODES:
return (
DENY_UNKNOWN_MODE,
f"Mode {md!r} is not one of {', '.join(MODES)}.",
)
if ns.lower() in _FLEET_TOKENS:
return (
DENY_FLEET_SCOPE,
(
"Fleet-wide restart is an explicit non-goal: it reproduces the "
"blast radius of `pkill -f mcp_server.py` (#630). Restart one "
"namespace at a time."
),
)
if ns not in KNOWN_NAMESPACES:
return (
DENY_UNKNOWN_NAMESPACE,
f"Namespace {ns!r} is not a known MCP namespace.",
)
return None
# --- Host hook --------------------------------------------------------------
def restart_hook(env: dict[str, str] | None = None) -> dict[str, Any]:
"""Report the configured host-managed restart hook.
The hook is a reference the *host* resolves (a supervisor label), not a
command this process runs. ``configured=False`` fails restart closed.
"""
source = env if env is not None else os.environ
reference = _clean(source.get(RESTART_HOOK_ENV))
return {
"configured": bool(reference),
"reference": reference or None,
"source": RESTART_HOOK_ENV if reference else None,
"self_assertable": False,
"console_executes_process": False,
}
# --- Preview ----------------------------------------------------------------
def build_restart_preview(
namespace: str,
mode: str = MODE_RESTART,
*,
principal: console_authz.Principal | None = None,
env: dict[str, str] | None = None,
) -> dict[str, Any]:
"""Render the dry-run preview for a restart/reload request.
Read-only: no authorization is granted, no host is contacted, and the
result never contains a shell command.
"""
ns = _clean(namespace)
md = _clean(mode)
action_id = ACTION_FOR_MODE.get(md, ACTION_RESTART_NAMESPACE)
action = console_authz.get_action(action_id)
decision = console_authz.authorize(action_id, principal)
scope_error = _validate_scope(ns, md)
hook = restart_hook(env)
operator = runtime_recovery_guard.operator_authorization(env)
return {
"action_id": action_id,
"namespace": ns,
"mode": md,
"scope_valid": scope_error is None,
"scope_reason_code": scope_error[0] if scope_error else None,
"scope_detail": scope_error[1] if scope_error else None,
"required_role": action.minimum_role if action else None,
"required_permission": action.mcp_permission if action else None,
"action_class": action.action_class if action else None,
"dual_control": action.dual_control if action else True,
"break_glass": action.break_glass if action else True,
"requires_confirmation": True,
"confirmation_phrase": confirmation_phrase(ns, md),
"mutation_ledger": [asdict(entry) for entry in _mutation_ledger(ns, md)],
"authorization": decision.to_dict(),
"operator_authorization": operator,
"restart_hook": hook,
"execution_enabled": False,
"raw_process_kill_exposed": False,
"known_namespaces": list(KNOWN_NAMESPACES),
"post_restart_verification_required": True,
}
# --- Gate -------------------------------------------------------------------
def assess_restart_request(
namespace: str,
mode: str = MODE_RESTART,
*,
principal: console_authz.Principal | None = None,
confirmation: str | None = None,
contamination_marker: dict[str, Any] | None = None,
env: dict[str, str] | None = None,
) -> dict[str, Any]:
"""Decide whether a restart request may proceed to the host hook.
Every gate must pass. The first failure wins and is reported with a stable
reason code; a pass never means "restarted", only "may be handed to the
configured host hook".
"""
ns = _clean(namespace)
md = _clean(mode)
action_id = ACTION_FOR_MODE.get(md, ACTION_RESTART_NAMESPACE)
preview = build_restart_preview(ns, md, principal=principal, env=env)
def refuse(reason_code: str, detail: str) -> dict[str, Any]:
return {
"allowed": False,
"gates_passed": False,
"reason_code": reason_code,
"detail": detail,
"action_id": action_id,
"namespace": ns,
"mode": md,
"preview": preview,
"execution_enabled": False,
}
scope_error = _validate_scope(ns, md)
if scope_error is not None:
return refuse(*scope_error)
# Authority is checked as an authorization decision, not an execution
# grant. ``for_execution=True`` asks "may the console perform this write?",
# and the answer here is permanently no: step 2 of the ledger is a request
# to the host supervisor, so the console's Phase 2 execution gate is not
# the relevant gate. Every branch below keeps ``execution_enabled`` False
# and :func:`execute_restart` never touches a process.
decision = console_authz.authorize(action_id, principal)
if not decision.allowed:
return refuse(DENY_UNAUTHORIZED, decision.detail)
if not _clean(confirmation):
return refuse(
DENY_CONFIRMATION_MISSING,
(
"Type the confirmation phrase "
f"{preview['confirmation_phrase']!r} to proceed."
),
)
if not confirmation_matches(ns, md, confirmation):
return refuse(
DENY_CONFIRMATION_MISMATCH,
(
"Confirmation does not name this namespace and mode; expected "
f"{preview['confirmation_phrase']!r}."
),
)
operator = preview["operator_authorization"]
if not operator["authorized"]:
return refuse(
DENY_OPERATOR_AUTHORIZATION,
(
"Host daemon maintenance requires out-of-band operator "
"authorization via "
f"{runtime_recovery_guard.OPERATOR_AUTHORIZATION_ENV}."
),
)
# #630's task-scoped gate deliberately lets a contaminated worker keep
# commenting and handing off. Restart is stricter and unconditional: a
# runtime already contaminated by a manual kill must be reconciled before
# it is restarted again, or the restart just launders the contamination.
if contamination_marker and not contamination_marker.get(
"cleared_by_reconciler"
):
return refuse(
DENY_CONTAMINATED_RUNTIME,
(
"A live contamination marker is present; clear it through the "
"reconciler path before restarting."
),
)
hook = preview["restart_hook"]
if not hook["configured"]:
return refuse(
DENY_HOOK_NOT_CONFIGURED,
(
"No host-managed restart hook is configured "
f"({RESTART_HOOK_ENV}). The console will not fall back to a "
"process kill."
),
)
return {
"allowed": True,
"gates_passed": True,
"reason_code": ALLOW_HOST_ACTION_REQUIRED,
"detail": (
"Every gate passed. The restart must be performed by the "
"configured host supervisor; the console does not signal the "
"process."
),
"action_id": action_id,
"namespace": ns,
"mode": md,
"preview": preview,
"execution_enabled": False,
"console_executes": False,
"console_active_phase": console_authz.ACTIVE_PHASE,
}
# --- Execution --------------------------------------------------------------
def execute_restart(
namespace: str,
mode: str = MODE_RESTART,
*,
principal: console_authz.Principal | None = None,
confirmation: str | None = None,
contamination_marker: dict[str, Any] | None = None,
env: dict[str, str] | None = None,
request_id: str | None = None,
session_id: str | None = None,
) -> dict[str, Any]:
"""Run every gate, audit the outcome, and hand off to the host.
This function never spawns a process, never sends a signal, and never
builds a command line. ``success`` is False in both directions: a refused
request is refused, and an authorized request still requires the host
supervisor to act.
"""
assessment = assess_restart_request(
namespace,
mode,
principal=principal,
confirmation=confirmation,
contamination_marker=contamination_marker,
env=env,
)
action_id = assessment["action_id"]
allowed = assessment["allowed"]
audit = console_audit.record_event(
action_id=action_id,
result=(
console_audit.RESULT_ALLOWED if allowed
else console_audit.RESULT_DENIED
),
principal=principal,
target={"namespace": assessment["namespace"], "mode": assessment["mode"]},
reason_code=assessment["reason_code"],
detail=assessment["detail"],
request_id=request_id,
session_id=session_id,
metadata={
"gates_passed": assessment["gates_passed"],
"process_kill_executed": False,
"post_restart_verification_required": True,
},
)
return {
"success": False,
"outcome": (
ALLOW_HOST_ACTION_REQUIRED if allowed else assessment["reason_code"]
),
"allowed": allowed,
"detail": assessment["detail"],
"namespace": assessment["namespace"],
"mode": assessment["mode"],
"action_id": action_id,
"process_kill_executed": False,
"host_hook": assessment["preview"]["restart_hook"],
"next_action": (
"Have the host supervisor perform the restart, then call "
"verify_post_restart_health with live client-namespace evidence "
"before claiming a clean session."
if allowed
else assessment["detail"]
),
"assessment": assessment,
"audit": audit,
}
# --- Contamination classification -------------------------------------------
def classify_restart_command(
command: str | None,
*,
mcp_pids: list[Any] | tuple[Any, ...] | None = None,
session_id: str | None = None,
remote: str | None = None,
role: str | None = None,
) -> dict[str, Any]:
"""Classify an operator-proposed recovery command (#630 AC2).
A manual ``pkill``/``kill`` of the MCP daemon is contamination, not a
restart. When contaminating, a durable marker is returned so downstream
gated mutations and clean claims fail closed.
"""
classification = runtime_recovery_guard.classify_recovery_command(
command, mcp_pids=mcp_pids
)
contaminating = bool(classification.get("contamination"))
marker = None
if contaminating:
marker = runtime_recovery_guard.build_contamination_record(
reason_class=(
classification.get("reason_class")
or runtime_recovery_guard.REASON_MANUAL_DAEMON_KILL
),
command_redacted=classification.get("redacted_command"),
session_id=session_id,
remote=remote,
role=role,
detail=(
"Manual daemon kill is forbidden; use the sanctioned "
f"{ACTION_RESTART_NAMESPACE} gated action instead."
),
)
return {
"contamination": contaminating,
"sanctioned": not contaminating and not classification.get("process_kill"),
"clean_claim_allowed": not contaminating,
"reason_class": classification.get("reason_class"),
"redacted_command": classification.get("redacted_command"),
"classification": classification,
"contamination_marker": marker,
"sanctioned_alternative": ACTION_RESTART_NAMESPACE,
}
# --- Post-restart health verification ---------------------------------------
def verify_post_restart_health(
namespace: str,
*,
probe_result: dict[str, Any] | None = None,
probe_source: str | None = None,
registered_tools: list[str] | tuple[str, ...] | None = None,
required_tool: str | None = None,
profile: str | None = None,
) -> dict[str, Any]:
"""Require live proof a namespace is callable before any clean claim (AC3).
Static registration is not proof and neither is an offline subprocess
probe: only ``probe_source=client_namespace`` evidence can clear a
post-restart session for mutations.
"""
ns = _clean(namespace)
health = mcp_namespace_health.classify_namespace_probe(
ns,
required_tool=required_tool,
registered_tools=registered_tools,
probe_result=probe_result,
profile=profile,
probe_source=probe_source,
)
healthy = bool(health.get("healthy"))
proven = bool(health.get("ide_namespace_proven"))
if healthy and proven:
status = HEALTH_CLEAN
elif healthy:
status = HEALTH_UNPROVEN
else:
status = HEALTH_UNHEALTHY
reasons = list(health.get("reasons") or [])
if status == HEALTH_UNPROVEN:
reasons.append(
"Namespace reported healthy without live client-namespace "
"evidence; a post-restart clean claim requires "
f"probe_source={mcp_namespace_health.PROBE_SOURCE_CLIENT!r}."
)
return {
"namespace": ns,
"status": status,
"healthy": healthy,
"ide_namespace_proven": proven,
"clean_claim_allowed": status == HEALTH_CLEAN,
"mutations_allowed": status == HEALTH_CLEAN,
"reasons": reasons,
"health": health,
}
# --- Policy surface ---------------------------------------------------------
def restart_policy() -> dict[str, Any]:
"""Machine-readable description of the sanctioned restart contract."""
return {
"policy_version": 1,
"modes": list(MODES),
"actions": [ACTION_RESTART_NAMESPACE, ACTION_RELOAD_NAMESPACE],
"known_namespaces": list(KNOWN_NAMESPACES),
"fleet_scope_permitted": False,
"console_executes_process_kill": False,
"raw_kill_exposed": False,
"requires_confirmation": True,
"confirmation_binds_namespace": True,
"operator_authorization_env": (
runtime_recovery_guard.OPERATOR_AUTHORIZATION_ENV
),
"restart_hook_env": RESTART_HOOK_ENV,
"manual_kill_classified_as": runtime_recovery_guard.CONTAMINATION_KIND,
"post_restart_clean_claim_requires": (
mcp_namespace_health.PROBE_SOURCE_CLIENT
),
"audit_required": True,
"silent_auto_restart_permitted": False,
}
+682
View File
@@ -0,0 +1,682 @@
"""Read-only system-health model for the operator console API (#634).
`/health` answers liveness only. Operators automating readiness checks need a
structured view of *why* the control plane is or is not usable: which
dependencies answered, how long they took, what version of the code is running,
and whether the runtime is stale relative to its remote.
Three rules shape this module.
* **Read-only.** Every probe opens its subject read-only. The control-plane
database is opened through a ``mode=ro`` URI so a health check can never
create or migrate a schema, and no probe writes, restarts, or reloads
anything restart controls are Phase 2, and #630 forbids process-kill
recovery outright.
* **Fail-soft.** A dependency that is unreachable is a *status*, not an
exception. Probes catch their own failures and report them as a degraded or
down entry carrying a reason.
* **Never claim more than was proven.** Readiness is derived only from probes
that actually ran, ``mutation_safe`` stays false unless the parity commits are
known and equal, and an MCP namespace is reported unproven because a web
process cannot exercise the IDE-managed client path (#543).
"""
from __future__ import annotations
import os
import re
import sqlite3
import subprocess
import time
from dataclasses import dataclass
from datetime import datetime, timezone
from pathlib import Path
from typing import Any, Callable
from urllib.parse import urlsplit, urlunsplit
import control_plane_db
import mcp_namespace_health
from gitea_auth import api_request, get_auth_header, gitea_url
from webui.project_registry import load_registry
SERVICE_NAME = "mcp-control-plane-webui"
API_PATH = "/api/v1/system/health"
STATUS_OK = "ok"
STATUS_DEGRADED = "degraded"
STATUS_DOWN = "down"
STATUS_SKIPPED = "skipped"
STATUS_UNPROVEN = "unproven"
# Statuses that count as a healthy answer from a probe.
_HEALTHY_STATUSES = frozenset({STATUS_OK})
# Statuses meaning "this probe did not run", as opposed to "it ran and failed".
_NOT_RUN_STATUSES = frozenset({STATUS_SKIPPED})
_DEEP_PROBE_TTL_ENV = "WEBUI_HEALTH_PROBE_TTL_SECONDS"
_DEFAULT_DEEP_PROBE_TTL = 15.0
_GITEA_PROBE_TIMEOUT_SECONDS = 5.0
_OFFLINE_ENV = "WEBUI_TEST_OFFLINE"
# Credential-shaped material that must never reach the browser, mirroring the
# forbidden client patterns in webui/deployment_boundary.py.
_SECRET_RE = re.compile(
r"(?i)\b(token|password|passwd|secret|authorization|bearer)\b\s*[:=]?\s*\S+"
)
_LONG_OPAQUE_RE = re.compile(r"\b[A-Za-z0-9_\-]{32,}\b")
# Captured once at import so uptime measures this process, not the request.
_STARTED_AT = datetime.now(timezone.utc)
_STARTED_MONOTONIC = time.monotonic()
# TTL cache for the expensive (network) probe only.
_deep_cache: dict[str, tuple[float, "DependencyProbe"]] = {}
@dataclass(frozen=True)
class DependencyProbe:
"""One dependency check, fail-soft, with its own latency."""
name: str
kind: str
status: str
detail: str
required: bool
latency_ms: float | None = None
metadata: dict[str, Any] | None = None
@property
def healthy(self) -> bool:
return self.status in _HEALTHY_STATUSES
@property
def ran(self) -> bool:
return self.status not in _NOT_RUN_STATUSES
@dataclass(frozen=True)
class VersionInfo:
git_sha: str | None
git_describe: str | None
control_plane_schema_version: int | None
python_version: str
known: bool
@dataclass(frozen=True)
class StaleRuntime:
"""Parity between the running code, the checkout, and the remote.
``mutation_safe`` is deliberately conservative: unknown is not safe.
"""
daemon_head: str | None
checkout_head: str | None
remote_head: str | None
stale: bool
determinable: bool
mutation_safe: bool
reasons: tuple[str, ...]
@dataclass(frozen=True)
class SystemHealthSnapshot:
status: str
ready: bool
readiness_complete: bool
readiness_reasons: tuple[str, ...]
service: str
mode: str
version: VersionInfo
started_at: str
uptime_seconds: float
timestamp: str
deep_probes_requested: bool
dependencies: tuple[DependencyProbe, ...]
mcp_namespaces: tuple[dict[str, Any], ...]
stale_runtime: StaleRuntime
probe_errors: tuple[str, ...] = ()
def process_uptime() -> tuple[str, float]:
"""Process start timestamp and uptime — in-memory, safe for `/health`."""
return _STARTED_AT.isoformat(), round(time.monotonic() - _STARTED_MONOTONIC, 3)
def _offline() -> bool:
return (os.environ.get(_OFFLINE_ENV) or "").strip().lower() in {"1", "true", "yes"}
def _repo_root() -> Path:
override = (os.environ.get("WEBUI_REPO_ROOT") or "").strip()
if override:
return Path(override).resolve()
return Path(__file__).resolve().parent.parent
def _deep_probe_ttl() -> float:
raw = (os.environ.get(_DEEP_PROBE_TTL_ENV) or "").strip()
if not raw:
return _DEFAULT_DEEP_PROBE_TTL
try:
value = float(raw)
except ValueError:
return _DEFAULT_DEEP_PROBE_TTL
return value if value >= 0 else _DEFAULT_DEEP_PROBE_TTL
def redact(text: str) -> str:
"""Strip credential-shaped material from operator-visible probe text.
Probe details carry exception strings, and an exception raised by an HTTP
client can quote the request that failed. Redaction happens here, at the
boundary where those strings become part of a browser-bound payload.
"""
if not text:
return ""
cleaned = _redact_urls(text)
cleaned = _SECRET_RE.sub(lambda m: f"{m.group(1)}=[redacted]", cleaned)
return _LONG_OPAQUE_RE.sub("[redacted]", cleaned)
def _redact_urls(text: str) -> str:
return re.sub(r"https?://\S+", lambda m: redact_url(m.group(0)), text)
def redact_url(url: str) -> str:
"""Reduce a URL to scheme://host/path — no userinfo, no query, no fragment."""
try:
parts = urlsplit(url)
except ValueError:
return "[redacted-url]"
if not parts.scheme or not parts.hostname:
return "[redacted-url]"
netloc = parts.hostname
if parts.port:
netloc = f"{netloc}:{parts.port}"
return urlunsplit((parts.scheme, netloc, parts.path, "", ""))
def _git(repo: Path, *args: str) -> str | None:
try:
completed = subprocess.run(
["git", "-C", str(repo), *args],
capture_output=True,
text=True,
check=False,
timeout=10,
)
except (OSError, subprocess.SubprocessError):
return None
if completed.returncode != 0:
return None
return (completed.stdout or "").strip() or None
def _load_version(repo: Path, *, schema_version: int | None) -> VersionInfo:
import platform
git_sha = None if _offline() else _git(repo, "rev-parse", "HEAD")
describe = None if _offline() else _git(repo, "describe", "--tags", "--always")
return VersionInfo(
git_sha=git_sha,
git_describe=describe,
control_plane_schema_version=schema_version,
python_version=platform.python_version(),
known=bool(git_sha),
)
# ---------------------------------------------------------------------------
# Dependency probes
# ---------------------------------------------------------------------------
def _elapsed_ms(started: float) -> float:
return round((time.monotonic() - started) * 1000, 3)
def probe_control_plane_db(db_path: str | None = None) -> DependencyProbe:
"""Read-only reachability check for the control-plane SQLite substrate.
Opened through a ``mode=ro`` URI on purpose: ``ControlPlaneDB.__init__``
creates directories and runs schema migrations, which a health check must
never do.
"""
path = (db_path or control_plane_db.default_db_path()).strip()
started = time.monotonic()
metadata: dict[str, Any] = {"path": path}
def _result(status: str, detail: str) -> DependencyProbe:
return DependencyProbe(
name="control_plane_db",
kind="sqlite",
status=status,
detail=detail,
required=True,
latency_ms=_elapsed_ms(started),
metadata=metadata,
)
if not path or not os.path.exists(path):
return _result(STATUS_DOWN, "control-plane database file does not exist yet")
try:
conn = sqlite3.connect(f"file:{path}?mode=ro", uri=True, timeout=5)
try:
row = conn.execute(
"SELECT value FROM schema_meta WHERE key = 'schema_version'"
).fetchone()
leases = conn.execute(
"SELECT COUNT(*) FROM leases WHERE status = 'active'"
).fetchone()
finally:
conn.close()
except sqlite3.Error as exc:
return _result(STATUS_DOWN, redact(f"control-plane database unreadable: {exc}"))
schema_version = int(row[0]) if row and str(row[0]).isdigit() else None
metadata["schema_version"] = schema_version
metadata["active_leases"] = int(leases[0]) if leases else None
if schema_version is None:
return _result(
STATUS_DEGRADED, "control-plane database has no recorded schema version"
)
if schema_version != control_plane_db.SCHEMA_VERSION:
return _result(
STATUS_DEGRADED,
f"control-plane schema version {schema_version} does not match the "
f"version this code expects ({control_plane_db.SCHEMA_VERSION})",
)
return _result(STATUS_OK, f"schema v{schema_version} readable")
def probe_repository(repo: Path) -> DependencyProbe:
"""Local checkout reachability — required, cheap, no network."""
started = time.monotonic()
metadata: dict[str, Any] = {"repo_root": str(repo)}
if _offline():
return DependencyProbe(
name="repository",
kind="git",
status=STATUS_SKIPPED,
detail=f"{_OFFLINE_ENV} is set; git probe skipped",
required=True,
latency_ms=_elapsed_ms(started),
metadata=metadata,
)
head = _git(repo, "rev-parse", "HEAD")
if not head:
return DependencyProbe(
name="repository",
kind="git",
status=STATUS_DOWN,
detail=f"HEAD could not be read at {repo}",
required=True,
latency_ms=_elapsed_ms(started),
metadata=metadata,
)
branch = _git(repo, "rev-parse", "--abbrev-ref", "HEAD")
metadata["head"] = head
metadata["branch"] = branch
return DependencyProbe(
name="repository",
kind="git",
status=STATUS_OK,
detail=f"checkout readable at {branch or 'detached HEAD'}",
required=True,
latency_ms=_elapsed_ms(started),
metadata=metadata,
)
def probe_gitea(host: str) -> DependencyProbe:
"""Live Gitea reachability. Expensive (network), so opt-in via ``deep``.
Optional by design: the console stays useful for local inventory when the
remote is unreachable, so a failure here degrades status without claiming
the process itself is unready.
"""
started = time.monotonic()
metadata: dict[str, Any] = {"host": host}
def _failure(status: str, detail: str) -> DependencyProbe:
return DependencyProbe(
name="gitea",
kind="http",
status=status,
detail=detail,
required=False,
latency_ms=_elapsed_ms(started),
metadata=metadata,
)
if not host:
return _failure(STATUS_DEGRADED, "no Gitea host is configured in the registry")
try:
auth = get_auth_header(host)
except Exception as exc: # noqa: BLE001 — credential guards are a status here
return _failure(STATUS_DEGRADED, redact(f"credential lookup refused: {exc}"))
if not auth:
return _failure(STATUS_DEGRADED, f"no credentials available for {host}")
url = gitea_url(host, "/api/v1/version")
metadata["endpoint"] = redact_url(url)
try:
data = api_request("GET", url, auth, timeout=_GITEA_PROBE_TIMEOUT_SECONDS)
except Exception as exc: # noqa: BLE001 — a down dependency is a status
return _failure(STATUS_DOWN, redact(f"Gitea probe failed: {exc}"))
if isinstance(data, dict) and data.get("version"):
metadata["gitea_version"] = str(data["version"])
return DependencyProbe(
name="gitea",
kind="http",
status=STATUS_OK,
detail=f"{host} reachable",
required=False,
latency_ms=_elapsed_ms(started),
metadata=metadata,
)
def _skipped_gitea(host: str) -> DependencyProbe:
return DependencyProbe(
name="gitea",
kind="http",
status=STATUS_SKIPPED,
detail="network probe not requested; call with ?deep=1 to run it",
required=False,
latency_ms=None,
metadata={"host": host},
)
def namespace_summaries() -> tuple[dict[str, Any], ...]:
"""Declared MCP namespaces, each honestly reported as unproven.
The web process runs outside the IDE-managed MCP client, so it cannot
invoke a namespace tool. Per #543 only a ``client_namespace`` probe proves
that path, and inventing a healthy verdict here is exactly the false claim
the mutation gates exist to prevent.
"""
rows: list[dict[str, Any]] = []
for namespace, required_tool in sorted(
mcp_namespace_health.REQUIRED_NAMESPACE_TOOLS.items()
):
classification = mcp_namespace_health.classify_namespace_probe(
namespace,
required_tool=required_tool,
probe_result=None,
probe_source=mcp_namespace_health.PROBE_SOURCE_UNKNOWN,
)
rows.append(
{
"namespace": namespace,
"required_tool": required_tool,
"status": STATUS_UNPROVEN,
"ide_namespace_proven": bool(classification.get("ide_namespace_proven")),
"reason": (
"the web console cannot invoke the IDE-managed MCP client; "
"namespace health must be proven with a client_namespace "
"probe (#543)"
),
"error_type": classification.get("error_type"),
}
)
return tuple(rows)
def assess_stale_runtime(
repo: Path,
*,
daemon_head: str | None = None,
git_reader: Callable[..., str | None] | None = None,
) -> StaleRuntime:
"""Three-way parity view: running code, local checkout, remote-tracking ref.
``mutation_safe`` requires all three to be known and equal. Anything less
including "the remote ref was never fetched" is reported as not safe with
a reason, so an operator never reads an unproven green.
"""
reader = git_reader or (lambda *args: _git(repo, *args))
reasons: list[str] = []
# The offline switch suppresses real subprocess calls; an explicitly
# injected reader is already a substitute for them and is always used.
offline = _offline() and git_reader is None
checkout_head = None if offline else reader("rev-parse", "HEAD")
remote_head = None if offline else reader("rev-parse", "@{upstream}")
if offline:
reasons.append(f"{_OFFLINE_ENV} is set; parity commits were not read")
else:
if checkout_head is None:
reasons.append("local checkout HEAD could not be read")
if remote_head is None:
reasons.append(
"no remote-tracking commit is known for the current branch; "
"remote staleness is indeterminate (no fetch is performed here)"
)
effective_daemon = daemon_head if daemon_head is not None else checkout_head
if daemon_head is None:
reasons.append(
"the running MCP daemon's startup commit is not observable from the "
"web process; the checkout commit is reported in its place"
)
determinable = bool(checkout_head and remote_head and effective_daemon)
stale = bool(
determinable and len({checkout_head, remote_head, effective_daemon}) > 1
)
if stale:
reasons.append(
"runtime, checkout, and remote commits disagree; restart the MCP "
"server after updating the checkout before trusting capability gates"
)
return StaleRuntime(
daemon_head=effective_daemon,
checkout_head=checkout_head,
remote_head=remote_head,
stale=stale,
determinable=determinable,
mutation_safe=bool(determinable and not stale),
reasons=tuple(reasons),
)
# ---------------------------------------------------------------------------
# Snapshot assembly
# ---------------------------------------------------------------------------
def _default_host() -> str:
registry = load_registry()
if not registry.projects:
return ""
raw = registry.projects[0].remote_host
parts = urlsplit(raw.strip())
return parts.netloc or raw.strip().rstrip("/")
def _aggregate(
probes: tuple[DependencyProbe, ...],
) -> tuple[str, bool, bool, tuple[str, ...]]:
"""Fold probe results into overall status and readiness.
Required probes drive readiness; optional probes can only degrade status.
A probe that did not run leaves readiness incomplete rather than passing.
"""
reasons: list[str] = []
required = [probe for probe in probes if probe.required]
unrun_required = [probe for probe in required if not probe.ran]
failed_required = [probe for probe in required if probe.ran and not probe.healthy]
failed_optional = [
probe
for probe in probes
if not probe.required and probe.ran and not probe.healthy
]
for probe in unrun_required:
reasons.append(
f"required dependency '{probe.name}' was not probed: {probe.detail}"
)
for probe in failed_required:
reasons.append(
f"required dependency '{probe.name}' is {probe.status}: {probe.detail}"
)
for probe in failed_optional:
reasons.append(
f"optional dependency '{probe.name}' is {probe.status}: {probe.detail}"
)
readiness_complete = not unrun_required
ready = readiness_complete and not failed_required
if any(probe.status == STATUS_DOWN for probe in failed_required):
status = STATUS_DOWN
elif failed_required or failed_optional or unrun_required:
status = STATUS_DEGRADED
else:
status = STATUS_OK
return status, ready, readiness_complete, tuple(reasons)
def load_system_health(
*,
deep: bool = False,
host: str | None = None,
probes: tuple[DependencyProbe, ...] | None = None,
daemon_head: str | None = None,
use_cache: bool = True,
) -> SystemHealthSnapshot:
"""Assemble the read-only system-health snapshot.
``deep=True`` adds the network probe against Gitea; its result is cached for
a short TTL so repeated dashboard polls do not amplify into remote load.
"""
repo = _repo_root()
probe_errors: list[str] = []
if probes is None:
collected: list[DependencyProbe] = []
for probe_fn in (
lambda: probe_control_plane_db(),
lambda: probe_repository(repo),
):
try:
collected.append(probe_fn())
except Exception as exc: # noqa: BLE001 — a probe must not 500 the API
probe_errors.append(redact(f"probe raised: {exc}"))
resolved_host = host if host is not None else _default_host()
if deep and not _offline():
collected.append(_cached_gitea_probe(resolved_host, use_cache=use_cache))
else:
collected.append(_skipped_gitea(resolved_host))
probes = tuple(collected)
status, ready, readiness_complete, reasons = _aggregate(probes)
stale = assess_stale_runtime(repo, daemon_head=daemon_head)
if stale.stale:
if status == STATUS_OK:
status = STATUS_DEGRADED
reasons = reasons + (
"runtime is stale relative to its remote-tracking commit",
)
db_probe = next((p for p in probes if p.name == "control_plane_db"), None)
schema_version = None
if db_probe and db_probe.metadata:
schema_version = db_probe.metadata.get("schema_version")
return SystemHealthSnapshot(
status=status,
ready=ready,
readiness_complete=readiness_complete,
readiness_reasons=reasons,
service=SERVICE_NAME,
mode="read-only",
version=_load_version(repo, schema_version=schema_version),
started_at=_STARTED_AT.isoformat(),
uptime_seconds=round(time.monotonic() - _STARTED_MONOTONIC, 3),
timestamp=datetime.now(timezone.utc).isoformat(),
deep_probes_requested=deep,
dependencies=probes,
mcp_namespaces=namespace_summaries(),
stale_runtime=stale,
probe_errors=tuple(probe_errors),
)
def _cached_gitea_probe(host: str, *, use_cache: bool = True) -> DependencyProbe:
ttl = _deep_probe_ttl()
now = time.monotonic()
if use_cache and ttl > 0:
cached = _deep_cache.get(host)
if cached and (now - cached[0]) < ttl:
return cached[1]
probe = probe_gitea(host)
if use_cache and ttl > 0:
_deep_cache[host] = (now, probe)
return probe
def clear_probe_cache() -> None:
"""Drop cached deep-probe results (tests and operator-forced refresh)."""
_deep_cache.clear()
def probe_to_dict(probe: DependencyProbe) -> dict[str, Any]:
return {
"name": probe.name,
"kind": probe.kind,
"status": probe.status,
"detail": probe.detail,
"required": probe.required,
"healthy": probe.healthy,
"latency_ms": probe.latency_ms,
"metadata": dict(probe.metadata or {}),
}
def snapshot_to_dict(snapshot: SystemHealthSnapshot) -> dict[str, Any]:
return {
"status": snapshot.status,
"service": snapshot.service,
"mode": snapshot.mode,
"api": API_PATH,
"timestamp": snapshot.timestamp,
"readiness": {
"ready": snapshot.ready,
"complete": snapshot.readiness_complete,
"reasons": list(snapshot.readiness_reasons),
},
"version": {
"git_sha": snapshot.version.git_sha,
"git_describe": snapshot.version.git_describe,
"control_plane_schema_version": (
snapshot.version.control_plane_schema_version
),
"python_version": snapshot.version.python_version,
"known": snapshot.version.known,
},
"process": {
"started_at": snapshot.started_at,
"uptime_seconds": snapshot.uptime_seconds,
},
"deep_probes_requested": snapshot.deep_probes_requested,
"dependencies": [probe_to_dict(probe) for probe in snapshot.dependencies],
"mcp_namespaces": [dict(row) for row in snapshot.mcp_namespaces],
"stale_runtime": {
"daemon_head": snapshot.stale_runtime.daemon_head,
"checkout_head": snapshot.stale_runtime.checkout_head,
"remote_head": snapshot.stale_runtime.remote_head,
"stale": snapshot.stale_runtime.stale,
"determinable": snapshot.stale_runtime.determinable,
"mutation_safe": snapshot.stale_runtime.mutation_safe,
"reasons": list(snapshot.stale_runtime.reasons),
},
"probe_errors": list(snapshot.probe_errors),
}
+307
View File
@@ -0,0 +1,307 @@
"""HTML views for the system-health dashboard (#639).
Renders the read-only :class:`~webui.system_health.SystemHealthSnapshot`
produced by the Phase 1 system-health API (#634). The page offers no restart,
reload, or process-kill control: those are Phase 2 work, and manual process
kills are the contamination path #630 exists to prevent.
Every free-text field passes through :func:`webui.system_health.redact` before
it reaches HTML, so a probe detail that captured a token or a credentialed URL
cannot leak through the dashboard even though the API redacts it already.
"""
from __future__ import annotations
import html
from webui.system_health import (
STATUS_DEGRADED,
STATUS_DOWN,
STATUS_OK,
STATUS_SKIPPED,
STATUS_UNPROVEN,
DependencyProbe,
SystemHealthSnapshot,
redact,
)
_STATUS_BADGE_CLASS = {
STATUS_OK: "badge-health-ok",
STATUS_DEGRADED: "badge-health-degraded",
STATUS_DOWN: "badge-health-down",
STATUS_SKIPPED: "badge-health-skipped",
STATUS_UNPROVEN: "badge-health-unproven",
}
def _safe(value: object) -> str:
"""Escape free text for HTML after redacting anything secret-shaped.
Use this for every value that can carry arbitrary text probe details,
reasons, probe errors because those are where a credential could ride
along.
"""
return html.escape(redact(str(value)))
def _esc(value: object) -> str:
"""Escape a structured field for HTML without redacting it.
Commit SHAs, probe names, statuses, and timestamps are enumerated or
machine-generated, never credential-bearing. They must not go through
:func:`redact`: its opaque-token rule matches any 32-plus-character run,
so a 40-character git SHA would render as ``[redacted]`` and the parity
view the one thing an operator reads this page for would be blank.
"""
return html.escape(str(value))
def _status_badge(status: str) -> str:
css = _STATUS_BADGE_CLASS.get(status, "badge-health-unproven")
return f'<span class="badge {css}">{_esc(status)}</span>'
def _reason_list(reasons: tuple[str, ...], *, empty: str) -> str:
if not reasons:
return f"<p class='muted'>{html.escape(empty)}</p>"
items = "".join(f"<li>{_safe(reason)}</li>" for reason in reasons)
return f"<ul class='reasons'>{items}</ul>"
def _readiness_card(snapshot: SystemHealthSnapshot) -> str:
"""Overall readiness.
``ready`` and ``readiness_complete`` are shown separately on purpose: a
snapshot whose required probes never ran is not the same as one that ran
them and passed, and collapsing the two would render an unproven green.
"""
if snapshot.ready and snapshot.readiness_complete:
headline = "Ready"
elif snapshot.ready:
headline = "Ready (incomplete evidence)"
else:
headline = "Not ready"
return (
"<section class='health-card'>"
f"<h3>Readiness {_status_badge(snapshot.status)}</h3>"
f"<p class='health-headline'>{html.escape(headline)}</p>"
"<table class='detail'>"
f"<tr><th>Service</th><td><code>{_esc(snapshot.service)}</code></td></tr>"
f"<tr><th>Mode</th><td>{_esc(snapshot.mode)}</td></tr>"
f"<tr><th>Ready</th><td>{_esc(snapshot.ready)}</td></tr>"
"<tr><th>Readiness evidence complete</th>"
f"<td>{_esc(snapshot.readiness_complete)}</td></tr>"
"<tr><th>Deep probes requested</th>"
f"<td>{_esc(snapshot.deep_probes_requested)}</td></tr>"
f"<tr><th>Observed at</th><td><code>{_esc(snapshot.timestamp)}</code></td></tr>"
"</table>"
"<h4>Readiness reasons</h4>"
f"{_reason_list(snapshot.readiness_reasons, empty='No readiness objections recorded.')}"
"</section>"
)
def _version_card(snapshot: SystemHealthSnapshot) -> str:
version = snapshot.version
uptime_hours = snapshot.uptime_seconds / 3600.0
known = (
"resolved"
if version.known
else "unresolved — version fields could not be read from the checkout"
)
schema = version.control_plane_schema_version
return (
"<section class='health-card'>"
"<h3>Version and uptime</h3>"
"<table class='detail'>"
f"<tr><th>Git SHA</th><td><code>{_esc(version.git_sha or 'unknown')}</code></td></tr>"
"<tr><th>Git describe</th>"
f"<td><code>{_esc(version.git_describe or 'unknown')}</code></td></tr>"
"<tr><th>Control-plane schema</th>"
f"<td>{_esc(schema if schema is not None else 'unknown')}</td></tr>"
f"<tr><th>Python</th><td><code>{_esc(version.python_version)}</code></td></tr>"
f"<tr><th>Version status</th><td>{html.escape(known)}</td></tr>"
f"<tr><th>Started at</th><td><code>{_esc(snapshot.started_at)}</code></td></tr>"
"<tr><th>Uptime</th>"
f"<td>{snapshot.uptime_seconds:.3f}s ({uptime_hours:.2f}h)</td></tr>"
"</table>"
"</section>"
)
def _dependency_rows(probes: tuple[DependencyProbe, ...]) -> str:
if not probes:
return "<p class='muted'>No dependency probes were reported.</p>"
rows = []
for probe in probes:
latency = (
f"{probe.latency_ms:.1f} ms" if probe.latency_ms is not None else "n/a"
)
rows.append(
"<tr>"
f"<td><code>{_esc(probe.name)}</code></td>"
f"<td>{_esc(probe.kind)}</td>"
f"<td>{_status_badge(probe.status)}</td>"
f"<td>{_esc('required' if probe.required else 'optional')}</td>"
f"<td>{html.escape(latency)}</td>"
f"<td>{_safe(probe.detail)}</td>"
"</tr>"
)
return (
"<table class='registry'><thead><tr>"
"<th>Dependency</th><th>Kind</th><th>Status</th><th>Requirement</th>"
"<th>Latency</th><th>Detail</th>"
"</tr></thead><tbody>"
f"{''.join(rows)}</tbody></table>"
)
def _dependency_card(snapshot: SystemHealthSnapshot) -> str:
degraded = [probe for probe in snapshot.dependencies if probe.ran and not probe.healthy]
not_run = [probe for probe in snapshot.dependencies if not probe.ran]
banner = ""
if degraded:
names = ", ".join(sorted(probe.name for probe in degraded))
banner += (
"<div class='stub health-degraded'><p><strong>Degraded dependencies:</strong> "
f"{_esc(names)}</p></div>"
)
if not_run:
names = ", ".join(sorted(probe.name for probe in not_run))
banner += (
"<div class='stub'><p><strong>Not probed:</strong> "
f"{_esc(names)} — these contribute no evidence and are not counted "
"as healthy.</p></div>"
)
return (
"<section class='health-card'>"
"<h3>Dependencies</h3>"
f"{banner}"
f"{_dependency_rows(snapshot.dependencies)}"
"<p class='muted'>Details are redacted at the API boundary and again "
"before rendering; credentials are never displayed.</p>"
"</section>"
)
def _namespace_card(snapshot: SystemHealthSnapshot) -> str:
if not snapshot.mcp_namespaces:
body = "<p class='muted'>No MCP namespaces are declared.</p>"
else:
rows = []
for entry in snapshot.mcp_namespaces:
rows.append(
"<tr>"
f"<td><code>{_esc(entry.get('namespace'))}</code></td>"
f"<td><code>{_esc(entry.get('required_tool'))}</code></td>"
f"<td>{_status_badge(str(entry.get('status') or STATUS_UNPROVEN))}</td>"
f"<td>{_esc(entry.get('ide_namespace_proven'))}</td>"
f"<td>{_safe(entry.get('reason'))}</td>"
"</tr>"
)
body = (
"<table class='registry'><thead><tr>"
"<th>Namespace</th><th>Required tool</th><th>Status</th>"
"<th>IDE-proven</th><th>Reason</th>"
"</tr></thead><tbody>"
f"{''.join(rows)}</tbody></table>"
)
return (
"<section class='health-card'>"
"<h3>MCP namespaces</h3>"
f"{body}"
"<p class='muted'>The web process runs outside the IDE-managed MCP "
"client, so namespace health is reported as unproven rather than "
"guessed (#543).</p>"
"</section>"
)
def _stale_runtime_card(snapshot: SystemHealthSnapshot) -> str:
stale = snapshot.stale_runtime
if stale.stale:
warning = (
"<div class='stub health-stale'><p><strong>Stale runtime:</strong> "
"the running code, the checkout, and the remote-tracking commit "
"disagree. Capability gates may be evaluating obsolete code — "
"do not treat this runtime as mutation-safe.</p></div>"
)
elif not stale.determinable:
warning = (
"<div class='stub health-stale'><p><strong>Staleness "
"indeterminate:</strong> parity could not be proven, so this "
"runtime is not reported as mutation-safe.</p></div>"
)
else:
warning = ""
return (
"<section class='health-card'>"
"<h3>Stale-runtime parity</h3>"
f"{warning}"
"<table class='detail'>"
"<tr><th>Daemon head</th>"
f"<td><code>{_esc(stale.daemon_head or 'unknown')}</code></td></tr>"
"<tr><th>Checkout head</th>"
f"<td><code>{_esc(stale.checkout_head or 'unknown')}</code></td></tr>"
"<tr><th>Remote head</th>"
f"<td><code>{_esc(stale.remote_head or 'unknown')}</code></td></tr>"
f"<tr><th>Stale</th><td>{_esc(stale.stale)}</td></tr>"
f"<tr><th>Determinable</th><td>{_esc(stale.determinable)}</td></tr>"
f"<tr><th>Mutation safe</th><td>{_esc(stale.mutation_safe)}</td></tr>"
"</table>"
f"{_reason_list(stale.reasons, empty='Runtime, checkout, and remote agree.')}"
"</section>"
)
def _probe_error_card(snapshot: SystemHealthSnapshot) -> str:
if not snapshot.probe_errors:
return ""
return (
"<section class='health-card'>"
"<h3>Probe errors</h3>"
f"{_reason_list(snapshot.probe_errors, empty='')}"
"</section>"
)
def _recovery_card() -> str:
"""Sanctioned recovery pointers only — never a manual process kill (#630)."""
return (
"<section class='health-card'>"
"<h3>Recovery</h3>"
"<p class='muted'>This dashboard is read-only. Restart and reload "
"controls arrive in Phase 2 (#642); until then recovery runs through "
"the sanctioned client reconnect / operator restart path.</p>"
"<ul class='reasons'>"
"<li><a href='/runtime'>Runtime and session view</a> — active profile, "
"workflow hashes, and shell health.</li>"
"<li>Reconnect the MCP client from the IDE, then re-run the blocked "
"cycle. Never kill the daemon process manually: unmanaged kills are "
"recorded as runtime contamination (#630).</li>"
"<li>See <code>docs/webui-local-dev.md</code> for the documented "
"recovery sequence.</li>"
"</ul>"
"</section>"
)
def render_system_health_page(snapshot: SystemHealthSnapshot) -> str:
"""Render the full system-health dashboard body."""
return (
"<h2>System health</h2>"
"<p class='meta'>Read-only view of the Phase 1 system-health API "
"(<code>/api/v1/system/health</code>). Reload this page to refresh; "
"nothing here polls or mutates on your behalf.</p>"
f"{_readiness_card(snapshot)}"
f"{_stale_runtime_card(snapshot)}"
f"{_version_card(snapshot)}"
f"{_dependency_card(snapshot)}"
f"{_namespace_card(snapshot)}"
f"{_probe_error_card(snapshot)}"
f"{_recovery_card()}"
)
+906
View File
@@ -0,0 +1,906 @@
"""Workflow-event and conversation timeline model (#637, Phase 1).
Operators cannot browse a unified timeline of workflow events, decisions,
tool calls, and handoffs: the evidence is scattered across control-plane
events, Gitea canonical handoff comments, and local logs. This module defines
one durable, versioned event schema and per-source adapters that normalise
those scattered records into a single ``WorkflowEvent`` stream, plus a
read-only query layer (filter by issue / PR / session, stable ordering,
pagination) that the ``/api/v1/timeline`` route serves.
Design rules honoured here:
- **Read-only.** Sources are read; nothing is mutated. The control-plane
database is opened through a ``mode=ro`` URI so a missing or unwritable DB
degrades to a reason instead of creating directories or running migrations.
- **Fail-soft per source.** An unavailable source degrades to a status with a
reason rather than raising, and a source that could not run is never
rendered as an empty-and-healthy timeline.
- **Answerable filters only.** Each source declares which filter dimensions it
can actually answer. A filter dimension no source that ran can carry is
refused with an explicit reason rather than silently matching nothing: an
empty page from an unanswerable filter reads to an operator as "no such
activity", which is a different — and false — statement.
- **Redaction at the boundary, fail closed.** Every free-text field (event
messages, redacted tool arguments, decision/proof text) is run through the
console redaction policy before it leaves this module, and *before* any
structured value is derived from it evidence references are extracted from
redacted text, then independently revalidated before serialization. An
unredactable value becomes the placeholder, and a value that cannot be proven
safe is dropped an unredacted payload is never emitted, and a generation
error never drops raw data to a caller or a log.
- **Stable ordering.** Events sort by ``(timestamp, source_rank, event_key)``
with a deterministic tiebreak, so pagination is stable across calls and
events with equal or missing timestamps keep a fixed order.
Non-goals (from the issue): no full chat replay, no mutation of historical
events, no unredacted tool-argument storage.
"""
from __future__ import annotations
import re
import sqlite3
from dataclasses import dataclass, replace
from datetime import datetime, timezone
from typing import Any, Callable, Iterable
import control_plane_db
from webui import console_redaction
# The schema is versioned so consumers can branch on shape. Bump on any
# breaking change to WorkflowEvent's serialized form.
TIMELINE_SCHEMA_VERSION = 1
# Known event sources and their deterministic ordering rank. When two events
# carry the same timestamp, the source rank breaks the tie before the
# per-source event key, so a control-plane event and a handoff comment minted
# in the same second always sort in a fixed order.
SOURCE_CONTROL_PLANE = "control_plane"
SOURCE_GITEA_HANDOFF = "gitea_handoff"
_SOURCE_RANK = {
SOURCE_CONTROL_PLANE: 0,
SOURCE_GITEA_HANDOFF: 1,
}
# The filter dimensions the query layer accepts.
FILTER_ISSUE = "issue"
FILTER_PR = "pr"
FILTER_SESSION = "session"
# Which dimensions each source can actually answer. This is a property of the
# underlying records, not of the query code: the control-plane ``events`` table
# is (event_id, work_item_id, event_type, message, created_at) and carries no
# session identity at all, so no control-plane event can ever match a session
# filter. A CTH handoff comment can declare its session as a field, so the
# handoff source answers all three. Filtering on a dimension the surviving
# sources cannot carry is refused in ``load_timeline`` rather than answered
# with an empty page.
_SOURCE_FILTER_SUPPORT: dict[str, tuple[str, ...]] = {
SOURCE_CONTROL_PLANE: (FILTER_ISSUE, FILTER_PR),
SOURCE_GITEA_HANDOFF: (FILTER_ISSUE, FILTER_PR, FILTER_SESSION),
}
# Why a source cannot answer a dimension, for the refusal reason an operator reads.
_SOURCE_FILTER_LIMITS: dict[tuple[str, str], str] = {
(SOURCE_CONTROL_PLANE, FILTER_SESSION): (
"control-plane events carry no session identity "
"(the events table has no session column)"
),
}
# A timestamp far in the future so events with no parseable timestamp sort
# last (after everything real) instead of first, without raising.
_MISSING_TS_SORT = "9999-12-31T23:59:59Z"
def _parse_ts(value: str | None) -> str | None:
"""Normalise a timestamp to ``...Z`` UTC, or None when unparseable."""
if not value:
return None
text = str(value).strip()
if not text:
return None
candidate = text[:-1] + "+00:00" if text.endswith("Z") else text
try:
parsed = datetime.fromisoformat(candidate)
except ValueError:
return None
if parsed.tzinfo is None:
parsed = parsed.replace(tzinfo=timezone.utc)
return parsed.astimezone(timezone.utc).replace(microsecond=0).isoformat().replace("+00:00", "Z")
def _redact(value: Any) -> Any:
"""Redact a single free-text field, failing closed to the placeholder."""
if value is None:
return None
return console_redaction.redact_text(str(value))
@dataclass(frozen=True)
class WorkflowEvent:
"""One normalised timeline event.
Every field is optional except ``source``/``event_type``/``event_key``
because sources carry different subsets. The class is frozen so an adapted
event is an immutable record; a consumer that needs a variant builds a new
one rather than mutating history.
"""
source: str
event_type: str
event_key: str
timestamp: str | None = None
actor: str | None = None
role: str | None = None
issue_number: int | None = None
pr_number: int | None = None
session_id: str | None = None
tool_name: str | None = None
decision: str | None = None
message: str | None = None
correlation_id: str | None = None
evidence_refs: tuple[str, ...] = ()
sensitive: bool = False
def sort_key(self) -> tuple[str, int, str]:
return (
self.timestamp or _MISSING_TS_SORT,
_SOURCE_RANK.get(self.source, 99),
self.event_key,
)
def to_dict(self) -> dict[str, Any]:
return {
"source": self.source,
"event_type": self.event_type,
"event_key": self.event_key,
"timestamp": self.timestamp,
"actor": self.actor,
"role": self.role,
"issue_number": self.issue_number,
"pr_number": self.pr_number,
"session_id": self.session_id,
"tool_name": self.tool_name,
"decision": self.decision,
"message": self.message,
"correlation_id": self.correlation_id,
"evidence_refs": list(self.evidence_refs),
"sensitive": self.sensitive,
}
# --------------------------------------------------------------------------- #
# Adapters — pure functions from a source's raw records to WorkflowEvents. #
# Each is total: a malformed record is skipped, never raised on. #
# --------------------------------------------------------------------------- #
# Event types whose payload is treated as sensitive and always redaction-hard
# (they can carry lease/session provenance or tool arguments).
_SENSITIVE_EVENT_HINTS = ("lease", "capability", "token", "auth", "secret")
# Reference tokens (issue/PR/comment ids) and SHAs parsed out of proof text.
_EVIDENCE_REF_RE = re.compile(r"(?:#|PR\s*#?|issue\s*#?|comment\s*#?)(\d+)", re.IGNORECASE)
# A commit reference is only recognised when the text *declares* it as one.
# A bare lowercase hex run is not evidence of anything: at 40 characters it is
# exactly the shape of a Gitea personal access token, and at 7 it also matches
# ordinary words such as "defaced". Requiring an anchoring keyword keeps real
# references ("commit abc1234", "at head a209756...", "base caaae9b6") usable
# while refusing to lift an undeclared secret-shaped run out of free text.
_SHA_RE = re.compile(
r"(?i:\b(?:commit|sha|head|base|parent|revision|rev|merge[- ]base)\b[\s:=@#]*)"
r"([0-9a-f]{7,40})\b"
)
# Shapes a serialized evidence reference is allowed to take. Anything else is
# dropped rather than emitted.
_REF_ISSUE_SHAPE = re.compile(r"^#[0-9]{1,9}$")
_REF_SHA_SHAPE = re.compile(r"^[0-9a-f]{7,40}$")
# A long undelimited hex run with no declaring context is treated as credential
# material wherever it appears, never as an identifier.
_BARE_SECRET_SHAPE = re.compile(r"^[0-9a-f]{32,}$")
# An event type reads like an identifier, but a stored one is externally
# influenced: any producer that writes the control-plane ``events`` table
# chooses the string. It reaches ``to_dict`` verbatim, so it is validated here
# rather than trusted because of where it came from.
_CP_EVENT_TYPE_SHAPE = re.compile(r"^[A-Za-z][A-Za-z0-9._:+-]{0,63}$")
# Emitted in place of a value that cannot be proven safe. Deliberately not a
# plausible workflow type: an unsafe value is refused, never quietly rewritten
# into a different valid-looking one that would misdescribe the record.
UNSAFE_EVENT_TYPE = "unsafe:redacted"
# Emitted for a CTH heading that is not a declared member of ``CTH_TYPES``. The
# contract is enforced on write (``format_cth_body``) and on assess; the read
# path the timeline uses enforces it too rather than assuming it was.
UNKNOWN_HANDOFF_EVENT_TYPE = "handoff:unrecognized"
# A source record id is a plain integer in both sources it comes from: the
# control-plane ``events`` primary key and a Gitea comment id. ``event_key`` is
# serialized verbatim and is the pagination tiebreak, so anything else is
# refused rather than interpolated into it.
_RECORD_ID_SHAPE = re.compile(r"^[0-9]{1,19}$")
def _kind_to_numbers(kind: str | None, number: int | None) -> tuple[int | None, int | None]:
"""Map a control-plane work-item (kind, number) to (issue_no, pr_no)."""
if number is None:
return (None, None)
if kind == "pr":
return (None, int(number))
if kind == "issue":
return (int(number), None)
return (None, None)
def _correlation_for(kind: str | None, number: int | None) -> str | None:
if number is None or kind not in ("issue", "pr"):
return None
return f"{kind}#{number}"
def _extract_evidence_refs(*texts: str | None) -> tuple[str, ...]:
"""Extract issue/PR and declared-commit references from **redacted** text.
Callers must pass text that has already been through :func:`_redact`; this
function derives a structured field from its input, so extracting ahead of
redaction would republish whatever redaction was about to remove. Every
reference is revalidated by :func:`_validated_evidence_refs` before it is
serialized.
"""
refs: list[str] = []
for text in texts:
if not text:
continue
for match in _EVIDENCE_REF_RE.finditer(text):
token = f"#{match.group(1)}"
if token not in refs:
refs.append(token)
for match in _SHA_RE.finditer(text):
token = match.group(1)
if token not in refs:
refs.append(token)
return tuple(refs)
def _validated_evidence_refs(refs: Iterable[str]) -> tuple[tuple[str, ...], bool]:
"""Independently revalidate references immediately before serialization.
Extraction is not trusted on its own. A reference survives only when it has
a known reference shape and is unchanged by a second redaction pass a
value the redaction policy would alter is credential material that must not
be emitted as a structured field. A full 40-character SHA stays usable
because extraction only accepts a hex run the source text explicitly
declared as a commit. Returns ``(safe_refs, dropped_any)``; ``dropped_any``
marks the event sensitive so the drop is visible rather than silent.
"""
safe: list[str] = []
dropped = False
for ref in refs or ():
try:
token = str(ref).strip()
if not token:
continue
recognised = bool(_REF_ISSUE_SHAPE.match(token) or _REF_SHA_SHAPE.match(token))
if not recognised:
dropped = True
continue
if _redact(token) != token:
dropped = True
continue
if token not in safe:
safe.append(token)
except Exception:
# Fail closed: a reference that cannot be proven safe is dropped.
dropped = True
continue
return (tuple(safe), dropped)
def _safe_session_id(value: Any) -> str | None:
"""Return a session identifier only when it is safe to emit.
The value is authoritative source data a session the record names for
itself but it is still free text. It is dropped when redaction alters it
or when it is a bare secret-shaped hex run, so a credential parked in a
session field can never reach the payload or be echoed back by a filter.
"""
if value is None:
return None
text = str(value).strip()
if not text:
return None
if _BARE_SECRET_SHAPE.match(text):
return None
return text if _redact(text) == text else None
def _safe_record_id(value: Any) -> str | None:
"""Return a source record id only when it is a plain numeric identifier.
``event_key`` is serialized verbatim and is the deterministic pagination
tiebreak, so an id is interpolated into it only when it has the shape both
real sources actually produce. A record whose identity cannot be trusted is
refused by the caller rather than keyed on.
"""
if value is None or isinstance(value, bool):
return None
if isinstance(value, int):
return str(value)
text = str(value).strip()
return text if _RECORD_ID_SHAPE.match(text) else None
def _safe_cp_event_type(value: Any) -> tuple[str, bool]:
"""Validate a stored control-plane event type. Returns ``(type, unsafe)``.
The stored value is externally influenced whichever producer wrote the
``events`` row chose the string and ``to_dict`` serializes it verbatim, so
it passes a boundary of its own instead of relying on the one ``message``
passes. A value survives only when it is an ordinary identifier, is not a
bare secret-shaped hex run, and is unchanged by a redaction pass. Anything
else fails closed to :data:`UNSAFE_EVENT_TYPE`: the record stays visible as
an audit entry, but the value itself is never republished not verbatim,
not partially sanitized, and not rewritten into some other valid-looking
type that would misdescribe what happened.
"""
text = ("" if value is None else str(value)).strip()
if not text:
return ("", False)
if _BARE_SECRET_SHAPE.match(text):
return (UNSAFE_EVENT_TYPE, True)
if not _CP_EVENT_TYPE_SHAPE.match(text):
return (UNSAFE_EVENT_TYPE, True)
if _redact(text) != text:
return (UNSAFE_EVENT_TYPE, True)
return (text, False)
def _safe_echo(value: Any) -> Any:
"""Guard a scalar that is echoed back rather than derived from a record.
Query scope and filter values are caller-supplied and are reflected in the
response so an operator can see what was asked. Reflection is still
emission: a value redaction would alter, or a bare secret-shaped hex run, is
replaced by the placeholder instead of being echoed verbatim. Ordinary
scope and filter values pass through untouched.
"""
if value is None or isinstance(value, (int, bool)):
return value
text = str(value)
if _BARE_SECRET_SHAPE.match(text.strip()):
return console_redaction.REDACTED
return _redact(text)
def adapt_cp_events(rows: Iterable[dict[str, Any]]) -> list[WorkflowEvent]:
"""Adapt control-plane ``events`` rows (joined to work_items) into events.
Each row is expected to carry ``event_id``, ``event_type``, ``message``,
``created_at`` and the joined work-item ``kind``/``number``. Rows missing
an id or type are skipped so a partially written table never raises.
"""
events: list[WorkflowEvent] = []
for row in rows or []:
try:
event_id = _safe_record_id(row.get("event_id"))
raw_event_type = (row.get("event_type") or "").strip()
if event_id is None or not raw_event_type:
continue
# The stored type is source data, not a trusted constant: validate
# it before it is serialized, exactly as `message` below is redacted
# before it is serialized.
event_type, event_type_unsafe = _safe_cp_event_type(raw_event_type)
kind = row.get("kind")
number = row.get("number")
issue_no, pr_no = _kind_to_numbers(kind, number)
sensitive = event_type_unsafe or any(
hint in raw_event_type.lower() for hint in _SENSITIVE_EVENT_HINTS
)
events.append(
WorkflowEvent(
source=SOURCE_CONTROL_PLANE,
event_type=event_type,
event_key=f"cp:{event_id}",
timestamp=_parse_ts(row.get("created_at")),
issue_number=issue_no,
pr_number=pr_no,
# No session_id: the control-plane events table is
# (event_id, work_item_id, event_type, message, created_at)
# and records no session. Inventing one from the work item
# or the message text would be a guess, so this source
# declares the session dimension unsupported instead
# (_SOURCE_FILTER_SUPPORT) and the query layer refuses a
# session filter it cannot honestly answer.
message=_redact(row.get("message")),
correlation_id=_correlation_for(kind, number),
sensitive=sensitive,
)
)
except Exception:
# A single malformed row must not sink the whole adaptation.
continue
return events
def adapt_cth_comments(
comments: Iterable[dict[str, Any]],
*,
kind: str,
number: int,
) -> list[WorkflowEvent]:
"""Adapt Gitea Canonical Thread Handoff (CTH) comments into events.
Only comments that parse as a CTH (``canonical_thread_handoff.parse_cth_comment``)
become events; ordinary comments are ignored. ``kind``/``number`` scope the
events to the issue or PR the comments belong to.
"""
# Imported lazily so this module has no import-time dependency on the
# handoff parser when only the control-plane adapter is used.
from canonical_thread_handoff import is_known_cth_type, parse_cth_comment
# ``kind``/``number`` are interpolated into event_key and correlation_id, so
# they are normalised once here. A scope this adapter cannot express is
# refused outright rather than serialized into an identifier.
kind = (kind or "").strip().lower()
if kind not in ("issue", "pr"):
return []
try:
number = int(number)
except (TypeError, ValueError):
return []
issue_no, pr_no = _kind_to_numbers(kind, number)
correlation = _correlation_for(kind, number)
events: list[WorkflowEvent] = []
for comment in comments or []:
try:
body = comment.get("body") or ""
parsed = parse_cth_comment(body)
if not parsed:
continue
fields = parsed.get("fields") or {}
cth_type = parsed.get("cth_type") or ""
comment_id = _safe_record_id(comment.get("id"))
if comment_id is None:
continue
# The CTH heading is free text: the parser accepts whatever follows
# "## CTH:", and only the write and assess paths check it against
# the contract. Check it here too — an unrecognised heading is
# reported as such rather than serialized into event_type, so
# arbitrary, malformed, or secret-shaped heading content has no way
# through. Declared types are preserved exactly.
cth_type_known = is_known_cth_type(cth_type)
# Redaction runs first, and every derived value is taken from the
# redacted text — deriving evidence refs from the raw proof would
# re-emit exactly what redaction was about to remove.
decision = _redact(fields.get("decision"))
proof = _redact(fields.get("proof"))
next_action = _redact(fields.get("next action"))
refs, refs_dropped = _validated_evidence_refs(
_extract_evidence_refs(proof, decision)
)
events.append(
WorkflowEvent(
source=SOURCE_GITEA_HANDOFF,
event_type=(
f"handoff:{cth_type.strip()}"
if cth_type_known
else UNKNOWN_HANDOFF_EVENT_TYPE
),
event_key=f"cth:{kind}:{number}:{comment_id}",
timestamp=_parse_ts(comment.get("created_at")),
actor=_redact((comment.get("user") or {}).get("login")),
role=_redact(fields.get("next owner")),
issue_number=issue_no,
pr_number=pr_no,
# A CTH names its own session when the producer records one;
# it is read from that declared field, never inferred from
# unrelated text.
session_id=_safe_session_id(fields.get("session")),
decision=decision,
message=next_action or _redact(fields.get("status")),
correlation_id=correlation,
evidence_refs=refs,
sensitive=refs_dropped or not cth_type_known,
)
)
except Exception:
continue
return events
# --------------------------------------------------------------------------- #
# Read-only control-plane event source. #
# --------------------------------------------------------------------------- #
_CP_EVENTS_QUERY = """
SELECT e.event_id AS event_id,
e.event_type AS event_type,
e.message AS message,
e.created_at AS created_at,
w.kind AS kind,
w.number AS number
FROM events e
JOIN work_items w ON e.work_item_id = w.work_item_id
WHERE w.remote = ? AND w.org = ? AND w.repo = ?
"""
@dataclass(frozen=True)
class SourceStatus:
"""Fail-soft status for one timeline source.
``supported_filters`` states which filter dimensions this source's records
can carry; ``unsupported_filters`` names the requested dimensions it cannot,
so an operator can see *why* a source contributed nothing rather than being
left to read an empty list as an absence of activity.
"""
name: str
ok: bool
reason: str | None = None
count: int = 0
supported_filters: tuple[str, ...] = ()
unsupported_filters: tuple[str, ...] = ()
def to_dict(self) -> dict[str, Any]:
return {
"name": self.name,
"ok": self.ok,
"reason": self.reason,
"count": self.count,
"supported_filters": list(self.supported_filters),
"unsupported_filters": list(self.unsupported_filters),
}
def _cp_status(*, ok: bool, reason: str | None = None, count: int = 0) -> SourceStatus:
return SourceStatus(
SOURCE_CONTROL_PLANE,
ok=ok,
# A failure reason is serialized like any other field and is often an
# exception string carrying a path or a transport error, so it crosses
# the redaction boundary too. Static reasons pass through unchanged.
reason=_redact(reason),
count=count,
supported_filters=_SOURCE_FILTER_SUPPORT[SOURCE_CONTROL_PLANE],
)
def _handoff_status(*, ok: bool, reason: str | None = None, count: int = 0) -> SourceStatus:
return SourceStatus(
SOURCE_GITEA_HANDOFF,
ok=ok,
# Same boundary as the control-plane status: this reason can quote an
# error raised by a live authenticated fetch.
reason=_redact(reason),
count=count,
supported_filters=_SOURCE_FILTER_SUPPORT[SOURCE_GITEA_HANDOFF],
)
def read_cp_events(
*,
remote: str,
org: str,
repo: str,
db_path: str | None = None,
) -> tuple[list[WorkflowEvent], SourceStatus]:
"""Read scoped control-plane events read-only. Never creates the DB.
Opens the SQLite file through a ``mode=ro`` URI: a health/timeline read
must never create directories or run the schema migration that
``ControlPlaneDB()`` performs on construction. A missing or unreadable DB
degrades to a status with a reason.
"""
path = (db_path or control_plane_db.default_db_path()).strip()
conn: sqlite3.Connection | None = None
try:
conn = sqlite3.connect(f"file:{path}?mode=ro", uri=True)
conn.row_factory = sqlite3.Row
cursor = conn.execute(_CP_EVENTS_QUERY, (remote, org, repo))
rows = [dict(r) for r in cursor.fetchall()]
except sqlite3.OperationalError as exc:
return ([], _cp_status(ok=False, reason=f"control-plane DB unavailable: {exc}"))
except sqlite3.Error as exc:
return ([], _cp_status(ok=False, reason=f"control-plane read failed: {exc}"))
finally:
if conn is not None:
conn.close()
events = adapt_cp_events(rows)
return (events, _cp_status(ok=True, count=len(events)))
# --------------------------------------------------------------------------- #
# Filter, sort, paginate. #
# --------------------------------------------------------------------------- #
def filter_events(
events: Iterable[WorkflowEvent],
*,
issue: int | None = None,
pr: int | None = None,
session: str | None = None,
) -> list[WorkflowEvent]:
"""Filter events by issue number, PR number, and/or session id.
Filters are conjunctive. A filter that names a dimension an event does not
carry excludes that event (an issue filter excludes PR-only events).
"""
out: list[WorkflowEvent] = []
for ev in events:
if issue is not None and ev.issue_number != issue:
continue
if pr is not None and ev.pr_number != pr:
continue
if session is not None and ev.session_id != session:
continue
out.append(ev)
return out
def sort_events(events: Iterable[WorkflowEvent]) -> list[WorkflowEvent]:
"""Return events in stable timeline order (ascending)."""
return sorted(events, key=lambda ev: ev.sort_key())
@dataclass(frozen=True)
class TimelinePage:
"""One page of the sorted, filtered timeline."""
events: tuple[WorkflowEvent, ...]
total: int
limit: int
offset: int
@property
def next_offset(self) -> int | None:
nxt = self.offset + len(self.events)
return nxt if nxt < self.total else None
def to_dict(self) -> dict[str, Any]:
return {
"events": [ev.to_dict() for ev in self.events],
"pagination": {
"total": self.total,
"limit": self.limit,
"offset": self.offset,
"returned": len(self.events),
"next_offset": self.next_offset,
"has_more": self.next_offset is not None,
},
}
_MAX_LIMIT = 500
_DEFAULT_LIMIT = 50
def _coerce_bounds(limit: int | None, offset: int | None) -> tuple[int, int]:
try:
lim = int(limit) if limit is not None else _DEFAULT_LIMIT
except (TypeError, ValueError):
lim = _DEFAULT_LIMIT
try:
off = int(offset) if offset is not None else 0
except (TypeError, ValueError):
off = 0
lim = max(1, min(lim, _MAX_LIMIT))
off = max(0, off)
return (lim, off)
def paginate(events: list[WorkflowEvent], *, limit: int | None, offset: int | None) -> TimelinePage:
lim, off = _coerce_bounds(limit, offset)
window = events[off : off + lim]
return TimelinePage(events=tuple(window), total=len(events), limit=lim, offset=off)
# --------------------------------------------------------------------------- #
# Composition — load_timeline aggregates all sources, fail-soft. #
# --------------------------------------------------------------------------- #
# A comment source is a callable that, given (kind, number), returns the raw
# Gitea comment list for that issue/PR. The route supplies a live fail-soft
# fetcher; tests supply a fixture. When None, the handoff source is reported as
# not-run (never silently empty-and-healthy).
CommentSource = Callable[[str, int], list[dict[str, Any]]]
@dataclass(frozen=True)
class TimelineSnapshot:
"""One answered timeline query.
``ok`` is False when the query could not be answered as asked currently
when a requested filter dimension no surviving source can carry was
supplied. The page is then empty *and* the snapshot says so, because an
``ok`` empty page is a claim that no such activity exists.
"""
schema_version: int
remote: str
org: str
repo: str
filters: dict[str, Any]
page: TimelinePage
sources: tuple[SourceStatus, ...]
ok: bool = True
error: dict[str, Any] | None = None
def to_dict(self) -> dict[str, Any]:
return {
"ok": self.ok,
"error": self.error,
"schema_version": self.schema_version,
# Scope and filters are echoed caller input, not derived record
# data. Reflecting a value is still emitting it, so both cross the
# same boundary; ordinary scope and filter values are unchanged.
"scope": {
"remote": _safe_echo(self.remote),
"org": _safe_echo(self.org),
"repo": _safe_echo(self.repo),
},
"filters": {key: _safe_echo(value) for key, value in self.filters.items()},
"sources": [s.to_dict() for s in self.sources],
**self.page.to_dict(),
}
def _unanswerable_reasons(
statuses: Iterable[SourceStatus], unanswerable: Iterable[str]
) -> list[dict[str, str]]:
"""Explain, per source, why each unanswerable dimension went unanswered."""
out: list[dict[str, str]] = []
for status in statuses:
for dim in unanswerable:
if dim not in status.supported_filters:
reason = _SOURCE_FILTER_LIMITS.get(
(status.name, dim), f"this source's records carry no {dim} identity"
)
elif not status.ok:
reason = (
f"this source can carry {dim} but did not run: "
f"{status.reason or 'unavailable'}"
)
else:
continue
out.append({"source": status.name, "filter": dim, "reason": reason})
return out
def load_timeline(
*,
remote: str,
org: str,
repo: str,
issue: int | None = None,
pr: int | None = None,
session: str | None = None,
limit: int | None = None,
offset: int | None = None,
db_path: str | None = None,
comment_source: CommentSource | None = None,
) -> TimelineSnapshot:
"""Aggregate every timeline source into one filtered, paginated snapshot.
Sources are read independently and fail soft: an unavailable source
contributes a ``SourceStatus`` with ``ok=False`` and a reason, and never
collapses the whole timeline. The handoff source only runs when a specific
issue or PR is requested (a handoff comment belongs to one thread) and a
``comment_source`` is available; otherwise it is reported as ``not run``
rather than as an empty-and-healthy source.
A filter dimension that no surviving source can carry a ``session``
filter when the only source that ran is the control plane, whose events
record no session is refused with ``ok=False`` and a structured error
instead of being answered with an empty page.
"""
all_events: list[WorkflowEvent] = []
statuses: list[SourceStatus] = []
cp_events, cp_status = read_cp_events(remote=remote, org=org, repo=repo, db_path=db_path)
all_events.extend(cp_events)
statuses.append(cp_status)
# Gitea handoff comments are thread-scoped: only fetch when the caller
# narrowed to one issue or PR, and only when a source was provided.
handoff_target: tuple[str, int] | None = None
if pr is not None:
handoff_target = ("pr", pr)
elif issue is not None:
handoff_target = ("issue", issue)
if handoff_target is None:
statuses.append(
_handoff_status(
ok=False,
reason="not run: handoff comments are thread-scoped; filter by issue or pr to include them",
)
)
elif comment_source is None:
statuses.append(
_handoff_status(
ok=False,
reason="not run: no comment source configured for this timeline read",
)
)
else:
kind, number = handoff_target
try:
comments = comment_source(kind, number) or []
handoff_events = adapt_cth_comments(comments, kind=kind, number=number)
all_events.extend(handoff_events)
statuses.append(_handoff_status(ok=True, count=len(handoff_events)))
except Exception as exc: # fail soft: a fetch/parse error degrades this source only
statuses.append(_handoff_status(ok=False, reason=f"handoff source failed: {exc}"))
requested = tuple(
name
for name, value in ((FILTER_ISSUE, issue), (FILTER_PR, pr), (FILTER_SESSION, session))
if value is not None
)
statuses = [
replace(
status,
unsupported_filters=tuple(
dim for dim in requested if dim not in status.supported_filters
),
)
for status in statuses
]
filters = {"issue": issue, "pr": pr, "session": session}
# A dimension is answerable only if a source that actually ran can carry it.
# If none can, refuse: an empty page would assert "no such activity", which
# is a claim this timeline is not in a position to make.
answerable: set[str] = set()
for status in statuses:
if status.ok:
answerable.update(status.supported_filters)
unanswerable = tuple(dim for dim in requested if dim not in answerable)
if unanswerable:
return TimelineSnapshot(
schema_version=TIMELINE_SCHEMA_VERSION,
remote=remote,
org=org,
repo=repo,
filters=filters,
page=paginate([], limit=limit, offset=offset),
sources=tuple(statuses),
ok=False,
error={
"code": "filter_not_supported",
"unsupported_filters": list(unanswerable),
"detail": (
"no timeline source that ran can answer "
+ ", ".join(f"'{dim}'" for dim in unanswerable)
+ "; the result is refused rather than returned empty"
),
"sources": _unanswerable_reasons(statuses, unanswerable),
},
)
filtered = filter_events(all_events, issue=issue, pr=pr, session=session)
ordered = sort_events(filtered)
page = paginate(ordered, limit=limit, offset=offset)
return TimelineSnapshot(
schema_version=TIMELINE_SCHEMA_VERSION,
remote=remote,
org=org,
repo=repo,
filters=filters,
page=page,
sources=tuple(statuses),
)
def snapshot_to_dict(snapshot: TimelineSnapshot) -> dict[str, Any]:
return snapshot.to_dict()
+271 -4
View File
@@ -34,7 +34,11 @@ import subprocess
from datetime import datetime, timezone
from typing import Any
from merged_cleanup_reconcile import branch_worktree_folder, read_local_worktree_state
from merged_cleanup_reconcile import (
branch_worktree_folder,
is_head_ancestor_of_ref,
read_local_worktree_state,
)
from reviewer_worktree import parse_dirty_tracked_files, REVIEW_WORKTREE_RE
PROTECTED_BRANCHES = frozenset({"master", "main", "dev"})
@@ -67,6 +71,14 @@ REMOVABLE_CLASSES = frozenset(
{CLASS_CLEAN_STALE_REMOVABLE, CLASS_DETACHED_REVIEW_LEFTOVER}
)
# Merged-PR linkage outcomes for issue worktrees (#858). Only ``LINKAGE_MERGED``
# is ownership proof; every other outcome leaves the worktree protected.
LINKAGE_MERGED = "merged_pr"
LINKAGE_OPEN = "open_pr"
LINKAGE_NONE = "no_owning_pr"
LINKAGE_AMBIGUOUS = "ambiguous"
LINKAGE_UNKNOWN = "unknown"
_ISSUE_REF_RE = re.compile(r"issue-(\d+)", re.IGNORECASE)
_ISSUE_BRANCH_PREFIXES = ("feat/", "fix/", "docs/", "chore/")
@@ -169,6 +181,186 @@ def is_ttl_expired(
return (now_dt - last).total_seconds() > ttl_hours * 3600.0
def build_pr_index(prs: list[dict[str, Any]] | None) -> dict[str, list[dict[str, Any]]]:
"""Index PR records by head branch for deterministic worktree linkage (#858).
Accepts Gitea PR payloads (``head`` as a dict) and pre-flattened records
(``head_branch``/``head_sha``). Records without a usable head branch or
number are dropped rather than guessed at, so a branch is only ever linked
to a PR the caller actually proved.
"""
index: dict[str, list[dict[str, Any]]] = {}
for pr in prs or []:
head = pr.get("head")
if isinstance(head, dict):
head_branch = head.get("ref")
head_sha = head.get("sha")
else:
head_branch = pr.get("head_branch") or (head if isinstance(head, str) else None)
head_sha = pr.get("head_sha")
number = pr.get("number")
if not head_branch or number is None:
continue
try:
pr_number = int(number)
except (TypeError, ValueError):
continue
index.setdefault(str(head_branch).strip(), []).append(
{
"pr_number": pr_number,
"head_branch": str(head_branch).strip(),
"head_sha": head_sha,
"merged": bool(pr.get("merged") or pr.get("merged_at")),
"state": pr.get("state"),
}
)
return index
def resolve_owning_pr(
*,
branch: str | None,
pr_index: dict[str, list[dict[str, Any]]] | None,
) -> dict[str, Any]:
"""Resolve the single PR that owns ``branch``, failing closed when unclear.
Ownership is only ``LINKAGE_MERGED`` when exactly one PR claims the branch
and that PR is merged. Several distinct PRs on one branch is a competing
claim (``LINKAGE_AMBIGUOUS``), and a still-open owner is reported as
``LINKAGE_OPEN`` both keep the worktree protected while still exposing
the PR number the audit resolved.
"""
if pr_index is None:
return {
"status": LINKAGE_UNKNOWN,
"pr_number": None,
"candidate_pr_numbers": [],
"reasons": ["live PR state was not supplied; ownership unproven"],
}
branch_name = (branch or "").strip()
if not branch_name:
return {
"status": LINKAGE_UNKNOWN,
"pr_number": None,
"candidate_pr_numbers": [],
"reasons": ["worktree has no attached branch; ownership unproven"],
}
candidates = list(pr_index.get(branch_name) or [])
numbers = sorted({c["pr_number"] for c in candidates})
if not candidates:
return {
"status": LINKAGE_NONE,
"pr_number": None,
"candidate_pr_numbers": [],
"reasons": [f"no PR claims branch '{branch_name}'"],
}
if len(numbers) > 1:
return {
"status": LINKAGE_AMBIGUOUS,
"pr_number": None,
"candidate_pr_numbers": numbers,
"reasons": [
f"branch '{branch_name}' is claimed by competing PRs {numbers}; "
"ownership is ambiguous"
],
}
owner = candidates[0]
pr_number = owner["pr_number"]
if owner.get("head_branch") != branch_name:
return {
"status": LINKAGE_UNKNOWN,
"pr_number": pr_number,
"candidate_pr_numbers": numbers,
"reasons": [
f"PR #{pr_number} head branch '{owner.get('head_branch')}' does not "
f"match worktree branch '{branch_name}'"
],
}
if not owner.get("merged"):
return {
"status": LINKAGE_OPEN,
"pr_number": pr_number,
"candidate_pr_numbers": numbers,
"pr_head_sha": owner.get("head_sha"),
"reasons": [f"owning PR #{pr_number} is not merged"],
}
return {
"status": LINKAGE_MERGED,
"pr_number": pr_number,
"candidate_pr_numbers": numbers,
"pr_head_sha": owner.get("head_sha"),
"reasons": [],
}
def assess_merged_pr_worktree_cleanup(
*,
linkage: dict[str, Any] | None,
head_sha: str | None,
head_in_master: bool | None,
is_dirty: bool,
has_open_pr: bool,
has_active_lease: bool,
has_active_issue_lock: bool,
is_protected: bool,
has_live_session: bool = False,
) -> dict[str, Any]:
"""Decide whether a merged issue worktree satisfies the full cleanup policy.
Every condition must be independently proven: conclusive merged-PR
ownership, agreement between the worktree branch and the PR head branch,
containment of the worktree head in authoritative master (which is what
proves no unmerged commits remain), absence of any open/competing PR,
lease, issue lock, or live session, a clean tree, and a worktree that is
not the protected control checkout. Anything unknown blocks.
"""
link = linkage or {
"status": LINKAGE_UNKNOWN,
"pr_number": None,
"reasons": ["no linkage assessment supplied"],
}
status = link.get("status")
reasons: list[str] = []
if status != LINKAGE_MERGED:
reasons.extend(
link.get("reasons") or ["owning PR could not be conclusively identified"]
)
if is_protected:
reasons.append("worktree is protected or the stable control checkout")
if is_dirty:
reasons.append("worktree has uncommitted changes")
if has_open_pr:
reasons.append("worktree branch has an open PR")
if has_active_lease:
reasons.append("worktree has an active lease")
if has_active_issue_lock:
reasons.append("an active issue lock references this branch")
if has_live_session:
reasons.append("a live process or session is using this worktree")
if not head_sha:
reasons.append("worktree head sha is unknown")
if head_in_master is None:
reasons.append("containment of the worktree head in master is unknown")
elif not head_in_master:
reasons.append(
"worktree head is not contained in authoritative master "
"(unmerged commits remain)"
)
proven = not reasons
return {
"linkage_status": status,
"pr_number": link.get("pr_number"),
"pr_head_sha": link.get("pr_head_sha"),
"head_in_master": head_in_master,
"proven": proven,
"block_reasons": reasons,
}
def classify_worktree(
*,
workflow_type: str,
@@ -181,6 +373,8 @@ def classify_worktree(
ttl_expired: bool = False,
is_protected: bool = False,
metadata_known: bool = True,
merged_pr_cleanup: dict[str, Any] | None = None,
has_live_session: bool = False,
) -> str:
"""Classify a worktree, safety-first: any preservation signal wins.
@@ -199,6 +393,8 @@ def classify_worktree(
return CLASS_ACTIVE_ISSUE_WORK # never auto-deleted (criterion 8)
if has_active_issue_lock:
return CLASS_ACTIVE_ISSUE_WORK
if has_live_session:
return CLASS_ACTIVE_ISSUE_WORK # a live session still owns this tree
if not metadata_known or workflow_type == WORKFLOW_UNKNOWN:
return CLASS_UNSAFE_UNKNOWN # never auto-deleted without proof
@@ -207,7 +403,15 @@ def classify_worktree(
if is_detached or branch_gone:
return CLASS_DETACHED_REVIEW_LEFTOVER
return CLASS_CLEAN_STALE_REMOVABLE
# issue_work / conflict_fix: only removable once the TTL has expired.
if workflow_type == WORKFLOW_ISSUE_WORK:
# #858: an issue worktree becomes removable only on authoritative
# merged-PR evidence satisfying the whole cleanup policy. Age alone
# never proves the branch landed, so TTL cannot qualify one by itself
# — otherwise a worktree holding unmerged commits would be reclaimed.
if (merged_pr_cleanup or {}).get("proven"):
return CLASS_CLEAN_STALE_REMOVABLE
return CLASS_ACTIVE_ISSUE_WORK
# conflict_fix: only removable once the TTL has expired.
if ttl_expired:
return CLASS_CLEAN_STALE_REMOVABLE
return CLASS_ACTIVE_ISSUE_WORK
@@ -400,6 +604,20 @@ def remove_worktree(project_root: str, path: str) -> dict[str, Any]:
}
def head_contained_in_ref(
project_root: str, head_sha: str | None, ref: str | None
) -> bool | None:
"""Return True when ``head_sha`` is already contained in ``ref``.
Shares :mod:`merged_cleanup_reconcile`'s ancestry check so the audit and
the PR-scoped reconciler agree on what "already landed" means (#858).
Returns None when containment cannot be determined, which fails closed.
"""
if not head_sha or not ref:
return None
return is_head_ancestor_of_ref(project_root, head_sha, ref)
def _is_under_branches(project_root: str, path: str) -> bool:
branches_root = os.path.join(os.path.abspath(project_root), "branches")
return os.path.abspath(path or "").startswith(branches_root + os.sep)
@@ -413,16 +631,30 @@ def audit_branches_directory(
active_issue_branches: set[str] | None = None,
now: datetime | str | None = None,
ttl_hours: float = DEFAULT_TTL_HOURS,
pr_index: dict[str, list[dict[str, Any]]] | None = None,
leased_issue_numbers: set[int] | None = None,
live_session_paths: set[str] | None = None,
master_ref: str | None = None,
) -> dict[str, Any]:
"""Classify every session-owned worktree under ``branches/``.
Read-only: shells out to git for discovery and dirty state, then applies
the pure classifier. Returns per-worktree classifications, counts, the
list of removable candidates, and the ``git worktree list`` proof.
``pr_index`` (see :func:`build_pr_index`) supplies the authoritative PR
ownership used to link issue worktrees to their merged PR (#858).
``master_ref`` is the ref a worktree head must be contained in before it
can be considered landed. Both are optional and their absence only ever
fails closed: without them no issue worktree becomes removable.
"""
open_pr_branches = open_pr_branches or set()
leased_branches = leased_branches or set()
active_issue_branches = active_issue_branches or set()
leased_issue_numbers = leased_issue_numbers or set()
live_session_paths = {
os.path.abspath(p) for p in (live_session_paths or set()) if p
}
worktrees: list[dict[str, Any]] = []
for entry in list_worktrees(project_root):
@@ -433,12 +665,42 @@ def audit_branches_directory(
)
dirty_state = read_worktree_dirty(path)
is_dirty = bool(dirty_state.get("dirty"))
head_sha = entry.get("head")
linkage = resolve_owning_pr(branch=branch, pr_index=pr_index)
metadata = build_worktree_metadata(
path=path, branch=branch, head_sha=entry.get("head")
path=path,
branch=branch,
head_sha=head_sha,
pr_number=linkage.get("pr_number"),
)
has_open_pr = bool(branch) and branch in open_pr_branches
has_active_lease = bool(branch) and branch in leased_branches
# A lease on issue N protects that issue's own work worktree. It must
# not incidentally protect a baseline/review scratch tree that merely
# carries the same issue marker in its name, which would change the
# classification of worktrees this policy does not own.
has_active_lease = (bool(branch) and branch in leased_branches) or (
metadata["workflow_type"] == WORKFLOW_ISSUE_WORK
and metadata.get("issue_number") is not None
and metadata["issue_number"] in leased_issue_numbers
)
has_active_lock = bool(branch) and branch in active_issue_branches
has_live_session = bool(path) and os.path.abspath(path) in live_session_paths
head_in_master = (
head_contained_in_ref(project_root, head_sha, master_ref)
if master_ref
else None
)
merged_pr_cleanup = assess_merged_pr_worktree_cleanup(
linkage=linkage,
head_sha=head_sha,
head_in_master=head_in_master,
is_dirty=is_dirty,
has_open_pr=has_open_pr,
has_active_lease=has_active_lease,
has_active_issue_lock=has_active_lock,
is_protected=is_protected,
has_live_session=has_live_session,
)
ttl_expired = is_ttl_expired(
last_used_at=metadata.get("last_used_at"), now=now, ttl_hours=ttl_hours
)
@@ -452,6 +714,8 @@ def audit_branches_directory(
branch_gone=branch is None and not entry.get("detached"),
ttl_expired=ttl_expired,
is_protected=is_protected,
merged_pr_cleanup=merged_pr_cleanup,
has_live_session=has_live_session,
)
metadata["cleanup_eligibility"] = classification
worktrees.append(
@@ -463,7 +727,10 @@ def audit_branches_directory(
"has_open_pr": has_open_pr,
"has_active_lease": has_active_lease,
"has_active_issue_lock": has_active_lock,
"has_live_session": has_live_session,
"is_protected": is_protected,
"merged_pr_linkage": linkage,
"merged_pr_cleanup": merged_pr_cleanup,
"classification": classification,
"removable": is_removable(classification),
}