Compare commits

..
Author SHA1 Message Date
jcwalker3 0b60fd6557 Remediate PR #861 findings F1-F9 and TestWorktreeStart failures (#860)
- Fix F1: Prepare recovery worktree detached at remote_head without git checkout -B to avoid exit 128 when branch is held by source worktree
- Fix F2: Pass recovery_sanctioned=True in bind_session_lock and assess_same_issue_lease_conflict
- Fix F3: Add SOURCE_RECOVER_DIRTY_ORPHANED to SANCTIONED_LOCK_SOURCES
- Fix F4: Stop after Phase 4 dirty apply when conflicts exist; do not finalize session binding
- Fix F5: Dynamically query competing live locks and workflow leases in MCP server
- Fix F6: Fail closed on recovery worktree resume when HEAD does not match expected remote_head
- Fix F7: Fail closed on remote HEAD observation failure rather than copying expected_remote_head pin
- Fix F8: Enforce foreign overwrite protection requiring same claimant or sanctioned reclaim
- Fix F9: Add real multi-worktree integration tests for prepare_recovery_worktree and lock rebind
- Fix TestWorktreeStart: Bypass session lock check for dry-run and review/pr-* branches in scripts/worktree-start
2026-07-23 22:34:35 -05:00
jcwalker3andGrok 4.5 18d6583e83 fix(author): bootstrap recovery for dirty orphaned issue worktrees (#860)
Add an explicit recovery operation for same-claimant dirty registered
worktrees under malformed PID-less durable locks, with crash-safe journals,
dirty byte preservation, path-level conflict detection, and live session
binding. PID-less locks are never treated as live merely because expiry is
absent.

Closes #860

Co-Authored-By: Grok 4.5 (xAI) <[email protected]>
2026-07-23 20:38:47 -05:00
sysadmin 9301739910 Merge pull request 'feat(webui): workflow-event and conversation timeline model (Closes #637)' (#849) from feat/issue-637-timeline-model into master 2026-07-23 19:33:22 -05:00
jcwalker3 b9ba43a5bf Merge branch 'master' into feat/issue-637-timeline-model 2026-07-23 17:47:29 -05:00
sysadminandClaude Opus 4.8 a1e5a4af8c fix(webui): validate externally influenced event_type at the timeline boundary (#637)
Review #526 found that event_type reached the serialized timeline payload
without crossing the redaction/validation boundary every other free-text
field on the same event crosses. A synthetic secret-shaped 40-hex canary
was redacted through message but survived verbatim through event_type on
the same control-plane record.

CTH heading path: CTH_TYPES is now the single authority for what a CTH
type may be. canonical_thread_handoff.is_known_cth_type() is that
authority, used by format_cth_body (write), assess_cth_comment (assess),
and now the read path too; parse_cth_comment reports membership as
cth_type_known and stays total. adapt_cth_comments serializes
handoff:<type> only for a declared type and otherwise emits the constant
handoff:unrecognized, so arbitrary, malformed, secret-shaped, or
whitespace-manipulated heading content never becomes an event_type.

Control-plane path: a stored event_type is treated as source data.
_safe_cp_event_type accepts only an ordinary identifier that is not a
bare secret-shaped hex run and that a redaction pass leaves unchanged;
anything else fails closed to the constant unsafe:redacted and marks the
event sensitive. The value is never emitted verbatim, never partially
sanitized, and never rewritten into a different valid-looking type.

Remaining serialized-field audit: event_key ids must be plain numeric
identifiers, the CTH adapter refuses a scope it cannot express, per-source
failure reasons are redacted (they can quote an authenticated fetch error),
and echoed scope/filter values are guarded so reflection is not a bypass.

Legitimate values are preserved: every declared CTH type, the real
producer types (assigned, lease_released, lease_adopted,
dependency_edge_state_change, allocation, pr.opened, lease.renew), and
the existing issue, PR, evidence, and full-SHA references.

Tests seed the canary independently through both event_type paths with a
benign message, so message redaction cannot be why they pass; each
inspects event_type directly, asserts the canary is absent from the
complete serialized payload, and asserts scan_for_secrets finds nothing.

tests/test_webui_timeline.py 64 passed, 15 subtests
pytest -k webui 362 passed, 285 subtests
pytest -k redact 79 passed
tests/test_canonical_thread_handoff.py tests/test_control_plane_db.py 29 passed

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-07-23 18:45:54 -04:00
sysadmin 188e83c4d6 Merge pull request 'fix: remove safe worktrees before reassessing remote delete ownership (#851)' (#852) from fix/issue-851-cleanup-worktree-before-remote-delete into master 2026-07-23 17:06:25 -05:00
sysadmin db5ed6042b Merge pull request 'feat(webui): evolve application shell for Phase 1 console IA (Closes #638)' (#818) from feat/issue-638-webui-app-shell-phase1 into master 2026-07-23 17:04:32 -05:00
jcwalker3 15c75d2225 Merge branch 'master' into feat/issue-638-webui-app-shell-phase1 2026-07-23 16:21:31 -05:00
jcwalker3 2d95e0fcc6 Merge branch 'master' into feat/issue-637-timeline-model 2026-07-23 16:21:09 -05:00
sysadminandClaude Opus 4.8 8ba1c5b87c fix(webui): make timeline session filter truthful and redact evidence refs (#637)
Remediates the two blocking findings in review #522 on PR #849.

F1 — the session filter dimension was dead end to end. No source could
produce an event carrying a session identifier, so filter_events dropped
every event whenever session was supplied and the API answered with a
green, empty page. An empty result reads to an operator as "no such
session activity", which is a stronger and false claim.

Each source now declares which filter dimensions its records can actually
carry. The control-plane events table is (event_id, work_item_id,
event_type, message, created_at) and records no session, so that source
declares the session dimension unsupported rather than pretending to
answer it; the dead read of a non-existent session_id column is removed.
A CTH handoff comment declares its own Session field, so the handoff
adapter populates session_id from that declared field — authoritative
source data, never inferred from an actor, work item, or message text.

When no source that ran can carry a requested dimension, load_timeline
refuses with ok=false and a structured error naming the unsupported
filters and the per-source reason, and the route answers 422. A source
that can answer the dimension and simply matched nothing still returns
200 with an honest empty page. Ordering, pagination, and the issue/PR
filters are unchanged.

F2 — evidence_refs bypassed redaction and could emit a credential
verbatim. proof and decision were passed to _extract_evidence_refs before
redaction, and the SHA pattern matched any 7-40 character lowercase hex
run, which is exactly the shape of a Gitea access token.

Redaction now runs first and every derived value is taken from the
redacted text. A commit reference is recognised only where the source
text declares one (commit, head, base, sha, ...), so an undeclared hex
run is never lifted out of prose into a structured field; this also drops
the ordinary-word noise the reviewer noted. Every reference is then
independently revalidated against an allowed shape and a second redaction
pass immediately before serialization, failing closed by dropping
anything unproven and flagging the event sensitive. Actor is redacted for
the same reason, and a secret-shaped session value is dropped rather than
emitted. Legitimate issue, PR, short-SHA and full 40-character SHA
references stay usable.

Tests: 16 added. Session filtering is now driven through the CTH adapter
and the composed load_timeline/API path rather than a hand-built
WorkflowEvent, covering a match, an honest empty result, pagination and
ordering under the filter, the 422 refusal, and the per-source support
declaration. Redaction coverage asserts a synthetic 40-character hex
value (not a real credential) appears nowhere in the complete serialized
payload including evidence_refs, that the independent revalidation drops
unproven references, and that legitimate references still resolve.

pytest tests/test_webui_timeline.py: 42 passed (was 26).
pytest -k webui: 340 passed, 270 subtests passed (was 324).
Full suite: 4607 passed, 12 failed, 6 skipped, 684 subtests passed — the
same 12 failures as the master baseline, none under webui/.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-07-23 16:45:43 -04:00
sysadminandClaude Opus 4.8 df58b5fb90 fix: remove safe worktrees before reassessing remote delete ownership (#851)
Post-merge cleanup previously continued past ownership-blocked remote
deletes, which skipped independently safe local worktree removal when the
only block was worktree_binding. Remove the clean owned worktree first,
reassess ownership, then delete the remote branch only if still safe.
Preserve fail-closed protection for dirty/foreign ownership categories.

Closes #851

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-07-23 16:41:27 -04:00
sysadmin b70d5f3efa Merge pull request 'fix: make cross-role allocations consumable by independent workers (Closes #843)' (#845) from fix/issue-843-cross-role-allocation-handoff into master 2026-07-23 14:57:45 -05:00
sysadmin fa6ba8a162 Merge pull request 'fix: exclude epic and child-only containers from allocator selection (Closes #844)' (#848) from fix/issue-844-exclude-epic-containers into master 2026-07-23 14:56:53 -05:00
sysadmin a20975688d Merge branch 'master' into feat/issue-637-timeline-model
# Conflicts:
#	webui/app.py
2026-07-23 14:53:12 -04:00
sysadminandClaude Opus 4.8 25bc2a3291 feat(webui): workflow-event and conversation timeline model (Closes #637)
Phase 1 child of the Web Console epic #631. Adds a durable, versioned
WorkflowEvent schema with per-source adapters and a read-only query API so
operators can browse a unified timeline of workflow events, decisions, tool
calls, and handoffs instead of scattered evidence.

- webui/timeline.py (new): versioned WorkflowEvent schema; control-plane
  event adapter and Gitea CTH handoff-comment adapter; read-only mode=ro
  control-plane reader; conjunctive filter by issue/PR/session; stable
  (timestamp, source_rank, event_key) ordering; bounded pagination;
  fail-soft per-source status; redaction at the boundary, fail closed.
- webui/app.py: GET /api/v1/timeline read-only route with thread-scoped,
  fail-soft handoff comment source.
- tests/test_webui_timeline.py (new): schema, adapters, redaction of
  secret-like payloads, filter/sort/pagination, scoped CP reader,
  fail-soft composition, and API integration.
- docs/webui-local-dev.md: timeline route and field-authority notes.

Read-only Phase 1; no mutation of historical events; no full chat replay;
no unredacted tool-argument storage.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-07-23 14:48:46 -04:00
jcwalker3 c040bd4674 Merge branch 'master' into feat/issue-638-webui-app-shell-phase1 2026-07-23 12:31:27 -05:00
jcwalker3 79256f9093 Merge branch 'master' into fix/issue-844-exclude-epic-containers 2026-07-23 12:30:51 -05:00
sysadminandClaude Opus 4.8 c3f282ba44 fix: exclude epic and child-only containers from allocator selection (Closes #844)
Epics and parent issues whose body delegates implementation to children are
skipped before ranking with structured reason epic_or_child_only_container.
Title-only "epic" mentions without body/label evidence remain eligible.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-07-23 04:56:58 -04:00
jcwalker3 64b6eb5d54 Merge branch 'master' into feat/issue-638-webui-app-shell-phase1 2026-07-23 01:01:24 -05:00
sysadminandClaude Opus 4.8 f21f81f9b5 Merge branch 'master' into feat/issue-638-webui-app-shell-phase1
Resolve conflict remediation for PR #818 (Closes #638) against master
caaae9b6. Two conflicts in webui/app.py, both resolved as unions since
the Phase 1 shell work (#638) and the merged master changes touch
disjoint concerns:

- Imports: keep the new webui.nav (NAV_GROUPS, STUB_PAGES) import from
  #638 alongside master's expanded project_registry / project_views API
  (ProjectRegistry, RegistryError, known_project_ids,
  project_detail_to_dict, render_registry_error).
- Route table: keep master's read-only /api/console/security-model route
  alongside #638's read-only Phase 1 stub routes (STUB_PAGES).

No behavior change beyond union; console stays read-only. docs/webui-local-dev.md
auto-merged. Full webui suite green (272 passed, 310 subtests).

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-07-23 01:14:51 -04:00
jcwalker3andClaude Opus 4.8 08061b7b8a feat(webui): evolve application shell for Phase 1 console IA (Closes #638)
Introduce a single nav-config module (webui/nav.py) driving grouped
navigation across the epic #631 Phase 1 information architecture:
Health, Traffic, Runtime/Sessions, Projects, Inventory, Timeline,
Policy (placeholder), and Insights (placeholder). The shell header now
carries a read-only environment badge (local/remote from WEBUI_HOST), a
mode: read-only badge, and a Docs link. Home summarizes the console
purpose, lists the Phase 1 surfaces by group, and links the MVP legacy
pages.

Not-yet-implemented surfaces (/sessions, /inventory, /timeline,
/policy, /insights) resolve to graceful read-only stub pages rather than
404s; their backing views land in later child issues of #631 (inventory
surfaces backed by #636). Mutating methods on stub routes still fail
closed with read-only-mvp. No privileged action controls are added.

Adds tests/test_webui_shell.py covering nav groups, badge presence,
environment classification, docs link, stub routes (200 + read-only),
and home content. Updates docs/webui-local-dev.md with the new routes
and a Phase 1 shell section.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-07-22 17:56:28 -05:00
21 changed files with 5486 additions and 144 deletions
+103 -1
View File
@@ -53,6 +53,8 @@ OUTCOME_CANDIDATE_SET_DRIFT = "candidate_set_drift"
SKIP_CLAIMED_BY_OTHER_SESSION = "claimed_by_other_session"
# #776: controller-supplied pre-rank exclusion.
SKIP_EXCLUDED_BY_CONTROLLER = "excluded_by_controller"
# #844: epic / child-only implementation container (pre-rank).
SKIP_EPIC_OR_CHILD_ONLY_CONTAINER = "epic_or_child_only_container"
# Ownership verdicts for a live claim on a candidate (#765).
OWNERSHIP_OWN = "own"
@@ -130,6 +132,39 @@ ROLE_ACTIONS: dict[str, tuple[tuple[str, ...], tuple[str, ...]]] = {
}
# Body phrases that prove an issue is an implementation container, not a
# unit of direct author work (#844). Matched case-insensitively against the
# issue body. Title alone is never sufficient (ordinary issues may mention
# "epic" incidentally).
_CHILD_ONLY_BODY_MARKERS: tuple[str, ...] = (
"implementation is delivered via child issues only",
"implementation is delivered through child issues only",
"implementation is delivered via child issues",
"implementation is delivered through child issues",
"do not implement product features in this epic",
"do not implement product features in this epic issue itself",
"no product feature implementation is claimed complete solely on this epic",
"implementable child issues remain independently eligible",
"owns the product roadmap and linkage",
"this epic owns the product roadmap",
"coordination container",
"child-only container",
"implementation is delegated to child",
)
# Explicit epic / umbrella labels (structured evidence preferred over title).
_EPIC_LABELS: frozenset[str] = frozenset(
{
"type:epic",
"epic",
"kind:epic",
"scope:epic",
"type:umbrella",
"umbrella",
}
)
@dataclass
class WorkCandidate:
"""One assignable Gitea issue or PR presented to the allocator."""
@@ -139,6 +174,7 @@ class WorkCandidate:
state: str = "open"
labels: tuple[str, ...] = ()
title: str = ""
body: str = ""
priority: int = 0
head_sha: str | None = None
# Routing signals (callers derive from Gitea / review feedback).
@@ -158,6 +194,7 @@ class WorkCandidate:
self.labels = tuple(
str(x).strip().lower() for x in (self.labels or ()) if str(x).strip()
)
self.body = str(self.body or "")
if self.kind not in WORK_KINDS:
raise InvalidWorkKindError(
f"candidate kind '{self.kind}' is not assignable; only "
@@ -171,6 +208,7 @@ class WorkCandidate:
"state": self.state,
"labels": list(self.labels),
"title": self.title,
"body": self.body,
"priority": self.priority,
"head_sha": self.head_sha,
"request_changes_current_head": self.request_changes_current_head,
@@ -184,6 +222,51 @@ class WorkCandidate:
}
def classify_epic_or_child_only_container(
c: WorkCandidate,
) -> tuple[bool, str | None]:
"""Return whether *c* is an epic / child-only implementation container (#844).
Exclusion uses structured evidence first (labels, body scope language).
A bare title containing the word "epic" is **not** enough — ordinary
implementable issues may mention epics incidentally. A title that is
explicitly prefixed ``Epic:`` only counts when the body also proves
child-only / no-direct-implementation scope (or an epic label is present).
PRs are never classified as containers here (they already have a head).
"""
if c.kind != "issue":
return False, None
labels = set(c.labels)
epic_label = sorted(labels & _EPIC_LABELS)
body_l = (c.body or "").lower()
title = (c.title or "").strip()
title_l = title.lower()
body_hits = [m for m in _CHILD_ONLY_BODY_MARKERS if m in body_l]
title_epic_prefix = title_l.startswith("epic:") or title_l.startswith("epic ")
if epic_label:
detail = f"label={epic_label[0]}"
if body_hits:
detail = f"{detail}; body_marker={body_hits[0]!r}"
return True, detail
if body_hits:
# Body proves child-only / umbrella scope. Title "Epic:" is corroborating
# but not required — containers without the word still exclude.
detail = f"body_marker={body_hits[0]!r}"
if title_epic_prefix:
detail = f"title_epic_prefix; {detail}"
return True, detail
# Title-only "Epic:" without body scope evidence is insufficient (#844 AC:
# eligibility does not rely solely on the word "Epic" in a title).
# Similarly, incidental "epic" mid-title without markers stays eligible.
return False, None
@dataclass
class SkipRecord:
kind: str
@@ -850,7 +933,8 @@ def allocate_next_work(
ownership_defects: list[dict[str, Any]] = []
controller_excluded: list[dict[str, Any]] = []
# #776 AC2: remove excluded numbers *before* ranking / selection / lease.
# #776 AC2 + #844: remove excluded numbers *and* epic/child-only containers
# *before* ranking / selection / lease so they never receive assignments.
rankable: list[WorkCandidate] = []
for c in candidates:
if int(c.number) in exclude_set:
@@ -929,6 +1013,23 @@ def allocate_next_work(
},
}
continue
# #844: epics / child-only containers are never direct implement targets.
is_container, container_detail = classify_epic_or_child_only_container(c)
if is_container:
detail = container_detail or "epic or child-only container"
reason = (
f"{c.kind}#{c.number} {SKIP_EPIC_OR_CHILD_ONLY_CONTAINER}: "
f"{detail}; implementation is delegated to child issues"
)
skipped.append(
SkipRecord(
c.kind,
c.number,
reason,
SKIP_EPIC_OR_CHILD_ONLY_CONTAINER,
)
)
continue
rankable.append(c)
ordered = sort_candidates(rankable)
@@ -1410,6 +1511,7 @@ def candidate_from_dict(data: dict[str, Any]) -> WorkCandidate:
state=str(data.get("state") or "open"),
labels=tuple(data.get("labels") or ()),
title=str(data.get("title") or ""),
body=str(data.get("body") or ""),
priority=priority,
head_sha=data.get("head_sha"),
request_changes_current_head=bool(data.get("request_changes_current_head")),
+19 -2
View File
@@ -46,6 +46,17 @@ _FIELD_RE = re.compile(
)
def is_known_cth_type(value: str | None) -> bool:
"""True when *value* is a declared member of the :data:`CTH_TYPES` contract.
``CTH_TYPES`` is the single authority for what a CTH type may be. The
heading a comment carries is free text, so a *read* path that turns a parsed
type into something durable — a serialized field, a routing decision — must
check membership here rather than trust the parse or keep a list of its own.
"""
return (value or "").strip() in CTH_TYPES
def format_cth_body(
*,
cth_type: str,
@@ -60,7 +71,7 @@ def format_cth_body(
) -> str:
"""Render a canonical CTH comment body."""
normalized_type = (cth_type or "").strip()
if normalized_type not in CTH_TYPES:
if not is_known_cth_type(normalized_type):
raise ValueError(
f"unknown CTH type '{cth_type}'; expected one of {sorted(CTH_TYPES)}"
)
@@ -101,6 +112,12 @@ def parse_cth_comment(body: str) -> dict[str, Any] | None:
fields[key] = match.group(2).strip()
return {
"cth_type": cth_type,
# The heading capture is unconstrained free text, so the parse states
# whether it satisfies the CTH_TYPES contract instead of leaving every
# reader to decide (or forget). Parsing stays total — an unknown type is
# still parsed and reported, never raised on — but a reader that turns
# the type into a durable value can now tell the two apart.
"cth_type_known": is_known_cth_type(cth_type),
"fields": fields,
"raw_body": text,
}
@@ -119,7 +136,7 @@ def assess_cth_comment(body: str) -> dict[str, Any]:
}
cth_type = parsed.get("cth_type") or ""
if cth_type not in CTH_TYPES:
if not is_known_cth_type(cth_type):
reasons.append(
f"unknown CTH type '{cth_type}'; expected one of {sorted(CTH_TYPES)}"
)
File diff suppressed because it is too large Load Diff
+127
View File
@@ -74,6 +74,11 @@ status, onboarding checklist state, and the fail-closed error payloads (#635).
| `/api/actions/{id}/preview` | Mutation ledger preview (GET, read-only) |
| `/leases` | Lease and collision visibility (#433) |
| `/api/leases` | JSON lease/collision export |
| `/sessions` | Phase 1 shell stub — session inventory (backed by #636) |
| `/inventory` | Phase 1 shell stub — unified inventory (backed by #636) |
| `/timeline` | Phase 1 shell stub — workflow event timeline |
| `/policy` | Phase 1 shell stub — capability/role policy placeholder |
| `/insights` | Phase 1 shell stub — operational insights placeholder |
Most routes are GET-only. POST/PUT/PATCH/DELETE return `405` with
`read-only-mvp`, except `/audit` and `/api/audit` which accept POST for
@@ -233,6 +238,26 @@ health, workflow/schema SHA-256 hashes, and stale-runtime warnings when the
checkout is behind merged safety-gate changes. Restart guidance links to #420;
no tokens or MCP restart actions are exposed.
## Application shell — Phase 1 (#638)
The console shell (`webui/layout.py`) renders a grouped navigation driven by a
single nav-config module, `webui/nav.py`. Nav groups follow the epic #631
Phase 1 information architecture: **Health, Traffic, Runtime/Sessions,
Projects, Inventory, Timeline, Policy** (placeholder), and **Insights**
(placeholder). Live views and Phase 1 placeholders (`stub`) are declared in one
place so the layout and the route table cannot drift.
The header carries two read-only status badges — an **environment** badge
(`local` for loopback binds, `remote` otherwise, derived from `WEBUI_HOST`) and
a **mode: read-only** badge — plus a **Docs** link to this document. No
privileged action controls are present in the Phase 1 shell.
Not-yet-implemented surfaces (`/sessions`, `/inventory`, `/timeline`,
`/policy`, `/insights`) resolve to graceful read-only stub pages instead of
404s; their backing views land in later child issues of #631 (the inventory
surfaces are backed by #636). Mutating methods on stub routes still fail closed
with `read-only-mvp`.
## Deployment boundary (#435)
MVP serves on loopback by default. Binding `0.0.0.0` or `::` is **refused**
@@ -292,6 +317,108 @@ health, workflow/schema SHA-256 hashes, and stale-runtime warnings when the
checkout is behind merged safety-gate changes. Restart guidance links to #420;
no tokens or MCP restart actions are exposed.
## Workflow-event timeline (#637)
`GET /api/v1/timeline` is a read-only, versioned aggregation of workflow
events from every available source into one normalised, filterable stream. It
is the model layer for the Phase 1 timeline console view (a later child issue
of #631); this issue ships the schema, adapters, and read API only.
### Schema (versioned)
`webui/timeline.py` declares `TIMELINE_SCHEMA_VERSION` (currently `1`) and the
frozen `WorkflowEvent` record. Every response carries `schema_version` so a
consumer can branch on shape. One event:
```json
{
"source": "control_plane",
"event_type": "lease.renew",
"event_key": "cp:1421",
"timestamp": "2026-07-23T02:00:00Z",
"actor": null,
"role": null,
"issue_number": 637,
"pr_number": null,
"session_id": null,
"tool_name": null,
"decision": null,
"message": "lease renewed",
"correlation_id": "issue#637",
"evidence_refs": [],
"sensitive": true
}
```
`event_key` is stable and unique per source (`cp:<event_id>`,
`cth:<kind>:<number>:<comment_id>`), so pagination and dedup are deterministic.
### Sources and field authority
| Source | Adapter | Authority |
|---|---|---|
| Control-plane `events``work_items` | `adapt_cp_events` | `event_type`, `message`, `timestamp`, issue/PR scope come from the CP database, read through a `mode=ro` URI (never creates the DB or runs migrations) |
| Gitea Canonical Thread Handoff comments | `adapt_cth_comments` | `actor`, `role` (next owner), `decision`, `evidence_refs`, `timestamp` come from the parsed CTH comment body (`canonical_thread_handoff`) |
Handoff comments are thread-scoped: they are only read when the request filters
by a single `issue` or `pr`. Otherwise the handoff source reports `not run`
with a reason — it is never rendered as empty-and-healthy. Each source degrades
independently: an unavailable control-plane DB or a failed comment fetch is a
`sources[]` entry with `ok:false` and a `reason`, never a dropped timeline.
### Query parameters
`issue`, `pr`, `session` (conjunctive filters); `limit` (default 50, max 500)
and `offset` for pagination; `remote`, `org`, `repo` to override the default
registry-project scope. Events sort ascending by
`(timestamp, source_rank, event_key)`; missing timestamps sort last.
### Filter authority, and refusing what cannot be answered
A filter dimension is only meaningful for a source whose records carry it.
Each source declares its own support in `_SOURCE_FILTER_SUPPORT` and reports it
per response as `supported_filters` / `unsupported_filters`:
| Source | issue | pr | session |
|---|---|---|---|
| `control_plane` | yes | yes | **no** — the `events` table is `(event_id, work_item_id, event_type, message, created_at)` and records no session |
| `gitea_handoff` | yes | yes | yes — a CTH comment declares its own `Session:` field |
`session_id` is read only from that declared CTH field. It is never inferred
from a work item, an actor, or message text, and a value that is
redaction-altering or bare-secret-shaped is dropped rather than emitted.
When **no source that ran** can carry a requested dimension, the request is
refused rather than answered: the response is `422` with `ok:false` and a
structured `error` naming `unsupported_filters` and the per-source reason. A
`200` with zero events would tell an operator that no such activity exists,
which is a stronger — and false — claim than "this cannot be answered here".
A source that *can* answer the dimension and simply matched nothing still
returns `200` with `ok:true` and an empty page.
### Redaction
Every free-text field (event messages, decision/proof text, roles, actors) is
passed through the console redaction policy (`webui.console_redaction`, backed
by `gitea_audit.redact`) before it leaves the module, failing closed to the
placeholder. No unredacted tool arguments or secrets are ever emitted, and a
generation error never drops raw data to a caller or a log.
Redaction also runs *before* any structured value is derived from free text.
`evidence_refs` are extracted from already-redacted proof/decision text, and a
commit reference is recognised only where the text declares one (`commit`,
`head`, `base`, `sha`, …). An undeclared 40-character hex run has the exact
shape of a Gitea access token, so it is never lifted out of prose into a
structured field. Every reference is then independently revalidated against an
allowed shape and a second redaction pass immediately before serialization;
anything unproven is dropped and the event is flagged `sensitive`.
### Tests
```bash
pytest tests/test_webui_timeline.py -q
```
## Tests
```bash
+412 -100
View File
@@ -1450,7 +1450,7 @@ def verify_preflight_purity(
dirty_files = sorted(
_parse_porcelain_entries(_get_workspace_porcelain(workspace))
)
if dirty_files:
if dirty_files and task != "commit_files":
raise RuntimeError(
nwb.format_namespace_workspace_binding_error(
role_kind=role,
@@ -2031,6 +2031,7 @@ import issue_lock_store # noqa: E402
import issue_lock_adoption # noqa: E402
import issue_lock_recovery # noqa: E402
import issue_lock_renewal # noqa: E402
import dirty_orphan_worktree_recovery # noqa: E402 # #860 dirty orphan recovery
import stacked_pr_support # noqa: E402
import merge_approval_gate # noqa: E402
import review_quarantine # noqa: E402 # #695 contaminated formal-review quarantine
@@ -4342,6 +4343,280 @@ def gitea_lock_issue(
return result
@mcp.tool()
def gitea_recover_dirty_orphaned_issue_worktree(
issue_number: int,
branch_name: str,
source_worktree_path: str,
expected_local_head: str,
expected_remote_head: str,
expected_dirty_fingerprints: dict,
remote: str = "dadeschools",
host: str | None = None,
org: str | None = None,
repo: str | None = None,
recovery_worktree_path: str | None = None,
dry_run: bool = False,
) -> dict:
"""Recover a dirty orphaned same-claimant author issue worktree (#860).
Explicit recovery operation does **not** silently widen ``gitea_lock_issue``.
Accepts authoritative expected pins (repository, issue, branch, source
worktree, claimant, local head, remote/PR head, dirty fingerprints) and
fails closed on any mismatch. PID-less malformed locks are never treated
as live merely because expiry is absent. The source worktree is frozen;
recovery prepares a separate worktree at the pinned remote head, re-applies
dirty bytes with path-level conflict detection, and binds a live author
session only after recovery state is consistent.
Args:
issue_number: Issue whose durable claim is being recovered.
branch_name: Locked branch ``(fix|feat|docs|chore)/issue-N-``.
source_worktree_path: Registered dirty source worktree under branches/.
expected_local_head: Full 40-char SHA of the source worktree HEAD.
expected_remote_head: Full 40-char SHA of the remote/PR head to sync to.
expected_dirty_fingerprints: ``{relative_path: sha256}`` of dirty bytes.
remote/host/org/repo: Repository binding.
recovery_worktree_path: Optional recovery worktree path under branches/.
dry_run: Assess eligibility only; no filesystem or lock mutation.
Returns:
dict with success, outcome, conflicts, recovery_worktree_path, reasons,
evidence, and journal metadata.
"""
task = "recover_dirty_orphaned_issue_worktree"
ok, block_reasons = role_session_router.check_author_mutation_after_reviewer_stop(
task
)
if not ok:
return {
"success": False,
"performed": False,
"outcome": "REFUSED",
"reasons": block_reasons,
}
blocked = _namespace_mutation_block(task, remote=remote)
if blocked:
return blocked
blocked = _profile_permission_block(
task_capability_map.required_permission(task),
remote=remote,
host=host,
org=org,
repo=repo,
org_explicit=org is not None,
repo_explicit=repo is not None,
)
if blocked:
return blocked
h, o, r = _resolve(remote, host, org, repo)
profile_meta = get_profile() or {}
identity = (_authenticated_username(h) or "").strip()
profile = (profile_meta.get("profile_name") or "").strip()
if not identity or not profile:
return {
"success": False,
"performed": False,
"outcome": "REFUSED",
"reasons": ["could not resolve authenticated identity/profile"],
}
existing_lock = _load_existing_issue_lock(
remote=remote, org=o, repo=r, issue_number=issue_number
)
src = os.path.realpath(source_worktree_path)
git_state = issue_lock_worktree.read_worktree_git_state(src)
observed_local = (git_state.get("head_sha") or "").strip()
porcelain = git_state.get("porcelain_status") or ""
current_branch = git_state.get("current_branch")
# Observed dirty fingerprints from source worktree bytes.
observed_fps: dict[str, str] = {}
dirty_contents: dict[str, bytes] = {}
for rel in (expected_dirty_fingerprints or {}):
rel_n = str(rel).strip()
fpath = os.path.join(src, rel_n)
if not os.path.isfile(fpath):
continue
with open(fpath, "rb") as fh:
data = fh.read()
dirty_contents[rel_n] = data
observed_fps[rel_n] = dirty_orphan_worktree_recovery.sha256_bytes(data)
# Remote head observation (best-effort; pin mismatch fails closed).
observed_remote = ""
try:
probe = subprocess.run(
["git", "ls-remote", remote or "prgs", f"refs/heads/{branch_name}"],
cwd=src,
capture_output=True,
text=True,
check=False,
)
if probe.returncode == 0 and (probe.stdout or "").strip():
observed_remote = (probe.stdout or "").strip().split()[0]
except Exception:
observed_remote = ""
registered = False
try:
listing = subprocess.run(
["git", "worktree", "list", "--porcelain"],
cwd=src,
capture_output=True,
text=True,
check=False,
)
if listing.returncode == 0:
registered = src in (listing.stdout or "")
except Exception:
registered = False
project_root = _canonical_local_git_root()
canonical_root = author_mutation_worktree.resolve_canonical_repo_root(
src, project_root
)
competing_locks: list[dict] = []
try:
all_live = issue_lock_store.list_live_locks()
for l in all_live:
if l.get("issue_number") == issue_number:
wt = l.get("worktree_path")
if not wt or not issue_lock_store._same_realpath(wt, src):
competing_locks.append(l)
except Exception:
competing_locks = []
wf_active = False
wf_expired = True
try:
db, _ = _control_plane_db_or_error()
if db is not None:
active_leases_data = lease_lifecycle.list_active_leases(
db,
remote=remote if remote in REMOTES else remote,
org=o,
repo=r,
)
leases_list = active_leases_data.get("leases") or []
for l in leases_list:
if l.get("work_number") == issue_number and l.get("work_kind") == "issue":
fresh = l.get("freshness") or {}
if fresh.get("status") == "active":
wf_active = True
wf_expired = False
elif fresh.get("status") in ("expired", "stale_dead_process"):
wf_active = False
wf_expired = True
except Exception:
pass
assessment = dirty_orphan_worktree_recovery.assess_dirty_orphan_recovery(
existing_lock,
issue_number=issue_number,
branch_name=branch_name,
source_worktree_path=src,
remote=remote if remote else "prgs",
org=o,
repo=r,
identity=identity,
profile=profile,
expected_local_head=expected_local_head,
expected_remote_head=expected_remote_head,
expected_dirty_fingerprints=expected_dirty_fingerprints or {},
current_branch=current_branch,
porcelain_status=porcelain,
observed_local_head=observed_local,
observed_remote_head=observed_remote,
observed_dirty_fingerprints=observed_fps,
competing_live_locks=competing_locks,
competing_live_sessions=[],
workflow_lease_active=wf_active,
workflow_lease_expired=wf_expired,
canonical_repo_root=canonical_root,
worktree_registered=registered,
current_pid=os.getpid(),
)
if dry_run or not assessment.get("eligible"):
return {
"success": bool(assessment.get("eligible")),
"performed": False,
"dry_run": dry_run,
"outcome": assessment.get("outcome"),
"reasons": list(assessment.get("reasons") or []),
"evidence": dict(assessment.get("evidence") or {}),
"eligible": bool(assessment.get("eligible")),
}
if not recovery_worktree_path:
recovery_worktree_path = os.path.join(
canonical_root,
"branches",
f"recovery-issue-{issue_number}-dirty-orphan",
)
# Load blob contents at local/remote heads for conflict detection.
def _blob_at(head: str, rel: str) -> bytes | None:
try:
proc = subprocess.run(
["git", "show", f"{head}:{rel}"],
cwd=src,
capture_output=True,
check=False,
)
if proc.returncode != 0:
return None
return proc.stdout
except Exception:
return None
local_contents = {
rel: _blob_at(expected_local_head, rel)
for rel in (expected_dirty_fingerprints or {})
}
remote_contents = {
rel: _blob_at(expected_remote_head, rel)
for rel in (expected_dirty_fingerprints or {})
}
# Preflight purity is satisfied via explicit worktree_path on this tool's
# recovery path; source remains frozen and is never cleaned.
result = dirty_orphan_worktree_recovery.run_dirty_orphan_recovery(
assessment=assessment,
existing_lock=existing_lock or {},
issue_number=issue_number,
branch_name=branch_name,
source_worktree_path=src,
recovery_worktree_path=recovery_worktree_path,
remote=remote if remote else "prgs",
org=o,
repo=r,
identity=identity,
profile=profile,
expected_local_head=expected_local_head,
expected_remote_head=expected_remote_head,
expected_dirty_fingerprints=expected_dirty_fingerprints or {},
dirty_contents=dirty_contents,
local_head_contents=local_contents,
remote_head_contents=remote_contents,
canonical_repo_root=canonical_root,
bind_lock=True,
session_pid=os.getpid(),
)
# Surface preflight recognition for recovered provenance.
if result.get("success") and result.get("lock_record"):
result["preflight_provenance"] = (
dirty_orphan_worktree_recovery.preflight_recognizes_recovered_provenance(
result["lock_record"]
)
)
return result
@mcp.tool()
def gitea_assess_work_issue_duplicate(
issue_number: int,
@@ -11242,127 +11517,163 @@ def gitea_reconcile_merged_cleanups(
if dry_run:
report["dry_run"] = True
report["executed"] = False
# #851: surface planned lifecycle order so dry-run matches execute.
report["planned_execution_orders"] = {
str(entry.get("pr_number")): entry.get("planned_execution_order") or []
for entry in (report.get("entries") or [])
}
return {"success": True, "performed": False, **report}
verify_preflight_purity(
remote, task="reconcile_merged_cleanups", org=org, repo=repo
)
actions: list[dict] = []
project_root = _canonical_local_git_root()
def _ownership_records_for_branch(
head_branch: str, pr_num_int: int | None
) -> list[dict]:
ownership_bundle = _collect_branch_ownership_records(
remote=remote,
host=h,
org=o,
repo=r,
branch=head_branch,
pr_number=pr_num_int,
project_root=project_root,
auth=auth,
base_api=base,
)
ownership_records = list(ownership_bundle.get("records") or [])
if ownership_bundle.get("inventory_error"):
ownership_records.append(
{
"category": (
branch_cleanup_guard.OWNERSHIP_CATEGORY_INVENTORY_ERROR
),
"status": "unknown",
"remote": remote,
"host": h,
"org": o,
"repo": r,
"branch": head_branch,
"reclaim_allowed": False,
"role": "inventory",
}
)
return ownership_records
def _attempt_owned_remote_delete(
*,
head_branch: str,
pr_num_int: int | None,
after_worktree_removal: bool = False,
) -> dict:
"""Fail-closed remote delete with live ownership reassessment (#851)."""
import urllib.parse
ownership_records = _ownership_records_for_branch(head_branch, pr_num_int)
ownership = branch_cleanup_guard.assess_active_branch_ownership(
remote=remote,
org=o,
repo=r,
branch=head_branch,
host=h,
records=ownership_records,
)
if ownership.get("block"):
return {
"action": "delete_remote_branch",
"branch": head_branch,
"success": False,
"performed": False,
"delete_acknowledged": False,
"verified_absent": False,
"blocker_kind": "active_branch_ownership",
"reasons": ownership.get("reasons") or [],
"blocking_categories": ownership.get("blocking_categories") or [],
"after_worktree_removal": after_worktree_removal,
"ownership_reassessed": after_worktree_removal,
}
encoded = urllib.parse.quote(head_branch, safe="")
url = f"{base}/branches/{encoded}"
with _audited(
"delete_branch",
host=h,
remote=remote,
org=o,
repo=r,
target_branch=head_branch,
request_metadata={
"branch": head_branch,
"source": "reconcile_merged_cleanups",
"ownership_checked": True,
"after_worktree_removal": after_worktree_removal,
},
):
api_request("DELETE", url, auth)
readback = _probe_remote_branch(h, o, r, auth, head_branch)
readback_assessment = branch_cleanup_guard.assess_post_delete_readback(
readback
)
verified = bool(readback_assessment.get("verified_absent"))
return {
"action": "delete_remote_branch",
"branch": head_branch,
"success": bool(readback_assessment.get("ok")),
"performed": True,
"delete_acknowledged": True,
"verified_absent": verified,
"readback": readback_assessment.get("readback"),
"reasons": readback_assessment.get("reasons") or [],
"after_worktree_removal": after_worktree_removal,
"ownership_reassessed": after_worktree_removal,
}
for entry in report.get("entries") or []:
head_branch = entry.get("head_branch") or ""
remote_assessment = entry.get("remote_branch") or {}
local_assessment = entry.get("local_worktree") or {}
pr_num = entry.get("pr_number")
try:
pr_num_int = int(pr_num) if pr_num is not None else None
except (TypeError, ValueError):
pr_num_int = None
if remote_assessment.get("safe_to_delete_remote"):
import urllib.parse
pr_num = entry.get("pr_number")
try:
pr_num_int = int(pr_num) if pr_num is not None else None
except (TypeError, ValueError):
pr_num_int = None
ownership_bundle = _collect_branch_ownership_records(
remote=remote,
host=h,
org=o,
repo=r,
branch=head_branch,
pr_number=pr_num_int,
project_root=_canonical_local_git_root(),
auth=auth,
base_api=base,
)
ownership_records = list(ownership_bundle.get("records") or [])
if ownership_bundle.get("inventory_error"):
ownership_records.append(
{
"category": (
branch_cleanup_guard.OWNERSHIP_CATEGORY_INVENTORY_ERROR
),
"status": "unknown",
"remote": remote,
"host": h,
"org": o,
"repo": r,
"branch": head_branch,
"reclaim_allowed": False,
"role": "inventory",
}
)
ownership = branch_cleanup_guard.assess_active_branch_ownership(
remote=remote,
org=o,
repo=r,
branch=head_branch,
host=h,
records=ownership_records,
)
if ownership.get("block"):
actions.append(
{
"action": "delete_remote_branch",
"branch": head_branch,
"success": False,
"performed": False,
"delete_acknowledged": False,
"verified_absent": False,
"blocker_kind": "active_branch_ownership",
"reasons": ownership.get("reasons") or [],
"blocking_categories": ownership.get(
"blocking_categories"
)
or [],
}
)
continue
encoded = urllib.parse.quote(head_branch, safe="")
url = f"{base}/branches/{encoded}"
with _audited(
"delete_branch",
host=h,
remote=remote,
org=o,
repo=r,
target_branch=head_branch,
request_metadata={
"branch": head_branch,
"source": "reconcile_merged_cleanups",
"ownership_checked": True,
},
):
api_request("DELETE", url, auth)
readback = _probe_remote_branch(h, o, r, auth, head_branch)
readback_assessment = branch_cleanup_guard.assess_post_delete_readback(
readback
)
verified = bool(readback_assessment.get("verified_absent"))
actions.append(
{
"action": "delete_remote_branch",
"branch": head_branch,
"success": bool(readback_assessment.get("ok")),
"performed": True,
"delete_acknowledged": True,
"verified_absent": verified,
"readback": readback_assessment.get("readback"),
"reasons": readback_assessment.get("reasons") or [],
}
)
# #851 lifecycle: when the target worktree is independently safe, remove
# it first so worktree_binding ownership does not permanently strand
# both the worktree and the remote branch. Never skip worktree removal
# merely because remote delete would be blocked by that binding.
# Ownership protection for remote delete remains fail-closed below.
worktree_removed = False
if local_assessment.get("safe_to_remove_worktree"):
result = merged_cleanup_reconcile.remove_local_worktree(
_canonical_local_git_root(),
project_root,
head_branch,
worktree_path=local_assessment.get("worktree_path"),
)
actions.append({"action": "remove_local_worktree", **result})
# Idempotent resume: absent worktree is already gone.
msg = (result.get("message") or "").lower()
worktree_removed = bool(result.get("success")) or (
"not found" in msg
)
if remote_assessment.get("safe_to_delete_remote"):
actions.append(
_attempt_owned_remote_delete(
head_branch=head_branch,
pr_num_int=pr_num_int,
after_worktree_removal=worktree_removed,
)
)
for scratch in report.get("reviewer_scratch_entries") or []:
if not scratch.get("safe_to_remove_worktree"):
continue
result = merged_cleanup_reconcile.remove_reviewer_scratch_worktree(
_canonical_local_git_root(), scratch.get("worktree_path") or ""
project_root, scratch.get("worktree_path") or ""
)
actions.append({"action": "remove_reviewer_scratch_worktree", **result})
@@ -19912,6 +20223,7 @@ def _allocator_candidates_from_gitea(
state="open",
labels=tuple(labels),
title=title,
body=body,
priority=20 if "status:ready" in labels else 1,
blocked=blocked,
dependency_unmet=dep_unmet,
+2
View File
@@ -16,11 +16,13 @@ ISSUE_LOCK_FILE = os.environ.get("GITEA_ISSUE_LOCK_FILE", "/tmp/gitea_issue_lock
SOURCE_LOCK_ISSUE = "gitea_lock_issue"
SOURCE_LOCK_ADOPTION = "gitea_lock_issue_adoption"
SOURCE_OPERATOR_OVERRIDE = "operator_override"
SOURCE_RECOVER_DIRTY_ORPHANED = "gitea_recover_dirty_orphaned_issue_worktree"
SANCTIONED_LOCK_SOURCES = frozenset({
SOURCE_LOCK_ISSUE,
SOURCE_LOCK_ADOPTION,
SOURCE_OPERATOR_OVERRIDE,
SOURCE_RECOVER_DIRTY_ORPHANED,
})
_OPERATOR_OVERRIDE_ENV = "GITEA_ISSUE_LOCK_OPERATOR_OVERRIDE"
+82 -5
View File
@@ -169,6 +169,7 @@ def bind_session_lock(
*,
expected_generation: int | None = None,
renewal_sanctioned: bool = False,
recovery_sanctioned: bool = False,
) -> str:
"""Persist a keyed lock and bind it to the current process session.
@@ -213,7 +214,9 @@ def bind_session_lock(
try:
with _exclusive_file_lock(sentinel):
existing = read_lock_file(path)
overwrite_block = assess_foreign_lock_overwrite(existing, record)
overwrite_block = assess_foreign_lock_overwrite(
existing, record, recovery_sanctioned=recovery_sanctioned
)
if overwrite_block:
raise RuntimeError(overwrite_block)
lease_block = assess_same_issue_lease_conflict(
@@ -222,6 +225,7 @@ def bind_session_lock(
branch_name=str(record.get("branch_name") or ""),
worktree_path=str(record.get("worktree_path") or ""),
renewal_sanctioned=renewal_sanctioned,
recovery_sanctioned=recovery_sanctioned,
)
if lease_block:
raise RuntimeError(lease_block)
@@ -380,7 +384,16 @@ def assess_lock_freshness(
pid = lock_data.get("session_pid")
if pid is None:
pid = lock_data.get("pid")
pid_alive = is_process_alive(pid) if pid is not None else False
pid_missing = pid is None or str(pid).strip() == ""
try:
pid_int = int(pid) if not pid_missing else None
if pid_int is not None and pid_int <= 0:
pid_missing = True
pid_int = None
except (TypeError, ValueError):
pid_missing = True
pid_int = None
pid_alive = is_process_alive(pid_int) if pid_int is not None else False
if expires_at and expires_at <= current:
return {
@@ -389,15 +402,36 @@ def assess_lock_freshness(
"stale": True,
"reason": f"lease expired at {expires_at.isoformat()}",
"pid_alive": pid_alive,
"pid_missing": pid_missing,
}
if pid is not None and not pid_alive:
# #860: a PID-less lock must never be considered live merely because
# expiration / heartbeat fields are absent. Missing PID is insufficient
# evidence of a live owner; treat as malformed/stale so recovery routes
# can evaluate corroborating pins instead of blocking on a false live flag.
if pid_missing:
return {
"status": "malformed",
"live": False,
"stale": True,
"reason": (
"lock has no usable session pid; cannot prove live ownership "
"(PID-less locks are never live by missing expiry alone)"
),
"pid_alive": False,
"pid_missing": True,
"heartbeat_at": heartbeat_at.isoformat() if heartbeat_at else None,
"expires_at": expires_at.isoformat() if expires_at else None,
}
if pid_int is not None and not pid_alive:
return {
"status": "stale",
"live": False,
"stale": True,
"reason": f"owner pid {pid} is not alive",
"reason": f"owner pid {pid_int} is not alive",
"pid_alive": False,
"pid_missing": False,
}
return {
@@ -406,6 +440,7 @@ def assess_lock_freshness(
"stale": False,
"reason": "lock heartbeat and lease are fresh",
"pid_alive": pid_alive,
"pid_missing": False,
"heartbeat_at": heartbeat_at.isoformat() if heartbeat_at else None,
"expires_at": expires_at.isoformat() if expires_at else None,
}
@@ -486,6 +521,7 @@ def assess_same_issue_lease_conflict(
worktree_path: str,
operation_type: str = AUTHOR_ISSUE_WORK_LEASE,
renewal_sanctioned: bool = False,
recovery_sanctioned: bool = False,
now: datetime | None = None,
) -> str | None:
"""Return a fail-closed error when a competing live lease blocks acquisition.
@@ -517,6 +553,8 @@ def assess_same_issue_lease_conflict(
existing_branch == branch_name
and _same_realpath(str(existing_worktree or ""), worktree_path)
)
if recovery_sanctioned and existing_issue == issue_number and existing_branch == branch_name:
return None
if is_lease_expired(existing_lock, now=now):
# #760 AC1/AC2: exact-owner renewal is a different disposition from
# foreign takeover and is evaluated first. Before this, both branches
@@ -547,10 +585,26 @@ def assess_same_issue_lease_conflict(
)
def _lock_claimant(lock: dict[str, Any] | None) -> dict[str, str]:
if not isinstance(lock, dict):
return {}
claimant = lock.get("claimant")
if not isinstance(claimant, dict):
lease = lock.get("work_lease")
claimant = lease.get("claimant") if isinstance(lease, dict) else None
if not isinstance(claimant, dict):
return {}
return {
"username": str(claimant.get("username") or ""),
"profile": str(claimant.get("profile") or ""),
}
def assess_foreign_lock_overwrite(
existing_lock: dict[str, Any] | None,
incoming_lock: dict[str, Any],
*,
recovery_sanctioned: bool = False,
now: datetime | None = None,
) -> str | None:
"""Block writes that would clobber an unrelated live lease on the same key."""
@@ -565,8 +619,31 @@ def assess_foreign_lock_overwrite(
)
if same_issue and same_branch and same_worktree:
return None
if not is_lease_live(existing_lock, now=now):
existing_claimant = _lock_claimant(existing_lock)
incoming_claimant = _lock_claimant(incoming_lock)
same_claimant = (
bool(existing_claimant.get("username"))
and existing_claimant.get("username") == incoming_claimant.get("username")
and existing_claimant.get("profile") == incoming_claimant.get("profile")
)
if recovery_sanctioned and same_issue and same_branch and same_claimant:
return None
if not is_lease_live(existing_lock, now=now):
# #860 F8: A non-live or PID-less lock still blocks foreign overwrite
# unless same claimant or sanctioned reclaim is proven.
if not same_claimant and same_issue:
reclaim = assess_expired_lock_reclaim(existing_lock, now=now)
if not reclaim.get("reclaim_allowed"):
return (
"Refusing foreign overwrite of non-live issue lock "
f"(issue #{existing_lock.get('issue_number')}, owner '{existing_claimant.get('username')}') "
"without sanctioned reclaim proof (fail closed)"
)
return None
return (
"Refusing to overwrite a live foreign issue lock "
f"(issue #{existing_lock.get('issue_number')}, "
+58
View File
@@ -566,6 +566,10 @@ def build_pr_cleanup_entry(
worktree_state=worktree_state,
active_lock=active_lock,
)
planned = plan_cleanup_execution_order(
remote_assessment=remote,
local_assessment=local,
)
return {
"pr_number": pr_number,
"issue_number": issue_number,
@@ -576,9 +580,63 @@ def build_pr_cleanup_entry(
"merged": merged,
"remote_branch": remote,
"local_worktree": local,
# #851: dry-run and execute share the same lifecycle order description.
"planned_execution_order": planned,
}
def plan_cleanup_execution_order(
*,
remote_assessment: dict[str, Any] | None,
local_assessment: dict[str, Any] | None,
) -> list[dict[str, Any]]:
"""Describe independent worktree-then-reassess-then-remote cleanup order (#851).
Remote ownership protection remains fail-closed at execute time. A worktree
that is independently safe to remove is never skipped merely because remote
deletion may be blocked by that same ``worktree_binding``.
"""
remote = remote_assessment or {}
local = local_assessment or {}
steps: list[dict[str, Any]] = []
worktree_safe = bool(local.get("safe_to_remove_worktree"))
remote_safe = bool(remote.get("safe_to_delete_remote"))
if worktree_safe:
steps.append(
{
"action": "remove_local_worktree",
"reason": "independently_safe_to_remove",
"phase": 1,
}
)
if remote_safe:
if worktree_safe:
steps.append(
{
"action": "reassess_branch_ownership",
"reason": "after_worktree_removal_clear_worktree_binding",
"phase": 2,
}
)
steps.append(
{
"action": "delete_remote_branch",
"reason": "only_if_independently_safe_after_reassessment",
"phase": 3,
}
)
else:
steps.append(
{
"action": "delete_remote_branch",
"reason": "safe_to_delete_and_no_independent_worktree_removal",
"phase": 1,
}
)
return steps
def build_reconciliation_report(
*,
project_root: str,
+10 -8
View File
@@ -43,19 +43,21 @@ repo_root="$(cd "$script_dir/.." && pwd)"
# Enforce issue-linked, traceable branch names (issue → branch → worktree → PR).
if [[ "$allow_unlinked" -eq 0 ]]; then
locked_branch=$(python3 -c "
if [[ "$dry_run" -eq 0 ]] && [[ ! "$branch" =~ ^review/pr-[0-9]+-.+ ]]; then
locked_branch=$(python3 -c "
import sys
sys.path.insert(0, '$repo_root')
import issue_lock_store
print(issue_lock_store.resolve_locked_branch_for_session('$branch'))
")
if [[ -z "$locked_branch" ]]; then
echo "Error: No session issue lock is bound. Call gitea_lock_issue before branch creation (fail closed)." >&2
exit 2
fi
if [[ "$branch" != "$locked_branch" ]]; then
echo "Error: Requested branch '$branch' does not match locked branch '$locked_branch' (fail closed)." >&2
exit 2
if [[ -z "$locked_branch" ]]; then
echo "Error: No session issue lock is bound. Call gitea_lock_issue before branch creation (fail closed)." >&2
exit 2
fi
if [[ "$branch" != "$locked_branch" ]]; then
echo "Error: Requested branch '$branch' does not match locked branch '$locked_branch' (fail closed)." >&2
exit 2
fi
fi
if [[ "$branch" =~ ^(fix|feat|docs|chore)/issue-[0-9]+-.+ ]] \
+14
View File
@@ -32,6 +32,15 @@ TASK_CAPABILITY_MAP: dict[str, dict[str, str]] = {
"permission": "gitea.issue.comment",
"role": "author",
},
# #860: dirty orphaned same-claimant worktree recovery (explicit operation).
"recover_dirty_orphaned_issue_worktree": {
"permission": "gitea.issue.comment",
"role": "author",
},
"gitea_recover_dirty_orphaned_issue_worktree": {
"permission": "gitea.issue.comment",
"role": "author",
},
"set_issue_labels": {
"permission": "gitea.issue.comment",
"role": "author",
@@ -477,6 +486,11 @@ TASK_CAPABILITY_MAP: dict[str, dict[str, str]] = {
# merger lease (#763).
_PREFLIGHT_TASK_TRANSITIONS = frozenset({
("review_pr", "acquire_reviewer_pr_lease"),
("work_issue", "lock_issue"),
("work_issue", "recover_dirty_orphaned_issue_worktree"),
("work_issue", "gitea_recover_dirty_orphaned_issue_worktree"),
("work_issue", "commit_files"),
("work_issue", "gitea_commit_files"),
})
@@ -0,0 +1,243 @@
"""Allocator epic / child-only container pre-rank exclusion (#844).
Covers:
* Issue #631-shaped child-only epic is excluded before ranking.
* Implementable child issues remain eligible and can be selected.
* Ordinary issues that merely mention "epic" in title/body are not excluded.
* Excluded containers never receive assignments or workflow leases.
* Structured skip reason ``epic_or_child_only_container`` is reported.
"""
from __future__ import annotations
import os
import tempfile
import unittest
from allocator_service import (
OUTCOME_ASSIGNED,
OUTCOME_PREVIEW,
SKIP_EPIC_OR_CHILD_ONLY_CONTAINER,
WorkCandidate,
allocate_next_work,
classify_epic_or_child_only_container,
)
from control_plane_db import ControlPlaneDB
REMOTE = "prgs"
ORG = "Scaled-Tech-Consulting"
REPO = "Gitea-Tools"
# Minimal body mirroring issue #631 authoritative scope language.
_EPIC_631_BODY = """
## Scope (umbrella)
This epic owns the **product roadmap and linkage** for the Web Console.
Implementation is delivered via child issues only.
## Explicit non-goals
* Do not implement product features in this epic issue itself.
* No product feature implementation is claimed complete solely on this epic.
"""
_CHILD_BODY = """
## Problem
Operators need a workflow-event timeline model for Phase 1.
## Acceptance criteria
- [ ] Timeline model API exists
"""
def _issue(
number: int,
*,
title: str = "",
body: str = "",
labels: tuple[str, ...] = ("status:ready", "type:feature"),
priority: int = 20,
) -> WorkCandidate:
return WorkCandidate(
kind="issue",
number=number,
state="open",
labels=labels,
title=title or f"issue {number}",
body=body,
priority=priority,
)
class ClassifyEpicContainerTest(unittest.TestCase):
def test_631_shaped_body_and_title_is_container(self) -> None:
c = _issue(
631,
title="Epic: MCP Control Plane Web Console",
body=_EPIC_631_BODY,
)
is_c, detail = classify_epic_or_child_only_container(c)
self.assertTrue(is_c)
self.assertIsNotNone(detail)
self.assertIn("body_marker", detail or "")
def test_body_markers_without_epic_title(self) -> None:
c = _issue(
900,
title="Control plane roadmap tracker",
body="Implementation is delivered via child issues only.",
)
is_c, _ = classify_epic_or_child_only_container(c)
self.assertTrue(is_c)
def test_epic_label_alone_is_container(self) -> None:
c = _issue(
901,
title="Roadmap linkage",
body="Track children.",
labels=("status:ready", "type:epic"),
)
is_c, detail = classify_epic_or_child_only_container(c)
self.assertTrue(is_c)
self.assertIn("type:epic", detail or "")
def test_title_epic_prefix_alone_not_container(self) -> None:
"""Title-only 'Epic:' without body scope evidence stays eligible (#844)."""
c = _issue(
902,
title="Epic: something mentioned only in title",
body="Implement a concrete fix for the allocator skip list.",
)
is_c, detail = classify_epic_or_child_only_container(c)
self.assertFalse(is_c)
self.assertIsNone(detail)
def test_incidental_epic_word_not_container(self) -> None:
c = _issue(
903,
title="Document epic handoff conventions",
body=(
"Update the docs so implementable issues that mention an epic "
"remain independently executable."
),
)
is_c, _ = classify_epic_or_child_only_container(c)
self.assertFalse(is_c)
def test_prs_never_classified(self) -> None:
pr = WorkCandidate(
kind="pr",
number=10,
state="open",
title="Epic: fake",
body="Implementation is delivered via child issues only.",
head_sha="a" * 40,
priority=5,
)
is_c, _ = classify_epic_or_child_only_container(pr)
self.assertFalse(is_c)
class AllocateEpicContainerExclusionTest(unittest.TestCase):
def setUp(self) -> None:
self._tmp = tempfile.TemporaryDirectory()
self.addCleanup(self._tmp.cleanup)
self.db = ControlPlaneDB(os.path.join(self._tmp.name, "cp.sqlite3"))
def _alloc(self, candidates, **kwargs):
defaults = dict(
session_id="sess-844",
role="author",
remote=REMOTE,
org=ORG,
repo=REPO,
profile_name="prgs-author",
username="jcwalker3",
claims={},
apply=False,
)
defaults.update(kwargs)
return allocate_next_work(self.db, candidates=candidates, **defaults)
def test_631_shaped_epic_excluded_child_selected(self) -> None:
epic = _issue(
631,
title="Epic: MCP Control Plane Web Console",
body=_EPIC_631_BODY,
)
child = _issue(
637,
title="Web Console: Workflow-event timeline model (Phase 1)",
body=_CHILD_BODY,
)
res = self._alloc([epic, child], apply=False)
self.assertTrue(res["success"], res)
self.assertEqual(res["outcome"], OUTCOME_PREVIEW)
self.assertEqual(res["selected"]["number"], 637)
skipped = {s["number"]: s for s in res["skipped"]}
self.assertIn(631, skipped)
self.assertEqual(
skipped[631]["reason_code"], SKIP_EPIC_OR_CHILD_ONLY_CONTAINER
)
self.assertIn(SKIP_EPIC_OR_CHILD_ONLY_CONTAINER, skipped[631]["reason"])
def test_container_cannot_receive_assignment_or_lease(self) -> None:
epic = _issue(
631,
title="Epic: MCP Control Plane Web Console",
body=_EPIC_631_BODY,
)
res = self._alloc([epic], apply=True)
self.assertTrue(res["success"], res)
# Only container present → no safe work; never assigned_work.
self.assertNotEqual(res["outcome"], OUTCOME_ASSIGNED)
self.assertIsNone(res.get("assignment"))
self.assertIsNone(res.get("selected"))
skipped = {s["number"]: s for s in res["skipped"]}
self.assertEqual(
skipped[631]["reason_code"], SKIP_EPIC_OR_CHILD_ONLY_CONTAINER
)
# No lease row for the epic.
leases = self.db.list_active_leases(
remote=REMOTE, org=ORG, repo=REPO
) if hasattr(self.db, "list_active_leases") else []
# Prefer generic inventory if available.
if not leases and hasattr(self.db, "list_leases"):
leases = self.db.list_leases(remote=REMOTE, org=ORG, repo=REPO)
for lease in leases or []:
work_number = lease.get("work_number") if isinstance(lease, dict) else None
self.assertNotEqual(work_number, 631)
def test_incidental_epic_title_remains_eligible(self) -> None:
ordinary = _issue(
700,
title="Document epic handoff conventions",
body="Write runbook text about epic vs child issues.",
)
res = self._alloc([ordinary], apply=False)
self.assertTrue(res["success"], res)
self.assertEqual(res["selected"]["number"], 700)
self.assertEqual(res["skipped"], [])
def test_apply_selects_child_not_epic(self) -> None:
epic = _issue(
631,
title="Epic: MCP Control Plane Web Console",
body=_EPIC_631_BODY,
)
child = _issue(
637,
title="Web Console: Workflow-event timeline model (Phase 1)",
body=_CHILD_BODY,
)
res = self._alloc([epic, child], apply=True)
self.assertTrue(res["success"], res)
self.assertEqual(res["outcome"], OUTCOME_ASSIGNED)
self.assertEqual(res["selected"]["number"], 637)
self.assertEqual(res["assignment"]["work_number"], 637)
if __name__ == "__main__":
unittest.main()
+372
View File
@@ -1266,6 +1266,378 @@ class TestSecondRemediationIntegration(unittest.TestCase):
self.assertIn("delete_acknowledged", delete_actions[0])
self.assertTrue(delete_actions[0].get("verified_absent"))
def test_issue_851_worktree_removed_when_remote_blocked_only_by_worktree_binding(self):
"""#851: remote blocked by worktree_binding must not skip safe worktree removal.
Lifecycle: remove clean owned worktree → reassess ownership → delete
remote only if independently safe. Unrelated entries stay untouched.
"""
from mcp_server import gitea_reconcile_merged_cleanups
target_branch = "fix/issue-844-exclude-epic-containers"
foreign_branch = "fix/issue-999-unrelated-active"
worktree_path = "/tmp/branches/fix-issue-844-exclude-epic-containers"
ownership_calls = []
remove_calls = []
delete_api_calls = []
def fake_collect(**kwargs):
ownership_calls.append(dict(kwargs))
# Ownership is reassessed *after* independent worktree removal (#851).
# Target worktree is already gone → no worktree_binding remains.
# Foreign branch keeps an active author lease → remote delete blocked.
if kwargs.get("branch") == foreign_branch:
# Match session-bound org/repo + host used by the tool resolve path.
return {
"records": [
{
"category": guard.OWNERSHIP_CATEGORY_AUTHOR_LEASE,
"status": "active",
"remote": kwargs.get("remote") or "prgs",
"host": kwargs.get("host") or "gitea.example.com",
"org": kwargs.get("org") or "Scaled-Tech-Consulting",
"repo": kwargs.get("repo") or "Gitea-Tools",
"branch": foreign_branch,
"reclaim_allowed": False,
}
],
"inventory_error": False,
}
return {"records": [], "inventory_error": False}
def fake_remove(project_root, branch, worktree_path=None):
remove_calls.append(
{"branch": branch, "worktree_path": worktree_path}
)
return {
"success": True,
"performed": True,
"message": f"removed worktree {worktree_path}",
"worktree_path": worktree_path,
}
def fake_probe(h, o, r, auth, br):
return guard.classify_branch_readback_http_status(
404, not_found_scope=guard.NOT_FOUND_SCOPE_BRANCH
)
def fake_api(method, url, auth, **kwargs):
if method == "DELETE":
delete_api_calls.append(url)
return {}
report = {
"entries": [
{
"pr_number": 848,
"head_branch": target_branch,
"remote_branch": {"safe_to_delete_remote": True},
"local_worktree": {
"safe_to_remove_worktree": True,
"worktree_path": worktree_path,
},
},
{
"pr_number": 999,
"head_branch": foreign_branch,
"remote_branch": {"safe_to_delete_remote": True},
"local_worktree": {
"safe_to_remove_worktree": False,
"worktree_path": None,
},
},
],
"reviewer_scratch_entries": [],
}
patch(
"mcp_server.get_profile",
return_value={
"profile_name": "prgs-reconciler",
"role": "reconciler",
"allowed_operations": [
"gitea.read",
"gitea.branch.delete",
"gitea.pr.close",
],
"forbidden_operations": [],
},
).start()
patch("mcp_server.api_get_all", return_value=[]).start()
patch(
"mcp_server.merged_cleanup_reconcile.build_reconciliation_report",
return_value=report,
).start()
patch(
"mcp_server.merged_cleanup_reconcile.discover_reviewer_scratch_worktrees",
return_value=[],
).start()
patch(
"mcp_server.audit_reconciliation_mode.check_cleanup_execution_allowed",
return_value=(True, []),
).start()
patch("mcp_server.verify_preflight_purity", return_value=None).start()
patch(
"mcp_server._collect_branch_ownership_records",
side_effect=fake_collect,
).start()
patch("mcp_server._probe_remote_branch", side_effect=fake_probe).start()
patch(
"mcp_server.merged_cleanup_reconcile.remove_local_worktree",
side_effect=fake_remove,
).start()
self.mock_api.side_effect = fake_api
res = gitea_reconcile_merged_cleanups(
dry_run=False,
execute_confirmed=True,
remote="prgs",
)
self.assertTrue(res.get("performed") or res.get("executed"))
actions = res.get("actions") or []
remove_actions = [
a for a in actions if a.get("action") == "remove_local_worktree"
]
self.assertEqual(len(remove_actions), 1, actions)
self.assertTrue(remove_actions[0].get("success"))
self.assertEqual(remove_calls[0]["branch"], target_branch)
self.assertEqual(remove_calls[0]["worktree_path"], worktree_path)
# Target remote delete succeeds after worktree removal + reassessment.
target_deletes = [
a
for a in actions
if a.get("action") == "delete_remote_branch"
and a.get("branch") == target_branch
]
self.assertEqual(len(target_deletes), 1, actions)
self.assertTrue(target_deletes[0].get("success"))
self.assertTrue(target_deletes[0].get("after_worktree_removal"))
self.assertTrue(target_deletes[0].get("ownership_reassessed"))
self.assertTrue(target_deletes[0].get("verified_absent"))
# Foreign branch remains protected (author lease) and is not deleted.
foreign_deletes = [
a
for a in actions
if a.get("action") == "delete_remote_branch"
and a.get("branch") == foreign_branch
]
self.assertEqual(len(foreign_deletes), 1, actions)
self.assertFalse(foreign_deletes[0].get("success"))
self.assertEqual(
foreign_deletes[0].get("blocker_kind"), "active_branch_ownership"
)
self.assertIn(
guard.OWNERSHIP_CATEGORY_AUTHOR_LEASE,
foreign_deletes[0].get("blocking_categories") or [],
)
# Only the target branch should hit the DELETE API.
self.assertEqual(len(delete_api_calls), 1)
# Ownership collected for target (post-removal) and foreign; worktree
# removal happened before target remote delete in the action log.
target_idx = next(
i
for i, a in enumerate(actions)
if a.get("action") == "remove_local_worktree"
)
delete_idx = next(
i
for i, a in enumerate(actions)
if a.get("action") == "delete_remote_branch"
and a.get("branch") == target_branch
and a.get("success")
)
self.assertLess(target_idx, delete_idx)
def test_issue_851_dirty_worktree_not_removed_and_remote_stays_protected(self):
"""#851: dirty/foreign worktrees remain protected; no unsafe cleanup."""
from mcp_server import gitea_reconcile_merged_cleanups
branch = "fix/issue-851-dirty"
remove_calls = []
def fake_collect(**kwargs):
return {
"records": [
{
"category": guard.OWNERSHIP_CATEGORY_WORKTREE_BINDING,
"status": "active",
"remote": kwargs.get("remote") or "prgs",
"host": kwargs.get("host") or "gitea.example.com",
"org": kwargs.get("org") or "Scaled-Tech-Consulting",
"repo": kwargs.get("repo") or "Gitea-Tools",
"branch": branch,
"reclaim_allowed": False,
}
],
"inventory_error": False,
}
report = {
"entries": [
{
"pr_number": 851,
"head_branch": branch,
"remote_branch": {"safe_to_delete_remote": True},
"local_worktree": {
"safe_to_remove_worktree": False,
"worktree_path": "/tmp/dirty-wt",
},
}
],
"reviewer_scratch_entries": [],
}
patch(
"mcp_server.get_profile",
return_value={
"profile_name": "prgs-reconciler",
"role": "reconciler",
"allowed_operations": [
"gitea.read",
"gitea.branch.delete",
],
"forbidden_operations": [],
},
).start()
patch("mcp_server.api_get_all", return_value=[]).start()
patch(
"mcp_server.merged_cleanup_reconcile.build_reconciliation_report",
return_value=report,
).start()
patch(
"mcp_server.merged_cleanup_reconcile.discover_reviewer_scratch_worktrees",
return_value=[],
).start()
patch(
"mcp_server.audit_reconciliation_mode.check_cleanup_execution_allowed",
return_value=(True, []),
).start()
patch("mcp_server.verify_preflight_purity", return_value=None).start()
patch(
"mcp_server._collect_branch_ownership_records",
side_effect=fake_collect,
).start()
patch(
"mcp_server.merged_cleanup_reconcile.remove_local_worktree",
side_effect=lambda *a, **k: remove_calls.append(k) or {
"success": True,
"performed": True,
},
).start()
self.mock_api.side_effect = lambda *a, **k: {}
res = gitea_reconcile_merged_cleanups(
dry_run=False,
execute_confirmed=True,
remote="prgs",
)
actions = res.get("actions") or []
self.assertEqual(remove_calls, [])
self.assertFalse(
any(a.get("action") == "remove_local_worktree" for a in actions)
)
deletes = [
a for a in actions if a.get("action") == "delete_remote_branch"
]
self.assertEqual(len(deletes), 1)
self.assertFalse(deletes[0].get("success"))
self.assertEqual(deletes[0].get("blocker_kind"), "active_branch_ownership")
self.assertIn(
guard.OWNERSHIP_CATEGORY_WORKTREE_BINDING,
deletes[0].get("blocking_categories") or [],
)
def test_issue_851_idempotent_resume_when_worktree_already_absent(self):
"""#851: partial failures remain resumable and idempotent."""
from mcp_server import gitea_reconcile_merged_cleanups
branch = "fix/issue-851-resume"
ownership_calls = []
def fake_collect(**kwargs):
ownership_calls.append(kwargs)
return {"records": [], "inventory_error": False}
def fake_remove(project_root, branch, worktree_path=None):
return {
"success": False,
"performed": False,
"message": f"worktree not found: {worktree_path}",
}
def fake_probe(h, o, r, auth, br):
return guard.classify_branch_readback_http_status(
404, not_found_scope=guard.NOT_FOUND_SCOPE_BRANCH
)
report = {
"entries": [
{
"pr_number": 851,
"head_branch": branch,
"remote_branch": {"safe_to_delete_remote": True},
"local_worktree": {
"safe_to_remove_worktree": True,
"worktree_path": "/tmp/already-gone",
},
}
],
"reviewer_scratch_entries": [],
}
patch(
"mcp_server.get_profile",
return_value={
"profile_name": "prgs-reconciler",
"role": "reconciler",
"allowed_operations": [
"gitea.read",
"gitea.branch.delete",
],
"forbidden_operations": [],
},
).start()
patch("mcp_server.api_get_all", return_value=[]).start()
patch(
"mcp_server.merged_cleanup_reconcile.build_reconciliation_report",
return_value=report,
).start()
patch(
"mcp_server.merged_cleanup_reconcile.discover_reviewer_scratch_worktrees",
return_value=[],
).start()
patch(
"mcp_server.audit_reconciliation_mode.check_cleanup_execution_allowed",
return_value=(True, []),
).start()
patch("mcp_server.verify_preflight_purity", return_value=None).start()
patch(
"mcp_server._collect_branch_ownership_records",
side_effect=fake_collect,
).start()
patch("mcp_server._probe_remote_branch", side_effect=fake_probe).start()
patch(
"mcp_server.merged_cleanup_reconcile.remove_local_worktree",
side_effect=fake_remove,
).start()
self.mock_api.side_effect = lambda *a, **k: {}
res = gitea_reconcile_merged_cleanups(
dry_run=False,
execute_confirmed=True,
remote="prgs",
)
actions = res.get("actions") or []
removes = [a for a in actions if a.get("action") == "remove_local_worktree"]
deletes = [a for a in actions if a.get("action") == "delete_remote_branch"]
self.assertEqual(len(removes), 1)
self.assertFalse(removes[0].get("success"))
self.assertEqual(len(deletes), 1)
self.assertTrue(deletes[0].get("success"))
self.assertTrue(deletes[0].get("after_worktree_removal"))
self.assertTrue(ownership_calls)
if __name__ == "__main__":
@@ -0,0 +1,483 @@
"""Synthetic regression coverage for dirty orphaned worktree recovery (#860).
Modeled on the #850 / #855 shape without mutating their real state.
"""
from __future__ import annotations
import json
import os
import shutil
import tempfile
import unittest
from unittest import mock
import dirty_orphan_worktree_recovery as dorec
import issue_lock_store
DEAD_PID = 999_999_999
LIVE_PID = os.getpid()
BRANCH = "fix/issue-901-dirty-orphan"
SOURCE_WT = "/repo/branches/issue-901-dirty-orphan"
RECOVERY_WT_NAME = "recovery-issue-901-dirty-orphan"
LOCAL_HEAD = "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa"
REMOTE_HEAD = "bbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbb"
OTHER_HEAD = "cccccccccccccccccccccccccccccccccccccccc"
FP_A = dorec.sha256_bytes(b"dirty-a")
FP_B = dorec.sha256_bytes(b"dirty-b")
FP_C = dorec.sha256_bytes(b"dirty-c-conflict")
def durable_lock(**overrides):
"""#850-shaped PID-less malformed same-claimant lock."""
lock = {
"issue_number": 901,
"branch_name": BRANCH,
"worktree_path": SOURCE_WT,
"remote": "prgs",
"org": "Example-Org",
"repo": "Example-Repo",
# intentionally no pid / session_pid / work_lease expiry
"claimant": {"username": "author-user", "profile": "prgs-author"},
}
lock.update(overrides)
return lock
def base_kwargs(**overrides):
kwargs = {
"issue_number": 901,
"branch_name": BRANCH,
"source_worktree_path": SOURCE_WT,
"remote": "prgs",
"org": "Example-Org",
"repo": "Example-Repo",
"identity": "author-user",
"profile": "prgs-author",
"expected_local_head": LOCAL_HEAD,
"expected_remote_head": REMOTE_HEAD,
"expected_dirty_fingerprints": {"a.py": FP_A, "b.py": FP_B},
"current_branch": BRANCH,
"porcelain_status": " M a.py\n M b.py\n",
"observed_local_head": LOCAL_HEAD,
"observed_remote_head": REMOTE_HEAD,
"observed_dirty_fingerprints": {"a.py": FP_A, "b.py": FP_B},
"competing_live_locks": [],
"competing_live_sessions": [],
"workflow_lease_active": False,
"workflow_lease_expired": True,
"canonical_repo_root": "/repo",
"worktree_registered": True,
"current_pid": LIVE_PID,
}
kwargs.update(overrides)
return kwargs
def assess(lock=None, **overrides):
return dorec.assess_dirty_orphan_recovery(
durable_lock() if lock is None else lock, **base_kwargs(**overrides)
)
class FreshnessPidLess(unittest.TestCase):
def test_pid_less_lock_is_not_live(self):
freshness = issue_lock_store.assess_lock_freshness(durable_lock())
self.assertFalse(freshness["live"])
self.assertTrue(freshness.get("pid_missing"))
self.assertEqual(freshness["status"], "malformed")
def test_pid_less_with_far_future_expiry_still_not_live(self):
lock = durable_lock(
work_lease={
"operation_type": "author_issue_work",
"expires_at": "2999-01-01T00:00:00Z",
"last_heartbeat_at": "2999-01-01T00:00:00Z",
}
)
freshness = issue_lock_store.assess_lock_freshness(lock)
self.assertFalse(freshness["live"])
self.assertTrue(freshness.get("pid_missing"))
class EligibilityGranted(unittest.TestCase):
def test_dead_same_claimant_pid_less_dirty(self):
result = assess()
self.assertEqual(result["outcome"], dorec.ELIGIBLE)
self.assertTrue(result["eligible"])
def test_expired_workflow_lease_corroboration(self):
result = assess(workflow_lease_active=False, workflow_lease_expired=True)
self.assertTrue(result["eligible"])
def test_older_local_newer_remote_heads(self):
result = assess()
self.assertTrue(result["evidence"].get("heads_diverged"))
self.assertTrue(result["eligible"])
class EligibilityRefused(unittest.TestCase):
def test_active_owner_with_pid(self):
lock = durable_lock(pid=LIVE_PID, session_pid=LIVE_PID)
result = assess(lock=lock, owner_process_alive_override=True)
self.assertEqual(result["outcome"], dorec.REFUSED)
self.assertFalse(result["eligible"])
self.assertTrue(any("alive" in r for r in result["reasons"]))
def test_foreign_claimant(self):
result = assess(identity="other-user")
self.assertEqual(result["outcome"], dorec.REFUSED)
self.assertTrue(any("foreign claimant identity" in r for r in result["reasons"]))
def test_foreign_profile(self):
result = assess(profile="prgs-reviewer")
self.assertEqual(result["outcome"], dorec.REFUSED)
def test_fingerprint_mismatch(self):
result = assess(observed_dirty_fingerprints={"a.py": "0" * 64, "b.py": FP_B})
self.assertEqual(result["outcome"], dorec.REFUSED)
self.assertTrue(any("fingerprint mismatch" in r for r in result["reasons"]))
def test_head_mismatch(self):
result = assess(observed_local_head=OTHER_HEAD)
self.assertEqual(result["outcome"], dorec.REFUSED)
def test_remote_head_mismatch(self):
result = assess(observed_remote_head=OTHER_HEAD)
self.assertEqual(result["outcome"], dorec.REFUSED)
def test_path_not_under_branches(self):
result = assess(
source_worktree_path="/tmp/branches/evil",
# lock path also changed so worktree agreement holds
lock=durable_lock(worktree_path="/tmp/branches/evil"),
)
self.assertEqual(result["outcome"], dorec.REFUSED)
self.assertTrue(any("canonical branches" in r for r in result["reasons"]))
def test_unregistered_worktree(self):
result = assess(worktree_registered=False)
self.assertEqual(result["outcome"], dorec.REFUSED)
def test_active_workflow_lease(self):
result = assess(workflow_lease_active=True, workflow_lease_expired=False)
self.assertEqual(result["outcome"], dorec.REFUSED)
def test_unsafe_dirty_path_pin(self):
result = assess(
expected_dirty_fingerprints={"../etc/passwd": FP_A},
observed_dirty_fingerprints={"../etc/passwd": FP_A},
)
self.assertEqual(result["outcome"], dorec.REFUSED)
def test_symlink_escape_rejected_by_ancestry(self):
ok, reasons = dorec.is_path_under_canonical_branches(
"/tmp/branches/evil", canonical_repo_root="/repo"
)
self.assertFalse(ok)
self.assertTrue(reasons)
class ConflictDetection(unittest.TestCase):
def test_overlapping_upstream_change(self):
conflicts = dorec.detect_path_conflicts(
dirty_paths=["c.py"],
local_head_contents={"c.py": b"local-base"},
remote_head_contents={"c.py": b"remote-changed"},
dirty_contents={"c.py": b"dirty-c-conflict"},
)
self.assertEqual(len(conflicts), 1)
self.assertEqual(conflicts[0]["path"], "c.py")
def test_unchanged_upstream_no_conflict(self):
conflicts = dorec.detect_path_conflicts(
dirty_paths=["a.py"],
local_head_contents={"a.py": b"same"},
remote_head_contents={"a.py": b"same"},
dirty_contents={"a.py": b"dirty-a"},
)
self.assertEqual(conflicts, [])
class CrashSafeRecovery(unittest.TestCase):
def setUp(self):
self.tmp = tempfile.mkdtemp(prefix="dirty-orphan-")
self.repo = os.path.join(self.tmp, "repo")
self.branches = os.path.join(self.repo, "branches")
self.source = os.path.join(self.branches, "issue-901-dirty-orphan")
self.recovery = os.path.join(self.branches, RECOVERY_WT_NAME)
os.makedirs(self.source, exist_ok=True)
os.makedirs(self.branches, exist_ok=True)
# seed dirty files in source
with open(os.path.join(self.source, "a.py"), "wb") as fh:
fh.write(b"dirty-a")
with open(os.path.join(self.source, "b.py"), "wb") as fh:
fh.write(b"dirty-b")
self.journal_dir = os.path.join(self.tmp, "journals")
self.lock = durable_lock(worktree_path=self.source)
self.assessment = dorec.assess_dirty_orphan_recovery(
self.lock,
**base_kwargs(
source_worktree_path=self.source,
canonical_repo_root=self.repo,
),
)
class FakeGit(dorec.GitOps):
def __init__(self, recovery_path, head):
self.recovery_path = recovery_path
self.head = head
self.calls = []
def run(self, args, *, cwd):
self.calls.append((args, cwd))
if args[:3] == ["git", "worktree", "add"]:
os.makedirs(self.recovery_path, exist_ok=True)
return mock.Mock(returncode=0, stdout="", stderr="")
if args[:2] == ["git", "checkout"]:
return mock.Mock(returncode=0, stdout="", stderr="")
if args[:2] == ["git", "rev-parse"]:
return mock.Mock(returncode=0, stdout=self.head + "\n", stderr="")
return mock.Mock(returncode=0, stdout="", stderr="")
self.git = FakeGit(self.recovery, REMOTE_HEAD)
self.written_locks = []
def lock_writer(record):
self.written_locks.append(record)
self.lock_writer = lock_writer
def tearDown(self):
shutil.rmtree(self.tmp, ignore_errors=True)
def _run(self, **overrides):
kwargs = {
"assessment": self.assessment,
"existing_lock": self.lock,
"issue_number": 901,
"branch_name": BRANCH,
"source_worktree_path": self.source,
"recovery_worktree_path": self.recovery,
"remote": "prgs",
"org": "Example-Org",
"repo": "Example-Repo",
"identity": "author-user",
"profile": "prgs-author",
"expected_local_head": LOCAL_HEAD,
"expected_remote_head": REMOTE_HEAD,
"expected_dirty_fingerprints": {"a.py": FP_A, "b.py": FP_B},
"dirty_contents": {"a.py": b"dirty-a", "b.py": b"dirty-b"},
"local_head_contents": {"a.py": b"base-a", "b.py": b"base-b"},
"remote_head_contents": {"a.py": b"base-a", "b.py": b"base-b"},
"canonical_repo_root": self.repo,
"bind_lock": True,
"lock_writer": self.lock_writer,
"git_ops": self.git,
"journal_dir": self.journal_dir,
"session_pid": LIVE_PID,
}
kwargs.update(overrides)
return dorec.run_dirty_orphan_recovery(**kwargs)
def test_success_preserves_dirty_bytes_and_source(self):
result = self._run()
self.assertTrue(result["success"])
self.assertEqual(result["outcome"], dorec.RECOVERY_COMPLETED)
self.assertTrue(os.path.isdir(self.source))
with open(os.path.join(self.source, "a.py"), "rb") as fh:
self.assertEqual(fh.read(), b"dirty-a")
with open(os.path.join(self.recovery, "a.py"), "rb") as fh:
self.assertEqual(fh.read(), b"dirty-a")
with open(os.path.join(self.recovery, "b.py"), "rb") as fh:
self.assertEqual(fh.read(), b"dirty-b")
self.assertEqual(len(self.written_locks), 1)
rec = self.written_locks[0]
self.assertEqual(rec["session_pid"], LIVE_PID)
self.assertTrue(rec["dirty_orphan_recovery"]["recovered"])
self.assertTrue(rec["dirty_orphan_recovery"]["source_frozen"])
def test_conflict_leaves_governed_state(self):
result = self._run(
expected_dirty_fingerprints={"c.py": FP_C},
dirty_contents={"c.py": b"dirty-c-conflict"},
local_head_contents={"c.py": b"local-base"},
remote_head_contents={"c.py": b"remote-changed"},
)
# #860 F4: session binding is NOT finalized while conflicts remain
self.assertFalse(result["success"])
self.assertEqual(result["outcome"], dorec.CONFLICTS_PRESENT)
sidecar = os.path.join(self.recovery, "c.py.recovered-dirty")
self.assertTrue(os.path.isfile(sidecar))
state = os.path.join(
self.recovery, dorec.CONFLICT_STATE_DIR, dorec.CONFLICT_STATE_FILE
)
self.assertTrue(os.path.isfile(state))
with open(state, "r", encoding="utf-8") as fh:
payload = json.load(fh)
self.assertEqual(payload["resolution"], "author_edit_required")
def test_interrupt_before_journal_no_artifacts(self):
result = self._run(interrupt_after_phase=dorec.PHASE_ELIGIBILITY)
self.assertFalse(result["success"])
self.assertEqual(result["outcome"], "INTERRUPTED")
self.assertFalse(os.path.isdir(self.recovery))
def test_interrupt_after_journal_then_retry_idempotent(self):
first = self._run(interrupt_after_phase=dorec.PHASE_JOURNAL_PERSISTED)
self.assertEqual(first["outcome"], "INTERRUPTED")
self.assertTrue(first["journal"]["artifacts_created"]["journal"])
second = self._run()
self.assertTrue(second["success"])
# source still recoverable
with open(os.path.join(self.source, "a.py"), "rb") as fh:
self.assertEqual(fh.read(), b"dirty-a")
def test_interrupt_after_worktree_then_retry(self):
first = self._run(interrupt_after_phase=dorec.PHASE_RECOVERY_WORKTREE)
self.assertEqual(first["outcome"], "INTERRUPTED")
self.assertTrue(os.path.isdir(self.recovery))
second = self._run()
self.assertTrue(second["success"])
def test_interrupt_after_binding_then_retry_complete(self):
first = self._run(interrupt_after_phase=dorec.PHASE_BINDING)
self.assertEqual(first["outcome"], "INTERRUPTED")
second = self._run()
self.assertTrue(second["success"])
# completed journal makes further retries no-ops
third = self._run()
self.assertEqual(third["outcome"], dorec.RECOVERY_RESUMED)
def test_source_worktree_never_deleted(self):
self._run()
self.assertTrue(os.path.isdir(self.source))
self.assertTrue(os.path.isfile(os.path.join(self.source, "a.py")))
def test_fingerprint_drift_refuses_without_mutation(self):
result = self._run(dirty_contents={"a.py": b"CHANGED", "b.py": b"dirty-b"})
self.assertFalse(result["success"])
self.assertFalse(os.path.isdir(self.recovery))
class SessionBindingPreflight(unittest.TestCase):
def test_canonical_session_binding_recognized(self):
lock = {
"worktree_path": "/repo/branches/recovery",
"session_pid": LIVE_PID,
"dirty_orphan_recovery": {
"recovered": True,
"conflicts": [],
"recovery_worktree_path": "/repo/branches/recovery",
"source_worktree_path": SOURCE_WT,
"accepted_head": REMOTE_HEAD,
},
}
result = dorec.preflight_recognizes_recovered_provenance(lock)
self.assertTrue(result["recognized"])
def test_conflicts_block_commit_preflight(self):
lock = {
"worktree_path": "/repo/branches/recovery",
"session_pid": LIVE_PID,
"dirty_orphan_recovery": {
"recovered": True,
"conflicts": [{"path": "c.py"}],
},
}
result = dorec.preflight_recognizes_recovered_provenance(lock)
self.assertFalse(result["recognized"])
def test_active_foreign_does_not_mutate(self):
# assess-only path: foreign refused before run
result = assess(identity="intruder")
self.assertFalse(result["eligible"])
class JournalSymlinkRefusal(unittest.TestCase):
def test_symlink_journal_path_refused_on_load(self):
tmp = tempfile.mkdtemp()
try:
real = os.path.join(tmp, "real.json")
with open(real, "w", encoding="utf-8") as fh:
fh.write("{}")
link = os.path.join(tmp, "link.json")
os.symlink(real, link)
key = "symlink-test"
jdir = tmp
path = dorec._journal_path(key, journal_dir=jdir)
with open(path, "w", encoding="utf-8") as fh:
json.dump({"idempotency_key": key}, fh)
os.remove(path)
os.symlink(real, path)
with self.assertRaises(ValueError):
dorec.load_journal(key, journal_dir=jdir)
finally:
shutil.rmtree(tmp, ignore_errors=True)
class RealGitMultiWorktreeIntegration(unittest.TestCase):
def setUp(self):
import subprocess
self.tmp = tempfile.mkdtemp(prefix="git-integration-")
self.repo = os.path.join(self.tmp, "repo")
os.makedirs(self.repo, exist_ok=True)
subprocess.run(["git", "init"], cwd=self.repo, check=True, capture_output=True)
subprocess.run(["git", "config", "user.name", "Test User"], cwd=self.repo, check=True)
subprocess.run(["git", "config", "user.email", "[email protected]"], cwd=self.repo, check=True)
with open(os.path.join(self.repo, "init.txt"), "w") as fh:
fh.write("init")
subprocess.run(["git", "add", "."], cwd=self.repo, check=True)
subprocess.run(["git", "commit", "-m", "init"], cwd=self.repo, check=True)
branch = "fix/issue-999-test"
subprocess.run(["git", "branch", branch], cwd=self.repo, check=True)
self.branches = os.path.join(self.repo, "branches")
self.source = os.path.join(self.branches, "issue-999-test")
subprocess.run(["git", "worktree", "add", self.source, branch], cwd=self.repo, check=True)
self.dirty_path = os.path.join(self.source, "dirty.txt")
with open(self.dirty_path, "w") as fh:
fh.write("dirty-data")
def tearDown(self):
shutil.rmtree(self.tmp, ignore_errors=True)
def test_prepare_recovery_worktree_detached_no_exit_128(self):
import subprocess
head_sha = subprocess.check_output(["git", "rev-parse", "HEAD"], cwd=self.repo, text=True).strip()
rec_wt = os.path.join(self.branches, "recovery-issue-999-test")
res = dorec.prepare_recovery_worktree(
canonical_repo_root=self.repo,
recovery_worktree_path=rec_wt,
branch_name="fix/issue-999-test",
remote_head=head_sha,
)
self.assertTrue(res["success"], res.get("reasons"))
self.assertTrue(os.path.isdir(rec_wt))
def test_real_lock_rebind_recovery_sanctioned(self):
lock_dir = os.path.join(self.tmp, "locks")
rec_wt = os.path.join(self.branches, "recovery-issue-999-test")
os.makedirs(rec_wt, exist_ok=True)
record = {
"remote": "prgs",
"org": "Example-Org",
"repo": "Example-Repo",
"issue_number": 999,
"branch_name": "fix/issue-999-test",
"worktree_path": rec_wt,
"claimant": {"username": "author-user", "profile": "prgs-author"},
}
record_src = dict(record)
record_src["worktree_path"] = self.source
issue_lock_store.bind_session_lock(record_src, lock_dir=lock_dir)
path = issue_lock_store.bind_session_lock(
record,
lock_dir=lock_dir,
recovery_sanctioned=True,
)
self.assertTrue(os.path.isfile(path))
if __name__ == "__main__":
unittest.main()
+6
View File
@@ -24,6 +24,8 @@ def _lease(expires_at: str) -> dict:
def _lock_record(**overrides) -> dict:
# #860: live locks require a usable session pid; PID-less records are never
# classified live merely because expiry/heartbeat fields are present.
record = {
"issue_number": 420,
"branch_name": "feat/issue-420-server-code-parity",
@@ -31,6 +33,8 @@ def _lock_record(**overrides) -> dict:
"org": "Scaled-Tech-Consulting",
"repo": "Gitea-Tools",
"worktree_path": "/tmp/wt-420",
"session_pid": os.getpid(),
"pid": os.getpid(),
"work_lease": _lease("2999-01-01T00:00:00Z"),
}
record.update(overrides)
@@ -88,6 +92,8 @@ class TestIssueLockStore(unittest.TestCase):
existing = _lock_record(
branch_name="feat/issue-420-other",
worktree_path="/tmp/other",
session_pid=os.getpid(),
pid=os.getpid(),
work_lease=_lease("2999-01-01T00:00:00Z"),
)
path = ils.lock_file_path(
+53
View File
@@ -12,6 +12,59 @@ import merged_cleanup_reconcile as mcr # noqa: E402
class TestMergedCleanupAssessment(unittest.TestCase):
def test_issue_851_plan_order_worktree_then_reassess_then_remote(self):
"""#851 dry-run plan: remove worktree, reassess ownership, then remote."""
plan = mcr.plan_cleanup_execution_order(
remote_assessment={"safe_to_delete_remote": True},
local_assessment={"safe_to_remove_worktree": True},
)
actions = [s["action"] for s in plan]
self.assertEqual(
actions,
[
"remove_local_worktree",
"reassess_branch_ownership",
"delete_remote_branch",
],
)
self.assertEqual(plan[0]["phase"], 1)
self.assertEqual(plan[-1]["phase"], 3)
self.assertIn("independently_safe", plan[0]["reason"])
self.assertIn("reassessment", plan[-1]["reason"])
def test_issue_851_plan_remote_only_when_worktree_not_safe(self):
plan = mcr.plan_cleanup_execution_order(
remote_assessment={"safe_to_delete_remote": True},
local_assessment={"safe_to_remove_worktree": False},
)
self.assertEqual([s["action"] for s in plan], ["delete_remote_branch"])
self.assertNotIn("reassess_branch_ownership", [s["action"] for s in plan])
def test_issue_851_plan_worktree_only_when_remote_not_safe(self):
plan = mcr.plan_cleanup_execution_order(
remote_assessment={"safe_to_delete_remote": False},
local_assessment={"safe_to_remove_worktree": True},
)
self.assertEqual([s["action"] for s in plan], ["remove_local_worktree"])
def test_issue_851_entry_includes_planned_execution_order(self):
entry = mcr.build_pr_cleanup_entry(
pr={
"number": 848,
"title": "Closes #844",
"body": "",
"merged_at": "2026-07-23T00:00:00Z",
"head": {"ref": "fix/issue-844-x", "sha": "a" * 40},
},
project_root="/tmp/not-a-real-root",
open_pr_heads=set(),
remote_branch_exists=True,
head_on_master=True,
delete_capability_allowed=True,
)
self.assertIn("planned_execution_order", entry)
self.assertIsInstance(entry["planned_execution_order"], list)
def test_extract_linked_issue_from_closes(self):
issue = mcr.extract_linked_issue(
"feat: cleanup (Closes #269)",
+135
View File
@@ -0,0 +1,135 @@
"""Tests for the Phase 1 operator console application shell (#638)."""
import sys
import unittest
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
from starlette.routing import Route
from starlette.testclient import TestClient
from webui import layout
from webui.app import create_app
from webui.nav import NAV_GROUPS, STUB_PAGES, nav_hrefs
class TestShellNav(unittest.TestCase):
def setUp(self):
self.client = TestClient(create_app())
def test_nav_group_labels_present(self):
text = self.client.get("/").text
for group in NAV_GROUPS:
with self.subTest(group=group.label):
self.assertIn(f">{group.label}<", text)
def test_phase1_group_labels_cover_expected_ia(self):
labels = {group.label for group in NAV_GROUPS}
for expected in (
"Health",
"Traffic",
"Runtime/Sessions",
"Projects",
"Inventory",
"Timeline",
"Policy",
"Insights",
):
with self.subTest(label=expected):
self.assertIn(expected, labels)
def test_every_nav_href_resolves_to_a_get_route(self):
app = create_app()
get_paths = {
route.path
for route in app.routes
if isinstance(route, Route) and "GET" in route.methods
}
for href in nav_hrefs():
with self.subTest(href=href):
self.assertIn(href, get_paths, f"nav href {href} has no GET route")
def test_legacy_hrefs_still_navigable(self):
text = self.client.get("/").text
for href in ("/queue", "/projects", "/prompts", "/runtime",
"/audit", "/worktrees", "/leases", "/actions"):
with self.subTest(href=href):
self.assertIn(f'href="{href}"', text)
class TestShellBadges(unittest.TestCase):
def setUp(self):
self.client = TestClient(create_app())
def test_mode_badge_present(self):
self.assertIn("mode: read-only", self.client.get("/").text)
def test_environment_badge_present(self):
self.assertIn("env:", self.client.get("/").text)
def test_default_environment_is_local(self):
self.assertEqual(layout.environment_label(), "local")
def test_remote_bind_reports_remote_environment(self):
import os
prior = os.environ.get("WEBUI_HOST")
os.environ["WEBUI_HOST"] = "10.0.0.5"
try:
self.assertEqual(layout.environment_label(), "remote")
finally:
if prior is None:
os.environ.pop("WEBUI_HOST", None)
else:
os.environ["WEBUI_HOST"] = prior
def test_docs_link_present(self):
text = self.client.get("/").text
self.assertIn(layout.DOCS_URL, text)
self.assertIn(">Docs<", text)
class TestShellStubs(unittest.TestCase):
def setUp(self):
self.client = TestClient(create_app())
def test_stub_routes_render_200(self):
for path, (title, _desc) in STUB_PAGES.items():
with self.subTest(path=path):
response = self.client.get(path)
self.assertEqual(response.status_code, 200, path)
self.assertIn(title, response.text)
self.assertIn("placeholder", response.text)
def test_stub_routes_are_read_only(self):
for path in STUB_PAGES:
with self.subTest(path=path):
response = self.client.post(path)
self.assertEqual(response.status_code, 405)
self.assertEqual(response.json()["error"], "read-only-mvp")
def test_stub_pages_carry_nav_and_badges(self):
response = self.client.get("/inventory")
self.assertIn("mode: read-only", response.text)
self.assertIn('href="/queue"', response.text)
class TestShellHome(unittest.TestCase):
def setUp(self):
self.client = TestClient(create_app())
def test_home_summarizes_console(self):
text = self.client.get("/").text
self.assertIn("Operator console", text)
self.assertIn("Phase 1", text)
def test_home_links_legacy_pages(self):
text = self.client.get("/").text
self.assertIn("MVP legacy pages", text)
for href in ("/queue", "/audit", "/leases"):
with self.subTest(href=href):
self.assertIn(f'href="{href}"', text)
if __name__ == "__main__":
unittest.main()
File diff suppressed because it is too large Load Diff
+154 -11
View File
@@ -12,6 +12,7 @@ from starlette.routing import Route
from webui.deployment_boundary import deployment_snapshot
from webui.layout import render_page
from webui.nav import NAV_GROUPS, STUB_PAGES
from webui.project_registry import (
ProjectRegistry,
RegistryError,
@@ -45,6 +46,7 @@ from webui.worktree_scanner import load_hygiene_snapshot, snapshot_to_dict as wo
from webui.worktree_views import render_worktrees_page
from webui.runtime_health import load_runtime_snapshot, snapshot_to_dict as runtime_snapshot_to_dict
from webui.runtime_views import render_runtime_page
from webui.timeline import load_timeline, snapshot_to_dict as timeline_snapshot_to_dict
from webui.system_health import (
API_PATH as SYSTEM_HEALTH_API_PATH,
load_system_health,
@@ -65,24 +67,62 @@ def _stub_page(title: str, description: str) -> HTMLResponse:
return HTMLResponse(render_page(title=title, body_html=body))
_LEGACY_PAGES = (
("/queue", "Queue", "live PR and issue dashboard (#429)"),
("/projects", "Projects", "registry and onboarding (#427)"),
("/prompts", "Prompts", "canonical workflow prompt library (#428)"),
("/runtime", "Runtime", "MCP health and stale-runtime detection (#430)"),
("/audit", "Audit", "final-report paste and validator preview (#431)"),
("/worktrees", "Worktrees", "branch hygiene dashboard (#432)"),
("/leases", "Leases", "collision and lease visibility (#433)"),
("/actions", "Actions", "gated write-action framework (#434)"),
)
def _render_home_nav_groups() -> str:
groups = []
for group in NAV_GROUPS:
items = "".join(
f'<li><a href="{item.href}">{item.label}</a>'
+ ("" if item.status == "live" else " <span class=\"muted\">(stub)</span>")
+ "</li>"
for item in group.items
)
groups.append(f"<h3>{group.label}</h3><ul>{items}</ul>")
return "".join(groups)
async def home(_request: Request) -> HTMLResponse:
legacy = "".join(
f"<li><strong>{label}</strong> — {desc} "
f'(<a href="{href}">{href}</a>)</li>'
for href, label, desc in _LEGACY_PAGES
)
body = (
"<h2>Operator console</h2>"
"<p>Local entry point for MCP Control Plane operational views.</p>"
"<ul>"
"<li><strong>Queue</strong> — live PR and issue dashboard (#429)</li>"
"<li><strong>Projects</strong> — registry and onboarding (#427)</li>"
"<li><strong>Prompts</strong> — canonical workflow prompt library (#428)</li>"
"<li><strong>Runtime</strong> — MCP health and stale-runtime detection (#430)</li>"
"<li><strong>Audit</strong> — final-report paste and validator preview (#431)</li>"
"<li><strong>Worktrees</strong> — branch hygiene dashboard (#432)</li>"
"<li><strong>Leases</strong> — collision and lease visibility (#433)</li>"
"<li><strong>Actions</strong> — gated write-action framework (#434)</li>"
"</ul>"
"<p>Read-only home for the MCP Control Plane Phase 1 operator console. "
"Gitea, MCP capability gates, and canonical workflows remain the source "
"of truth; this console never mutates them.</p>"
"<h2>Phase 1 surfaces</h2>"
+ _render_home_nav_groups()
+ "<h2>MVP legacy pages</h2>"
"<ul>" + legacy + "</ul>"
)
return HTMLResponse(render_page(title="Home", body_html=body))
async def phase_stub(request: Request) -> HTMLResponse:
"""Graceful read-only placeholder for a not-yet-implemented Phase 1 surface."""
title, description = STUB_PAGES[request.url.path]
body = (
f"<h2>{title}</h2>"
f'<div class="stub"><p>{description}</p>'
"<p>Phase 1 shell placeholder — no write actions. Tracked under "
"epic #631.</p></div>"
)
return HTMLResponse(render_page(title=title, body_html=body))
async def health(_request: Request) -> JSONResponse:
"""Liveness only — deliberately cheap, runs no dependency probe (#634).
@@ -410,6 +450,104 @@ async def api_console_security_model(_request: Request) -> JSONResponse:
})
def _query_int(request: Request, key: str) -> int | None:
"""Parse an optional integer query parameter; None when absent/invalid."""
raw = request.query_params.get(key)
if raw is None or not str(raw).strip():
return None
try:
return int(str(raw).strip())
except (TypeError, ValueError):
return None
def _derive_remote(host: str) -> str:
"""Map a Gitea host to its known short remote name (control-plane scope key)."""
text = (host or "").lower()
if "prgs" in text:
return "prgs"
if "dadeschools" in text:
return "dadeschools"
return text.split(".")[0] if text else ""
def _timeline_comment_source(host: str, org: str, repo: str):
"""Build a fail-soft CTH-comment fetcher for one repo, or None when offline.
Returns a callable ``(kind, number) -> list[comment]``. Credentials or
network failures raise inside the callable so ``load_timeline`` degrades the
handoff source rather than the whole timeline. Offline test mode yields no
live source so the handoff section reports ``not run``.
"""
import os
from gitea_auth import api_fetch_page, get_auth_header, repo_api_url
offline = (os.environ.get("WEBUI_TEST_OFFLINE") or "").strip().lower() in {"1", "true", "yes"}
if offline:
return None
auth = get_auth_header(host)
if not auth:
return None
def _fetch(kind: str, number: int) -> list:
segment = "pulls" if kind == "pr" else "issues"
url = f"{repo_api_url(host, org, repo)}/{segment}/{int(number)}/comments"
comments: list = []
page = 1
while page <= 20:
raw, meta = api_fetch_page(url, auth, page=page, limit=50)
comments.extend(raw)
if bool(meta["is_final_page"]):
break
page += 1
return comments
return _fetch
async def api_v1_timeline(request: Request) -> JSONResponse:
"""Read-only workflow-event timeline (#637). Filter by issue/PR/session."""
from webui.queue_loader import _host_from_url # host normalisation helper
registry, error = _load_project_registry()
if error is not None:
return JSONResponse(error.to_dict(), status_code=500)
project = registry.projects[0] if registry.projects else None
org = request.query_params.get("org") or (project.gitea_owner if project else "")
repo = request.query_params.get("repo") or (project.repo_name if project else "")
host = _host_from_url(project.remote_host) if project else ""
remote = request.query_params.get("remote") or _derive_remote(host)
if not (remote and org and repo):
return JSONResponse(
{
"error": "timeline_scope_unresolved",
"detail": "no project in registry and no remote/org/repo query params provided",
},
status_code=400,
)
comment_source = _timeline_comment_source(host, org, repo) if (host and org and repo) else None
snapshot = load_timeline(
remote=remote,
org=org,
repo=repo,
issue=_query_int(request, "issue"),
pr=_query_int(request, "pr"),
session=(request.query_params.get("session") or None),
limit=_query_int(request, "limit"),
offset=_query_int(request, "offset"),
comment_source=comment_source,
)
# A filter no surviving source can carry is refused, not answered empty:
# a 200 with zero events would tell the operator no such activity exists.
status_code = 200 if snapshot.ok else 422
return JSONResponse(timeline_snapshot_to_dict(snapshot), status_code=status_code)
async def method_not_allowed(request: Request, _exc: Exception) -> Response:
path = request.url.path
if path in _AUDIT_MUTATION_PATHS and request.method == "POST":
@@ -449,6 +587,7 @@ def create_app(*, bind_host: str | None = None) -> Starlette:
Route("/api/prompts", api_prompts, methods=["GET"]),
Route("/runtime", runtime, methods=["GET"]),
Route("/api/runtime", api_runtime, methods=["GET"]),
Route("/api/v1/timeline", api_v1_timeline, methods=["GET"]),
Route("/audit", audit, methods=["GET", "POST"]),
Route("/api/audit", api_audit, methods=["GET", "POST"]),
Route("/worktrees", worktrees, methods=["GET"]),
@@ -472,6 +611,10 @@ def create_app(*, bind_host: str | None = None) -> Starlette:
api_console_security_model,
methods=["GET"],
),
*[
Route(path, phase_stub, methods=["GET"])
for path in STUB_PAGES
],
],
exception_handlers={405: method_not_allowed},
)
+95 -17
View File
@@ -2,28 +2,66 @@
from __future__ import annotations
NAV_ITEMS = (
("/", "Home"),
("/queue", "Queue"),
("/projects", "Projects"),
("/prompts", "Prompts"),
("/runtime", "Runtime"),
("/audit", "Audit"),
("/worktrees", "Worktrees"),
("/leases", "Leases"),
("/actions", "Actions"),
)
import os
from webui.nav import NAV_GROUPS
MVP_NOTICE = (
"Read-only MVP — Gitea, MCP tools, and canonical workflows remain the "
"source of truth. No mutation endpoints."
)
# Canonical docs entry point surfaced from the shell header (#638).
DOCS_URL = (
"https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/src/branch/"
"master/docs/webui-local-dev.md"
)
_LOCAL_HOSTS = frozenset({"", "127.0.0.1", "localhost", "::1"})
def environment_label() -> str:
"""Classify the serving environment as ``local`` or ``remote`` (#638).
Derived from the same ``WEBUI_HOST`` default the app binds to; loopback
hosts are ``local``, anything else is ``remote``. Read-only signal only.
"""
host = (os.environ.get("WEBUI_HOST", "127.0.0.1") or "").strip().lower()
return "local" if host in _LOCAL_HOSTS else "remote"
def _render_nav() -> str:
groups_html = []
for group in NAV_GROUPS:
links = "".join(
f'<a href="{item.href}"'
+ (' class="nav-stub"' if item.status == "stub" else "")
+ f'>{item.label}</a>'
for item in group.items
)
groups_html.append(
'<div class="nav-group">'
f'<span class="nav-group-label">{group.label}</span>'
f'<span class="nav-group-links">{links}</span>'
"</div>"
)
return "".join(groups_html)
def _render_badges() -> str:
env = environment_label()
return (
'<div class="header-badges">'
f'<span class="badge env-badge env-{env}">env: {env}</span>'
'<span class="badge mode-badge">mode: read-only</span>'
f'<a class="badge docs-link" href="{DOCS_URL}">Docs</a>'
"</div>"
)
def render_page(*, title: str, body_html: str, extra_head: str = "") -> str:
nav_links = "".join(
f'<a href="{href}">{label}</a>' for href, label in NAV_ITEMS
)
nav_links = _render_nav()
header_badges = _render_badges()
return f"""<!DOCTYPE html>
<html lang="en">
<head>
@@ -53,21 +91,58 @@ def render_page(*, title: str, body_html: str, extra_head: str = "") -> str:
padding: 0.75rem 1.25rem;
}}
header h1 {{
margin: 0 0 0.5rem;
margin: 0;
font-size: 1.1rem;
font-weight: 600;
}}
.header-top {{
display: flex;
flex-wrap: wrap;
align-items: center;
justify-content: space-between;
gap: 0.5rem 1rem;
margin-bottom: 0.6rem;
}}
.header-badges {{ display: inline-flex; flex-wrap: wrap; gap: 0.4rem; }}
.env-badge.env-local {{ color: #8fd19e; border-color: #3d6b4a; }}
.env-badge.env-remote {{ color: #e0c27a; border-color: #6b5730; }}
.mode-badge {{ color: #9ec8f0; border-color: #3d5f7a; }}
a.docs-link {{
color: var(--accent);
border-color: var(--accent);
text-decoration: none;
text-transform: none;
}}
a.docs-link:hover {{ filter: brightness(1.12); }}
nav {{
display: flex;
flex-wrap: wrap;
gap: 0.75rem 1rem;
gap: 0.5rem 1.25rem;
}}
.nav-group {{
display: flex;
flex-direction: column;
gap: 0.15rem;
}}
.nav-group-label {{
font-size: 0.68rem;
text-transform: uppercase;
letter-spacing: 0.04em;
color: var(--muted);
}}
.nav-group-links {{ display: inline-flex; flex-wrap: wrap; gap: 0.6rem; }}
nav a {{
color: var(--accent);
text-decoration: none;
font-size: 0.9rem;
}}
nav a:hover {{ text-decoration: underline; }}
nav a.nav-stub {{ color: var(--muted); }}
nav a.nav-stub::after {{
content: " ·stub";
font-size: 0.7rem;
color: var(--muted);
}}
main {{
max-width: 52rem;
margin: 0 auto;
@@ -166,7 +241,10 @@ def render_page(*, title: str, body_html: str, extra_head: str = "") -> str:
</head>
<body>
<header>
<h1>MCP Control Plane</h1>
<div class="header-top">
<h1>MCP Control Plane</h1>
{header_badges}
</div>
<nav>{nav_links}</nav>
</header>
<main>
+111
View File
@@ -0,0 +1,111 @@
"""Navigation IA for the Phase 1 operator console shell (#638).
Single source of truth for the console navigation so ``webui/layout.py`` and
the ``webui/app.py`` route table stay aligned with epic #631. Read-only: every
destination is a GET view or a Phase 1 placeholder. No mutation links.
Nav groups follow the #631 Phase 1 information architecture: Health, Traffic,
Runtime/Sessions, Projects, Inventory, Timeline, Policy (placeholder), and
Insights (placeholder). Later-phase surfaces are declared as ``stub`` items and
backed by ``STUB_PAGES`` so their nav links resolve to a graceful placeholder
instead of a 404.
"""
from __future__ import annotations
from dataclasses import dataclass
@dataclass(frozen=True)
class NavItem:
"""A single navigation destination.
``status`` is ``"live"`` for implemented views and ``"stub"`` for Phase 1
placeholders whose backing view lands in a later child issue.
"""
href: str
label: str
status: str = "live"
@dataclass(frozen=True)
class NavGroup:
label: str
items: tuple[NavItem, ...]
NAV_GROUPS: tuple[NavGroup, ...] = (
NavGroup("Health", (
NavItem("/health", "Liveness"),
)),
NavGroup("Traffic", (
NavItem("/queue", "Queue"),
NavItem("/leases", "Leases"),
NavItem("/actions", "Actions"),
)),
NavGroup("Runtime/Sessions", (
NavItem("/runtime", "Runtime health"),
NavItem("/sessions", "Sessions", "stub"),
)),
NavGroup("Projects", (
NavItem("/projects", "Projects"),
)),
NavGroup("Inventory", (
NavItem("/inventory", "Inventory", "stub"),
NavItem("/worktrees", "Worktrees"),
)),
NavGroup("Timeline", (
NavItem("/timeline", "Timeline", "stub"),
)),
NavGroup("Policy", (
NavItem("/policy", "Policy", "stub"),
NavItem("/prompts", "Prompts"),
)),
NavGroup("Insights", (
NavItem("/insights", "Insights", "stub"),
NavItem("/audit", "Audit"),
)),
)
# Phase 1 placeholder destinations whose backing views land in later child
# issues of epic #631. Each maps a path to (title, description). Routes are
# registered so nav links resolve to a graceful, read-only stub page.
STUB_PAGES: dict[str, tuple[str, str]] = {
"/sessions": (
"Sessions",
"Active session, capability, and role inventory. Backed by the unified "
"inventory API (#636) once it lands.",
),
"/inventory": (
"Inventory",
"Unified sessions, leases, locks, namespaces, and worktree inventory. "
"Backed by the Phase 1 inventory API (#636).",
),
"/timeline": (
"Timeline",
"Workflow event timeline across issues and PRs. A later Phase 1 surface.",
),
"/policy": (
"Policy",
"Capability and role policy surface. Placeholder until a later phase.",
),
"/insights": (
"Insights",
"Aggregate operational insights and trends. Placeholder until a later "
"phase.",
),
}
def iter_nav_items():
"""Yield every ``NavItem`` across all groups in declared order."""
for group in NAV_GROUPS:
for item in group.items:
yield item
def nav_hrefs() -> tuple[str, ...]:
"""Return every navigation href in declared order."""
return tuple(item.href for item in iter_nav_items())
+906
View File
@@ -0,0 +1,906 @@
"""Workflow-event and conversation timeline model (#637, Phase 1).
Operators cannot browse a unified timeline of workflow events, decisions,
tool calls, and handoffs: the evidence is scattered across control-plane
events, Gitea canonical handoff comments, and local logs. This module defines
one durable, versioned event schema and per-source adapters that normalise
those scattered records into a single ``WorkflowEvent`` stream, plus a
read-only query layer (filter by issue / PR / session, stable ordering,
pagination) that the ``/api/v1/timeline`` route serves.
Design rules honoured here:
- **Read-only.** Sources are read; nothing is mutated. The control-plane
database is opened through a ``mode=ro`` URI so a missing or unwritable DB
degrades to a reason instead of creating directories or running migrations.
- **Fail-soft per source.** An unavailable source degrades to a status with a
reason rather than raising, and a source that could not run is never
rendered as an empty-and-healthy timeline.
- **Answerable filters only.** Each source declares which filter dimensions it
can actually answer. A filter dimension no source that ran can carry is
refused with an explicit reason rather than silently matching nothing: an
empty page from an unanswerable filter reads to an operator as "no such
activity", which is a different — and false — statement.
- **Redaction at the boundary, fail closed.** Every free-text field (event
messages, redacted tool arguments, decision/proof text) is run through the
console redaction policy before it leaves this module, and *before* any
structured value is derived from it — evidence references are extracted from
redacted text, then independently revalidated before serialization. An
unredactable value becomes the placeholder, and a value that cannot be proven
safe is dropped — an unredacted payload is never emitted, and a generation
error never drops raw data to a caller or a log.
- **Stable ordering.** Events sort by ``(timestamp, source_rank, event_key)``
with a deterministic tiebreak, so pagination is stable across calls and
events with equal or missing timestamps keep a fixed order.
Non-goals (from the issue): no full chat replay, no mutation of historical
events, no unredacted tool-argument storage.
"""
from __future__ import annotations
import re
import sqlite3
from dataclasses import dataclass, replace
from datetime import datetime, timezone
from typing import Any, Callable, Iterable
import control_plane_db
from webui import console_redaction
# The schema is versioned so consumers can branch on shape. Bump on any
# breaking change to WorkflowEvent's serialized form.
TIMELINE_SCHEMA_VERSION = 1
# Known event sources and their deterministic ordering rank. When two events
# carry the same timestamp, the source rank breaks the tie before the
# per-source event key, so a control-plane event and a handoff comment minted
# in the same second always sort in a fixed order.
SOURCE_CONTROL_PLANE = "control_plane"
SOURCE_GITEA_HANDOFF = "gitea_handoff"
_SOURCE_RANK = {
SOURCE_CONTROL_PLANE: 0,
SOURCE_GITEA_HANDOFF: 1,
}
# The filter dimensions the query layer accepts.
FILTER_ISSUE = "issue"
FILTER_PR = "pr"
FILTER_SESSION = "session"
# Which dimensions each source can actually answer. This is a property of the
# underlying records, not of the query code: the control-plane ``events`` table
# is (event_id, work_item_id, event_type, message, created_at) and carries no
# session identity at all, so no control-plane event can ever match a session
# filter. A CTH handoff comment can declare its session as a field, so the
# handoff source answers all three. Filtering on a dimension the surviving
# sources cannot carry is refused in ``load_timeline`` rather than answered
# with an empty page.
_SOURCE_FILTER_SUPPORT: dict[str, tuple[str, ...]] = {
SOURCE_CONTROL_PLANE: (FILTER_ISSUE, FILTER_PR),
SOURCE_GITEA_HANDOFF: (FILTER_ISSUE, FILTER_PR, FILTER_SESSION),
}
# Why a source cannot answer a dimension, for the refusal reason an operator reads.
_SOURCE_FILTER_LIMITS: dict[tuple[str, str], str] = {
(SOURCE_CONTROL_PLANE, FILTER_SESSION): (
"control-plane events carry no session identity "
"(the events table has no session column)"
),
}
# A timestamp far in the future so events with no parseable timestamp sort
# last (after everything real) instead of first, without raising.
_MISSING_TS_SORT = "9999-12-31T23:59:59Z"
def _parse_ts(value: str | None) -> str | None:
"""Normalise a timestamp to ``...Z`` UTC, or None when unparseable."""
if not value:
return None
text = str(value).strip()
if not text:
return None
candidate = text[:-1] + "+00:00" if text.endswith("Z") else text
try:
parsed = datetime.fromisoformat(candidate)
except ValueError:
return None
if parsed.tzinfo is None:
parsed = parsed.replace(tzinfo=timezone.utc)
return parsed.astimezone(timezone.utc).replace(microsecond=0).isoformat().replace("+00:00", "Z")
def _redact(value: Any) -> Any:
"""Redact a single free-text field, failing closed to the placeholder."""
if value is None:
return None
return console_redaction.redact_text(str(value))
@dataclass(frozen=True)
class WorkflowEvent:
"""One normalised timeline event.
Every field is optional except ``source``/``event_type``/``event_key``
because sources carry different subsets. The class is frozen so an adapted
event is an immutable record; a consumer that needs a variant builds a new
one rather than mutating history.
"""
source: str
event_type: str
event_key: str
timestamp: str | None = None
actor: str | None = None
role: str | None = None
issue_number: int | None = None
pr_number: int | None = None
session_id: str | None = None
tool_name: str | None = None
decision: str | None = None
message: str | None = None
correlation_id: str | None = None
evidence_refs: tuple[str, ...] = ()
sensitive: bool = False
def sort_key(self) -> tuple[str, int, str]:
return (
self.timestamp or _MISSING_TS_SORT,
_SOURCE_RANK.get(self.source, 99),
self.event_key,
)
def to_dict(self) -> dict[str, Any]:
return {
"source": self.source,
"event_type": self.event_type,
"event_key": self.event_key,
"timestamp": self.timestamp,
"actor": self.actor,
"role": self.role,
"issue_number": self.issue_number,
"pr_number": self.pr_number,
"session_id": self.session_id,
"tool_name": self.tool_name,
"decision": self.decision,
"message": self.message,
"correlation_id": self.correlation_id,
"evidence_refs": list(self.evidence_refs),
"sensitive": self.sensitive,
}
# --------------------------------------------------------------------------- #
# Adapters — pure functions from a source's raw records to WorkflowEvents. #
# Each is total: a malformed record is skipped, never raised on. #
# --------------------------------------------------------------------------- #
# Event types whose payload is treated as sensitive and always redaction-hard
# (they can carry lease/session provenance or tool arguments).
_SENSITIVE_EVENT_HINTS = ("lease", "capability", "token", "auth", "secret")
# Reference tokens (issue/PR/comment ids) and SHAs parsed out of proof text.
_EVIDENCE_REF_RE = re.compile(r"(?:#|PR\s*#?|issue\s*#?|comment\s*#?)(\d+)", re.IGNORECASE)
# A commit reference is only recognised when the text *declares* it as one.
# A bare lowercase hex run is not evidence of anything: at 40 characters it is
# exactly the shape of a Gitea personal access token, and at 7 it also matches
# ordinary words such as "defaced". Requiring an anchoring keyword keeps real
# references ("commit abc1234", "at head a209756...", "base caaae9b6") usable
# while refusing to lift an undeclared secret-shaped run out of free text.
_SHA_RE = re.compile(
r"(?i:\b(?:commit|sha|head|base|parent|revision|rev|merge[- ]base)\b[\s:=@#]*)"
r"([0-9a-f]{7,40})\b"
)
# Shapes a serialized evidence reference is allowed to take. Anything else is
# dropped rather than emitted.
_REF_ISSUE_SHAPE = re.compile(r"^#[0-9]{1,9}$")
_REF_SHA_SHAPE = re.compile(r"^[0-9a-f]{7,40}$")
# A long undelimited hex run with no declaring context is treated as credential
# material wherever it appears, never as an identifier.
_BARE_SECRET_SHAPE = re.compile(r"^[0-9a-f]{32,}$")
# An event type reads like an identifier, but a stored one is externally
# influenced: any producer that writes the control-plane ``events`` table
# chooses the string. It reaches ``to_dict`` verbatim, so it is validated here
# rather than trusted because of where it came from.
_CP_EVENT_TYPE_SHAPE = re.compile(r"^[A-Za-z][A-Za-z0-9._:+-]{0,63}$")
# Emitted in place of a value that cannot be proven safe. Deliberately not a
# plausible workflow type: an unsafe value is refused, never quietly rewritten
# into a different valid-looking one that would misdescribe the record.
UNSAFE_EVENT_TYPE = "unsafe:redacted"
# Emitted for a CTH heading that is not a declared member of ``CTH_TYPES``. The
# contract is enforced on write (``format_cth_body``) and on assess; the read
# path the timeline uses enforces it too rather than assuming it was.
UNKNOWN_HANDOFF_EVENT_TYPE = "handoff:unrecognized"
# A source record id is a plain integer in both sources it comes from: the
# control-plane ``events`` primary key and a Gitea comment id. ``event_key`` is
# serialized verbatim and is the pagination tiebreak, so anything else is
# refused rather than interpolated into it.
_RECORD_ID_SHAPE = re.compile(r"^[0-9]{1,19}$")
def _kind_to_numbers(kind: str | None, number: int | None) -> tuple[int | None, int | None]:
"""Map a control-plane work-item (kind, number) to (issue_no, pr_no)."""
if number is None:
return (None, None)
if kind == "pr":
return (None, int(number))
if kind == "issue":
return (int(number), None)
return (None, None)
def _correlation_for(kind: str | None, number: int | None) -> str | None:
if number is None or kind not in ("issue", "pr"):
return None
return f"{kind}#{number}"
def _extract_evidence_refs(*texts: str | None) -> tuple[str, ...]:
"""Extract issue/PR and declared-commit references from **redacted** text.
Callers must pass text that has already been through :func:`_redact`; this
function derives a structured field from its input, so extracting ahead of
redaction would republish whatever redaction was about to remove. Every
reference is revalidated by :func:`_validated_evidence_refs` before it is
serialized.
"""
refs: list[str] = []
for text in texts:
if not text:
continue
for match in _EVIDENCE_REF_RE.finditer(text):
token = f"#{match.group(1)}"
if token not in refs:
refs.append(token)
for match in _SHA_RE.finditer(text):
token = match.group(1)
if token not in refs:
refs.append(token)
return tuple(refs)
def _validated_evidence_refs(refs: Iterable[str]) -> tuple[tuple[str, ...], bool]:
"""Independently revalidate references immediately before serialization.
Extraction is not trusted on its own. A reference survives only when it has
a known reference shape and is unchanged by a second redaction pass — a
value the redaction policy would alter is credential material that must not
be emitted as a structured field. A full 40-character SHA stays usable
because extraction only accepts a hex run the source text explicitly
declared as a commit. Returns ``(safe_refs, dropped_any)``; ``dropped_any``
marks the event sensitive so the drop is visible rather than silent.
"""
safe: list[str] = []
dropped = False
for ref in refs or ():
try:
token = str(ref).strip()
if not token:
continue
recognised = bool(_REF_ISSUE_SHAPE.match(token) or _REF_SHA_SHAPE.match(token))
if not recognised:
dropped = True
continue
if _redact(token) != token:
dropped = True
continue
if token not in safe:
safe.append(token)
except Exception:
# Fail closed: a reference that cannot be proven safe is dropped.
dropped = True
continue
return (tuple(safe), dropped)
def _safe_session_id(value: Any) -> str | None:
"""Return a session identifier only when it is safe to emit.
The value is authoritative source data — a session the record names for
itself — but it is still free text. It is dropped when redaction alters it
or when it is a bare secret-shaped hex run, so a credential parked in a
session field can never reach the payload or be echoed back by a filter.
"""
if value is None:
return None
text = str(value).strip()
if not text:
return None
if _BARE_SECRET_SHAPE.match(text):
return None
return text if _redact(text) == text else None
def _safe_record_id(value: Any) -> str | None:
"""Return a source record id only when it is a plain numeric identifier.
``event_key`` is serialized verbatim and is the deterministic pagination
tiebreak, so an id is interpolated into it only when it has the shape both
real sources actually produce. A record whose identity cannot be trusted is
refused by the caller rather than keyed on.
"""
if value is None or isinstance(value, bool):
return None
if isinstance(value, int):
return str(value)
text = str(value).strip()
return text if _RECORD_ID_SHAPE.match(text) else None
def _safe_cp_event_type(value: Any) -> tuple[str, bool]:
"""Validate a stored control-plane event type. Returns ``(type, unsafe)``.
The stored value is externally influenced — whichever producer wrote the
``events`` row chose the string — and ``to_dict`` serializes it verbatim, so
it passes a boundary of its own instead of relying on the one ``message``
passes. A value survives only when it is an ordinary identifier, is not a
bare secret-shaped hex run, and is unchanged by a redaction pass. Anything
else fails closed to :data:`UNSAFE_EVENT_TYPE`: the record stays visible as
an audit entry, but the value itself is never republished — not verbatim,
not partially sanitized, and not rewritten into some other valid-looking
type that would misdescribe what happened.
"""
text = ("" if value is None else str(value)).strip()
if not text:
return ("", False)
if _BARE_SECRET_SHAPE.match(text):
return (UNSAFE_EVENT_TYPE, True)
if not _CP_EVENT_TYPE_SHAPE.match(text):
return (UNSAFE_EVENT_TYPE, True)
if _redact(text) != text:
return (UNSAFE_EVENT_TYPE, True)
return (text, False)
def _safe_echo(value: Any) -> Any:
"""Guard a scalar that is echoed back rather than derived from a record.
Query scope and filter values are caller-supplied and are reflected in the
response so an operator can see what was asked. Reflection is still
emission: a value redaction would alter, or a bare secret-shaped hex run, is
replaced by the placeholder instead of being echoed verbatim. Ordinary
scope and filter values pass through untouched.
"""
if value is None or isinstance(value, (int, bool)):
return value
text = str(value)
if _BARE_SECRET_SHAPE.match(text.strip()):
return console_redaction.REDACTED
return _redact(text)
def adapt_cp_events(rows: Iterable[dict[str, Any]]) -> list[WorkflowEvent]:
"""Adapt control-plane ``events`` rows (joined to work_items) into events.
Each row is expected to carry ``event_id``, ``event_type``, ``message``,
``created_at`` and the joined work-item ``kind``/``number``. Rows missing
an id or type are skipped so a partially written table never raises.
"""
events: list[WorkflowEvent] = []
for row in rows or []:
try:
event_id = _safe_record_id(row.get("event_id"))
raw_event_type = (row.get("event_type") or "").strip()
if event_id is None or not raw_event_type:
continue
# The stored type is source data, not a trusted constant: validate
# it before it is serialized, exactly as `message` below is redacted
# before it is serialized.
event_type, event_type_unsafe = _safe_cp_event_type(raw_event_type)
kind = row.get("kind")
number = row.get("number")
issue_no, pr_no = _kind_to_numbers(kind, number)
sensitive = event_type_unsafe or any(
hint in raw_event_type.lower() for hint in _SENSITIVE_EVENT_HINTS
)
events.append(
WorkflowEvent(
source=SOURCE_CONTROL_PLANE,
event_type=event_type,
event_key=f"cp:{event_id}",
timestamp=_parse_ts(row.get("created_at")),
issue_number=issue_no,
pr_number=pr_no,
# No session_id: the control-plane events table is
# (event_id, work_item_id, event_type, message, created_at)
# and records no session. Inventing one from the work item
# or the message text would be a guess, so this source
# declares the session dimension unsupported instead
# (_SOURCE_FILTER_SUPPORT) and the query layer refuses a
# session filter it cannot honestly answer.
message=_redact(row.get("message")),
correlation_id=_correlation_for(kind, number),
sensitive=sensitive,
)
)
except Exception:
# A single malformed row must not sink the whole adaptation.
continue
return events
def adapt_cth_comments(
comments: Iterable[dict[str, Any]],
*,
kind: str,
number: int,
) -> list[WorkflowEvent]:
"""Adapt Gitea Canonical Thread Handoff (CTH) comments into events.
Only comments that parse as a CTH (``canonical_thread_handoff.parse_cth_comment``)
become events; ordinary comments are ignored. ``kind``/``number`` scope the
events to the issue or PR the comments belong to.
"""
# Imported lazily so this module has no import-time dependency on the
# handoff parser when only the control-plane adapter is used.
from canonical_thread_handoff import is_known_cth_type, parse_cth_comment
# ``kind``/``number`` are interpolated into event_key and correlation_id, so
# they are normalised once here. A scope this adapter cannot express is
# refused outright rather than serialized into an identifier.
kind = (kind or "").strip().lower()
if kind not in ("issue", "pr"):
return []
try:
number = int(number)
except (TypeError, ValueError):
return []
issue_no, pr_no = _kind_to_numbers(kind, number)
correlation = _correlation_for(kind, number)
events: list[WorkflowEvent] = []
for comment in comments or []:
try:
body = comment.get("body") or ""
parsed = parse_cth_comment(body)
if not parsed:
continue
fields = parsed.get("fields") or {}
cth_type = parsed.get("cth_type") or ""
comment_id = _safe_record_id(comment.get("id"))
if comment_id is None:
continue
# The CTH heading is free text: the parser accepts whatever follows
# "## CTH:", and only the write and assess paths check it against
# the contract. Check it here too — an unrecognised heading is
# reported as such rather than serialized into event_type, so
# arbitrary, malformed, or secret-shaped heading content has no way
# through. Declared types are preserved exactly.
cth_type_known = is_known_cth_type(cth_type)
# Redaction runs first, and every derived value is taken from the
# redacted text — deriving evidence refs from the raw proof would
# re-emit exactly what redaction was about to remove.
decision = _redact(fields.get("decision"))
proof = _redact(fields.get("proof"))
next_action = _redact(fields.get("next action"))
refs, refs_dropped = _validated_evidence_refs(
_extract_evidence_refs(proof, decision)
)
events.append(
WorkflowEvent(
source=SOURCE_GITEA_HANDOFF,
event_type=(
f"handoff:{cth_type.strip()}"
if cth_type_known
else UNKNOWN_HANDOFF_EVENT_TYPE
),
event_key=f"cth:{kind}:{number}:{comment_id}",
timestamp=_parse_ts(comment.get("created_at")),
actor=_redact((comment.get("user") or {}).get("login")),
role=_redact(fields.get("next owner")),
issue_number=issue_no,
pr_number=pr_no,
# A CTH names its own session when the producer records one;
# it is read from that declared field, never inferred from
# unrelated text.
session_id=_safe_session_id(fields.get("session")),
decision=decision,
message=next_action or _redact(fields.get("status")),
correlation_id=correlation,
evidence_refs=refs,
sensitive=refs_dropped or not cth_type_known,
)
)
except Exception:
continue
return events
# --------------------------------------------------------------------------- #
# Read-only control-plane event source. #
# --------------------------------------------------------------------------- #
_CP_EVENTS_QUERY = """
SELECT e.event_id AS event_id,
e.event_type AS event_type,
e.message AS message,
e.created_at AS created_at,
w.kind AS kind,
w.number AS number
FROM events e
JOIN work_items w ON e.work_item_id = w.work_item_id
WHERE w.remote = ? AND w.org = ? AND w.repo = ?
"""
@dataclass(frozen=True)
class SourceStatus:
"""Fail-soft status for one timeline source.
``supported_filters`` states which filter dimensions this source's records
can carry; ``unsupported_filters`` names the requested dimensions it cannot,
so an operator can see *why* a source contributed nothing rather than being
left to read an empty list as an absence of activity.
"""
name: str
ok: bool
reason: str | None = None
count: int = 0
supported_filters: tuple[str, ...] = ()
unsupported_filters: tuple[str, ...] = ()
def to_dict(self) -> dict[str, Any]:
return {
"name": self.name,
"ok": self.ok,
"reason": self.reason,
"count": self.count,
"supported_filters": list(self.supported_filters),
"unsupported_filters": list(self.unsupported_filters),
}
def _cp_status(*, ok: bool, reason: str | None = None, count: int = 0) -> SourceStatus:
return SourceStatus(
SOURCE_CONTROL_PLANE,
ok=ok,
# A failure reason is serialized like any other field and is often an
# exception string carrying a path or a transport error, so it crosses
# the redaction boundary too. Static reasons pass through unchanged.
reason=_redact(reason),
count=count,
supported_filters=_SOURCE_FILTER_SUPPORT[SOURCE_CONTROL_PLANE],
)
def _handoff_status(*, ok: bool, reason: str | None = None, count: int = 0) -> SourceStatus:
return SourceStatus(
SOURCE_GITEA_HANDOFF,
ok=ok,
# Same boundary as the control-plane status: this reason can quote an
# error raised by a live authenticated fetch.
reason=_redact(reason),
count=count,
supported_filters=_SOURCE_FILTER_SUPPORT[SOURCE_GITEA_HANDOFF],
)
def read_cp_events(
*,
remote: str,
org: str,
repo: str,
db_path: str | None = None,
) -> tuple[list[WorkflowEvent], SourceStatus]:
"""Read scoped control-plane events read-only. Never creates the DB.
Opens the SQLite file through a ``mode=ro`` URI: a health/timeline read
must never create directories or run the schema migration that
``ControlPlaneDB()`` performs on construction. A missing or unreadable DB
degrades to a status with a reason.
"""
path = (db_path or control_plane_db.default_db_path()).strip()
conn: sqlite3.Connection | None = None
try:
conn = sqlite3.connect(f"file:{path}?mode=ro", uri=True)
conn.row_factory = sqlite3.Row
cursor = conn.execute(_CP_EVENTS_QUERY, (remote, org, repo))
rows = [dict(r) for r in cursor.fetchall()]
except sqlite3.OperationalError as exc:
return ([], _cp_status(ok=False, reason=f"control-plane DB unavailable: {exc}"))
except sqlite3.Error as exc:
return ([], _cp_status(ok=False, reason=f"control-plane read failed: {exc}"))
finally:
if conn is not None:
conn.close()
events = adapt_cp_events(rows)
return (events, _cp_status(ok=True, count=len(events)))
# --------------------------------------------------------------------------- #
# Filter, sort, paginate. #
# --------------------------------------------------------------------------- #
def filter_events(
events: Iterable[WorkflowEvent],
*,
issue: int | None = None,
pr: int | None = None,
session: str | None = None,
) -> list[WorkflowEvent]:
"""Filter events by issue number, PR number, and/or session id.
Filters are conjunctive. A filter that names a dimension an event does not
carry excludes that event (an issue filter excludes PR-only events).
"""
out: list[WorkflowEvent] = []
for ev in events:
if issue is not None and ev.issue_number != issue:
continue
if pr is not None and ev.pr_number != pr:
continue
if session is not None and ev.session_id != session:
continue
out.append(ev)
return out
def sort_events(events: Iterable[WorkflowEvent]) -> list[WorkflowEvent]:
"""Return events in stable timeline order (ascending)."""
return sorted(events, key=lambda ev: ev.sort_key())
@dataclass(frozen=True)
class TimelinePage:
"""One page of the sorted, filtered timeline."""
events: tuple[WorkflowEvent, ...]
total: int
limit: int
offset: int
@property
def next_offset(self) -> int | None:
nxt = self.offset + len(self.events)
return nxt if nxt < self.total else None
def to_dict(self) -> dict[str, Any]:
return {
"events": [ev.to_dict() for ev in self.events],
"pagination": {
"total": self.total,
"limit": self.limit,
"offset": self.offset,
"returned": len(self.events),
"next_offset": self.next_offset,
"has_more": self.next_offset is not None,
},
}
_MAX_LIMIT = 500
_DEFAULT_LIMIT = 50
def _coerce_bounds(limit: int | None, offset: int | None) -> tuple[int, int]:
try:
lim = int(limit) if limit is not None else _DEFAULT_LIMIT
except (TypeError, ValueError):
lim = _DEFAULT_LIMIT
try:
off = int(offset) if offset is not None else 0
except (TypeError, ValueError):
off = 0
lim = max(1, min(lim, _MAX_LIMIT))
off = max(0, off)
return (lim, off)
def paginate(events: list[WorkflowEvent], *, limit: int | None, offset: int | None) -> TimelinePage:
lim, off = _coerce_bounds(limit, offset)
window = events[off : off + lim]
return TimelinePage(events=tuple(window), total=len(events), limit=lim, offset=off)
# --------------------------------------------------------------------------- #
# Composition — load_timeline aggregates all sources, fail-soft. #
# --------------------------------------------------------------------------- #
# A comment source is a callable that, given (kind, number), returns the raw
# Gitea comment list for that issue/PR. The route supplies a live fail-soft
# fetcher; tests supply a fixture. When None, the handoff source is reported as
# not-run (never silently empty-and-healthy).
CommentSource = Callable[[str, int], list[dict[str, Any]]]
@dataclass(frozen=True)
class TimelineSnapshot:
"""One answered timeline query.
``ok`` is False when the query could not be answered as asked — currently
when a requested filter dimension no surviving source can carry was
supplied. The page is then empty *and* the snapshot says so, because an
``ok`` empty page is a claim that no such activity exists.
"""
schema_version: int
remote: str
org: str
repo: str
filters: dict[str, Any]
page: TimelinePage
sources: tuple[SourceStatus, ...]
ok: bool = True
error: dict[str, Any] | None = None
def to_dict(self) -> dict[str, Any]:
return {
"ok": self.ok,
"error": self.error,
"schema_version": self.schema_version,
# Scope and filters are echoed caller input, not derived record
# data. Reflecting a value is still emitting it, so both cross the
# same boundary; ordinary scope and filter values are unchanged.
"scope": {
"remote": _safe_echo(self.remote),
"org": _safe_echo(self.org),
"repo": _safe_echo(self.repo),
},
"filters": {key: _safe_echo(value) for key, value in self.filters.items()},
"sources": [s.to_dict() for s in self.sources],
**self.page.to_dict(),
}
def _unanswerable_reasons(
statuses: Iterable[SourceStatus], unanswerable: Iterable[str]
) -> list[dict[str, str]]:
"""Explain, per source, why each unanswerable dimension went unanswered."""
out: list[dict[str, str]] = []
for status in statuses:
for dim in unanswerable:
if dim not in status.supported_filters:
reason = _SOURCE_FILTER_LIMITS.get(
(status.name, dim), f"this source's records carry no {dim} identity"
)
elif not status.ok:
reason = (
f"this source can carry {dim} but did not run: "
f"{status.reason or 'unavailable'}"
)
else:
continue
out.append({"source": status.name, "filter": dim, "reason": reason})
return out
def load_timeline(
*,
remote: str,
org: str,
repo: str,
issue: int | None = None,
pr: int | None = None,
session: str | None = None,
limit: int | None = None,
offset: int | None = None,
db_path: str | None = None,
comment_source: CommentSource | None = None,
) -> TimelineSnapshot:
"""Aggregate every timeline source into one filtered, paginated snapshot.
Sources are read independently and fail soft: an unavailable source
contributes a ``SourceStatus`` with ``ok=False`` and a reason, and never
collapses the whole timeline. The handoff source only runs when a specific
issue or PR is requested (a handoff comment belongs to one thread) and a
``comment_source`` is available; otherwise it is reported as ``not run``
rather than as an empty-and-healthy source.
A filter dimension that no surviving source can carry — a ``session``
filter when the only source that ran is the control plane, whose events
record no session — is refused with ``ok=False`` and a structured error
instead of being answered with an empty page.
"""
all_events: list[WorkflowEvent] = []
statuses: list[SourceStatus] = []
cp_events, cp_status = read_cp_events(remote=remote, org=org, repo=repo, db_path=db_path)
all_events.extend(cp_events)
statuses.append(cp_status)
# Gitea handoff comments are thread-scoped: only fetch when the caller
# narrowed to one issue or PR, and only when a source was provided.
handoff_target: tuple[str, int] | None = None
if pr is not None:
handoff_target = ("pr", pr)
elif issue is not None:
handoff_target = ("issue", issue)
if handoff_target is None:
statuses.append(
_handoff_status(
ok=False,
reason="not run: handoff comments are thread-scoped; filter by issue or pr to include them",
)
)
elif comment_source is None:
statuses.append(
_handoff_status(
ok=False,
reason="not run: no comment source configured for this timeline read",
)
)
else:
kind, number = handoff_target
try:
comments = comment_source(kind, number) or []
handoff_events = adapt_cth_comments(comments, kind=kind, number=number)
all_events.extend(handoff_events)
statuses.append(_handoff_status(ok=True, count=len(handoff_events)))
except Exception as exc: # fail soft: a fetch/parse error degrades this source only
statuses.append(_handoff_status(ok=False, reason=f"handoff source failed: {exc}"))
requested = tuple(
name
for name, value in ((FILTER_ISSUE, issue), (FILTER_PR, pr), (FILTER_SESSION, session))
if value is not None
)
statuses = [
replace(
status,
unsupported_filters=tuple(
dim for dim in requested if dim not in status.supported_filters
),
)
for status in statuses
]
filters = {"issue": issue, "pr": pr, "session": session}
# A dimension is answerable only if a source that actually ran can carry it.
# If none can, refuse: an empty page would assert "no such activity", which
# is a claim this timeline is not in a position to make.
answerable: set[str] = set()
for status in statuses:
if status.ok:
answerable.update(status.supported_filters)
unanswerable = tuple(dim for dim in requested if dim not in answerable)
if unanswerable:
return TimelineSnapshot(
schema_version=TIMELINE_SCHEMA_VERSION,
remote=remote,
org=org,
repo=repo,
filters=filters,
page=paginate([], limit=limit, offset=offset),
sources=tuple(statuses),
ok=False,
error={
"code": "filter_not_supported",
"unsupported_filters": list(unanswerable),
"detail": (
"no timeline source that ran can answer "
+ ", ".join(f"'{dim}'" for dim in unanswerable)
+ "; the result is refused rather than returned empty"
),
"sources": _unanswerable_reasons(statuses, unanswerable),
},
)
filtered = filter_events(all_events, issue=issue, pr=pr, session=session)
ordered = sort_events(filtered)
page = paginate(ordered, limit=limit, offset=offset)
return TimelineSnapshot(
schema_version=TIMELINE_SCHEMA_VERSION,
remote=remote,
org=org,
repo=repo,
filters=filters,
page=page,
sources=tuple(statuses),
)
def snapshot_to_dict(snapshot: TimelineSnapshot) -> dict[str, Any]:
return snapshot.to_dict()