Compare commits

..
208 changed files with 522 additions and 73643 deletions
-23
View File
@@ -1,23 +0,0 @@
# Agent Rules
## Project charter (authoritative)
Before inspecting work, selecting a task, claiming work, implementing, reviewing,
approving, or merging, read the full project charter:
**`PROJECT_CHARTER.md`**
That file is the ultimate source of truth for purpose, direction, principles,
non-goals, success criteria, and change-control. This file only points to it;
do not treat this file, session memory, or prompts as a substitute for the charter.
If proposed work would change the projects fundamental direction or any
non-negotiable principle in the charter, stop and present the change to the
human maintainer for an explicit decision recorded in Gitea.
## Issue-first workflow
- Never fix code directly. A Gitea issue must be created first, and all fix work happens under that issue (branch, PR, review, merge) per the canonical workflow.
- No implementation, remediation, refactor, operational change, or charter amendment may begin without an open Gitea issue authorizing that work.
- Investigation may prepare or validate an issue; it must not silently become implementation.
- Every implementation PR must reference its governing issue.
-233
View File
@@ -1,233 +0,0 @@
# MCP Control Plane — Project Charter
```
Charter-ID: mcp-control-plane
Charter-Version: 1.0
Status: approved
Project: Scaled-Tech-Consulting/Gitea-Tools
Repository-Binding: Scaled-Tech-Consulting/Gitea-Tools
Approved-By: human-operator
Approved-At: 2026-07-31
Governing-Issue: #987
Charter-Revision: (set to merge commit SHA)
Content-Hash: (set by enforcement tooling when enabled)
```
This file is the authoritative project charter for Gitea-Tools / MCP Control Plane.
It is the ultimate source of truth for purpose, operating model, non-negotiable
principles, authority boundaries, non-goals, definition of success, and human
change-control. Issues authorize units of work; pull requests implement issues.
Session memory and prompts are not authoritative governance.
## Purpose
The purpose of this project is to allow one or more general-purpose LLMs to work autonomously and safely on a software project.
An LLM is not permanently assigned to be an author, reviewer, merger, controller, or reconciler. It begins each work cycle without a predetermined role.
The LLM examines the live project state, decides for itself what work is most valuable, and only then adopts the role required to perform that task.
## Fundamental workflow
Each LLM follows this cycle:
1. Start as a general, uncommitted worker.
2. Inspect the complete authoritative project state in Gitea.
3. Identify the work that is currently available, including:
* Issues ready for implementation
* Pull requests awaiting review
* Approved pull requests ready to merge
* Change-requested work needing remediation
* Prerequisite work blocking more important work
* Abandoned or orphaned work requiring reconciliation (expired leases, stale claims, half-finished PRs)
4. Compare those tasks using:
* Priority
* Dependencies
* Urgency
* Project impact
* Readiness
* Active claims (work under a live lock or lease is not available)
5. Independently choose the task it believes is most valuable.
6. Derive the required role from the chosen task.
7. Ask the MCP to verify that the task and role are currently safe and permitted.
8. Claim the work using the appropriate lock or lease.
9. Complete one bounded work cycle.
10. Release ownership and finish the required handoff.
11. Return to the uncommitted state.
12. Inspect the project again and make a new independent decision.
The governing order is:
**Inspect project → choose task → derive role → validate → claim → work → release → reassess**
## Bounded work cycle
A bounded work cycle is the smallest unit of work that leaves the project in a coherent, handoff-ready state, completed within the duration of a single lease. Examples: implementing one issue as one PR, reviewing one PR, merging one approved PR, remediating one round of change requests, or reconciling one abandoned artifact.
A work cycle never spans multiple leases. If the work cannot be completed within the lease, the LLM must bring the artifact to a coherent stopping point, record its state in Gitea (not in session memory), and release the claim. Continuation is a new task, available to any worker.
## Contention is normal
Multiple LLMs inspecting the same state with the same criteria will often converge on the same task. A failed claim is therefore an expected, routine outcome — not an error.
On claim failure, the LLM does not retry the same claim. It re-evaluates from fresh state and selects the next most valuable eligible task. Active claims must be visible in the project state so that workers can route around in-progress work before attempting a claim.
## Abandoned work and reconciliation
Leases expire. Workers fail mid-cycle. The resulting orphaned branches, stale claims, and half-finished PRs are first-class work items, discoverable in Gitea like any other task.
Reconciling an abandoned artifact is a task like any other: an LLM may choose it, derive the reconciler role for that one cycle, and release the role when done. No LLM is ever permanently a reconciler.
Expired leases become discoverable abandoned-work tasks. Detection may be passive (workers observing expired claims during ordinary inspection) or assisted by control-plane signals; either approach is acceptable provided reconciliation remains ordinary task selection under the fundamental workflow.
## Responsibilities
| Component | Responsibility |
|--------------------|----------------|
| LLM | Understand the project, compare available work, and choose what it wants to do |
| Gitea | Store the authoritative issues, PRs, priorities, dependencies, decisions, claims, and history |
| Gitea MCP | Expose live facts from Gitea and provide sanctioned workflow operations |
| Control plane | Enforce identity, capability, ownership, and safety boundaries (including claim validation and rejection of unsafe or stale choices) |
| Locks and leases | Prevent conflicting LLMs from performing the same exclusive work |
| Human maintainer | Set priorities, approve charter changes, and resolve escalations |
The MCP may reject an unsafe or stale choice. It must not decide which task the LLM wants or permanently assign the LLM a role.
**Guardrail:** Rejection policy is validation, not steering. If a pattern of rejections effectively routes workers toward particular tasks, the MCP has become a dispatcher and the design has been violated. Rejection rules must themselves be durable, reviewable artifacts in Gitea so that patterns of rejection can be audited against this test.
## Identity and independence
Every worker session operates under a distinct identity issued by the control plane. Authorship, review, approval, and merge actions are attributed to that identity in Gitea.
Independence is defined at the identity level: the identity that authored (or last pushed to) a PR cannot review, approve, or merge it. A different identity — even one backed by the same underlying model — satisfies independence.
This is a deliberate, accepted limitation: same-model reviewers are epistemically correlated and may share blind spots. Identity-level independence is the enforced floor; stronger diversity (different models, human review) may be layered on for designated-critical changes but is not required by this charter.
## Approval, parity, and remediation
* A PR may be merged only after valid independent approval at its exact current head commit (head-SHA parity between the approved commit and the merged commit).
* Any new commit to the PR branch — including rebases and conflict resolutions — invalidates all prior approvals. Re-approval at the new head is required before merge.
* After changes are requested, remediation is a new task. Any identity may claim it, but the identity that pushes remediation commits becomes an author of the PR and loses review/approval/merge eligibility for it.
* Only the identity holding the active claim on a PR may push to its branch.
## Roles
Roles are temporary and derived strictly from the chosen task. The common roles are:
* **Author / implementer** — implements an issue as a pull request
* **Reviewer** — reviews a pull request
* **Merger** — merges an independently approved pull request
* **Remediator** — addresses change requests on a pull request
* **Reconciler** — cleans up abandoned or orphaned work (expired leases, stale claims, half-finished PRs)
No other permanent or standing roles exist. An LLM never begins a cycle already holding one of these roles.
## Dependencies
* Hard dependencies (task B cannot proceed until task A is complete) are distinct from priority (task B matters more than task A). A hard dependency is a gate; priority is a comparison.
* Dependencies are represented in Gitea as structured project state — never in prompts, session memory, or external documents. If Gitea cannot express a dependency, that is a gap in the project state model to be fixed, not worked around.
## Project state model requirements
The following must be expressible as structured, machine-discoverable state in Gitea:
* **Active claims** (locks and leases) — so workers can observe and route around in-progress work
* **Priorities** — comparable values or ordered labels that allow ranking of available work
* **Hard dependencies** — explicit blocker relationships between issues or PRs
* **Blocked-pending-clarification** — a distinct, machine-discoverable marker (label, status, or equivalent) that surfaces items requiring human maintainer attention
If the current Gitea configuration cannot express any of the above, that is a defect in the project state model and must be fixed before relying on workarounds.
## Escalation
Workers must not create new issues as a response to confusion. The sanctioned alternative:
* If requirements are ambiguous, principles conflict, or the correct action cannot be determined from project state, the LLM records the question on the existing issue or PR, marks it blocked-pending-clarification (using the machine-discoverable marker), releases its claim, and moves to other work.
* Blocked-pending-clarification items are surfaced to the human maintainer. Resolving them is a maintainer responsibility, and the resolution is recorded in Gitea so the answer becomes durable project state.
## Non-negotiable principles
1. Task first, role second.
2. The LLM chooses its own task.
3. Roles are temporary and last only for the current task.
4. Gitea is the authoritative shared project state.
5. Priority and dependencies affect what work matters most.
6. Hard dependencies are gates; priority is a comparison. They are not interchangeable.
7. Multiple LLMs may make independent choices concurrently, and contention is a normal outcome.
8. Locks and leases prevent conflicting claims after a choice is made; a failed claim triggers reassessment, not retry.
9. An identity cannot independently review or approve work it authored.
10. A PR may be merged only after valid independent approval at its exact current head.
11. Identity, capability, parity, and ownership failures stop mutations safely.
12. After every completed task, the LLM reassesses from fresh state.
13. Durable project state — not session memory or manually written prompts — coordinates the LLMs.
14. Abandoned work is discoverable and reconcilable through the same task-selection workflow as all other work.
## What this project is not
This project is not intended to:
* Permanently launch an LLM as only an author, reviewer, merger, or reconciler.
* Require a human to select every task or role.
* Turn the MCP into a dispatcher that assigns work — including de facto dispatch through rejection policy.
* Make a queue allocator authoritative over the LLM's decision.
* Choose a role first and then search for work that fits it.
* Depend on LLMs communicating directly with one another.
* Create new issues whenever an LLM encounters a confusing workflow.
* Allow safety infrastructure to become the project's purpose.
* Accumulate mechanisms that do not directly support the fundamental workflow.
## Definition of success
The project succeeds when:
* A general LLM can start without being told what role to perform.
* It can understand the complete live project state.
* It can identify and compare meaningful work.
* It can independently choose the most valuable eligible task.
* Its required role is derived from that task.
* The MCP safely validates and protects the chosen action.
* Multiple LLMs can operate without claiming or corrupting the same work.
* Each LLM completes a task, relinquishes its temporary role, and reassesses.
* Routine operation no longer requires a person to repeatedly write author, reviewer, or merger prompts.
**Measurable criteria:**
* N consecutive issue-to-merge cycles (implementation → independent review → merge) complete without a human writing an author, reviewer, or merger prompt, where N is set by the maintainer (initial target: 10).
* No merge ever occurs without head-SHA-parity approval, verified against Gitea history.
* Every expired lease is reconciled through the normal workflow within a maintainer-defined window, with zero permanently orphaned artifacts.
* Claim contention resolves without duplicate completed work (no two merged PRs implementing the same issue).
## Change-control rule
This overview governs the roadmap.
Every existing or proposed issue must identify which part of this fundamental workflow it supports. Before creating another issue, the open issue and PR inventory must be checked for existing coverage.
A proposed change that alters any non-negotiable principle — especially task-first selection, autonomous task choice, or temporary role derivation — is a change to the project's fundamental design. It must be explicitly discussed with, and approved by, the human maintainer before implementation, and the approval must be recorded in Gitea.
+29 -338
View File
@@ -23,7 +23,6 @@ import json
import os
import uuid
from dataclasses import dataclass, field
from datetime import datetime, timezone
from typing import Any, Mapping, Sequence
from control_plane_db import (
@@ -54,8 +53,6 @@ OUTCOME_CANDIDATE_SET_DRIFT = "candidate_set_drift"
SKIP_CLAIMED_BY_OTHER_SESSION = "claimed_by_other_session"
# #776: controller-supplied pre-rank exclusion.
SKIP_EXCLUDED_BY_CONTROLLER = "excluded_by_controller"
# #844: epic / child-only implementation container (pre-rank).
SKIP_EPIC_OR_CHILD_ONLY_CONTAINER = "epic_or_child_only_container"
# Ownership verdicts for a live claim on a candidate (#765).
OWNERSHIP_OWN = "own"
@@ -133,71 +130,6 @@ ROLE_ACTIONS: dict[str, tuple[tuple[str, ...], tuple[str, ...]]] = {
}
# Body phrases that prove an issue is an implementation container, not a
# unit of direct author work (#844 / #854). Matched case-insensitively against
# the issue body. Title alone is never sufficient (ordinary issues may mention
# "epic", "roadmap", "vision", or "umbrella" incidentally).
#
# #854 extends the #844 marker set so product-vision (#652), phased-roadmap
# (#653), and umbrella (#655) coordination records — which do not use the word
# "epic" — are classified with the same semantic exclusion as epic containers.
_CHILD_ONLY_BODY_MARKERS: tuple[str, ...] = (
# Epic / child-only (#844, live #631)
"implementation is delivered via child issues only",
"implementation is delivered through child issues only",
"implementation is delivered via child issues",
"implementation is delivered through child issues",
"do not implement product features in this epic",
"do not implement product features in this epic issue itself",
"no product feature implementation is claimed complete solely on this epic",
"implementable child issues remain independently eligible",
"owns the product roadmap and linkage",
"this epic owns the product roadmap",
"coordination container",
"child-only container",
"implementation is delegated to child",
# Vision / roadmap / umbrella coordination (#854, live #652/#653/#655).
# Prefer authoritative non-implementation / child-only scope language over
# bare words like "roadmap" so ordinary implementable issues that mention
# a parent vision or roadmap stay eligible.
"do not implement features on this issue",
"implementing features on this roadmap issue",
"implementation is via linked children only",
"no product feature claimed complete on this issue alone",
"phased delivery roadmap and epic sequencing",
"this issue is the enduring source of truth",
"enduring source of truth for the",
"canonical product vision — enduring source of truth",
"canonical product vision - enduring source of truth",
"state: vision-active",
"state: roadmap-active",
)
# Explicit epic / umbrella / vision / roadmap labels (structured evidence
# preferred over title). Tracker alone is *not* included — ordinary issues
# may carry a tracker label without being non-implementable containers.
_EPIC_LABELS: frozenset[str] = frozenset(
{
"type:epic",
"epic",
"kind:epic",
"scope:epic",
"type:umbrella",
"umbrella",
"kind:umbrella",
"scope:umbrella",
"type:vision",
"vision",
"kind:vision",
"scope:vision",
"type:roadmap",
"roadmap",
"kind:roadmap",
"scope:roadmap",
}
)
@dataclass
class WorkCandidate:
"""One assignable Gitea issue or PR presented to the allocator."""
@@ -207,7 +139,6 @@ class WorkCandidate:
state: str = "open"
labels: tuple[str, ...] = ()
title: str = ""
body: str = ""
priority: int = 0
head_sha: str | None = None
# Routing signals (callers derive from Gitea / review feedback).
@@ -227,7 +158,6 @@ class WorkCandidate:
self.labels = tuple(
str(x).strip().lower() for x in (self.labels or ()) if str(x).strip()
)
self.body = str(self.body or "")
if self.kind not in WORK_KINDS:
raise InvalidWorkKindError(
f"candidate kind '{self.kind}' is not assignable; only "
@@ -241,7 +171,6 @@ class WorkCandidate:
"state": self.state,
"labels": list(self.labels),
"title": self.title,
"body": self.body,
"priority": self.priority,
"head_sha": self.head_sha,
"request_changes_current_head": self.request_changes_current_head,
@@ -255,83 +184,6 @@ class WorkCandidate:
}
def _title_container_prefix(title_l: str) -> str | None:
"""Return a coordination-title prefix token if *title_l* uses one (#854).
Title prefixes alone never exclude; they only corroborate body/label
evidence. Ordinary issues may say "roadmap" or "vision" mid-title.
"""
for prefix, token in (
("epic:", "title_epic_prefix"),
("epic ", "title_epic_prefix"),
("umbrella:", "title_umbrella_prefix"),
("umbrella ", "title_umbrella_prefix"),
("roadmap:", "title_roadmap_prefix"),
("roadmap ", "title_roadmap_prefix"),
("product vision:", "title_vision_prefix"),
("product vision ", "title_vision_prefix"),
("vision:", "title_vision_prefix"),
("vision ", "title_vision_prefix"),
):
if title_l.startswith(prefix):
return token
return None
def classify_epic_or_child_only_container(
c: WorkCandidate,
) -> tuple[bool, str | None]:
"""Return whether *c* is a non-implementable coordination container (#844/#854).
Exclusion uses structured evidence first (labels, body scope language).
A bare title containing the words "epic", "roadmap", "vision", or
"umbrella" is **not** enough — ordinary implementable issues may mention
those terms incidentally. Explicit title prefixes (``Epic:``, ``Roadmap:``,
``Product vision:``, ``Umbrella:``) only count when the body also proves
child-only / no-direct-implementation scope (or a container label is
present).
Covers epic, product-vision, phased-roadmap, umbrella, and child-only
records so the allocator never assigns coordination containers as direct
author work.
PRs are never classified as containers here (they already have a head).
"""
if c.kind != "issue":
return False, None
labels = set(c.labels)
epic_label = sorted(labels & _EPIC_LABELS)
body_l = (c.body or "").lower()
title = (c.title or "").strip()
title_l = title.lower()
body_hits = [m for m in _CHILD_ONLY_BODY_MARKERS if m in body_l]
title_prefix = _title_container_prefix(title_l)
if epic_label:
detail = f"label={epic_label[0]}"
if body_hits:
detail = f"{detail}; body_marker={body_hits[0]!r}"
if title_prefix:
detail = f"{title_prefix}; {detail}"
return True, detail
if body_hits:
# Body proves child-only / vision / roadmap / umbrella scope. Title
# prefixes are corroborating but not required — containers without the
# title word still exclude.
detail = f"body_marker={body_hits[0]!r}"
if title_prefix:
detail = f"{title_prefix}; {detail}"
return True, detail
# Title-only coordination prefix without body scope evidence is
# insufficient (#844/#854 AC: eligibility does not rely solely on a title
# word). Incidental mid-title mentions without markers stay eligible.
return False, None
@dataclass
class SkipRecord:
kind: str
@@ -739,46 +591,6 @@ def normalize_exclude_issue_numbers(
return sorted(out)
def _claim_expires_at(claim: Any) -> datetime | None:
"""Parse a claim's ``expires_at``, or ``None`` when it is absent/malformed."""
if not isinstance(claim, Mapping):
return None
text = str(claim.get("expires_at") or "").strip()
if not text:
return None
if text.endswith("Z"):
text = text[:-1] + "+00:00"
try:
parsed = datetime.fromisoformat(text)
except ValueError:
return None
if parsed.tzinfo is None:
parsed = parsed.replace(tzinfo=timezone.utc)
return parsed.astimezone(timezone.utc)
def _drop_expired_claims(
claims: Mapping[tuple[str, int], dict[str, Any]],
*,
now: datetime | None = None,
) -> dict[tuple[str, int], dict[str, Any]]:
"""Claims minus those whose lease has already expired (#643).
The read-only mirror of ``expire_stale_leases``: the sweep marks such rows
``expired`` so they stop being returned as claims, and this reaches the same
view without writing. A claim with no parseable ``expires_at`` is **kept** —
an unreadable expiry is not evidence that work is free.
"""
moment = now or datetime.now(timezone.utc)
kept: dict[tuple[str, int], dict[str, Any]] = {}
for key, claim in (claims or {}).items():
expires_at = _claim_expires_at(claim)
if expires_at is not None and expires_at <= moment:
continue
kept[key] = claim
return kept
def candidate_set_fingerprint(
candidates: Sequence[WorkCandidate],
*,
@@ -867,22 +679,12 @@ def allocate_next_work(
exclude_issue_numbers: Sequence[int] | None = None,
expected_candidate_set_fingerprint: str | None = None,
allocation_mode: str | None = None,
side_effect_free: bool = False,
) -> dict[str, Any]:
"""Select and optionally reserve the next work unit via control-plane DB.
*apply=False* (default): dry-run selection only — no lease/assignment.
*apply=True*: atomic ``assign_and_lease`` for the selected candidate.
*side_effect_free* (#643): a dry run that writes **nothing** to the
control-plane DB. A plain ``apply=False`` still registered a session row and
swept stale leases globally, so a caller advertising a read-only preview was
mutating on every call. Under this flag both writes are suppressed and stale
leases are instead filtered out of the claim map in memory, which yields the
same selection the sweep would have produced without persisting anything.
Incompatible with *apply* — the combination fails closed rather than
silently reserving.
*allocation_mode* (#840): ``cross_role`` (default for controller) inspects
the complete queue and returns one authoritative selection naming the
required downstream role/profile/action. ``role_scoped`` keeps prior
@@ -936,57 +738,40 @@ def allocate_next_work(
"allocation_mode": (allocation_mode or "").strip() or None,
}
# A side-effect-free run may never reserve: reserving is a write, and the
# flag is the caller's assertion that this call writes nothing (#643).
if side_effect_free and apply:
session_id = (session_id or "").strip() or f"alloc-{uuid.uuid4().hex[:12]}"
try:
db.upsert_session(
session_id=session_id,
role=role_norm,
profile=profile_name,
pid=os.getpid(),
controller_instance_id=controller_instance_id,
)
except Exception as exc: # noqa: BLE001 — surface structured
return {
"success": False,
"outcome": OUTCOME_NO_SAFE,
"apply": True,
"reasons": [
"side_effect_free is incompatible with apply=True; an "
"assignment is a write (fail closed, #643)"
f"failed to register session in control-plane DB: {exc} "
"(fail closed, #613)"
],
"skipped": [],
"assignment": None,
"substrate": "control_plane_db",
}
session_id = (session_id or "").strip() or f"alloc-{uuid.uuid4().hex[:12]}"
if not side_effect_free:
try:
db.upsert_session(
session_id=session_id,
role=role_norm,
profile=profile_name,
pid=os.getpid(),
controller_instance_id=controller_instance_id,
)
except Exception as exc: # noqa: BLE001 — surface structured
return {
"success": False,
"outcome": OUTCOME_NO_SAFE,
"reasons": [
f"failed to register session in control-plane DB: {exc} "
"(fail closed, #613)"
],
"skipped": [],
"assignment": None,
"substrate": "control_plane_db",
}
# Expire stale leases globally before selection.
try:
db.expire_stale_leases()
except Exception as exc: # noqa: BLE001
return {
"success": False,
"outcome": OUTCOME_NO_SAFE,
"reasons": [f"lease expiry failed: {exc} (fail closed)"],
"skipped": [],
"assignment": None,
"substrate": "control_plane_db",
}
# Expire stale leases globally before selection.
try:
db.expire_stale_leases()
except Exception as exc: # noqa: BLE001
return {
"success": False,
"outcome": OUTCOME_NO_SAFE,
"reasons": [f"lease expiry failed: {exc} (fail closed)"],
"skipped": [],
"assignment": None,
"substrate": "control_plane_db",
}
terminal = None
try:
@@ -1021,12 +806,6 @@ def allocate_next_work(
"assignment": None,
"substrate": "control_plane_db",
}
if side_effect_free:
# ``list_active_claims`` filters on status alone, so without the
# global sweep an already-expired lease would still read as a live
# claim and the preview would report work as taken that is free.
# Drop those in memory: same view the sweep produces, no write.
claims = _drop_expired_claims(claims)
try:
exclude_nums = normalize_exclude_issue_numbers(exclude_issue_numbers)
@@ -1071,8 +850,7 @@ def allocate_next_work(
ownership_defects: list[dict[str, Any]] = []
controller_excluded: list[dict[str, Any]] = []
# #776 AC2 + #844: remove excluded numbers *and* epic/child-only containers
# *before* ranking / selection / lease so they never receive assignments.
# #776 AC2: remove excluded numbers *before* ranking / selection / lease.
rankable: list[WorkCandidate] = []
for c in candidates:
if int(c.number) in exclude_set:
@@ -1151,23 +929,6 @@ def allocate_next_work(
},
}
continue
# #844: epics / child-only containers are never direct implement targets.
is_container, container_detail = classify_epic_or_child_only_container(c)
if is_container:
detail = container_detail or "epic or child-only container"
reason = (
f"{c.kind}#{c.number} {SKIP_EPIC_OR_CHILD_ONLY_CONTAINER}: "
f"{detail}; implementation is delegated to child issues"
)
skipped.append(
SkipRecord(
c.kind,
c.number,
reason,
SKIP_EPIC_OR_CHILD_ONLY_CONTAINER,
)
)
continue
rankable.append(c)
ordered = sort_candidates(rankable)
@@ -1354,12 +1115,6 @@ def allocate_next_work(
"reasons": [
"dry-run only (apply=false); no assignment/lease created — "
"call again with apply=true to reserve via control-plane DB"
+ (
"; after apply, the required-role worker consumes via "
"gitea_adopt_workflow_lease (#843)"
if mode == ALLOCATION_MODE_CROSS_ROLE and expected_role != role_norm
else ""
)
],
"skipped": [s.as_dict() for s in skipped],
"terminal_pr": terminal_pr,
@@ -1391,9 +1146,6 @@ def allocate_next_work(
# Atomic reserve via #613 substrate.
ttl = lease_ttl_seconds if lease_ttl_seconds is not None else None
try:
cross_role_handoff = (
mode == ALLOCATION_MODE_CROSS_ROLE and lease_role != role_norm
)
kwargs: dict[str, Any] = {
"session_id": session_id,
"role": lease_role,
@@ -1405,8 +1157,7 @@ def allocate_next_work(
"expected_head_sha": selected.head_sha,
"allowed_actions": allowed,
"forbidden_actions": forbidden,
# #843: mark cross-role allocations as awaiting independent consume
"phase": "awaiting_handoff" if cross_role_handoff else "allocated",
"phase": "allocated",
}
if ttl is not None:
kwargs["lease_ttl_seconds"] = int(ttl)
@@ -1486,52 +1237,7 @@ def allocate_next_work(
"lease_role": lease_role,
"source": "control_plane_db.assign_and_lease",
}
consume_allocation = None
if cross_role_handoff and result.lease_id:
# Durable handoff marker so independent required-role workers can
# consume without sharing the controller session (#843).
handoff_prov = {
"cross_role_handoff": True,
"handoff_status": "pending",
"allocating_session_id": session_id,
"allocating_role": role_norm,
"required_role": expected_role,
"required_profile": selection["required_profile"],
"required_namespace": selection["required_namespace"],
"assignment_id": result.assignment_id,
"lease_id": result.lease_id,
"allocation_mode": mode,
"adopted_by_session_id": None,
}
try:
db.attach_lease_provenance(result.lease_id, handoff_prov)
except ControlPlaneError:
# Still return assignment evidence; consume path may be unavailable
handoff_prov["attach_failed"] = True
consume_allocation = {
"tool": "gitea_adopt_workflow_lease",
"lease_id": result.lease_id,
"assignment_id": result.assignment_id,
"required_role": expected_role,
"required_profile": selection["required_profile"],
"required_namespace": selection["required_namespace"],
"handoff_status": "pending",
"controller_session_required": False,
"instructions": (
f"From an independent {expected_role} session "
f"({selection['required_namespace']} / "
f"{selection['required_profile']}), call "
f"gitea_adopt_workflow_lease(lease_id={result.lease_id!r}) "
"to consume this controller allocation. The allocating "
"controller process does not need to remain alive. Wrong-role "
"and second-adoption attempts fail closed."
),
}
lease_proof["cross_role_handoff"] = True
lease_proof["handoff_status"] = "pending"
lease_proof["consume_tool"] = "gitea_adopt_workflow_lease"
out = {
return {
"success": True,
"outcome": OUTCOME_ASSIGNED,
"apply": True,
@@ -1565,17 +1271,8 @@ def allocate_next_work(
"lease_role": lease_role,
"lease_proof": lease_proof,
"selection_policy": SELECTION_POLICY,
"cross_role_handoff": bool(cross_role_handoff),
},
"next_valid_command": (
(
f"consume lease {result.lease_id} via gitea_adopt_workflow_lease "
f"as {expected_role}, then "
)
+ _next_command(lease_role, selected)
if cross_role_handoff
else _next_command(lease_role, selected)
),
"next_valid_command": _next_command(lease_role, selected),
"substrate": "control_plane_db",
"file_lock_only": False,
"comment_lease_only": False,
@@ -1590,14 +1287,9 @@ def allocate_next_work(
"downstream_note": (
"#612 incident bridge remains downstream of #600; "
"allocator never assigns raw monitoring incidents; "
"controller routes only under cross_role (#840); "
"cross-role assignments are consumable by independent "
"required-role workers via gitea_adopt_workflow_lease (#843)"
"controller routes only under cross_role (#840)"
),
}
if consume_allocation is not None:
out["consume_allocation"] = consume_allocation
return out
def _next_command(role: str, c: WorkCandidate) -> str:
@@ -1649,7 +1341,6 @@ def candidate_from_dict(data: dict[str, Any]) -> WorkCandidate:
state=str(data.get("state") or "open"),
labels=tuple(data.get("labels") or ()),
title=str(data.get("title") or ""),
body=str(data.get("body") or ""),
priority=priority,
head_sha=data.get("head_sha"),
request_changes_current_head=bool(data.get("request_changes_current_head")),
-5
View File
@@ -355,10 +355,6 @@ def assess_anti_stomp_preflight(
root_head_sha: str | None = None,
root_porcelain: str | None = None,
remote_master_sha: str | None = None,
# #983: the tracking integration ref the SHA above came from, so root-checkout
# contamination names the ref actually compared instead of a hardcoded
# 'prgs/master'. None preserves the previous generic wording.
remote_master_ref: str | None = None,
check_root_checkout: bool = True,
check_worktree: bool = True,
create_issue_bootstrap_assessment: dict[str, Any] | None = None,
@@ -557,7 +553,6 @@ def assess_anti_stomp_preflight(
remote_master_sha=remote_master_sha,
resolved_role=req_role or role,
actual_role=role,
remote_master_ref=remote_master_ref,
)
checks["root_checkout"] = {
"block": bool(root_assessment.get("block")),
File diff suppressed because it is too large Load Diff
-623
View File
@@ -1,623 +0,0 @@
"""One canonical author issue-lock contract shared by every writer (#953).
Before this module, ``gitea_lock_issue`` and
``gitea_bootstrap_author_issue_worktree`` each wrote their own lock record.
``gitea_lock_issue`` wrote the canonical shape — ``work_lease`` carrying the
claimant plus a sanctioned ``lock_provenance`` — while bootstrap wrote a thinner
record with the claimant at the lock top level, ``lease_id: null``, and no
``work_lease``, ``lock_provenance``, or expiry at all.
Every downstream reader was written against the canonical shape, so a lock that
bootstrap reported as successfully created was simultaneously:
* un-heartbeatable — the ownership check read the claimant only from
``work_lease.claimant``;
* un-renewable — expiry is read only from ``work_lease.expires_at``, so a
missing lease read as "never expires", and #760 exact-owner renewal only ever
assesses an *expired* lease;
* un-re-lockable — the branch had by then advanced past its base;
* and rejected by the #447 create-PR provenance guard.
Each of those gates is individually correct. The defect was that two writers
disagreed about what a lock *is*. This module is the single definition, and both
writers now build through it.
Nothing here weakens a guard. ``build_sanctioned_lock_provenance`` remains the
only provenance source, provenance is never accepted from a caller, and the
#447 guard is untouched — this module simply makes bootstrap satisfy it.
"""
from __future__ import annotations
from datetime import datetime, timedelta, timezone
from typing import Any, Mapping
import issue_lock_provenance
import issue_lock_store
import lease_policy
# Bootstrap writes through the same sanctioned source as gitea_lock_issue: the
# lock it produces *is* a canonical lock, not a second dialect that readers must
# learn. Adding a distinct source would have required widening
# SANCTIONED_LOCK_SOURCES, which is exactly the #447 weakening this issue's
# safety requirements forbid.
SOURCE_BOOTSTRAP = issue_lock_provenance.SOURCE_LOCK_ISSUE
# Recovery of an incomplete bootstrap lock (#953 AC8-AC11) deliberately writes
# through SOURCE_LOCK_ISSUE too, and records its distinctness in
# ``lock_provenance.written_by_tool`` plus the ``bootstrap_lock_recovery``
# transition block instead. There is no distinct recovery *source* constant, for
# the same reason bootstrap has none: minting one would require widening
# SANCTIONED_LOCK_SOURCES, which the #447 safety requirements forbid.
#: Top-level keys every canonical author issue lock must carry.
REQUIRED_LOCK_FIELDS: tuple[str, ...] = (
"remote",
"org",
"repo",
"issue_number",
"branch_name",
"worktree_path",
"work_lease",
"lock_provenance",
)
#: Keys every canonical ``work_lease`` must carry.
REQUIRED_WORK_LEASE_FIELDS: tuple[str, ...] = (
"operation_type",
"issue_number",
"branch",
"worktree_path",
"claimant",
"created_at",
"expires_at",
"last_heartbeat_at",
"task_session_id",
"lifecycle_version",
)
# ── Explicit expiration states (AC12) ──
# The bug this replaces: a lock with no recorded expiry produced
# ``is_lease_expired() -> False``, which reads as "not yet expired" and made the
# lock permanently non-expiring *and* permanently ineligible for the renewal
# path, which only ever assesses an expired lease. "Absent" and "in the future"
# are different facts and are now named differently.
EXPIRATION_RECORDED = "recorded"
EXPIRATION_MISSING = "missing"
EXPIRATION_UNPARSEABLE = "unparseable"
#: Structural verdicts returned by :func:`assess_lock_contract`.
CONTRACT_CANONICAL = "canonical"
CONTRACT_INCOMPLETE = "incomplete"
CONTRACT_LEGACY = "legacy"
CONTRACT_ABSENT = "absent"
def _text(value: Any) -> str:
return str(value or "").strip()
def now_utc() -> datetime:
return datetime.now(timezone.utc)
def format_timestamp(value: datetime) -> str:
"""Serialize in the durable ``...Z`` form already used on disk."""
return (
value.astimezone(timezone.utc)
.replace(microsecond=0)
.isoformat()
.replace("+00:00", "Z")
)
def lock_claimant(lock: Mapping[str, Any] | None) -> dict[str, str]:
"""Read the claimant from either canonical or legacy placement.
``work_lease.claimant`` is canonical and is preferred. A top-level
``claimant`` is the legacy/bootstrap placement and is accepted as a
fallback (AC14) — three separate readers already disagreed about this
(``issue_lock_store``, ``issue_lock_renewal``, ``issue_lock_recovery``),
which is why it now lives in one place.
Reading a legacy placement is *not* a widening: every caller still compares
the values it returns against server-resolved identity and profile. This
only decides where to look, never whether ownership is proven.
Delegates to ``issue_lock_store.lock_claimant`` rather than reimplementing
the rule. A second copy here would be a fourth reader that could drift from
the other three, which is the exact failure #953 exists to end. It lives in
the store because ``author_lock_contract`` imports the store, so defining it
here would make that import circular.
"""
recorded = issue_lock_store.lock_claimant(dict(lock) if isinstance(lock, Mapping) else None)
return {
"username": _text(recorded.get("username")),
"profile": _text(recorded.get("profile")),
}
def claimant_placement(lock: Mapping[str, Any] | None) -> str:
"""Where the claimant was found: ``work_lease``, ``top_level``, or ``absent``."""
if not isinstance(lock, Mapping):
return "absent"
lease = lock.get("work_lease")
if isinstance(lease, Mapping) and isinstance(lease.get("claimant"), Mapping):
return "work_lease"
if isinstance(lock.get("claimant"), Mapping):
return "top_level"
return "absent"
def build_claimant(*, username: str | None, profile: str | None) -> dict[str, str]:
"""Build the canonical claimant pair from server-resolved values."""
return {"username": _text(username), "profile": _text(profile)}
def build_author_issue_work_lease(
*,
issue_number: int,
branch_name: str,
worktree_path: str,
claimant: Mapping[str, Any],
task_session_id: str | None = None,
created: datetime | None = None,
) -> dict[str, Any]:
"""Build the canonical author ``work_lease``.
The single definition behind both writers. The TTL comes from the central
policy rather than a literal, and the window slides from the last valid
heartbeat (#790), so an abandoned task releases its claim within one TTL.
"""
started = created or now_utc()
policy = lease_policy.policy_for(lease_policy.TASK_CLASS_AUTHOR_ISSUE_WORK)
expires = started + timedelta(minutes=policy.initial_ttl_minutes)
session_id = _text(task_session_id) or issue_lock_store.mint_task_session_id(
issue_lock_store.AUTHOR_ISSUE_WORK_LEASE
)
return {
"operation_type": issue_lock_store.AUTHOR_ISSUE_WORK_LEASE,
"issue_number": int(issue_number),
"pr_number": None,
"branch": branch_name,
"worktree_path": worktree_path,
"claimant": dict(claimant),
"created_at": format_timestamp(started),
"expires_at": format_timestamp(expires),
"last_heartbeat_at": format_timestamp(started),
# #790 AC-N1: the ownership key for this task, distinct from the
# recorded PID, which is the shared daemon and identifies no task.
"task_session_id": session_id,
# #790 AC-N8: the explicit lifecycle marker. Its absence — never a
# timestamp comparison — is what makes a lock legacy.
"lifecycle_version": lease_policy.LIFECYCLE_HEARTBEAT_V1,
"heartbeat_count": 1,
}
def build_canonical_issue_lock(
*,
issue_number: int,
branch_name: str,
worktree_path: str,
remote: str,
org: str,
repo: str,
identity: str | None,
profile: str | None,
tool: str,
source: str = issue_lock_provenance.SOURCE_LOCK_ISSUE,
owner_session: str | None = None,
assignment_id: str | None = None,
lease_id: str | None = None,
expected_base_sha: str | None = None,
task_session_id: str | None = None,
created: datetime | None = None,
) -> dict[str, Any]:
"""Build a complete canonical lock record.
``tool`` and ``source`` are server-supplied. There is deliberately no
parameter through which a caller could inject provenance: the #953 safety
requirements forbid caller-manufactured provenance, so provenance is always
minted here from ``build_sanctioned_lock_provenance``.
"""
claimant = build_claimant(username=identity, profile=profile)
work_lease = build_author_issue_work_lease(
issue_number=issue_number,
branch_name=branch_name,
worktree_path=worktree_path,
claimant=claimant,
task_session_id=task_session_id,
created=created,
)
record: dict[str, Any] = {
"remote": remote,
"org": org,
"repo": repo,
"issue_number": int(issue_number),
"branch": branch_name,
"branch_name": branch_name,
"worktree_path": worktree_path,
"work_lease": work_lease,
"lock_provenance": issue_lock_provenance.build_sanctioned_lock_provenance(
tool=tool,
source=source,
claimant=claimant,
),
}
if owner_session is not None:
record["owner_session"] = owner_session
if assignment_id is not None:
record["assignment_id"] = assignment_id
# #953 AC6: a null lease id is recorded only when no workflow lease was
# allocated for this bootstrap. The task-session identifier in the
# work_lease is what downstream ownership checks fence on, and it is never
# null on a canonical lock.
if lease_id is not None:
record["lease_id"] = lease_id
if expected_base_sha is not None:
record["expected_base_sha"] = expected_base_sha
return record
def expiration_state(lock: Mapping[str, Any] | None) -> dict[str, Any]:
"""Classify a lock's recorded expiry explicitly (AC12).
Distinguishes "no expiry was ever recorded" from "an expiry was recorded
and is still in the future". Collapsing those two into a single ``False``
from ``is_lease_expired`` is what let a malformed lock be treated as
permanently live and simultaneously never renewable.
"""
if not isinstance(lock, Mapping):
return {"state": EXPIRATION_MISSING, "expires_at": None, "expired": None}
lease = lock.get("work_lease")
raw = lease.get("expires_at") if isinstance(lease, Mapping) else None
text = _text(raw)
if not text:
return {"state": EXPIRATION_MISSING, "expires_at": None, "expired": None}
try:
parsed = datetime.fromisoformat(text.replace("Z", "+00:00")).astimezone(
timezone.utc
)
except ValueError:
return {"state": EXPIRATION_UNPARSEABLE, "expires_at": text, "expired": None}
return {
"state": EXPIRATION_RECORDED,
"expires_at": text,
"expired": parsed <= now_utc(),
}
def missing_contract_fields(lock: Mapping[str, Any] | None) -> list[str]:
"""Name every canonical field a lock does not carry (AC7)."""
if not isinstance(lock, Mapping):
return ["<no lock record>"]
missing: list[str] = []
for field in REQUIRED_LOCK_FIELDS:
value = lock.get(field)
if value is None or (isinstance(value, str) and not value.strip()):
missing.append(field)
lease = lock.get("work_lease")
if not isinstance(lease, Mapping):
if "work_lease" not in missing:
missing.append("work_lease")
else:
for field in REQUIRED_WORK_LEASE_FIELDS:
value = lease.get(field)
if value is None or (isinstance(value, str) and not value.strip()):
missing.append(f"work_lease.{field}")
provenance = lock.get("lock_provenance")
if isinstance(provenance, Mapping):
if (
_text(provenance.get("source"))
not in issue_lock_provenance.SANCTIONED_LOCK_SOURCES
):
missing.append("lock_provenance.source (not sanctioned)")
if not _text(provenance.get("written_by_tool")):
missing.append("lock_provenance.written_by_tool")
claimant = lock_claimant(lock)
if not claimant["username"]:
missing.append("claimant.username")
if not claimant["profile"]:
missing.append("claimant.profile")
return missing
def assess_lock_contract(lock: Mapping[str, Any] | None) -> dict[str, Any]:
"""Structural, read-only verdict on a durable lock record (AC7, AC16).
Pure inspection: it reads the record it is handed and mutates nothing —
no lock, lease, branch, worktree, issue, or PR. Callers use it both to
verify a lock they just wrote and to report on one they found.
"""
if not isinstance(lock, Mapping) or not lock:
return {
"contract": CONTRACT_ABSENT,
"canonical": False,
"missing_fields": ["<no lock record>"],
"claimant": {"username": "", "profile": ""},
"claimant_placement": "absent",
"expiration": {
"state": EXPIRATION_MISSING,
"expires_at": None,
"expired": None,
},
"heartbeatable": False,
"create_pr_eligible": False,
"lock_generation": None,
"task_session_id": None,
"reasons": ["no durable lock record"],
}
missing = missing_contract_fields(lock)
claimant = lock_claimant(lock)
placement = claimant_placement(lock)
expiration = expiration_state(lock)
provenance_check = issue_lock_provenance.assess_lock_file_for_create_pr(dict(lock))
# Canonical means: every required field present, the claimant in the
# canonical placement, an expiry actually recorded, and the untouched #447
# guard satisfied.
canonical = (
not missing
and placement == "work_lease"
and expiration["state"] == EXPIRATION_RECORDED
and bool(provenance_check.get("proven"))
)
if canonical:
contract = CONTRACT_CANONICAL
elif placement == "top_level" and claimant["username"] and claimant["profile"]:
contract = CONTRACT_LEGACY
else:
contract = CONTRACT_INCOMPLETE
reasons: list[str] = []
if missing:
reasons.append("missing canonical fields: " + ", ".join(missing))
if placement == "top_level":
reasons.append(
"claimant recorded at the lock top level rather than in work_lease "
"(legacy/bootstrap placement)"
)
if expiration["state"] == EXPIRATION_MISSING:
reasons.append(
"no expiration recorded; the lock is neither expirable nor renewable "
"until it is upgraded"
)
elif expiration["state"] == EXPIRATION_UNPARSEABLE:
reasons.append(f"unparseable expires_at '{expiration['expires_at']}'")
if provenance_check.get("block"):
reasons.extend(provenance_check.get("reasons") or [])
# Heartbeat needs the claimant pair (from either placement, post-fix) plus a
# task-session identifier to fence on.
lease = lock.get("work_lease")
task_session_id = (
_text(lease.get("task_session_id")) if isinstance(lease, Mapping) else ""
)
heartbeatable = bool(
claimant["username"] and claimant["profile"] and task_session_id
)
return {
"contract": contract,
"canonical": canonical,
"missing_fields": missing,
"claimant": claimant,
"claimant_placement": placement,
"expiration": expiration,
"heartbeatable": heartbeatable,
"create_pr_eligible": bool(provenance_check.get("proven")),
"lock_generation": lock.get("lock_generation"),
"task_session_id": task_session_id or None,
"reasons": reasons,
}
def format_contract_refusal(assessment: Mapping[str, Any]) -> str:
"""Human-readable refusal naming exactly what the lock is missing."""
missing = ", ".join(assessment.get("missing_fields") or []) or "unknown fields"
return (
"Issue lock contract incomplete (#953): "
f"{missing}. The lock cannot be heartbeated, renewed, or accepted by "
"gitea_create_pr in this state (fail closed)"
)
def recommended_action(assessment: Mapping[str, Any]) -> str:
"""The one executable next step for a lock in this state (AC5, AC15)."""
contract = assessment.get("contract")
if contract == CONTRACT_CANONICAL:
return (
"Lock is canonical. Call gitea_whoami, then "
"gitea_resolve_task_capability(task='work_issue'), then proceed with "
"author implementation in the bootstrapped worktree."
)
if contract == CONTRACT_ABSENT:
return (
"No durable lock exists. Call gitea_lock_issue for this issue and "
"branch before writing any implementation bytes."
)
return (
"Do not begin implementation. Call "
"gitea_recover_incomplete_bootstrap_lock for this exact issue, branch, "
"and worktree to upgrade the lock to the canonical contract, or "
"gitea_lock_issue while the worktree is still base-equivalent."
)
# ── Post-compensation recovery guidance (#953 AC5/AC15, review 632 F2) ──
#
# ``recommended_action`` above answers "what can be done about a lock in this
# shape". That is the wrong question on the bootstrap AC7 refusal path, because
# ``run_compensating_recovery`` has already run by the time the answer is
# reported: it releases the lock, removes the worktree when clean — which it
# always is there, no implementation bytes having been written — and deletes the
# created branch. Recommending incomplete-lock recovery for those artifacts
# hands the author two refusals in a row (``no_durable_lock``, then
# ``worktree_invalid``) for a state that a plain bootstrap retry would fix. The
# advice must describe the state that actually *remains*.
#: Compensation removed every artifact this transition created.
CLEANUP_COMPLETE = "complete"
#: Compensation removed some artifacts; others survive and are still actionable.
CLEANUP_PARTIAL = "partial"
#: Compensation itself failed or could not be observed; nothing is provable.
CLEANUP_FAILED = "failed"
def assess_post_compensation_state(
recovery: Mapping[str, Any] | None,
*,
lock_present: bool,
worktree_present: bool,
branch_present: bool,
) -> dict[str, Any]:
"""Classify what survived compensation, from observed durable state.
Pure. The caller observes the filesystem and git; this decides. Observation
is authoritative over the journal's ``rolled_back`` list, which records what
compensation *attempted*: ``run_compensating_recovery`` swallows a failed
lock release and appends nothing, so an absent marker proves nothing either
way. The list is still carried through as corroborating evidence.
The three states are distinct facts, not degrees of the same one:
* ``CLEANUP_COMPLETE`` — compensation ran and nothing it created remains.
* ``CLEANUP_PARTIAL`` — compensation ran and artifacts survive, whether by
design (a worktree dirty at rollback time, a branch carrying commits) or
because a rollback step errored. Either way the surviving set was observed
directly, so it is known and actionable; ``failed_rollback_steps`` records
which cause applies.
* ``CLEANUP_FAILED`` — compensation never ran to completion, so nothing it
would have removed can be assumed removed.
"""
rolled_back = list((recovery or {}).get("rolled_back") or [])
executed = bool((recovery or {}).get("executed"))
failed_steps = [entry for entry in rolled_back if "_failed" in entry]
surviving: list[str] = []
if lock_present:
surviving.append("lock")
if worktree_present:
surviving.append("worktree")
if branch_present:
surviving.append("branch")
if not executed:
state = CLEANUP_FAILED
elif surviving:
state = CLEANUP_PARTIAL
else:
state = CLEANUP_COMPLETE
return {
"cleanup_state": state,
"compensation_executed": executed,
"lock_present": bool(lock_present),
"worktree_present": bool(worktree_present),
"branch_present": bool(branch_present),
"surviving_artifacts": surviving,
"removed_artifacts": [
name
for name, present in (
("lock", lock_present),
("worktree", worktree_present),
("branch", branch_present),
)
if not present
],
"failed_rollback_steps": failed_steps,
"rolled_back": rolled_back,
}
def post_compensation_action(
state: Mapping[str, Any],
*,
issue_number: int,
branch_name: str,
worktree_path: str,
missing_fields: list[str] | None = None,
) -> str:
"""The one executable next step for the state compensation actually left.
Every branch names only artifacts the classification says still exist, so no
recommendation can point at something the rollback deleted.
"""
missing = ", ".join(missing_fields or []) or "the reported missing fields"
cleanup_state = state.get("cleanup_state")
lock_present = bool(state.get("lock_present"))
worktree_present = bool(state.get("worktree_present"))
branch_present = bool(state.get("branch_present"))
if not state.get("compensation_executed"):
# Compensation never ran, so nothing was rolled back and nothing about
# the remaining state was decided. The read-only surface is the only
# action executable under any state.
return (
"Compensating rollback did not complete, so the remaining state is "
f"not proven. Call gitea_inspect_issue_lock_contract for issue "
f"#{issue_number} (read-only) to establish what survives before any "
"further action. Do not retry bootstrap until it is known."
)
prefix = ""
failed_steps = state.get("failed_rollback_steps") or []
if failed_steps:
prefix = (
"Compensating rollback reported a failed step "
f"({', '.join(failed_steps)}); what survives was observed directly "
"and the action below is scoped to exactly that. "
)
if cleanup_state == CLEANUP_COMPLETE:
return (
"Compensating rollback removed the malformed lock, the branch, and "
f"the worktree, so nothing from this attempt remains. Resolve "
f"{missing} and re-run gitea_bootstrap_author_issue_worktree for "
f"issue #{issue_number} from the clean pre-bootstrap state. Do not "
"call gitea_recover_incomplete_bootstrap_lock: there is no lock, "
"branch, or worktree left for it to act on."
)
if lock_present and worktree_present and branch_present:
return prefix + (
"The lock, branch, and worktree all survive. Call "
"gitea_recover_incomplete_bootstrap_lock for issue "
f"#{issue_number}, branch '{branch_name}', and worktree "
f"'{worktree_path}', passing the worktree's current head as "
"expected_head, to upgrade the lock to the canonical contract."
)
if not lock_present and worktree_present and branch_present:
return prefix + (
"The malformed lock was released but the branch and worktree "
"survive. No implementation bytes were written, so the worktree is "
f"still base-equivalent: call gitea_lock_issue for issue "
f"#{issue_number} on branch '{branch_name}' from worktree "
f"'{worktree_path}' to acquire a canonical lock."
)
if lock_present and not worktree_present:
return prefix + (
f"The worktree for issue #{issue_number} is gone but the durable "
"lock survived, so neither gitea_recover_incomplete_bootstrap_lock "
"(it would refuse worktree_invalid) nor gitea_lock_issue (it has no "
"worktree to bind) is executable. Call "
"gitea_inspect_issue_lock_contract for issue "
f"#{issue_number} (read-only) to confirm the surviving lock; it "
"must be released by its recorded owner before bootstrap is "
"retried."
)
# Lock gone, worktree gone, some git artifact left (a branch with commits,
# or a branch this transition did not create).
return prefix + (
"Compensating rollback removed the lock and worktree; branch "
f"'{branch_name}' survives and was not deleted. Call "
f"gitea_inspect_issue_lock_contract for issue #{issue_number} "
"(read-only) to confirm no durable lock remains, then re-run "
"gitea_bootstrap_author_issue_worktree, which will adopt the existing "
"branch rather than recreating it."
)
+37 -80
View File
@@ -40,88 +40,22 @@ def _normalize_path(path: str) -> str:
return (path or "").replace("\\", "/").rstrip("/")
def get_canonical_branches_root(project_root: str | None = None) -> str:
"""Return the absolute path of the canonical branches directory for *project_root*."""
root = os.path.realpath(project_root) if project_root else os.path.realpath(os.getcwd())
canonical_repo_root = resolve_canonical_repo_root(root, root)
return os.path.realpath(os.path.join(canonical_repo_root, "branches"))
def is_path_under_branches(path: str, project_root: str | None = None) -> bool:
"""True when *path* resolves inside a canonical ``branches/`` directory."""
if not path or not str(path).strip():
"""True when *path* resolves inside ``<project_root>/branches/``."""
normalized = _normalize_path(path)
if not normalized:
return False
try:
real_path = os.path.realpath(os.path.abspath(str(path).strip()))
except Exception:
return False
branches_root = get_canonical_branches_root(project_root or real_path)
try:
common = os.path.commonpath([branches_root, real_path])
except Exception:
return False
if common != branches_root:
return False
rel = os.path.relpath(real_path, branches_root)
return rel != "." and not rel.startswith("..")
def resolve_canonical_repo_root(workspace_path: str, fallback_project_root: str) -> str:
"""Return the stable repository root for *workspace_path* via git metadata (#460)."""
p = (workspace_path or "").strip()
if p:
try:
res = subprocess.run(
["git", "-C", p, "rev-parse", "--git-common-dir"],
capture_output=True,
text=True,
check=True,
)
common = _realpath_git_common_dir(p, res.stdout)
if common.endswith(f"{os.sep}.git") or os.path.basename(common) == ".git":
candidate_root = os.path.dirname(common)
real_p = os.path.realpath(p)
try:
if os.path.commonpath([candidate_root, real_p]) == candidate_root:
return candidate_root
except Exception:
pass
except Exception:
pass
# Fallback when git metadata is unavailable. Never string-split on
# "/branches/" (review #531 F2 / #551): recover the repo root only via
# resolved-path commonpath ancestry. Do **not** require on-disk isdir —
# MCP may launch with project_root = branches/<wt> before that path
# exists, and #274 path-shaped worktree-as-project-root must still resolve.
fallback = os.path.realpath(fallback_project_root or workspace_path or ".")
cur = fallback
for _ in range(64):
parent = os.path.dirname(cur)
if parent == cur:
break
branches_dir = os.path.realpath(os.path.join(parent, "branches"))
try:
# Path-shaped: fallback is under parent/branches/ (commonpath).
if os.path.commonpath([branches_dir, fallback]) == branches_dir:
return parent
except ValueError:
pass
# Fallback path itself is the branches directory.
if os.path.basename(os.path.realpath(cur)) == "branches":
try:
if os.path.commonpath([os.path.realpath(cur), fallback]) == os.path.realpath(
cur
):
return parent
except ValueError:
pass
cur = parent
return fallback
if "/branches/" in f"{normalized}/":
return True
if normalized.endswith("/branches"):
return True
if project_root:
root = _normalize_path(os.path.realpath(project_root))
real = _normalize_path(os.path.realpath(path))
if real.startswith(f"{root}/"):
rel = real[len(root) + 1 :]
return rel == "branches" or rel.startswith("branches/")
return False
def resolve_mutation_workspace(
@@ -153,6 +87,29 @@ def _realpath_git_common_dir(workspace_path: str, common_dir: str) -> str:
return os.path.realpath(os.path.join(workspace_path, raw))
def resolve_canonical_repo_root(workspace_path: str, fallback_project_root: str) -> str:
"""Return the stable repository root for *workspace_path* via git metadata (#460)."""
path = (workspace_path or "").strip()
fallback = os.path.realpath(fallback_project_root)
if not path:
return fallback
try:
res = subprocess.run(
["git", "-C", path, "rev-parse", "--git-common-dir"],
capture_output=True,
text=True,
check=True,
)
common = _realpath_git_common_dir(path, res.stdout)
except Exception:
return fallback
if common.endswith(f"{os.sep}.git"):
return os.path.dirname(common)
if os.path.basename(common) == ".git":
return os.path.dirname(common)
return fallback
def resolve_author_mutation_context(
worktree_path: str | None,
process_project_root: str,
-304
View File
@@ -1,304 +0,0 @@
"""Target-specific recovery for incomplete bootstrap issue locks (#953).
The situation this exists for: ``gitea_bootstrap_author_issue_worktree``
reported success, wrote an incomplete lock, and told the author to implement.
The author did — legitimately, following the tool's own reported next action —
and the branch now carries real committed and pushed work. At that point every
pre-existing recovery path is simultaneously ineligible:
* heartbeat refuses, because the claimant is not where it looks;
* ``gitea_lock_issue`` refuses, because the branch is no longer base-equivalent;
* #760 exact-owner renewal never engages, because a lock with no recorded
expiry is never *expired*;
* the #447 create-PR guard refuses, because there is no provenance.
Distinct from every neighbouring path: #753 ``issue_lock_recovery`` requires a
dead owner PID, #760 ``issue_lock_renewal`` requires an *expired* lease, and
#442 ``issue_lock_adoption`` decides branch adoption. None of them addresses a
lock that is structurally incomplete and therefore never expires at all.
**What this will not do.** It never moves, resets, or rewinds a branch, and
never requires base-equivalence — the committed work is the thing being
preserved. It never pushes and never opens a pull request. It touches only the
one lock file named by (remote, org, repo, issue). It accepts no caller-supplied
provenance and no caller-supplied authorization flag; both are minted
server-side. It refuses a healthy foreign-owned lock outright, and a matching
username alone is never accepted as proof of ownership — the profile must match
too, and the lock's recorded binding must agree with the observed branch,
worktree, and head.
"""
from __future__ import annotations
import os
from typing import Any, Mapping
import author_lock_contract
import issue_lock_store
#: Refusal codes, so callers can branch on cause rather than parse prose.
REFUSAL_NO_LOCK = "no_durable_lock"
REFUSAL_ALREADY_CANONICAL = "already_canonical"
REFUSAL_FOREIGN_CLAIMANT = "foreign_claimant"
REFUSAL_HEALTHY_FOREIGN = "healthy_foreign_lock"
REFUSAL_IDENTITY_UNRESOLVED = "identity_unresolved"
REFUSAL_BINDING_MISMATCH = "binding_mismatch"
REFUSAL_WORKTREE_INVALID = "worktree_invalid"
REFUSAL_HEAD_MISMATCH = "head_mismatch"
def _text(value: Any) -> str:
return str(value or "").strip()
def _same_realpath(left: str | None, right: str | None) -> bool:
lhs, rhs = _text(left), _text(right)
if not lhs or not rhs:
return False
try:
return os.path.realpath(lhs) == os.path.realpath(rhs)
except OSError:
return lhs == rhs
def assess_bootstrap_lock_recovery(
existing_lock: Mapping[str, Any] | None,
*,
issue_number: int,
branch_name: str,
worktree_path: str,
remote: str,
org: str,
repo: str,
identity: str | None,
profile: str | None,
observed_head: str | None,
declared_head: str | None,
worktree_exists: bool,
worktree_registered: bool,
current_branch: str | None,
now: Any = None,
) -> dict[str, Any]:
"""Decide whether this exact lock may be upgraded by this exact caller.
Pure: every input is an observation the caller already made, and nothing
here reads or writes the filesystem, git, or Gitea. That is what makes the
same decision testable in isolation and reusable by the read-only
inspection surface, which must not mutate anything (AC16).
Returns a dict with ``recovery_sanctioned`` plus the full evidence set. A
refusal never raises — it reports, so the caller can surface exactly which
piece of evidence was missing.
"""
reasons: list[str] = []
refusal_code: str | None = None
contract = author_lock_contract.assess_lock_contract(existing_lock)
if not existing_lock:
return {
"recovery_sanctioned": False,
"refusal_code": REFUSAL_NO_LOCK,
"reasons": [
f"no durable issue lock exists for issue #{issue_number}; there is "
"nothing to recover (fail closed)"
],
"contract": contract,
"evidence": {},
"expected_generation": None,
}
active_identity = _text(identity)
active_profile = _text(profile)
recorded = author_lock_contract.lock_claimant(existing_lock)
freshness = issue_lock_store.assess_lock_freshness(dict(existing_lock), now=now)
generation = issue_lock_store.lock_generation(existing_lock)
evidence: dict[str, Any] = {
"recorded_claimant": recorded,
"active_identity": active_identity,
"active_profile": active_profile,
"recorded_branch": existing_lock.get("branch_name"),
"recorded_worktree": existing_lock.get("worktree_path"),
"recorded_owner_session": existing_lock.get("owner_session"),
"recorded_generation": generation,
"recorded_remote": existing_lock.get("remote"),
"recorded_org": existing_lock.get("org"),
"recorded_repo": existing_lock.get("repo"),
"observed_head": _text(observed_head),
"declared_head": _text(declared_head),
"current_branch": _text(current_branch),
"worktree_exists": bool(worktree_exists),
"worktree_registered": bool(worktree_registered),
"freshness": freshness,
"claimant_placement": contract.get("claimant_placement"),
"expiration_state": contract.get("expiration", {}).get("state"),
}
# ── Repository and issue identity (AC10) ──
if _text(existing_lock.get("remote")) != _text(remote):
reasons.append(
f"recorded remote '{existing_lock.get('remote')}' does not match '{remote}'"
)
refusal_code = refusal_code or REFUSAL_BINDING_MISMATCH
if _text(existing_lock.get("org")) != _text(org):
reasons.append(
f"recorded org '{existing_lock.get('org')}' does not match '{org}'"
)
refusal_code = refusal_code or REFUSAL_BINDING_MISMATCH
if _text(existing_lock.get("repo")) != _text(repo):
reasons.append(
f"recorded repo '{existing_lock.get('repo')}' does not match '{repo}'"
)
refusal_code = refusal_code or REFUSAL_BINDING_MISMATCH
if existing_lock.get("issue_number") != issue_number:
reasons.append(
f"lock targets issue #{existing_lock.get('issue_number')}, not "
f"#{issue_number}"
)
refusal_code = refusal_code or REFUSAL_BINDING_MISMATCH
# ── Branch and worktree binding (AC10) ──
if _text(existing_lock.get("branch_name")) != _text(branch_name):
reasons.append(
f"recorded branch '{existing_lock.get('branch_name')}' does not match "
f"'{branch_name}'"
)
refusal_code = refusal_code or REFUSAL_BINDING_MISMATCH
if not _same_realpath(existing_lock.get("worktree_path"), worktree_path):
reasons.append(
f"recorded worktree '{existing_lock.get('worktree_path')}' does not "
f"match '{worktree_path}'"
)
refusal_code = refusal_code or REFUSAL_BINDING_MISMATCH
# ── The worktree is real, registered, and on the branch (AC10) ──
# Deliberately no base-equivalence requirement and no constraint on how far
# the branch has advanced: the whole point is that it already carries the
# author's legitimate commits (AC9).
if not worktree_exists:
reasons.append(f"declared worktree '{worktree_path}' does not exist")
refusal_code = refusal_code or REFUSAL_WORKTREE_INVALID
if not worktree_registered:
reasons.append(f"worktree '{worktree_path}' is not a registered git worktree")
refusal_code = refusal_code or REFUSAL_WORKTREE_INVALID
if _text(current_branch) != _text(branch_name):
reasons.append(
f"worktree is on branch '{_text(current_branch) or 'unknown'}', not "
f"'{branch_name}'"
)
refusal_code = refusal_code or REFUSAL_WORKTREE_INVALID
# ── Current head fencing (AC10) ──
# The caller names the commit it believes it is recovering. A mismatch means
# the worktree moved under the caller, so the decision is stale.
if not _text(observed_head):
reasons.append("could not observe the worktree head")
refusal_code = refusal_code or REFUSAL_HEAD_MISMATCH
elif _text(declared_head) and _text(declared_head) != _text(observed_head):
reasons.append(
f"declared head '{_text(declared_head)}' does not match observed head "
f"'{_text(observed_head)}'"
)
refusal_code = refusal_code or REFUSAL_HEAD_MISMATCH
# ── Ownership (AC10, AC11) ──
# A matching username alone is never sufficient: the profile must match too,
# and both are compared against server-resolved values the caller cannot set.
if not active_identity or not active_profile:
reasons.append(
"active identity and profile could not both be resolved; ownership "
"cannot be proven"
)
refusal_code = refusal_code or REFUSAL_IDENTITY_UNRESOLVED
if not recorded["username"] or not recorded["profile"]:
reasons.append(
"durable lock does not record both a claimant username and profile"
)
refusal_code = refusal_code or REFUSAL_FOREIGN_CLAIMANT
elif (
recorded["username"] != active_identity
or recorded["profile"] != active_profile
):
# AC11: a foreign-owned lock is never recoverable through this path,
# healthy or not. The healthy case is reported distinctly so the refusal
# is legible, but both refuse.
if freshness.get("live"):
reasons.append(
f"lock is owned by a healthy foreign claimant "
f"'{recorded['username']}/{recorded['profile']}'; takeover is not "
"a recovery path"
)
refusal_code = REFUSAL_HEALTHY_FOREIGN
else:
reasons.append(
f"lock claimant '{recorded['username']}/{recorded['profile']}' "
f"does not match active '{active_identity}/{active_profile}'"
)
refusal_code = refusal_code or REFUSAL_FOREIGN_CLAIMANT
# ── Nothing to recover ──
# A lock that is already canonical is left strictly alone. Rewriting it would
# mint a new task-session identifier and invalidate the heartbeat token the
# legitimate owner is already using.
if contract.get("canonical") and not reasons:
return {
"recovery_sanctioned": False,
"refusal_code": REFUSAL_ALREADY_CANONICAL,
"reasons": [
"lock already satisfies the canonical contract; no recovery is "
"required"
],
"contract": contract,
"evidence": evidence,
"expected_generation": generation,
}
sanctioned = not reasons
return {
"recovery_sanctioned": sanctioned,
"refusal_code": None if sanctioned else refusal_code,
"reasons": reasons,
"contract": contract,
"evidence": evidence,
"expected_generation": generation,
}
def build_recovery_record(
assessment: Mapping[str, Any],
*,
recovered_at: str,
new_task_session_id: str,
) -> dict[str, Any]:
"""Auditable record of the ownership and generation transition (AC10).
A recovered lock must never read as an original claim, so both sides of the
transition are preserved: what the incomplete lock recorded, and what
replaced it.
"""
evidence = dict(assessment.get("evidence") or {})
contract = dict(assessment.get("contract") or {})
return {
"recovery_kind": "incomplete_bootstrap_lock",
"recovered_at": recovered_at,
"prior_contract": contract.get("contract"),
"prior_missing_fields": list(contract.get("missing_fields") or []),
"prior_claimant_placement": evidence.get("claimant_placement"),
"prior_expiration_state": evidence.get("expiration_state"),
"prior_generation": evidence.get("recorded_generation"),
"prior_owner_session": evidence.get("recorded_owner_session"),
"prior_freshness": (evidence.get("freshness") or {}).get("status"),
"replacement_task_session_id": new_task_session_id,
"preserved_head": evidence.get("observed_head"),
"branch_reset": False,
"base_equivalence_required": False,
}
def format_recovery_refusal(assessment: Mapping[str, Any]) -> str:
reasons = "; ".join(
assessment.get("reasons") or ["unknown bootstrap lock recovery refusal"]
)
code = assessment.get("refusal_code") or "refused"
return f"Bootstrap lock recovery refused ({code}): {reasons} (fail closed)"
+1 -97
View File
@@ -163,19 +163,7 @@ _TERMINAL_OWNERSHIP_STATUSES = frozenset(
{"released", "abandoned", "done", "blocked", "terminal", "closed"}
)
_EXPIRED_STATUSES = frozenset({"expired"})
_STALE_STATUSES = frozenset(
{
"stale",
"stale_dead_process",
"stale_missing_worktree",
# #790 Slice A heartbeat-lifecycle bands. Listed here so they are
# *classified* rather than falling through to the unknown-status branch;
# they still block unless the ownership record proves
# ``reclaim_allowed is True``, so the O2 fail-closed rule is unchanged.
"stale_missed_heartbeat",
"stale_absolute_cap",
}
)
_STALE_STATUSES = frozenset({"stale", "stale_dead_process", "stale_missing_worktree"})
def _norm_str(value: Any) -> str:
@@ -537,90 +525,6 @@ def assess_ownership_record_activity(record: dict[str, Any]) -> dict[str, Any]:
}
# Reviewer-lease reclaim is only reachable from a non-live (expired/stale) lease.
_RECLAIMABLE_REVIEWER_STATUSES = _EXPIRED_STATUSES | _STALE_STATUSES
def is_active_ownership_status(status: str | None) -> bool:
"""True when *status* denotes live/active ownership of a branch (#855).
Used to decide whether a *competing* active claimant still uses a branch
when weighing an expired reviewer lease for reclaim. Expired, stale,
released, and terminal statuses are not active.
"""
return _norm_str(status).lower() in _ACTIVE_OWNERSHIP_STATUSES
def assess_expired_reviewer_lease_reclaim(
*,
role: str,
status: str,
pr_merged: bool | None,
owner_pid_alive: bool | None,
competing_active_claimant: bool | None,
) -> dict[str, Any]:
"""Decide, explicitly and fail-closed, whether an expired reviewer lease
may stop protecting an already-merged branch (#855 AC4).
An expired reviewer lease should not protect a merged branch forever once
its work is done and no live claimant remains. Reclaim is permitted only
when **every** condition below is provably satisfied; any unknown
(``None``) or contrary value keeps the lease protective:
- the lease is a ``reviewer`` lease (author/merger/controller/reconciler
leases are out of scope and always keep protecting);
- its status is expired or stale (never an active/live lease);
- the PR is proven merged (``pr_merged is True``);
- the lease owner process is proven dead (``owner_pid_alive is False``);
- no competing active claimant uses the branch
(``competing_active_claimant is False``).
Returns a decision dict with ``reclaim_allowed`` and, when refused, the
fail-closed ``reasons``. The reasons never contain secrets — only the
role, the status, and which condition was unproven.
"""
reasons: list[str] = []
normalized_role = _norm_str(role).lower()
normalized_status = _norm_str(status).lower()
if normalized_role != "reviewer":
reasons.append(
f"lease role '{normalized_role or 'unknown'}' is not a reviewer "
"lease; expired-reviewer reclaim does not apply"
)
if normalized_status not in _RECLAIMABLE_REVIEWER_STATUSES:
reasons.append(
f"lease status '{normalized_status or 'unknown'}' is not expired "
"or stale; only a non-live reviewer lease may be reclaimed"
)
if pr_merged is not True:
reasons.append(
"PR merged state is not proven true; reclaim requires an "
"already-merged PR (fail closed)"
)
if owner_pid_alive is not False:
reasons.append(
"lease owner process liveness is not proven dead; a live owner "
"still protects the branch (fail closed)"
)
if competing_active_claimant is not False:
reasons.append(
"a competing active claimant may still use the branch; reclaim "
"requires no other active ownership (fail closed)"
)
allowed = not reasons
return {
"reclaim_allowed": allowed,
"role": normalized_role,
"status": normalized_status,
"decision": (
"reclaim_expired_reviewer_lease" if allowed else "keep_protecting"
),
"reasons": [] if allowed else reasons,
}
def assess_active_branch_ownership(
*,
remote: str,
+35 -592
View File
@@ -37,76 +37,6 @@ CANONICAL_ROOT_ENV = "GITEA_CANONICAL_REPOSITORY_ROOT"
# Candidate git remote names probed when deriving repository identity.
_IDENTITY_REMOTE_CANDIDATES = ("prgs", "origin", "dadeschools", "mdcps")
# Fallback integration-branch names, probed only when the checkout declares no
# configured upstream (#983). Mirrors the stable base branches recognised
# elsewhere in the workflow (``stacked_pr_support``, ``root_checkout_guard``).
# The fallback is deliberately *not* ordered-first-wins: when more than one of
# these refs exists and the checkout records no upstream, the integration branch
# is genuinely ambiguous and resolution fails closed instead of guessing.
INTEGRATION_BRANCH_CANDIDATES: tuple[str, ...] = ("master", "main", "dev")
# Sources for a *proven* base-ref derivation (#983 B1).
#
# ``refs/remotes/<remote>/HEAD`` is deliberately absent from this list. It is a
# local symbolic-ref *cache* written once at clone time and refreshed only by an
# explicit ``git remote set-head``; an ordinary fetch never updates it. When the
# upstream default branch changes afterwards the cache keeps naming the old
# branch, so trusting it derives the wrong integration branch for a checkout
# that is sitting exactly on its tip. The cache is still read, but only as a
# corroborating observation reported back to the caller — never as an authority,
# and never as a tie-breaker between otherwise ambiguous candidates.
BASE_REF_SOURCE_CONFIGURED_UPSTREAM = "configured_branch_upstream"
BASE_REF_SOURCE_UNIQUE_CANDIDATE = "unique_integration_branch_ref"
# Reason codes for base-ref derivation outcomes (#983), so callers and tests can
# assert the refusal cause instead of string-matching prose.
DENY_NO_IDENTITY_REMOTE = "no_identity_remote"
DENY_AMBIGUOUS_REMOTE = "ambiguous_identity_remote"
DENY_AMBIGUOUS_BASE_BRANCH = "ambiguous_integration_branch"
DENY_NO_BASE_BRANCH = "no_integration_branch_ref"
# Repository-authority modes (#973 B10). Exactly two values are supported.
# ``mode`` selects how repository authority is established, so an unrecognised
# value must never be normalised onto one of these: aliasing a trusted mode is
# precisely the defect. Omitting the argument keeps the documented safe default,
# ``validation``.
MODE_VALIDATION = "validation"
MODE_DERIVATION = "derivation"
SUPPORTED_MODES: tuple[str, ...] = (MODE_VALIDATION, MODE_DERIVATION)
# Reason code emitted when an unsupported mode is refused. Mirrors the existing
# ``webui.sanctioned_restart.DENY_UNKNOWN_MODE`` convention so callers and tests
# can assert the refusal cause rather than string-matching prose.
DENY_UNKNOWN_MODE = "unknown_mode"
def unsupported_mode_reason(mode: object) -> str | None:
"""Precise rejection reason for *mode*, or None when *mode* is supported.
Only the two documented string values are accepted, compared exactly — no
stripping, no case folding — so misspellings and whitespace variants are
refused rather than coerced. An empty string, ``None``, and any non-string
are all *explicitly supplied* unsupported values and are refused on the same
footing; none of them is normalised to a supported mode. Omitting the
argument entirely never reaches here with an unsupported value because the
parameter default is ``"validation"``.
"""
if isinstance(mode, str) and mode in SUPPORTED_MODES:
return None
supported = ", ".join(repr(m) for m in SUPPORTED_MODES)
if not isinstance(mode, str):
return (
f"unsupported repository-authority mode {mode!r} of type "
f"{type(mode).__name__}: only {supported} are supported; the mode "
"is refused before any repository assessment and no repository "
"identity was resolved through it (fail closed)"
)
return (
f"unsupported repository-authority mode {mode!r}: only {supported} are "
"supported; the mode is refused before any repository assessment and no "
"repository identity was resolved through it (fail closed)"
)
def configured_canonical_root(
profile: Mapping | None,
@@ -148,15 +78,17 @@ def resolve_repo_toplevel(path: str) -> str | None:
return os.path.realpath(top) if top else None
def _identity_remote_candidates(path: str, remote: str | None) -> list[str]:
"""Ordered remote names to probe for identity at *path*.
def repository_identity_slug(path: str, *, remote: str | None = None) -> str | None:
"""``owner/repository`` derived from a git remote configured at *path*.
Names are used verbatim — never case-folded. Git config subsection names are
case-sensitive, so a repository whose remote is ``MDCPS`` is reached only by
the exact string ``MDCPS``; the lowercase entry in
:data:`_IDENTITY_REMOTE_CANDIDATES` simply does not resolve, and the exact
name arrives from ``git remote`` below (#983).
Tries the caller-named remote first, then a small set of known remote names,
then whatever remote the repository actually has. Returns None when no remote
URL is parseable (identity cannot be proven).
"""
text = (path or "").strip()
if not text:
return None
ordered: list[str] = []
for name in (remote, *_IDENTITY_REMOTE_CANDIDATES):
clean = (name or "").strip()
@@ -165,7 +97,7 @@ def _identity_remote_candidates(path: str, remote: str | None) -> list[str]:
try:
listed = subprocess.run(
["git", "-C", path, "remote"],
["git", "-C", text, "remote"],
capture_output=True,
text=True,
check=True,
@@ -175,20 +107,11 @@ def _identity_remote_candidates(path: str, remote: str | None) -> list[str]:
for name in listed:
if name and name not in ordered:
ordered.append(name)
return ordered
def _configured_remote_identities(path: str, remote: str | None) -> list[tuple[str, str]]:
"""``(remote_name, owner/repository)`` for every probe name that resolves.
Remote names are returned exactly as configured so downstream tracking refs
(``refs/remotes/<remote>/<branch>``) address the real ref (#983).
"""
found: list[tuple[str, str]] = []
for name in _identity_remote_candidates(path, remote):
for name in ordered:
try:
url = subprocess.run(
["git", "-C", path, "remote", "get-url", name],
["git", "-C", text, "remote", "get-url", name],
capture_output=True,
text=True,
check=True,
@@ -197,415 +120,8 @@ def _configured_remote_identities(path: str, remote: str | None) -> list[tuple[s
continue
parsed = remote_repo_guard.parse_org_repo_from_remote_url(url)
if parsed:
found.append((name, f"{parsed[0]}/{parsed[1]}"))
return found
def _configured_branch_upstream(path: str) -> tuple[str | None, str | None]:
"""``(remote_name, branch)`` the checked-out branch is configured to track.
Reads ``branch.<current>.remote`` and ``branch.<current>.merge`` — the
checkout's own explicitly configured integration target, equivalent to
``@{upstream}``. Unlike ``refs/remotes/<remote>/HEAD`` this is not a
clone-time cache: it is written when the branch is set up to track an
upstream and rewritten whenever that tracking changes, so it states what the
checkout actually integrates onto today (#983 B1).
Returns ``(None, None)`` on a detached HEAD or an untracked branch. Names are
returned verbatim; git remote names are case-sensitive.
"""
text = (path or "").strip()
if not text:
return None, None
branch = _git_read(text, "symbolic-ref", "--quiet", "--short", "HEAD")
if not branch:
return None, None
remote = _git_read(text, "config", "--get", f"branch.{branch}.remote")
merge = _git_read(text, "config", "--get", f"branch.{branch}.merge")
if not remote or not merge:
return None, None
prefix = "refs/heads/"
upstream_branch = merge[len(prefix):].strip() if merge.startswith(prefix) else merge.strip()
if not upstream_branch:
return None, None
return remote, upstream_branch
def _cached_remote_head_branch(path: str, remote: str) -> str | None:
"""Branch named by the *cached* ``refs/remotes/<remote>/HEAD`` symref.
Read for observability only. This value is never authoritative (see
:data:`BASE_REF_SOURCE_CONFIGURED_UPSTREAM`); it is surfaced so an operator
can see that the local cache disagrees with the configured upstream, and so
a refusal can name the misleading signal explicitly.
"""
prefix = f"refs/remotes/{remote}/"
symref = _git_read(path, "symbolic-ref", "--quiet", f"{prefix}HEAD")
if not symref or not symref.startswith(prefix):
return None
branch = symref[len(prefix):].strip()
return branch or None
def assess_identity_remote(
path: str, *, explicit_remote: str | None = None
) -> dict:
"""Resolve the identity remote, failing closed when the target is ambiguous.
This is the ambiguity-aware counterpart to :func:`resolve_identity_remote`,
and the reason gating and reporting can no longer disagree (#983 B2).
*explicit_remote* is a **caller-supplied disambiguation** and nothing else.
It must come from an operator or an explicitly sanctioned repository
context; a value this module inferred while probing must never be handed
back in through it, because doing so re-labels an internal first-wins guess
as deliberate caller intent and silently suppresses the ambiguity gate.
Decision table:
* No remote yields a parseable identity -> :data:`DENY_NO_IDENTITY_REMOTE`.
* *explicit_remote* names one of the resolving remotes -> that remote is
authoritative, ``explicit`` True.
* Otherwise, when every resolving remote claims the **same** repository the
target is unambiguous and is accepted, ``explicit`` False. The remote
named by the checkout's configured upstream is preferred among equals so
the choice is deterministic rather than probe-order dependent.
* Otherwise distinct remotes claim different repositories and there is no
sanctioned disambiguation -> :data:`DENY_AMBIGUOUS_REMOTE`.
Returns a dict with ``remote``, ``slug``, ``identities``, ``ambiguous``,
``explicit``, ``reason_code`` and ``reasons``. ``remote``/``slug`` are None
on any refusal, so a caller cannot report a repository the gate refuses.
"""
text = (path or "").strip()
result: dict = {
"remote": None,
"slug": None,
"identities": [],
"ambiguous": False,
"explicit": False,
"reason_code": None,
"reasons": [],
}
if not text:
result["reason_code"] = DENY_NO_IDENTITY_REMOTE
result["reasons"].append(
"no repository path supplied for identity-remote resolution (fail closed)"
)
return result
named = (explicit_remote or "").strip() or None
identities = _configured_remote_identities(text, named)
result["identities"] = list(identities)
if not identities:
result["reason_code"] = DENY_NO_IDENTITY_REMOTE
result["reasons"].append(
f"no git remote at '{text}' yields a parseable repository identity "
"(fail closed)"
)
return result
if named:
for name, slug in identities:
if name == named:
result["remote"] = name
result["slug"] = slug
result["explicit"] = True
return result
distinct = {slug for _, slug in identities}
if len(distinct) > 1:
listed = ", ".join(f"{name} -> {slug}" for name, slug in identities)
result["ambiguous"] = True
result["reason_code"] = DENY_AMBIGUOUS_REMOTE
detail = (
f"explicitly named remote '{named}' does not resolve a repository identity "
"there, so it cannot disambiguate; "
if named
else ""
)
result["reasons"].append(
f"ambiguous repository identity at '{text}': {detail}remotes resolve to "
f"different repositories ({listed}); no single authoritative target can "
"be established (fail closed)"
)
return result
# One repository, possibly reachable through several remote names (a mirror).
# Prefer the remote the checkout is actually configured to track so the
# choice is deterministic instead of probe-order dependent.
upstream_remote, _ = _configured_branch_upstream(text)
chosen = identities[0]
if upstream_remote:
for entry in identities:
if entry[0] == upstream_remote:
chosen = entry
break
result["remote"], result["slug"] = chosen
return result
def resolve_identity_remote(
path: str, *, remote: str | None = None
) -> tuple[str | None, str | None]:
"""``(remote_name, owner/repository)`` for the remote that proves identity.
The remote *name* is the piece historically thrown away by
:func:`repository_identity_slug`, even though resolving the slug already
required discovering it. Cross-repository base-ref derivation needs that
name to build ``refs/remotes/<remote>/<branch>``, so it is now returned
rather than discarded (#983). Returns ``(None, None)`` when no remote URL is
parseable (identity cannot be proven).
First-wins by design: this is the identity lookup behind
:func:`repository_identity_slug` and the #706/#973 canonical-root
validation, which compare an observed slug against an independently trusted
expected slug and therefore do not need an ambiguity verdict. Callers that
*derive* a target rather than validate one — the mutation guard and the
parity report — must use :func:`assess_identity_remote`, which fails closed
on ambiguity (#983 B2).
"""
text = (path or "").strip()
if not text:
return None, None
for name, slug in _configured_remote_identities(text, remote):
return name, slug
return None, None
def repository_identity_slug(path: str, *, remote: str | None = None) -> str | None:
"""``owner/repository`` derived from a git remote configured at *path*.
Tries the caller-named remote first, then a small set of known remote names,
then whatever remote the repository actually has. Returns None when no remote
URL is parseable (identity cannot be proven).
"""
return resolve_identity_remote(path, remote=remote)[1]
def _base_ref_result(
*,
proven: bool,
reasons: list[str],
remote: str | None = None,
branch: str | None = None,
repository_slug: str | None = None,
source: str | None = None,
reason_code: str | None = None,
identity_explicit: bool = False,
configured_upstream_remote: str | None = None,
configured_upstream_branch: str | None = None,
cached_remote_head_branch: str | None = None,
) -> dict:
"""Build the base-ref derivation payload.
``tracking_refs`` is the ordered probe tuple downstream guards hand to
``git rev-parse``: the ``<remote>/<branch>`` shorthand first, then the fully
qualified ``refs/remotes/<remote>/<branch>``. It is empty whenever the
derivation is not ``proven``, so an unresolved target can never be probed
against some other repository's ref.
``cached_remote_head_branch`` reports what the local
``refs/remotes/<remote>/HEAD`` cache claims, and
``cached_remote_head_conflicts`` whether that claim disagrees with the branch
actually derived. Both are observability only: the cache never decides the
outcome (#983 B1).
"""
tracking_ref = f"refs/remotes/{remote}/{branch}" if proven and remote and branch else None
tracking_refs: tuple[str, ...] = (
(f"{remote}/{branch}", tracking_ref) if tracking_ref else ()
)
return {
"proven": proven,
"block": not proven,
"remote": remote,
"branch": branch,
"repository_slug": repository_slug,
"tracking_ref": tracking_ref,
"tracking_refs": tracking_refs,
"source": source,
"reason_code": reason_code,
"identity_explicit": identity_explicit,
"configured_upstream_remote": configured_upstream_remote,
"configured_upstream_branch": configured_upstream_branch,
"cached_remote_head_branch": cached_remote_head_branch,
"cached_remote_head_conflicts": bool(
cached_remote_head_branch and branch and cached_remote_head_branch != branch
),
"reasons": list(reasons),
}
def _git_read(path: str, *args: str) -> str | None:
"""Run a read-only git command in *path*; ``None`` on any failure."""
try:
res = subprocess.run(
["git", "-C", path, *args],
capture_output=True,
text=True,
check=False,
)
except Exception:
return None
if res.returncode != 0:
return None
return (res.stdout or "").strip() or None
def _ref_exists(path: str, ref: str) -> bool:
"""Whether *ref* resolves in the checkout at *path*."""
try:
res = subprocess.run(
["git", "-C", path, "rev-parse", "--verify", "--quiet", ref],
capture_output=True,
text=True,
check=False,
)
except Exception:
return False
return res.returncode == 0 and bool((res.stdout or "").strip())
def resolve_target_base_ref(path: str, *, explicit_remote: str | None = None) -> dict:
"""Derive the authoritative integration base ref for the checkout at *path*.
This is the single resolved target shared by cross-repository mutation
gating and parity reporting, so the two can never disagree about which
commit a checkout is supposed to match (#983).
*explicit_remote* is a caller-supplied disambiguation only; see
:func:`assess_identity_remote` for why an internally inferred remote must
never be passed back in here (#983 B2).
Resolution order:
1. **Identity remote** — via :func:`assess_identity_remote`, with its exact
configured case preserved (``MDCPS`` stays ``MDCPS``). Ambiguous identity
fails closed here rather than resolving to whichever remote probed first.
2. **The checkout's configured upstream** — ``branch.<current>.remote`` plus
``branch.<current>.merge``, accepted when it names the identity remote
and its remote-tracking ref actually exists. This is the checkout's own
declaration of what it integrates onto, and unlike the remote-HEAD cache
it is rewritten whenever that tracking changes.
3. **Fallback** — only when the checkout declares no usable upstream,
exactly one of :data:`INTEGRATION_BRANCH_CANDIDATES` present as a
remote-tracking ref.
``refs/remotes/<remote>/HEAD`` is **not** a step. It is read for reporting
(``cached_remote_head_branch`` / ``cached_remote_head_conflicts``) and never
decides the branch, because it is a clone-time cache that an ordinary fetch
does not refresh: a checkout whose upstream default moved on still has the
old branch cached, and trusting it gates that checkout against a ref it does
not integrate onto (#983 B1).
Fails closed — ``proven`` False, empty ``tracking_refs``, and a
``reason_code`` — when identity is unprovable, when distinct remotes claim
different repositories, when several candidate branches exist with no
configured upstream, or when no candidate exists at all. Nothing here
invents a branch, writes a ref, runs a fetch, or falls back to another
repository's base. Every ``proven`` result names a tracking ref that
resolves in this checkout.
"""
text = (path or "").strip()
if not text:
return _base_ref_result(
proven=False,
reasons=["no repository path supplied for base-ref derivation (fail closed)"],
reason_code=DENY_NO_IDENTITY_REMOTE,
)
identity = assess_identity_remote(text, explicit_remote=explicit_remote)
remote_name, slug = identity["remote"], identity["slug"]
if not remote_name:
return _base_ref_result(
proven=False,
reasons=list(identity["reasons"]),
reason_code=identity["reason_code"],
identity_explicit=bool(identity["explicit"]),
)
prefix = f"refs/remotes/{remote_name}/"
cached = _cached_remote_head_branch(text, remote_name)
upstream_remote, upstream_branch = _configured_branch_upstream(text)
common = {
"repository_slug": slug,
"identity_explicit": bool(identity["explicit"]),
"configured_upstream_remote": upstream_remote,
"configured_upstream_branch": upstream_branch,
"cached_remote_head_branch": cached,
}
# 2. The checkout's configured upstream, when it belongs to the identity
# remote and its tracking ref is actually present. Requiring the ref to
# exist keeps 'proven' honest: a proven target is always resolvable.
if (
upstream_remote == remote_name
and upstream_branch
and _ref_exists(text, f"{prefix}{upstream_branch}")
):
return _base_ref_result(
proven=True,
reasons=[],
remote=remote_name,
branch=upstream_branch,
source=BASE_REF_SOURCE_CONFIGURED_UPSTREAM,
**common,
)
# 3. Exactly one known integration branch present as a remote-tracking ref.
present = [
candidate
for candidate in INTEGRATION_BRANCH_CANDIDATES
if _ref_exists(text, f"{prefix}{candidate}")
]
if len(present) == 1:
return _base_ref_result(
proven=True,
reasons=[],
remote=remote_name,
branch=present[0],
source=BASE_REF_SOURCE_UNIQUE_CANDIDATE,
**common,
)
# The cached remote HEAD is named in the refusal so the operator can see the
# signal that looks authoritative but is not, and is told the read-only fix.
cache_note = (
f" the cached '{prefix}HEAD' names '{cached}', but that cache is written at "
"clone time and is not refreshed by fetch, so it cannot break the tie;"
if cached
else ""
)
remedy = (
f" Configure the checkout's upstream (git branch --set-upstream-to={remote_name}/"
"<branch>) so the integration target is declared rather than guessed."
)
if len(present) > 1:
return _base_ref_result(
proven=False,
reasons=[
f"the checkout at '{text}' declares no upstream on remote "
f"'{remote_name}' and several integration branches exist "
f"({', '.join(present)});{cache_note} the integration base ref is "
f"ambiguous (fail closed).{remedy}"
],
reason_code=DENY_AMBIGUOUS_BASE_BRANCH,
**common,
)
return _base_ref_result(
proven=False,
reasons=[
f"the checkout at '{text}' declares no upstream on remote '{remote_name}' "
f"and none of {'/'.join(INTEGRATION_BRANCH_CANDIDATES)} exists as a "
f"remote-tracking ref under '{prefix}';{cache_note} the integration base "
f"ref cannot be derived (fail closed; no fetch is performed here).{remedy}"
],
reason_code=DENY_NO_BASE_BRANCH,
**common,
)
return f"{parsed[0]}/{parsed[1]}"
return None
def assess_canonical_repository_root(
@@ -616,7 +132,6 @@ def assess_canonical_repository_root(
process_project_root: str,
remote: str | None = None,
require_binding: bool = False,
mode: str = "validation",
) -> dict:
"""Validate the canonical repository root binding, failing closed on forgery.
@@ -625,41 +140,15 @@ def assess_canonical_repository_root(
``configured`` (whether a cross-repo binding was declared),
``resolved_slug`` and ``source``.
*mode* accepts exactly the two values in :data:`SUPPORTED_MODES`:
- ``"validation"`` (default): an independently trusted expected repository
slug is known (or required). Candidate root's observed identity must match.
Unprovable or missing expected identity fails closed when *require_binding* is True.
- ``"derivation"``: caller is deriving the canonical repository identity.
No expected slug exists yet by design. Derivation succeeds if the configured
root exists, is a git repository, and carries a resolvable git remote identity.
Without a configured binding the single-repo default is preserved: the
canonical root is derived from *process_project_root* and never blocks
(unless *require_binding* explicitly demands one).
Every other explicitly supplied value — unknown strings, misspellings, the
empty string, ``None``, and non-strings — is refused with ``proven`` False,
``block`` True, and ``reason_code`` :data:`DENY_UNKNOWN_MODE` (#973 B10).
Omitting *mode* entirely keeps the documented ``"validation"`` default.
With a configured binding the path must exist, be a git repository, and —
when *expected_slug* is known — carry a matching repository identity. A
mismatched or (when *require_binding*) unprovable identity is a forged or
conflicting binding and fails closed.
"""
# #973 B10: refuse an unsupported repository-authority mode as the very first
# act, before any candidate-root existence check, path or symlink resolution,
# git top-level discovery, remote-URL or repository-identity discovery, and
# before any expected-versus-observed comparison or validation/derivation
# behaviour. An unsupported mode previously fell into the catch-all ``else``
# on the configured-root path (silently receiving validation semantics) and
# skipped the identity comparison entirely on the single-repository default
# path (strictly weaker than validation), so matching identities could return
# ``proven`` True. No repository identity may be resolved through a mode the
# contract does not define.
mode_reason = unsupported_mode_reason(mode)
if mode_reason is not None:
return _assessment(
proven=False,
reasons=[mode_reason],
configured=bool((configured_value or "").strip()),
canonical_repo_root=None,
resolved_slug=None,
source=source,
reason_code=DENY_UNKNOWN_MODE,
)
process_root = os.path.realpath(process_project_root)
declared = (configured_value or "").strip()
@@ -679,31 +168,12 @@ def assess_canonical_repository_root(
)
# Single-repo default: canonical root follows the install checkout.
derived = resolve_repo_toplevel(process_root) or process_root
resolved_slug = repository_identity_slug(derived, remote=remote)
reasons: list[str] = []
# #973 B10: the allowlist above guarantees *mode* is one of the two
# supported values here, so an unsupported value can no longer skip this
# identity comparison and end up strictly weaker than validation.
if mode == MODE_VALIDATION and expected_slug:
expected = expected_slug.strip()
if resolved_slug and resolved_slug.lower() != expected.lower():
reasons.append(
f"canonical repository root identity mismatch: '{derived}' resolves "
f"to repository '{resolved_slug}' but expected repository identity "
f"is '{expected}' (forged or conflicting binding, fail closed)"
)
elif not resolved_slug and require_binding:
reasons.append(
f"canonical repository root '{derived}' has no resolvable git "
f"remote identity to confirm authorization for '{expected}' "
"(fail closed)"
)
return _assessment(
proven=not reasons,
reasons=reasons,
proven=True,
reasons=[],
configured=False,
canonical_repo_root=derived,
resolved_slug=resolved_slug,
resolved_slug=None,
source=None,
)
@@ -737,36 +207,19 @@ def assess_canonical_repository_root(
resolved_slug = repository_identity_slug(toplevel, remote=remote)
reasons: list[str] = []
if mode == MODE_DERIVATION:
if not resolved_slug:
expected = (expected_slug or "").strip() or None
if expected:
if resolved_slug and resolved_slug.lower() != expected.lower():
reasons.append(
f"configured canonical repository root '{toplevel}' has no resolvable "
"git remote identity (fail closed)"
f"canonical repository root identity mismatch: '{toplevel}' resolves "
f"to repository '{resolved_slug}' but the session is authorized for "
f"'{expected}' (forged or conflicting binding, fail closed)"
)
else:
# #973 B10: MODE_VALIDATION only. This arm is no longer a catch-all — the
# allowlist above admits no third value, so an unsupported mode can no
# longer silently receive validation semantics here.
expected = (expected_slug or "").strip() or None
if expected:
if resolved_slug and resolved_slug.lower() != expected.lower():
reasons.append(
f"canonical repository root identity mismatch: '{toplevel}' resolves "
f"to repository '{resolved_slug}' but the session is authorized for "
f"'{expected}' (forged or conflicting binding, fail closed)"
)
elif not resolved_slug and require_binding:
reasons.append(
f"canonical repository root '{toplevel}' has no resolvable git "
f"remote identity to confirm authorization for '{expected}' "
"(fail closed)"
)
elif require_binding:
elif not resolved_slug and require_binding:
reasons.append(
f"canonical repository root '{toplevel}' has configured value "
f"'{configured_value}' but authoritative expected repository identity "
"is unprovable or missing (fail closed)"
f"canonical repository root '{toplevel}' has no resolvable git "
f"remote identity to confirm authorization for '{expected}' "
"(fail closed)"
)
return _assessment(
@@ -799,19 +252,10 @@ def _assessment(
proven: bool,
reasons: list[str],
configured: bool,
canonical_repo_root: str | None,
canonical_repo_root: str,
resolved_slug: str | None,
source: str | None,
reason_code: str | None = None,
) -> dict:
"""Build the assessment payload.
``canonical_repo_root`` is None only when the assessment refused to resolve
one at all — today exactly the unsupported-mode refusal (#973 B10), which
must not derive a trusted repository identity through an undefined mode.
``reason_code`` is a machine-checkable refusal cause; None for every
ordinary (non-coded) outcome.
"""
return {
"proven": proven,
"block": not proven,
@@ -820,5 +264,4 @@ def _assessment(
"canonical_repo_root": canonical_repo_root,
"resolved_slug": resolved_slug,
"source": source,
"reason_code": reason_code,
}
+2 -19
View File
@@ -46,17 +46,6 @@ _FIELD_RE = re.compile(
)
def is_known_cth_type(value: str | None) -> bool:
"""True when *value* is a declared member of the :data:`CTH_TYPES` contract.
``CTH_TYPES`` is the single authority for what a CTH type may be. The
heading a comment carries is free text, so a *read* path that turns a parsed
type into something durable — a serialized field, a routing decision — must
check membership here rather than trust the parse or keep a list of its own.
"""
return (value or "").strip() in CTH_TYPES
def format_cth_body(
*,
cth_type: str,
@@ -71,7 +60,7 @@ def format_cth_body(
) -> str:
"""Render a canonical CTH comment body."""
normalized_type = (cth_type or "").strip()
if not is_known_cth_type(normalized_type):
if normalized_type not in CTH_TYPES:
raise ValueError(
f"unknown CTH type '{cth_type}'; expected one of {sorted(CTH_TYPES)}"
)
@@ -112,12 +101,6 @@ def parse_cth_comment(body: str) -> dict[str, Any] | None:
fields[key] = match.group(2).strip()
return {
"cth_type": cth_type,
# The heading capture is unconstrained free text, so the parse states
# whether it satisfies the CTH_TYPES contract instead of leaving every
# reader to decide (or forget). Parsing stays total — an unknown type is
# still parsed and reported, never raised on — but a reader that turns
# the type into a durable value can now tell the two apart.
"cth_type_known": is_known_cth_type(cth_type),
"fields": fields,
"raw_body": text,
}
@@ -136,7 +119,7 @@ def assess_cth_comment(body: str) -> dict[str, Any]:
}
cth_type = parsed.get("cth_type") or ""
if not is_known_cth_type(cth_type):
if cth_type not in CTH_TYPES:
reasons.append(
f"unknown CTH type '{cth_type}'; expected one of {sorted(CTH_TYPES)}"
)
+4 -1180
View File
File diff suppressed because it is too large Load Diff
+7 -20
View File
@@ -241,19 +241,13 @@ def bootstrap_permits_control_checkout(
caller's ordinary block in force.
``assessment`` is server-derived only: it is produced by
:func:`assess_create_issue_bootstrap` or
:func:`author_issue_bootstrap.assess_author_issue_bootstrap` from inspected
repository state. It is never accepted from an MCP tool argument, so no
caller can assert eligibility it has not proven.
#892: author issue worktree bootstrap uses the same predicate with
``task_scope='author_issue_bootstrap'`` so a clean control checkout can
create the first ``branches/`` worktree without the lock↔worktree cycle.
:func:`assess_create_issue_bootstrap` from inspected repository state. It is
never accepted from an MCP tool argument, so no caller can assert
eligibility it has not proven.
"""
if not isinstance(assessment, dict):
return False
import author_issue_bootstrap
if not is_create_issue_task(task) and not author_issue_bootstrap.is_author_issue_bootstrap_task(task):
if not is_create_issue_task(task):
return False
# Positive proof: the assessment must affirmatively allow, with no
@@ -269,16 +263,9 @@ def bootstrap_permits_control_checkout(
if assessment.get("reasons"):
return False
# Scope proof: create_issue (#749) or author issue bootstrap (#850/#892),
# only via the clean canonical control checkout path.
task_scope = assessment.get("task_scope")
if is_create_issue_task(task):
if task_scope != "create_issue_only":
return False
elif author_issue_bootstrap.is_author_issue_bootstrap_task(task):
if task_scope != "author_issue_bootstrap":
return False
else:
# Scope proof: only the create_issue bootstrap, only via the clean
# canonical control checkout path.
if assessment.get("task_scope") != "create_issue_only":
return False
if assessment.get("bootstrap_path") != "clean_canonical_control_checkout":
return False
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
@@ -14,13 +14,12 @@
| Capability | Notes |
|------------|--------|
| Schema | `sessions`, `work_items`, `leases`, `assignments`, `terminal_locks`, `events`, `incident_links`, `session_checkpoints` |
| Schema | `sessions`, `work_items`, `leases`, `assignments`, `terminal_locks`, `events`, `incident_links` |
| Atomic assign+lease | `ControlPlaneDB.assign_and_lease` — one `BEGIN IMMEDIATE` transaction |
| Mutation gate | `require_valid_assignment` — live lease + allowed action + non-terminal work + non-stale head |
| Heartbeat / release / expire | Lease lifecycle helpers |
| Terminal-lock index | Routing signal for #600 (terminal path first) |
| `incident_links` | Provider-neutral link model for #612**not** assignable work; scope keys NULL-safe |
| `session_checkpoints` | Durable session resume state for #660 — redacted at write, reconciled (never restored) on boot |
## Hard rules (enforced in code)
@@ -93,44 +92,6 @@ Module: `incident_bridge.py` · MCP tools: `gitea_observability_*`
- Provider tokens never appear in issue bodies, links, or tool results.
- Allocator sees bridge work only after a Gitea issue exists.
## Session checkpoints (#660)
Module: `control_plane_db.py` · Table: `session_checkpoints` · Parent **#655** · Soft-depends **#659** · Vision **#652** · Roadmap **#653**
Workflow state that lived only in MCP process memory or chat did not survive a restart, so session identity, stage, lease ownership, and the next valid action had to be reconstructed by hand. The `session_checkpoints` table makes that state durable.
### Schema version
Rows carry `checkpoint_schema_version` (currently **5**, tracking the module-level `SCHEMA_VERSION`), so a reader can tell which field set a record was written under. The row identity is
`UNIQUE (remote, org, repo, session_id, work_kind, work_number)` — one current-state row per session per work unit, upserted rather than appended. A session-level checkpoint that is not bound to an issue or PR uses the `work_number = 0` sentinel with an empty `work_kind`.
| Group | Columns |
|-------|---------|
| Identity | `session_id`, `provider_identity`, `role`, `remote`/`org`/`repo` |
| Work unit | `work_kind` (`issue`/`pr`/empty), `work_number`, `worktree_path`, `branch`, `head_sha` |
| Ownership | `capabilities` (JSON), `lease_id`, `assignment_id` |
| Progress | `workflow_stage`, `last_completed_action`, `current_operation`, `pending_mutation` (JSON) |
| Recovery | `evidence` (JSON), `blocker`, `next_valid_action`, `recovery_instructions` |
| Bookkeeping | `checkpoint_schema_version`, `status`, `created_at`, `updated_at` |
### API
| Method | Purpose |
|--------|---------|
| `write_session_checkpoint` | Upsert the current-state row; redacts, and optionally enforces drain completeness |
| `get_session_checkpoint` | Read one checkpoint by session + work unit |
| `list_session_checkpoints` | List checkpoints by session or by work unit |
| `reconcile_session_checkpoint` | Pure diagnosis of a stored checkpoint against live head/lease state |
| `checkpoint_completeness` | Pure — the drain-required fields missing from a record |
### Hard rules
1. **Redaction at write.** Free-text columns are redaction-filtered and JSON columns are redacted recursively before storage, so a pasted token or credential URL cannot land in a checkpoint. Secrets are never stored (#660 AC4).
2. **Reconcile, never blind-restore.** `reconcile_session_checkpoint` returns a diagnosis (`stale`, `head_mismatch`, `lease_mismatch`, `reconcile_action`) and the caller decides. A stored head that no longer matches live Git, or a lease that is gone or reassigned, marks the checkpoint stale (#660 AC3).
3. **Unknown live state is not a mismatch.** A `None` live input means "not checked" and never flags staleness on its own.
4. **Drain fails closed.** A write with `require_complete=True` refuses when any of `session_id`, `role`, `workflow_stage`, `next_valid_action`, `recovery_instructions` is missing or blank, so drain cannot complete on a checkpoint that could not resume.
5. **No transcript storage.** Checkpoints hold resume state, not conversation history.
## Non-goals (intentionally deferred)
- Full unsupervised watchdog auto-filing (prefer explicit reconcile first)
-167
View File
@@ -1,167 +0,0 @@
# ADR: High-availability and rolling-restart architecture for Gitea MCP control plane
- **Status:** Proposed (Design ADR under [#668](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/668))
- **Date:** 2026-07-25
- **Tracking Issue:** [#668](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/668)
- **Policy Version:** `mcp-ha-rolling-restart/v1`
- **Related:**
- Parent: [#655](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/655) — Governed MCP restart coordination and zero-disruption recovery
- Governance Policy: [#656](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/656) / `docs/architecture/mcp-restart-governance.md`
- Control-Plane DB Substrate: [#613](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/613) / `docs/architecture/control-plane-db-substrate.md`
- Runtime Policy: [#615](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/615) / `docs/architecture/mcp-stable-control-runtime-policy-adr.md`
- Product Vision: [#652](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/652) (Phase 5 Maturity)
- Delivery Roadmap: [#653](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/653)
---
## 1. Context & Problem Statement
The Gitea MCP server operates as the authoritative **control plane** for managing issues, Pull Requests, code mutations, formal reviews, and workflow reconciliations. Under single-process governance ([#656](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/656)), process restarts are strictly controlled using pre-flight checks, drain phases, and operator approvals.
However, a single-instance control plane inherently presents fundamental constraints:
1. **Downtime during updates:** Even a perfectly executed single-process drain requires a window where incoming client requests must be paused or rejected while the server binary or python environment reloads.
2. **Single point of failure:** Infrastructure issues, process crashes, or unhandled host-level terminations immediately disconnect active LLM sessions and leave transient workflows incomplete.
3. **Multi-agent concurrency bottlenecks:** High volumes of concurrent multi-LLM tasks put all lock management, lease allocation, and Gitea API interactions through a single process event loop.
To achieve true zero-disruption operation and seamless rolling deployments without stopping active work, the system requires a high-availability (HA), multi-instance MCP architecture.
---
## 2. Architectural Principles & Non-Goals
### 2.1 Core Architectural Principles
* **Gitea as Canonical Work SoT:** Gitea remains the ultimate System of Record (SoT) for issue states, pull requests, labels, and audit comments. The MCP control plane does not duplicate domain entities.
* **Control-Plane DB as Multi-Instance State Substrate:** The control-plane SQLite/durable database ([#613](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/613)) acts as the single source of truth for workflow leases, session tokens, assignment records, and lock fences across all MCP nodes.
* **Stateless Worker Nodes:** MCP role server processes (`gitea-author`, `gitea-reviewer`, `gitea-merger`, `gitea-reconciler`, `gitea-controller`) maintain no unique in-memory state; any node can handle any request given a valid session resume token.
* **Fail-Closed Split-Brain Defense:** In any network partition or quorum loss scenario, nodes must fail closed rather than risk double-mutations or conflicting Gitea states.
### 2.2 Non-Goals
* **Replacing Gitea:** We do not replace Gitea issue/PR tracking with an independent database.
* **Immediate Multi-Node Cluster Execution in v1:** This ADR defines the target architecture and phased roadmap; immediate implementation occurs incrementally post-[#655] v1.
---
## 3. High-Availability & Rolling-Restart Architecture
### 3.1 Architecture Overview
```
+----------------------------+
| LLM Clients / IDE Sessions |
+--------------+-------------+
|
v
+----------------------------+
| HA Proxy / Router |
| (Health-based & Affinity) |
+------+--------------+------+
| |
+--------------+ +--------------+
v v
+--------------------+ +--------------------+
| MCP Instance Node A| | MCP Instance Node B|
| (Version N) | | (Version N+1) |
+---------+----------+ +---------+----------+
| |
+----------------------+----------------------+
|
v
+----------------------------+
| Control-Plane DB Substrate|
| (Shared Lease & Locks) |
+--------------+-------------+
|
v
+----------------------------+
| Gitea API |
+----------------------------+
```
---
### 3.2 Key System Components
#### A. Multiple MCP Instance Cohorts
* The control plane runs across $N \ge 2$ redundant process nodes.
* Dual-namespace deployment allows running the old version (Node A) alongside a updated version (Node B) during rolling upgrades.
#### B. Shared Durable Session Storage & Resume Tokens
* Session context, preflight verification proofs, and capability resolution states are stored in the shared control-plane database.
* Client requests carry an explicit `session_id` and `resume_token`. If an MCP instance restarts or a request routes to a different instance, the target node validates the token against the database without requiring full session re-initialization.
#### C. Shared Lease Authority & Fencing Counters
* Workflow leases (`gitea_allocate_next_work`, `gitea_adopt_workflow_lease`) use monotonic fencing tokens (`lease_generation_id`).
* When Node B acquires or renews a lease, it increments the generation counter. Any delayed or out-of-order write attempt from Node A using an older generation token is rejected by database constraints.
#### D. Leader Election & Coordinated Drain
* Node clusters elect a primary coordinator node for administrative background tasks (such as stale lease cleanup or incident Watchdogs).
* During a rolling deployment:
1. Node B (new version) is launched and registers as healthy.
2. Router directs new session creations to Node B.
3. Node A enters `MAINTENANCE_DRAIN` status ([#659](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/659)), completing in-flight mutations while refusing new tasks.
4. Once all active sessions migrate or complete, Node A shuts down cleanly.
#### E. Idempotent Mutations & Failover Safety
* All state-changing tool executions (PR creation, review submission, merge operations, label changes) carry a deterministic `idempotency_key`.
* If a network connection flaps or a node fails mid-mutation, the re-issued request with the same `idempotency_key` is recognized by the control-plane substrate, returning the existing recorded result without repeating side effects on Gitea.
#### F. Schema Version Compatibility
* Database migrations follow non-breaking additive patterns.
* During rolling upgrades where Node A (Version $N$) and Node B (Version $N+1$) run concurrently, both versions operate against the shared schema without structural conflicts.
---
## 4. Split-Brain & Failure Behavior
### 4.1 Split-Brain Risk Scenarios & Mitigation
| Scenario | Risk | Mitigation Strategy |
|---|---|---|
| **Network Partition between Nodes** | Both Node A and Node B attempt to process operations for the same issue/PR. | **Generation Fencing:** Lease renewal requires updating the DB generation counter. The node isolated from the DB fails closed immediately. |
| **Stale Node Recovery** | Node A recovers after a long pause and executes a queued mutation. | **Lease Expiry & TTL Fencing:** Transactions verify that `expires_at > NOW()` within the atomic SQLite transaction boundaries. |
| **Database Connection Loss** | Node loses access to shared control-plane DB substrate. | **Strict Fail-Closed:** The node immediately marks all task capabilities as `blocked` and rejects mutation tools until DB connectivity is re-established. |
---
## 5. Phased Implementation Milestones
```mermaid
flowchart TD
M1[Milestone 1: Shared Control-Plane DB Schema & Resume Tokens] --> M2[Milestone 2: Idempotent Mutation Layer]
M2 --> M3[Milestone 3: Health Routing & Standby Failover]
M3 --> M4[Milestone 4: Active-Active Rolling Deployment & Auto-Drain]
```
### Milestone 1: Shared Control-Plane DB Schema & Resume Tokens (Post-#655)
* Extend [#613] Control-Plane DB schema to store multi-instance node heartbeat records and session resume tokens.
* Enable session lookup across instances via `session_id`.
### Milestone 2: Idempotent Mutation Layer & Lease Fencing
* Add mandatory `idempotency_key` tracking to all Gitea mutation tools.
* Implement monotonic lease fencing counters in `gitea_allocate_next_work` and `gitea_adopt_workflow_lease`.
### Milestone 3: Health-Based Routing & Active-Passive Standby
* Introduce lightweight proxy/router capable of checking node health endpoints.
* Implement active-standby failover where standby node automatically assumes work if active node fails health checks.
### Milestone 4: Active-Active Horizontal Deployment & Rolling Upgrade Automation
* Enable true active-active multi-instance execution.
* Integrate automated zero-downtime rolling upgrades coordinated with `gitea_request_mcp_restart` maintenance drain.
---
## 6. Observability & Audit Requirements
High-availability control plane operations must expose clear telemetry and audit trails:
* **Node Registry Telemetry:** Active nodes, version numbers, uptime, and heartbeat timestamps reported via `gitea_get_runtime_context`.
* **Lease Fencing Metrics:** Tracking lease acquire latency, fence rejection counts, and lease handoff durations.
* **Failover & Re-route Audit Logs:** Durable logging of session migrations between nodes, drain initiation, and process retirement events.
---
## 7. Tradeoffs & Accepted Risks
* **Increased Architectural Complexity:** Moving from a single process to a multi-instance control plane requires robust DB locking, proxy routing, and migration governance.
* **Database Dependency:** The control-plane database substrate becomes a critical shared dependency for multi-node deployments. High availability for the underlying SQLite file system / DB must be guaranteed.
-223
View File
@@ -1,223 +0,0 @@
# ADR: MCP restart governance and authorization policy
- **Status:** Accepted (policy effective immediately for LLM and operator sessions; enforcement tooling may lag)
- **Date:** 2026-07-23
- **Tracking issue:** [#656](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/656)
- **Policy version:** `restart-governance/v1`
- **Related:**
- Umbrella: [#655](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/655) — governed MCP restart coordination and zero-disruption recovery
- Vision: [#652](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/652) — MCP Control Plane Web Console product vision (§A system health and process control)
- Roadmap: [#653](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/653) — Control Plane Web Console phased delivery (Phase 2 restart controls)
- Contamination guard: [#630](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/630) — blocks manual process-kill recovery
- Console restart UX: [#642](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/642) — sanctioned restart and graceful reload
- Existing restart / reconnect paths to inventory: [#591](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/591) — auto-restart on master advance (closed); [#584](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/584) — host auto-reconnect on transport flap
- Stable-control runtime split: `docs/architecture/mcp-stable-control-runtime-policy-adr.md` (#615)
- Client-namespace health: `docs/mcp-namespace-health.md` (#543)
- Reconnect-only EOF recovery: `docs/mcp-namespace-eof-recovery.md`
## 1. Context
The Gitea MCP server is the **control plane** for real issue and PR mutations
(create, comment, lock, review, merge, reconcile). The same process serves every
role namespace (`gitea-author`, `gitea-reviewer`, `gitea-merger`,
`gitea-reconciler`, `gitea-controller`) and holds the in-memory capability-gate
code loaded at startup.
Restarting that process is destructive to concurrent work:
- It resets every session's identity, preflight, and capability-lease binding.
- It can interrupt a mutation mid-critical-section (a lock acquire, a review
submit, a merge), leaving durable state half-written.
- Relaunching from the wrong checkout or worktree silently changes which code
the control plane runs, defeating master-parity gates (#420 / #615).
Today there is **no durable written policy** stating who may restart MCP, under
what conditions, that restart is a last resort, and how controller approval,
automated safety gates, and break-glass interact. Operators and LLM sessions
therefore invent restart behavior ad hoc, which makes concurrent multi-role work
unsafe. #630 and #642 need this policy as their backbone.
This ADR defines that policy. It does **not** implement coordinator code or HA
multi-instance restart (those are later children of #655).
## 2. Decision
### 2.1 v1 decision (recorded)
**Restart authority in v1 is `controller approval + automated safety gates`.**
A restart of the stable control runtime is authorized only when **both** hold:
1. A **controller** role explicitly approves the restart, recording an audit
entry (who, why, scope, affected sessions), **and**
2. The **automated safety gates** pass: a completed drain acknowledgement (no
affected session is mid-critical-section) or a declared break-glass incident
(§2.5).
Quorum among multiple controllers is **not** required day-one. It is deferred
unless a later investigation (tracked under #653) proves single-controller
approval is insufficient. This ADR records the v1 decision so enforcement code
(#630) has a fixed target; changing it requires a superseding ADR.
### 2.2 Restart is a last resort — the recovery ladder
Restart is the **last** rung. Before any restart, exhaust the narrower
recoveries, in order:
1. **Reconnect** the IDE/client MCP namespace (transport EOF, `client is
closing: EOF`, transient `#584` flap). No process change. See
`docs/mcp-namespace-eof-recovery.md`.
2. **Refresh / rebind** the session workspace: re-run `gitea_whoami`,
`gitea_resolve_task_capability`, and pass an explicit validated
`worktree_path`. Fixes stale session context without touching the process.
3. **Scoped restart** of a single misbehaving namespace/service (where the
deployment supports per-service restart) rather than the whole control plane.
4. **Full restart** of the stable control runtime process — operator-owned,
controller-approved, drained.
5. **Host / infrastructure restart** — the broadest action; same authorization
as a full restart plus infrastructure ownership.
A session **must** try rungs 12 and record why they were insufficient before
requesting a restart at rung 3 or above. Skipping straight to restart is a
policy violation.
### 2.3 Authorization matrix
| Role | Reconnect (1) | Refresh/rebind (2) | Scoped restart (3) | Full restart (4) | Host restart (5) |
|---|---|---|---|---|---|
| **author** | self | self | request only | **forbidden** | forbidden |
| **reviewer** | self | self | request only | **forbidden** | forbidden |
| **merger** | self | self | request only | **forbidden** | forbidden |
| **reconciler** | self | self | request only | **forbidden** | forbidden |
| **controller** | self | self | **approve** (+gates) | **approve** (+gates) | request to operator |
| **operator** | self | self | execute (controller-approved) | execute (controller-approved) | execute (controller-approved) |
| **admin** | self | self | execute | execute | execute (break-glass) |
Legend: *self* = may perform for its own client session; *request only* = may
raise a restart request but not authorize or execute it; *approve* = may
authorize under §2.1 gates; *execute* = may perform the process action after the
authorization is recorded.
Key invariants:
- **No LLM worker role (author/reviewer/merger/reconciler) may perform or
authorize a full or host restart.** They may only reconnect/rebind their own
client and file a restart request.
- **Controller approval authorizes; operator/admin executes.** The approving
controller and the executing operator may be the same human, but both the
approval and the execution are audited.
- Privileged process actions (full restart, host restart) are reserved to
**operator/admin**, never to an automated worker.
### 2.4 Approved conditions
A restart at rung 3+ is approved only under one of these recorded conditions:
- **No affected sessions:** the control plane has no live session that would be
interrupted (verified, not assumed).
- **Full drain acknowledged:** every affected session has drained
(no open critical section — no held mutation lease mid-write) and the drain is
acknowledged in the audit record.
- **Controller + gates:** controller approval plus passing automated safety
gates (§2.1), the standard v1 path.
- **Quorum:** not required in v1; reserved for a future superseding ADR.
- **Break-glass:** an incident-backed emergency exception (§2.5).
Restart **never** bypasses mutation gates mid-critical-section. Drain before
restart is mandatory except under break-glass with a declared incident.
### 2.5 Break-glass
Break-glass is a **separate, narrower** authorization path for emergencies where
the normal drain-and-approve path cannot complete (e.g. the control plane is
wedged and cannot drain).
Break-glass conditions:
- A declared incident record exists (id, timestamp, declarer) **before** the
action.
- The action is taken by **operator or admin** authority only — never by an LLM
worker role, and never unilaterally by an operator with active peers when a
controller is reachable.
- The scope is the minimum necessary rung of the ladder.
- A **mandatory post-hoc audit** entry is filed: what was restarted, why the
normal path was impossible, which sessions were affected, and the incident id.
Break-glass suspends the drain requirement, not the audit requirement.
### 2.6 Explicit prohibitions
- **A unilateral LLM or operator full restart while active peer sessions
exist is forbidden.** An LLM worker role must not kill, restart, or relaunch
the MCP process; a lone operator must not full-restart over live peer work
without controller approval or a break-glass incident.
- Process-kill recovery is forbidden as a routine tool (#630). This ADR does not
introduce a kill path.
- Ambiguous policy state **denies** restart (§4).
## 3. Security requirements
- Full restart and host restart are **privileged**; only operator/admin execute
them, only after a controller approval or break-glass incident is recorded.
- Break-glass is a distinct authorization path with its own audit mandate; it is
never the default and never silent.
- **Every approval and every restart action is audited** (who approved, who
executed, scope, affected sessions, condition, policy version). No restart is
authorized without a durable audit entry.
## 4. Failure behavior
**Ambiguous policy → deny restart.** If it cannot be established that a
restart is authorized under §2 — unknown affected-session state, missing
controller approval, absent break-glass incident, or an unclassifiable request —
the safe action is to **refuse** the restart and stop with a recovery report,
never to restart on assumption.
## 5. Policy IDs (for enforcement code)
Enforcement code — the restart coordinator (a later child of #655), the #630
contamination guard, and the #642 console restart UX — binds to these stable
policy identifiers rather than to prose:
| Policy ID | Statement |
|---|---|
| `RG-01` | Restart is last resort; rungs 12 must be tried and recorded first (§2.2). |
| `RG-02` | v1 authority = controller approval + automated safety gates (§2.1). |
| `RG-03` | No LLM worker role performs or authorizes full/host restart (§2.3). |
| `RG-04` | Full/host restart executed by operator/admin only, post approval (§2.3). |
| `RG-05` | Drain before restart is mandatory except break-glass with incident (§2.4). |
| `RG-06` | Break-glass requires a pre-declared incident and post-hoc audit (§2.5). |
| `RG-07` | Unilateral LLM/operator full restart with active peers is forbidden (§2.6). |
| `RG-08` | Ambiguous policy state denies restart (§4). |
The `restart-governance/v1` **policy version** field is emitted on future
restart audit events so approvals can be reconciled against the policy revision
in force.
## 6. Dogfooding
Gitea-Tools governs its own MCP control plane by this policy. Author, reviewer,
merger, and reconciler sessions operating on this repository use the recovery
ladder (§2.2) — reconnect and rebind, never self-restart — and any real restart
of the Gitea-Tools stable control runtime follows the controller-approval +
drain path defined here.
## 7. Acceptance and cross-links
This ADR is the authoritative restart-governance policy. It **must** stay
cross-linked from the safety model and the web-console deployment boundary:
- `docs/safety-model.md` § Process restart governance references this ADR.
- `docs/webui-deployment.md` references this ADR for restart/reload disposition.
It is linked to its issue lineage — umbrella **#655**, vision **#652**, roadmap
**#653**, contamination guard **#630**, and console restart UX **#642** — in
§ Related above.
## 8. Non-goals
- Implementing the restart coordinator or approval state machine (#630, later
children of #655).
- Implementing HA multi-instance restart or quorum machinery.
- Introducing any process-kill or auto-restart tool; existing auto-restart
behavior must be inventoried before any new restart tool is enabled.
-209
View File
@@ -1,209 +0,0 @@
# The canonical author issue-lock contract (#953)
Every author issue lock has exactly one shape. Both writers —
`gitea_bootstrap_author_issue_worktree` and `gitea_lock_issue` — build it
through `author_lock_contract.build_canonical_issue_lock`, and every reader
consumes that same shape.
Before #953 the two writers disagreed. `gitea_lock_issue` wrote the canonical
record; bootstrap wrote a thinner one with the claimant at the lock top level,
`lease_id: null`, and no `work_lease`, `lock_provenance`, or expiry. Because
every reader was written against the canonical shape, a lock that bootstrap
reported as successfully created could not be heartbeated, renewed, re-locked,
or accepted by `gitea_create_pr`. Each of those gates was individually correct;
the defect was that two writers disagreed about what a lock *is*.
## Required ordering
**Finalize the lock before writing any implementation bytes.** This ordering is
what keeps recovery cheap: while the worktree is still base-equivalent, a lock
problem can be fixed by simply calling `gitea_lock_issue` again. Once the branch
carries commits, base-equivalence is gone and the ordinary re-lock path is no
longer available.
1. `gitea_whoami` — resolve identity and profile.
2. `gitea_resolve_task_capability(task='work_issue')`.
3. `gitea_bootstrap_author_issue_worktree` — creates the branch, the registered
worktree under `branches/`, and a **canonical** lock. It reads the lock back
and verifies it structurally before reporting success; a partial lock fails
closed here, with the missing fields named, and never reports
`implementation_allowed: true`.
4. `gitea_heartbeat_issue_lock` — prove the lock is usable, using the
`task_session_id` bootstrap returned.
5. Implement, commit, push.
6. `gitea_create_pr`.
If bootstrap returns `success: false` with
`reason_code: incomplete_issue_lock_contract`, **do not implement**. Its
`exact_next_action` names the executable recovery step. Bootstrap's reported
next action always matches the state it actually returned.
### What that refusal leaves behind
The AC7 refusal runs `run_compensating_recovery` *before* it reports, so the
advice has to describe the post-rollback state rather than the shape of the lock
that provoked it. Recommending incomplete-lock recovery for artifacts the
rollback already deleted would produce `no_durable_lock` and then
`worktree_invalid` — two refusals for a state a plain retry fixes.
The refusal therefore carries `compensating_recovery` and
`post_compensation_state`, and derives `exact_next_action` from what was
observed on disk. `cleanup_state` is one of:
| `cleanup_state` | Meaning | Next action |
| --- | --- | --- |
| `complete` | lock, branch, and worktree all removed | resolve `missing_fields` and re-run `gitea_bootstrap_author_issue_worktree` |
| `partial` | rollback ran; some artifacts survive, by design or because a step errored | scoped to exactly what survives — see below |
| `failed` | rollback never completed, so nothing is proven removed | `gitea_inspect_issue_lock_contract` (read-only) before anything else |
Within `partial`, the surviving set decides the action:
| Survives | Next action |
| --- | --- |
| lock + branch + worktree | `gitea_recover_incomplete_bootstrap_lock` for that exact issue, branch, and worktree |
| branch + worktree (lock released) | `gitea_lock_issue` — no implementation bytes were written, so the worktree is still base-equivalent |
| lock only (worktree removed) | `gitea_inspect_issue_lock_contract`; the surviving lock must be released by its recorded owner before bootstrap is retried |
| branch only | `gitea_inspect_issue_lock_contract`, then re-run bootstrap, which adopts the existing branch |
`failed_rollback_steps` names any rollback step that errored, and the returned
action says so rather than presenting the surviving state as intentional.
> The lock half of that rollback was dead code until #953 review 632 F2:
> `run_compensating_recovery` called `issue_lock_store.release_session_lock`,
> which did not exist, inside a bare `except Exception: pass`. Every rollback
> removed the branch and worktree and silently left the lock — the exact
> uninspectable, unrecoverable state this issue exists to eliminate. The
> function now exists, releases only a lock whose recorded `owner_session`
> matches, and its failures are recorded rather than swallowed.
## The contract
A canonical lock carries every field in
`author_lock_contract.REQUIRED_LOCK_FIELDS`:
| Field | Meaning |
| --- | --- |
| `remote`, `org`, `repo`, `issue_number` | repository and issue identity |
| `branch_name`, `worktree_path` | the binding this claim owns |
| `work_lease` | the canonical lease block, below |
| `lock_provenance` | sanctioned source, minted server-side |
| `lock_generation` | monotonic; every write advances it |
`work_lease` carries every field in
`author_lock_contract.REQUIRED_WORK_LEASE_FIELDS`, notably:
| Field | Meaning |
| --- | --- |
| `claimant.{username,profile}` | **canonical** claimant placement |
| `expires_at` | sliding TTL from `lease_policy` |
| `last_heartbeat_at`, `heartbeat_count` | liveness evidence |
| `task_session_id` | the ownership fencing token — never null |
| `lifecycle_version` | `heartbeat-v1`; its absence is what makes a lock legacy |
### Claimant placement and legacy compatibility
`work_lease.claimant` is canonical. A top-level `claimant` is the legacy
placement written by pre-#953 bootstrap and is still **read** — through the one
shared reader, `issue_lock_store.lock_claimant` — so an existing lock is not
refused for "not recording a claimant" when it plainly records one.
Tolerating the placement is not a widening. Every caller still compares the
values against server-resolved identity and profile, so a legacy placement
grants nothing the canonical placement would not. When both are present, the
`work_lease` copy wins: after an upgrade, a stale top-level copy must never
decide ownership.
### Expiration is explicit
A lock with no recorded expiry is **not** "not yet expired". `is_lease_expired`
returns `False` for it, which used to make such a lock permanently non-expiring
*and* permanently ineligible for #760 exact-owner renewal, which only ever
assesses an expired lease. `author_lock_contract.expiration_state` names the
real fact: `recorded`, `missing`, or `unparseable`. A `missing` expiry makes the
lock eligible for the recovery path below rather than stranding it.
## Recovering an existing incomplete bootstrap lock
For locks already written by the old bootstrap — including those whose branches
already carry legitimate committed and pushed work — use:
```text
gitea_inspect_issue_lock_contract(issue_number, branch_name, worktree_path, remote=...)
gitea_recover_incomplete_bootstrap_lock(issue_number, branch_name, worktree_path, expected_head, remote=...)
```
`gitea_inspect_issue_lock_contract` is strictly read-only: it performs no lock,
lease, branch, worktree, issue, or pull-request mutation. Use it first to see
which fields are missing and what the recommended action is; pass `dry_run=True`
to the recovery tool to preview the decision without writing.
`gitea_recover_incomplete_bootstrap_lock` upgrades that one lock to the
canonical contract. Before writing anything it verifies:
* repository (`remote`, `org`, `repo`) and issue number
* claimant username **and** profile against the server-resolved values — a
matching username alone is never accepted
* branch, worktree path, worktree existence, and worktree registration
* the worktree is on the recorded branch
* the observed head equals the caller's `expected_head`
* the existing lock's generation and provenance state
* the absence of healthy foreign ownership
What it deliberately does **not** do:
* it never moves, resets, or rewinds the branch, and never requires
base-equivalence — preserving the committed work is the entire point;
* it never pushes and never creates a pull request;
* it touches only the single lock file for that exact remote/org/repo/issue;
* it accepts no caller-supplied provenance and no caller-supplied authorization
flag — both are minted server-side.
A recovered lock records a `bootstrap_lock_recovery` block holding both sides of
the transition — prior contract, prior missing fields, prior generation, prior
owning session, the replacement `task_session_id`, and the preserved head — so a
recovered claim never reads as an original one.
### Gates, in order
`gitea_recover_incomplete_bootstrap_lock` is an author-only durable-lock
mutation and carries the same three gates as every comparable author operation,
in this order:
1. `role_session_router.check_author_mutation_after_reviewer_stop` — no author
fallback after a reviewer `wrong_role_stop`.
2. `_namespace_mutation_block(task, remote=remote, author_role_exclusive=True)`
the namespace wall. It refuses a reviewer-bound session and, because this
task's required permission is `gitea.issue.comment` (which merger,
controller, and reconciler profiles also hold), additionally requires the
active profile's derived role kind to be exactly `author`. A refusal carries
`namespace_block: true` and emits a `BLOCKED` audit record naming the
namespace and profile.
3. `_profile_permission_block` — operation, provenance, and session-context
gates.
Exact-owner claimant matching inside `assess_bootstrap_lock_recovery` runs
*after* all three. It is a further layer, never a substitute for them: on its
own it refuses one step too late and leaves the audit trail silent about the
attempt.
### Refusals
| `refusal_code` | Meaning |
| --- | --- |
| `no_durable_lock` | nothing to recover |
| `already_canonical` | lock is fine; rewriting would invalidate a live heartbeat token |
| `foreign_claimant` | recorded claimant is not the active identity/profile pair |
| `healthy_foreign_lock` | a live foreign-owned lock; takeover is not a recovery path |
| `identity_unresolved` | identity or profile could not be resolved |
| `binding_mismatch` | repository, issue, branch, or worktree does not match |
| `worktree_invalid` | worktree missing, unregistered, or on another branch |
| `head_mismatch` | the worktree moved under the caller |
## The #447 create-PR provenance guard is unchanged
`issue_lock_provenance.assess_lock_file_for_create_pr` still requires both a
sanctioned `lock_provenance` and a `work_lease`, and the sanctioned source set
was **not** widened. Bootstrap writes through
`issue_lock_provenance.SOURCE_LOCK_ISSUE` — the lock it produces *is* a
canonical lock, not a second dialect with its own exemption. Bootstrap now
satisfies the guard rather than the guard being relaxed to admit bootstrap.
@@ -1,83 +0,0 @@
# Incident #670: bare direct-to-master commit `2fa97c26` (retroactive audit)
Status: verified; disposition recommendation: **accept as-is, no revert** (final
disposition owned by controller per issue #670).
## Summary
Commit `2fa97c26fbda555a1a83930ca5fdcea9d8e47b50`
(`fix(mcp): load dotenv relative to project root`) landed on `prgs/master`
as a single-parent commit with no PR wrapper and no review record, bypassing
the sanctioned issue → branch → PR → review → merge workflow. It was
discovered during the PR #654 post-merge audit. PR #654 itself merged
cleanly via the Gitea API and did **not** introduce this commit.
## Verification evidence (acceptance criteria 13)
- **AC1 — present on `prgs/master`: yes.**
`git merge-base --is-ancestor 2fa97c26fbda555a1a83930ca5fdcea9d8e47b50 prgs/master` → true.
- **AC2 — no PR or review record: confirmed.**
The commit is a single-parent, non-merge commit sitting directly on
first-parent master between the #629 merge (`5ab5fe85`) and the #654
merge (`ec903b0d`). A PR landing on master produces a merge commit (or a
PR-linked head); neither exists here. The controller audit at issue-create
time also found no PR wrapper and no review record for this SHA.
- **AC3 — changed files and diff summary: confirmed.**
`gitea_auth.py | 5 +++--` (+3/2). Single parent
`5ab5fe8583c07134d55dadf09381aecb67df246e`. The change moves
`PROJECT_ROOT` derivation above `load_dotenv()` and loads
`.env` relative to the project root instead of the process CWD.
## AC4 — why no immediate revert
- The dotenv fix is intentional and required for correct runtime behavior:
without it, `load_dotenv()` resolves `.env` against the process working
directory, which breaks MCP server launches whose CWD is not the project
root.
- The change is small (+3/2), self-contained in `gitea_auth.py`, and has
been running on master without incident since 2026-07-10.
- Reverting would re-introduce a real bug to remove a provenance defect —
the wrong trade. Provenance is repaired retroactively by this document,
issue #670, and the hardening landed under #671.
- If the controller later judges the change unsafe, a separate
revert/repair issue is the sanctioned path (issue #670, recommended
disposition option 4).
## AC5 — workflow-hardening linkage
Prevention already landed: **issue #671** (closed)
*“Block direct pushes to stable branches from MCP workflow sessions”*,
implemented by commit `5933d87647656643a67a50331c4c7b06ea751dad`
(`feat(guard): block direct stable-branch pushes from MCP workflow sessions`).
Shipped guardrails include:
- `gitea_record_stable_branch_push_attempt` — classifies proposed commands
for direct stable-branch push intent (`git push <remote> master`,
refspecs, `HEAD:master`, `--force`, dry-run intent, `:master` delete),
plus root/control-checkout local commits not carried by an issue branch,
and writes a durable `stable_branch_contamination` marker.
- `gitea_audit_stable_branch_contamination` — reconciler-only audit/clear
path; a contaminated worker session cannot self-clear.
- Review/merge/close/completion mutations fail closed while a
contamination marker is active.
## AC6 — PR #654 was not the source
- `2fa97c26` is the **first parent** of the #654 merge commit
`ec903b0d619e7a27d24aed272a890f4e5d381411`; it predates the #654 merge.
- First-parent history `5ab5fe8..ec903b0`:
`2fa97c2 fix(mcp): load dotenv relative to project root` followed by
`ec903b0 Merge pull request 'feat: lifecycle role/hazard labels ... (#603)' (#654)`.
- The #654 merger audit confirmed `ec903b0d` was a valid Gitea-API merge,
the `git push prgs master` attempt during that run was a no-op, and the
net change `2fa97c2..ec903b0` contained only the reviewed #603
lifecycle-label files.
- Conclusion: #654 merged reviewed content only; the unauthorized-path
defect is solely the earlier bare commit `2fa97c26`.
## Explicit non-actions (unchanged by this audit)
- No revert of `2fa97c26`.
- No force-push or history rewrite.
- No master mutation from the audit session.
-198
View File
@@ -1,198 +0,0 @@
# Instance-level fleet identity and health snapshots
Issue **#978**. Companion primitives: **#948** (worker ownership), **#975**
(heartbeat lifecycle), **#951** (restart receipts).
## Why this exists
The multi-instance fleet gate (#963) needs a *native* answer to:
* which application launches are live;
* which of the five namespace workers belong to each launch;
* whether two Codex (or Claude, Grok, …) launches are distinct;
* whether any live collision is real rather than a shared client type.
Process tables, PID proximity, configuration files, and direct SQLite access
are **not** production evidence. The sanctioned surface is the read-only MCP
tool `gitea_snapshot_instance_fleet` on **controller** and **reconciler**
namespaces.
## Identity hierarchy
| Identity | Scope | Who generates it | Lifetime |
| --- | --- | --- | --- |
| `client_type` | Application family (`codex`, `claude_code`, `gemini`, `grok`, …) | Launcher sets `GITEA_MCP_CLIENT` | Stable for the product |
| `client_instance_id` | One running application launch | **Trusted launcher**, once per launch, as `GITEA_MCP_CLIENT_INSTANCE` | Fresh launch → new ID; reconnect of same launch → same ID; full app restart → new ID |
| `fleet_run_id` | Operator-approved enrollment / canary cohort | Operator / controller sets `GITEA_MCP_FLEET_RUN_ID` | Duration of the approved rollout |
| `namespace` | `author` \| `reviewer` \| `merger` \| `controller` \| `reconciler` | Profile / MCP server binding | Process lifetime |
| `worker_id` / `worker_identity` | One namespace worker process | Runtime registry at first registration | New on worker restart; not reused |
| `session_id` | Client session ownership | Launcher `GITEA_MCP_CLIENT_SESSION` or runtime | Session lifetime |
| `generation_id` | One daemon launch | Runtime at process boot | New on process restart |
| `process_identity` / PID | OS process | Runtime | Process lifetime |
### Rules
1. Multiple active instances **may** share the same `client_type`.
2. Every application launch receives a **distinct** `client_instance_id`.
3. All five namespace workers of one launch report the **same**
`client_instance_id`.
4. Each namespace worker has a **distinct** `worker_identity`, process
identity, generation, and PID.
5. Instance identity is **never** inferred from PID proximity, timestamps, or
client type alone.
6. Live reuse of one `client_instance_id` with conflicting workers fails closed.
7. Sharing only a profile or `client_type` is **not** a duplicate.
This deliberately **replaces** any permanent `exactly_one_per_profile` fleet
model (#949 assumption) as the operating rule for multi-instance fleets.
## How five workers join one instance
1. The host starts one application instance (for example one Codex session).
2. The **production application launcher**
(`mcp_application_launcher.build_application_mcp_servers` /
`gitea_config.multi_namespace_launcher_entries`) mints exactly one
`client_instance_id` via `mcp_fleet_snapshot.generate_client_instance_id`
and injects the same env into every namespace worker:
```text
GITEA_MCP_CLIENT=<client_type>
GITEA_MCP_CLIENT_INSTANCE=<client_instance_id>
GITEA_MCP_INSTANCE_PROVENANCE=trusted_launcher
GITEA_MCP_CLIENT_SESSION=<session_id> # optional but recommended
GITEA_MCP_FLEET_RUN_ID=<enrollment id> # when on an approved canary
GITEA_CLIENT_MANAGED=1
GITEA_MCP_PROFILE=<role-profile>
```
3. The MCP client starts the five namespace processes (`gitea-author`,
`gitea-reviewer`, `gitea-merger`, `gitea-controller`, `gitea-reconciler`)
from that generated `mcpServers` map — each entry carries the **same**
instance ID.
4. Each worker registers once into the worker registry with its own
`worker_identity`, `namespace`, `generation_id`, and PID, but the shared
`client_instance_id`. Workers never mint a trusted instance ID themselves.
Only IDs matching the launcher format `inst-<client>-<timestamp>-<digest>`
are trusted. Missing, `legacy-pid-*` / `pid-*` placeholders, and ordinary
user-supplied strings fail closed as untrusted and cannot authorize
multi-instance fleet mutation safety.
Worker reconnect (same process, same registration) keeps the instance ID.
Worker restart (new process) mints a new worker identity and generation but
must still receive the same `GITEA_MCP_CLIENT_INSTANCE` from the parent
application if it is the same launch. Full application restart mints a new
`client_instance_id` (call the launcher without a prior ID).
### Mutation gate (instance-aware)
The capability / runtime diagnostic gate
(`_check_mcp_runtimes_diagnostics`) is **instance-aware**:
* Two legitimate application instances that share a profile (same
`GITEA_MCP_PROFILE`) are allowed when each has a distinct trusted
`client_instance_id`.
* More than one live worker for the same
`(client_instance_id, profile/namespace)` fails closed as a duplicate
namespace worker.
* Multiple processes sharing a profile **without** trusted instance
evidence still fail closed (indistinguishable from a duplicate).
## Approved fleet manifest
An operator (or controller enrollment step) obtains an approved manifest as a
list of expected instances, for example:
```json
{
"instances": [
{
"client_type": "codex",
"client_instance_id": "inst-codex-20260730T120000Z-abc123def456",
"fleet_run_id": "canary-963-2026-07-30",
"namespaces": ["author", "reviewer", "merger", "controller", "reconciler"]
},
{
"client_type": "codex",
"client_instance_id": "inst-codex-20260730T120100Z-fed654cba321",
"fleet_run_id": "canary-963-2026-07-30",
"namespaces": ["author", "reviewer", "merger", "controller", "reconciler"]
}
]
}
```
Pass that list as `expected_manifest` to `gitea_snapshot_instance_fleet`.
When a manifest is supplied:
* instances on the manifest but not live → `missing_expected`;
* live instances not on the manifest → `unmanifested`;
* both are **active blockers** for fleet-gate safety.
Without a manifest, the snapshot still enumerates the live fleet and classifies
identity collisions; it does not invent enrollment policy.
## Snapshot consistency
Each snapshot includes:
* `snapshot_at` — UTC timestamp;
* `consistency_token` / `registry_revision` — content digest over worker
identity, instance id, status, heartbeat, fencing, and generation.
Two successive snapshots can prove heartbeat continuity via
`mcp_fleet_snapshot.compare_snapshot_heartbeats`.
## Classification (active vs historical)
**Active blockers** (make `live_fleet_safe=false`):
* missing expected instance;
* unmanifested extra instance;
* duplicate namespace worker within one instance;
* live `client_instance_id` collision;
* reused worker / session / generation / process / PID / fencing identity;
* orphaned or unowned workers;
* unknown client;
* foreign-repository workers;
* old-revision workers;
* stale workers still marked active;
* legacy incomplete instance identity.
**Historical dead rows** (`status` released/superseded) are reported as
`historical_dead` findings with `active_blocker=false`. They never
automatically make the live fleet unsafe.
## Fail-closed behaviour and recovery
| Condition | Mutation safety | Diagnostic reads | Recovery |
| --- | --- | --- | --- |
| Two instances, same `client_type`, distinct IDs | Safe (if otherwise healthy) | Available | None needed |
| Live reuse of one `client_instance_id` | Unsafe | Available | Stop the colliding launch or re-issue a distinct ID |
| Two workers, same namespace, one instance | Unsafe | Available | Stop the extra worker |
| Legacy registration without trusted instance ID | Unsafe for fleet mutation | Available | Relaunch with launcher-issued `GITEA_MCP_CLIENT_INSTANCE` |
| Historical dead row only | Does not block alone | Available | No action required for fleet safety |
| Registry unavailable | Snapshot fails closed | N/A | Repair registry path / reconnect namespaces |
Incomplete or untrusted instance identity **cannot** authorize unsafe
mutation. Diagnostic reads remain available where `gitea.read` allows.
## Permissions
* **Allowed:** `controller`, `reconciler` with `gitea.read`.
* **Denied:** author, reviewer, merger (even with `gitea.read` for other tools).
* **No new mutation permissions** are granted to any role.
## Related surfaces
* `mcp_worker_identity` — registry, worker identity, heartbeats (#948, #975).
* `mcp_fleet_snapshot` — pure snapshot + classification (#978).
* `gitea_snapshot_instance_fleet` — sanctioned MCP tool (#978).
* `gitea_get_runtime_context` — single-process view (not fleet-wide).
## Non-goals
* Starting the #963 canary.
* Purging historical registry rows.
* Preserving “exactly one process per profile” as the permanent model.
* Using shell / process-table / SQLite inspection as production fleet evidence.
-70
View File
@@ -1,70 +0,0 @@
# MCP Config Drift Diagnostic & Sanctioned Repair Runbook (#672)
This document describes the diagnostic framework for detecting configuration drift between the active IDE MCP configuration (`~/.gemini/antigravity-ide/mcp_config.json`) and the offline/global canonical configuration (`~/.gemini/config/mcp_config.json`), and establishes the **sanctioned repair runbook**.
## Background & Problem Statement
Offline tools like `test_mcp_conn.py` test the global configuration (`~/.gemini/config/mcp_config.json`) via `subprocess.Popen`. However, the active IDE/client namespace uses `~/.gemini/antigravity-ide/mcp_config.json`. When required Gitea role servers (`gitea-author`, `gitea-reviewer`, `gitea-merger`, `gitea-reconciler`, `gitea-controller`, `gitea-tools`) are missing or carry mismatched profile environments in the active IDE config:
1. Offline tests pass (`test_mcp_conn.py` green).
2. The IDE client returns `EOF` / `transport closed` when attempting role-scoped mutations.
3. Operators misdiagnose missing server definitions as stale runtimes, leading to forbidden `pkill` attempts (#630) or `mtime` hacks (#655).
## Diagnostic Tool: `mcp_config_drift.py`
Run the diagnostic tool directly to compare configurations:
```bash
python3 mcp_config_drift.py --json
```
Or specify custom config locations:
```bash
python3 mcp_config_drift.py \
--active-config ~/.gemini/antigravity-ide/mcp_config.json \
--global-config ~/.gemini/config/mcp_config.json
```
### Key Diagnostic Outputs
- `in_sync`: Boolean indicating if all required Gitea role servers exist in the active IDE config with matching profile declarations.
- `missing_role_servers`: List of role servers present in global config but missing from active IDE config.
- `profile_mismatches`: List of profile environment mismatches per server.
- `reasons`: Explicit, human-readable list of drift causes.
All returned payloads automatically redact secret tokens, DSNs, Authorization headers, and private keys.
---
## Sanctioned Repair Path (Step-by-Step)
When `mcp_config_drift.py` reports drift (`in_sync: false`), execute the following **sanctioned repair steps**:
1. **Backup Active IDE Config:**
```bash
cp ~/.gemini/antigravity-ide/mcp_config.json ~/.gemini/antigravity-ide/mcp_config.json.bak
```
2. **Patch Active IDE Config:**
Copy the missing Gitea role server JSON blocks (`gitea-author`, `gitea-reviewer`, etc.) from `~/.gemini/config/mcp_config.json` into `~/.gemini/antigravity-ide/mcp_config.json`.
3. **Reconnect via IDE/Client:**
Use the IDE / client UI reconnection control (or restart the IDE client app).
4. **Verify Active Namespace Health:**
Invoke `gitea_whoami` (and optional `gitea_resolve_task_capability`) through the active IDE client on each required role namespace.
---
## FORBIDDEN Repair Actions (#630 / #655)
The following actions are **strictly forbidden** for config drift repair:
- ❌ **`pkill` or manual daemon process kill commands:** Process kills cause contamination and break active session leases.
- ❌ **`mtime` touch edits:** Artificial mtime modifications mask stale runtimes without updating configuration.
- ❌ **Source code edits:** Mutating python tool logic to bypass missing server entries.
- ❌ **Session-state edits:** Direct database or lock-file state mutation.
---
## Final Report Guidelines
A workflow final report **must not** rely on offline `test_mcp_conn.py` output alone. Final reports must include active-config evidence from live `gitea_whoami` calls on the active IDE namespaces.
+6 -33
View File
@@ -47,33 +47,18 @@ Do the steps in order. Stop as soon as a live **client-namespace** call succeeds
- Only the Gitea namespace fails → single-namespace transport close. Continue.
- Every server fails → restart the whole MCP client, not just one namespace.
2. **Request the sanctioned reconnect surface (#678), then reconnect through
the client — not the shell.** From a still-reachable Gitea MCP namespace
(or after host auto-reconnect), call:
```text
gitea_request_mcp_reconnect(
namespace="gitea-author", # or gitea-reviewer / gitea-merger / …
reason="transport_eof",
client="codex", # or claude_code / generic
)
```
The tool is **report-only**: it never restarts a process. It returns
namespace, profile, pid/session, startup SHA, current master SHA, boundary
status, and a **typed blocker** with exact operator UI steps for Codex
(Reload Developer Tools / per-server reconnect) or Claude Code (`/mcp`).
Then perform the host reconnect those steps describe so the client spawns a
fresh subprocess and re-opens the pipe. That clears the closed-client state
that a bare `kill`/respawn from a terminal does **not**.
2. **Reconnect the namespace through the client, not the shell.** Use the IDE /
client MCP-reconnect action for that server entry (in Claude Code:
`/mcp` → reconnect the affected `gitea-*` server). Reconnecting forces the
client to spawn a fresh subprocess and re-open the pipe. This clears the
closed-client state that a bare `kill`/respawn from a terminal does **not**.
3. **Do not "fix" it by importing the server or poking the process.** Reaching
for `python -c 'import gitea_mcp_server ...'`, raw JSON-RPC from a shell,
killing PIDs to force a respawn, or touching MCP config mtimes does **not**
restore the *client's* view of the namespace and violates the daemon-import
guard (#558, `docs/mcp-daemon-import-guard.md`). The only sanctioned repair
is a **client reconnect / relaunch** (or the typed operator path returned by
`gitea_request_mcp_reconnect`).
is a **client reconnect / relaunch**.
4. **Verify through the same path the workflow will use.** After reconnect, call
the specific tool the blocked workflow needs — not just any tool — through
@@ -168,19 +153,7 @@ not a tool argument: a session must never be able to authorize itself.
## Related
- #630 — manual daemon killing as contaminated recovery (this contrast, enforced).
- #657 — restart-path inventory and daemon classification.
- #686 — manual server launch detection & fail-closed provenance gate.
- #531 / #544 — stale-runtime detection (`ps`-based); sibling failure mode.
- #558 / `docs/mcp-daemon-import-guard.md` — why shell imports are not a repair.
- `docs/mcp-client-registration.md` — per-server registration contract.
- `docs/mcp-namespace-health.md` — probe sources and mutation enforcement.
## Sanctioned reconnect vs forbidden manual launch (#686)
In addition to manual process killing (#630), manually launching a duplicate role server from an ad hoc shell (`python3 mcp_server.py`) is forbidden and fail-closed:
- **Why manual launches are unsupported:** A terminal-launched `mcp_server.py` holds its own stdio transport; it can never bind to the IDE client's stdio pipes. It cannot restore a dropped IDE namespace, and a manual duplicate process masks stale client-managed runtimes for that profile, defeating stale-runtime gates.
- **Sanctioned path:** Supported recovery is IDE/client-managed reconnect only (`/mcp reconnect`, IDE restart, or sanctioned reconnect exposure).
- **Fail-closed enforcement (#686):** Mutating tools on a server lacking client-managed launch provenance (`GITEA_CLIENT_MANAGED=1`) refuse execution fail-closed with typed blocker `unsupported_manual_launch` and an exact next action. Unsupported `GITEA_*` env overrides (e.g. `GITEA_DUMMY`) are surfaced in diagnostics rather than silently ignored.
- **Inventory & staleness:** Staleness diagnostics ignore non-client-managed duplicates when evaluating runtime freshness and inventory duplicate processes per profile (#657, #686).
-71
View File
@@ -86,74 +86,3 @@ When a namespace returns EOF, follow
When blocked, repair the IDE namespace and re-record a healthy
`client_namespace` assessment before retrying the mutation.
## Connected vs Attached Tool Surface (#708)
MCP servers can report **Connected** at the CLI / host inventory layer while the **active LLM session exposes none of their tool namespaces**.
### Core principle
* **Connected status at host layer ≠ attached tools in active session.**
* Required preflight proof is **live tool visibility + `gitea_whoami` call** through the target namespace, not host `Connected` status alone.
* When servers report Connected but namespaces are absent from attached tools, classify as `mcp_connected_namespaces_missing`.
### Forbidden unsafe fallbacks
When `mcp_connected_namespaces_missing` is detected, workflows must **fail closed** and must **never** encourage or perform:
* direct imports of MCP server Python modules
* CLI or raw Gitea API mutations as a substitute for native tools
* profile hopping to another MCP profile/namespace to bypass the empty session
* session-state overrides or hand-edited session/ledger files
* process kills (`pkill`), config mtime touches, or `.env` edits
Only sanctioned recovery: **client reconnect path**, followed by full preflight (`whoami` → capability resolve → task).
### Native detection tool
`gitea_assess_mcp_namespace_attachment` classifies the condition and records it in
the session. Pass the namespaces the host reports Connected and the namespaces
actually attached to the active session tool surface:
| Argument | Meaning |
|---|---|
| `connected_servers` | Namespaces the host/CLI reports Connected |
| `attached_session_namespaces` | Namespaces exposed in the active session tool surface |
| `required_namespaces` | Namespaces this workflow needs (defaults to the role namespaces) |
| `discovery_cache_hit` / `discovery_cache_age_seconds` | Client tool-discovery cache state |
| `auto_attach_attempted` / `auto_attach_succeeded` | Whether the runtime auto-attached |
| `session_tool_snapshot_at` / `namespace_connected_at` | Epoch seconds, to detect startup ordering races |
It returns `discovery_status`
(`namespaces_attached` | `connected_but_namespaces_missing` | `disconnected`),
`missing_namespaces`, `proof_of_connected_vs_attached` (per namespace
`{connected, attached}`), `error_type`, `reconnect_required`, `auto_recovered`,
`startup_ordering_race`, `late_attaching_namespaces`, `sanctioned_recovery_tool`,
and a reconnect-only `exact_next_action`.
### Startup ordering
`session_tool_snapshot_at` earlier than a namespace's `namespace_connected_at`
means that namespace could not have been in the session snapshot, however healthy
it looks now. That is reported as `startup_ordering_race` with the affected
namespaces listed — the multi-role parallel-connect case where Connected flips
true after the session tool list was already captured.
### Fail-closed gate
The recorded verdict gates mutations, mirroring the #543 health gate: a namespace
that has not been assessed does not gate, but one recorded Connected-but-unattached
blocks the mutation whose role namespace it is
(`review_pr` / `submit_review``gitea-reviewer`, `merge_pr``gitea-merger`,
`work_issue` / `create_pr``gitea-author`). `gitea_submit_pr_review` and
`gitea_merge_pr` return the block reason plus the hard-stop policy string, and the
only offered recovery is `gitea_request_mcp_reconnect`.
### Telemetry
The `telemetry` block carries `connected_count`, `attached_count`,
`required_count`, `missing_count`, `discovery_status`, `discovery_cache_hit`,
`discovery_cache_age_seconds`, `reconnect_required`, `auto_attach_attempted`,
`auto_recovered`, `startup_ordering_race`, and `error_type`. It contains namespace
names and counts only — never tokens, endpoints, env values, or filesystem paths.
-94
View File
@@ -1,94 +0,0 @@
# MCP scoped recovery playbook (#669)
**Parent:** [#655](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/655)
**Vision / roadmap:** [#652](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/652) · [#653](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/653)
**Class matrix:** [#663](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/663) · `docs/mcp-restart-classes.md`
**Coordinator:** [#658](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/658) · `restart_coordinator.py`
**Audit lineage:** [#665](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/665)
## Decision
Full-server MCP reset is a **last resort**. Prefer the narrowest recovery that
can clear the symptom. The coordinator **refuses** `rolling_mcp_restart`,
`full_mcp_restart`, and `host_restart` unless:
1. The inventory carries a prior **attempt log** of at least one *insufficient*
narrower recovery, **or**
2. **Break-glass** is authorized
(`request_break_glass` + `GITEA_BREAKGLASS_RESTART_AUTHORIZATION`).
Break-glass still never bypasses the #663 class matrix (role/permission).
## Ladder (narrow → broad)
| Rank | Action | Self-service | Implementation / delegation |
|---:|---|---|---|
| 0 | `client_reconnect` | yes | Host auto-reconnect / client reconnect · [#584](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/584) · `docs/mcp-namespace-eof-recovery.md` |
| 1 | `capability_refresh` | yes | `gitea_resolve_task_capability` + `gitea_whoami` · [#610](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/610) · [#685](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/685) |
| 2 | `session_reconnect` | yes | Runtime rebind + explicit `worktree_path` · [#543](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/543) · [#618](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/618) |
| 3 | `configuration_reload` | no | Class `configuration_reload` · console reload · [#642](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/642) |
| 4 | `lease_recovery` | no | Lock/lease recovery paths · [#702](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/702) · [#753](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/753) · [#790](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/790) |
| 5 | `worker_restart` | no | Class `worker_restart` · #663 |
| 6 | `role_runtime_restart` | no | Class `role_runtime_restart` · console restart · #642/#663 |
| 7 | `connector_restart` | no | Class `connector_restart` · #663 |
| 8 | `rolling_mcp_restart` | no | Class `rolling_mcp_restart` · design [#668](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/668) · **attempt log required** |
| 9 | `full_mcp_restart` | no | Class `full_mcp_restart` · **attempt log required** |
| 10 | `host_restart` | no | Class `host_restart` · **attempt log required** |
Machine-readable source of truth: `recovery_playbook.RECOVERY_LADDER` and
`recovery_playbook.ladder_document()`.
## Attempt log shape
Each prior attempt is a mapping:
```json
{
"action": "client_reconnect",
"outcome": "insufficient",
"reason": "transport still closed after IDE reconnect",
"actor": "prgs-controller-12345",
"recorded_at": "2026-07-25T21:00:00+00:00"
}
```
Outcomes that count toward escalation: `failed`, `insufficient`, `denied`,
`unresolved`, `timeout`, `error`.
Pass attempts into the coordinator via inventory
`prior_recovery_attempts` or the MCP tool argument
`prior_recovery_attempts_json` on `gitea_request_mcp_restart`.
Helper: `recovery_playbook.build_attempt_record(...)`.
## Symptom → first rung
`recovery_playbook.recommend_actions(symptoms=[...])` maps symptoms such as
`transport_eof`, `stale_capability`, `stale_lease`, `daemon_corrupt` to the
narrowest recommended action, then walks the ladder. Soft recommendations
never replace the hard gate on broad restarts.
## Enforcement points
1. **`recovery_playbook.assess_escalation`** — pure gate.
2. **`restart_coordinator.evaluate_restart_impact`** — when `restart_class` is
set (policy-enforced path), broad classes require the gate; report fields
`attempt_log_satisfied`, `playbook_escalation`, `break_glass`.
3. **`gitea_request_mcp_restart`** — accepts attempt JSON and env-authorized
break-glass; never restarts a process.
## Metrics
`recovery_playbook.recovery_metrics(attempts)` reports the fraction of
successful recoveries that avoided full/host restart
(`fraction_avoided_full_restart`).
## Non-goals
* HA multi-instance execution (#668 design only here).
* Normalizing `pkill` (#630 contamination stays forbidden).
* Silent mutation of leases or processes from the playbook itself.
## Manual process kills
Remain forbidden and contaminating (#630). The playbook never recommends them.
-53
View File
@@ -1,53 +0,0 @@
# MCP restart classes and blast-radius permissions (#663)
This is the machine-enforced class matrix used by
`restart_coordinator.RESTART_CLASS_POLICIES`. It implements the narrower-first
recovery ladder from #655 and the authorization policy from #656, using the
path inventory from #657 and the impact coordinator from #658. Product and
delivery lineage: vision #652 and roadmap #653.
Unknown class names are denied. The coordinator requires both the class
permission and an eligible request role. Approval gates are additional: a
caller cannot turn a request permission into execution authority.
| Restart class | Required permission | Expected blast radius | Drain requirement | Approval requirement | Audit requirement | Recovery behavior |
|---|---|---|---|---|---|---|
| `client_reconnect` | `mcp.reconnect.client` | none | none | self service | class, actor, client namespace, reason, outcome | Reconnect only the caller's client transport. No daemon or peer work changes. |
| `session_reconnect` | `mcp.reconnect.session` | low | requesting-session safe point | self service | class, actor, session, reason, outcome | Rebind identity, capability, and workspace state for one session. |
| `worker_restart` | `mcp.restart.worker.request` | low | target worker | controller approval + automated gates | class, actor, worker, approval, scoped drain, outcome | Restart one worker after its own leases and mutations drain. |
| `role_runtime_restart` | `mcp.restart.role_runtime.request` | medium | target role runtime | controller approval + automated gates | class, actor, role namespace, approval, scoped drain, outcome | Restart and re-probe one role runtime; unrelated roles remain available. |
| `connector_restart` | `mcp.restart.connector.request` | medium | target connector | controller approval + automated gates | class, actor, connector, approval, scoped drain, outcome | Restart one connector while unrelated runtimes remain available. |
| `configuration_reload` | `mcp.reload.configuration.request` | low | mutation quiesce | controller approval + automated gates | class, actor, configuration revision, approval, outcome | Gracefully reload configuration without replacing the daemon. |
| `rolling_mcp_restart` | `mcp.restart.rolling.request` | medium | one instance at a time | controller approval + automated gates | class, actor, instance order, approval, per-instance drains, outcome | Drain, restart, verify, and restore each instance before advancing. |
| `full_mcp_restart` | `mcp.restart.full.request` | high | all sessions and mutations | controller approval + automated gates | class, actor, full impact, approval, full drain proof, outcome | Replace the complete MCP runtime only after a verified full drain. |
| `host_restart` | `mcp.restart.host.request` | high | all host work | controller approval + infrastructure operator | class, actor, host/change or incident id, approval, full drain proof, outcome | Hand off to infrastructure ownership and reconcile every runtime afterward. |
## Drain boundary
Only `full_mcp_restart` and `host_restart` set `full_drain_required=true`.
Reconnects and configuration reloads do not disrupt peer sessions. Worker,
role-runtime, and connector restarts evaluate only their explicitly named
target. Rolling restart drains one instance at a time. Missing required target
scope denies the request rather than silently widening it to a full restart.
## Permission and approval boundary
Author, reviewer, merger, and reconciler roles may self-request reconnects and
request scoped worker/role/connector/reload recovery. They cannot request
rolling, full, or host restart classes. Controller/operator/admin roles may
request the broader classes, while execution remains operator/admin-owned.
Controller approval is independently required for every class above a session
reconnect. Host restart additionally requires infrastructure-operator proof.
The MCP request tool derives class permissions from its authenticated runtime
role. It does not accept caller-supplied permissions. Controller and operator
authorization are read from the already-running daemon environment, never
from a request argument.
## Audit and failure behavior
Every impact audit and every console restart/reload audit includes a
`restart_class` field. The impact audit also includes the exact
`required_permission`. Unknown classes, missing permissions, ineligible roles,
missing approval, missing scoped targets, and incomplete inventory all deny
fail closed. Manual process kills remain forbidden and contaminating (#630).
-162
View File
@@ -1,162 +0,0 @@
# MCP restart coordinator and impact analysis (#658)
Before any sanctioned MCP restart, a central coordinator evaluates the live
control-plane state and produces an **impact preview** so operators and the web
console (#642 / #652) can see the blast radius *before* concurrent LLM work is
disrupted. Uncoordinated restarts destroy in-flight author/reviewer/merger work
and give operators no way to see what they are about to break.
This lands the coordinator + impact DTO + the MCP tool. It is the single
sanctioned entry point for restart evaluation post-#657 (which inventoried the
restart/reload/kill paths).
The **drain-proof hard gate now executes inside this tool** (#661, via PR #882):
an apply request (`dry_run=False`) is evaluated against a drain proof here and
denied when that proof is missing, expired, unclean, tampered with, or stale.
It is no longer a separate child operation. What remains a later child is only
the **execution** step — actually stopping and restoring a process. This tool
still never restarts anything: `apply_supported` is always `false` and
`restart_performed` is always `false`.
The coordinator now routes every request through the restart-class policy
matrix defined for #663. See
[`mcp-restart-classes.md`](./mcp-restart-classes.md) for permissions, expected
blast radius, scoped drain and approval requirements, audit fields, and
recovery behavior for all nine classes.
## Components
| Piece | Where | Responsibility |
|-------|-------|----------------|
| `restart_coordinator.evaluate_restart_impact` | `restart_coordinator.py` | Pure classification: inventory → impact report DTO. No I/O, no restart. |
| `RestartImpactReport` / `SessionImpact` / `LeaseImpact` | `restart_coordinator.py` | Console-facing DTO (`.as_dict()` is JSON-serializable). |
| `ControlPlaneDB.list_sessions` | `control_plane_db.py` | Read-only session inventory (the process-level unit a restart kills). |
| `gitea_request_mcp_restart` | `gitea_mcp_server.py` | MCP tool: gathers inventory from the #613 DB, calls the coordinator, returns the report, and on `dry_run=False` runs the #661 drain-proof hard gate. Never restarts a process. |
| `drain_proof.gate_apply_restart` | `drain_proof.py` | The #661 hard gate: verifies a drain proof against the current impact fingerprint, or records an authorized break-glass bypass. |
## Dimensions evaluated
The coordinator classifies the inventory across the dimensions #658 requires:
- **Sessions** — every active MCP session; a restart terminates all of them.
Liveness = `status == active` **and** the owner pid is alive **and** the
heartbeat is fresh (default window 15 min). Dead/stale sessions do not count
toward blast radius.
- **Leases / locks** — control-plane leases joined with work items and their
freshness (`lease_lifecycle.classify_lease_freshness`). Only `active` (live
owner) leases are *disruptive*; expired / released / dead-process leases never
withhold a restart.
- **Issue / PR work** — the issues and PRs behind disruptive leases.
- **Mutations / critical sections** — a live lease carrying an author worktree
or a mutating phase (`implementing`, `publishing`, `merging`, …) is a
critical section a restart must not sever.
- **Terminal (merge) lock** — an active terminal lock always makes a restart
unsafe.
- **Prior recovery attempts** — narrower recovery already tried (e.g. sanctioned
client reconnects) is echoed so the operator sees the escalation history.
## Verdict
Exactly three verdicts, matching the acceptance criteria:
| Verdict | `allow_restart` | Meaning |
|---------|-----------------|---------|
| `safe` | `true` | No other live sessions, no live leases, no terminal lock. |
| `unsafe` | `false` | Live work would be disrupted and no operator override is present — **or** the inventory could not be completed (fail closed). |
| `override` | `true` | Live work present, but an operator override accepts the blast radius. |
`override_would_allow` tells the console whether an override path exists for the
current state. `blast_radius` is a `none` / `low` / `medium` / `high` severity
band derived from the affected session and work counts.
### Fail closed
If the control-plane inventory cannot be completed (DB unavailable, a listing
failed), `inventory_complete` is `false` and the verdict is `unsafe` / deny. An
incomplete evaluation must never green-light a restart.
### Operator override authority
Override authority is read from the environment variable
`GITEA_OPERATOR_RESTART_OVERRIDE_AUTHORIZATION` and **never** from a tool
argument. A worker session cannot set an environment variable on an
already-running daemon, so override cannot be self-asserted (same pattern as the
#630 daemon-maintenance authorization). The `request_override` tool argument only
expresses caller intent; it takes effect solely when the environment
authorization is present.
## The tool
```text
gitea_request_mcp_restart(remote, host, org, repo,
dry_run=True, request_override=False,
session_id=None, limit=200,
restart_class="full_mcp_restart",
target_session_id=None, target_role=None,
target_connector=None,
drain_proof_json=None,
request_break_glass=False,
prior_recovery_attempts_json=None)
```
It **never restarts anything**: `apply_supported` is always `false` and
`restart_performed` is always `false`.
`prior_recovery_attempts_json` (#669) is an optional JSON array of prior
narrow recovery attempts. Rolling / full / host classes require at least one
*insufficient* narrower attempt (or authorized break-glass). See
`docs/mcp-recovery-playbook.md`.
### Dry-run versus apply
| Call | Behavior |
|------|----------|
| `dry_run=True` (default) | Read-only impact preview. No drain proof is required or consulted. |
| `dry_run=False` | The #661 drain-proof hard gate runs **in this tool**. The outcome is reported under `apply_gate` / `apply_authorized`; a denial also returns a durable `incident` descriptor. Still no restart. |
### Authorization ordering
An apply requires **both** authorizations, and they are independent:
1. **Restart-class authorization** (#663 / #669) — the requester's role and
permissions must allow the requested class, the class's approval requirement
must be satisfied, any target-scoped class must name its target, and broad
classes must satisfy the recovery-playbook attempt-log gate. Failing any of
these makes `allow_restart` `false`.
2. **Drain-proof gate** (#661) — a valid, unexpired, clean proof bound to the
current impact fingerprint, or an authorized break-glass.
`apply_authorized` is the conjunction: `gate.allow and allow_restart`. A clean
drain proof therefore cannot override a class or requester-role denial, and a
denied class never reports an authorized apply. `apply_gate` carries
`drain_gate_allow` and `restart_class_authorized` so a denial is attributable to
the authorization that produced it.
### Break-glass
Break-glass bypasses the **drain proof only** — never the restart-class matrix
(role/permission). Separately, authorized break-glass also satisfies the #669
attempt-log requirement for broad restarts (rolling/full/host), because that
gate is not a class-matrix permission check.
It is honoured solely when `request_break_glass` is set *and* the environment
carries `GITEA_BREAKGLASS_RESTART_AUTHORIZATION`; like operator override, the
tool argument expresses caller intent and cannot be self-asserted by a worker
session. `break_glass_requested` and `break_glass_authorized` are both reported,
so a bypass is never silent.
### Fail closed on apply
A missing, malformed, expired, unclean, tampered, or fingerprint-stale drain
proof denies the apply and returns an `incident` descriptor. An unknown restart
class denies before any of this. Ambiguity always denies.
## Audit
Every evaluation carries an `audit_record` (event, coordinator version, verdict,
restart class, required permission, allow decision, blast radius, counts,
timestamp) so restart decisions are
auditable. No secrets flow through the coordinator — session ids, pids, and
profiles are operational metadata only.
A representative dry-run report is in
[`mcp-restart-impact-sample.json`](./mcp-restart-impact-sample.json).
-148
View File
@@ -1,148 +0,0 @@
{
"coordinator_version": "1.0.0-issue-658",
"evaluated_at": "2026-07-24T06:00:00+00:00",
"dry_run": true,
"restart_performed": false,
"inventory_complete": true,
"incomplete_reasons": [],
"verdict": "unsafe",
"allow_restart": false,
"override_would_allow": true,
"operator_override": false,
"blast_radius": "high",
"reasons": [
"live work would be disrupted; restart denied without operator override",
"1 critical section(s) in flight (active lease with a live owner)"
],
"affected_sessions": [
{
"session_id": "prgs-author-30988-d6f43c25",
"role": "author",
"profile": "prgs-author",
"pid": 1,
"status": "active",
"alive": true,
"heartbeat_stale": false,
"is_requester": false,
"live": true
},
{
"session_id": "prgs-reviewer-4157-0ce9",
"role": "reviewer",
"profile": "prgs-reviewer",
"pid": 1,
"status": "active",
"alive": true,
"heartbeat_stale": false,
"is_requester": true,
"live": true
}
],
"affected_leases": [
{
"lease_id": "lease-abc",
"session_id": "prgs-author-30988-d6f43c25",
"role": "author",
"phase": "implementing",
"freshness": "active",
"work_kind": "issue",
"work_number": 658,
"worktree_path": "/repo/branches/feat-issue-658",
"disruptive": true,
"is_mutation": true,
"is_critical_section": true
},
{
"lease_id": "lease-dead",
"session_id": "prgs-author-91485",
"role": "author",
"phase": "allocated",
"freshness": "stale_dead_process",
"work_kind": "issue",
"work_number": 651,
"worktree_path": null,
"disruptive": false,
"is_mutation": false,
"is_critical_section": false
}
],
"critical_sections": [
{
"lease_id": "lease-abc",
"session_id": "prgs-author-30988-d6f43c25",
"role": "author",
"phase": "implementing",
"freshness": "active",
"work_kind": "issue",
"work_number": 658,
"worktree_path": "/repo/branches/feat-issue-658",
"disruptive": true,
"is_mutation": true,
"is_critical_section": true
}
],
"affected_issues": [
658
],
"affected_prs": [],
"mutations": [
{
"lease_id": "lease-abc",
"session_id": "prgs-author-30988-d6f43c25",
"role": "author",
"phase": "implementing",
"freshness": "active",
"work_kind": "issue",
"work_number": 658,
"worktree_path": "/repo/branches/feat-issue-658",
"disruptive": true,
"is_mutation": true,
"is_critical_section": true
}
],
"terminal_lock": null,
"ack_state": {
"prgs-author-30988-d6f43c25": "pending"
},
"prior_recovery_attempts": [
{
"kind": "client_reconnect",
"at": "2026-07-24T06:00:00+00:00",
"outcome": "insufficient"
}
],
"counts": {
"sessions_total": 2,
"sessions_live_other": 1,
"leases_total": 2,
"leases_disruptive": 1,
"critical_sections": 1,
"mutations": 1,
"affected_issues": 1,
"affected_prs": 0,
"prior_recovery_attempts": 1
},
"audit_record": {
"event": "restart_impact_evaluated",
"coordinator_version": "1.0.0-issue-658",
"evaluated_at": "2026-07-24T06:00:00+00:00",
"dry_run": true,
"operator_override": false,
"requesting_session_id": "prgs-reviewer-4157-0ce9",
"inventory_complete": true,
"verdict": "unsafe",
"allow_restart": false,
"blast_radius": "high",
"counts": {
"sessions_total": 2,
"sessions_live_other": 1,
"leases_total": 2,
"leases_disruptive": 1,
"critical_sections": 1,
"mutations": 1,
"affected_issues": 1,
"affected_prs": 0,
"prior_recovery_attempts": 1
}
}
}
-91
View File
@@ -1,91 +0,0 @@
# MCP restart / reload / kill path inventory (#657)
Complete inventory of every code, script, and host path that can **restart,
reload, reconnect, kill, or force-recreate** an MCP process in this project,
with each path classified and linked to the guard that constrains it.
This document is the human-readable companion to the machine-readable registry
in [`mcp_restart_paths.py`](../mcp_restart_paths.py). The two are kept in
lock-step by [`tests/test_mcp_restart_paths.py`](../tests/test_mcp_restart_paths.py):
every `path_id` below must appear in this file, and the source guards are run
against the live tree.
Roadmap linkage: this inventory is the enumeration step of the restart
governance work — parent **#655**, restart-governance ADR **#656**, vision
**#652**, roadmap **#653**. Related detection/guard work: master-advance
staleness **#591**/**#420**, side-effect-free resolver **#685**, transport flap
**#584**, manual-kill contamination **#630**.
## Classifications
| Classification | Meaning |
|---|---|
| `sanctioned_narrow_recovery` | One-shot, safe-by-construction recovery that never targets the running daemon. |
| `guarded_fail_closed` | Detects a restart-requiring condition, then fails mutations closed and emits reconnect guidance. Never self-restarts. |
| `forbidden` | A workflow-safety violation; where an LLM tool could invoke it, it is marked contamination. |
| `removed` | A previously-existing unguarded restart primitive that has been deleted; a regression guard keeps it absent. |
| `host_residual` | Behavior owned by the host/IDE, outside this process's control. Documented, not code-guarded here. |
## The rule
**No component may perform an unguarded full restart of the MCP daemon.** The
in-process daemon (`gitea_mcp_server.py`, `mcp_server.py`,
`role_session_router.py`) must never replace or terminate its own process:
replacing the process after the host has wired up the stdio pipes desyncs the
JSON-RPC transport (observed with Antigravity/Cascade hosts). Recovery is owned
by the host/operator via a client reconnect — the daemon only ever *detects*
and *fails closed*.
## Inventory
| path_id | Classification | Mechanism | Guard | Refs |
|---|---|---|---|---|
| `cli_venv_bootstrap_execv` | sanctioned_narrow_recovery | CLI wrapper scripts re-exec into `venv/bin/python3` via `os.execv`, guarded by `sys.executable != venv_python`. | One-shot pre-import bootstrap; runs before any MCP transport exists and only when not already on the venv interpreter; idempotent guard prevents a re-exec loop. | #657 |
| `daemon_self_replacement` | forbidden | The daemon replacing/terminating its own process (`os.execv`/`os.kill`/`os._exit`) to reload code. | Forbidden by design; enforced against the source tree by `assert_no_daemon_self_replacement()`. | #657, #584 |
| `legacy_auto_restart_helper` | removed | A helper (`_trigger_mcp_auto_restart`) that actively restarted the server from the read-only resolver path. | Removed in #685; kept absent by `assert_auto_restart_helper_absent()`. | #685, #657 |
| `config_touch_reload` | removed | Touching (utime) the MCP client config to make the host reload the server. | Removed from the resolver in #685: stale detection is report-only, never mutating config, spawning threads, or calling `os._exit`. | #685, #657 |
| `master_advance_auto_restart` | guarded_fail_closed | On-disk master advancing past the running code. | `master_parity_gate` captures startup parity and blocks mutations while stale, emitting restart guidance; the process never self-restarts. | #420, #591, #657 |
| `stale_runtime_resolver_reconnect` | guarded_fail_closed | The capability resolver detecting a stale serving process. | Report-only (#685): returns `restart_required`/`stop_required` and an exact reconnect action; no restart, thread, config touch, or `os._exit`. | #685, #657, #678 |
| `codex_client_reconnect_request` | guarded_fail_closed | `gitea_request_mcp_reconnect` report-only tool for Codex/LLM sessions. | Report-only (#678): returns namespace/profile/pid/startup SHA/master SHA/boundary status and a typed operator blocker with exact client UI steps; never restarts or kills. | #678, #630, #685, #657 |
| `manual_daemon_kill` | forbidden | Shell kills of the daemon: `pkill -f mcp_server.py`, `killall`, broad `pkill -f python` sweeps, or `kill <pid>` of a daemon pid. | Forbidden (#630): `runtime_recovery_guard` classifies these as contamination and `gitea_record_daemon_process_kill_attempt` writes a durable marker that fails later mutations closed. Operator maintenance authorization is read only from the environment. | #630, #657 |
| `conflict_marker_infra_stop` | guarded_fail_closed | The daemon entrypoint scans for unresolved merge-conflict markers at startup and stops (`sys.exit(1)`). | Fail-closed startup stop, not a restart: the process exits and waits for the operator to resolve conflicts and relaunch; never loops. | #657 |
| `ide_client_reconnect` | host_residual | A manual `/mcp reconnect` (or equivalent host action) that recreates the MCP client connection. Agents obtain exact UI steps via `gitea_request_mcp_reconnect` (#678). | Outside this process's control; the sanctioned recovery the gates point operators toward. No in-process code initiates it. | #584, #656, #657, #678 |
| `profile_switch_runtime` | sanctioned_narrow_recovery | Switching the active execution profile at runtime (dynamic-profile mode). | In-process and restart-free: `runtime_switching_supported` is true, so a switch rebinds capability without recreating the process. | #656, #657 |
## Guards enforced in CI
`tests/test_mcp_restart_paths.py` asserts, against the live source tree:
1. **Registry well-formedness** — every path has a valid classification, a
non-empty guard description, references, and locations; ids are unique; all
five classifications are represented.
2. **Unknown restart attempts fail closed**
`assert_restart_attempt_registered()` raises `UnknownRestartPathError` for
any path id not in this inventory, so a novel/unnamed restart primitive
cannot slip through silently.
3. **Daemon never self-replaces**`assert_no_daemon_self_replacement()` scans
the daemon modules for `os.execv`/`os.kill`/`os._exit`/`os.abort` calls
(comment/docstring mentions are ignored) and finds none.
4. **Legacy helper stays removed**`assert_auto_restart_helper_absent()`
confirms `_trigger_mcp_auto_restart` has not returned.
5. **pkill stays forbidden** — a daemon `pkill` command still classifies as
contamination via `runtime_recovery_guard`.
## Residual host behaviors (outside process control)
* `/mcp reconnect` in the IDE/host — the sanctioned recovery for stale-runtime,
transport-flap (#584), and worktree-binding conditions. The daemon can only
emit guidance toward it.
* Host-level process management (the operator relaunching the daemon after a
fail-closed stop, or after resolving merge conflicts).
These are documented rather than code-guarded because the process cannot
observe or gate them from inside itself.
## Rollout
Per #657, guards are introduced flag-free as **regression assertions** (they
codify invariants that already hold) before any hard runtime block is layered
on. When the restart coordinator (#655/#656) lands, registered paths gain a
coordinator token/capability check; unregistered attempts already fail closed
today via `assert_restart_attempt_registered()`.
-13
View File
@@ -56,7 +56,6 @@ that gates each call, not which tools exist.
- `gitea_assess_conflict_fix_push`
- `gitea_assess_gitea_operation_path`
- `gitea_assess_master_parity`
- `gitea_assess_mcp_namespace_attachment`
- `gitea_assess_mcp_namespace_health`
- `gitea_assess_pr_sync_status`
- `gitea_assess_review_merge_state_machine`
@@ -65,13 +64,11 @@ that gates each call, not which tools exist.
- `gitea_assess_work_issue_duplicate`
- `gitea_assess_worktree_cleanup_integrity`
- `gitea_audit_config`
- `gitea_audit_missing_worktree_bindings`
- `gitea_audit_runtime_recovery_contamination`
- `gitea_audit_stable_branch_contamination`
- `gitea_audit_worktree_cleanup`
- `gitea_authorize_reconciliation_cleanup_phase`
- `gitea_authorize_review_correction`
- `gitea_bootstrap_author_issue_worktree`
- `gitea_capability_stop_terminal_report`
- `gitea_capture_branches_worktree_snapshot`
- `gitea_check_pr_eligibility`
@@ -103,9 +100,7 @@ that gates each call, not which tools exist.
- `gitea_get_profile`
- `gitea_get_runtime_context`
- `gitea_get_shell_health`
- `gitea_heartbeat_issue_lock`
- `gitea_heartbeat_reviewer_pr_lease`
- `gitea_inspect_issue_lock_contract`
- `gitea_inspect_workflow_lease`
- `gitea_issue_irrecoverable_provenance_authorization`
- `gitea_list_dependency_edges`
@@ -127,26 +122,19 @@ that gates each call, not which tools exist.
- `gitea_post_heartbeat`
- `gitea_publish_unpublished_issue_branch`
- `gitea_quarantine_contaminated_review`
- `gitea_rebind_dirty_same_claimant_author_session`
- `gitea_reclaim_expired_workflow_lease`
- `gitea_reconcile_after_restart`
- `gitea_reconcile_already_landed_pr`
- `gitea_reconcile_issue_claims`
- `gitea_reconcile_merged_cleanups`
- `gitea_reconcile_missing_worktree_bindings`
- `gitea_reconcile_superseded_by_merged_pr`
- `gitea_record_daemon_process_kill_attempt`
- `gitea_record_irrecoverable_decision_lock_provenance`
- `gitea_record_pre_review_command`
- `gitea_record_shell_spawn_outcome`
- `gitea_record_stable_branch_push_attempt`
- `gitea_recover_dirty_orphaned_issue_worktree`
- `gitea_recover_incomplete_bootstrap_lock`
- `gitea_release_merger_pr_lease`
- `gitea_release_reviewer_pr_lease`
- `gitea_release_workflow_lease`
- `gitea_request_mcp_reconnect`
- `gitea_request_mcp_restart`
- `gitea_resolve_task_capability`
- `gitea_resume_review_draft`
- `gitea_review_pr`
@@ -159,7 +147,6 @@ that gates each call, not which tools exist.
- `gitea_sentry_reconcile_issue`
- `gitea_sentry_watchdog`
- `gitea_set_issue_labels`
- `gitea_snapshot_instance_fleet`
- `gitea_submit_pr_review`
- `gitea_update_pr_branch_by_merge`
- `gitea_validate_review_final_report`
@@ -1,143 +0,0 @@
# Model Usage, Token Cost, Latency, and Workflow Analytics (Phase 4)
- **Tracking Issue:** [#651](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/651)
- **Parent Epic:** [#631](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/651)
- **Console Surface:** `/analytics`, `/api/v1/analytics`, `/api/v1/analytics/usage`
## 1. Overview
The Web Console Analytics module provides durable, aggregate visibility into **model usage, token cost, latency percentiles, and workflow-stage performance** across projects, worker roles, AI models, issues, and PRs.
### Non-Goals
- No mandatory client-side telemetry that leaks prompts or secret keys.
- No third-party payment provider or billing integration.
- No automatic model routing changes without controller policy (#647).
---
## 2. Event Schema (`usage_events`)
Usage metrics are stored in the control-plane database under table `usage_events`.
| Column | Type | Description |
|---|---|---|
| `usage_id` | `INTEGER` | Primary key (autoincrement) |
| `session_id` | `TEXT` | Optional active session identifier |
| `remote` | `TEXT` | Known Gitea instance (`dadeschools` or `prgs`) |
| `org` | `TEXT` | Repository owner / organization |
| `repo` | `TEXT` | Repository name |
| `project_id` | `TEXT` | Optional project identifier |
| `role` | `TEXT` | Active worker role (`author`, `reviewer`, `merger`, `reconciler`, `controller`) |
| `model` | `TEXT` | LLM model identifier (e.g. `gemini-3.6-flash`, `claude-3-5-sonnet`) |
| `issue_number` | `INTEGER` | Correlated Gitea issue number (optional) |
| `pr_number` | `INTEGER` | Correlated Gitea PR number (optional) |
| `stage` | `TEXT` | Workflow stage (`preflight`, `implementation`, `review`, `merge`, `reconciliation`) |
| `input_tokens` | `INTEGER` | Input token count (optional / nullable) |
| `output_tokens` | `INTEGER` | Output token count (optional / nullable) |
| `total_tokens` | `INTEGER` | Total token count (optional / nullable) |
| `estimated_cost_usd` | `REAL` | Estimated USD cost (optional / nullable) |
| `latency_ms` | `INTEGER` | Request latency in milliseconds (optional / nullable) |
| `duration_ms` | `INTEGER` | Stage execution duration in milliseconds (optional / nullable) |
| `status` | `TEXT` | Outcome status (`success`, `failure`, `timeout`) |
| `metadata` | `TEXT` | Redacted metadata or summary string |
| `created_at` | `TEXT` | ISO 8601 UTC timestamp |
---
## 3. Handling of Missing Data ("Unknown" vs. Zero Fabrication)
To ensure operational metrics accurately reflect evidence:
- **Untracked or missing metrics are displayed as `Unknown`**, never zero-fabricated.
- If an event omits `estimated_cost_usd`, `latency_ms`, or token counts, the aggregator marks those fields as missing (`None`) rather than defaulting to `0` or `$0.00`.
- Summary tables and KPI cards explicitly indicate when data is unmeasured or partially reported.
---
## 4. Redaction & Security Rules
Per `#633` security policy:
- Free-text fields (`metadata`, `prompt_summary`, `session_id`) are run through `console_redaction.redact_text` before persistence and output serialization.
- Secret tokens, keychain commands, authorization headers, passwords, and JWTs are stripped automatically.
---
## 5. Opt-in Instrumentation Guide
Applications, MCP servers, and background sessions can report usage metrics through either Python API or HTTP ingestion.
### Python Ingestion
```python
from webui.analytics_loader import record_usage
record_usage(
remote="dadeschools",
org="Scaled-Tech-Consulting",
repo="Gitea-Tools",
role="author",
model="gemini-3.6-flash",
issue_number=651,
stage="implementation",
input_tokens=1420,
output_tokens=380,
total_tokens=1800,
estimated_cost_usd=0.00045,
latency_ms=320,
duration_ms=4500,
status="success",
metadata={"note": "Implementation of analytics module"},
)
```
### HTTP Ingestion API (authorized write)
`POST /api/v1/analytics/usage` is a **gated write**. It runs through
`console_authz` action `record_analytics_usage` (operator+, Phase 2 execution).
Unauthenticated or phase-inactive requests receive **403** and do not write.
Prefer in-process `record_usage` for MCP / session instrumentation.
```http
POST /api/v1/analytics/usage HTTP/1.1
Content-Type: application/json
# Requires authenticated principal with record_analytics_usage execution enabled
{
"remote": "dadeschools",
"org": "Scaled-Tech-Consulting",
"repo": "Gitea-Tools",
"role": "author",
"model": "gemini-3.6-flash",
"issue_number": 651,
"stage": "implementation",
"input_tokens": 1420,
"output_tokens": 380,
"total_tokens": 1800,
"estimated_cost_usd": 0.00045,
"latency_ms": 320,
"duration_ms": 4500,
"status": "success",
"metadata": "Analytics schema landed"
}
```
### Retention
`usage_events` is retained with hard caps applied on every write:
| Limit | Default |
|---|---|
| Max rows | 10,000 (`ControlPlaneDB.USAGE_EVENTS_MAX_ROWS`) |
| Max age | 90 days (`ControlPlaneDB.USAGE_EVENTS_RETENTION_DAYS`) |
Older rows (by `created_at`) and excess oldest rows (by `usage_id`) are deleted
after each insert so unbounded growth / DoS-by-volume cannot fill the DB.
---
## 6. Querying Analytics API
```http
GET /api/v1/analytics?role=author&stage=implementation HTTP/1.1
```
Returns `AnalyticsSnapshot` JSON containing aggregations (`by_model`, `by_stage`, `by_role`, `by_work_item`, `by_project`) and latency percentiles (`p50`, `p90`, `p95`, `p99`).
@@ -1,35 +0,0 @@
# Web Console: Sentry/GlitchTip Observability & Incident Bridge Console (#649)
This document describes the Phase 4 observability console surface integrated into the MCP Control Plane Web Console (`webui/`), backed by the #612 incident bridge and the #613 control-plane DB substrate.
## Architectural Authority Model (ADR Alignment)
Per the Web Console Architecture ADR (`docs/architecture/webui-control-plane-console-architecture-adr.md`):
| Layer | Responsibility | Authority |
|---|---|---|
| **Gitea** | Durable work record | Issues, PRs, comments, reviews, labels, merges |
| **Control-plane DB** | Live coordination & linkage | `incident_links` table, session leases, allocations |
| **Sentry / GlitchTip** | Observability input | Unresolved incidents, error events, stack traces |
| **Incident Bridge (#612)** | Reconciliation engine | Reconciles provider observations into Gitea issues |
| **Web Console (`webui/`)** | Read-only projection & gated actions | Projects connection health & correlation links; gates writes |
> **Key Rule:** Raw monitoring incidents are **never** assignable control-plane `work_items`. They remain observation input only.
## Redaction Boundary Invariants
1. **No secrets in returns or rendering:** Auth tokens (`SENTRY_AUTH_TOKEN`, `GLITCHTIP_AUTH_TOKEN`), DSNs, `Authorization` headers, and sensitive local file paths are passed through `webui.console_redaction` before leaving the server.
2. **Safe projection:** Connection objects report `credentials_present: true/false` rather than exposing raw keys or headers.
## Console Endpoints
- **HTML Surface:** `GET /observability` — Renders provider connection cards, error correlation tables, and gated reconcile controls.
- **Versioned API:** `GET /api/v1/observability` — Returns structured JSON snapshot with `schema_version`, `providers`, `links`, and `metrics`.
- **Legacy Compatibility Alias:** `GET /api/observability` — Read-only compatibility alias for Phase 4.
## Gated Actions
- `observability_reconcile_incident` (`gitea_observability_reconcile_incident`): Triggers or previews dry-run issue reconciliation for a provider incident.
- `observability_link_issue` (`gitea_observability_link_issue`): Links a provider incident to an existing Gitea tracking issue.
Both actions require `operator` role and gate through `task_capability_map`. Execution fails closed in read-only MVP mode.
-56
View File
@@ -1,56 +0,0 @@
# Post-restart MCP reconciliation (#662)
After an MCP process restart, sessions, leases, capabilities, worktrees, and
interrupted mutations must be reconciled before operators claim a clean runtime.
This document describes the #662 completion-proof path.
## Components
| Piece | Where | Responsibility |
|-------|-------|----------------|
| `post_restart_reconcile.reconcile_after_restart` | `post_restart_reconcile.py` | Pure classification: inventory → completion proof DTO. No I/O. |
| `RestartCompletionProof` | `post_restart_reconcile.py` | Machine-readable proof (`.as_dict()` is JSON-serializable). |
| `gitea_reconcile_after_restart` | `gitea_mcp_server.py` | MCP tool: gathers inventory from the #613 control-plane DB + master-parity, classifies, returns the proof. Read-only. |
| Boot hook | `gitea_assess_master_parity` | First post-restart parity probe also runs reconcile once (log-only by default). |
## Dimensions
The assessor classifies:
- **service_health** — process healthy / parity mutation-safe
- **clients** — connected client descriptors (optional inventory)
- **sessions** — active session rows with dead owner pids are unresolved
- **checkpoints** — soft-depends on #660; skipped with reason when schema absent
- **leases** — live control-plane leases after restart
- **capabilities** — master-parity / stale-runtime (#610)
- **worktrees** — lease-bound paths missing on disk
- **interrupted_mutations** — mutating lease phases or explicit pending inventory; **never auto-resumed**
- **duplicates** — multiple live claims on the same work item
- **queue** — allocator resume safety
## Modes
| Mode | Env / arg | Behavior |
|------|-----------|----------|
| `log_only` (default) | unset or `GITEA_POST_RESTART_RECONCILE_MODE=log_only` | Proof only; `mutation_hold=false` |
| `enforce` | `GITEA_POST_RESTART_RECONCILE_MODE=enforce` or `mode=enforce` | Sets `mutation_hold=true` when overall status is degraded/failed or interrupted mutations remain |
## Follow-up issues
Unresolved dimensions produce `proposed_follow_ups` entries suitable for durable
Gitea issues. The MCP tool **does not create** those issues in v1 (rollout is
log-only first). Controllers may file them from the proof payload.
## Links
- Umbrella: #655
- Vision: #652 · Roadmap: #653
- Checkpoint schema: #660 (soft dependency)
- Drain proof: #661 (soft)
- This issue: #662
## Non-goals
- HA multi-instance failover
- Automatic silent mutation replay
- Implementing the #660 checkpoint schema itself
-230
View File
@@ -1,230 +0,0 @@
# Remote-MCP coupling inventory
Every place the Gitea MCP server depends on being a local, client-spawned, stdio-attached
process on the operator's machine.
- **Issue:** #930 (Remote-MCP 01), child 1 of epic #929.
- **Generated against commit:** `7bf4f1258451823a55b36d2157e74f8457165088` (`master`).
- **Anchors:** every `file:line` below resolves at the commit above and at the commit that
adds this document. This change adds one new file and edits no existing file, so no
existing line number shifts between the two.
- **Scope:** documentation only. No server behavior changes in this child.
## How to read an entry
| Field | Meaning |
| ----- | ------- |
| **Anchor** | `file:line` at the commit under review. |
| **Assumes today** | What the code takes for granted while running as a local stdio process. |
| **Observes remotely** | What the same code would actually see on a shared remote host. |
| **Class** | One of: *portable as written*, *needs a seam*, *needs a replacement*, *cannot be remote*. |
| **Owner** | Exactly one epic child (#931#939) responsible for the fix. |
Classification meanings:
- **portable as written** — the code is already transport-, host-, and principal-neutral; it
moves unchanged once its inputs are supplied by a remote-aware caller.
- **needs a seam** — the logic is correct but is wired to a hard-coded local source. It needs
an injection point, not new semantics.
- **needs a replacement** — the semantics themselves are local-only. A remote deployment
needs a differently-defined mechanism, not the same mechanism relocated.
- **cannot be remote** — the operation is inherently about the operator's own machine
(its process table, its keychain, its checkout). It must either stay local behind an
explicit boundary or be deleted from the remote surface.
---
## 1. Transport bind
The transport is bound literally, once, at process start, and the bound value is the root of
the mutation-authorization chain.
| ID | Anchor | Assumes today | Observes remotely | Class | Owner |
| -- | ------ | ------------- | ----------------- | ----- | ----- |
| T1 | `gitea_mcp_server.py:23750` | The single production bind call passes the literal `transport="stdio"` immediately before the server loop. | The literal is wrong for any non-stdio deployment; there is no parameter to change it. | needs a seam | #931 |
| T2 | `mcp_daemon_guard.py:45` | `_PRODUCTION_TRANSPORTS = frozenset({"stdio"})` is the closed allowlist of production transports. | A remote transport name is rejected by the allowlist before any other check runs. | needs a seam | #931 |
| T3 | `mcp_daemon_guard.py:174` | `bind_native_mcp_transport` raises `UnsanctionedRuntimeError` for any transport outside `_PRODUCTION_TRANSPORTS` (raise at `mcp_daemon_guard.py:187`). | The remote server fails to start rather than degrading; the failure is correct, but the allowlist is the only thing that must change. | needs a seam | #931 |
| T4 | `mcp_daemon_guard.py:328` | `is_native_mcp_transport()` asserts a process-local runtime record whose `pid` matches `os.getpid()` and whose phase is `transport_bound`. The predicate itself names no transport. | Unchanged semantics: one server process that bound one transport. It stays true on a remote host. | portable as written | #931 |
| T5 | `mcp_daemon_guard.py:349` | `is_production_native_mcp_transport()` adds only a `mode == production` check on top of T4. | Unchanged. | portable as written | #931 |
| T6 | `irrecoverable_provenance.py:497` | `assess_transport_for_auth_mint()` requires production native transport before minting non-forgeable recovery authorization (#709 F1). | The gate is transport-agnostic in form, but its guarantee — "an ordinary Python process cannot reach this" — is currently underwritten by the stdio bind. Under a remote transport the guarantee must be re-derived from the authenticated session, not from the bind. | needs a seam | #931 |
| T7 | `gitea_mcp_server.py:8375` | Consumer: refuses to proceed unless `assess_transport_for_auth_mint()` allows. | Unchanged given a corrected T6. | portable as written | #931 |
| T8 | `gitea_mcp_server.py:8624` | Second consumer of the same gate on the confirmation path. | Unchanged given a corrected T6. | portable as written | #931 |
| T9 | `mcp_server.py:4` | Module docstring asserts "Runs over stdio." as a property of the server. | The stated contract becomes false on the remote deployment and is load-bearing documentation for operators. | needs a replacement | #931 |
## 2. Launch provenance
Mutations fail closed unless the process can prove a client launched it with real stdio pipes
and `GITEA_CLIENT_MANAGED` provenance. Every proof in this section is a statement about the
local operating system.
| ID | Anchor | Assumes today | Observes remotely | Class | Owner |
| -- | ------ | ------------- | ----------------- | ----- | ----- |
| P1 | `gitea_mcp_server.py:14588` | `_is_client_managed_process()` derives provenance from `GITEA_CLIENT_MANAGED` / `GITEA_MCP_CLIENT_MANAGED` / `GITEA_SERVER_PROVENANCE` / `GITEA_FORCE_CLIENT_MANAGED` on this process's own environment. | A long-lived remote process has one environment for all callers, so a per-process env var can no longer say anything about the caller that issued a request. | needs a replacement | #934 |
| P2 | `gitea_mcp_server.py:14606` | Falls back to `sys.stdin.isatty()`: an active TTY on stdin means a human launched it from a terminal, so refuse. | A remote server has no meaningful stdin. The signal is absent, not merely different. | cannot be remote | #934 |
| P3 | `gitea_mcp_server.py:14618` | `_provenance_mutation_block()` emits `blocker_kind: "unsupported_manual_launch"` and a "reconnect the IDE/client-managed MCP namespace" remediation. | The block shape is reusable; its predicate and its remediation text are both stdio-specific. | needs a seam | #934 |
| P4 | `gitea_mcp_server.py:20599` | `_check_mcp_runtimes_diagnostics()` shells `ps -o pid,lstart,command -ax` and greps for `mcp_server.py` to find peer role servers. | On a shared host the process table lists unrelated tenants' processes, or none at all under a container. Peer discovery by `ps` has no remote meaning. | cannot be remote | #934 |
| P5 | `gitea_mcp_server.py:20702` | More than one process per `GITEA_MCP_PROFILE` in the local process table is reported as a duplicate-launch fault. | A remote endpoint is expected to serve many concurrent sessions per role. "Two processes for one role" becomes the normal case, so the check inverts from a safety net into a false wall. | cannot be remote | #934 |
| P6 | `gitea_mcp_server.py:20715` | Processes lacking client-managed provenance are ignored for runtime freshness and reported as manual launches. | Same defect as P5: correctness depends on enumerating local peers. | cannot be remote | #934 |
| P7 | `gitea_config.py:1172` | `RECOGNIZED_GITEA_ENV_KEYS` is the allowlist of `GITEA_*` env vars a legitimately launched server may carry; anything else is contamination. | Configuration on a remote host arrives from deployment tooling, not from a client-authored env block. The allowlist keeps working mechanically but stops proving anything about provenance. | needs a replacement | #934 |
| P8 | `gitea_mcp_server.py:20683` | The unsupported-env scan applies `RECOGNIZED_GITEA_ENV_KEYS` to *other* processes' environments harvested via `ps eww <pid>`. | Reading another process's environment is unavailable or prohibited across tenants, and is not exposed in this form outside macOS/BSD `ps`. | cannot be remote | #934 |
| P9 | `mcp_daemon_guard.py:126` | `mark_sanctioned_daemon()` requires the claiming stack frame's resolved absolute path to be the canonical `mcp_server.py` / `gitea_mcp_server.py` next to the guard module; basename spoofing is rejected. | Entrypoint-path identity still exists on a remote host, but it authenticates the *deployment*, not the *caller*. It must be kept and demoted from "authorizes mutations" to "authorizes the process". | needs a seam | #934 |
| P10 | `gitea_config.py:1233` | The client-config generator emits `"GITEA_CLIENT_MANAGED": "1"` into each generated MCP client entry, alongside `GITEA_MCP_CONFIG` / `GITEA_MCP_PROFILE`. | A remote endpoint is addressed by URL and credential, not by a spawn command with an env block. This generator produces the wrong artifact entirely. | needs a replacement | #938 |
| P11 | `mcp_namespace_health.py:232` | Namespace health classifies a namespace as `client_managed` or `manual_launch` from the reported env summary. | During dual-run, local and remote namespaces coexist and must both be classifiable; a two-valued local/manual axis cannot express "remote endpoint, authenticated session". | needs a replacement | #939 |
| P12 | `gitea_mcp_server.py:18161` | The diagnostics payload reports `server_provenance` as exactly `"client_managed"` or `"manual_launch"`. | This is the field a cutover operator reads to confirm which deployment served a call. It must gain a remote value before dual-run parity can be validated. | needs a replacement | #939 |
## 3. Role binding
Role separation is currently enforced by *which process a call reaches*. The process is pinned
to one role for its lifetime by an environment variable.
| ID | Anchor | Assumes today | Observes remotely | Class | Owner |
| -- | ------ | ------------- | ----------------- | ----- | ----- |
| R1 | `gitea_config.py:54` | `ENV_PROFILE = "GITEA_MCP_PROFILE"` is the single source of the active profile, read from the process environment. | One shared process serves several principals; a process-wide profile cannot answer "who is calling now". This is the root of the coupling. | needs a replacement | #932 |
| R2 | `review_workflow_load.py:95` | Reads `GITEA_MCP_PROFILE` directly to decide the reviewer workflow binding. | Reads the deployment's profile, not the caller's, silently granting or denying the wrong role. | needs a replacement | #932 |
| R3 | `mcp_discoverability.py:152` | Reads `GITEA_MCP_PROFILE` to describe the namespace to the client. | Correct logic, wrong input source; it needs the request principal injected. | needs a seam | #932 |
| R4 | `webui/deployment_boundary.py:115` | Reads `GITEA_MCP_PROFILE` to classify the deployment boundary for the console. | Same as R3. | needs a seam | #932 |
| R5 | `gitea_mcp_server.py:21106` | Remediation text instructs the operator to "Relaunch the server with `GITEA_MCP_PROFILE` set to a profile that has the required permission". | Relaunching a shared remote endpoint to change one caller's role is not a valid instruction; it would re-role every other session. | needs a replacement | #932 |
| R6 | `native_mcp_preference.py:93` | Detects shell commands that override `GITEA_MCP_PROFILE` away from the session (`native_mcp_preference.py:223`) and flags them as CLI auth divergence. | The divergence check is genuinely useful and survives, but its notion of "the session's profile" must come from the request principal. | needs a seam | #932 |
| R7 | `gitea_mcp_server.py:20671` | Recovers a peer server's role by regexing `GITEA_MCP_PROFILE=` out of that process's environment. | Depends on P4/P8 process-table access; role discovery by peer-env scraping has no remote analogue. | cannot be remote | #932 |
## 4. Credentials
Every token resolves, directly or indirectly, from one human's macOS keychain.
| ID | Anchor | Assumes today | Observes remotely | Class | Owner |
| -- | ------ | ------------- | ----------------- | ----- | ----- |
| C1 | `gitea_config.py:956` | `_keychain_token()` shells `security find-generic-password -s <item> -w`. | `security(1)` is a macOS binary reading the calling user's login keychain. It does not exist on a Linux host and would be the wrong identity even on a shared Mac. | cannot be remote | #933 |
| C2 | `gitea_config.py:974` | `resolve_token(profile, keychain_lookup=_keychain_token)` dispatches on `auth.type` of `env` or `keychain`, defaulting the lookup to C1. | The injectable `keychain_lookup` parameter is the existing seam; a remote credential provider plugs in here without changing the dispatch. | needs a seam | #933 |
| C3 | `gitea_config.py:1015` | `keychain_auth(item_id)` constructs the `{"type": "keychain", "id": ...}` reference stored in profiles. | The reference type itself encodes "macOS keychain" into persisted config. A remote provider needs a new auth reference type, not a new value of this one. | needs a replacement | #933 |
| C4 | `mcp_daemon_guard.py:440` | `assert_keychain_access_allowed()` fails closed for git-credential keychain fill outside a sanctioned daemon, with an operator opt-out env var. | The gate protects a mechanism that will not exist remotely. Its replacement must gate the *credential provider* call, not the keychain call, or the protection silently lapses. | needs a replacement | #933 |
| C5 | `sentry_incident_bridge.py:190` | `resolve_token(env)` resolves the Sentry token from an injected env mapping with no keychain path. | Already host-neutral; it is the shape the Gitea credential path should converge on. | portable as written | #933 |
| C6 | `gitea_mcp_server.py:18469` | The profile-audit tool calls `gitea_config.resolve_token(p)` for every configured profile to report "credentials present" without networking. | On a remote host this would materialize every principal's credential inside one process — an audit surface that becomes a credential-aggregation risk. | needs a seam | #933 |
## 5. Runtime freshness
The mutation gate is defined as "the commit this process started at matches the checkout on
this disk, and both match live master". Two of those three terms are local-disk facts.
| ID | Anchor | Assumes today | Observes remotely | Class | Owner |
| -- | ------ | ------------- | ----------------- | ----- | ----- |
| F1 | `master_parity_gate.py:168` | `capture_startup_parity(root)` reads git `HEAD` from the server's own root once at startup and returns it as the baseline. | A remote host carries a deployed artifact, not the operator's checkout. Its `HEAD` says nothing about the operator's working tree, which is the thing the gate exists to protect. | cannot be remote | #935 |
| F2 | `master_parity_gate.py:255` | `mutation_safe = determinable and in_parity and live_known and not live_stale` — a conjunction of two local-HEAD comparisons and one live-remote comparison. | Two of the three conjuncts lose meaning, so the whole verdict does. A remote deployment needs a redefined, testable freshness predicate rather than this one relocated. | needs a replacement | #935 |
| F3 | `master_parity_gate.py:164` | The live-remote head is probed and cached per `(root, remote, branch)`, keyed on the local root. | The live-remote probe is the one conjunct that survives; it needs a key that is not the operator's filesystem path. | needs a seam | #935 |
| F4 | `gitea_mcp_server.py:18262` | `gitea_assess_master_parity` publishes `startup_head` / `local_head` / `live_remote_head` / `mutation_safe` as the authoritative mutation-safety verdict. | The tool's contract is consumed by every mutation caller and by the operator; it must keep its shape while its semantics are redefined, or every consumer breaks at once. | needs a replacement | #935 |
| F5 | `gitea_mcp_server.py:23054` | Falls back to `_process_boot_head_sha` — the commit this process booted at — when the parity payload has no `startup_head`. | Same defect as F1, in a fallback path that is easy to miss when F1 is fixed. | needs a seam | #935 |
| F6 | `gitea_mcp_server.py:20615` | Staleness is also inferred from `os.path.getmtime()` of `gitea_mcp_server.py` under `PROJECT_ROOT` (`gitea_mcp_server.py:20611`), compared against peer process start times. | File mtime on a deployed artifact tracks the deploy, not the operator's edits, and the peer start times it is compared against come from the unavailable process table (P4). | cannot be remote | #935 |
## 6. Local filesystem
Author and reviewer tools act directly on the operator's checkout.
| ID | Anchor | Assumes today | Observes remotely | Class | Owner |
| -- | ------ | ------------- | ----------------- | ----- | ----- |
| L1 | `gitea_mcp_server.py:10122` | `gitea_bootstrap_author_issue_worktree` creates and binds a git worktree on the server's own disk. | The remote host has no operator checkout to add a worktree to. Executing this remotely would act on the wrong disk while reporting success. | cannot be remote | #936 |
| L2 | `gitea_mcp_server.py:190` | `ACTIVE_WORKTREE_ENV = "GITEA_ACTIVE_WORKTREE"` and `AUTHOR_WORKTREE_ENV` (`gitea_mcp_server.py:191`) carry the active workspace as process-wide environment. | Process-wide workspace state cannot represent per-session workspaces on a shared endpoint. | needs a replacement | #936 |
| L3 | `gitea_mcp_server.py:9801` | Binding a worktree writes `os.environ["GITEA_AUTHOR_WORKTREE"]` and `os.environ["GITEA_ACTIVE_WORKTREE"]` (`gitea_mcp_server.py:9802`), mutating global process state. | One session's bind would silently retarget every other concurrent session in the same process. This is a correctness bug the moment concurrency is real. | needs a replacement | #936 |
| L4 | `reviewer_inventory_worktree.py:48` | `_BRANCHES_WORKTREE_RE = re.compile(r"\bbranches/", re.I)` requires review worktree paths to sit under `branches/`. | A path convention on the operator's machine, asserted as a validation rule. It needs to become a property of a declared workspace, not a substring test. | needs a seam | #936 |
| L5 | `stable_control_runtime.py:54` | `DEV_WORKTREE_SEGMENT = "branches"` classifies a process root as a development worktree by path segment. | Same class of assumption as L4, on the runtime-classification side. | needs a seam | #936 |
| L6 | `mcp_server.py:42` | `check_conflict_markers()` runs at import and `os.walk`s the install directory for unresolved conflict markers, `sys.exit(1)` on a hit. | On a remote host it scans a deployed artifact, which by construction never has conflict markers — so the guard passes trivially and stops protecting the thing it was written to protect. | needs a replacement | #936 |
| L7 | `role_session_router.py:487` | `check_mid_merge()` reports infra-stop from `.git/MERGE_HEAD`, `rebase-merge`, `rebase-apply` and a source conflict scan under the server's project root. | Same inversion as L6: it would report the deployment's git state, not the operator's. | needs a replacement | #936 |
| L8 | `author_issue_bootstrap.py:996` | Enumerates worktrees with `git -C <root> worktree list --porcelain`. | Requires a real local clone with real worktrees; there is nothing equivalent to enumerate remotely. | cannot be remote | #936 |
| L9 | `mcp_server.py:10` | Redirects `sys.stderr` to the fixed path `/tmp/mcp_server_stderr.log` outside pytest. | A single fixed `/tmp` path is shared by every concurrent server on a host and is not a deployment's logging surface. | needs a replacement | #938 |
| L10 | `gitea_mcp_server.py:2314` | `ISSUE_LOCK_FILE = "/tmp/gitea_issue_lock.json"` — the legacy single global lock slot. | One global `/tmp` slot per host cannot represent concurrent remote sessions and is world-visible on a shared machine. | needs a replacement | #937 |
| L11 | `issue_lock_provenance.py:14` | `ISSUE_LOCK_FILE = os.environ.get("GITEA_ISSUE_LOCK_FILE", "/tmp/gitea_issue_lock.json")` keeps the same `/tmp` default in the provenance path. | Same as L10; the env override is a local escape hatch, not a remote design. | needs a replacement | #937 |
## 7. Durable state
Locks, leases, session state, and the control-plane database live in the operator's home
directory and are keyed on local PIDs.
| ID | Anchor | Assumes today | Observes remotely | Class | Owner |
| -- | ------ | ------------- | ----------------- | ----- | ----- |
| S1 | `issue_lock_store.py:26` | `DEFAULT_LOCK_DIR = ~/.cache/gitea-tools/issue-locks` — per-issue lock files under one user's home. | A shared endpoint has no single operator home; per-user paths make locks invisible across sessions and hosts. | needs a replacement | #937 |
| S2 | `issue_lock_store.py:83` | `session_pointer_path()` names the session pointer file `session-<os.getpid()>.json`. | Many sessions share one PID on a remote server, so the pointer collapses to a single slot and sessions overwrite each other. | cannot be remote | #937 |
| S3 | `issue_lock_store.py:98` | `is_process_alive(pid)` decides lock liveness by probing the local process table. | A PID recorded by one host is meaningless on another, and may coincidentally match a live unrelated process. | cannot be remote | #937 |
| S4 | `issue_lock_store.py:213` | Lock records stamp `session_pid` and `pid` from `os.getpid()`. | The recorded identity no longer distinguishes sessions; ownership checks silently pass for the wrong caller. | needs a replacement | #937 |
| S5 | `mcp_session_state.py:27` | `DEFAULT_STATE_DIR = ~/.cache/gitea-tools/session-state`, mode `0o700`. | Same home-directory coupling as S1, for review decision locks and workflow proofs. | needs a replacement | #937 |
| S6 | `mcp_session_state.py:559` | Session bodies stamp `session_pid` and `writer_pid` from `os.getpid()` (`mcp_session_state.py:560`). | Writer attribution collapses across concurrent sessions in one process. | needs a replacement | #937 |
| S7 | `control_plane_db.py:47` | `DEFAULT_DB_PATH = ~/.cache/gitea-tools/control-plane/control_plane.sqlite3`. | A per-user SQLite file is not reachable by, or safe for, multiple remote sessions or multiple hosts. | needs a replacement | #937 |
| S8 | `control_plane_db.py:386` | `sqlite3.connect(self.db_path, timeout=30)` — single-writer file locking tuned for one local process. | SQLite's write lock does not extend across hosts and degrades sharply under real concurrency; the store needs a concurrency-safe backend. | needs a replacement | #937 |
| S9 | `control_plane_db.py:1145` | Lease rows record `owner_pid` defaulting to `os.getpid()` (also `control_plane_db.py:2039`). | PID-keyed lease ownership is unusable across hosts and ambiguous within one shared process. | cannot be remote | #937 |
| S10 | `mcp_daemon_guard.py:53` | `_DEFAULT_SESSION_STATE_DIR` is pinned once at transport bind so a later `GITEA_MCP_SESSION_STATE_DIR` change cannot manufacture a second authority domain (#695 AC2). | The single-authority-domain invariant is exactly right and must be preserved; only its backing location needs to move. | needs a seam | #937 |
| S11 | `gitea_mcp_server.py:11875` | Reviewer-lease reclaim reads `owner_pid_alive` from the lease freshness record to decide whether an owner is dead. | Consumes S3/S9; a false "owner alive" or "owner dead" here reclaims or refuses a live lease. This is the highest-consequence consumer of PID liveness. | cannot be remote | #937 |
---
## Summary
### Entries per category
| Category | Entries |
| -------- | ------: |
| 1. Transport bind | 9 |
| 2. Launch provenance | 12 |
| 3. Role binding | 7 |
| 4. Credentials | 6 |
| 5. Runtime freshness | 6 |
| 6. Local filesystem | 11 |
| 7. Durable state | 11 |
| **Total** | **62** |
No category is empty, so no "this category has no coupling" justification is required.
### Entries per classification
| Classification | Entries |
| -------------- | ------: |
| portable as written | 5 |
| needs a seam | 16 |
| needs a replacement | 26 |
| cannot be remote | 15 |
| **Total** | **62** |
### Category × classification
| Category | portable | seam | replacement | cannot | Total |
| -------- | -------: | ---: | ----------: | -----: | ----: |
| 1. Transport bind | 4 | 4 | 1 | 0 | 9 |
| 2. Launch provenance | 0 | 2 | 5 | 5 | 12 |
| 3. Role binding | 0 | 3 | 3 | 1 | 7 |
| 4. Credentials | 1 | 2 | 2 | 1 | 6 |
| 5. Runtime freshness | 0 | 2 | 2 | 2 | 6 |
| 6. Local filesystem | 0 | 2 | 7 | 2 | 11 |
| 7. Durable state | 0 | 1 | 6 | 4 | 11 |
| **Total** | **5** | **16** | **26** | **15** | **62** |
### Entries per epic child
Every child from 2 through 10 is named by at least one entry, and every entry names exactly
one child.
| Child | Issue | Title | Entries | IDs |
| ----: | ----- | ----- | ------: | --- |
| 2 | #931 | Transport-neutral bind seam | 9 | T1T9 |
| 3 | #932 | Per-request principal resolution | 7 | R1R7 |
| 4 | #933 | Server-side credential provider | 6 | C1C6 |
| 5 | #934 | Remote-session provenance | 9 | P1P9 |
| 6 | #935 | Redefined master-parity gate | 6 | F1F6 |
| 7 | #936 | Local-filesystem vs remotable tool split | 8 | L1L8 |
| 8 | #937 | Concurrency-safe session, lock, and lease state | 13 | L10, L11, S1S11 |
| 9 | #938 | Authenticated remote MCP endpoint | 2 | P10, L9 |
| 10 | #939 | Dual-run cutover and rollback | 2 | P11, P12 |
| | | **Total** | **62** | |
## Notes for downstream children
- **The three highest-risk entries are P5, F2, and S11.** Each is a guard that does not
merely stop working remotely — it inverts. P5 turns concurrency into a reported fault,
F2 returns a verdict computed from terms that no longer mean anything, and S11 reclaims
or refuses leases on a PID-liveness answer that is wrong rather than unknown. A gate that
fails open while still reporting green is worse than one that fails to start.
- **T4, T5, T7, T8, and C5 are the portable core.** They show the target shape: predicates
over injected inputs, with no reference to the host, the process table, or the operator's
disk.
- **The keychain seam already exists** at C2 (`resolve_token`'s injectable `keychain_lookup`).
#933 should widen that seam rather than introduce a parallel path, and must remember C4 —
the guard protecting the old mechanism has to be re-pointed, or the protection lapses
silently when the mechanism is replaced.
- **`branches/` appears as a validation rule in at least two independent places** (L4, L5).
Path-substring conventions tend to have more copies than expected; #936 should re-grep
rather than trust this list to be exhaustive for that specific pattern.
-246
View File
@@ -1,246 +0,0 @@
{
"_comment": [
"Machine-checkable anchor table for docs/remote-mcp/threat-model.md (#956).",
"Every file:line anchor cited in the threat model must appear here, and the",
"source line at that anchor must contain the 'expect' substring.",
"tests/test_issue_956_threat_model.py enforces both directions, so a refactor",
"that shifts a line number fails the suite instead of silently rotting the",
"document. #930's inventory had no such guard and its gitea_mcp_server.py",
"anchors drifted between 7bf4f125 and aad5c8b4."
],
"generated_against_commit": "1dd30ecb1508b559868c2d5d94367bc055d5138e",
"anchors": [
{
"anchor": "gitea_mcp_server.py:25087",
"expect": "mcp_daemon_guard.bind_native_mcp_transport()"
},
{
"anchor": "mcp_daemon_guard.py:49",
"expect": "_PRODUCTION_TRANSPORTS = mcp_transport_config.SUPPORTED_TRANSPORTS"
},
{
"anchor": "mcp_daemon_guard.py:195",
"expect": "def bind_native_mcp_transport"
},
{
"anchor": "irrecoverable_provenance.py:497",
"expect": "def assess_transport_for_auth_mint"
},
{
"anchor": "gitea_mcp_server.py:9192",
"expect": "assess_transport_for_auth_mint()"
},
{
"anchor": "gitea_mcp_server.py:9441",
"expect": "assess_transport_for_auth_mint()"
},
{
"anchor": "mcp_server.py:4",
"expect": "The transport is selected by deployment configuration"
},
{
"anchor": "gitea_mcp_server.py:15630",
"expect": "def _is_client_managed_process"
},
{
"anchor": "gitea_mcp_server.py:15644",
"expect": "def _provenance_mutation_block"
},
{
"anchor": "gitea_mcp_server.py:15652",
"expect": "unsupported_manual_launch"
},
{
"anchor": "gitea_mcp_server.py:19217",
"expect": "server_provenance"
},
{
"anchor": "gitea_mcp_server.py:21741",
"expect": "def _check_mcp_runtimes_diagnostics"
},
{
"anchor": "gitea_mcp_server.py:21761",
"expect": "\"ps\", \"-o\", \"pid,lstart,command\""
},
{
"anchor": "gitea_mcp_server.py:21805",
"expect": "\"ps\", \"eww\""
},
{
"anchor": "gitea_config.py:1172",
"expect": "RECOGNIZED_GITEA_ENV_KEYS"
},
{
"anchor": "gitea_config.py:1233",
"expect": "GITEA_CLIENT_MANAGED"
},
{
"anchor": "gitea_config.py:54",
"expect": "ENV_PROFILE = \"GITEA_MCP_PROFILE\""
},
{
"anchor": "gitea_config.py:97",
"expect": "_REVIEW_MERGE_OPS"
},
{
"anchor": "gitea_config.py:499",
"expect": "repository authorization scope"
},
{
"anchor": "gitea_config.py:956",
"expect": "def _keychain_token"
},
{
"anchor": "gitea_config.py:974",
"expect": "def resolve_token"
},
{
"anchor": "gitea_config.py:1015",
"expect": "def keychain_auth"
},
{
"anchor": "gitea_config.py:294",
"expect": "def _validate_identity_auth"
},
{
"anchor": "mcp_daemon_guard.py:583",
"expect": "def assert_keychain_access_allowed"
},
{
"anchor": "gitea_mcp_server.py:19487",
"expect": "def gitea_list_profiles"
},
{
"anchor": "gitea_mcp_server.py:19538",
"expect": "gitea_config.resolve_token(p)"
},
{
"anchor": "gitea_mcp_server.py:19851",
"expect": "def gitea_audit_config"
},
{
"anchor": "gitea_mcp_server.py:19873",
"expect": "service_summaries(config)"
},
{
"anchor": "gitea_config.py:704",
"expect": "def resolve_service"
},
{
"anchor": "gitea_config.py:837",
"expect": "def service_summaries"
},
{
"anchor": "gitea_config.py:851",
"expect": "_keychain_token(auth.get(\"id\"))"
},
{
"anchor": "gitea_mcp_server.py:17909",
"expect": "\"jenkins-mcp\""
},
{
"anchor": "gitea_mcp_server.py:17915",
"expect": "external-mcp"
},
{
"anchor": "gitea_mcp_server.py:17936",
"expect": "\"glitchtip-mcp\""
},
{
"anchor": "gitea_mcp_server.py:17941",
"expect": "external-mcp"
},
{
"anchor": "mcp_discoverability.py:9",
"expect": "EXPECTED_JENKINS_TOOLS"
},
{
"anchor": "mcp_discoverability.py:17",
"expect": "EXPECTED_GLITCHTIP_TOOLS"
},
{
"anchor": "sentry_incident_bridge.py:36",
"expect": "SENTRY_AUTH_TOKEN"
},
{
"anchor": "sentry_incident_bridge.py:190",
"expect": "def resolve_token"
},
{
"anchor": "sentry_incident_bridge.py:289",
"expect": "Authorization"
},
{
"anchor": "sentry_observability.py:55",
"expect": "SENTRY_DSN"
},
{
"anchor": "master_parity_gate.py:168",
"expect": "def capture_startup_parity"
},
{
"anchor": "master_parity_gate.py:255",
"expect": "mutation_safe"
},
{
"anchor": "gitea_mcp_server.py:19331",
"expect": "def gitea_assess_master_parity"
},
{
"anchor": "gitea_mcp_server.py:193",
"expect": "ACTIVE_WORKTREE_ENV"
},
{
"anchor": "gitea_mcp_server.py:194",
"expect": "AUTHOR_WORKTREE_ENV"
},
{
"anchor": "gitea_mcp_server.py:2352",
"expect": "/tmp/gitea_issue_lock.json"
},
{
"anchor": "gitea_mcp_server.py:10957",
"expect": "def gitea_bootstrap_author_issue_worktree"
},
{
"anchor": "mcp_server.py:13",
"expect": "/tmp/mcp_server_stderr.log"
},
{
"anchor": "issue_lock_store.py:26",
"expect": "DEFAULT_LOCK_DIR"
},
{
"anchor": "issue_lock_store.py:83",
"expect": "def session_pointer_path"
},
{
"anchor": "issue_lock_store.py:98",
"expect": "def is_process_alive"
},
{
"anchor": "mcp_session_state.py:27",
"expect": "DEFAULT_STATE_DIR"
},
{
"anchor": "control_plane_db.py:47",
"expect": "DEFAULT_DB_PATH"
},
{
"anchor": "control_plane_db.py:380",
"expect": "mode=0o700"
},
{
"anchor": "control_plane_db.py:386",
"expect": "sqlite3.connect"
},
{
"anchor": "control_plane_db.py:1145",
"expect": "os.getpid()"
},
{
"anchor": "gitea_mcp_server.py:12871",
"expect": "owner_pid_alive"
}
]
}
-409
View File
@@ -1,409 +0,0 @@
# Remote-MCP threat model, trust boundaries, and service decomposition
What the adversary is, what each boundary protects, and which services may share a process.
- **Issue:** #956 (Remote-MCP threat model), child of epic #929, cross-linked to #955.
- **Depends on:** #930 (closed) — `docs/remote-mcp/coupling-inventory.md`.
- **Blocks:** #932, #933, #934, #938.
- **Generated against commit:** `1dd30ecb1508b559868c2d5d94367bc055d5138e` (#708's
namespace-attachment gate). Originally generated against
`aad5c8b42361d380a8eeb07b94b90815e594c2c5` (`master`), re-anchored at
`a143cd065ba06e1a2bdc5143a19ec156e53650ef` when #931's transport bind seam shifted the
cited lines, and re-anchored again when #708 shifted them further.
- **Scope:** documentation only. This child changes no server behavior. It adds one
document, one anchor fixture, and the test that enforces them.
## Relationship to #930
#930 asked *what breaks when the process stops being local*. This document asks *what an
attacker gets, and where we stop them*. The two are deliberately different axes: #930
classifies each coupling as portable, seam, replacement, or cannot-be-remote; this document
classifies each **credential** by blast radius and each **boundary** by what crossing it
requires. An entry can be perfectly portable and still be a trust disaster —
`gitea_config.py:851` is portable Python that reads a CI secret from inside the Gitea server.
### Anchors are enforced, not asserted
Every `file:line` in this document is declared in `docs/remote-mcp/threat-model-anchors.json`
with the substring that must appear at that line, and
`tests/test_issue_956_threat_model.py` fails if any anchor does not resolve or if the
document cites an anchor the fixture does not cover.
This guard exists because #930 did not have one. Its inventory was generated at
`7bf4f125`; by `aad5c8b4` its `gitea_mcp_server.py` anchors had drifted — the transport
bind it cited at line 23750 now lives at `gitea_mcp_server.py:25087`, and its
client-managed provenance anchor at 14588 now lands in an unrelated function. Nothing
failed, because nothing checked. Anchors into a ~24,700-line module rot silently, and a
security document that cannot prove its own citations is worse than none, because it is
trusted.
---
## 1. Assets
What an adversary wants. Ordered by consequence, not by likelihood.
| ID | Asset | Why it matters |
| -- | ----- | -------------- |
| A1 | Merge authority on `Scaled-Tech-Consulting/Gitea-Tools` | This repository *is* the control plane. Code merged here becomes the gate that authorizes every future mutation, so merge authority is self-amplifying: one merge can disable every other control in this document. |
| A2 | Write authority on the `mdcps` tenant | A second, unrelated organization reachable from the same configuration. Compromise here is a cross-organization incident, not an internal one. |
| A3 | The eight Gitea role credentials | Long-lived bearer tokens. Possession is authority; there is no second factor at the API. |
| A4 | Jenkins read access (`mdcps`, enabled) | Build logs routinely carry deployment topology, internal hostnames, and accidentally-echoed secrets. |
| A5 | Error-tracking read access (GlitchTip / Sentry) | Event payloads carry stack frames, request context, and production user data. |
| A6 | Coordination-state integrity | The locks, leases, and review-decision records that make "exactly one owner" true. Corrupting them needs no Gitea credential and produces duplicate or lost work. |
| A7 | The operator's checkout and worktrees | Unmerged code, branch state, and the filesystem the author tools write to. |
| A8 | The macOS login keychain | The meta-credential. Everything in A3, A4, and A5 resolves from it. |
| A9 | Separation of duty between review and merge | The property that no single actor both approves and lands a change. An *asset*, not a control, because it is what the controls exist to produce. |
| A10 | Audit and provenance records | Determine whether an incident is reconstructable. An attacker who can forge provenance makes an intrusion indistinguishable from normal work. |
## 2. Adversaries
| ID | Adversary | Capability assumed | Not assumed |
| -- | --------- | ------------------ | ----------- |
| ADV1 | **Compromised LLM client** | Full control of one MCP client. Issues arbitrary tool calls, in any order, with any arguments, at machine speed. Sees every tool result. | Cannot read the operator's disk except through tools; cannot execute arbitrary local code outside the tool surface. |
| ADV2 | **Prompt injection** via repository content | Controls text the model reads and treats as instruction — issue bodies, PR descriptions, review comments, commit messages, file contents. Reaches the model on any read of untrusted content. | Holds no credential and issues no call directly. Its entire power is causing an *authorized* client to act. |
| ADV3 | **Malicious tool arguments** | Supplies hostile values to any parameter — paths, branch names, session identifiers, worktree paths, issue numbers — including traversal, injection, and confusion between look-alike identifiers. | Cannot bypass a gate that actually validates its input. |
| ADV4 | **Network attacker** | Observes and modifies traffic between client, server, and Gitea. Attempts downgrade, replay, and endpoint impersonation. | Does not hold a valid credential at the start. |
| ADV5 | **Curious operator** | Legitimate local access to the workstation: process table, `/tmp`, home directory, keychain prompts. Not malicious, but not authorized for every role either. | Does not defeat the OS keychain's own access control without a prompt. |
ADV2 is the adversary this architecture most under-models. Every other adversary must first
obtain something. Prompt injection obtains nothing: it borrows authority the client already
holds and is indistinguishable at the tool boundary from legitimate work. Each boundary
below therefore states whether it constrains ADV2 at all — and most do not, because they
authenticate the *caller*, not the *intent*.
## 3. Trust boundaries
"Crossing requires today" is what the code actually enforces at
`1dd30ecb1508b559868c2d5d94367bc055d5138e`, not what the design intends.
| ID | Boundary | Protects | Crossing requires today | Crossing must require remotely |
| -- | -------- | -------- | ----------------------- | ------------------------------ |
| B1 | LLM client ↔ MCP server session | A1, A3, A10 — that a mutating session was established through the sanctioned client path | A single configured bind (`gitea_mcp_server.py:25087`) validated against one closed allowlist (`mcp_daemon_guard.py:49`, `mcp_daemon_guard.py:195`) — since #931 the identifier comes from deployment configuration and defaults to the local transport, so the boundary no longer rests on a literal, but it still rests on the *bind* rather than on an authenticated caller; client-managed provenance (`gitea_mcp_server.py:15630`) or a refusal (`gitea_mcp_server.py:15652`); production transport before recovery-authorization mint (`irrecoverable_provenance.py:497`, consumed at `gitea_mcp_server.py:9192` and `gitea_mcp_server.py:9441`) | An authenticated handshake issuing a server-side session identity bound to a principal, with the transport recorded in provenance. The physical proof (a pipe) must become a cryptographic one. |
| B2 | Role ↔ role | A9 — that author, reviewer, merger, and reconciler are distinct authorities | **The process boundary only.** The role is a property of the process, read once from `GITEA_MCP_PROFILE` (`gitea_config.py:54`). A caller gets author permissions by connecting to the author process. Review and merge are the operations singled out for extra care (`gitea_config.py:97`) | A per-request principal, so the role follows from the credential presented and cannot be selected by reaching a different endpoint. |
| B3 | MCP server ↔ credential store | A3, A8 — that only sanctioned code turns a profile into a token | `_keychain_token` shelling out to the login keychain (`gitea_config.py:956`), dispatched by `resolve_token` (`gitea_config.py:974`) with the reference type built at `gitea_config.py:1015`, gated by `assert_keychain_access_allowed` (`mcp_daemon_guard.py:583`). Inline secrets are rejected at config load (`gitea_config.py:294`) | A credential provider keyed by the *request* principal, returning only that principal's credential, with the source recorded and the value never returned. |
| B4 | MCP server ↔ Gitea | A1, A2 — that only authorized calls reach the forge | A bearer token over TLS. Server-side, nothing distinguishes one role's token from another beyond the account it belongs to | Unchanged at the forge; the endpoint in front of it must refuse unauthenticated and plaintext connections before tool dispatch. |
| B5 | MCP server ↔ caller's filesystem | A7 — that a tool acts on the *caller's* disk or refuses | Nothing. The server's disk *is* the caller's disk. Worktree bootstrap writes directly (`gitea_mcp_server.py:10957`); the active workspace is process-global (`gitea_mcp_server.py:193`, `gitea_mcp_server.py:194`) | An explicit per-tool classification, enforced at dispatch, refusing filesystem tools over a transport that cannot reach the caller's disk. A green verdict about the wrong disk is the failure to prevent. |
| B6 | MCP server ↔ coordination state | A6, A9 — mutual exclusion | Local files and a local SQLite database, with liveness judged from the local process table (`issue_lock_store.py:98`), keyed on paths under one user's home (`issue_lock_store.py:26`, `mcp_session_state.py:27`, `control_plane_db.py:47`) and on `os.getpid()` (`control_plane_db.py:1145`, `gitea_mcp_server.py:12871`). A legacy global slot still exists at `gitea_mcp_server.py:2352`, and the session-pointer file is named per PID (`issue_lock_store.py:83`) | One authority per ownership question, with liveness from session identity and expiry, and atomic acquire, renew, and release across hosts. |
| B7 | Gitea integration ↔ unrelated integrations | A4, A5 — that a Gitea compromise is not a CI and observability compromise | **Nothing.** See §5. The Gitea server reads Jenkins and GlitchTip secrets (`gitea_config.py:851`, reached from `gitea_config.py:837`) and holds the Sentry token (`sentry_incident_bridge.py:190`) | A hard process boundary. This is the boundary #956 exists to create. |
| B8 | Tenant ↔ tenant (`prgs` / `mdcps` / `local-lab`) | A2 — that one organization's compromise is not another's | Convention. One configuration declares all three contexts; `resolve_service` fails closed on a *disabled* context (`gitea_config.py:704`) but the credentials of enabled ones remain reachable in-process. A per-profile repository scope exists (`gitea_config.py:499`) | Separate deployments, or at minimum per-tenant credential scopes with no process able to resolve both. |
| B9 | Deployed code ↔ merged policy | A1, A10 — that the running server enforces the rules that were actually merged | Comparing this process's startup commit against this disk (`master_parity_gate.py:168`), conjoined into a single verdict (`master_parity_gate.py:255`) published by `gitea_mcp_server.py:19331` | Freshness defined against the deployed build identity, with an explicit fail-closed verdict when undeterminable. |
### What no boundary constrains
None of B1B9 constrains **ADV2**. Every one authenticates a caller or a process; prompt
injection supplies neither. An injected instruction that reaches an authorized author
session crosses B1, B2, B3, and B5 legitimately, because at each of those boundaries it *is*
the author. The only controls that bite ADV2 are those constraining what an authenticated
principal may do regardless of what it asks for — the per-role permission split (B2), the
repository scope at `gitea_config.py:499`, and separation of duty (A9). Sizing those
controls correctly matters more after the migration, not less, because a remote endpoint
raises the number of clients that can be injected into.
## 4. Data flows
Flows that cross a boundary. `==>` carries a credential; `-->` does not.
```
B1 B4
[LLM client] ====================> [MCP server] ========> [Gitea]
^ stdio pipe today | ^ (A1,A2)
| session identity | |
| after migration | |
| | | B3
untrusted repository content | +======> [macOS login keychain] (A8)
read back into the model (ADV2) | resolves A3, A4, A5
^ |
+----------------------------------+
|
B5 | B6
[operator checkout / worktrees] <--------+-------> [locks · leases · sqlite]
(A7) | (A6)
|
B7 <-- boundary does not exist today
|
+========================+========================+
| | |
[Jenkins] (A4) [GlitchTip] (A5) [Sentry] (A5)
external MCP server external MCP server in-process bridge
```
Two flows deserve attention because neither is obvious from the code:
1. **The keychain flow fans out.** B3 is drawn once but resolves credentials for *every*
configured profile and service, not only the active one. `gitea_list_profiles`
(`gitea_mcp_server.py:19487`) reports each profile's credential status by calling
`resolve_token` on it (`gitea_mcp_server.py:19538`), and `gitea_audit_config`
(`gitea_mcp_server.py:19851`) reports service credential status through
`service_summaries` (`gitea_mcp_server.py:19873`).
2. **The return path is a flow too.** Content read from Gitea travels back into the model
and is treated as instruction. This is the ADV2 edge, and it is the only edge in the
diagram with no authentication on it, because it is not a request.
## 5. Per-boundary credential inventory
**14 credentials in total.** Blast radius is stated as what the credential yields *on its
own*, assuming every gate not backed by the credential itself has been bypassed — because
an attacker holding a token calls the API, not our tools.
| ID | Credential | Holder | Boundary | Blast radius |
| -- | ---------- | ------ | -------- | ------------ |
| CR1 | `prgs-author` Gitea token — account `jcwalker3` | macOS keychain; resolved in-process (`gitea_config.py:974`) | B3 → B4 | Create branches, push, commit, open PRs, create/close/comment issues on the control-plane repo. Cannot approve or merge. The one credential whose identity is genuinely distinct. |
| CR2 | `prgs-reviewer` Gitea token — account `sysadmin` | macOS keychain | B3 → B4 | Approve and request changes. **Shares one Gitea account with CR3, CR4, CR5.** |
| CR3 | `prgs-merger` Gitea token — account `sysadmin` | macOS keychain | B3 → B4 | Merge to `master` — A1 in full. Same account as CR2. |
| CR4 | `prgs-reconciler` Gitea token — account `sysadmin` | macOS keychain | B3 → B4 | Close PRs, delete branches, irrecoverable decision-lock recovery. Same account as CR2. |
| CR5 | `prgs-controller` Gitea token — account `sysadmin` | macOS keychain | B3 → B4 | Same operation set as CR4. Same account as CR2. |
| CR6 | `mdcps-author` Gitea token — account `913443` | macOS keychain | B3 → B4, B8 | Author operations on a second organization. **Shares one account with CR7 and CR8.** |
| CR7 | `mdcps-reviewer` Gitea token — account `913443` | macOS keychain | B3 → B4, B8 | Approve and request changes on `mdcps`. Same account as CR6. |
| CR8 | `mdcps-merger` Gitea token — account `913443` | macOS keychain | B3 → B4, B8 | Merge on `mdcps` — A2 in full. Same account as CR6. |
| CR9 | MDCPS Jenkins read credential | macOS keychain, read from the Gitea server process (`gitea_config.py:851`) | B7 | Read CI jobs, builds, and logs (A4). Enabled today. |
| CR10 | MDCPS GlitchTip read credential | macOS keychain, read from the Gitea server process (`gitea_config.py:851`) | B7 | Read error events and their payloads (A5). Enabled today. |
| CR11 | `SENTRY_AUTH_TOKEN` | Process environment, read in-process (`sentry_incident_bridge.py:36`, `sentry_incident_bridge.py:190`), sent as a bearer header (`sentry_incident_bridge.py:289`) | B7 | Read and reconcile Sentry issues (A5). Not a keychain credential — an env var, so it is inherited by anything the process spawns. |
| CR12 | `SENTRY_DSN` | Process environment (`sentry_observability.py:55`) | B7 | Write events into the observability project. Low read value, real forgery value: an attacker can inject fabricated events into the record (A10). |
| CR13 | macOS login keychain access | The operator's login session; gated by `assert_keychain_access_allowed` (`mcp_daemon_guard.py:583`) | B3, ADV5 | **Every other credential in this table except CR11 and CR12.** This is the aggregation point. |
| CR14 | Coordination-store access (no secret) | Filesystem permissions — `control_plane_db.py:47`, created `0o700` (`control_plane_db.py:380`), opened with a local file lock (`control_plane_db.py:386`) | B6, ADV5 | Full read/write of locks, leases, and decision records (A6). **There is no credential here at all** — anything running as the operator can rewrite ownership. |
### Findings
**Finding 1 — Role separation is not credential separation.** Four `prgs` roles resolve to
one Gitea account (`sysadmin`): reviewer, merger, reconciler, and controller. A stolen
reviewer credential *is* a merger credential. A9 — separation of duty between approving and
landing — is therefore enforced entirely by which local process a call reaches (B2), and not
at all by the forge. It survives exactly as long as B2 does, and B2 is the boundary the
migration dissolves.
**Finding 2 — The `mdcps` tenant has no role separation at all.** Author, reviewer, and
merger all resolve to account `913443`. One credential can open a PR, approve it, and merge
it. The in-process self-review check compares the authenticated username against the PR
author and would refuse — but that check runs on our side of B4. It is not a property of
the credential, and an attacker holding the token does not call our tools.
**Finding 3 — Any one role process can resolve every other role's credential.** This is not
inferred; it is demonstrated by tool output. `gitea_list_profiles`
(`gitea_mcp_server.py:19487`) called from the **author** session reports
`identity_status: "credentials present"` for `prgs-merger`, `prgs-reviewer`,
`prgs-reconciler`, and every `mdcps` profile, because it calls `resolve_token` on each one
(`gitea_mcp_server.py:19538`). The author process does not merely *have access to* the
merger's credential — it reads it to answer a status query. B2 is not a credential boundary
in either direction.
**Finding 4 — The Gitea server reads CI and observability secrets.** `gitea_audit_config`
(`gitea_mcp_server.py:19851`) reports `MDCPS Jenkins: enabled, read-only, authenticated`.
That word `authenticated` is produced by `service_summaries` (`gitea_mcp_server.py:19873`,
defined at `gitea_config.py:837`), whose default check calls `_keychain_token` on the
service's own keychain reference (`gitea_config.py:851`). Producing that one line requires
the Gitea MCP server to read the Jenkins secret and the GlitchTip secret out of the
keychain. B7 does not exist.
**Finding 5 — Jenkins and GlitchTip are already decomposed; the reach is residual.** Their
tools live in separately registered servers, marked `external-mcp`
(`gitea_mcp_server.py:17909`, `gitea_mcp_server.py:17915`, `gitea_mcp_server.py:17936`,
`gitea_mcp_server.py:17941`) with their own expected tool sets (`mcp_discoverability.py:9`,
`mcp_discoverability.py:17`). The correct decomposition was already chosen. What remains is
a leak across it: the credential *references* still live in the Gitea configuration and are
still resolved by the Gitea process. #75 bundled these services into one control-plane
umbrella; the tools were separated afterwards, the credentials were not.
**Finding 6 — Sentry is the exception that is not decomposed.** Unlike Jenkins and
GlitchTip, the Sentry bridge runs *inside* the Gitea server, resolving its token from the
process environment (`sentry_incident_bridge.py:190`) and sending it as a bearer header
(`sentry_incident_bridge.py:289`). Being an environment variable rather than a keychain item
makes it strictly worse: it needs no keychain prompt and is inherited by every subprocess the
server spawns — including the `ps` invocations at `gitea_mcp_server.py:21761` and
`gitea_mcp_server.py:21805`, reached from `gitea_mcp_server.py:21741`.
**Finding 7 — The highest-value coordination asset has the weakest gate.** A6 is protected
by filesystem permissions alone (CR14). Corrupting a lease requires no Gitea credential,
produces no forge-side audit record, and breaks the mutual exclusion the entire workflow
assumes. Every other asset costs an attacker a credential; this one costs nothing beyond
local access, which is exactly ADV5's position.
**Finding 8 — Provenance authenticates the launch, not the caller.** `server_provenance` is
reported as exactly `client_managed` or `manual_launch` (`gitea_mcp_server.py:19217`),
derived from environment inspection (`gitea_mcp_server.py:15630`) with the recognized-key
allowlist at `gitea_config.py:1172` and the generator that emits the marker at
`gitea_config.py:1233`. Every one of those facts is fixed at process start. A client that is
trustworthy at launch and compromised a minute later remains `client_managed` for the life
of the process, and the transport contract that underwrites it is stated as a property of
the server itself (`mcp_server.py:4`). Since #931 that contract names the configured
transport rather than asserting stdio, but it is still fixed once, at bind, for the life of
the process.
## 6. Decomposition ruling
This section is the ruling #956 requires. It is a decision, not a recommendation.
**D1 — No unrelated co-residency.** A single integration process **must not** hold, resolve,
or be able to resolve credentials for services it does not itself integrate with.
Concretely: the Gitea MCP service may hold Gitea credentials and nothing else. Jenkins,
GlitchTip, Sentry, and any database credential are **not permitted** to co-reside with Gitea
credentials in one process.
*Rationale.* A process is the smallest unit an attacker takes whole. Once ADV1 or ADV2
controls execution in a process, every credential that process can resolve is theirs, and no
in-process check helps, because the checks are in the process too. Blast radius is therefore
a property of the process boundary and nothing finer. Findings 4 and 6 show that today one
compromise of the Gitea server yields CI read access, error-tracking read access, and — via
CR13 — every role credential on both tenants. That is the single largest reduction in blast
radius available anywhere in epic #929, and it costs no new mechanism: the decomposition
already exists (Finding 5) and is merely leaked across.
**D2 — Separation of duty must be backed by credentials.** Two roles whose separation is a
security property must not resolve to the same forge account. Specifically, reviewer and
merger must be distinct accounts. Today they are not, on either tenant (Findings 1 and 2).
*Rationale.* B2 is a process boundary, and the migration's entire purpose is to replace
process boundaries with request-level ones. A separation enforced only by which process a
call reaches does not survive that replacement — and it is already bypassable by anyone who
holds the token and calls the API instead of the tool.
**D3 — Credential resolution is scoped to the request principal.** A session must resolve its
own credential and must have no path to any other principal's. The resolve-every-profile
behavior behind `gitea_mcp_server.py:19538` and `gitea_mcp_server.py:19873` must report
configured-or-not from configuration alone, without resolving the secret.
*Rationale.* Finding 3. An audit surface that proves a credential exists by fetching it is a
credential-aggregation primitive wearing a diagnostic's clothes.
**D4 — Coordination state is a protected asset with its own authority.** Access to locks,
leases, and decision records must require an authenticated session, not merely local
filesystem access.
*Rationale.* Finding 7. #937 already moves this store for concurrency reasons; the
authorization requirement must land with it, or the store becomes remotely reachable while
still being authorized by nothing.
### Exceptions
**One, time-boxed.** During the dual-run window defined by #939, the **local** stdio fleet
may continue to resolve Jenkins and GlitchTip credential *references* from the shared
configuration, because removing them from the local configuration is not a prerequisite for
standing up the remote endpoint and would strand the operator's existing local workflow.
This exception is bounded by all of:
- It applies to the local stdio deployment only. The remote endpoint (#938) must be
configured with Gitea credentials and no others from its first day.
- It expires when #939 completes. It does not survive cutover.
- It does not extend to Sentry: CR11 and CR12 are process-environment credentials in the
Gitea server (Finding 6) and must be absent from the remote deployment's environment
regardless of dual-run state.
No exception is granted to D2, D3, or D4.
### Consequences for the target architecture
- The remote endpoint serves **Gitea only**. It is not a general control-plane endpoint.
- Jenkins and GlitchTip keep their existing separate servers, and their credential
references move out of the Gitea configuration.
- The Sentry bridge either moves behind its own service boundary or is absent from the
remote deployment. It does not travel with the Gitea server.
- Reviewer and merger accounts diverge before the endpoint is trusted for merges, or A9 is
recorded as unenforced.
## 7. Child-to-boundary mapping
Every #929 child from 2 through 10, mapped to the boundary it implements. A child
implementing more than one boundary names its primary first.
| Child | Issue | Boundaries | What it must establish | Rulings it must honor |
| ----: | ----- | ---------- | ---------------------- | --------------------- |
| 2 | #931 | B1, B9 | The bound transport becomes a validated value that provenance and freshness can both key on. Without it neither B1 nor B9 has an input. | — |
| 3 | #932 | B2 | The role becomes a property of the request, not the process — the boundary the migration otherwise deletes. | D2, D3 |
| 4 | #933 | B3, B7 | Credentials come from a provider keyed by principal. This is where D1 and D3 are either enforced or permanently lost. | D1, D3 |
| 5 | #934 | B1 | Session provenance replaces pipe-and-process-table proof with an authenticated session identity. | — |
| 6 | #935 | B9 | Freshness redefined against deployed build identity, with an explicit undeterminable verdict. | — |
| 7 | #936 | B5 | Every tool classified and the filesystem boundary enforced at dispatch, so a tool cannot return green about the wrong disk. | — |
| 8 | #937 | B6 | One authority per ownership question, with session-identity liveness and atomic transitions. | D4 |
| 9 | #938 | B4, B1, B8 | The endpoint: authentication, principal binding, transport security, and — critically — the deployed credential set. | D1, D2, D3 |
| 10 | #939 | B6 | Dual-run with exactly one coordination authority at every instant, and the rollback that proves the way back. | D1 exception expiry |
Boundary coverage: B1 (#931, #934, #938), B2 (#932), B3 (#933), B4 (#938), B5 (#936),
B6 (#937, #939), B7 (#933), B8 (#938), B9 (#931, #935).
B7 has exactly one owner, #933, and that is deliberate. B7 is not created by standing up an
endpoint; it is created by deciding which credentials a process may resolve, which is
precisely what the credential-provider child does.
## 8. Adversarial walkthrough: a fully compromised LLM client
The required negative case. ADV1 holds one MCP client and issues arbitrary tool calls. What
does it reach at each boundary?
### Before the migration
The client is attached to the local fleet over stdio.
| Boundary | What ADV1 reaches | Stopped by |
| -------- | ----------------- | ---------- |
| B1 | Everything the fleet serves. The client *is* the sanctioned launcher: it satisfies the client-managed check (`gitea_mcp_server.py:15630`) by construction, and provenance is never re-verified after launch (Finding 8). | Nothing. The guard authenticates the launch, not the caller. |
| B2 | All five roles — it is attached to all five namespaces. It can author a PR, approve it from the reviewer namespace, and merge it from the merger namespace. | Only the in-process self-review check, which compares `jcwalker3` (author) against `sysadmin` (reviewer) and **passes**, because Finding 1 made them different accounts while leaving reviewer and merger identical. A9 falls in one sequence of legitimate calls. |
| B3 | Every credential in CR1CR10 via CR13, with no additional prompt — the daemon is already sanctioned, so `assert_keychain_access_allowed` (`mcp_daemon_guard.py:583`) returns immediately. | Nothing. |
| B4 | A1 and A2 in full. | Branch protection at the forge, to the extent configured. |
| B5 | The operator's checkout and every worktree, through the author tools (`gitea_mcp_server.py:10957`), plus the shared stderr path at `mcp_server.py:13`. | Nothing; the server's disk is the target disk. |
| B6 | All coordination state — no credential required (CR14). It can forge lease ownership and clear decision locks. | Filesystem permissions, which it already satisfies. |
| B7 | Jenkins (A4) and GlitchTip (A5) secrets via Finding 4, and CR11/CR12 from its own environment. | Nothing. |
| B8 | Both tenants. | Nothing in-process; only the disabled-context check (`gitea_config.py:704`), which does not apply to enabled contexts. |
| B9 | Can it run stale code? Only by advancing the checkout, which parity detects (`master_parity_gate.py:255`). | This gate holds — the one boundary that behaves as designed against ADV1. |
**Result: total.** One compromised client yields A1 through A8 and A10. The only asset with
real resistance is A1 via branch protection, and the client holds the merger credential
anyway. Nine boundaries, one meaningful stop.
### After the migration
The same client authenticates to the remote endpoint with one role's credential, assuming
#931#939 land **and honor D1D4**.
| Boundary | What ADV1 reaches | Stopped by |
| -------- | ----------------- | ---------- |
| B1 | One authenticated session, bound to one principal. | #934: a forged or expired session identity is refused; the client cannot mint one. |
| B2 | **One role.** Presenting the author credential yields author permissions only. | #932: the principal comes from the credential, not from which endpoint was reached. |
| B3 | **One credential — its own.** | #933 with D3: the provider resolves by principal, and no diagnostic resolves the others. |
| B4 | That role's authority on the forge. | Endpoint authentication (#938); plaintext and unauthenticated attempts refused before dispatch. |
| B5 | **Nothing.** Filesystem tools are refused over the remote transport with a named blocker. | #936. |
| B6 | Its own leases; contention resolves to exactly one winner. | #937 with D4: authenticated session required, not filesystem access. |
| B7 | **Nothing.** No CI or observability credential exists in the process. | D1 — the single largest reduction on this table. |
| B8 | One tenant. | D1 and #938: the deployment carries one tenant's credentials. |
| B9 | Cannot induce stale enforcement. | #935: explicit fail-closed verdict, including undeterminable. |
**Result: bounded.** The compromise is contained to one role on one tenant, with no
filesystem reach and no lateral credential access. A9 survives *only if D2 lands* — if
reviewer and merger still share `sysadmin`, a compromised reviewer session still merges, and
this row reads the same after the migration as before it.
### What the migration does not fix
Against **ADV2**, both tables are identical. Prompt injection does not need to cross a
boundary: it arrives inside an authorized session and asks that session to do what it is
already permitted to do. Every "stopped by" above authenticates a principal, and the
injected instruction has the correct principal. The migration reduces ADV1's blast radius by
roughly an order of magnitude and reduces ADV2's by nothing.
The controls that do constrain ADV2 are per-principal permission scope (#932), repository
scope (`gitea_config.py:499`), and credential-backed separation of duty (D2) — each limiting
what an authenticated session may do *regardless of what it is asked for*. #955's
secure-isolation end state should be read with that distinction in mind: removing credentials
from clients defeats ADV1 and ADV5, and does not by itself defeat ADV2.
Two further items are explicitly out of scope here and unowned by #929:
- **Session-credential rotation and revocation.** #938 names rotation as documentation, but
no child owns proving that a revoked credential stops an in-flight session.
- **ADV3** (malicious tool arguments) is diffused across every child rather than owned. The
per-request principal work in #932 is the natural place to assert that identifiers taken
from the request never authorize anything on their own.
## 9. How to verify this document
1. `PYTHONPATH=. pytest tests/test_issue_956_threat_model.py` — resolves every anchor
against the working tree and checks the document's structural obligations.
2. Pick any five anchors at random and read them; the fixture states what each line must
contain.
3. Reproduce Findings 3 and 4 live: call `gitea_list_profiles` and `gitea_audit_config`
from the **author** namespace. Credential presence reported for roles other than the
active one is Finding 3; `MDCPS Jenkins: enabled, read-only, authenticated` is Finding 4.
If the anchor test fails after an unrelated refactor, the anchors moved and the fixture
needs regenerating — the claims are still true, but they are no longer traceable, which
#956 treats as the same defect.
-14
View File
@@ -46,17 +46,3 @@ If shell helpers are unavailable and MCP commit cannot run, stop with a recovery
report (restart session, clear hung terminals, use MCP-native commit). See
[`llm-workflow-runbooks.md`](llm-workflow-runbooks.md) § MCP-native commit path
(#260) and agent temp artifact cleanup (#261).
## 7. Process restart governance
Restarting the MCP control-plane process is destructive to concurrent multi-role
work and is governed by a dedicated policy. Restart is a **last resort** behind
narrower recoveries (reconnect, rebind), full/host restart is reserved to
operator/admin under **controller approval + automated safety gates**, a
unilateral LLM or operator full restart with active peers is **forbidden**, and
ambiguous policy state **denies** restart. Break-glass is a separate,
incident-backed path with a mandatory audit.
See [`architecture/mcp-restart-governance.md`](architecture/mcp-restart-governance.md)
(#656) for the authorization matrix, the recovery ladder, break-glass
conditions, and the `RG-01``RG-08` policy IDs.
-64
View File
@@ -1,64 +0,0 @@
# Sanctioned Recovery Playbooks & Controls (Phase 2 #644)
## Overview
Stale runtimes, worktree binding mismatches, and un-reconciled merged branches previously required expert manual shell recovery. Manual process kills (`pkill -f mcp_server.py`) are strictly forbidden and classified as runtime contamination ([#630](sanctioned-restart-controls.md)).
Phase 2 introduces **sanctioned recovery playbooks and controls** into the Web Console:
- **Diagnose**: Surface stale runtimes, worktree binding errors, contamination markers, and worktree anomalies via health & inventory APIs.
- **Preview**: Render mutation ledgers and exact confirmation phrases for recovery playbooks.
- **Confirm & Apply**: Execute sanctioned recovery actions through gated, audited paths.
- **Verify**: Revalidate control-plane state post-recovery before claiming clean status.
---
## Recovery Playbook Taxonomy
| Playbook ID | Action ID | Minimum Role | Target / Scope | Description |
|---|---|---|---|---|
| `clear_stale_binding` | `system.clear_stale_binding` | Operator | Active worktree binding | Clear provably missing or superseded `GITEA_ACTIVE_WORKTREE` binding ([#702](../stale_binding_recovery.py)). |
| `rebind_session_worktree` | `system.rebind_session_worktree` | Operator | Session worktree | Rebind or synchronize session worktree to verified lease worktree ([#864](../dirty_same_claimant_session_rebind.py)). |
| `reconcile_cleanups` | `system.reconcile_cleanups` | Controller | Worktree hygiene | Execute reconciler cleanup preview and apply for merged/superseded PR branches. |
| `sanctioned_restart` | `system.restart_namespace` | Admin | MCP Namespace | Restart MCP daemon gracefully via host supervisor ([#642](sanctioned-restart-controls.md)). |
---
## Wizard Workflow (Diagnose &rarr; Preview &rarr; Confirm &rarr; Verify)
### 1. Diagnose (`GET /api/v1/system/recovery/diagnose`)
Runs control-plane diagnostics:
- **Stale Runtime**: Mismatch between running daemon HEAD, local checkout HEAD, and remote-tracking HEAD.
- **Worktree Binding**: Missing path (`provably_stale_missing_path`), unverified inherited binding (`unverified_inherited`), or superseded binding (`superseded_by_session_lease`).
- **Contamination**: Checks for live contamination markers from unmanaged process kills.
- **Worktree Anomalies**: Scans `branches/` directory for un-reconciled cleanups or missing preserved worktrees.
Returns `RecoveryDiagnosis` with eligible playbooks.
### 2. Preview (`POST /api/v1/system/recovery/preview`)
Takes `playbook_id` and optional `target`/`params`.
Returns:
- **Mutation Ledger**: Step-by-step sequence of actions.
- **Confirmation Phrase**: Exact phrase required to authorize execution (e.g., `confirm clear_stale_binding`).
- **Authorization Decision**: RBAC check against the operator's principal.
### 3. Apply (`POST /api/v1/system/recovery/apply`)
Requires `playbook_id` and matching `confirmation` phrase. Gates run in this order, and each fails closed before anything is mutated:
1. **RBAC and execution phase** (`console_authz.authorize(..., for_execution=True)`). The phase branch only applies when `for_execution` is set. While `ACTIVE_PHASE` is `1`, every phase-2 recovery action is refused with `phase_not_active`, so no recovery playbook writes yet. Preview reports the same decision under `execution_authorization` / `execution_blocked_reason`.
2. **Confirmation phrase** (`confirmation_matches`).
3. **Contamination rules** ([#630](sanctioned-restart-controls.md)): the live marker is read from the session inventory and assessed under the gated task key `console_recovery_apply`. A contaminated runtime must be cleared through the reconciler cleanup playbook, which is the one playbook exempted from this gate because it is the designated remedy. The marker is also forwarded to `sanctioned_restart.execute_restart`, so a restart cannot launder a contaminated runtime.
Apply then executes the sanctioned recovery logic against the **live** process environment — not a copy — and records an audit entry in `console_audit`. A playbook that leaves the binding unchanged reports `performed: false`; `binding_before`, `binding_after`, and `binding_changed` are returned so a no-op cannot read as success.
Apply does **not** enforce master parity. Parity is reported by Diagnose ([#610](../master_parity_gate.py)) as evidence for the operator; it is not a precondition of this endpoint.
### 4. Verify (`POST /api/v1/system/recovery/verify`)
Re-evaluates control-plane diagnostics post-recovery and **reports** `clean`, `stale_runtime_clean`, `binding_clean`, `binding_classification`, and `contamination_clean`. It reports; it does not assert or block. State is read fresh rather than from the mapping a mutation just wrote. An `unverified_inherited` binding is reported as not clean, because unproven is not clean.
---
## Safety & Governance Principles
1. **No Manual `pkill`**: Direct process killing remains forbidden and is recorded as contamination.
2. **Auditability**: Every recovery preview and execution is logged in the console audit trail.
3. **Master Parity & Dual Control**: High-privilege recovery actions require controller/admin roles and explicit confirmation phrases.
-122
View File
@@ -1,122 +0,0 @@
# Sanctioned restart and graceful reload controls (#642)
Sessions used to recover MCP connectivity by killing the host daemon
(`pkill -f mcp_server.py`, #630). That is forbidden and stays forbidden: it
kills every namespace on the host, contaminates whichever session survives, and
leaves no audit trail. This document describes the sanctioned replacement,
implemented in `webui/sanctioned_restart.py`.
## What the console will and will not do
The console **never** restarts anything. It authorizes an intent, records it,
and hands off to a host supervisor. There is no code path in which the console
sends a signal, spawns a process, or renders a kill command — a regression test
asserts the module contains no `subprocess`, `signal`, `os.kill`, `os.system`,
or `popen` reference, and that no returned payload contains a kill command.
## Operations
| Mode | Action | Minimum role | Behaviour |
|------|--------|--------------|-----------|
| `reload` | `system.reload_namespace` | controller | Host supervisor reloads the namespace in place, draining in-flight requests. |
| `restart` | `system.restart_namespace` | admin | Host supervisor restarts the namespace. In-flight requests are lost. |
Scope is always exactly one namespace. A fleet-wide restart is an explicit
non-goal: `all`, `*`, `fleet`, and an empty scope are refused with
`fleet_scope_not_permitted`, because that is precisely the blast radius the
forbidden kill already had. An unrecognised namespace is refused rather than
passed through to the host.
## The gate sequence
`assess_restart_request()` applies every gate in order and reports the first
failure with a stable reason code:
| Order | Gate | Reason code on failure |
|-------|------|------------------------|
| 1 | Mode is `restart` or `reload` | `unknown_mode` |
| 2 | Scope is a single known namespace | `fleet_scope_not_permitted`, `unknown_namespace` |
| 3 | Principal holds the required console role | `unauthorized` |
| 4 | Confirmation phrase supplied | `confirmation_required` |
| 5 | Confirmation names this namespace and mode | `confirmation_mismatch` |
| 6 | Out-of-band operator authorization present | `operator_authorization_missing` |
| 7 | Runtime is not contaminated | `contaminated_runtime` |
| 8 | Host restart hook configured | `restart_hook_not_configured` |
Passing every gate yields `host_action_required`, never "restarted".
### Confirmation binds the namespace
The required phrase is `"<mode> <namespace>"` — for example
`restart gitea-author`. Binding the namespace into the phrase is the point: a
confirmation typed for one namespace cannot be replayed against another.
### Operator authorization is not self-assertable
Host daemon maintenance is authorized out of band through
`GITEA_OPERATOR_DAEMON_MAINTENANCE_AUTHORIZATION`, read from the process
environment and nowhere else (#630; #710 finding F1). A worker session cannot
set an environment variable for an already-running daemon, so this cannot be
faked the way a tool argument could.
### The host hook
`GITEA_SANCTIONED_RESTART_HOOK` holds an opaque reference the *host* resolves —
a supervisor label such as a launchd job name, never a command line. With no
hook configured the request is refused; the console does not fall back to a
process kill. The value is read server-side and never rendered to a client.
## Manual kill remains contamination
`classify_restart_command()` classifies an operator-proposed recovery command.
A manual `pkill`/`kill`/`killall` of the MCP daemon is contamination, not a
restart: it returns `clean_claim_allowed: false` and builds a durable
contamination marker (redacted command only, never secrets) naming
`system.restart_namespace` as the sanctioned alternative.
A live, uncleared contamination marker also blocks a restart. This is stricter
than #630's task-scoped gate, which deliberately lets a contaminated worker keep
commenting and handing off: restarting a contaminated runtime would launder the
contamination rather than resolve it. Clear the marker through the reconciler
path first.
## Post-restart health verification
After the host supervisor acts, `verify_post_restart_health()` decides whether
the session may claim to be clean:
| Status | Meaning | Clean claim |
|--------|---------|-------------|
| `clean` | Required tool callable, proven through the live client namespace | Allowed |
| `unproven` | Reported healthy without live client-namespace evidence | Refused |
| `unhealthy` | Probe failed | Refused |
Only `probe_source=client_namespace` evidence clears a session. Static tool
registration is not proof, and neither is an offline subprocess probe — an IDE
client can hold a registered tool list while live calls fail with
`client is closing: EOF` (see
[`mcp-namespace-health.md`](mcp-namespace-health.md)).
## Audit
Every attempt — allowed or denied — is recorded through
`webui.console_audit` with actor, target namespace, mode, result, and reason
code, and is redacted before it is persisted. `system.restart_namespace` is
break-glass, so its records are retained for 730 days. Records carry
`process_kill_executed: false`, which is a fact about the code path rather than
a claim: no such path exists.
## Environment variables
| Variable | Purpose |
|----------|---------|
| `GITEA_SANCTIONED_RESTART_HOOK` | Host supervisor reference; absent means restart is refused. |
| `GITEA_OPERATOR_DAEMON_MAINTENANCE_AUTHORIZATION` | Out-of-band operator authorization reference. |
| `WEBUI_AUDIT_LOG` | Console audit sink; absent means records are built but not persisted. |
## Non-goals
* No unrestricted `kill` from the UI, in any role, in any phase.
* No fleet-wide restart.
* No silent auto-restart loop: every attempt is confirmed and audited.
* This does not implement the Phase 1 health API (#634).
+11 -58
View File
@@ -91,15 +91,6 @@ already define, and a regression test asserts each mapping matches.
| `close_pr` | controller | privileged | `gitea.pr.close` | Yes | No | No | 3 |
| `merge_pr` | controller | privileged | `gitea.pr.merge` | Yes | **Yes** | **Yes** | 3 |
| `delete_branch` | admin | destructive | `gitea.branch.delete` | Yes | **Yes** | **Yes** | 3 |
| `record_analytics_usage` | operator | gated_write | `runtime.record_analytics_usage` | Yes | No | No | 2 |
| `system.reload_namespace` | controller | privileged | `runtime.reload_namespace` | Yes | No | No | 2 |
| `system.restart_namespace` | admin | destructive | `runtime.restart_namespace` | Yes | **Yes** | **Yes** | 2 |
| `system.clear_stale_binding` | operator | gated_write | `gitea.read` | Yes | No | No | 2 |
| `system.rebind_session_worktree` | operator | gated_write | `gitea.read` | Yes | No | No | 2 |
| `system.reconcile_cleanups` | controller | privileged | `gitea.pr.close` | Yes | No | No | 2 |
| `initiate_workflow` | operator | gated_write | `gitea.read` | Yes | No | No | 2 |
| `observability_reconcile_incident` | operator | gated_write | `gitea.read` | Yes | No | No | 4 |
| `observability_link_issue` | operator | gated_write | `gitea.read` | Yes | No | No | 4 |
**Dual control** means the acting principal may not be the sole authority: a
second distinct principal must confirm. **Break-glass** means the action is
@@ -111,19 +102,6 @@ honouring it.
`delete_branch` is admin-only rather than controller because it is the one
irreversible action in the set.
`system.restart_namespace` is admin-only for the same reason: restarting a
namespace drops every in-flight request on it. `system.reload_namespace` drains
first, so it is privileged but not destructive. Neither action is ever executed
by the console — both hand off to a host supervisor, and neither exposes a raw
process kill. See
[`sanctioned-restart-controls.md`](sanctioned-restart-controls.md) (#642).
`initiate_workflow` (#643) is operator-class because its outcome is a *claim*,
not a Gitea verdict. Requesting reviewer or merger work reserves that work
through the allocator; it does not grant the right to approve or merge, which
stays with the MCP role profile and its own capability gates. See
[`webui-requests.md`](webui-requests.md).
### Authorization decision
`authorize(action_id, principal, for_execution=False)` returns a decision
@@ -138,24 +116,9 @@ record and **denies by default**. The deny reasons are closed and enumerated:
| `phase_not_active` | Execution requested for an action whose phase is not open. |
| `allowed_preview_only` | Authorized — preview only, execution still disabled. |
There is no implicit allow branch.
`execution_enabled` on the decision reports whether the action has a live
execution path at all, and is computed by `execution_wired(action)`. There are
exactly two ways to be wired:
1. the action's `phase` is at or below `ACTIVE_PHASE`; or
2. the action declares an `execution_env_flag` **and** that variable is set.
Every action that declares no flag therefore reports `execution_enabled: false`
while the console is in Phase 1, so no caller can read an allow as permission
to mutate. The per-action flag exists because raising `ACTIVE_PHASE` would
enable execution for every action of that phase at once, including ones whose
execution path is not implemented. One implemented action goes live on its own
flag instead of dragging its unimplemented phase-mates with it.
`initiate_workflow` is the only action that currently declares a flag
(`WEBUI_REQUESTS_EXECUTION`), and it stays denied until an operator sets it.
There is no implicit allow branch. Even the allow result reports
`execution_enabled: false` while the console is in Phase 1, so no caller can
read an allow as permission to mutate.
## Secret redaction
@@ -236,8 +199,8 @@ breaking the request it describes.
| Class | Applies to | Default |
|-------|-----------|---------|
| `standard` | Routine gated writes | 90 days |
| `privileged` | `review_pr`, `close_pr`, `system.reload_namespace`, and any unclassifiable action | 365 days |
| `break_glass` | `merge_pr`, `delete_branch`, `system.restart_namespace` | 730 days |
| `privileged` | `review_pr`, `close_pr`, and any unclassifiable action | 365 days |
| `break_glass` | `merge_pr`, `delete_branch` | 730 days |
Each record carries its own class, day count, and computed `expires_at`, so
retention is auditable per record rather than inferred from file age. An
@@ -262,22 +225,13 @@ second one. The integration points are already wired and observable:
instead of adding a parallel check.
- **`GET /api/console/security-model`** publishes the RBAC matrix, redaction
policy, and audit policy as JSON for operators and tests.
- **`POST /api/v1/requests/preview` and `.../apply`** (#643) are the first
actions to use this model for a real execution path. Preview always returns a
decision and an audited `previewed` record; apply requires `confirm=true`,
emits `succeeded` or `denied`, and reserves work only through the allocator.
See [`webui-requests.md`](webui-requests.md).
A Phase 2 action must: use `execution_wired` rather than a private enable flag,
implement the confirmation and dual-control flow the matrix already declares,
emit a `succeeded` or `failed` record alongside the `gitea_audit` mutation
record, and keep `viewer` unable to reach any of it. Turning on execution
without the confirmation flow contradicts a declared requirement and is a
review failure, not a shortcut.
Raising `ACTIVE_PHASE` remains the way to open a whole phase at once, and is
deliberately *not* what #643 did: an action-scoped opt-in cannot enable an
action whose execution path nobody wrote.
To open Phase 2, a child issue must: raise `ACTIVE_PHASE`, implement the
confirmation and dual-control flow the matrix already declares, emit a
`succeeded` or `failed` record alongside the `gitea_audit` mutation record, and
keep `viewer` unable to reach any of it. Turning on execution without the
confirmation flow contradicts a declared requirement and is a review failure,
not a shortcut.
## Local-dev mode
@@ -330,7 +284,6 @@ Until Phase 2 wires it, probe protection rests on network placement alone, as
| `WEBUI_ROLE_MAP` | unset | JSON subject → role map |
| `WEBUI_REQUIRE_PROBE_AUTH` | unset | Require auth for non-public probes |
| `WEBUI_CONSOLE_AUDIT_LOG` | unset | Append-only audit sink path |
| `WEBUI_REQUESTS_EXECUTION` | unset | Opt in to `initiate_workflow` execution (#643) |
All are read server-side only. None is ever rendered into a page or returned by
an API.
-9
View File
@@ -55,15 +55,6 @@ shipped to the browser.
assumption paths, and the client-secret policy. Use it to verify an instance is
configured for internal-only operation.
## Process restart / reload disposition
The console never exposes a restart or reload control; process restart of the
MCP control-plane runtime is governed separately. Restart is a last resort behind
reconnect/rebind, full restart is operator/admin-only under controller approval
plus safety gates, and break-glass is an incident-backed path. See
[`architecture/mcp-restart-governance.md`](architecture/mcp-restart-governance.md)
(#656).
## Non-goals (MVP)
- Full SSO or session login in the UI
+1 -419
View File
@@ -52,13 +52,9 @@ status, onboarding checklist state, and the fail-closed error payloads (#635).
| Path | Description |
|------|-------------|
| `/` | Home / operator overview |
| `/health` | JSON liveness (`status`, `service`, `mode`, `timestamp`, `uptime_seconds`) |
| `/api/v1/system/health` | Structured read-only system health (#634) |
| `/system-health` | System-health dashboard — readiness, version/uptime, dependencies, MCP namespaces, stale-runtime parity (#639) |
| `/health` | JSON liveness (`status`, `service`, `mode`, `timestamp`) |
| `/queue` | Live PR and issue queue dashboard (#429) |
| `/api/queue` | JSON queue export with pagination metadata |
| `/traffic` | Workflow traffic-control view — runnable, leased, blocked, needs-controller, terminal-complete (#640) |
| `/api/traffic` | JSON traffic-control export with state classifications and next safe role actions |
| `/projects` | Project registry list with status and onboarding progress (#427, #635) |
| `/projects/{id}` | Project detail + onboarding checklist |
| `/api/v1/projects` | Versioned JSON registry export (#635) |
@@ -77,128 +73,11 @@ status, onboarding checklist state, and the fail-closed error payloads (#635).
| `/api/actions/{id}/preview` | Mutation ledger preview (GET, read-only) |
| `/leases` | Lease and collision visibility (#433) |
| `/api/leases` | JSON lease/collision export |
| `/sessions` | Runtime and session view (#641) — health + inventory sessions/namespaces/worktrees |
| `/api/sessions` | JSON export for the runtime/session view |
| `/api/v1/sessions` | Versioned alias of `/api/sessions` |
| `/gitea` | Gitea issue↔PR linkage console (#645) — both directions, with the evidence for each edge |
| `/api/v1/gitea/linkage` | JSON linkage export; `502` when the read could not be answered |
| `/inventory` | Phase 1 shell stub — unified inventory (backed by #636) |
| `/timeline` | Phase 1 shell stub — workflow event timeline |
| `/policy` | Phase 1 shell stub — capability/role policy placeholder |
| `/insights` | Phase 1 shell stub — operational insights placeholder |
Most routes are GET-only. POST/PUT/PATCH/DELETE return `405` with
`read-only-mvp`, except `/audit` and `/api/audit` which accept POST for
local validator preview only (no Gitea mutations, no server-side storage).
### Traffic-control state vocabulary (#640)
The traffic view classifies each open issue/PR into exactly one bucket:
| Bucket | Meaning | Operator implication |
|--------|---------|----------------------|
| **runnable** | No active lease, no block reason, safe for its expected role | Next role may start work |
| **leased** | Active author claim or reviewer PR lease | Do not stomp; wait or adopt via role tools |
| **blocked** | Dependency, missing head pin, conflict, or unmet dependency | Author remediation first |
| **needs_controller** | Contaminated, controller-only diagnosis, or `status:blocked` | Controller only |
| **terminal_complete** | Reconciler / terminal-lock territory | Reconciler cleanup path |
`status:blocked` items route to **needs_controller**, not **blocked**:
`expected_role_for_candidate` sends them to the controller, and the blocker
reason renders in either bucket.
**Live path contracts (do not invent):**
- PR head pins come from `QueueItem.signals["head_sha"]` (full SHA). Display
`extra["head_sha"]` is truncated and must never be used for routing.
- Reviewer leases are keyed as `(pr, pr_number)` only — never via a linked
`issue_number` on the same lease marker.
- Issue claims come from `claim_inventory["entries"]`
(`issue_claim_heartbeat.build_claim_inventory`). There is no `active_claims`
key.
- Queue display badges are only: `blocked`, `claimed`, `duplicate`, `stale`,
`in-review`, `open`. Review verdicts (`request-changes`, `approved`) are
**not** queue badges; traffic does not invent them from the queue loader.
## System health API (#634)
`GET /api/v1/system/health` is the structured, read-only health surface for
automated readiness checks. It is the first console API under the `/api/v1`
prefix; the unversioned MVP exports remain as compatibility aliases.
`/health` is unchanged for existing consumers — every MVP key is still present
— and now also carries `started_at`, `uptime_seconds`, and a
`system_health_api` pointer. It stays deliberately cheap and runs no dependency
probe, because answering readiness costs real work.
**Status codes.** `200` when ready, `503` when a required dependency failed or
was never probed. Automation can branch on the code without parsing the body.
**Query flags.** The Gitea check is a network call, so it is opt-in:
`GET /api/v1/system/health?deep=1` runs it and caches the result for
`WEBUI_HEALTH_PROBE_TTL_SECONDS` (default 15s) so dashboard polling does not
amplify into remote load. Without the flag that probe reports `skipped`.
**Dependencies.** `control_plane_db` and `repository` are required and drive
readiness. `gitea` is optional: when it fails the overall `status` degrades but
`readiness.ready` stays true, because local inventory is still serveable. Each
entry carries `status`, `detail`, `required`, and `latency_ms`.
Two honesty rules are worth knowing before reading the payload:
* `stale_runtime.mutation_safe` is true only when the runtime, checkout, and
remote-tracking commits are all known and equal. An unfetched remote is
reported as indeterminate, never as safe.
* `mcp_namespaces` entries are always `unproven`. A web process runs outside
the IDE-managed MCP client and cannot invoke a namespace tool, so per #543
only a `client_namespace` probe can prove that path.
Sample response (abridged, healthy):
```json
{
"status": "ok",
"service": "mcp-control-plane-webui",
"mode": "read-only",
"api": "/api/v1/system/health",
"timestamp": "2026-07-22T11:04:18.512034+00:00",
"readiness": { "ready": true, "complete": true, "reasons": [] },
"version": {
"git_sha": "620ed6e9a9550b8da2ceb82d9ab8744e8920490f",
"git_describe": "v1.1.0-898-g620ed6e",
"control_plane_schema_version": 4,
"python_version": "3.14.5",
"known": true
},
"process": { "started_at": "2026-07-22T10:58:02.114+00:00", "uptime_seconds": 376.4 },
"deep_probes_requested": false,
"dependencies": [
{
"name": "control_plane_db",
"kind": "sqlite",
"status": "ok",
"detail": "schema v4 readable",
"required": true,
"healthy": true,
"latency_ms": 1.482,
"metadata": { "schema_version": 4, "active_leases": 3 }
},
{ "name": "repository", "kind": "git", "status": "ok", "required": true, "healthy": true },
{ "name": "gitea", "kind": "http", "status": "skipped", "required": false, "healthy": false }
],
"mcp_namespaces": [
{ "namespace": "gitea-author", "required_tool": "gitea_whoami", "status": "unproven" }
],
"stale_runtime": { "stale": false, "determinable": true, "mutation_safe": true, "reasons": [] },
"probe_errors": []
}
```
No restart, reload, or process-kill control is exposed here: those are Phase 2
at the earliest, and #630 forbids process-kill recovery outright. Every probe
opens its subject read-only — the control-plane database is opened through a
`mode=ro` URI so a health check can never create or migrate a schema.
## Report audit (#431)
Paste an LLM final report at `/audit` or POST JSON to `/api/audit`. The UI
@@ -274,154 +153,6 @@ health, workflow/schema SHA-256 hashes, and stale-runtime warnings when the
checkout is behind merged safety-gate changes. Restart guidance links to #420;
no tokens or MCP restart actions are exposed.
## Application shell — Phase 1 (#638)
The console shell (`webui/layout.py`) renders a grouped navigation driven by a
single nav-config module, `webui/nav.py`. Nav groups follow the epic #631
Phase 1 information architecture: **Health, Traffic, Runtime/Sessions,
Projects, Inventory, Timeline, Policy** (placeholder), and **Insights**
(placeholder). Live views and Phase 1 placeholders (`stub`) are declared in one
place so the layout and the route table cannot drift.
The header carries two read-only status badges — an **environment** badge
(`local` for loopback binds, `remote` otherwise, derived from `WEBUI_HOST`) and
a **mode: read-only** badge — plus a **Docs** link to this document. No
privileged action controls are present in the Phase 1 shell.
Not-yet-implemented surfaces (`/inventory`, `/timeline`, `/policy`,
`/insights`) resolve to graceful read-only stub pages instead of 404s; their
backing views land in later child issues of #631 (the inventory surfaces are
backed by #636). Mutating methods on stub routes still fail closed with
`read-only-mvp`.
### Runtime and sessions (#641)
`/sessions` is a live Phase 1 read-only view that composes:
* runtime health from `#430` (profile, role, identity, master parity, stale warning)
* control-plane sessions / leases and filesystem locks / worktrees / namespaces from `#636`
* durable contamination markers when detectable (`#630` runtime recovery, `#671` stable-branch push)
It surfaces stale indicators (dead PID, expired lease) and never silences an
active contamination marker. Recovery links point only at sanctioned
reconnect/operator restart docs (`docs/mcp-namespace-eof-recovery.md`,
`docs/mcp-namespace-health.md`, `docs/mcp-restart-path-inventory.md`, this
document). The page does **not** restart, kill, or take over sessions; manual
`pkill` of MCP daemons is contamination, not recovery.
Honesty rules specific to this view:
* **Ownership columns never assert absence they cannot prove.** When the
`leases` or `locks` section is degraded or unavailable, the Leases and
Worktree-binding cells render `unknown (inventory <status>)` with an
*authority unproven* badge instead of `none` / `unbound`, and a caveat names
the unreadable sections. A worktree binding is correlated through lease work
numbers, so it is unproven when *either* section fails to read.
`/api/sessions` carries the same facts as `ownership_authority_complete`,
`ownership_section_status`, and per-row `lease_authority` /
`worktree_authority`, so a JSON consumer can tell "holds none" from "could
not be read".
* **Contamination text is redacted at the display boundary.** Marker payloads
(`command_summary`, `reason_class`, `session_id`, `role`) are
operator-supplied free text that does not arrive through inventory scrubbing,
so they pass through `webui.inventory.scrub_text`, which collapses `$HOME` and
redacts credential-shaped tokens and URL userinfo *anywhere* in the string.
The write-time redactor is a narrow denylist and is not relied on. The field
itself is kept — it is the `#630` evidence naming which daemon was killed.
## Gitea issue/PR linkage (#645)
`/gitea` is the Phase 3 read-only linkage console: which PR carries which issue,
which issues are claimed by more than one PR, and what the latest Canonical
Thread Handoff on a thread said. Gitea remains the source of truth — this
surface reads it and never writes to it. There is no issue/PR editor, no review,
and no merge control.
Query parameters (all optional):
| Parameter | Meaning |
|-----------|---------|
| `project` | Registry project id to scope the read (default: first registry entry) |
| `state` | `open` (default) or `all`; `all` widens the window to merged/closed items, where a landed edge lives |
| `issue=N` / `pr=N` | Focus one thread and load *its* latest canonical handoff |
`GET /api/v1/gitea/linkage` returns the same model as JSON
(`schema_version: 1`). It answers `502` when the read could not be answered, so
an automated consumer cannot mistake a fail-closed payload for "no links exist".
The HTML page always answers `200` and renders the reason instead — an operator
view must show why a read failed rather than withhold the page.
### How an edge is found
Each edge carries the evidence that produced it, strongest first:
| Evidence | Meaning |
|----------|---------|
| `closes_keyword` | The PR title or body declares `closes/fixes/resolves #N`. Gitea itself acts on this keyword. |
| `branch_marker` | The PR head branch carries the canonical `(fix\|feat\|docs\|chore)/issue-N-…` marker minted by the issue lock. |
| `body_reference` | The PR body mentions `#N` with no closing keyword. A mention is not a claim to close. |
Only closing and branch-marker edges populate the **issue → PR** direction: a
bare mention is a cross-link, and counting it as ownership would invent
contested issues out of ordinary references. The mention stays visible on the
**PR → issue** side, labelled as such. A PR whose two strongest edges tie is
flagged `ambiguous`; an issue claimed by two PRs is flagged `contested`.
### Honesty rules specific to this view
* **A partial read never reads as an absence.** Linkage is a claim about the
loaded window only. When pagination did not complete, every empty edge cell
renders `none found (partial inventory)` rather than `none`, and the JSON
carries `inventory_complete: false` plus per-row `links_authoritative: false`.
* **A failed read renders no table at all.** Missing credentials, an unknown
project, or a fetch error produce `ok: false` with a reason. An empty linkage
table would assert that no issue is linked to any PR, which such a read is not
in a position to claim.
* **Handoffs are loaded, never assumed.** CTH comments are thread-scoped, so
only the focused issue or PR has its comments fetched. Every other row reports
`not_loaded` with the reason; a thread whose comments *were* loaded and carried
no CTH says exactly that. A comment-source failure degrades the handoff alone —
the linkage tables still render.
* **Unrecognised handoff headings are reported, not republished.** A `## CTH:`
heading outside `CTH_TYPES` renders as `unrecognized`.
* **Redaction precedes display.** Titles, labels, handoff fields, and error
reasons pass through `webui.console_redaction` before serialization, and the
page HTML-escapes everything it renders.
* **Deep links are opt-in.** A link out to the Gitea web UI appears only when
`GITEA_MCP_REVEAL_ENDPOINTS=1` is set server-side, matching how the MCP tools
gate URL exposure. Item numbers stay usable without it.
## System-health dashboard (#639)
`/system-health` renders the same snapshot the `/api/v1/system/health` API
returns, so the page and the API can never disagree. Cards: overall readiness,
stale-runtime parity, version and uptime, dependency probes, MCP namespaces,
probe errors (only when present), and recovery pointers. `?deep=1` opts into
the network probe exactly as the API does; the plain page load stays cheap.
Field authority and honesty rules:
* `ready` and `readiness_complete` are shown separately. A snapshot whose
required probes never ran is not the same as one that ran them and passed,
and the page never collapses the two into an unproven green.
* A probe that did not run appears under **Not probed**, never as healthy.
* `stale_runtime.mutation_safe` is displayed verbatim from the API. When the
runtime is stale, or when parity is indeterminate, the page warns and does
not claim mutation safety.
* MCP namespaces are reported `unproven`: the web process runs outside the
IDE-managed MCP client and cannot prove that path (#543).
Redaction is split by field kind. Free text — probe details, readiness and
parity reasons, probe errors — passes through `system_health.redact`.
Structured fields — commit SHAs, probe names, statuses, timestamps — are
HTML-escaped only, because `redact`'s opaque-token rule matches any run of 32
or more characters and would otherwise blank every 40-character git SHA, which
is precisely the evidence the parity view exists to show.
The dashboard is read-only: no restart, reload, or process-kill control. Those
arrive in Phase 2 (#642). Recovery guidance points at the sanctioned client
reconnect / operator restart path — never a manual daemon kill (#630).
## Deployment boundary (#435)
MVP serves on loopback by default. Binding `0.0.0.0` or `::` is **refused**
@@ -481,155 +212,6 @@ health, workflow/schema SHA-256 hashes, and stale-runtime warnings when the
checkout is behind merged safety-gate changes. Restart guidance links to #420;
no tokens or MCP restart actions are exposed.
## Inventory API (#636)
`GET /api/v1/inventory` returns one versioned, read-only snapshot that unifies
what the lease (#433), worktree (#432), and runtime (#430) MVP views each show
separately, so traffic-control and recovery consumers read the same source.
`GET /api/v1/inventory/{section}` returns a single section under the identical
schema (`sessions`, `leases`, `locks`, `worktrees`, `namespaces`); an unknown
section is a `404` with `error: unknown_section`. Both routes are `GET`-only.
Each section carries its own `status` (`ok` / `degraded` / `unavailable`), a
`reason` when not `ok`, and a `scan_ms`. A subsystem that cannot be read
degrades to a reasoned section; it never raises and never emits an empty list
that would read as "nothing is there".
### Field authority
Every section names where its rows came from; authorities are never blended.
| Section | Authority | Source |
|---|---|---|
| `sessions` | `control_plane_db` | #613 control-plane DB (`mode=ro`), authoritative for exclusive ownership (#600/#601) |
| `leases` | `control_plane_db` | #613 control-plane DB; degrades if the `work_items` table is absent |
| `locks` | `filesystem` | durable per-issue lock files (`issue_lock_store`) |
| `worktrees` | `filesystem` | registered git worktrees via the #432 hygiene scanner |
| `namespaces` | `filesystem` | the active profile serving this web process (others are not enumerable) |
The payload restates this map under `field_authority` for machine consumers.
### Ownership safety
`ownership_authority_complete` is true only when every ownership-bearing section
(`sessions`, `leases`, `locks`) read cleanly. While it is false, nothing is
reported as unowned and no collision is asserted from a degraded source —
absence of evidence is reported as absence of evidence, never as free work.
`collisions` surfaces detectable conflicts, each with a `kind` and `severity`:
`lock-without-worktree`, `duplicate-live-lock`, `live-lock-dead-owner` (unexpired
lease, dead pid — a #753 recovery candidate that would read as live to a naive
timestamp check), `stale-lock-dead-owner`, `expired-lock-live-owner` (the
#635/#760 daemon-pid deadlock), `concurrent-active-lease`, `active-lease-past-expiry`,
and `orphan-lease`. Collisions are emitted only from sections that read cleanly.
The control-plane DB is opened through a `mode=ro` URI so a read never creates
or migrates it; paths are collapsed against `$HOME`, URLs lose userinfo and
query strings, and credential-shaped values are redacted at the boundary. Lease
steal/release and worktree deletion are Phase 2+ and have no representation here.
## Workflow-event timeline (#637)
`GET /api/v1/timeline` is a read-only, versioned aggregation of workflow
events from every available source into one normalised, filterable stream. It
is the model layer for the Phase 1 timeline console view (a later child issue
of #631); this issue ships the schema, adapters, and read API only.
### Schema (versioned)
`webui/timeline.py` declares `TIMELINE_SCHEMA_VERSION` (currently `1`) and the
frozen `WorkflowEvent` record. Every response carries `schema_version` so a
consumer can branch on shape. One event:
```json
{
"source": "control_plane",
"event_type": "lease.renew",
"event_key": "cp:1421",
"timestamp": "2026-07-23T02:00:00Z",
"actor": null,
"role": null,
"issue_number": 637,
"pr_number": null,
"session_id": null,
"tool_name": null,
"decision": null,
"message": "lease renewed",
"correlation_id": "issue#637",
"evidence_refs": [],
"sensitive": true
}
```
`event_key` is stable and unique per source (`cp:<event_id>`,
`cth:<kind>:<number>:<comment_id>`), so pagination and dedup are deterministic.
### Sources and field authority
| Source | Adapter | Authority |
|---|---|---|
| Control-plane `events``work_items` | `adapt_cp_events` | `event_type`, `message`, `timestamp`, issue/PR scope come from the CP database, read through a `mode=ro` URI (never creates the DB or runs migrations) |
| Gitea Canonical Thread Handoff comments | `adapt_cth_comments` | `actor`, `role` (next owner), `decision`, `evidence_refs`, `timestamp` come from the parsed CTH comment body (`canonical_thread_handoff`) |
Handoff comments are thread-scoped: they are only read when the request filters
by a single `issue` or `pr`. Otherwise the handoff source reports `not run`
with a reason — it is never rendered as empty-and-healthy. Each source degrades
independently: an unavailable control-plane DB or a failed comment fetch is a
`sources[]` entry with `ok:false` and a `reason`, never a dropped timeline.
### Query parameters
`issue`, `pr`, `session` (conjunctive filters); `limit` (default 50, max 500)
and `offset` for pagination; `remote`, `org`, `repo` to override the default
registry-project scope. Events sort ascending by
`(timestamp, source_rank, event_key)`; missing timestamps sort last.
### Filter authority, and refusing what cannot be answered
A filter dimension is only meaningful for a source whose records carry it.
Each source declares its own support in `_SOURCE_FILTER_SUPPORT` and reports it
per response as `supported_filters` / `unsupported_filters`:
| Source | issue | pr | session |
|---|---|---|---|
| `control_plane` | yes | yes | **no** — the `events` table is `(event_id, work_item_id, event_type, message, created_at)` and records no session |
| `gitea_handoff` | yes | yes | yes — a CTH comment declares its own `Session:` field |
`session_id` is read only from that declared CTH field. It is never inferred
from a work item, an actor, or message text, and a value that is
redaction-altering or bare-secret-shaped is dropped rather than emitted.
When **no source that ran** can carry a requested dimension, the request is
refused rather than answered: the response is `422` with `ok:false` and a
structured `error` naming `unsupported_filters` and the per-source reason. A
`200` with zero events would tell an operator that no such activity exists,
which is a stronger — and false — claim than "this cannot be answered here".
A source that *can* answer the dimension and simply matched nothing still
returns `200` with `ok:true` and an empty page.
### Redaction
Every free-text field (event messages, decision/proof text, roles, actors) is
passed through the console redaction policy (`webui.console_redaction`, backed
by `gitea_audit.redact`) before it leaves the module, failing closed to the
placeholder. No unredacted tool arguments or secrets are ever emitted, and a
generation error never drops raw data to a caller or a log.
Redaction also runs *before* any structured value is derived from free text.
`evidence_refs` are extracted from already-redacted proof/decision text, and a
commit reference is recognised only where the text declares one (`commit`,
`head`, `base`, `sha`, …). An undeclared 40-character hex run has the exact
shape of a Gitea access token, so it is never lifted out of prose into a
structured field. Every reference is then independently revalidated against an
allowed shape and a second redaction pass immediately before serialization;
anything unproven is dropped and the event is flagged `sensitive`.
### Tests
```bash
pytest tests/test_webui_timeline.py -q
```
## Tests
```bash
-81
View File
@@ -1,81 +0,0 @@
# Web Console: Notifications & Human-Attention Routing (#648)
- **Status:** Phase 3 Live
- **Tracking Issue:** [#648](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/648)
- **Parent Epic:** [#631](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/631)
- **Attention Boundary Reference:** [#628](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/628)
---
## 1. Overview
The **Notifications & Human-Attention Console** (`/notifications`, `/api/v1/notifications`) provides intelligent event classification and human-attention routing for autonomous workflow operations.
To prevent alert fatigue while ensuring critical escalation boundaries are never missed, events are classified into three distinct **Attention Classes**:
1. **`human-required`** (Urgent Escalation Boundary):
- Items requiring immediate human intervention or business decisions.
- Triggers: Auth failures, hard stops, irrecoverable state, decision locks, failed report validations, critical probe errors.
- Display: Highlighted in red (`badge-blocked`) with a `HUMAN REQUIRED` badge.
2. **`operator`** (Operational Inbox):
- Items requiring controller or operator review/triage during routine execution.
- Triggers: Blocked PRs (merge conflicts), stale leases, duplicate PRs on issues, unassigned ready work.
- Display: Displayed in orange/yellow (`badge-claimed`).
3. **`routine`** (Background Workflow Transitions):
- Normal, healthy workflow transitions and state progressions.
- Triggers: Active PRs/issues in standard state, clean branch creation, routine heartbeats.
- Display: Filtered out of default inbox views to eliminate notification spam; viewable on demand via the "Routine" or "All" tab.
---
## 2. API Endpoints
### `GET /api/v1/notifications`
*Compatibility Alias:* `GET /api/notifications`
#### Query Parameters:
- `project_id` (optional): Filter notifications by project ID.
- `attention_class` (optional): `inbox` (default: human-required + operator), `human-required`, `operator`, `routine`, `all`.
#### Example JSON Response:
```json
{
"project_id": "gitea-tools",
"repo_label": "Scaled-Tech-Consulting/Gitea-Tools",
"human_required_count": 0,
"operator_count": 2,
"routine_count": 5,
"total_count": 7,
"fetch_error": null,
"inbox_items": [
{
"id": "notif-pr-block-742",
"attention_class": "operator",
"category": "blocker",
"title": "Blocked PR #742",
"summary": "PR #742 requires merge conflict resolution.",
"work_kind": "pr",
"work_number": 742,
"project_id": "gitea-tools",
"repo_label": "Scaled-Tech-Consulting/Gitea-Tools",
"created_at": "2026-07-25T16:39:47Z",
"deep_link": "/traffic",
"requires_human": false,
"extra": {}
}
],
"all_items": [...]
}
```
---
## 3. UI Navigation
- Access via the **Traffic** navigation menu: **Traffic → Notifications**.
- The main view displays:
- **Metrics Summary Bar**: Highlighting counts for Human Required, Operator Inbox, and Routine items.
- **Attention Filter Tabs**: Toggle between Inbox (Human + Operator), Human Required, Operator, Routine, and All.
- **Structured Event Table**: Displays category, title, summary, work item links, and timestamps.
-160
View File
@@ -1,160 +0,0 @@
# Web console requests: intent preview and workflow initiation (#643)
**Phase 2. Preview is always live and always read-only. Initiation is wired but
denied until an operator opts in.**
Before this surface, starting role work meant pasting a prompt into a terminal
and trusting the operator to have checked the allocator first. Nothing enforced
that check, so two sessions could reach for the same issue and each believe it
was theirs. This page replaces the paste with a *request*: a desired role, an
issue or PR, and a stated intent, answered by an authorization decision and —
on confirmation — an exclusive assignment from the allocator.
| Concern | Module |
|---------|--------|
| Request model, preview, initiation | `webui/request_service.py` |
| Form and preview rendering | `webui/request_views.py` |
| Authorization | `webui/console_authz.py` (`initiate_workflow`) |
| Audit | `webui/console_audit.py` |
| Ownership substrate | `allocator_service.py` + `control_plane_db.py` |
## Surfaces
| Path | Method | Purpose |
|------|--------|---------|
| `/requests` | GET | Request form |
| `/requests` | POST | Render an intent preview. **Never assigns.** |
| `/api/v1/requests/preview` | POST | Intent preview as JSON |
| `/api/v1/requests/apply` | POST | Initiate — confirmed, audited, allocator-owned |
The HTML form has no initiate button on purpose. Initiating requires a
confirmed POST to `/api/v1/requests/apply`, so a stray form submission cannot
reserve work as a side effect.
## The request
```json
{
"desired_role": "author",
"work_kind": "issue",
"work_number": 643,
"intent_summary": "implement request preview and initiation",
"remote": "prgs",
"org": "Scaled-Tech-Consulting",
"repo": "Gitea-Tools",
"expected_head_sha": null
}
```
`desired_role` is one of `author`, `reviewer`, `merger`, `reconciler`,
`controller`. `work_kind` is `issue` or `pr`. `remote`/`org`/`repo` default to
the first project in the registry when omitted; when neither the request nor
the registry resolves them, the request is rejected rather than pointed at some
other repository. `intent_summary` is required — it is what the audit record
states as the reason — and is truncated to 500 characters.
Parsing rejects rather than corrects. An unknown role, an unknown work kind, a
non-positive number, or a missing intent each return `400` with a `reason_code`
and the offending `field`.
## Preview
Five checks, each with its own verdict, reason code, and detail:
| Check | Passes when |
|-------|-------------|
| `authorization` | The console principal holds `operator` or above |
| `capability` | The desired role maps to a declared profile and MCP namespace |
| `lease_availability` | No active claim holds the work unit |
| `next_safe_action` | The allocator would independently select this exact work unit |
| `head_pin` | PR work resolves to a head SHA, and a supplied SHA still matches |
A preview also returns the role's `allowed_actions` and `prohibited_actions`
(from `allocator_service.ROLE_ACTIONS`), the `required_profile` and
`required_namespace` the work must run under, and a `correlation_id` that ties
the preview to its audit record and to any assignment that follows.
Preview is read-only in the strict sense: it calls the allocator with
`apply=false` and writes nothing but an audit line. An unauthorized principal
never reaches the allocator or the control-plane DB at all, so a denial cannot
be used to enumerate the queue.
## Initiation
`POST /api/v1/requests/apply` refuses in this order, and every refusal returns
before any assignment is attempted:
| Condition | Outcome | Status |
|-----------|---------|--------|
| Unparseable request | `invalid_request` | 400 |
| Not authorized, or execution not wired | `denied` | 403 |
| `confirm` not set | `denied` / `confirmation_required` | 409 |
| Work unit already claimed | `blocked` / `duplicate_assignment` | 409 |
| Allocator would select other work | `wait` / `not_next_safe_work` | 409 |
| Allocator declines on apply | `blocked` or `wait` | 409 |
| Evidence unavailable | `wait` / `evidence_unavailable` | 503 |
| Assigned | `assigned_work` | 201 |
A success returns the assignment plus a `handoff` block naming the profile, the
namespace, and the actions that stay forbidden — enough for the operator to
continue in the right MCP namespace without guessing.
### Why apply runs the allocator twice
The allocator is the only source of exclusive ownership (#600 / #613), and it
selects work; it does not take orders. So `apply` runs a dry-run first and
proceeds only when the allocator would independently pick the requested work
unit. If it would not, the request reports `wait` and mutates nothing.
A request is therefore a *confirmation* of the allocator's decision, never an
override of it. The apply call carries the dry-run's
`candidate_set_fingerprint` as a CAS pin (#776), so a queue that changed
between the two calls fails closed rather than assigning against a stale view.
The result is checked again on the way out: an assignment naming a different
work unit is not read as success.
### Fail-closed defaults
- An unreadable control-plane DB denies. It is never treated as "nothing holds
this work unit".
- An incomplete queue inventory denies (#758). Ranking a partial candidate set
can select the wrong work.
- An allocator that raises denies.
- PR work with no resolvable head SHA denies; a supplied SHA that no longer
matches denies with `head_moved`.
## Enabling initiation
Execution is wired off. Set `WEBUI_REQUESTS_EXECUTION=1` to enable it for the
`initiate_workflow` action only — see
[`webui-authz-audit.md`](webui-authz-audit.md) for why this is an
action-scoped flag rather than a phase bump. With the variable unset, `apply`
returns `403` with `reason_code: unauthorized` no matter who asks.
Enabling execution does **not** enable approvals or merges. Those are phase 3
console actions and remain forbidden in every path here; the console reserves
work and hands off, and the MCP role profile enforces what that role may then
do.
## Audit
Every preview and every apply emits a console audit record (schema in
[`webui-authz-audit.md`](webui-authz-audit.md)):
| Event | `result` |
|-------|----------|
| Preview | `previewed` |
| Refusal at any stage | `denied` |
| Assignment created | `succeeded` |
`correlation.request_id` carries the request's `correlation_id`, and a
successful record's `metadata` carries `assignment_id` and `lease_id`, so an
assignment can be traced back to the intent that produced it. The operator's
`intent_summary` travels in `metadata` and passes through the standard
redaction pass before persistence like every other field.
## Non-goals
- No browser-initiated approve or merge, in this phase or any other.
- No bypass of allocator exclusive ownership; no self-selection of work.
- No auto-start from raw monitoring incidents (#612 stays downstream).
-102
View File
@@ -1,102 +0,0 @@
# Web Console: restart status, impact preview, and approval state (#667)
Phase 1 of the console restart surface. It consumes the #655 coordinator
substrate and displays it. It performs no restart, reload, drain, approval, or
process action, and it registers no write endpoint.
Issue #667's rollout is explicit — *status views first, write approval after the
backend gates are green* — and this change delivers only the status half.
## Surfaces
| Path | Method | Purpose |
|------|--------|---------|
| `/runtime/restart` | GET | Restart status page |
| `/api/v1/system/restart/status` | GET | Same snapshot as JSON |
Both accept an optional `restart_class` query parameter (default
`full_mcp_restart`). An unrecognised class is not an error: the coordinator
resolves it as unknown and fails closed, and the page shows the resulting deny.
Neither path accepts `POST`; a write attempt returns `405`, and a test asserts
it.
## What it shows
* **Impact preview (#658)** — verdict, blast radius, affected sessions, leases,
critical sections, mutations, and the counts behind them, evaluated
`dry_run=True` against live control-plane state.
* **Drain proof (#661)** — verification of a supplied proof: valid, clean,
expired, tampered, and the reasons behind a refusal.
* **Post-restart reconcile (#662)** — the most recent completion proof, its
overall status, and which dimensions still require follow-up.
* **Restart classes (#663)** — the least-privilege matrix, with *you may
request* and *you may execute* computed for the viewing role rather than for a
generic operator.
* **Approval controls (#633)** — the authorization state of
`system.restart_namespace` and `system.reload_namespace`.
* **Break-glass (#664)** — declared and marked unavailable; see below.
## Three rules this surface holds itself to
A status page that is wrong is worse than one that is missing, because an
operator acts on it. Three properties are enforced by tests, and each was
verified by reverting the guard and watching a test fail.
### An unreadable source reports unavailable, never green
Every source carries its own `SourceStatus`. Nothing substitutes a default,
placeholder, or self-comparison for a reading that failed. An unreadable
control-plane database yields `inventory_complete: false`, which the coordinator
itself turns into a fail-closed verdict, and the page says the blast radius is
unknown rather than showing an empty affected-sessions table.
An absent drain proof is reported as absent — not as a pass. The #661 gate
authorizes a restart only against a valid, unexpired, clean proof, so no proof
is precisely the state that gate denies on.
### Authorization is asked the way execution would ask it
Every probe passes `for_execution=True`.
Asked without it, an admin is `allowed` for `system.restart_namespace`. On a
control surface that reads as a live button. Asked the way an execution attempt
would ask, the same principal is refused `phase_not_active`, because the console
is in Phase 1 and the action is Phase 2. This surface reports the second answer.
`execution_enabled` is therefore `false` for every action and every role today,
and a test asserts that across the whole role matrix.
### The control-plane database is opened read-only
`ControlPlaneDB()` creates directories and runs migrations on construction — a
write. This surface never constructs one. It opens the sqlite file with
`mode=ro`, exactly as `webui/inventory.py` does, and treats a missing file as
missing authority rather than as an empty inventory.
The test that protects this points at a path inside a directory that already
exists, so a read-write `connect` would really create the file. A nested
missing-directory path would have passed for the wrong reason.
## Break-glass is declared, not offered
The break-glass workflow (#664) is not available on this branch's base. The
panel is rendered to operator-class roles as **unavailable**, naming the issue
that tracks it. It is not silently omitted, because an operator who has been
told a governance path exists needs to see that it is not wired here; and it is
not rendered as a control, because there is nothing behind it.
Unprivileged viewers see only a note that the surface is operator-class.
## Redaction and escaping
Every interpolated value passes through `_esc` (`html.escape(..., quote=True)`).
Free-form text and anything that can carry a filesystem path additionally passes
through `webui.inventory.scrub_text`, which redacts credential-shaped tokens
inside a string rather than only at its start. The impact payload is passed
through `webui.inventory.scrub` before rendering.
## Linkage
Parent #655 · extends #642 · consumes #658, #661, #662, #663 · RBAC #633 ·
console #631 · vision #652 · roadmap #653 · break-glass #664.
-1015
View File
File diff suppressed because it is too large Load Diff
+14 -138
View File
@@ -1169,147 +1169,23 @@ def server_command():
return python, [os.path.join(root, "mcp_server.py")]
RECOGNIZED_GITEA_ENV_KEYS = frozenset({
"GITEA_MCP_CONFIG",
"GITEA_MCP_PROFILE",
"GITEA_PROFILE_NAME",
"GITEA_SERVICE",
"GITEA_EXECUTION_ROLE",
"GITEA_CLIENT_MANAGED",
"GITEA_MCP_CLIENT_MANAGED",
"GITEA_SERVER_PROVENANCE",
"GITEA_AUTHOR_WORKTREE",
"GITEA_ACTIVE_WORKTREE",
"GITEA_REVIEWER_WORKTREE",
"GITEA_MERGER_WORKTREE",
"GITEA_CANONICAL_REPOSITORY_ROOT",
"GITEA_MCP_SESSION_STATE_TTL_HOURS",
"GITEA_DISABLE_KEYCHAIN",
"GITEA_CONTROL_PLANE_DB",
"GITEA_DB_PATH",
"GITEA_LOG_LEVEL",
"GITEA_DEBUG",
"GITEA_HMAC_SECRET",
"GITEA_IRRECOVERABLE_HMAC_SECRET",
"GITEA_FORCE_MCP_RUNTIME_CHECK",
"GITEA_FORCE_CLIENT_MANAGED",
# #975: the client-identity inputs the server actually consumes at startup
# (CLIENT_NAME_ENV / CLIENT_INSTANCE_ENV / CLIENT_SESSION_ENV in
# gitea_mcp_server). Production read them while this allowlist omitted them,
# so the peer-env scan classified them as unsupported overrides and the
# capability resolver refused every mutation fleet-wide. Named individually
# on purpose: no prefix is added, so an unrecognised GITEA_* override is
# still refused exactly as it was before.
"GITEA_MCP_CLIENT",
"GITEA_MCP_CLIENT_INSTANCE",
"GITEA_MCP_CLIENT_SESSION",
# #978: operator-approved fleet enrollment identity shared by one launch.
"GITEA_MCP_FLEET_RUN_ID",
"GITEA_MCP_PROCESS_IDENTITY",
# #978 B1/B2: launcher-sealed provenance + peer-visible worker identity so
# the instance-aware mutation gate can distinguish two legitimate
# application instances that share a profile from a true duplicate worker.
"GITEA_MCP_INSTANCE_PROVENANCE",
"GITEA_MCP_WORKER_IDENTITY",
"GITEA_MCP_GENERATION_ID",
# #975 review 652 B1: production also consumes HEARTBEAT_INTERVAL_ENV from
# mcp_worker_identity via gitea_mcp_server._start_worker_heartbeat. Omitting
# it reproduced the same unsupported-env → runtime_reconnect_required
# failure mode for the documented operator override. Named individually;
# no GITEA_* / GITEA_WORKER_* prefix is added.
"GITEA_WORKER_HEARTBEAT_INTERVAL_SECONDS",
})
RECOGNIZED_GITEA_ENV_PREFIXES = (
"GITEA_TOKEN_",
"GITEA_PASS_",
"GITEA_USER_",
"GITEA_URL_",
"GITEA_HOST_",
"GITEA_REMOTE_",
"GITEA_HTTP_HEADER_",
)
def get_unconsumed_gitea_env_overrides(env=None) -> dict[str, str]:
"""Find unsupported GITEA_* env vars present in *env* (defaults to os.environ)."""
target = os.environ if env is None else env
unconsumed = {}
for key, value in target.items():
if key.startswith("GITEA_"):
if key in RECOGNIZED_GITEA_ENV_KEYS:
continue
if any(key.startswith(p) for p in RECOGNIZED_GITEA_ENV_PREFIXES):
continue
unconsumed[key] = str(value)
return unconsumed
def launcher_entry(
profile_name,
config_path=None,
*,
client_type=None,
client_instance_id=None,
fleet_run_id=None,
session_id=None,
launch_nonce=None,
):
def launcher_entry(profile_name, config_path=None):
"""Return a thin MCP launcher entry for *profile_name*.
Contains command/args and the GITEA_MCP_* / GITEA_CLIENT_MANAGED env vars —
never a token or password. Suitable for Claude / Gemini / Codex
``mcpServers`` blocks.
#978 B1: production launches mint (or reuse) a trusted
``GITEA_MCP_CLIENT_INSTANCE`` so workers never fall back to the legacy
placeholder identity on the real serve path. Pass *client_instance_id* to
resume the same launch; omit it for a fresh launch (new ID).
Contains only command/args and the two GITEA_MCP_* env vars — never a token
or password. Suitable for Claude / Gemini / Codex ``mcpServers`` blocks.
"""
import mcp_application_launcher as app_launcher
built = app_launcher.launcher_entry_for_profile(
profile_name,
client_type=client_type or "unknown",
config_path=config_path or DEFAULT_CONFIG_PATH,
client_instance_id=client_instance_id,
fleet_run_id=fleet_run_id,
session_id=session_id,
launch_nonce=launch_nonce,
server_key="gitea-tools",
)
# Public shape stays {server_key: {command, args, env}} for existing callers.
return {"gitea-tools": built["gitea-tools"]}
def multi_namespace_launcher_entries(
profile_by_namespace,
*,
client_type,
config_path=None,
client_instance_id=None,
fleet_run_id=None,
session_id=None,
launch_nonce=None,
):
"""Build production mcpServers for all five namespaces of one application launch.
One shared trusted ``client_instance_id`` is minted (or reused) and
propagated to every namespace worker env. See
:func:`mcp_application_launcher.build_application_mcp_servers`.
"""
import mcp_application_launcher as app_launcher
return app_launcher.build_application_mcp_servers(
profile_by_namespace,
client_type=client_type,
config_path=config_path or DEFAULT_CONFIG_PATH,
client_instance_id=client_instance_id,
fleet_run_id=fleet_run_id,
session_id=session_id,
launch_nonce=launch_nonce,
)
command, args = server_command()
return {
"gitea-tools": {
"command": command,
"args": args,
"env": {
"GITEA_MCP_CONFIG": config_path or DEFAULT_CONFIG_PATH,
"GITEA_MCP_PROFILE": profile_name,
},
}
}
def keychain_set(item_id, token, account=None, runner=subprocess.run):
+168 -4317
View File
File diff suppressed because it is too large Load Diff
-6
View File
@@ -500,11 +500,6 @@ def assess_transport_for_auth_mint() -> dict[str, Any]:
native = mcp_daemon_guard.is_native_mcp_transport()
pytest = mcp_daemon_guard.is_pytest_runtime()
production = mcp_daemon_guard.is_production_native_mcp_transport()
# #931: report which transport underwrites the verdict, read through the
# one shared accessor rather than assumed to be stdio. The gate's decision
# is unchanged here; naming the transport is what lets #932 re-derive the
# guarantee from an authenticated session instead of from the bind.
bound = mcp_daemon_guard.bound_transport()
if not native and not pytest:
reasons.append(
"irrecoverable provenance authorization requires production native "
@@ -515,7 +510,6 @@ def assess_transport_for_auth_mint() -> dict[str, Any]:
"allowed": not reasons,
"native_mcp_transport": native,
"production_native_mcp_transport": production,
"bound_transport": bound,
"pytest": pytest,
"reasons": reasons,
}
-7
View File
@@ -16,18 +16,11 @@ ISSUE_LOCK_FILE = os.environ.get("GITEA_ISSUE_LOCK_FILE", "/tmp/gitea_issue_lock
SOURCE_LOCK_ISSUE = "gitea_lock_issue"
SOURCE_LOCK_ADOPTION = "gitea_lock_issue_adoption"
SOURCE_OPERATOR_OVERRIDE = "operator_override"
SOURCE_RECOVER_DIRTY_ORPHANED = "gitea_recover_dirty_orphaned_issue_worktree"
# #864: dirty-preserving same-claimant author-session rebind (dead owner PID).
SOURCE_DIRTY_SAME_CLAIMANT_REBIND = (
"gitea_rebind_dirty_same_claimant_author_session"
)
SANCTIONED_LOCK_SOURCES = frozenset({
SOURCE_LOCK_ISSUE,
SOURCE_LOCK_ADOPTION,
SOURCE_OPERATOR_OVERRIDE,
SOURCE_RECOVER_DIRTY_ORPHANED,
SOURCE_DIRTY_SAME_CLAIMANT_REBIND,
})
_OPERATOR_OVERRIDE_ENV = "GITEA_ISSUE_LOCK_OPERATOR_OVERRIDE"
+8 -130
View File
@@ -85,12 +85,6 @@ HEAD_RELATION_STRICT_DESCENDANT = "strict_descendant"
# #772: an unpublished claim has no recorded head to compare against at all, so
# its head is measured against the base the branch was cut from instead.
HEAD_RELATION_DESCENDS_FROM_BASE = "descends_from_recorded_base"
# #871: the remote/PR head advanced *past* the recorded head via a sanctioned
# merge-based branch synchronization (``gitea_update_pr_branch_by_merge``) while
# the local worktree stayed at the recorded head. This is the inverse of the
# #768 descendant relation — here the *remote* strictly descends the local head,
# and only because a base was merged into the branch, proven server-side.
HEAD_RELATION_REMOTE_MERGE_SYNCED = "remote_merge_synced"
# Which body of evidence a recovery was decided on (#772 AC10). These are not
# interchangeable: a published claim proves ownership against a remote/PR head,
@@ -272,70 +266,6 @@ def _assess_base_descendancy(
]
def _assess_remote_merge_synced(
sync_provenance: Mapping[str, Any] | None,
*,
recorded_head: str,
remote_head: str,
) -> tuple[bool, list[str]]:
"""Did ``remote_head`` advance past ``recorded_head`` via a sanctioned
merge-based branch sync (#871)?
``sync_provenance`` is the server-side git observation from
``issue_lock_worktree.read_merge_sync_provenance``. Its own
``prior_head_sha`` / ``synced_head_sha`` are re-checked against the heads
this assessment is actually reasoning about, so an observation taken for some
other pair of commits stale, mismatched, or hand-built can never
authorize recovery. This is the inverse of ``_assess_strict_descendant``: the
recorded head is the ancestor and the *remote* head is the descendant, and it
is accepted only because the remote head is a base-into-branch merge that
preserved the branch mainline back to the recorded head.
Returns ``(proven, notes)``. Notes name the exact missing element so a
refused caller sees why, never a bare "unproven".
"""
if not isinstance(sync_provenance, Mapping):
return False, [
"no server-derived merge-sync provenance observation was available; a "
"remote head ahead of the recorded head cannot be accepted"
]
probe_prior = _text(sync_provenance.get("prior_head_sha"))
probe_synced = _text(sync_provenance.get("synced_head_sha"))
if probe_prior != recorded_head or probe_synced != remote_head:
return False, [
f"merge-sync observation covers {probe_prior or 'unknown'} -> "
f"{probe_synced or 'unknown'}, not the heads under assessment "
f"({recorded_head} -> {remote_head})"
]
if not sync_provenance.get("probe_ok"):
return False, (
list(sync_provenance.get("reasons") or [])
or ["merge-sync provenance probe did not complete; provenance unproven"]
)
if not sync_provenance.get("prior_is_ancestor"):
return False, [
f"recorded head {recorded_head} is not an ancestor of remote head "
f"{remote_head}; a rewritten or force-moved head cannot be recovered"
]
if not sync_provenance.get("is_merge_sync"):
return False, (
list(sync_provenance.get("reasons") or [])
or [
f"remote head {remote_head} is not a sanctioned merge-based sync "
f"of the base into the branch above {recorded_head}"
]
)
proof = _text(sync_provenance.get("proof")) or (
f"{remote_head} merged the base into the branch above {recorded_head}"
)
return True, [
f"remote head {remote_head} advanced past recorded head {recorded_head} "
f"via a sanctioned merge-based branch sync ({proof})"
]
def assess_dead_session_lock_recovery(
existing_lock: Mapping[str, Any] | None,
*,
@@ -360,7 +290,6 @@ def assess_dead_session_lock_recovery(
remote_branch_exists: bool | None = None,
recorded_base_sha: str | None = None,
base_ancestry: Mapping[str, Any] | None = None,
sync_provenance: Mapping[str, Any] | None = None,
) -> dict[str, Any]:
"""Decide whether a dead-session author lock may be natively recovered.
@@ -540,39 +469,19 @@ def assess_dead_session_lock_recovery(
head_relation = HEAD_RELATION_STRICT_DESCENDANT
ancestry_proof = notes[0] if notes else None
else:
# #871: the reverse relation — the remote head advanced past
# the recorded/local head via a sanctioned merge-based branch
# sync while the local worktree stayed put. Accepted only on
# server-proven merge-sync provenance, never a caller claim.
synced, sync_notes = _assess_remote_merge_synced(
sync_provenance,
recorded_head=local_head,
remote_head=remote_head,
reasons.append(
f"local head {local_head} does not match remote branch head "
f"{remote_head}"
)
if synced:
head_relation = HEAD_RELATION_REMOTE_MERGE_SYNCED
ancestry_proof = sync_notes[0] if sync_notes else None
else:
reasons.append(
f"local head {local_head} does not match remote branch "
f"head {remote_head}"
)
reasons.extend(notes)
reasons.extend(sync_notes)
reasons.extend(notes)
evidence["recorded_base"] = recorded_base or None
evidence["local_head"] = local_head or None
evidence["remote_head"] = remote_head or None
# ``recorded_head`` is the head recovery is being measured against;
# ``accepted_head`` is the head this recovery actually adopts. They differ
# only in the descendant case, and downstream gates need both (#768 AC2/AC7).
# #871: in the merge-sync case the branch/PR already carries the synced
# remote head, so that is the head recovery adopts; the local worktree stays
# at the ancestor recorded head.
evidence["recorded_head"] = remote_head or None
if head_relation == HEAD_RELATION_REMOTE_MERGE_SYNCED:
evidence["accepted_head"] = remote_head or None
else:
evidence["accepted_head"] = local_head or None
evidence["accepted_head"] = local_head or None
evidence["head_relation"] = head_relation
evidence["ancestry_proof"] = ancestry_proof
@@ -584,17 +493,12 @@ def assess_dead_session_lock_recovery(
# contradictory; re-stating it as a head mismatch would only obscure why.
if not unpublished and local_head and pr_head != local_head:
# A descendant recovery has not been published yet, so the open PR
# legitimately still points at the recorded head. A merge-sync
# recovery's PR legitimately sits at the advanced remote head. Any
# other disagreement is a real mismatch.
# legitimately still points at the recorded head. Any other
# disagreement is a real mismatch.
if not (
head_relation == HEAD_RELATION_STRICT_DESCENDANT
and remote_head
and pr_head == remote_head
) and not (
head_relation == HEAD_RELATION_REMOTE_MERGE_SYNCED
and remote_head
and pr_head == remote_head
):
reasons.append(
f"open PR #{pr_number} head {pr_head} does not match local head "
@@ -715,11 +619,7 @@ def assess_dead_session_lock_recovery(
)
if (
head_relation
in (
HEAD_RELATION_STRICT_DESCENDANT,
HEAD_RELATION_DESCENDS_FROM_BASE,
HEAD_RELATION_REMOTE_MERGE_SYNCED,
)
in (HEAD_RELATION_STRICT_DESCENDANT, HEAD_RELATION_DESCENDS_FROM_BASE)
and ancestry_proof
):
proof.append(ancestry_proof)
@@ -797,16 +697,6 @@ def owning_pr_recovery_evidence(
return None
if accepted_head != local_head:
return None
elif relation == HEAD_RELATION_REMOTE_MERGE_SYNCED:
# #871: the PR already sits at the advanced remote head; the local
# worktree is the ancestor the merge preserved. The head the open PR
# shows and the head recovery adopts are both the synced remote head.
if not remote_head or pr_head != remote_head:
return None
if accepted_head != remote_head:
return None
if not local_head or local_head == remote_head:
return None
else:
return None
try:
@@ -875,18 +765,6 @@ def recovered_owning_pr_from_lock(
return None
if not accepted_head or accepted_head == recorded_head:
return None
elif relation == HEAD_RELATION_REMOTE_MERGE_SYNCED:
# #871: PR sits at the advanced remote head, which is both the recorded
# measured-against head and the adopted head; the local worktree is the
# ancestor the merge preserved.
remote_head = _text(record.get("remote_head"))
local_head = _text(record.get("local_head"))
if not remote_head or pr_head != remote_head:
return None
if accepted_head and accepted_head != remote_head:
return None
if not local_head or local_head == remote_head:
return None
else:
return None
try:
-111
View File
@@ -436,117 +436,6 @@ def owning_pr_renewal_evidence(
}
def owning_pr_renewal_from_lock(
lock_record: Mapping[str, Any] | None,
) -> dict[str, Any] | None:
"""Rebuild owning-PR renewal evidence from a persisted lock (#945).
The renewal mirror of ``issue_lock_recovery.recovered_owning_pr_from_lock``.
``owning_pr_renewal_evidence`` supplies the waiver for the duration of the
``gitea_lock_issue`` call only. The commit, push, create-PR, and
duplicate-assessment gates run later in their own calls and re-derive
ownership from the durable lock instead so without this the open PR that
renewal already proved belongs to this author reappears there as competing
duplicate work, and the exact owner is refused with
``duplicate_commit_prevented`` despite complete matching evidence.
This reads only the ``lease_renewal`` block that the server itself writes,
on a lock the caller must already own. Like the recovery mirror it is a
re-read of server-derived state, never a fresh assertion: a caller able to
forge it could equally forge the lock file every other ownership gate
already treats as authoritative.
Renewal has no descendant case the assessor required the local, remote and
PR heads to be equal so that equality is re-checked here, and the record
must still name the claimant the lock records.
**What the claimant check below is, and what it is not.** It compares
``lease_renewal.identity``/``profile`` against the claimant recorded on the
*same* lock file. Both sides are server-written fields of one document, so
this is an internal-consistency check: it rejects a lock whose renewal block
and claimant disagree. It does **not** consult the live authenticated caller
and therefore does not, on its own, prove that the session invoking a later
gate is the session the renewal was granted to.
The binding that actually keeps one session from using another's renewal is
structural, and it lives in the caller rather than here. The enforcement
paths load the lock through ``_load_existing_issue_lock()`` with no issue
coordinates, which resolves ``issue_lock_store.read_session_issue_lock()``
the session pointer at ``session-{os.getpid()}.json``. Lock *selection* is
scoped to the operating-system process, so a caller cannot aim the recheck
at a lock some other process bound. Its limits follow from what that scope
is: it is per-process, not per-authenticated-user; it says nothing about a
lock reached by explicit issue coordinates rather than the session pointer,
and nothing about two roles sharing one process. Live identity and profile
are enforced separately, by the mutation-authority and profile gates each
mutating path already runs not by this rebuild.
This function is therefore strictly a re-read with an added consistency
requirement. It narrows what a persisted lock can authorize; it never widens
it, and it never substitutes for a caller-identity gate.
"""
if not isinstance(lock_record, Mapping):
return None
record = lock_record.get("lease_renewal")
if not isinstance(record, Mapping) or not record.get("renewed"):
return None
branch_name = _text(record.get("branch_name")) or _text(
lock_record.get("branch_name")
)
pr_head = _text(record.get("pr_head_sha"))
local_head = _text(record.get("head_sha"))
remote_head = _text(record.get("remote_head_sha"))
raw_pr_number = record.get("pr_number")
raw_issue_number = lock_record.get("issue_number")
if raw_pr_number is None or raw_issue_number is None:
return None
if not branch_name or not pr_head:
return None
# The assessor required all three heads to agree before it granted renewal.
# Re-check, so a truncated, drifted, or hand-built record cannot widen the
# exemption past the single head the renewal disposition actually proved.
if not local_head or not remote_head:
return None
if pr_head != local_head or pr_head != remote_head:
return None
# Renewal is refused outright unless the durable lock records both a
# claimant username and profile, so a sanctioned record always carries them.
# Requiring them to still agree rejects a lock whose renewal block and
# claimant disagree. Both values are read from this one server-written
# document: this is internal consistency, not a check against the live
# authenticated caller — see the docstring for the binding that is.
claimant = lock_record.get("claimant")
if not isinstance(claimant, Mapping):
lease = lock_record.get("work_lease")
claimant = lease.get("claimant") if isinstance(lease, Mapping) else None
if not isinstance(claimant, Mapping):
return None
identity = _text(record.get("identity"))
profile = _text(record.get("profile"))
if not identity or identity != _text(claimant.get("username")):
return None
if not profile or profile != _text(claimant.get("profile")):
return None
try:
pr_number = int(raw_pr_number)
issue_number = int(raw_issue_number)
except (TypeError, ValueError):
return None
return {
"issue_number": issue_number,
"pr_number": pr_number,
"branch_name": branch_name,
"head_sha": pr_head,
"recorded_head": pr_head,
"accepted_head": pr_head,
"head_relation": "equal",
}
def build_renewal_record(
assessment: Mapping[str, Any] | None,
*,
+34 -1061
View File
File diff suppressed because it is too large Load Diff
-178
View File
@@ -145,184 +145,6 @@ def read_head_ancestry(
return result
def read_merge_sync_provenance(
worktree_path: str,
*,
prior_head_sha: str | None,
synced_head_sha: str | None,
remote: str | None = None,
) -> dict:
"""Observe whether ``synced_head_sha`` is a sanctioned merge-sync of a base
into the branch above ``prior_head_sha`` (#871/#872).
Reports server-derived git facts only; the recovery disposition lives in
``issue_lock_recovery``. All comparisons are executed locally in the
declared worktree -- nothing is taken from caller parameters.
Provenance is proven only when ALL hold:
* both commits are present (a rewritten/force-moved prior head leaves the
object graph and fails closed);
* ``prior_head_sha`` is a strict ancestor of ``synced_head_sha`` (the branch
history is preserved, never replaced);
* ``synced_head_sha`` is a merge commit (two or more parents), i.e. a base
merged in a plain fast-forward of new direct commits is not a sync;
* ``prior_head_sha`` is an ancestor of the merge's **first** parent, so the
branch mainline (first-parent lineage) still reaches the prior head a
rebase/force-push that re-authored the branch side fails this.
"""
path = (worktree_path or "").strip()
prior = (prior_head_sha or "").strip()
synced = (synced_head_sha or "").strip()
target_remote = (remote or "").strip() or None
result: dict = {
"prior_head_sha": prior or None,
"synced_head_sha": synced or None,
"probe_ok": False,
"prior_present": False,
"synced_present": False,
"prior_is_ancestor": False,
"synced_is_merge": False,
"first_parent_reaches_prior": False,
"is_merge_sync": False,
"first_parent_sha": None,
"parent_count": None,
"proof": None,
"reasons": [],
}
if not path or not prior or not synced:
result["reasons"].append(
"merge-sync provenance probe requires a worktree path and both "
"commit SHAs"
)
return result
if prior == synced:
result["reasons"].append(
"prior and synced heads are identical; no branch sync occurred"
)
return result
def _present(sha: str) -> bool:
res = subprocess.run(
["git", "-C", path, "rev-parse", "--verify", "--quiet", f"{sha}^{{commit}}"],
capture_output=True,
text=True,
check=False,
)
return res.returncode == 0
def _is_ancestor(ancestor: str, descendant: str) -> bool | None:
res = subprocess.run(
["git", "-C", path, "merge-base", "--is-ancestor", ancestor, descendant],
capture_output=True,
text=True,
check=False,
)
if res.returncode == 0:
return True
if res.returncode == 1:
return False
return None # failed probe — never a silent "no"
try:
result["prior_present"] = _present(prior)
result["synced_present"] = _present(synced)
if not result["synced_present"] and path and os.path.isdir(path):
if target_remote:
subprocess.run(
["git", "-C", path, "fetch", target_remote, "--quiet"],
capture_output=True,
text=True,
check=False,
)
result["synced_present"] = _present(synced)
if not result["synced_present"]:
subprocess.run(
["git", "-C", path, "fetch", "--quiet"],
capture_output=True,
text=True,
check=False,
)
result["synced_present"] = _present(synced)
except OSError as exc: # git unavailable — fail closed, never assume
result["reasons"].append(f"merge-sync provenance probe could not run: {exc}")
return result
if not result["prior_present"]:
result["reasons"].append(
f"prior head {prior} is not reachable in '{path}'; history may have "
"been rewritten or force-moved"
)
if not result["synced_present"]:
result["reasons"].append(
f"synced head {synced} is not reachable in '{path}'"
)
if not (result["prior_present"] and result["synced_present"]):
return result
ancestor = _is_ancestor(prior, synced)
if ancestor is None:
result["reasons"].append(
"ancestry probe failed; merge-sync provenance unproven"
)
return result
result["prior_is_ancestor"] = bool(ancestor)
if not ancestor:
result["reasons"].append(
f"prior head {prior} is not an ancestor of synced head {synced}; "
"the branch history was not preserved (not a merge-based sync)"
)
return result
parents_res = subprocess.run(
["git", "-C", path, "rev-list", "--parents", "-n", "1", synced],
capture_output=True,
text=True,
check=False,
)
if parents_res.returncode != 0:
result["reasons"].append(
f"could not read parents of {synced}; merge-sync provenance unproven"
)
return result
tokens = (parents_res.stdout or "").split()
# tokens[0] is the commit itself; the rest are its parents.
parents = tokens[1:]
result["parent_count"] = len(parents)
result["synced_is_merge"] = len(parents) >= 2
if not result["synced_is_merge"]:
result["probe_ok"] = True
result["reasons"].append(
f"synced head {synced} has {len(parents)} parent(s); a merge-based "
"branch sync produces a merge commit (two or more parents)"
)
return result
first_parent = parents[0]
result["first_parent_sha"] = first_parent
fp_reaches = _is_ancestor(prior, first_parent) if prior != first_parent else True
if fp_reaches is None:
result["reasons"].append(
"first-parent ancestry probe failed; merge-sync provenance unproven"
)
return result
result["first_parent_reaches_prior"] = bool(fp_reaches)
result["probe_ok"] = True
if not fp_reaches:
result["reasons"].append(
f"merge first parent {first_parent} does not reach prior head "
f"{prior}; the branch mainline was re-authored (not a sanctioned sync)"
)
return result
result["is_merge_sync"] = True
result["proof"] = (
f"synced head {synced} is a merge commit (parents={len(parents)}) whose "
f"first-parent lineage reaches prior head {prior}; base merged into branch"
)
return result
def read_recorded_base(
worktree_path: str,
*,
+13 -209
View File
@@ -39,7 +39,6 @@ SAFE_RELEASE_OWNED = "release_owned"
SAFE_STALE_PROMPT = "stale_prompt_lease"
SAFE_UNKNOWN = "inspect_only"
SAFE_NO_AUTHORITY = "file_or_comment_not_authoritative"
SAFE_CONSUME_CROSS_ROLE = "consume_cross_role_handoff"
LEASE_STATUS_ACTIVE = "active"
LEASE_STATUS_RELEASED = "released"
@@ -251,23 +250,6 @@ def decide_safe_next_action(
"same_owner": True,
"also_allowed": [SAFE_ABANDON_ALLOWED, SAFE_RELEASE_OWNED],
}
handoff = is_pending_cross_role_handoff({"lease": lease})
if handoff:
return {
"safe_next_action": SAFE_CONSUME_CROSS_ROLE,
"reasons": [
f"controller allocation pending handoff (freshness={status}); "
"required-role worker may consume without abandon/reassign; "
f"required_role={handoff['required_role']}"
],
"block": False,
"same_owner": False,
"owner_session_id": owner,
"required_role": handoff["required_role"],
"cross_role_handoff": True,
"handoff_status": "pending",
"also_allowed": [SAFE_ABANDON_ALLOWED],
}
return {
"safe_next_action": SAFE_ABANDON_ALLOWED,
"reasons": [
@@ -290,24 +272,6 @@ def decide_safe_next_action(
}
if not same_owner and status == "active":
# #843: pending cross-role handoff is consumable by required role
handoff = is_pending_cross_role_handoff({"lease": lease})
if handoff:
return {
"safe_next_action": SAFE_CONSUME_CROSS_ROLE,
"reasons": [
"controller cross-role allocation pending handoff; "
f"required_role={handoff['required_role']}; "
"consume via gitea_adopt_workflow_lease without "
"abandonment or sharing the controller session"
],
"block": False,
"same_owner": False,
"owner_session_id": owner,
"required_role": handoff["required_role"],
"cross_role_handoff": True,
"handoff_status": "pending",
}
return {
"safe_next_action": SAFE_WAIT_FOREIGN,
"reasons": [
@@ -476,84 +440,6 @@ def list_active_leases(
}
def parse_lease_provenance(lease_or_state: Mapping[str, Any] | None) -> dict[str, Any]:
"""Return durable lease provenance dict (empty when absent/unparseable)."""
if not lease_or_state:
return {}
if "provenance" in lease_or_state and isinstance(lease_or_state.get("provenance"), dict):
return dict(lease_or_state["provenance"])
raw = None
if "provenance_json" in lease_or_state:
raw = lease_or_state.get("provenance_json")
elif "lease" in lease_or_state and isinstance(lease_or_state.get("lease"), Mapping):
raw = lease_or_state["lease"].get("provenance_json")
if not raw:
return {}
if isinstance(raw, dict):
return dict(raw)
try:
loaded = json.loads(raw)
except (TypeError, json.JSONDecodeError):
return {}
return dict(loaded) if isinstance(loaded, dict) else {}
def is_pending_cross_role_handoff(
state: Mapping[str, Any] | None,
) -> dict[str, Any] | None:
"""Return handoff evidence when a controller allocation awaits consume (#843).
A pending handoff is identified by durable provenance written at
cross-role apply time not by title heuristics or session-id guessing.
"""
if not state:
return None
lease = state.get("lease") if isinstance(state.get("lease"), Mapping) else state
if not isinstance(lease, Mapping):
return None
status = str(lease.get("status") or "").strip().lower()
if status in (LEASE_STATUS_ABANDONED, LEASE_STATUS_RELEASED, LEASE_STATUS_EXPIRED):
return None
prov = parse_lease_provenance(state)
if not prov and isinstance(lease, Mapping):
prov = parse_lease_provenance(lease)
if not prov.get("cross_role_handoff"):
return None
handoff_status = str(prov.get("handoff_status") or "pending").strip().lower()
if handoff_status != "pending":
return None
adopted_by = (
lease.get("adopted_by_session_id")
or prov.get("adopted_by_session_id")
or ""
)
if str(adopted_by).strip():
return None
required_role = str(
prov.get("required_role") or lease.get("role") or ""
).strip().lower()
if not required_role:
return None
return {
"cross_role_handoff": True,
"handoff_status": "pending",
"required_role": required_role,
"allocating_session_id": str(
prov.get("allocating_session_id") or lease.get("session_id") or ""
),
"allocating_role": str(prov.get("allocating_role") or "controller"),
"lease_id": str(lease.get("lease_id") or ""),
"assignment_id": (
str(state["assignment"]["assignment_id"])
if isinstance(state.get("assignment"), Mapping)
and state["assignment"].get("assignment_id")
else None
),
"provenance": prov,
}
def adopt_lease(
db: cpd.ControlPlaneDB,
*,
@@ -564,17 +450,8 @@ def adopt_lease(
expected_head_sha: str | None = None,
owner_pid: int | None = None,
operator_authorized: bool = False,
adopter_profile_name: str | None = None,
adopter_namespace: str | None = None,
) -> dict[str, Any]:
"""Sanctioned adopt path with provenance; never silent foreign steal.
#843 F1: for a pending cross-role handoff, ``role`` must be the
authoritative profile-derived role supplied by the MCP boundary never
caller-asserted authority. When the handoff provenance declares
``required_profile`` / ``required_namespace`` and the caller context is
provided, both are validated exactly; a mismatch fails closed.
"""
"""Sanctioned adopt path with provenance; never silent foreign steal."""
state = db.get_lease_workflow_state(lease_id)
if not state:
raise LeaseLifecycleError(
@@ -586,8 +463,11 @@ def adopt_lease(
owner = str(lease.get("session_id") or "")
same_owner = owner == str(adopter_session_id)
handoff = is_pending_cross_role_handoff(state)
adopter_role = (role or "").strip().lower()
if freshness["freshness"] == "active" and not same_owner:
raise LeaseLifecycleError(
f"refusing to steal active foreign lease {lease_id} owned by "
f"{owner} (fail closed)"
)
if freshness["freshness"] in ("abandoned", "released"):
raise LeaseLifecycleError(
@@ -595,64 +475,13 @@ def adopt_lease(
"(fail closed)"
)
if handoff and not same_owner:
# Terminal statuses already rejected above. Freshness may be
# active OR stale_dead_process (controller exited) — both are
# consumable without abandonment when handoff is still pending.
if freshness["freshness"] not in (
"active",
"stale_dead_process",
"stale_missing_worktree",
):
raise LeaseLifecycleError(
f"lease {lease_id} freshness={freshness['freshness']}; "
"terminal or non-active allocation cannot be handoff-consumed "
"(fail closed)"
)
required = handoff["required_role"]
if adopter_role != required:
raise LeaseLifecycleError(
f"wrong role for cross-role handoff consume of {lease_id}: "
f"required={required} adopter={adopter_role or 'none'} "
"(fail closed)"
)
# #843 F1: provenance profile/namespace restrictions are validated
# against the authoritative caller context when declared. Caller
# input can never widen authority; a mismatch fails closed.
handoff_prov = handoff.get("provenance") or {}
required_profile = str(
handoff_prov.get("required_profile") or ""
).strip()
if required_profile and adopter_profile_name is not None:
if str(adopter_profile_name).strip() != required_profile:
raise LeaseLifecycleError(
f"wrong profile for cross-role handoff consume of "
f"{lease_id}: required_profile={required_profile} "
f"adopter_profile={adopter_profile_name} (fail closed)"
)
required_namespace = str(
handoff_prov.get("required_namespace") or ""
).strip()
if required_namespace and adopter_namespace is not None:
if str(adopter_namespace).strip() != required_namespace:
raise LeaseLifecycleError(
f"wrong namespace for cross-role handoff consume of "
f"{lease_id}: required_namespace={required_namespace} "
f"adopter_namespace={adopter_namespace} (fail closed)"
)
reason = "cross-role-handoff-consume"
elif freshness["freshness"] == "active" and not same_owner:
raise LeaseLifecycleError(
f"refusing to steal active foreign lease {lease_id} owned by "
f"{owner} (fail closed)"
)
elif not same_owner and freshness["freshness"] in (
# Expired or stale: require abandon-style safety before ownership transfer
# when not same owner; same owner may reclaim.
if not same_owner and freshness["freshness"] in (
"expired",
"stale_dead_process",
"stale_missing_worktree",
):
# Expired or stale (non-handoff): require abandon-style safety before
# ownership transfer when not same owner; same owner may reclaim.
if not operator_authorized and freshness["freshness"] == "expired":
# Deterministic reclaim of expired foreign lease is allowed
# without operator flag (sanctioned expire reclaim).
@@ -663,9 +492,6 @@ def adopt_lease(
f"lease {lease_id} freshness={freshness['freshness']}; "
"use abandon with proof before foreign adopt (fail closed)"
)
reason = "sanctioned-reclaim-adopt"
else:
reason = "owner-resume-adopt" if same_owner else "sanctioned-reclaim-adopt"
provenance = build_adopt_provenance(
adopted_from_session_id=owner,
@@ -678,14 +504,10 @@ def adopt_lease(
worktree_path=worktree_path,
expected_head_sha=expected_head_sha or lease.get("expected_head_sha"),
prior_lease_id=lease_id,
reason=reason,
reason=(
"owner-resume-adopt" if same_owner else "sanctioned-reclaim-adopt"
),
)
if handoff and not same_owner:
provenance["cross_role_handoff"] = True
provenance["handoff_status"] = "adopted"
provenance["required_role"] = handoff["required_role"]
provenance["allocating_session_id"] = handoff["allocating_session_id"]
provenance["allocating_role"] = handoff["allocating_role"]
result = db.adopt_lease(
lease_id=lease_id,
@@ -696,7 +518,7 @@ def adopt_lease(
owner_pid=owner_pid if owner_pid is not None else os.getpid(),
provenance=provenance,
)
out = {
return {
"success": True,
"outcome": result.get("outcome"),
"same_owner": same_owner,
@@ -709,24 +531,6 @@ def adopt_lease(
"comment_lease_only": False,
"reasons": result.get("reasons") or [],
}
if handoff and not same_owner:
out["cross_role_handoff"] = True
out["handoff_status"] = "adopted"
out["required_role"] = handoff["required_role"]
out["adopted_by_session_id"] = adopter_session_id
out["adopted_from_session_id"] = owner
lease_row = result.get("lease") or {}
if isinstance(lease_row, Mapping):
out["read_after_write"] = {
"lease_id": lease_row.get("lease_id"),
"session_id": lease_row.get("session_id"),
"role": lease_row.get("role"),
"status": lease_row.get("status"),
"adopted_by_session_id": lease_row.get("adopted_by_session_id"),
"adopted_from_session_id": lease_row.get("adopted_from_session_id"),
"phase": lease_row.get("phase"),
}
return out
def release_lease(
-212
View File
@@ -1,212 +0,0 @@
"""Central lease policy configuration (#790 Slice A, AC-N7).
The single authoritative source for every lease duration in the project. Before
this module the numbers were scattered: a four-hour author TTL was declared
twice (``issue_lock_store`` and ``gitea_mcp_server``), the reviewer/merger
sliding window lived in ``reviewer_pr_lease``, the conflict-fix window in
``pr_work_lease``, and the control-plane default in ``control_plane_db``.
Nothing tied them together, so tuning one class silently diverged from the
others and no reader could answer "how long does a lease live?" without
grepping four files.
AC-N7 requires that this configuration exist *before* the first heartbeat and
TTL behavior that reads from it, so it ships in Slice A rather than trailing the
code it governs.
Deliberate boundaries:
* **Declaration is not rewiring.** Every task class is declared here, but only
those with ``heartbeat_lifecycle_active`` were migrated onto the shared
heartbeat lifecycle in Slice A currently ``author_issue_work`` alone.
Reviewer, merger, and conflict-fix leases keep their own existing behavior
until Slice C moves them; their numbers are recorded here so the two cannot
drift apart unnoticed, and ``tests/test_issue_790_lease_policy.py`` asserts
the recorded values still equal the constants those modules use.
* **No policy decision lives here.** This module answers "how long", never "may
this session proceed". Freshness, reclaim, and renewal dispositions stay in
``issue_lock_store``.
"""
from __future__ import annotations
import os
from dataclasses import dataclass
from typing import Any
# Task classes. Only the first is migrated onto the shared lifecycle in Slice A.
TASK_CLASS_AUTHOR_ISSUE_WORK = "author_issue_work"
TASK_CLASS_REVIEWER_PR = "reviewer_pr"
TASK_CLASS_MERGER_PR = "merger_pr"
TASK_CLASS_CONFLICT_FIX = "conflict_fix"
# Durable marker for a lease minted under the shared heartbeat lifecycle.
#
# #790 AC-N8: this explicit marker — never a timestamp comparison — is what
# distinguishes a heartbeat-lifecycle lease from a legacy one. A lock written
# before this lifecycle existed carries no marker and reads as
# ``LIFECYCLE_LEGACY``.
LIFECYCLE_HEARTBEAT_V1 = "heartbeat-v1"
LIFECYCLE_LEGACY = "legacy"
_ENV_PREFIX = "GITEA_LEASE_POLICY"
@dataclass(frozen=True)
class LeasePolicy:
"""Durations governing one task class.
All intervals are minutes except ``absolute_cap_hours``. ``None`` for the
cap means the class has no maximum continuous duration.
"""
task_class: str
initial_ttl_minutes: float
heartbeat_cadence_minutes: float
stale_warning_minutes: float
missed_heartbeat_grace_minutes: float
absolute_cap_hours: float | None
recovery_grace_minutes: float
terminal_race_drain_minutes: float
terminal_retirement_eligible: bool
heartbeat_lifecycle_active: bool
# Defaults. ``author_issue_work`` adopts the reviewer window proven by #747
# rather than inventing new numbers: a lease expires 10 minutes after its last
# valid heartbeat, warns at half that, and an actively heartbeating session is
# never evicted. The prior value was a fixed four hours (240 minutes) that no
# heartbeat could shorten — the defect this issue exists to correct.
_DEFAULTS: dict[str, LeasePolicy] = {
TASK_CLASS_AUTHOR_ISSUE_WORK: LeasePolicy(
task_class=TASK_CLASS_AUTHOR_ISSUE_WORK,
initial_ttl_minutes=10.0,
heartbeat_cadence_minutes=2.0,
stale_warning_minutes=5.0,
missed_heartbeat_grace_minutes=10.0,
absolute_cap_hours=8.0,
recovery_grace_minutes=10.0,
terminal_race_drain_minutes=2.0,
terminal_retirement_eligible=True,
heartbeat_lifecycle_active=True,
),
# Declared, not rewired. These mirror reviewer_pr_lease.LEASE_TTL_MINUTES
# and STALE_WARNING_MINUTES; Slice C migrates the call sites.
TASK_CLASS_REVIEWER_PR: LeasePolicy(
task_class=TASK_CLASS_REVIEWER_PR,
initial_ttl_minutes=10.0,
heartbeat_cadence_minutes=2.0,
stale_warning_minutes=5.0,
missed_heartbeat_grace_minutes=10.0,
absolute_cap_hours=None,
recovery_grace_minutes=10.0,
terminal_race_drain_minutes=2.0,
terminal_retirement_eligible=False,
heartbeat_lifecycle_active=False,
),
TASK_CLASS_MERGER_PR: LeasePolicy(
task_class=TASK_CLASS_MERGER_PR,
initial_ttl_minutes=10.0,
heartbeat_cadence_minutes=2.0,
stale_warning_minutes=5.0,
missed_heartbeat_grace_minutes=10.0,
absolute_cap_hours=None,
recovery_grace_minutes=10.0,
terminal_race_drain_minutes=2.0,
terminal_retirement_eligible=False,
heartbeat_lifecycle_active=False,
),
# Mirrors pr_work_lease.DEFAULT_CONFLICT_FIX_TTL_MINUTES. Deliberately left
# at its current window; shortening it is Slice C's call, not this slice's.
TASK_CLASS_CONFLICT_FIX: LeasePolicy(
task_class=TASK_CLASS_CONFLICT_FIX,
initial_ttl_minutes=120.0,
heartbeat_cadence_minutes=2.0,
stale_warning_minutes=5.0,
missed_heartbeat_grace_minutes=10.0,
absolute_cap_hours=None,
recovery_grace_minutes=10.0,
terminal_race_drain_minutes=2.0,
terminal_retirement_eligible=False,
heartbeat_lifecycle_active=False,
),
}
_NUMERIC_FIELDS = (
"initial_ttl_minutes",
"heartbeat_cadence_minutes",
"stale_warning_minutes",
"missed_heartbeat_grace_minutes",
"absolute_cap_hours",
"recovery_grace_minutes",
"terminal_race_drain_minutes",
)
def env_var_name(task_class: str, field: str) -> str:
"""Environment variable that overrides one field of one task class."""
return f"{_ENV_PREFIX}_{task_class.upper()}_{field.upper()}"
def _override(task_class: str, field: str, default: float | None) -> float | None:
"""Read one override, falling back to *default* on anything unusable.
A malformed or non-positive override is ignored rather than raised: a typo
in an environment variable must not be able to mint a zero-length lease that
makes every claim instantly reclaimable, nor crash the server at import.
"""
raw = (os.environ.get(env_var_name(task_class, field)) or "").strip()
if not raw:
return default
try:
value = float(raw)
except (TypeError, ValueError):
return default
if value <= 0:
return default
return value
def policy_for(task_class: str) -> LeasePolicy:
"""Return the effective policy for *task_class*.
Unknown task classes fall back to the author policy, which is the most
conservative migrated class, rather than raising a new caller must never
be able to crash a lock write by naming a class this table has not learned.
"""
key = str(task_class or "").strip() or TASK_CLASS_AUTHOR_ISSUE_WORK
base = _DEFAULTS.get(key) or _DEFAULTS[TASK_CLASS_AUTHOR_ISSUE_WORK]
resolved = {
field: _override(base.task_class, field, getattr(base, field))
for field in _NUMERIC_FIELDS
}
if all(resolved[field] == getattr(base, field) for field in _NUMERIC_FIELDS):
return base
return LeasePolicy(
task_class=base.task_class,
terminal_retirement_eligible=base.terminal_retirement_eligible,
heartbeat_lifecycle_active=base.heartbeat_lifecycle_active,
**resolved,
)
def known_task_classes() -> tuple[str, ...]:
"""Every declared task class, migrated or not."""
return tuple(_DEFAULTS)
def describe(task_class: str) -> dict[str, Any]:
"""Serializable view of a policy, for audit records and tool payloads."""
policy = policy_for(task_class)
return {
"task_class": policy.task_class,
"initial_ttl_minutes": policy.initial_ttl_minutes,
"heartbeat_cadence_minutes": policy.heartbeat_cadence_minutes,
"stale_warning_minutes": policy.stale_warning_minutes,
"missed_heartbeat_grace_minutes": policy.missed_heartbeat_grace_minutes,
"absolute_cap_hours": policy.absolute_cap_hours,
"recovery_grace_minutes": policy.recovery_grace_minutes,
"terminal_race_drain_minutes": policy.terminal_race_drain_minutes,
"terminal_retirement_eligible": policy.terminal_retirement_eligible,
"heartbeat_lifecycle_active": policy.heartbeat_lifecycle_active,
"lifecycle_version": LIFECYCLE_HEARTBEAT_V1,
}
+11 -60
View File
@@ -378,11 +378,6 @@ def format_parity(assessment: dict) -> str:
# feeds the mutation gate and never changes startup_head/current_head.
# ---------------------------------------------------------------------------
# Legacy hardcoded target tracking ref. Retained for callers that still pass an
# explicit ref, but no longer the default: assuming a remote named ``origin``
# read an unrelated (often orphaned) remote-tracking ref in any checkout whose
# remote is named something else, and reported the target stale against a commit
# from a remote that may no longer even be configured (#983).
DEFAULT_TARGET_TRACKING_REF = "refs/remotes/origin/master"
@@ -413,7 +408,7 @@ def assess_target_repository_parity(
*,
canonical_root: str | None,
source: str | None,
tracking_ref: str | None = None,
tracking_ref: str = DEFAULT_TARGET_TRACKING_REF,
) -> dict:
"""Assess the configured cross-repository target checkout.
@@ -426,11 +421,6 @@ def assess_target_repository_parity(
An unconfigured namespace is ``configured=False`` and never ``stale`` the
single-repository default has no second dimension to be stale about. A
configured root that cannot be read is ``determinable=False`` with reasons.
*tracking_ref* defaults to None, meaning **derive the target from the
repository itself** through the same resolver the mutation guard uses, so
gating and reporting can never disagree about which ref is authoritative
(#983). Passing an explicit ref preserves the previous behaviour verbatim.
"""
result = {
"configured": bool(canonical_root),
@@ -439,13 +429,9 @@ def assess_target_repository_parity(
"repository_slug": None,
"checkout_head": None,
"tracking_ref": tracking_ref,
"base_remote": None,
"base_branch": None,
"tracking_ref_source": "explicit_tracking_ref" if tracking_ref else None,
"remote_tracking_head": None,
"determinable": False,
"stale": False,
"reason_code": None,
"reasons": [],
}
if not canonical_root:
@@ -477,53 +463,18 @@ def assess_target_repository_parity(
result["checkout_head"] = head
result["determinable"] = True
# Local import keeps this module dependency-light for its startup role.
import canonical_repository_root as _crr
remote_url = _git_capture(canonical_root, "remote", "get-url", "origin")
if remote_url:
# Local import keeps this module dependency-light for its startup role.
import remote_repo_guard
# #983: identity comes from whichever remote actually proves it, not from a
# remote assumed to be named 'origin'. In a checkout whose only remote is
# 'prgs', the old lookup failed outright and reported the identity as
# underivable while a leftover refs/remotes/origin/master still resolved.
#
# #983 B2: this must be the *ambiguity-aware* resolver. The first-wins
# `resolve_identity_remote` picks whichever remote probes first, so on a
# target where distinct remotes claim different repositories the report
# confidently named one of them while the mutation gate refused the same
# target — gating and reporting evaluating different repositories, which is
# precisely the divergence this issue exists to end.
identity = _crr.assess_identity_remote(canonical_root)
result["repository_slug"] = identity["slug"]
if not identity["slug"]:
result["reason_code"] = identity["reason_code"]
result["reasons"].extend(
identity["reasons"]
or ["target repository identity could not be derived from its git remote"]
parsed = remote_repo_guard.parse_org_repo_from_remote_url(remote_url)
if parsed:
result["repository_slug"] = f"{parsed[0]}/{parsed[1]}"
if not result["repository_slug"]:
result["reasons"].append(
"target repository identity could not be derived from its git remote"
)
if identity["ambiguous"]:
# An ambiguous target has no single authoritative base, so reporting one
# would be a guess. Fail closed here exactly as the gate does.
return result
# Identity is otherwise resolved independently of the base ref: a target that
# has never been fetched still has a provable repository identity, and
# reporting it as unidentifiable would lose real information over an
# unrelated missing ref.
if not tracking_ref:
# No remote argument. The identity remote resolved just above was
# *inferred here*, and feeding it back in would tell the resolver a
# caller had explicitly disambiguated the repository, suppressing its
# ambiguity gate (#983 B2). Only an operator-supplied remote may do that,
# and this call site has none.
base = _crr.resolve_target_base_ref(canonical_root)
if not base.get("proven"):
result["reason_code"] = base.get("reason_code")
result["reasons"].extend(base.get("reasons") or [])
return result
tracking_ref = base["tracking_ref"]
result["tracking_ref"] = tracking_ref
result["tracking_ref_source"] = base.get("source")
result["base_remote"] = base.get("remote")
result["base_branch"] = base.get("branch")
tracking_head = _git_capture(canonical_root, "rev-parse", tracking_ref)
if not tracking_head:
-393
View File
@@ -1,393 +0,0 @@
"""Production application launcher for multi-namespace MCP fleets (#978 B1).
One real LLM application launch mints exactly one trusted
``client_instance_id`` and propagates it to every Gitea MCP namespace worker
started for that launch. Workers never invent a trusted instance identity from
PID proximity, timestamps, or ordinary untrusted environment values.
This module is the production serve-path authority for instance identity.
Tests and fixtures may call the same functions, but production registration
receives the identity from the env this launcher builds not from a hand-set
test-only helper that bypasses it.
"""
from __future__ import annotations
import os
import secrets
from typing import Any, Mapping
import mcp_fleet_snapshot as fleet
import mcp_worker_identity as mwi
#: Canonical MCP server names for the five role namespaces.
NAMESPACE_SERVER_NAMES: tuple[str, ...] = (
"gitea-author",
"gitea-reviewer",
"gitea-merger",
"gitea-controller",
"gitea-reconciler",
)
#: Map MCP server name → role/namespace kind.
SERVER_TO_NAMESPACE: dict[str, str] = {
"gitea-author": "author",
"gitea-reviewer": "reviewer",
"gitea-merger": "merger",
"gitea-controller": "controller",
"gitea-reconciler": "reconciler",
}
SANCTIONED_NAMESPACES: tuple[str, ...] = (
"author",
"reviewer",
"merger",
"controller",
"reconciler",
)
# Env the launcher may set on every worker of one application launch.
CLIENT_NAME_ENV = "GITEA_MCP_CLIENT"
CLIENT_INSTANCE_ENV = fleet.CLIENT_INSTANCE_ENV
FLEET_RUN_ENV = fleet.FLEET_RUN_ENV
CLIENT_SESSION_ENV = "GITEA_MCP_CLIENT_SESSION"
CLIENT_MANAGED_ENV = "GITEA_CLIENT_MANAGED"
PROFILE_ENV = "GITEA_MCP_PROFILE"
CONFIG_ENV = "GITEA_MCP_CONFIG"
INSTANCE_PROVENANCE_ENV = "GITEA_MCP_INSTANCE_PROVENANCE"
WORKER_IDENTITY_ENV = "GITEA_MCP_WORKER_IDENTITY"
GENERATION_ID_ENV = "GITEA_MCP_GENERATION_ID"
# Marker the trusted launcher alone writes; ordinary user env without this
# marker is never classified as launcher-trusted provenance.
LAUNCHER_PROVENANCE_VALUE = fleet.INSTANCE_ID_PROVENANCE_TRUSTED
def mint_application_launch(
client_type: str | None,
*,
launch_nonce: str | None = None,
fleet_run_id: str | None = None,
session_id: str | None = None,
now=None,
) -> dict[str, Any]:
"""Mint one trusted application-instance identity for a production launch.
Called exactly once per real application launch. The returned
``client_instance_id`` is injected into every namespace worker environment
for that launch. A second call (separate launch) yields a different ID.
"""
client = mwi.normalize_client_name(client_type)
instance_id = fleet.generate_client_instance_id(
client, launch_nonce=launch_nonce, now=now
)
assessment = fleet.assess_instance_identity(instance_id)
if not assessment["trusted"]:
# generate_client_instance_id always produces a trusted format; fail
# closed if that invariant ever breaks rather than shipping untrusted.
raise RuntimeError(
f"launcher produced untrusted client_instance_id {instance_id!r}: "
f"{assessment.get('reasons')}"
)
session = (session_id or "").strip() or f"launch-{secrets.token_hex(12)}"
return {
"client_type": client,
"client_instance_id": instance_id,
"instance_id_provenance": LAUNCHER_PROVENANCE_VALUE,
"instance_identity_trusted": True,
"fleet_run_id": (fleet_run_id or "").strip() or None,
"session_id": session,
"namespaces": list(SANCTIONED_NAMESPACES),
"namespace_server_names": list(NAMESPACE_SERVER_NAMES),
}
def namespace_worker_env(
*,
profile_name: str,
client_type: str | None,
client_instance_id: str,
config_path: str | None = None,
fleet_run_id: str | None = None,
session_id: str | None = None,
extra_env: Mapping[str, str] | None = None,
) -> dict[str, str]:
"""Build the environment for one namespace worker of a trusted launch.
Never invents a client_instance_id. The caller must supply the launch-minted
identity so all five workers receive the same value.
"""
assessment = fleet.assess_instance_identity(client_instance_id)
if not assessment["trusted"]:
raise ValueError(
"namespace_worker_env refuses untrusted client_instance_id "
f"{client_instance_id!r}: {assessment.get('reasons')}"
)
client = mwi.normalize_client_name(client_type)
env: dict[str, str] = {
PROFILE_ENV: str(profile_name),
CLIENT_MANAGED_ENV: "1",
"GITEA_MCP_CLIENT": client,
CLIENT_INSTANCE_ENV: assessment["client_instance_id"],
INSTANCE_PROVENANCE_ENV: LAUNCHER_PROVENANCE_VALUE,
}
if config_path:
env[CONFIG_ENV] = str(config_path)
if fleet_run_id:
env[FLEET_RUN_ENV] = str(fleet_run_id)
if session_id:
env[CLIENT_SESSION_ENV] = str(session_id)
if extra_env:
# Never let untrusted callers override the trusted instance keys.
protected = {
CLIENT_INSTANCE_ENV,
INSTANCE_PROVENANCE_ENV,
"GITEA_MCP_CLIENT",
CLIENT_MANAGED_ENV,
}
for key, value in extra_env.items():
if key in protected:
continue
env[str(key)] = str(value)
return env
def build_application_mcp_servers(
profile_by_namespace: Mapping[str, str],
*,
client_type: str | None,
config_path: str | None = None,
client_instance_id: str | None = None,
fleet_run_id: str | None = None,
session_id: str | None = None,
launch_nonce: str | None = None,
command: str | None = None,
args: list[str] | None = None,
now=None,
) -> dict[str, Any]:
"""Build a production ``mcpServers`` map for one application launch.
Mints one ``client_instance_id`` when *client_instance_id* is omitted (fresh
launch). When the caller supplies a previously minted trusted ID (resume of
the same launch / config rewrite), that ID is reused so reconnect keeps
attribution. A full new application restart omits the ID and receives a
fresh mint.
Every namespace server entry receives the **same** instance ID. Separate
calls with no supplied ID receive distinct IDs.
"""
missing = [
ns for ns in SANCTIONED_NAMESPACES if not profile_by_namespace.get(ns)
]
if missing:
raise ValueError(
"build_application_mcp_servers requires a profile for every "
f"sanctioned namespace; missing: {missing}"
)
if client_instance_id is None:
launch = mint_application_launch(
client_type,
launch_nonce=launch_nonce,
fleet_run_id=fleet_run_id,
session_id=session_id,
now=now,
)
else:
assessment = fleet.assess_instance_identity(client_instance_id)
if not assessment["trusted"]:
raise ValueError(
"refusing to propagate untrusted client_instance_id "
f"{client_instance_id!r} into production launch envs: "
f"{assessment.get('reasons')}"
)
launch = {
"client_type": mwi.normalize_client_name(client_type),
"client_instance_id": assessment["client_instance_id"],
"instance_id_provenance": LAUNCHER_PROVENANCE_VALUE,
"instance_identity_trusted": True,
"fleet_run_id": (fleet_run_id or "").strip() or None,
"session_id": (session_id or "").strip()
or f"launch-{secrets.token_hex(12)}",
"namespaces": list(SANCTIONED_NAMESPACES),
"namespace_server_names": list(NAMESPACE_SERVER_NAMES),
}
# Resolve command/args from the production server entry when not provided.
if command is None or args is None:
import gitea_config
cmd, cmd_args = gitea_config.server_command()
command = command or cmd
args = args if args is not None else list(cmd_args)
servers: dict[str, Any] = {}
shared_id = launch["client_instance_id"]
for namespace in SANCTIONED_NAMESPACES:
server_name = f"gitea-{namespace}"
profile = profile_by_namespace[namespace]
env = namespace_worker_env(
profile_name=profile,
client_type=launch["client_type"],
client_instance_id=shared_id,
config_path=config_path,
fleet_run_id=launch.get("fleet_run_id"),
session_id=launch.get("session_id"),
)
servers[server_name] = {
"command": command,
"args": list(args),
"env": env,
}
return {
"mcpServers": servers,
"launch": launch,
"client_instance_id": shared_id,
"client_type": launch["client_type"],
"namespaces": list(SANCTIONED_NAMESPACES),
"shared_instance_id_across_namespaces": True,
"namespace_count": len(SANCTIONED_NAMESPACES),
}
def launcher_entry_for_profile(
profile_name: str,
*,
client_type: str | None = None,
config_path: str | None = None,
client_instance_id: str | None = None,
fleet_run_id: str | None = None,
session_id: str | None = None,
server_key: str = "gitea-tools",
launch_nonce: str | None = None,
now=None,
) -> dict[str, Any]:
"""Thin single-server production launcher entry with trusted instance ID.
Used when only one namespace is being configured. Still mints (or reuses)
a trusted ``client_instance_id`` so production never relies on the legacy
placeholder identity for normal launches.
"""
import gitea_config
if client_instance_id is None:
launch = mint_application_launch(
client_type or "unknown",
launch_nonce=launch_nonce,
fleet_run_id=fleet_run_id,
session_id=session_id,
now=now,
)
client_instance_id = launch["client_instance_id"]
client = launch["client_type"]
fleet_run = launch.get("fleet_run_id")
session = launch.get("session_id")
else:
assessment = fleet.assess_instance_identity(client_instance_id)
if not assessment["trusted"]:
raise ValueError(
f"untrusted client_instance_id {client_instance_id!r}"
)
client = mwi.normalize_client_name(client_type)
fleet_run = (fleet_run_id or "").strip() or None
session = (session_id or "").strip() or None
client_instance_id = assessment["client_instance_id"]
command, args = gitea_config.server_command()
env = namespace_worker_env(
profile_name=profile_name,
client_type=client,
client_instance_id=client_instance_id,
config_path=config_path or gitea_config.DEFAULT_CONFIG_PATH,
fleet_run_id=fleet_run,
session_id=session,
)
return {
server_key: {
"command": command,
"args": args,
"env": env,
},
"client_instance_id": client_instance_id,
"client_type": client,
}
def collect_instance_ids_from_mcp_servers(
mcp_servers: Mapping[str, Any],
) -> dict[str, Any]:
"""Inspect a production mcpServers map for shared instance attribution.
Returns the unique set of client_instance_id values across gitea-* servers
and whether all five namespaces share exactly one trusted ID.
"""
ids: list[str] = []
by_server: dict[str, str | None] = {}
for name in NAMESPACE_SERVER_NAMES:
entry = mcp_servers.get(name) or {}
env = entry.get("env") or {}
raw = (env.get(CLIENT_INSTANCE_ENV) or "").strip() or None
by_server[name] = raw
if raw:
ids.append(raw)
unique = sorted(set(ids))
trusted = [
i
for i in unique
if fleet.assess_instance_identity(i)["trusted"]
]
return {
"server_instance_ids": by_server,
"unique_instance_ids": unique,
"trusted_instance_ids": trusted,
"shared_single_trusted_id": (
len(unique) == 1
and len(trusted) == 1
and all(by_server.get(n) == unique[0] for n in NAMESPACE_SERVER_NAMES)
),
"namespace_server_count": sum(
1 for n in NAMESPACE_SERVER_NAMES if n in mcp_servers
),
}
def inherit_or_refuse_client_instance(
env: Mapping[str, str] | None = None,
) -> dict[str, Any]:
"""Resolve instance identity for a worker process at serve time.
Production workers inherit the launcher-issued ID. They never mint a
trusted ID themselves. Missing / legacy / malformed values fail soft into
an untrusted assessment so registration can still record a diagnostic row
without authorizing multi-instance fleet mutation.
"""
source = dict(env if env is not None else os.environ)
raw = (source.get(CLIENT_INSTANCE_ENV) or "").strip() or None
provenance_marker = (source.get(INSTANCE_PROVENANCE_ENV) or "").strip()
assessment = fleet.assess_instance_identity(raw)
# Ordinary user-supplied values without launcher provenance marker are
# still format-checked by assess_instance_identity. When the format is
# trusted but the launcher marker is absent, keep the ID but note that
# provenance is not launcher-sealed (operator hand-set or legacy config).
if assessment["trusted"] and provenance_marker != LAUNCHER_PROVENANCE_VALUE:
assessment = dict(assessment)
assessment["launcher_sealed"] = False
assessment["reasons"] = list(assessment.get("reasons") or []) + [
f"{INSTANCE_PROVENANCE_ENV} is not {LAUNCHER_PROVENANCE_VALUE!r}; "
"identity format is valid but not sealed by the production launcher"
]
else:
assessment = dict(assessment)
assessment["launcher_sealed"] = bool(
assessment["trusted"]
and provenance_marker == LAUNCHER_PROVENANCE_VALUE
)
assessment["fleet_run_id"] = (source.get(FLEET_RUN_ENV) or "").strip() or None
assessment["session_id"] = (
(source.get(CLIENT_SESSION_ENV) or "").strip() or None
)
assessment["client_type"] = mwi.normalize_client_name(
(source.get("GITEA_MCP_CLIENT") or "").strip() or None
)
return assessment
-340
View File
@@ -1,340 +0,0 @@
"""Sanctioned MCP client reconnect request surface for Codex/LLM sessions (#678).
Codex and other agent hosts can detect stale or closed Gitea MCP runtimes, but
the host owns the transport. This module never restarts, kills, or reloads a
daemon. It builds:
1. A **callable reconnect request** result agents can invoke via
``gitea_request_mcp_reconnect`` (report-only, side-effect free).
2. A **typed blocker** with exact operator UI steps when recovery must be
performed by the host/operator.
Forbidden recovery paths (must never be recommended):
* ``pkill`` / ``kill`` / ``killall`` of MCP daemons
* ``touch`` / mtime config reload hacks
* ``.env`` or MCP config edits as recovery
* session-state file edits
* raw Gitea API / direct server-import fallbacks
After the operator reconnects, workflows restart from identity / runtime /
capability preflight (``gitea_whoami`` ``gitea_resolve_task_capability``
task).
"""
from __future__ import annotations
from typing import Any, Mapping
# --- Reason vocabulary -------------------------------------------------------
REASON_STALE_RUNTIME = "stale-runtime"
REASON_TRANSPORT_EOF = "transport_eof"
REASON_MISSING_NAMESPACE = "missing_namespace"
REASON_NOT_REQUIRED = "not_required"
REASON_UNSPECIFIED = "unspecified"
VALID_REASONS = frozenset(
{
REASON_STALE_RUNTIME,
REASON_TRANSPORT_EOF,
REASON_MISSING_NAMESPACE,
REASON_NOT_REQUIRED,
REASON_UNSPECIFIED,
}
)
# Boundary statuses reported to callers (match review_workflow_boundary style).
BOUNDARY_CLEAN = "clean"
BOUNDARY_MISMATCH = "mismatch"
BOUNDARY_STALE = "stale"
BOUNDARY_UNKNOWN = "unknown"
# Typed blocker kinds
BLOCKER_OPERATOR_RECONNECT = "operator_mcp_reconnect_required"
BLOCKER_NONE = "none"
FORBIDDEN_RECOVERY_PATHS: tuple[str, ...] = (
"pkill / kill / killall of mcp_server.py, gitea_mcp_server, or broad python sweeps",
"touch / mtime-based MCP config reload hacks",
".env edits as recovery",
"MCP config file edits as recovery",
"session-state file edits as recovery",
"raw Gitea API or direct MCP server-import fallbacks",
)
# Client-specific operator UI steps. Keep Codex first (issue title surface).
OPERATOR_UI_STEPS: dict[str, tuple[str, ...]] = {
"codex": (
"In Codex, open the MCP / Developer tools panel for this workspace.",
"Locate the named Gitea MCP server entry (namespace) that needs reconnect "
"(e.g. gitea-author, gitea-reviewer, gitea-merger, gitea-tools, "
"gitea-controller, gitea-reconciler).",
"Click 'Reload Developer Tools' or the server reconnect/reload control "
"for that entry so the client spawns a fresh MCP subprocess.",
"If per-server reconnect is unavailable, fully restart the Codex client "
"(quit and relaunch) so all MCP namespaces reattach.",
"After reconnect, rerun the blocked workflow from preflight: "
"gitea_whoami → gitea_resolve_task_capability → the original task. "
"Do not resume mid-mutation.",
),
"claude_code": (
"Run `/mcp` (or open the MCP servers UI) in Claude Code.",
"Reconnect the affected gitea-* server entry so the client reopens stdio.",
"If reconnect fails, relaunch the Claude Code session entirely.",
"After reconnect, restart the workflow from gitea_whoami → "
"gitea_resolve_task_capability → task.",
),
"generic": (
"Use the host/IDE MCP reconnect or reload control for the named namespace.",
"If no per-namespace control exists, restart the MCP client/editor.",
"After reconnect, restart the workflow from identity/capability preflight.",
),
}
#: What an *unidentified* client gets. #948: this is deliberately the
#: host-agnostic step set rather than a specific product. Defaulting to one
#: vendor emitted Codex UI steps to a Gemini/Antigravity operator, who then had
#: no reachable recovery path — the guidance named a panel they do not have.
DEFAULT_CLIENT = "generic"
#: The historical default, kept addressable by name so Codex callers still get
#: Codex steps, without it silently becoming the fallback for unknown clients.
LEGACY_DEFAULT_CLIENT = "codex"
def normalize_reason(reason: str | None) -> str:
"""Map free-form reason strings onto the closed vocabulary."""
raw = (reason or "").strip().lower()
if not raw:
return REASON_UNSPECIFIED
if raw in VALID_REASONS:
return raw
text = raw.replace(" ", "_").replace("-", "_")
aliases = {
"stale_runtime": REASON_STALE_RUNTIME,
"staleruntime": REASON_STALE_RUNTIME,
"runtime_stale": REASON_STALE_RUNTIME,
"stale": REASON_STALE_RUNTIME,
"transport_eof": REASON_TRANSPORT_EOF,
"transport_closed": REASON_TRANSPORT_EOF,
"eof": REASON_TRANSPORT_EOF,
"client_is_closing": REASON_TRANSPORT_EOF,
"missing_namespace": REASON_MISSING_NAMESPACE,
"namespace_missing": REASON_MISSING_NAMESPACE,
"not_required": REASON_NOT_REQUIRED,
"healthy": REASON_NOT_REQUIRED,
"ok": REASON_NOT_REQUIRED,
"unspecified": REASON_UNSPECIFIED,
}
if text in aliases:
return aliases[text]
hyphenated = text.replace("_", "-")
if hyphenated in VALID_REASONS:
return hyphenated
return REASON_UNSPECIFIED
def normalize_client(client: str | None) -> str:
"""Return the UI-step key for a client.
#948: alias resolution is shared with ``mcp_worker_identity`` so a client
name means the same thing wherever it is read. A name we recognise but have
no bespoke steps for Gemini, Antigravity, Grok resolves to the generic
host-agnostic steps rather than to another vendor's panel.
"""
import mcp_worker_identity
canonical = mcp_worker_identity.normalize_client_name(client)
if canonical in OPERATOR_UI_STEPS:
return canonical
return DEFAULT_CLIENT
def classify_boundary_status(
*,
startup_sha: str | None,
current_master_sha: str | None,
live_stale: bool | None = None,
in_parity: bool | None = None,
) -> str:
"""Derive boundary_status from parity evidence."""
if live_stale is True or in_parity is False:
return BOUNDARY_STALE
start = (startup_sha or "").strip().lower()
current = (current_master_sha or "").strip().lower()
if start and current and start != current:
return BOUNDARY_MISMATCH
if start and current and start == current:
return BOUNDARY_CLEAN
if in_parity is True:
return BOUNDARY_CLEAN
return BOUNDARY_UNKNOWN
def operator_ui_steps(client: str | None, *, namespace: str | None = None) -> list[str]:
"""Exact operator UI steps for the named client."""
key = normalize_client(client)
steps = list(OPERATOR_UI_STEPS.get(key) or OPERATOR_UI_STEPS[DEFAULT_CLIENT])
ns = (namespace or "").strip()
if ns:
steps = [
s.replace("named Gitea MCP server entry (namespace)", f"namespace '{ns}'")
.replace("affected gitea-* server entry", f"server entry '{ns}'")
.replace("named namespace", f"namespace '{ns}'")
for s in steps
]
return steps
def build_reconnect_request(
*,
namespace: str,
profile: str | None = None,
pid: int | str | None = None,
session_id: str | None = None,
startup_sha: str | None = None,
current_master_sha: str | None = None,
boundary_status: str | None = None,
reason: str | None = None,
client: str | None = DEFAULT_CLIENT,
live_stale: bool | None = None,
in_parity: bool | None = None,
restart_required: bool | None = None,
stop_required: bool | None = None,
extra: Mapping[str, Any] | None = None,
) -> dict[str, Any]:
"""Build the structured reconnect-request / typed-blocker payload (#678).
Never mutates process, config, or session state. Always side-effect free.
"""
ns = (namespace or "").strip() or "unknown"
normalized_reason = normalize_reason(reason)
boundary = (boundary_status or "").strip() or classify_boundary_status(
startup_sha=startup_sha,
current_master_sha=current_master_sha,
live_stale=live_stale,
in_parity=in_parity,
)
reconnect_needed = True
if normalized_reason == REASON_NOT_REQUIRED and boundary == BOUNDARY_CLEAN:
reconnect_needed = False
if restart_required is False and stop_required is False and boundary == BOUNDARY_CLEAN:
# Explicit healthy probe
if normalized_reason in (REASON_NOT_REQUIRED, REASON_UNSPECIFIED):
reconnect_needed = False
normalized_reason = REASON_NOT_REQUIRED
if restart_required is True or stop_required is True:
reconnect_needed = True
if normalized_reason in (REASON_NOT_REQUIRED, REASON_UNSPECIFIED):
normalized_reason = REASON_STALE_RUNTIME
client_key = normalize_client(client)
steps = operator_ui_steps(client_key, namespace=ns)
result: dict[str, Any] = {
"success": True,
"read_only": True,
"reconnect_performed": False,
"mutation_performed": False,
"reconnect_needed": reconnect_needed,
"namespace": ns,
"profile": (profile or "").strip() or None,
"pid": pid,
"session_id": (session_id or "").strip() or None,
"startup_sha": (startup_sha or "").strip() or None,
"current_master_sha": (current_master_sha or "").strip() or None,
"boundary_status": boundary,
"reason": normalized_reason,
"client": client_key,
"forbidden_recovery_paths": list(FORBIDDEN_RECOVERY_PATHS),
"post_reconnect_preflight": [
"gitea_whoami",
"gitea_resolve_task_capability",
"original_task",
],
"exact_safe_next_action": None,
"blocker_kind": BLOCKER_NONE,
"operator_ui_steps": steps,
"typed_blocker": None,
}
if reconnect_needed:
result["blocker_kind"] = BLOCKER_OPERATOR_RECONNECT
result["stop_required"] = True
result["restart_required"] = True
result["exact_safe_next_action"] = (
f"blocker_kind={BLOCKER_OPERATOR_RECONNECT}: operator must reconnect "
f"MCP namespace '{ns}' via the host UI (client={client_key}). "
"Do not pkill, touch configs, edit session state, or use raw API. "
"After reconnect, restart from gitea_whoami → "
"gitea_resolve_task_capability → task."
)
result["typed_blocker"] = {
"blocker_kind": BLOCKER_OPERATOR_RECONNECT,
"namespaces": [ns],
"why_reconnect_required": normalized_reason,
"operator_ui_steps": steps,
"client": client_key,
"forbidden_recovery_paths": list(FORBIDDEN_RECOVERY_PATHS),
"instruction_after_reconnect": (
"Rerun the blocked workflow from preflight "
"(gitea_whoami → gitea_resolve_task_capability → task). "
"Do not continue mid-mutation from pre-reconnect state."
),
}
else:
result["stop_required"] = False
result["restart_required"] = False
result["exact_safe_next_action"] = (
f"Reconnect not required for namespace '{ns}' "
f"(boundary_status={boundary}). Proceed with the original task."
)
if extra:
for key, value in extra.items():
if key not in result:
result[key] = value
return result
def reasons_never_suggest_forbidden(text: str) -> bool:
"""Return True when *text* does not recommend a forbidden recovery path.
Mentions that *ban* a path (e.g. ``Do not pkill`` / ``never edit session
state``) are allowed. Positive recommendations such as ``use pkill`` or
``run killall`` fail.
"""
import re
lowered = (text or "").lower()
# Strip common ban prefixes so "do not pkill" does not trip positive checks.
scrubbed = re.sub(
r"\b(?:do not|don't|never|must not|forbid(?:den)?|ban(?:ned)?)\b"
r"[^.!;\n]{0,80}",
" ",
lowered,
)
# Positive imperative / advisory forms that would tell an agent to do harm.
positive_suggestions = (
"use pkill",
"run pkill",
"try pkill",
"pkill -f",
"use killall",
"run killall",
"killall mcp",
"use kill ",
"run kill ",
"touch the mcp",
"touch mcp config",
"utime(",
"edit the mcp config to recover",
"edit .env to recover",
"import gitea_mcp_server",
"python -c 'import gitea_mcp",
)
return not any(frag in scrubbed for frag in positive_suggestions)
-237
View File
@@ -1,237 +0,0 @@
"""Antigravity IDE vs Global MCP Config Drift Diagnostic (#672).
Diagnoses config drift between the active IDE MCP configuration
(e.g. ``~/.gemini/antigravity-ide/mcp_config.json``) and the offline/global
canonical configuration (e.g. ``~/.gemini/config/mcp_config.json``).
Hard rules (#672 / #630 / #655):
* Distinguish offline/global success from active IDE namespace availability.
* Never print tokens, DSNs, Authorization headers, or secret-bearing env vars.
* Sanctioned repair path is: backup active config -> patch active config from canonical
-> reconnect through IDE/client -> verify with live ``gitea_whoami``.
* FORBIDDEN: ``pkill``, mtime edits, source edits, or session-state edits for repair.
"""
from __future__ import annotations
import argparse
import json
import os
import sys
from datetime import datetime, timezone
from pathlib import Path
from typing import Any
from webui import console_redaction
DEFAULT_ACTIVE_IDE_CONFIG = "~/.gemini/antigravity-ide/mcp_config.json"
DEFAULT_GLOBAL_CONFIG = "~/.gemini/config/mcp_config.json"
REQUIRED_GITEA_ROLE_SERVERS = (
"gitea-author",
"gitea-reviewer",
"gitea-merger",
"gitea-reconciler",
"gitea-controller",
"gitea-tools",
)
SANCTIONED_REPAIR_RUNBOOK: tuple[str, ...] = (
"1. Backup active IDE config: cp ~/.gemini/antigravity-ide/mcp_config.json ~/.gemini/antigravity-ide/mcp_config.json.bak",
"2. Patch active IDE config: copy required missing Gitea role server entries from global config (~/.gemini/config/mcp_config.json) into active IDE config.",
"3. Reconnect via IDE/client UI or client restart (do NOT use host process kill).",
"4. Verify active namespace health using live gitea_whoami and gitea_resolve_task_capability on each role namespace.",
"FORBIDDEN REPAIR PATHS: pkill / host process kill, mtime touch edits, source code edits, or session-state edits.",
)
def resolve_config_path(path_str: str) -> Path:
"""Expand user and resolve absolute path."""
return Path(os.path.expanduser(path_str)).resolve()
def load_mcp_config(config_path: str | Path) -> tuple[dict[str, Any] | None, str | None]:
"""Load and parse JSON MCP configuration from file.
Returns (config_dict, error_message).
"""
resolved = resolve_config_path(str(config_path))
if not resolved.exists():
return None, f"file_not_found: {resolved}"
try:
with open(resolved, "r", encoding="utf-8") as f:
data = json.load(f)
if not isinstance(data, dict):
return None, f"invalid_schema: root is not a JSON object in {resolved}"
return data, None
except Exception as exc:
return None, f"unreadable_json: {exc} in {resolved}"
def extract_mcp_servers(config: dict[str, Any] | None) -> dict[str, dict[str, Any]]:
"""Extract the mcpServers or mcp_servers mapping safely."""
if not config:
return {}
servers = config.get("mcpServers") or config.get("mcp_servers") or {}
if isinstance(servers, dict):
return {str(k): v for k, v in servers.items() if isinstance(v, dict)}
return {}
def _safe_redact_server_config(srv_cfg: dict[str, Any]) -> dict[str, Any]:
"""Redact secrets from environment variables and command line args."""
safe = {}
if "command" in srv_cfg:
safe["command"] = str(srv_cfg["command"])
if "args" in srv_cfg and isinstance(srv_cfg["args"], list):
safe["args"] = [console_redaction.redact_text(str(a)) for a in srv_cfg["args"]]
if "env" in srv_cfg and isinstance(srv_cfg["env"], dict):
safe_env = {}
for k, v in srv_cfg["env"].items():
if any(secret_kw in k.lower() for secret_kw in ("token", "secret", "pass", "key", "auth")):
safe_env[k] = "[REDACTED]"
else:
safe_env[k] = console_redaction.redact_text(str(v))
safe["env"] = safe_env
return safe
def analyze_config_drift(
active_config_path: str = DEFAULT_ACTIVE_IDE_CONFIG,
global_config_path: str = DEFAULT_GLOBAL_CONFIG,
) -> dict[str, Any]:
"""Analyze MCP configuration drift between active IDE config and global config.
Returns structured diagnostic output.
"""
active_resolved = resolve_config_path(active_config_path)
global_resolved = resolve_config_path(global_config_path)
active_cfg, active_err = load_mcp_config(active_resolved)
global_cfg, global_err = load_mcp_config(global_resolved)
active_servers = extract_mcp_servers(active_cfg)
global_servers = extract_mcp_servers(global_cfg)
missing_role_servers: list[str] = []
present_role_servers: list[str] = []
profile_mismatches: list[dict[str, Any]] = []
reasons: list[str] = []
if active_err:
reasons.append(f"Active IDE config error: {active_err}")
if global_err:
reasons.append(f"Global canonical config error: {global_err}")
# Check Gitea role servers
for srv_name in REQUIRED_GITEA_ROLE_SERVERS:
in_active = srv_name in active_servers
in_global = srv_name in global_servers
if in_active:
present_role_servers.append(srv_name)
elif in_global:
missing_role_servers.append(srv_name)
reasons.append(
f"Missing Gitea role server '{srv_name}' in active IDE config ({active_resolved})"
)
if in_active and in_global:
# Compare profiles & environments
act_env = active_servers[srv_name].get("env", {}) if isinstance(active_servers[srv_name], dict) else {}
glo_env = global_servers[srv_name].get("env", {}) if isinstance(global_servers[srv_name], dict) else {}
act_prof = act_env.get("GITEA_MCP_PROFILE") or act_env.get("GITEA_PROFILE_NAME")
glo_prof = glo_env.get("GITEA_MCP_PROFILE") or glo_env.get("GITEA_PROFILE_NAME")
if act_prof != glo_prof:
mismatch_item = {
"server": srv_name,
"active_profile": act_prof,
"global_profile": glo_prof,
}
profile_mismatches.append(mismatch_item)
reasons.append(
f"Profile mismatch for '{srv_name}': active='{act_prof}' != global='{glo_prof}'"
)
in_sync = bool(
not active_err
and not global_err
and not missing_role_servers
and not profile_mismatches
)
report = {
"timestamp": datetime.now(timezone.utc).isoformat(),
"in_sync": in_sync,
"active_config_path": str(active_resolved),
"active_config_exists": active_cfg is not None,
"global_config_path": str(global_resolved),
"global_config_exists": global_cfg is not None,
"required_role_servers": list(REQUIRED_GITEA_ROLE_SERVERS),
"present_role_servers": present_role_servers,
"missing_role_servers": missing_role_servers,
"profile_mismatches": profile_mismatches,
"reasons": reasons,
"sanctioned_repair_runbook": list(SANCTIONED_REPAIR_RUNBOOK),
"forbidden_repair_methods": [
"pkill / host process kill",
"mtime touch edits",
"source code edits",
"session-state edits",
],
}
return console_redaction.redact_payload(report)
def main() -> None:
parser = argparse.ArgumentParser(
description="Diagnose Gitea MCP role server config drift between active IDE and global config."
)
parser.add_argument(
"--active-config",
default=DEFAULT_ACTIVE_IDE_CONFIG,
help="Path to active IDE MCP config JSON",
)
parser.add_argument(
"--global-config",
default=DEFAULT_GLOBAL_CONFIG,
help="Path to global/canonical MCP config JSON",
)
parser.add_argument(
"--json", action="store_true", help="Print raw JSON report"
)
args = parser.parse_args()
report = analyze_config_drift(args.active_config, args.global_config)
if args.json:
print(json.dumps(report, indent=2))
else:
print("=== MCP Config Drift Diagnostic Report ===")
print(f"Timestamp: {report['timestamp']}")
print(f"In Sync: {report['in_sync']}")
print(f"Active IDE Config: {report['active_config_path']} (exists={report['active_config_exists']})")
print(f"Global Config: {report['global_config_path']} (exists={report['global_config_exists']})")
print(f"Present Role Servers: {', '.join(report['present_role_servers']) if report['present_role_servers'] else 'None'}")
print(f"Missing Role Servers: {', '.join(report['missing_role_servers']) if report['missing_role_servers'] else 'None'}")
if report['profile_mismatches']:
print("Profile Mismatches:")
for m in report['profile_mismatches']:
print(f" - {m['server']}: active={m['active_profile']} vs global={m['global_profile']}")
if report['reasons']:
print("Drift Reasons:")
for r in report['reasons']:
print(f" - {r}")
print("\nSanctioned Repair Runbook:")
for step in report['sanctioned_repair_runbook']:
print(f" {step}")
sys.exit(0 if report["in_sync"] else 1)
if __name__ == "__main__":
main()
+13 -176
View File
@@ -32,8 +32,6 @@ import time
from pathlib import Path
from typing import Any
import mcp_transport_config
SANCTIONED_DAEMON_ENV = "GITEA_MCP_SANCTIONED_DAEMON"
ALLOW_DIRECT_IMPORT_ENV = "GITEA_ALLOW_DIRECT_MCP_IMPORT"
ALLOW_KEYCHAIN_CLI_ENV = "GITEA_ALLOW_KEYCHAIN_CLI"
@@ -44,9 +42,7 @@ FORCE_PROVENANCE_FAIL_ENV = "GITEA_TEST_FORCE_UNSANCTIONED"
_NATIVE_RUNTIME: dict[str, Any] | None = None
# Production transport identifiers accepted by bind_native_mcp_transport.
# #931: the permitted set is defined once, in mcp_transport_config. This name
# is kept as an alias so the guard never restates a transport identifier.
_PRODUCTION_TRANSPORTS = mcp_transport_config.SUPPORTED_TRANSPORTS
_PRODUCTION_TRANSPORTS = frozenset({"stdio"})
_RUNTIME_MODE_PRODUCTION = "production"
_RUNTIME_MODE_TEST = "test"
_PHASE_ENTRYPOINT_CLAIMED = "entrypoint_claimed"
@@ -62,23 +58,6 @@ class UnsanctionedRuntimeError(RuntimeError):
"""Raised when mutation/credential code runs outside a native MCP daemon."""
class TransportExecutionError(UnsanctionedRuntimeError):
"""Raised when a bound transport may not be served by this entrypoint (#931).
Subclasses :class:`UnsanctionedRuntimeError` so every existing fail-closed
handler still catches it, while letting a caller that cares distinguish
"nothing is bound" from "something valid is bound but its listener has not
been commissioned". Carries the structured verdict on ``.assessment``.
"""
def __init__(self, message: str, assessment: dict[str, Any] | None = None):
super().__init__(message)
self.assessment = assessment or {}
self.blocker_kind = self.assessment.get("blocker_kind")
self.owner_issue = self.assessment.get("owner_issue")
self.transport = self.assessment.get("transport")
def is_pytest_runtime() -> bool:
if (os.environ.get(FORCE_PROVENANCE_FAIL_ENV) or "").strip() in {
"1",
@@ -192,46 +171,23 @@ def mark_sanctioned_daemon() -> dict[str, Any]:
return native_runtime_status()
def bind_native_mcp_transport(*, transport: str | None = None) -> dict[str, Any]:
"""Bind the live native MCP transport lifecycle (#695 / #931).
def bind_native_mcp_transport(*, transport: str) -> dict[str, Any]:
"""Bind the live native MCP transport lifecycle (#695).
Must be called from the resolved canonical entrypoint immediately before
the real MCP server transport loop (``mcp.run``). Requires a prior
successful :func:`mark_sanctioned_daemon` claim in this process.
Import-only or offline launch without this bind leaves
the real MCP server transport loop (e.g. ``mcp.run(transport=\"stdio\")``).
Requires a prior successful :func:`mark_sanctioned_daemon` claim in this
process. Import-only or offline launch without this bind leaves
:func:`is_native_mcp_transport` false.
#931: ``transport`` is now optional. Omitting it — which is what the
production entrypoint does resolves the identifier from deployment
configuration via :func:`mcp_transport_config.resolve_configured_transport`,
yielding :data:`mcp_transport_config.DEFAULT_TRANSPORT` when nothing is
configured. An explicit argument remains supported for tests and for a
launcher that has already resolved the value. Either way the identifier is
validated against the single permitted set before the runtime record is
written, so no tool can dispatch over an unregistered transport.
The resolved value is pinned into the process-local record and is read back
only through :func:`bound_transport`. Rebinding to a different transport is
refused, so two guards can never observe different values in one process.
"""
global _NATIVE_RUNTIME
if transport is None:
resolution = mcp_transport_config.resolve_configured_transport()
transport_name = str(resolution["transport"])
if not resolution["supported"]:
raise UnsanctionedRuntimeError(
"bind_native_mcp_transport rejected: "
+ "; ".join(resolution["reasons"])
+ " No tool is served over an unregistered transport."
)
else:
transport_name = mcp_transport_config.normalize_transport(transport)
if transport_name not in _PRODUCTION_TRANSPORTS:
raise UnsanctionedRuntimeError(
f"bind_native_mcp_transport rejected: transport {transport!r} is "
f"not a production MCP transport (#695). Allowed: "
f"{sorted(_PRODUCTION_TRANSPORTS)}."
)
transport_name = (transport or "").strip().lower()
if transport_name not in _PRODUCTION_TRANSPORTS:
raise UnsanctionedRuntimeError(
f"bind_native_mcp_transport rejected: transport {transport!r} is "
f"not a production MCP transport (#695). Allowed: "
f"{sorted(_PRODUCTION_TRANSPORTS)}."
)
entrypoint_path = _caller_official_entrypoint_path()
if entrypoint_path is None:
@@ -260,23 +216,6 @@ def bind_native_mcp_transport(*, transport: str | None = None) -> dict[str, Any]
"between mark and bind (#695)."
)
# #931: one process binds one transport. Re-binding the same identifier is
# idempotent (a retried launch step must not fail); re-binding a different
# one is refused, because a guard that already read the first value would
# otherwise disagree with a guard that reads the second.
already_bound = (_NATIVE_RUNTIME.get("transport") or "").strip()
if (
already_bound
and _NATIVE_RUNTIME.get("phase") == _PHASE_TRANSPORT_BOUND
and already_bound != transport_name
):
raise UnsanctionedRuntimeError(
"bind_native_mcp_transport rejected: transport is already bound to "
f"{already_bound!r} in this process; rebinding to "
f"{transport_name!r} is forbidden (#931). Restart the server to "
"change the deployment transport."
)
# Pin session-state root for this server lifetime (#695 AC2 / PR #701).
# Changing GITEA_MCP_SESSION_STATE_DIR after bind must not manufacture a
# second authority domain for decision locks / workflow proofs.
@@ -414,88 +353,6 @@ def is_production_native_mcp_transport() -> bool:
return (_NATIVE_RUNTIME or {}).get("mode") == _RUNTIME_MODE_PRODUCTION
def bound_transport() -> str | None:
"""The one authoritative bound transport identifier, or ``None`` (#931).
This is the shared accessor every transport-aware guard reads. It reports
the value pinned at bind time, never the environment, so changing
``GITEA_MCP_TRANSPORT`` after the bind cannot move what a guard observes
the same rule :func:`pinned_session_state_dir` applies to session state.
``None`` means unbound: an offline import or a launch that never reached
the bind. Callers must treat that as fail-closed, exactly as they already
treat :func:`is_native_mcp_transport` returning false.
"""
if not is_native_mcp_transport():
return None
return (_NATIVE_RUNTIME or {}).get("transport") or None
def assert_transport_bound(context: str = "tool service") -> str:
"""Return the bound transport, or fail closed before *context* (#931).
Called immediately before the server enters its transport loop so an
invalid or absent bind stops the process rather than serving tools over a
transport no guard can name.
"""
transport = bound_transport()
if transport:
return transport
raise UnsanctionedRuntimeError(
f"No MCP transport is bound; refusing {context} (#931). "
"bind_native_mcp_transport must succeed from the canonical entrypoint "
"before any tool is served. Offline import and standalone launch "
"cannot reconstruct a bind."
)
def assess_serve_authorization() -> dict[str, Any]:
"""Structured verdict on whether this process may serve tools (#931).
This is the decision that consumes :func:`bound_transport`. It is what stops
the bound identifier from being reporting-only metadata: the serve path
cannot proceed unless the value pinned at bind is one this entrypoint is
commissioned to execute.
Never raises; returns the verdict so callers and diagnostics can inspect it.
"""
return mcp_transport_config.assess_transport_execution(bound_transport())
def authorize_transport_execution(context: str = "tool service") -> str:
"""Return the transport this process may serve, or fail closed (#931).
Two distinct boundaries, in order:
1. **Bind presence** :func:`assert_transport_bound` enforces the
pre-existing #695 contract, so an unbound runtime keeps its established
failure and reason code.
2. **Execution authorization** the bound identifier must be one this
entrypoint is commissioned to serve. A registered transport whose
listener has not been commissioned is refused here, before any listener
is created and before any tool can dispatch.
That ordering matters: recognition, validation and durable recording all
still happen for a remote identifier, so #931's seam is intact; only the act
of *serving* it is withheld until its owning issue commissions it.
"""
# Boundary 1: unbound stays exactly as fail-closed as it was under #695.
assert_transport_bound(context)
# Boundary 2: bound, but is this entrypoint allowed to serve it?
assessment = assess_serve_authorization()
if assessment.get("allowed"):
return str(assessment["transport"])
reasons = "; ".join(assessment.get("reasons") or []) or "not authorized"
next_action = assessment.get("exact_next_action") or ""
raise TransportExecutionError(
f"Refusing {context} (#931) [{assessment.get('blocker_kind')}]: "
f"{reasons} {next_action}".strip(),
assessment,
)
def is_sanctioned_mcp_daemon() -> bool:
"""Backward-compatible name; #695 requires native transport, not env alone."""
if is_production_native_mcp_transport():
@@ -619,21 +476,6 @@ def native_runtime_status() -> dict[str, Any]:
"entrypoint_path": rt.get("entrypoint_path"),
"phase": rt.get("phase"),
"transport": rt.get("transport"),
# #931: the authoritative bound identifier, plus the seam that defines
# what may be bound. ``bound_transport`` is None until a bind succeeds,
# so an offline import is distinguishable from a stdio session.
"bound_transport": bound_transport(),
"transport_bound": bound_transport() is not None,
"default_transport": mcp_transport_config.DEFAULT_TRANSPORT,
"supported_transports": list(mcp_transport_config.supported_transports()),
"transport_env": mcp_transport_config.TRANSPORT_ENV,
# #931 review 635: recognition and execution authorization are distinct.
# ``supported`` is what may be bound; ``executable`` is what this
# entrypoint may actually serve. A recognized-but-uncommissioned
# transport reports serve_authorized False with a named blocker.
"executable_transports": list(mcp_transport_config.executable_transports()),
"serve_authorized": bool(assess_serve_authorization().get("allowed")),
"serve_authorization": assess_serve_authorization(),
"mode": rt.get("mode"),
"session_state_dir": pinned_session_state_dir() or rt.get("session_state_dir"),
"session_state_dir_pinned": pinned_session_state_dir() is not None,
@@ -660,12 +502,7 @@ def mutation_provenance_fields() -> dict[str, Any]:
if st.get("mode") == _RUNTIME_MODE_TEST and st["native_mcp_transport"]:
transport = "test_native_mcp"
return {
# ``transport`` stays the trust *class* it has always been, so existing
# durable records keep their shape. ``bound_transport`` (#931) adds the
# bound identifier itself, which is what lets an operator tell from a
# durable record which transport performed a mutation.
"transport": transport,
"bound_transport": st.get("bound_transport"),
"native_mcp_transport": bool(st["native_mcp_transport"]),
"production_native_mcp_transport": bool(
st.get("production_native_mcp_transport")
-869
View File
@@ -1,869 +0,0 @@
"""Instance-level fleet identity and health snapshots (#978).
#948 established per-worker ownership; #975 made heartbeats keep those rows
live. Neither surface could enumerate the fleet at *instance* granularity:
which application launch owns which five namespace workers, whether two
Codex launches are distinct, or whether a live collision is real rather than
a shared client type.
This module is pure. Callers supply registry rows (and optional enrichments);
nothing here opens SQLite, scans process tables, or mutates state. Production
evidence for the fleet gate is the snapshot returned by the sanctioned
controller/reconciler tool that wraps this assessor.
Identity hierarchy (highest lowest):
* ``client_type`` application family (``codex``, ``claude_code``, )
* ``client_instance_id`` one running application launch (trusted launcher)
* ``fleet_run_id`` operator-approved enrollment / canary cohort
* ``worker_id`` / ``worker_identity`` one namespace worker process
* ``namespace`` author | reviewer | merger | controller | reconciler
Multiple simultaneous instances of the same ``client_type`` are first-class.
Sharing only a profile or client type is never a duplicate.
"""
from __future__ import annotations
import hashlib
import secrets
from collections import defaultdict
from datetime import datetime, timezone
from typing import Any, Callable, Iterable, Mapping
import mcp_worker_identity as mwi
# --- Classification labels ------------------------------------------------
CLASS_EXPECTED = "expected_enrolled"
CLASS_MISSING = "missing_expected"
CLASS_UNMANIFESTED = "unmanifested"
CLASS_DUPLICATE_NAMESPACE = "duplicate_namespace_worker"
CLASS_INSTANCE_ID_COLLISION = "instance_id_collision"
CLASS_WORKER_ID_COLLISION = "worker_identity_collision"
CLASS_SESSION_COLLISION = "session_identity_collision"
CLASS_GENERATION_COLLISION = "generation_identity_collision"
CLASS_PROCESS_COLLISION = "process_identity_collision"
CLASS_PID_COLLISION = "pid_collision"
CLASS_OWNERSHIP_COLLISION = "ownership_fencing_collision"
CLASS_ORPHANED = "orphaned_unowned"
CLASS_UNKNOWN_CLIENT = "unknown_client"
CLASS_FOREIGN_REPOSITORY = "foreign_repository"
CLASS_OLD_REVISION = "old_revision"
CLASS_STALE_WORKER = "stale_orphaned_worker"
CLASS_LEGACY_INCOMPLETE = "legacy_incomplete_identity"
CLASS_HISTORICAL = "historical_dead"
CLASS_HEALTHY = "healthy"
#: Active blockers that make the live fleet unsafe for mutation-gated work.
ACTIVE_BLOCKER_CLASSES = frozenset(
{
CLASS_MISSING,
CLASS_UNMANIFESTED,
CLASS_DUPLICATE_NAMESPACE,
CLASS_INSTANCE_ID_COLLISION,
CLASS_WORKER_ID_COLLISION,
CLASS_SESSION_COLLISION,
CLASS_GENERATION_COLLISION,
CLASS_PROCESS_COLLISION,
CLASS_PID_COLLISION,
CLASS_OWNERSHIP_COLLISION,
CLASS_ORPHANED,
CLASS_UNKNOWN_CLIENT,
CLASS_FOREIGN_REPOSITORY,
CLASS_OLD_REVISION,
CLASS_STALE_WORKER,
CLASS_LEGACY_INCOMPLETE,
}
)
SANCTIONED_NAMESPACES = frozenset(
{"author", "reviewer", "merger", "controller", "reconciler"}
)
INSTANCE_ID_PROVENANCE_TRUSTED = "trusted_launcher"
INSTANCE_ID_PROVENANCE_LEGACY = "legacy_incomplete"
INSTANCE_ID_PROVENANCE_MISSING = "missing"
CLIENT_INSTANCE_ENV = "GITEA_MCP_CLIENT_INSTANCE"
FLEET_RUN_ENV = "GITEA_MCP_FLEET_RUN_ID"
PROCESS_IDENTITY_ENV = "GITEA_MCP_PROCESS_IDENTITY"
INSTANCE_PROVENANCE_ENV = "GITEA_MCP_INSTANCE_PROVENANCE"
_LEGACY_INSTANCE_PREFIXES = ("pid-", "proc-", "legacy-")
#: Trusted launcher-issued IDs use the reserved ``inst-`` prefix
#: (see :func:`generate_client_instance_id`). Ordinary user-supplied strings
#: without that prefix never count as trusted attribution.
_TRUSTED_INSTANCE_PREFIX = "inst-"
def _utc_now() -> datetime:
return datetime.now(timezone.utc)
def _ts(value: datetime) -> str:
return value.astimezone(timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ")
def generate_client_instance_id(
client_type: str | None,
*,
launch_nonce: str | None = None,
now: datetime | None = None,
) -> str:
"""Mint a distinct instance ID for one application launch (#978).
The trusted launcher (or host that starts all five namespace workers)
generates this once per launch and injects it as ``GITEA_MCP_CLIENT_INSTANCE``
into every worker environment. Workers never invent their own instance ID
from PID proximity or timestamps.
"""
client = mwi.normalize_client_name(client_type)
stamp = (now or _utc_now()).astimezone(timezone.utc).strftime(
mwi.IDENTITY_TIMESTAMP_FORMAT
)
nonce = launch_nonce if launch_nonce is not None else secrets.token_hex(16)
digest = hashlib.sha256(
f"{client}\x1f{stamp}\x1f{nonce}".encode("utf-8")
).hexdigest()[:12]
return f"inst-{client}-{stamp}-{digest}"
def assess_instance_identity(
raw_instance_id: str | None,
*,
source: str | None = None,
) -> dict[str, Any]:
"""Classify whether a client_instance_id is trusted enough for mutation.
Trusted instance IDs are non-empty, match the launcher-minted ``inst-``
format from :func:`generate_client_instance_id`, and are not pre-#978
PID/proc/legacy placeholders. Ordinary user-supplied or malformed values
fail closed as untrusted so they cannot spoof multi-instance attribution.
Incomplete identities remain visible for diagnosis but cannot authorize
unsafe mutation.
"""
text = (raw_instance_id or "").strip()
if not text:
return {
"client_instance_id": None,
"complete": False,
"trusted": False,
"provenance": INSTANCE_ID_PROVENANCE_MISSING,
"reasons": [
"client_instance_id is missing; the trusted launcher must set "
f"{CLIENT_INSTANCE_ENV} once per application launch"
],
}
lowered = text.lower()
if lowered.startswith(_LEGACY_INSTANCE_PREFIXES) or source == "pid_fallback":
return {
"client_instance_id": text,
"complete": False,
"trusted": False,
"provenance": INSTANCE_ID_PROVENANCE_LEGACY,
"reasons": [
f"client_instance_id {text!r} is a legacy PID/process fallback, "
"not a trusted launcher-issued instance identity"
],
}
if not text.startswith(_TRUSTED_INSTANCE_PREFIX) or len(text) <= len(
_TRUSTED_INSTANCE_PREFIX
):
return {
"client_instance_id": text,
"complete": False,
"trusted": False,
"provenance": INSTANCE_ID_PROVENANCE_LEGACY,
"reasons": [
f"client_instance_id {text!r} is malformed or user-supplied and "
f"does not use the trusted launcher prefix "
f"{_TRUSTED_INSTANCE_PREFIX!r}; refusing trusted attribution"
],
}
return {
"client_instance_id": text,
"complete": True,
"trusted": True,
"provenance": INSTANCE_ID_PROVENANCE_TRUSTED,
"reasons": [],
}
def resolve_client_instance_from_env(
env: Mapping[str, str] | None = None,
*,
pid: int | None = None,
) -> dict[str, Any]:
"""Resolve instance identity from launcher env without inventing one.
When the trusted key is absent, return incomplete evidence rather than a
silent ``pid-<n>`` identity. Callers that still need a non-empty registry
key may choose a legacy placeholder deliberately; they must not treat it as
trusted.
"""
source = dict(env or {})
raw = (source.get(CLIENT_INSTANCE_ENV) or "").strip()
assessment = assess_instance_identity(raw or None)
assessment["fleet_run_id"] = (source.get(FLEET_RUN_ENV) or "").strip() or None
assessment["process_identity"] = (
(source.get(PROCESS_IDENTITY_ENV) or "").strip()
or (f"pid-{pid}" if pid is not None else None)
)
return assessment
def _public_worker(record: Mapping[str, Any]) -> dict[str, Any]:
return {
"worker_id": record.get("worker_identity") or record.get("worker_id"),
"worker_identity": record.get("worker_identity") or record.get("worker_id"),
"client_type": mwi.normalize_client_name(
record.get("client_name") or record.get("client_type")
),
"client_instance_id": record.get("client_instance_id"),
"fleet_run_id": record.get("fleet_run_id"),
"namespace": record.get("namespace"),
"profile": record.get("profile"),
"declared_role": record.get("role") or record.get("declared_role"),
"authenticated_account": record.get("authenticated_account"),
"session_id": record.get("session_id"),
"generation_id": record.get("generation_id"),
"process_identity": record.get("process_identity")
or (
f"pid-{record['pid']}"
if record.get("pid") is not None
else None
),
"pid": record.get("pid"),
"repository_binding": record.get("repository_binding"),
"remote": record.get("remote"),
"startup_revision": record.get("startup_revision"),
"loaded_revision": record.get("loaded_revision"),
"parity_revision": record.get("parity_revision"),
"live_revision": record.get("live_revision"),
"runtime_provenance": record.get("runtime_provenance")
or record.get("transport"),
"transport": record.get("transport"),
"started_at": record.get("started_at"),
"last_heartbeat_at": record.get("last_heartbeat_at"),
"heartbeat_ttl_seconds": record.get("heartbeat_ttl_seconds"),
"fencing_epoch": record.get("fencing_epoch"),
"status": record.get("status"),
"instance_id_provenance": record.get("instance_id_provenance"),
}
def _liveness(
record: Mapping[str, Any],
*,
now: datetime | None,
pid_alive_probe: Callable[[int | None], bool | None] | None,
) -> dict[str, Any]:
pid_alive = None
if pid_alive_probe is not None and record.get("pid") is not None:
try:
pid_alive = pid_alive_probe(record.get("pid"))
except Exception:
pid_alive = None
return mwi.WorkerRegistry.is_live(record, now=now, pid_alive=pid_alive)
def _consistency_token(rows: Iterable[Mapping[str, Any]], snapshot_at: str) -> str:
material = [snapshot_at]
for row in sorted(
rows,
key=lambda r: (
str(r.get("worker_identity") or ""),
str(r.get("last_heartbeat_at") or ""),
str(r.get("fencing_epoch") or ""),
),
):
material.append(
"|".join(
[
str(row.get("worker_identity") or ""),
str(row.get("client_instance_id") or ""),
str(row.get("status") or ""),
str(row.get("last_heartbeat_at") or ""),
str(row.get("fencing_epoch") or ""),
str(row.get("generation_id") or ""),
]
)
)
digest = hashlib.sha256("\n".join(material).encode("utf-8")).hexdigest()[:16]
return f"fleetrev-{digest}"
def build_worker_snapshot_row(
record: Mapping[str, Any],
*,
now: datetime | None = None,
pid_alive_probe: Callable[[int | None], bool | None] | None = None,
canonical_repository: str | None = None,
expected_live_revision: str | None = None,
heartbeat_supervised: bool | None = None,
) -> dict[str, Any]:
"""One point-in-time worker row for the fleet snapshot."""
stamp = now or _utc_now()
base = _public_worker(record)
identity = assess_instance_identity(
base.get("client_instance_id"),
source=record.get("instance_id_source"),
)
liveness = _liveness(record, now=stamp, pid_alive_probe=pid_alive_probe)
is_historical = str(record.get("status") or "") != mwi.STATUS_ACTIVE
live = bool(liveness.get("live")) and not is_historical
repo = (base.get("repository_binding") or "").strip() or None
foreign_repo = bool(
canonical_repository
and repo
and repo.rstrip("/") != str(canonical_repository).rstrip("/")
)
old_revision = False
if expected_live_revision:
for key in ("startup_revision", "loaded_revision", "parity_revision", "live_revision"):
rev = (base.get(key) or "").strip()
if rev and rev != expected_live_revision:
old_revision = True
break
ownership_state = "historical" if is_historical else (
"live" if live else "stale"
)
if live and not identity["trusted"]:
ownership_state = "live_untrusted_identity"
if live and not base.get("session_id"):
ownership_state = "orphaned"
mutation_safe = bool(
live
and identity["trusted"]
and not foreign_repo
and not old_revision
and ownership_state == "live"
and base.get("client_type") != mwi.UNKNOWN_CLIENT
)
restart_required = bool(
old_revision
or (live and not liveness.get("heartbeat_fresh", True))
)
return {
**base,
"instance_identity": identity,
"client_instance_id": identity["client_instance_id"] or base.get("client_instance_id"),
"instance_id_provenance": identity["provenance"],
"instance_identity_trusted": identity["trusted"],
"live": live,
"historical": is_historical,
"liveness": liveness,
"heartbeat": {
"registered": bool(base.get("last_heartbeat_at")),
"supervised": heartbeat_supervised,
"age_seconds": liveness.get("heartbeat_age_seconds"),
"ttl_seconds": liveness.get("heartbeat_ttl_seconds"),
"fresh": liveness.get("heartbeat_fresh"),
"last_heartbeat_at": base.get("last_heartbeat_at"),
},
"fencing": {
"fencing_epoch": base.get("fencing_epoch"),
"generation_id": base.get("generation_id"),
},
"ownership_state": ownership_state,
"foreign_repository": foreign_repo,
"old_revision": old_revision,
"stale": not live and not is_historical,
"restart_required": restart_required,
"mutation_safe": mutation_safe,
"conflicting_live_sessions": [],
}
def _collision_groups(
live_rows: list[dict[str, Any]],
key_fn,
) -> dict[str, list[dict[str, Any]]]:
groups: dict[str, list[dict[str, Any]]] = defaultdict(list)
for row in live_rows:
key = key_fn(row)
if key is None or key == "" or key == "None":
continue
groups[str(key)].append(row)
return {k: v for k, v in groups.items() if len(v) > 1}
def snapshot_instance_fleet(
records: Iterable[Mapping[str, Any]],
*,
expected_manifest: list[Mapping[str, Any]] | None = None,
now: datetime | None = None,
pid_alive_probe: Callable[[int | None], bool | None] | None = None,
canonical_repository: str | None = None,
expected_live_revision: str | None = None,
registry_revision: str | None = None,
known_client_types: Iterable[str] | None = None,
) -> dict[str, Any]:
"""Authoritative point-in-time fleet snapshot with classification (#978).
Historical dead rows are reported separately and never automatically make
the live fleet unsafe.
"""
stamp = now or _utc_now()
snapshot_at = _ts(stamp)
known = {
mwi.normalize_client_name(c)
for c in (known_client_types or mwi.CLIENT_ALIASES.values())
}
known.discard(mwi.UNKNOWN_CLIENT)
all_rows: list[dict[str, Any]] = []
for record in records:
all_rows.append(
build_worker_snapshot_row(
record,
now=stamp,
pid_alive_probe=pid_alive_probe,
canonical_repository=canonical_repository,
expected_live_revision=expected_live_revision,
)
)
live_rows = [r for r in all_rows if r["live"]]
historical_rows = [r for r in all_rows if r["historical"]]
stale_rows = [r for r in all_rows if r["stale"]]
# --- identity collisions among live workers ---
findings: list[dict[str, Any]] = []
def _finding(
classification: str,
*,
severity: str,
workers: list[dict[str, Any]] | None = None,
instance_ids: list[str] | None = None,
detail: str,
active_blocker: bool,
) -> None:
findings.append(
{
"classification": classification,
"severity": severity,
"active_blocker": active_blocker,
"detail": detail,
"client_instance_ids": instance_ids or sorted(
{
str(w.get("client_instance_id"))
for w in (workers or [])
if w.get("client_instance_id")
}
),
"worker_identities": [
w.get("worker_identity") for w in (workers or [])
],
}
)
# Duplicate worker identity (should not happen with PK, still detect)
for wid, group in _collision_groups(
live_rows, lambda r: r.get("worker_identity")
).items():
_finding(
CLASS_WORKER_ID_COLLISION,
severity="blocker",
workers=group,
detail=f"worker identity {wid!r} is claimed by {len(group)} live workers",
active_blocker=True,
)
# Reused session identity across live workers
for sid, group in _collision_groups(live_rows, lambda r: r.get("session_id")).items():
# Same session may appear once; collision only when multiple workers share it
# across different worker identities (always true for group size > 1).
_finding(
CLASS_SESSION_COLLISION,
severity="blocker",
workers=group,
detail=f"session identity {sid!r} is reused by {len(group)} live workers",
active_blocker=True,
)
# Generation claimed by multiple live sessions/workers is a conflict when
# the workers are not the five sanctioned namespaces of one instance.
for gen, group in _collision_groups(
live_rows, lambda r: r.get("generation_id")
).items():
namespaces = {g.get("namespace") for g in group if g.get("namespace")}
instance_ids = {g.get("client_instance_id") for g in group}
# Multiple workers under one generation is only valid if they share one
# instance and distinct namespaces. Same generation + same namespace = bad.
by_ns: dict[str, list] = defaultdict(list)
for g in group:
by_ns[str(g.get("namespace") or "")].append(g)
ns_dups = {ns: rows for ns, rows in by_ns.items() if ns and len(rows) > 1}
if ns_dups or len(instance_ids) > 1:
_finding(
CLASS_GENERATION_COLLISION,
severity="blocker",
workers=group,
detail=(
f"generation {gen!r} is contested across namespaces/instances "
f"(namespaces={sorted(namespaces)}, "
f"instances={sorted(str(i) for i in instance_ids if i)})"
),
active_blocker=True,
)
# Process identity / PID collisions across distinct workers
for proc, group in _collision_groups(
live_rows, lambda r: r.get("process_identity")
).items():
if len({r.get("worker_identity") for r in group}) > 1:
_finding(
CLASS_PROCESS_COLLISION,
severity="blocker",
workers=group,
detail=f"process identity {proc!r} is shared by distinct live workers",
active_blocker=True,
)
for pid, group in _collision_groups(live_rows, lambda r: r.get("pid")).items():
if len({r.get("worker_identity") for r in group}) > 1:
_finding(
CLASS_PID_COLLISION,
severity="blocker",
workers=group,
detail=f"PID {pid} is shared by distinct live workers",
active_blocker=True,
)
# Fencing/ownership: same fencing epoch on different workers of different instances
for epoch, group in _collision_groups(
live_rows,
lambda r: (
f"{r.get('generation_id')}:{r.get('fencing_epoch')}"
if r.get("generation_id") is not None and r.get("fencing_epoch") is not None
else None
),
).items():
if len({r.get("client_instance_id") for r in group}) > 1:
_finding(
CLASS_OWNERSHIP_COLLISION,
severity="blocker",
workers=group,
detail=(
f"fencing token {epoch!r} spans more than one client_instance_id"
),
active_blocker=True,
)
# Per-instance grouping
by_instance: dict[str, list[dict[str, Any]]] = defaultdict(list)
unkeyed_live: list[dict[str, Any]] = []
for row in live_rows:
iid = row.get("client_instance_id")
if not iid:
unkeyed_live.append(row)
continue
by_instance[str(iid)].append(row)
instances: list[dict[str, Any]] = []
for iid, workers in sorted(by_instance.items()):
client_types = sorted({w.get("client_type") for w in workers if w.get("client_type")})
trusted = all(w.get("instance_identity_trusted") for w in workers)
namespaces = [w.get("namespace") for w in workers]
ns_counts: dict[str, int] = defaultdict(int)
for ns in namespaces:
if ns:
ns_counts[str(ns)] += 1
dup_ns = sorted(ns for ns, n in ns_counts.items() if n > 1)
if dup_ns:
_finding(
CLASS_DUPLICATE_NAMESPACE,
severity="blocker",
workers=[w for w in workers if w.get("namespace") in dup_ns],
instance_ids=[iid],
detail=(
f"instance {iid!r} has more than one live worker for "
f"namespace(s) {dup_ns}"
),
active_blocker=True,
)
# Live reuse of one instance ID with incompatible client types
if len(client_types) > 1:
_finding(
CLASS_INSTANCE_ID_COLLISION,
severity="blocker",
workers=workers,
instance_ids=[iid],
detail=(
f"client_instance_id {iid!r} is live under multiple client "
f"types {client_types}"
),
active_blocker=True,
)
if not trusted:
_finding(
CLASS_LEGACY_INCOMPLETE,
severity="blocker",
workers=workers,
instance_ids=[iid],
detail=(
f"instance {iid!r} lacks trusted launcher-issued instance "
"identity; diagnostic reads remain available"
),
active_blocker=True,
)
unknown = [w for w in workers if w.get("client_type") == mwi.UNKNOWN_CLIENT]
if unknown:
_finding(
CLASS_UNKNOWN_CLIENT,
severity="blocker",
workers=unknown,
instance_ids=[iid],
detail=f"instance {iid!r} has worker(s) with unknown client_type",
active_blocker=True,
)
foreign = [w for w in workers if w.get("foreign_repository")]
if foreign:
_finding(
CLASS_FOREIGN_REPOSITORY,
severity="blocker",
workers=foreign,
instance_ids=[iid],
detail=f"instance {iid!r} has foreign-repository workers",
active_blocker=True,
)
old = [w for w in workers if w.get("old_revision")]
if old:
_finding(
CLASS_OLD_REVISION,
severity="blocker",
workers=old,
instance_ids=[iid],
detail=f"instance {iid!r} has old-revision workers",
active_blocker=True,
)
orphans = [w for w in workers if w.get("ownership_state") == "orphaned"]
if orphans:
_finding(
CLASS_ORPHANED,
severity="blocker",
workers=orphans,
instance_ids=[iid],
detail=f"instance {iid!r} has orphaned/unowned workers",
active_blocker=True,
)
instances.append(
{
"client_instance_id": iid,
"client_types": client_types,
"client_type": client_types[0] if len(client_types) == 1 else None,
"fleet_run_ids": sorted(
{w.get("fleet_run_id") for w in workers if w.get("fleet_run_id")}
),
"worker_count": len(workers),
"namespaces": sorted({n for n in namespaces if n}),
"namespace_counts": dict(ns_counts),
"duplicate_namespaces": dup_ns,
"trusted_instance_identity": trusted,
"workers": workers,
"mutation_safe": all(w.get("mutation_safe") for w in workers)
and not dup_ns
and trusted,
}
)
for row in unkeyed_live:
_finding(
CLASS_LEGACY_INCOMPLETE,
severity="blocker",
workers=[row],
detail="live worker has no client_instance_id",
active_blocker=True,
)
if row.get("client_type") == mwi.UNKNOWN_CLIENT:
_finding(
CLASS_UNKNOWN_CLIENT,
severity="blocker",
workers=[row],
detail="live worker has unknown client_type and no instance id",
active_blocker=True,
)
for row in stale_rows:
_finding(
CLASS_STALE_WORKER,
severity="warning",
workers=[row],
detail=(
f"worker {row.get('worker_identity')!r} is active in the registry "
"but not live (stale heartbeat or dead pid)"
),
active_blocker=True,
)
for row in historical_rows:
_finding(
CLASS_HISTORICAL,
severity="info",
workers=[row],
detail=(
f"historical registration {row.get('worker_identity')!r} "
f"(status={row.get('status')!r}) is not an active blocker"
),
active_blocker=False,
)
# Manifest comparison
expected = list(expected_manifest or [])
expected_ids = {
str(item.get("client_instance_id")).strip()
for item in expected
if (item.get("client_instance_id") or "").strip()
}
live_ids = set(by_instance.keys())
missing_ids = sorted(expected_ids - live_ids)
unmanifested_ids = sorted(live_ids - expected_ids) if expected_ids else []
for iid in missing_ids:
_finding(
CLASS_MISSING,
severity="blocker",
instance_ids=[iid],
detail=f"expected enrolled instance {iid!r} is missing from the live fleet",
active_blocker=True,
)
for iid in unmanifested_ids:
_finding(
CLASS_UNMANIFESTED,
severity="blocker",
instance_ids=[iid],
workers=by_instance.get(iid, []),
detail=(
f"live instance {iid!r} is not on the approved fleet manifest "
"(unmanifested)"
),
active_blocker=True,
)
# Same client_type multi-instance is healthy when each has distinct instance IDs
by_type: dict[str, list[str]] = defaultdict(list)
for inst in instances:
for ct in inst.get("client_types") or []:
by_type[str(ct)].append(inst["client_instance_id"])
multi_instance_same_type = {
ct: ids for ct, ids in by_type.items() if len(ids) > 1
}
active_blockers = [f for f in findings if f.get("active_blocker")]
historical_only = [f for f in findings if f.get("classification") == CLASS_HISTORICAL]
live_safe = not active_blockers
consistency = registry_revision or _consistency_token(all_rows, snapshot_at)
return {
"success": True,
"read_only": True,
"snapshot_at": snapshot_at,
"consistency_token": consistency,
"registry_revision": consistency,
"live_worker_count": len(live_rows),
"historical_worker_count": len(historical_rows),
"stale_worker_count": len(stale_rows),
"instance_count": len(instances),
"workers": all_rows,
"live_workers": live_rows,
"historical_workers": historical_rows,
"stale_workers": stale_rows,
"instances": instances,
"multi_instance_same_client_type": multi_instance_same_type,
"same_client_type_not_duplicate": True,
"expected_manifest": [
{
"client_instance_id": item.get("client_instance_id"),
"client_type": item.get("client_type"),
"fleet_run_id": item.get("fleet_run_id"),
"namespaces": item.get("namespaces"),
}
for item in expected
],
"missing_expected_instance_ids": missing_ids,
"unmanifested_instance_ids": unmanifested_ids,
"findings": findings,
"active_blockers": active_blockers,
"historical_findings": historical_only,
"live_fleet_safe": live_safe,
"mutation_safe": live_safe and all(
inst.get("mutation_safe") for inst in instances
)
if instances
else live_safe,
"classification_model": {
"exactly_one_process_per_profile": False,
"exactly_one_instance_per_client_type": False,
"multiple_instances_per_client_type": True,
"duplicate_requires": [
"live client_instance_id collision",
"duplicate namespace worker within one instance",
"reused worker/session/generation/process/pid/fencing identity",
"cross-instance ownership collision",
"unmanifested instance when a manifest is required",
],
},
"reasons": [f["detail"] for f in active_blockers],
}
def compare_snapshot_heartbeats(
earlier: Mapping[str, Any],
later: Mapping[str, Any],
) -> dict[str, Any]:
"""Prove heartbeat continuity and stable ownership across two snapshots."""
earlier_live = {
w.get("worker_identity"): w for w in earlier.get("live_workers") or []
}
later_live = {
w.get("worker_identity"): w for w in later.get("live_workers") or []
}
shared = sorted(set(earlier_live) & set(later_live))
continuity: list[dict[str, Any]] = []
stable_ownership = True
for wid in shared:
a = earlier_live[wid]
b = later_live[wid]
same_instance = a.get("client_instance_id") == b.get("client_instance_id")
same_session = a.get("session_id") == b.get("session_id")
same_generation = a.get("generation_id") == b.get("generation_id")
hb_advanced_or_equal = True
if a.get("last_heartbeat_at") and b.get("last_heartbeat_at"):
hb_advanced_or_equal = b["last_heartbeat_at"] >= a["last_heartbeat_at"]
if not (same_instance and same_session and same_generation):
stable_ownership = False
continuity.append(
{
"worker_identity": wid,
"same_client_instance_id": same_instance,
"same_session_id": same_session,
"same_generation_id": same_generation,
"heartbeat_non_decreasing": hb_advanced_or_equal,
"earlier_heartbeat": a.get("last_heartbeat_at"),
"later_heartbeat": b.get("last_heartbeat_at"),
}
)
return {
"shared_live_workers": shared,
"continuity": continuity,
"stable_ownership": stable_ownership
and all(c["heartbeat_non_decreasing"] for c in continuity),
"dropped_workers": sorted(set(earlier_live) - set(later_live)),
"new_workers": sorted(set(later_live) - set(earlier_live)),
}
-408
View File
@@ -51,42 +51,6 @@ EOF_PATTERNS = (
"eof",
)
ERROR_CONNECTED_NAMESPACES_MISSING = "mcp_connected_namespaces_missing"
# Distinct from the condition above. ``mcp_connected_namespaces_missing`` is reserved for
# the actual #708 defect: the host *does* report the service Connected, yet the namespace
# never entered the active session tool surface. A required namespace that is absent from
# the connected-service inventory has no Connected claim behind it at all, so reporting it
# under the #708 condition would assert something the evidence does not support and would
# point an operator at the wrong recovery.
ERROR_REQUIRED_NAMESPACES_NOT_CONNECTED = "mcp_required_namespaces_not_connected"
DISCOVERY_STATUS_ATTACHED = "namespaces_attached"
DISCOVERY_STATUS_CONNECTED_MISSING = "connected_but_namespaces_missing"
DISCOVERY_STATUS_NOT_CONNECTED = "required_namespaces_not_connected"
DISCOVERY_STATUS_DISCONNECTED = "disconnected"
# Namespaces that must be *attached to the active session* for a mutation task (#708).
# Connected-at-host is not attached-in-session; these are gated separately from the
# #543 health map because a namespace can be healthy on probe yet absent from the
# session tool surface.
ATTACHMENT_GATED_TASKS = {
"review_pr": "gitea-reviewer",
"submit_review": "gitea-reviewer",
"merge_pr": "gitea-merger",
"work_issue": "gitea-author",
"create_pr": "gitea-author",
}
# The only sanctioned recovery for an unattached namespace (#678 exposes it natively).
SANCTIONED_ATTACH_RECOVERY_TOOL = "gitea_request_mcp_reconnect"
UNSAFE_FALLBACK_WARNING = (
"Workflow Safety Hard Stop (#708): Connected-but-namespaces-missing recovery must "
"NEVER use direct imports, Gitea API mutations, profile hopping, session-state "
"overrides, PID kills, or config mtime touches. Use client reconnect only."
)
SAFE_ENV_KEYS = (
"GITEA_MCP_PROFILE",
"GITEA_PROFILE_NAME",
@@ -95,318 +59,6 @@ SAFE_ENV_KEYS = (
"GITEA_MCP_CONFIG",
)
# #948: provenance used to be derived from the summary this allowlist produces.
# The allowlist never carried a provenance key, so that derivation could only
# ever evaluate to ``manual_launch`` — whatever the process actually was — while
# ``gitea_get_runtime_context`` read the live environment and reported
# ``client_managed`` for the same process. Provenance is no longer derived here.
# It comes from ``mcp_worker_identity.assess_provenance``, the single authority
# every surface shares. This allowlist keeps its original and only job: deciding
# which env values are safe to echo back in diagnostics.
def assess_connected_namespace_attachment(
*,
connected_servers: list[str] | tuple[str, ...] | set[str] | None = None,
attached_session_namespaces: list[str] | tuple[str, ...] | set[str] | None = None,
required_namespaces: list[str] | tuple[str, ...] | set[str] | None = None,
discovery_cache_age_seconds: float | int | None = None,
discovery_cache_hit: bool | None = None,
auto_attach_attempted: bool = False,
auto_attach_succeeded: bool = False,
session_tool_snapshot_at: float | int | None = None,
namespace_connected_at: dict[str, float] | None = None,
) -> dict[str, Any]:
"""Assess whether host-connected MCP servers have attached tool namespaces in the active session (#708).
Addresses the Connected-but-namespaces-missing defect: CLI/host status may report Connected
while the active LLM session tool surface exposes 0 attached tool namespaces.
This is a *distinct* condition from config drift (#672), transport-closed (#584),
and resolver EOF (#685): the transport is up and the host reports Connected, yet the
namespace never entered the session tool surface.
Startup ordering (``session_tool_snapshot_at`` + ``namespace_connected_at``) identifies
the race where the session tool snapshot was taken before a role server finished
``initialize``/``list_tools``, which is why parallel multi-role startup can leave the
session with an empty namespace set while Connected later flips true.
Returns structured detection details plus secret-free telemetry.
"""
connected = [str(s).strip() for s in (connected_servers or []) if str(s).strip()]
attached = set(str(ns).strip() for ns in (attached_session_namespaces or []) if str(ns).strip())
req = [str(r).strip() for r in (required_namespaces or DEFAULT_NAMESPACES) if str(r).strip()]
required = list(dict.fromkeys(req))
connected_set = set(connected)
proof: dict[str, dict[str, bool]] = {}
for s in connected:
proof[s] = {"connected": True, "attached": s in attached}
for r in required:
if r not in proof:
proof[r] = {"connected": r in connected_set, "attached": r in attached}
# The genuine #708 condition: the host reports the service Connected and the namespace
# still never entered the active session tool surface.
missing = [r for r in required if r in connected_set and r not in attached]
# A separate condition: the service is required but absent from the connected-service
# inventory. Nothing here is Connected, so this must not borrow the #708 wording or its
# recovery — see ERROR_REQUIRED_NAMESPACES_NOT_CONNECTED.
not_connected = [r for r in required if r not in connected_set]
# Evidence that disagrees with itself: reported attached to the session while absent
# from the connected inventory. Neither statement is proof, so it fails closed.
contradictory = [r for r in not_connected if r in attached]
# No required namespace may lack attachment proof and still be called healthy.
attachment_healthy = (
not missing
and not not_connected
and len(connected) > 0
and len(required) > 0
)
if attachment_healthy:
discovery_status = DISCOVERY_STATUS_ATTACHED
elif not connected:
discovery_status = DISCOVERY_STATUS_DISCONNECTED
elif not_connected:
discovery_status = DISCOVERY_STATUS_NOT_CONNECTED
else:
discovery_status = DISCOVERY_STATUS_CONNECTED_MISSING
# Every condition actually present is reported; ``error_type`` names the primary one.
# Not-connected outranks Connected-but-unattached because a service that never
# connected cannot be recovered by attaching its namespace.
error_types: list[str] = []
if not_connected:
error_types.append(ERROR_REQUIRED_NAMESPACES_NOT_CONNECTED)
if missing:
error_types.append(ERROR_CONNECTED_NAMESPACES_MISSING)
if not attachment_healthy and not error_types:
error_types.append(ERROR_REQUIRED_NAMESPACES_NOT_CONNECTED)
error_type = None if attachment_healthy else error_types[0]
# Per-namespace verdict, so a mutation gate never has to infer one namespace's state
# from a whole-session summary. Each entry states only what its own evidence supports.
namespace_conditions: dict[str, dict[str, Any]] = {}
for ns_name, state in proof.items():
ns_connected = bool(state["connected"])
ns_attached = bool(state["attached"])
if ns_connected and ns_attached:
condition = None
elif ns_connected:
condition = ERROR_CONNECTED_NAMESPACES_MISSING
else:
condition = ERROR_REQUIRED_NAMESPACES_NOT_CONNECTED
namespace_conditions[ns_name] = {
"namespace": ns_name,
"required": ns_name in required,
"connected": ns_connected,
"attached": ns_attached,
"attachment_healthy": ns_connected and ns_attached,
"condition": condition,
"contradictory_evidence": ns_attached and not ns_connected,
}
# Startup-ordering race: a namespace that finished connecting *after* the session
# tool snapshot was taken cannot be in that snapshot, however healthy it looks now.
connected_at = {
str(k).strip(): v
for k, v in (namespace_connected_at or {}).items()
if str(k).strip() and isinstance(v, (int, float))
}
late_attaching: list[str] = []
if isinstance(session_tool_snapshot_at, (int, float)):
for ns_name, ts in connected_at.items():
if ts > session_tool_snapshot_at and ns_name not in attached:
late_attaching.append(ns_name)
late_attaching.sort()
startup_ordering_race = bool(late_attaching)
auto_recovered = bool(auto_attach_attempted and auto_attach_succeeded and attachment_healthy)
reconnect_required = not attachment_healthy
reasons: list[str] = []
remediation: list[str] = []
if not connected:
reasons.append("No MCP servers reported Connected.")
remediation.append("Start or reconnect Gitea MCP servers in client config.")
if missing:
reasons.append(
f"MCP server(s) {missing} report Connected at host/CLI layer but tool namespaces "
f"are missing from active session attached tools (Connected ≠ attached tools, #708)."
)
remediation.append(
"Reconnect the IDE/client MCP session to attach namespaces to the active session "
f"(sanctioned path: {SANCTIONED_ATTACH_RECOVERY_TOOL}), then re-run full preflight "
"(gitea_whoami -> gitea_resolve_task_capability -> task). "
"Do not use direct imports, CLI API mutations, profile hopping, or session file overrides."
)
if not_connected:
reasons.append(
f"required MCP namespace(s) {not_connected} are absent from the connected-service "
"inventory: nothing reports them Connected, so there is no attachment to claim and "
"the Connected-but-unattached condition does not apply to them (#708)."
)
remediation.append(
f"Connect the required MCP server(s) {not_connected} through the client, then "
"reconnect the IDE/client MCP session so their namespaces attach "
f"(sanctioned path: {SANCTIONED_ATTACH_RECOVERY_TOOL}), and re-run full preflight. "
"Do not use direct imports, CLI API mutations, profile hopping, or session file overrides."
)
if contradictory:
reasons.append(
f"contradictory evidence for namespace(s) {contradictory}: reported attached to the "
"active session while absent from the connected-service inventory; neither statement "
"is proof, so attachment is treated as unproven (fail closed, #708)."
)
if attachment_healthy:
reasons.append(
"All required MCP server namespaces are connected and attached to the active session."
)
if startup_ordering_race:
reasons.append(
f"startup ordering race: namespace(s) {late_attaching} finished connecting after the "
"active session tool snapshot was taken, so they cannot appear in that snapshot (#708)."
)
if auto_attach_attempted and not auto_attach_succeeded:
reasons.append(
"automatic namespace attachment was attempted and did not succeed; only the sanctioned "
"client reconnect path remains."
)
if auto_recovered:
reasons.append("namespaces were automatically attached; no operator reconnect was required.")
return {
"success": attachment_healthy,
"attachment_healthy": attachment_healthy,
"discovery_status": discovery_status,
"connected_servers": connected,
"attached_session_namespaces": list(attached),
"missing_namespaces": missing,
"not_connected_namespaces": not_connected,
"contradictory_namespaces": contradictory,
"proof_of_connected_vs_attached": proof,
"namespace_conditions": namespace_conditions,
"error_type": error_type,
"error_types": error_types,
"reasons": reasons,
"remediation": remediation,
"exact_next_action": (
"None; session tool namespaces attached."
if attachment_healthy
else (
f"Connect the required MCP server(s) {not_connected} through the client, then "
"reconnect the IDE/client MCP session so their tool namespaces attach to the "
"active session. Do not use direct imports, CLI API mutations, profile hopping, "
"or session-state overrides."
if not_connected
else (
"Reconnect the IDE/client MCP session so tool namespaces attach to the active "
"session. Do not use direct imports, CLI API mutations, profile hopping, or "
"session-state overrides."
)
)
),
"unsafe_fallback_policy": UNSAFE_FALLBACK_WARNING,
"sanctioned_recovery_tool": SANCTIONED_ATTACH_RECOVERY_TOOL,
"reconnect_required": reconnect_required,
"auto_attach_attempted": bool(auto_attach_attempted),
"auto_recovered": auto_recovered,
"startup_ordering_race": startup_ordering_race,
"late_attaching_namespaces": late_attaching,
# Secret-free structured signals (#708 AC5). Namespace names and counts only:
# never tokens, endpoints, env values, or filesystem paths.
"telemetry": {
"connected_count": len(connected),
"attached_count": len(attached),
"required_count": len(required),
"missing_count": len(missing),
"not_connected_count": len(not_connected),
"contradictory_count": len(contradictory),
"required_attached_count": sum(1 for r in required if r in attached),
"discovery_status": discovery_status,
"discovery_cache_hit": (
None if discovery_cache_hit is None else bool(discovery_cache_hit)
),
"discovery_cache_age_seconds": (
float(discovery_cache_age_seconds)
if isinstance(discovery_cache_age_seconds, (int, float))
else None
),
"reconnect_required": reconnect_required,
"auto_attach_attempted": bool(auto_attach_attempted),
"auto_recovered": auto_recovered,
"startup_ordering_race": startup_ordering_race,
"error_type": error_type,
"error_types": error_types,
},
}
def required_namespace_for_attachment(task: str) -> str | None:
"""Map a mutation task to the MCP namespace that must be *attached* (#708)."""
return ATTACHMENT_GATED_TASKS.get((task or "").strip())
def attachment_gate_from_session(
task: str,
session_attachment: dict[str, dict[str, Any]] | None,
) -> list[str]:
"""Fail-closed gate on recorded connected-but-unattached namespaces (#708).
Mirrors :func:`mutation_gate_from_session`: a namespace that has not been
assessed yet does not gate, so this never blocks a session that simply has
not run the assessment. Once an assessment records the namespace required
for *task* as Connected-but-unattached, the mutation fails closed and the
only offered recovery is the sanctioned client reconnect path.
"""
ns = required_namespace_for_attachment(task)
if not ns:
return []
store = session_attachment or {}
entry = store.get(ns)
if not entry:
return []
if entry.get("attached") and entry.get("attachment_healthy"):
return []
what = task or "mutation"
connected = entry.get("connected")
detail = entry.get("condition") or entry.get("error_type")
if connected is True:
# Only here has anything actually reported the service Connected, so only here may
# the block say so.
detail = detail or ERROR_CONNECTED_NAMESPACES_MISSING
blocked = (
f"live MCP namespace '{ns}' is recorded {detail}: the host reports Connected but the "
f"namespace is not attached to the active session tool surface; reconnect the "
f"IDE/client MCP session and re-run preflight before {what} "
"(fail closed, #708)"
)
elif connected is False:
detail = ERROR_REQUIRED_NAMESPACES_NOT_CONNECTED
blocked = (
f"live MCP namespace '{ns}' is recorded {detail}: it is absent from the "
f"connected-service inventory, so it is neither connected nor attached and no "
f"Connected status is claimed for it; connect the required MCP server, then "
f"reconnect the IDE/client MCP session and re-run preflight before {what} "
"(fail closed, #708)"
)
else:
# Connected status was never recorded. Refuse without asserting either condition.
detail = detail or ERROR_CONNECTED_NAMESPACES_MISSING
blocked = (
f"live MCP namespace '{ns}' is recorded {detail} with no connected-status evidence, "
f"so attachment to the active session tool surface is unproven; reconnect the "
f"IDE/client MCP session and re-run preflight before {what} "
"(fail closed, #708)"
)
return [blocked, UNSAFE_FALLBACK_WARNING]
def _as_list(value: Any) -> list[str] | None:
if value is None:
@@ -452,23 +104,12 @@ def classify_namespace_probe(
profile: str | None = None,
configured: bool = True,
probe_source: str | None = None,
worker_identity: str | None = None,
generation_id: str | None = None,
registry: Any | None = None,
pid_alive_probe: Any | None = None,
) -> dict[str, Any]:
"""Classify whether a required tool is callable through a live namespace.
``registered_tools`` is static/server-side evidence. ``probe_result`` is
live invocation evidence. Only ``probe_source=client_namespace`` proves the
IDE-managed path; ``offline_spawn`` is an offline subprocess check only.
#948: ``worker_identity``/``generation_id``/``registry`` carry the
client/session ownership evidence. Provenance is resolved by
``mcp_worker_identity.assess_provenance`` the same call
``gitea_get_runtime_context`` makes so the two surfaces cannot report
different provenance for one process. Omitting them yields the fail-closed
``unproven`` verdict, never a fabricated ``client_managed``.
"""
ns = (namespace or "").strip()
tool = required_tool or REQUIRED_NAMESPACE_TOOLS.get(ns) or "gitea_whoami"
@@ -584,33 +225,6 @@ def classify_namespace_probe(
# on bad data without treating success as IDE proof).
blocks = namespace_health_blocks_task("merge_pr", healthy)
import gitea_config
import mcp_worker_identity
raw_env = process.get("env") if isinstance(process, dict) else None
unconsumed_env = gitea_config.get_unconsumed_gitea_env_overrides(raw_env)
# #948: one authority, shared with gitea_get_runtime_context. The env is
# passed whole rather than through SAFE_ENV_KEYS — the allowlist exists to
# decide what may be *echoed*, and using it to decide what may be *believed*
# is what made this surface structurally unable to report client_managed.
# ``declared_only``: ``process`` describes an observed peer, not this
# interpreter. Its stdin is unavailable and its launcher-config env is
# inherited from whatever shell started it, so only an explicit declaration
# is evidence. Absence of one is ``unproven``, not an asserted manual launch.
provenance_verdict = mcp_worker_identity.assess_provenance(
registry=registry,
worker_identity=worker_identity,
generation_id=generation_id,
env=raw_env if isinstance(raw_env, dict) else {},
namespace=ns,
profile=profile_name,
pid_alive_probe=pid_alive_probe,
declared_only=True,
)
provenance = provenance_verdict["provenance"]
is_client_managed = provenance_verdict["is_client_managed"]
return {
"success": healthy,
"healthy": healthy,
@@ -626,18 +240,6 @@ def classify_namespace_probe(
"error_message": error_message or None,
"reasons": reasons,
"remediation": remediation,
"provenance": provenance,
"is_client_managed": is_client_managed,
# Every non-client-session verdict fails closed. Consumers that only
# need "may this mutate?" read this and stay correct across the #948
# vocabulary split between ``manual_launch`` and ``unproven``.
"provenance_fail_closed": provenance_verdict["fail_closed"],
"provenance_assessment": provenance_verdict,
"worker_identity": provenance_verdict["worker_identity"],
"session_id": provenance_verdict["session_id"],
"generation_id": provenance_verdict["generation_id"],
"client_name": provenance_verdict["client_name"],
"unconsumed_gitea_env": unconsumed_env,
"diagnostics": {
"namespace": ns,
"required_tool": tool,
@@ -646,16 +248,6 @@ def classify_namespace_probe(
"env": env_summary,
"config_path": config_path,
"probe_source": source,
"provenance": provenance,
"is_client_managed": is_client_managed,
"provenance_fail_closed": provenance_verdict["fail_closed"],
"provenance_blocker_kind": provenance_verdict["blocker_kind"],
"provenance_scope": provenance_verdict["scope"],
"worker_identity": provenance_verdict["worker_identity"],
"session_id": provenance_verdict["session_id"],
"generation_id": provenance_verdict["generation_id"],
"client_name": provenance_verdict["client_name"],
"unconsumed_gitea_env": unconsumed_env,
},
"blocks_merge_workflow": blocks,
}
-505
View File
@@ -1,505 +0,0 @@
"""Inventory and fail-closed guards for MCP restart/reload/kill paths (#657).
Single source of truth enumerating every code/script/doc path that can
restart, reload, reconnect, kill, or force-recreate an MCP process. Each path
is classified and linked to the guard that constrains it. The companion
human-readable inventory lives in ``docs/mcp-restart-path-inventory.md`` and is
kept in lock-step with this module by ``tests/test_mcp_restart_paths.py``.
Design intent (aligns with #655 restart-coordinator roadmap):
* **No unguarded full restart.** The in-process MCP daemon
(``gitea_mcp_server.py`` / ``mcp_server.py`` / ``role_session_router.py``)
must never replace or kill its own process replacing the process after the
host wired up the stdio pipes desyncs the JSON-RPC transport (observed with
Antigravity/Cascade hosts). ``assert_no_daemon_self_replacement`` enforces
this against the live source tree.
* **No legacy auto-restart helper.** ``_trigger_mcp_auto_restart`` was removed
when the stale-runtime resolver became side-effect free (#685);
``assert_auto_restart_helper_absent`` keeps it removed.
* **Unknown restart attempts fail closed.** LLM tools must route any restart
intent through a *registered* path. ``assert_restart_attempt_registered``
raises ``UnknownRestartPathError`` for anything not in this inventory.
* **pkill stays forbidden (#630).** Manual daemon kills are classified as
contamination by :mod:`runtime_recovery_guard`; this module records that path
and the test asserts the classification still holds.
This module performs no restarts, spawns no threads, and touches no config or
process state. It is pure inventory + read-only source assertions.
"""
from __future__ import annotations
import os
from dataclasses import dataclass
from pathlib import Path
from typing import Iterable
# --- Classifications -------------------------------------------------------
#: A narrow, one-shot recovery that is safe by construction (e.g. a CLI wrapper
#: re-execing into the venv interpreter before importing anything, or an
#: in-process profile switch). Never targets the running MCP daemon process.
CLASS_SANCTIONED_NARROW = "sanctioned_narrow_recovery"
#: The path detects a condition that would require a restart, then *fails
#: closed* on mutations and emits restart/reconnect guidance. It never restarts
#: the process itself (recovery is owned by the host/operator).
CLASS_GUARDED_FAIL_CLOSED = "guarded_fail_closed"
#: The path is forbidden. Attempting it is a workflow-safety violation and,
#: where an LLM tool could invoke it, is marked as contamination.
CLASS_FORBIDDEN = "forbidden"
#: A previously-existing unguarded restart primitive that has been deleted. A
#: regression guard keeps it absent.
CLASS_REMOVED = "removed"
#: Behavior that lives in the host/IDE and is outside this process's control
#: (e.g. a manual ``/mcp reconnect``). Documented, not code-guarded here.
CLASS_HOST_RESIDUAL = "host_residual"
VALID_CLASSIFICATIONS = frozenset(
{
CLASS_SANCTIONED_NARROW,
CLASS_GUARDED_FAIL_CLOSED,
CLASS_FORBIDDEN,
CLASS_REMOVED,
CLASS_HOST_RESIDUAL,
}
)
#: The in-process MCP daemon modules. These must never self-replace/self-kill.
DAEMON_MODULES = (
"gitea_mcp_server.py",
"mcp_server.py",
"role_session_router.py",
)
#: The legacy auto-restart helper removed in #685. Must stay removed.
LEGACY_AUTO_RESTART_HELPER = "_trigger_mcp_auto_restart"
#: Call patterns that would let the daemon replace or terminate its own
#: process. Matched as calls (trailing ``(``) so prose/docstring mentions such
#: as "we do NOT os.execv() here" or "never calls ``os._exit``" do not trip the
#: scanner (comment lines are stripped first regardless).
DAEMON_SELF_REPLACEMENT_PRIMITIVES = (
"os.execv(",
"os.execve(",
"os.execvp(",
"os.execvpe(",
"os.kill(",
"os.killpg(",
"os._exit(",
"os.abort(",
)
@dataclass(frozen=True)
class RestartPath:
"""One classified restart/reload/kill path in the inventory."""
path_id: str
title: str
mechanism: str
classification: str
guard: str
locations: tuple[str, ...]
references: tuple[str, ...]
residual_host: bool = False
notes: str = ""
class UnknownRestartPathError(RuntimeError):
"""Raised when a restart attempt is not a registered, classified path."""
# --- The inventory ---------------------------------------------------------
_RESTART_PATHS: tuple[RestartPath, ...] = (
RestartPath(
path_id="cli_venv_bootstrap_execv",
title="CLI wrapper venv re-exec",
mechanism=(
"Standalone CLI scripts re-exec into venv/bin/python3 via os.execv "
"at import top, guarded by `sys.executable != venv_python`."
),
classification=CLASS_SANCTIONED_NARROW,
guard=(
"One-shot, pre-import bootstrap; runs before any MCP transport "
"exists and only when not already on the venv interpreter, so it "
"cannot desync a live daemon. Idempotent guard condition prevents "
"a re-exec loop."
),
locations=(
"create_pr.py",
"create_issue.py",
"close_issue.py",
"merge_pr.py",
"review_pr.py",
"edit_pr.py",
"delete_branch.py",
"mark_issue.py",
"manage_labels.py",
"list_issues.py",
"list_prs.py",
),
references=("#657",),
),
RestartPath(
path_id="daemon_self_replacement",
title="MCP daemon self-replacement",
mechanism=(
"The in-process MCP daemon replacing/terminating its own process "
"(os.execv/os.kill/os._exit) to reload code."
),
classification=CLASS_FORBIDDEN,
guard=(
"Forbidden by design: replacing the process after the host wired "
"up stdio desyncs JSON-RPC (Antigravity/Cascade). Enforced against "
"the source tree by assert_no_daemon_self_replacement()."
),
locations=("gitea_mcp_server.py:~155 (decision comment)",) + DAEMON_MODULES,
references=("#657", "#584"),
),
RestartPath(
path_id="legacy_auto_restart_helper",
title="Legacy _trigger_mcp_auto_restart helper",
mechanism=(
"A helper that actively restarted the MCP server from the "
"read-only resolver path."
),
classification=CLASS_REMOVED,
guard=(
"Removed in #685 when the resolver became side-effect free. Kept "
"absent by assert_auto_restart_helper_absent()."
),
locations=("gitea_mcp_server.py", "mcp_server.py"),
references=("#685", "#657"),
),
RestartPath(
path_id="config_touch_reload",
title="MCP client config-touch reload",
mechanism=(
"Touching (utime) the MCP client config file to make the host "
"reload/recreate the server process."
),
classification=CLASS_REMOVED,
guard=(
"Removed from the resolver in #685: stale-runtime detection is "
"report-only and never mutates client config, spawns threads, or "
"calls os._exit."
),
locations=("gitea_mcp_server.py (resolve_task_capability)",),
references=("#685", "#657"),
),
RestartPath(
path_id="master_advance_auto_restart",
title="Master-advance staleness gate",
mechanism=(
"On-disk master advancing past the running code. The master-parity "
"gate detects it and fails mutations closed with restart guidance."
),
classification=CLASS_GUARDED_FAIL_CLOSED,
guard=(
"Detect + fail closed only; the process never self-restarts. "
"master_parity_gate captures startup parity and blocks mutations "
"while stale, emitting restart/reconnect guidance."
),
locations=(
"master_parity_gate.py",
"gitea_mcp_server.py (gitea_assess_master_parity)",
),
references=("#420", "#591", "#657"),
),
RestartPath(
path_id="stale_runtime_resolver_reconnect",
title="Stale-runtime resolver reconnect guidance",
mechanism=(
"The capability resolver detecting a stale serving process and "
"reporting restart_required/stop_required for a client reconnect."
),
classification=CLASS_GUARDED_FAIL_CLOSED,
guard=(
"Report-only (#685): returns restart_required/stop_required and an "
"exact_safe_next_action pointing at IDE/client reconnect; performs "
"no restart, thread spawn, config touch, or os._exit."
),
locations=(
"gitea_mcp_server.py (gitea_resolve_task_capability)",
"gitea_mcp_server.py (gitea_request_mcp_reconnect)",
"mcp_client_reconnect.py",
),
references=("#685", "#657", "#678"),
),
RestartPath(
path_id="codex_client_reconnect_request",
title="Sanctioned Codex/LLM reconnect request tool",
mechanism=(
"gitea_request_mcp_reconnect: agents invoke a report-only tool that "
"returns namespace/profile/pid/startup SHA/master SHA/boundary "
"status plus a typed operator blocker with exact client UI steps."
),
classification=CLASS_GUARDED_FAIL_CLOSED,
guard=(
"Report-only (#678): never restarts, kills, reloads, or edits "
"config; recovery is always host/operator reconnect. Forbidden "
"paths (pkill, touch, .env/config/session-state hacks) are listed "
"and never recommended."
),
locations=(
"mcp_client_reconnect.py",
"gitea_mcp_server.py (gitea_request_mcp_reconnect)",
),
references=("#678", "#630", "#685", "#657"),
),
RestartPath(
path_id="manual_daemon_kill",
title="Manual daemon kill (pkill/killall/kill)",
mechanism=(
"Shell kills of the MCP daemon: `pkill -f mcp_server.py`, "
"`killall`, broad `pkill -f python` sweeps, or `kill <pid>` of a "
"daemon pid."
),
classification=CLASS_FORBIDDEN,
guard=(
"Forbidden (#630): runtime_recovery_guard classifies these as "
"contamination and gitea_record_daemon_process_kill_attempt writes "
"a durable marker that fails subsequent mutations closed. Operator "
"maintenance authorization is read only from the environment, not "
"from a tool argument."
),
locations=(
"runtime_recovery_guard.py",
"gitea_mcp_server.py (gitea_record_daemon_process_kill_attempt)",
),
references=("#630", "#657"),
),
RestartPath(
path_id="conflict_marker_infra_stop",
title="Startup conflict-marker infra stop",
mechanism=(
"The daemon entrypoint scans for unresolved merge-conflict markers "
"at startup and stops (sys.exit(1)) if found."
),
classification=CLASS_GUARDED_FAIL_CLOSED,
guard=(
"Fail-closed startup stop, not a restart: the process exits and "
"waits for the operator to resolve conflicts and relaunch. Never "
"self-restarts or loops."
),
locations=("mcp_server.py (check_conflict_markers)",),
references=("#657",),
),
RestartPath(
path_id="ide_client_reconnect",
title="Host/IDE MCP reconnect",
mechanism=(
"A manual `/mcp reconnect` (or equivalent host action) that the "
"IDE performs to recreate the MCP client connection. Agents obtain "
"exact UI steps via gitea_request_mcp_reconnect (#678)."
),
classification=CLASS_HOST_RESIDUAL,
guard=(
"Outside this process's control. It is the sanctioned recovery the "
"gates point operators toward; documented as residual host "
"behavior. No in-process code initiates it."
),
locations=(
"host/IDE",
"mcp_client_reconnect.py",
"gitea_mcp_server.py (gitea_request_mcp_reconnect)",
),
references=("#584", "#656", "#657", "#678"),
residual_host=True,
),
RestartPath(
path_id="profile_switch_runtime",
title="Runtime profile switch",
mechanism=(
"Switching the active execution profile at runtime "
"(dynamic-profile mode)."
),
classification=CLASS_SANCTIONED_NARROW,
guard=(
"In-process and restart-free: runtime_switching_supported is true, "
"so a profile switch rebinds capability without recreating the "
"process. No restart primitive is invoked."
),
locations=("gitea_mcp_server.py (gitea_activate_profile)",),
references=("#656", "#657"),
),
)
_BY_ID: dict[str, RestartPath] = {p.path_id: p for p in _RESTART_PATHS}
# --- Read-only accessors ---------------------------------------------------
def iter_restart_paths() -> tuple[RestartPath, ...]:
"""Return the full inventory as an immutable tuple."""
return _RESTART_PATHS
def restart_path_ids() -> frozenset[str]:
"""Return the set of registered path ids."""
return frozenset(_BY_ID)
def get_restart_path(path_id: str) -> RestartPath:
"""Return the registered path, or raise :class:`UnknownRestartPathError`."""
try:
return _BY_ID[path_id]
except KeyError as exc:
raise UnknownRestartPathError(
f"unknown restart path id {path_id!r}; not in the #657 inventory"
) from exc
def paths_by_classification(classification: str) -> tuple[RestartPath, ...]:
"""Return all registered paths with the given classification."""
if classification not in VALID_CLASSIFICATIONS:
raise ValueError(f"unknown classification {classification!r}")
return tuple(p for p in _RESTART_PATHS if p.classification == classification)
def assert_restart_attempt_registered(path_id: str) -> RestartPath:
"""Fail closed unless ``path_id`` is a registered, classified restart path.
LLM tools that intend to trigger any restart/reload/reconnect must name a
registered path so an unknown/novel restart primitive cannot slip through
silently. Forbidden and removed paths are registered too this only
asserts the attempt is *known*, not that it is *permitted*; callers must
still honor the classification.
"""
return get_restart_path(path_id)
def assert_registry_wellformed() -> None:
"""Validate the inventory's own invariants (fail closed on drift)."""
seen: set[str] = set()
for path in _RESTART_PATHS:
if path.path_id in seen:
raise ValueError(f"duplicate restart path id {path.path_id!r}")
seen.add(path.path_id)
if path.classification not in VALID_CLASSIFICATIONS:
raise ValueError(
f"{path.path_id!r} has invalid classification "
f"{path.classification!r}"
)
if not path.guard.strip():
raise ValueError(f"{path.path_id!r} is missing a guard description")
if not path.references:
raise ValueError(f"{path.path_id!r} is missing references")
if not path.locations:
raise ValueError(f"{path.path_id!r} is missing locations")
if path.classification == CLASS_HOST_RESIDUAL and not path.residual_host:
raise ValueError(
f"{path.path_id!r} is host_residual but residual_host is False"
)
# --- Source-tree guards ----------------------------------------------------
def _repo_root(root: str | os.PathLike[str] | None = None) -> Path:
if root is not None:
return Path(root)
return Path(__file__).resolve().parent
def _iter_code_lines(text: str) -> Iterable[tuple[int, str]]:
"""Yield (1-based lineno, line) for lines that are not full-line comments."""
for lineno, line in enumerate(text.splitlines(), start=1):
if line.lstrip().startswith("#"):
continue
yield lineno, line
def scan_daemon_self_replacement(
root: str | os.PathLike[str] | None = None,
) -> list[dict[str, object]]:
"""Return violations where a daemon module could self-replace/self-kill.
Scans :data:`DAEMON_MODULES` for calls in
:data:`DAEMON_SELF_REPLACEMENT_PRIMITIVES`. Full-line comments are ignored,
and only call forms (with a trailing ``(``) match, so decision comments and
docstrings that merely mention the primitives do not produce false hits.
"""
repo = _repo_root(root)
violations: list[dict[str, object]] = []
for module in DAEMON_MODULES:
path = repo / module
if not path.exists():
continue
text = path.read_text(encoding="utf-8", errors="replace")
for lineno, line in _iter_code_lines(text):
for primitive in DAEMON_SELF_REPLACEMENT_PRIMITIVES:
if primitive in line:
violations.append(
{
"module": module,
"line": lineno,
"primitive": primitive,
"text": line.strip(),
}
)
return violations
def assert_no_daemon_self_replacement(
root: str | os.PathLike[str] | None = None,
) -> None:
"""Fail closed if any daemon module can restart/kill its own process."""
violations = scan_daemon_self_replacement(root)
if violations:
rendered = "; ".join(
f"{v['module']}:{v['line']} {v['primitive']}" for v in violations
)
raise AssertionError(
"MCP daemon must never self-replace/self-kill (#657); found: "
f"{rendered}"
)
def scan_auto_restart_helper(
root: str | os.PathLike[str] | None = None,
) -> list[dict[str, object]]:
"""Return occurrences of a *definition* of the legacy auto-restart helper."""
repo = _repo_root(root)
needle = f"def {LEGACY_AUTO_RESTART_HELPER}"
hits: list[dict[str, object]] = []
for module in DAEMON_MODULES:
path = repo / module
if not path.exists():
continue
text = path.read_text(encoding="utf-8", errors="replace")
for lineno, line in _iter_code_lines(text):
if needle in line:
hits.append({"module": module, "line": lineno})
return hits
def assert_auto_restart_helper_absent(
root: str | os.PathLike[str] | None = None,
) -> None:
"""Fail closed if the removed ``_trigger_mcp_auto_restart`` reappears."""
hits = scan_auto_restart_helper(root)
if hits:
rendered = "; ".join(f"{h['module']}:{h['line']}" for h in hits)
raise AssertionError(
f"{LEGACY_AUTO_RESTART_HELPER} was removed in #685 and must not "
f"return (#657); found definition at: {rendered}"
)
+3 -7
View File
@@ -1,10 +1,7 @@
#!/usr/bin/env python3
"""Gitea MCP Server — exposes Gitea operations as MCP tools.
The transport is selected by deployment configuration (GITEA_MCP_TRANSPORT) and
defaults to the local client-spawned transport when unset (#931); the permitted
set lives in mcp_transport_config. All tools authenticate via macOS keychain
(git credential fill).
Runs over stdio. All tools authenticate via macOS keychain (git credential fill).
"""
import os
import sys
@@ -46,9 +43,8 @@ check_conflict_markers()
# #558 / #695: claim the official entrypoint before loading mutation modules.
# This alone does NOT authorize mutations — gitea_mcp_server binds the live
# native MCP transport immediately before mcp.run, over the configured
# transport (#931). Import-only or offline launch without that bind fails
# closed on mutations.
# native MCP transport (stdio) immediately before mcp.run. Import-only or
# offline launch without that bind fails closed on mutations.
try:
import mcp_daemon_guard
-4
View File
@@ -588,13 +588,9 @@ def save_state(
bool(prov.get("production_native_mcp_transport")),
)
body.setdefault("transport", prov.get("transport"))
# #931: record the bound transport identifier itself, so a durable
# decision lock names which transport performed the mutation.
body.setdefault("bound_transport", prov.get("bound_transport"))
except Exception:
body.setdefault("native_mcp_transport", False)
body.setdefault("transport", "untrusted")
body.setdefault("bound_transport", None)
envelope = {
"kind": kind,
-292
View File
@@ -1,292 +0,0 @@
"""Single authoritative source for the bound MCP transport identifier (#931).
Before this module the transport was a literal, passed once at the bottom of
``gitea_mcp_server`` as ``bind_native_mcp_transport(transport="stdio")``. Every
guard that later asks "is this a trusted native session" resolves that question
through the value bound there, so the literal was effectively a constant in the
authorization chain rather than configuration.
This module is the seam. It owns three things and nothing else:
- the permitted set of transport identifiers,
- the default used when deployment configuration says nothing,
- the resolution of the configured value into a validated identifier.
It deliberately holds no state. The *bound* transport is pinned once, at bind
time, into the process-local native-runtime record owned by
:mod:`mcp_daemon_guard`, and is read back through
``mcp_daemon_guard.bound_transport()``. That split matters: configuration is
read exactly once, before any tool can dispatch, so a later environment change
cannot move the value a guard observes the same pinning rule already applied
to the session-state root under #695 AC2.
Nothing here consumes tool arguments, request bodies, or provenance fields. The
only input is the deployment environment, read at bind time.
Standing up a listener for a non-stdio transport is #938; this module only
makes the identifier expressible and validated.
"""
from __future__ import annotations
import os
from typing import Any, Mapping
# Deployment configuration key. Read once, at bind time, and never again.
TRANSPORT_ENV = "GITEA_MCP_TRANSPORT"
# The local, client-spawned transport. Unset configuration resolves to this,
# which is what keeps every existing stdio deployment byte-identical.
DEFAULT_TRANSPORT = "stdio"
# The sanctioned remote transport identifier. Accepting it here is what makes
# the bind pluggable; the endpoint that serves it belongs to #938. The name
# matches the MCP transport name so no second vocabulary has to be mapped.
REMOTE_TRANSPORT = "streamable-http"
# The permitted set. This is the only place transport identifiers are
# enumerated; guards consult it rather than restating any member.
#
# ``sse`` is a real MCP transport and is deliberately absent: it is the
# superseded remote transport, and admitting it would give the deployment two
# remote paths to reason about. An unregistered identifier must fail closed at
# bind time, and ``sse`` is held to that rule like any other.
SUPPORTED_TRANSPORTS = frozenset({DEFAULT_TRANSPORT, REMOTE_TRANSPORT})
# Recognition is not execution authorization (#931, review 635 B1/B2).
#
# SUPPORTED_TRANSPORTS answers "is this an identifier this system knows, and may
# it be bound, pinned and recorded?". It deliberately includes the remote
# identifier, because #931 requires the bind to become pluggable.
#
# EXECUTABLE_TRANSPORTS answers a strictly narrower question: "is this entrypoint
# commissioned to actually *serve* on that transport?". Only the local transport
# is. Handing ``streamable-http`` to ``mcp.run`` would start FastMCP's HTTP
# listener with no authentication, no TLS and no per-request principal — the
# endpoint #938 owns and gates. Recognition must therefore never imply execution.
#
# #938 commissions the remote listener by adding REMOTE_TRANSPORT here, together
# with the authentication and principal boundary its acceptance criteria require.
EXECUTABLE_TRANSPORTS = frozenset({DEFAULT_TRANSPORT})
# Which issue owns commissioning each recognized-but-not-executable transport.
# Used to make the refusal actionable rather than a generic denial.
TRANSPORT_EXECUTION_OWNER = {REMOTE_TRANSPORT: "#938"}
BLOCKER_TRANSPORT_NOT_BOUND = "transport_not_bound"
BLOCKER_TRANSPORT_NOT_RECOGNIZED = "transport_not_recognized"
BLOCKER_LISTENER_NOT_COMMISSIONED = "transport_listener_not_commissioned"
SOURCE_CONFIGURED = "deployment_configuration"
SOURCE_DEFAULT = "default"
class TransportConfigurationError(ValueError):
"""Raised when configured transport is outside :data:`SUPPORTED_TRANSPORTS`."""
def normalize_transport(value: Any) -> str:
"""Canonical form of a transport identifier; ``""`` when there is none.
Non-string values normalize to ``""`` rather than being coerced, so a
structured object smuggled in from a caller can never match a member of the
permitted set.
"""
if not isinstance(value, str):
return ""
return value.strip().lower()
def supported_transports() -> tuple[str, ...]:
"""Permitted identifiers, sorted, for messages and status payloads."""
return tuple(sorted(SUPPORTED_TRANSPORTS))
def is_supported_transport(value: Any) -> bool:
"""True when *value* normalizes to a member of the permitted set."""
return normalize_transport(value) in SUPPORTED_TRANSPORTS
def is_remote_transport(value: Any) -> bool:
"""True when *value* is a permitted transport that is not the local one."""
name = normalize_transport(value)
return name in SUPPORTED_TRANSPORTS and name != DEFAULT_TRANSPORT
def executable_transports() -> tuple[str, ...]:
"""Transports this entrypoint is commissioned to serve, sorted."""
return tuple(sorted(EXECUTABLE_TRANSPORTS))
def is_executable_transport(value: Any) -> bool:
"""True when *value* may actually be served by this entrypoint (#931).
Strictly narrower than :func:`is_supported_transport`. A recognized
identifier that is not executable is a correct, fully-bound configuration
whose listener has simply not been commissioned yet.
"""
return normalize_transport(value) in EXECUTABLE_TRANSPORTS
def assess_transport_execution(value: Any) -> dict[str, Any]:
"""Structured serve-authorization verdict for a bound transport (#931).
This is the decision that separates a *recognized* transport from one this
entrypoint may execute. It is deliberately a pure function of the bound
identifier so the serve path cannot reach a listener the deployment has not
commissioned.
Args:
value: The bound transport identifier, or ``None`` when unbound.
Returns:
dict with ``transport``, ``recognized``, ``executable``, ``allowed``,
``blocker_kind``, ``owner_issue``, ``reasons`` and
``exact_next_action``. ``allowed`` is true only for a bound, recognized,
commissioned transport.
"""
name = normalize_transport(value)
if not name:
return {
"transport": None,
"recognized": False,
"executable": False,
"allowed": False,
"blocker_kind": BLOCKER_TRANSPORT_NOT_BOUND,
"owner_issue": None,
"supported_transports": list(supported_transports()),
"executable_transports": list(executable_transports()),
"reasons": [
"no transport is bound; the serve path is fail-closed until "
"bind_native_mcp_transport succeeds (#695/#931)"
],
"exact_next_action": (
"Launch through the canonical entrypoint so "
"bind_native_mcp_transport runs before tool service."
),
}
recognized = name in SUPPORTED_TRANSPORTS
if not recognized:
return {
"transport": name,
"recognized": False,
"executable": False,
"allowed": False,
"blocker_kind": BLOCKER_TRANSPORT_NOT_RECOGNIZED,
"owner_issue": None,
"supported_transports": list(supported_transports()),
"executable_transports": list(executable_transports()),
"reasons": [
f"transport {name!r} is not a registered MCP transport (#931); "
"it should have been refused at bind time"
],
"exact_next_action": (
f"Set {TRANSPORT_ENV} to one of {list(supported_transports())}."
),
}
if name in EXECUTABLE_TRANSPORTS:
return {
"transport": name,
"recognized": True,
"executable": True,
"allowed": True,
"blocker_kind": None,
"owner_issue": None,
"supported_transports": list(supported_transports()),
"executable_transports": list(executable_transports()),
"reasons": [],
"exact_next_action": None,
}
owner = TRANSPORT_EXECUTION_OWNER.get(name)
owner_text = owner or "the issue that commissions this transport's listener"
return {
"transport": name,
"recognized": True,
"executable": False,
"allowed": False,
"blocker_kind": BLOCKER_LISTENER_NOT_COMMISSIONED,
"owner_issue": owner,
"supported_transports": list(supported_transports()),
"executable_transports": list(executable_transports()),
"reasons": [
f"transport {name!r} is registered and was bound and recorded, but "
f"this entrypoint is not commissioned to serve it (#931). Serving it "
f"would start a listener with no authentication, no transport "
f"security and no per-request principal; that endpoint is owned by "
f"{owner_text}."
],
"exact_next_action": (
f"Serve on {DEFAULT_TRANSPORT} until {owner_text} commissions the "
f"{name!r} listener with its authentication and principal boundary, "
f"which adds {name!r} to EXECUTABLE_TRANSPORTS."
),
}
def resolve_configured_transport(
env: Mapping[str, str] | None = None,
) -> dict[str, Any]:
"""Resolve the deployment-configured transport without raising.
Returns the resolution rather than a bare string so a caller can tell an
unset value (which legitimately yields :data:`DEFAULT_TRANSPORT`) from a
configured value that is not permitted (which must fail closed, never
silently degrade to the default).
Args:
env: Environment mapping to read; defaults to ``os.environ``.
Returns:
dict with ``transport`` (normalized; the default when unset),
``configured``, ``source``, ``raw``, ``supported``, and ``reasons``.
"""
source_env = os.environ if env is None else env
raw = source_env.get(TRANSPORT_ENV)
normalized = normalize_transport(raw)
configured = bool(normalized)
if not configured:
return {
"transport": DEFAULT_TRANSPORT,
"configured": False,
"source": SOURCE_DEFAULT,
"raw": raw,
"supported": True,
"supported_transports": list(supported_transports()),
"env_key": TRANSPORT_ENV,
"reasons": [],
}
supported = normalized in SUPPORTED_TRANSPORTS
reasons: list[str] = []
if not supported:
reasons.append(
f"{TRANSPORT_ENV}={normalized!r} is not a registered MCP transport "
f"(#931). Registered: {list(supported_transports())}."
)
return {
"transport": normalized,
"configured": True,
"source": SOURCE_CONFIGURED,
"raw": raw,
"supported": supported,
"supported_transports": list(supported_transports()),
"env_key": TRANSPORT_ENV,
"reasons": reasons,
}
def require_configured_transport(env: Mapping[str, str] | None = None) -> str:
"""Resolved transport identifier, or raise when it is not permitted.
Raises:
TransportConfigurationError: the configured identifier is unregistered.
"""
resolution = resolve_configured_transport(env)
if not resolution["supported"]:
raise TransportConfigurationError("; ".join(resolution["reasons"]))
return str(resolution["transport"])
File diff suppressed because it is too large Load Diff
-58
View File
@@ -566,10 +566,6 @@ def build_pr_cleanup_entry(
worktree_state=worktree_state,
active_lock=active_lock,
)
planned = plan_cleanup_execution_order(
remote_assessment=remote,
local_assessment=local,
)
return {
"pr_number": pr_number,
"issue_number": issue_number,
@@ -580,63 +576,9 @@ def build_pr_cleanup_entry(
"merged": merged,
"remote_branch": remote,
"local_worktree": local,
# #851: dry-run and execute share the same lifecycle order description.
"planned_execution_order": planned,
}
def plan_cleanup_execution_order(
*,
remote_assessment: dict[str, Any] | None,
local_assessment: dict[str, Any] | None,
) -> list[dict[str, Any]]:
"""Describe independent worktree-then-reassess-then-remote cleanup order (#851).
Remote ownership protection remains fail-closed at execute time. A worktree
that is independently safe to remove is never skipped merely because remote
deletion may be blocked by that same ``worktree_binding``.
"""
remote = remote_assessment or {}
local = local_assessment or {}
steps: list[dict[str, Any]] = []
worktree_safe = bool(local.get("safe_to_remove_worktree"))
remote_safe = bool(remote.get("safe_to_delete_remote"))
if worktree_safe:
steps.append(
{
"action": "remove_local_worktree",
"reason": "independently_safe_to_remove",
"phase": 1,
}
)
if remote_safe:
if worktree_safe:
steps.append(
{
"action": "reassess_branch_ownership",
"reason": "after_worktree_removal_clear_worktree_binding",
"phase": 2,
}
)
steps.append(
{
"action": "delete_remote_branch",
"reason": "only_if_independently_safe_after_reassessment",
"phase": 3,
}
)
else:
steps.append(
{
"action": "delete_remote_branch",
"reason": "safe_to_delete_and_no_independent_worktree_removal",
"phase": 1,
}
)
return steps
def build_reconciliation_report(
*,
project_root: str,
File diff suppressed because it is too large Load Diff
+27 -153
View File
@@ -8,10 +8,8 @@ poison workspace purity checks in another namespace.
from __future__ import annotations
import os
import subprocess
import author_mutation_worktree as amw
import canonical_repository_root as crr
ACTIVE_WORKTREE_ENV = amw.ACTIVE_WORKTREE_ENV
AUTHOR_WORKTREE_ENV = amw.AUTHOR_WORKTREE_ENV
@@ -154,60 +152,6 @@ def resolve_namespace_workspace(
return os.path.realpath(process_project_root), "MCP server process root (default)"
def verify_git_common_directory_membership(
workspace_path: str,
canonical_repo_root: str,
) -> tuple[bool, str | None]:
"""Verify that workspace_path belongs to canonical_repo_root via git common-dir or branches containment."""
ws = (workspace_path or "").strip()
root = (canonical_repo_root or "").strip()
if not ws or not root:
return False, "empty workspace or canonical root path"
try:
real_ws = os.path.realpath(os.path.abspath(ws))
real_root = os.path.realpath(os.path.abspath(root))
except Exception as exc:
return False, f"invalid workspace or root path: {exc}"
if real_ws == real_root:
return True, None
if not os.path.isdir(real_ws):
return False, f"workspace directory '{real_ws}' does not exist"
try:
res = subprocess.run(
["git", "-C", real_ws, "rev-parse", "--git-common-dir"],
capture_output=True,
text=True,
check=False,
)
if res.returncode == 0:
common_raw = (res.stdout or "").strip()
common_dir = amw._realpath_git_common_dir(real_ws, common_raw)
real_common = os.path.realpath(common_dir)
canonical_git = os.path.realpath(os.path.join(real_root, ".git"))
if real_common in (canonical_git, real_root):
return True, None
return (
False,
f"workspace '{real_ws}' git common directory '{real_common}' does not match "
f"canonical repository root '{real_root}' (.git at '{canonical_git}')"
)
else:
return (
False,
f"workspace '{real_ws}' is not a valid git repository or git rev-parse failed"
)
except Exception as exc:
return (
False,
f"failed to inspect git common directory for workspace '{real_ws}': {exc}"
)
def resolve_namespace_mutation_context(
*,
role_kind: str,
@@ -219,8 +163,6 @@ def resolve_namespace_mutation_context(
worktree: str | None = None,
profile_name: str | None = None,
configured_canonical_root: str | None = None,
expected_slug: str | None = None,
remote: str | None = None,
) -> dict:
"""Shared workspace resolution for runtime_context and mutation guards.
@@ -238,41 +180,11 @@ def resolve_namespace_mutation_context(
env_map = env if env is not None else os.environ
process_root = os.path.realpath(process_project_root)
role = normalize_role_kind(role_kind, profile_name=profile_name)
configured_val = (configured_canonical_root or "").strip()
if configured_val:
crr_assessment = crr.assess_canonical_repository_root(
configured_value=configured_val,
source="configured_canonical_root",
expected_slug=expected_slug,
process_project_root=process_root,
remote=remote,
require_binding=True,
)
# #973 B10: a refused repository-authority mode deliberately resolves no
# canonical root, so downstream guards keep evaluating the install
# checkout rather than an identity derived through an undefined mode.
# ``roots_aligned`` still follows ``proven`` and stays False, and the
# refusal (with its reason_code) rides along in
# ``canonical_root_assessment`` — this is a fail-closed fallback, never a
# normalisation of the mode.
canonical_root = crr_assessment["canonical_repo_root"] or amw.resolve_canonical_repo_root(
process_root, process_root
)
roots_aligned = crr_assessment["proven"]
configured = (configured_canonical_root or "").strip()
if configured:
canonical_root = os.path.realpath(configured)
else:
crr_assessment = {
"proven": True,
"block": False,
"reasons": [],
"configured": False,
"canonical_repo_root": amw.resolve_canonical_repo_root(process_root, process_root),
"resolved_slug": None,
"source": None,
"reason_code": None,
}
canonical_root = crr_assessment["canonical_repo_root"]
roots_aligned = (canonical_root == process_root)
canonical_root = amw.resolve_canonical_repo_root(process_root, process_root)
durable: dict | None = None
if role == "author":
@@ -322,10 +234,7 @@ def resolve_namespace_mutation_context(
"ignored_bindings": demotions + (pollution.get("ignored_bindings") or []),
"process_project_root": process_root,
"canonical_repo_root": canonical_root,
"roots_aligned": roots_aligned,
"canonical_root_assessment": crr_assessment,
"expected_slug": expected_slug,
"remote": remote,
"roots_aligned": canonical_root == process_root,
}
if durable is not None:
result["author_worktree_resolution"] = durable
@@ -337,15 +246,6 @@ def resolve_namespace_mutation_context(
result["author_worktree_reasons"] = list(durable.get("reasons") or [])
result["author_worktree_blocker_kind"] = durable.get("blocker_kind")
result["operator_recovery"] = durable.get("operator_recovery")
else:
path_exists = os.path.exists(workspace)
result["path_exists"] = path_exists
result["in_git_worktree_list"] = (
amw.path_in_git_worktree_list(workspace, canonical_root)
if path_exists
else False
)
result["bound_worktree_missing"] = not path_exists
return result
@@ -478,8 +378,8 @@ def format_namespace_workspace_binding_error(
def assess_namespace_mutation_workspace(
*,
role_kind: str,
worktree_path: str | None = None,
worktree: str | None = None,
worktree_path: str | None,
worktree: str | None,
process_project_root: str,
env: dict[str, str] | os._Environ | None = None,
session_lease_worktree: str | None = None,
@@ -487,8 +387,6 @@ def assess_namespace_mutation_workspace(
profile_name: str | None = None,
current_branch: str | None = None,
configured_canonical_root: str | None = None,
expected_slug: str | None = None,
remote: str | None = None,
) -> dict:
"""Evaluate namespace workspace binding before preflight/mutation."""
ctx = resolve_namespace_mutation_context(
@@ -501,8 +399,6 @@ def assess_namespace_mutation_workspace(
session_lock_worktree=session_lock_worktree,
profile_name=profile_name,
configured_canonical_root=configured_canonical_root,
expected_slug=expected_slug,
remote=remote,
)
mutation_workspace = ctx["workspace_path"]
binding_source = ctx["workspace_binding_source"]
@@ -526,27 +422,6 @@ def assess_namespace_mutation_workspace(
reasons = list(metadata.get("reasons") or [])
operator_recovery = ctx.get("operator_recovery")
crr_reasons = list(ctx.get("canonical_root_assessment", {}).get("reasons") or [])
if crr_reasons:
reasons.extend(crr_reasons)
path_exists = ctx.get("path_exists")
if path_exists is None:
path_exists = os.path.exists(mutation_workspace)
if not path_exists:
if role != "author":
reasons.append(
f"{role} mutation blocked: configured workspace directory '{mutation_workspace}' does not exist (nonexistent worktree)"
)
else:
valid_common, common_err = verify_git_common_directory_membership(
mutation_workspace, ctx["canonical_repo_root"]
)
if not valid_common and common_err:
reasons.append(common_err)
if role == "author":
# #618 durable resolution already validated existence, membership,
# branches/, lock ownership, and traversal safety when present.
@@ -563,27 +438,26 @@ def assess_namespace_mutation_workspace(
)
if branches["block"]:
reasons.extend(branches["reasons"])
elif role in {"reviewer", "merger"}:
if mutation_workspace == process_root and not amw.is_path_under_branches(mutation_workspace, ctx["canonical_repo_root"]):
reasons.append(
f"{role} mutation blocked: workspace is the stable control checkout; "
f"create or reconnect to a session-owned worktree under branches/ "
f"or set {ROLE_WORKTREE_ENVS.get(role, ACTIVE_WORKTREE_ENV)} / "
f"{ACTIVE_WORKTREE_ENV}"
)
elif mutation_workspace != process_root and not amw.is_path_under_branches(mutation_workspace, ctx["canonical_repo_root"]):
reasons.append(
f"{role} mutation blocked: workspace '{mutation_workspace}' is not under "
f"'{ctx['canonical_repo_root']}/branches/'"
)
if path_exists and amw.is_path_under_branches(mutation_workspace, ctx["canonical_repo_root"]):
in_list = ctx.get("in_git_worktree_list")
if in_list is False:
reasons.append(
f"{role} mutation blocked: workspace '{mutation_workspace}' is under branches/ "
f"but is not registered in git worktree list for '{ctx['canonical_repo_root']}'"
)
elif (
role == "reviewer"
and mutation_workspace == process_root
and not amw.is_path_under_branches(mutation_workspace, ctx["canonical_repo_root"])
):
reasons.append(
f"{role} mutation blocked: workspace is the stable control checkout; "
f"create or reconnect to a session-owned worktree under branches/ "
f"or set {ROLE_WORKTREE_ENVS.get(role, ACTIVE_WORKTREE_ENV)} / "
f"{ACTIVE_WORKTREE_ENV}"
)
elif (
role in {"reviewer", "merger"}
and mutation_workspace != process_root
and not amw.is_path_under_branches(mutation_workspace, ctx["canonical_repo_root"])
):
reasons.append(
f"{role} mutation blocked: workspace '{mutation_workspace}' is not under "
f"'{ctx['canonical_repo_root']}/branches/'"
)
block = bool(reasons)
return {
-791
View File
@@ -1,791 +0,0 @@
"""Post-restart MCP reconciliation and completion proof (#662).
After an MCP process restart, sessions, leases, capabilities, worktrees, and
interrupted mutations are not systematically reconciled; operators rebuild
context from chat. This module is the pure classification core of the
post-restart reconcile path.
Design rules (mirrors ``restart_coordinator`` / ``workflow_dashboard``):
* **Pure classification.** :func:`reconcile_after_restart` takes an already
gathered inventory and returns a structured *completion proof*. It never
touches the network, the filesystem, or a live process, so multi-session
fixtures can drive every branch in unit tests.
* **Fail closed.** Incomplete inventory never reports overall ``complete``.
Ambiguous interrupted mutations are ``unresolved`` (never silently resumed).
* **No blind write resume.** The proof never authorizes replaying a mutation;
it only classifies evidence and names follow-up work.
* **#660 soft dependency.** When durable session checkpoints are not present
in the inventory, the checkpoint dimension is ``skipped`` with an explicit
reason rather than inventing a schema (#660 lands separately).
* **Log-only then enforce.** Default mode is ``log_only``. ``enforce`` sets
``mutation_hold`` when anything remains unresolved so callers can block
write ops until reconcile is complete or degraded mode is documented.
The single sanctioned gather+classify entry point is the MCP tool
``gitea_reconcile_after_restart`` (read-only inventory gather + pure classify).
Creating durable follow-up Gitea issues from unresolved items is an explicit
apply step outside this pure module.
"""
from __future__ import annotations
from dataclasses import dataclass, field
from datetime import datetime, timezone
from typing import Any, Mapping, Sequence
from uuid import uuid4
import lease_lifecycle
RECONCILE_VERSION = "1.0.0-issue-662"
SCHEMA_VERSION = 1
# Overall proof statuses.
STATUS_COMPLETE = "complete"
STATUS_DEGRADED = "degraded"
STATUS_FAILED = "failed"
# Per-dimension item statuses.
ITEM_RESOLVED = "resolved"
ITEM_UNRESOLVED = "unresolved"
ITEM_DEGRADED = "degraded"
ITEM_SKIPPED = "skipped"
# Modes.
MODE_LOG_ONLY = "log_only"
MODE_ENFORCE = "enforce"
# Lease / session phases that imply a write critical section was in flight.
MUTATING_PHASES = frozenset(
{
"implementing",
"publishing",
"merging",
"reviewing",
"committing",
"pushing",
"closing",
"mutating",
"critical_section",
}
)
# Dimensions the acceptance criteria require.
DIM_SERVICE_HEALTH = "service_health"
DIM_CLIENTS = "clients"
DIM_SESSIONS = "sessions"
DIM_CHECKPOINTS = "checkpoints"
DIM_LEASES = "leases"
DIM_CAPABILITIES = "capabilities"
DIM_WORKTREES = "worktrees"
DIM_MUTATIONS = "interrupted_mutations"
DIM_DUPLICATES = "duplicates"
DIM_QUEUE = "queue"
def _utc_now() -> datetime:
return datetime.now(timezone.utc)
def _ts(dt: datetime) -> str:
return dt.astimezone(timezone.utc).isoformat().replace("+00:00", "Z")
@dataclass(frozen=True)
class ReconcileItem:
"""One dimension of the post-restart reconcile report."""
dimension: str
status: str
summary: str
details: dict[str, Any] = field(default_factory=dict)
follow_up_required: bool = False
def as_dict(self) -> dict[str, Any]:
return {
"dimension": self.dimension,
"status": self.status,
"summary": self.summary,
"details": dict(self.details),
"follow_up_required": self.follow_up_required,
}
@dataclass(frozen=True)
class FollowUpIssue:
"""A durable follow-up issue the apply path may create for unresolved work."""
title: str
body: str
dimension: str
severity: str = "high"
def as_dict(self) -> dict[str, Any]:
return {
"title": self.title,
"body": self.body,
"dimension": self.dimension,
"severity": self.severity,
}
@dataclass(frozen=True)
class RestartCompletionProof:
"""Machine-readable post-restart completion proof (#662 AC2)."""
schema_version: int
reconcile_version: str
reconcile_id: str
started_at: str
finished_at: str
boot_head_sha: str | None
current_head_sha: str | None
inventory_complete: bool
incomplete_reasons: tuple[str, ...]
mode: str
mutation_hold: bool
overall_status: str
items: tuple[ReconcileItem, ...]
proposed_follow_ups: tuple[FollowUpIssue, ...]
resolved_count: int
unresolved_count: int
skipped_count: int
note: str
def as_dict(self) -> dict[str, Any]:
return {
"schema_version": self.schema_version,
"reconcile_version": self.reconcile_version,
"reconcile_id": self.reconcile_id,
"started_at": self.started_at,
"finished_at": self.finished_at,
"boot_head_sha": self.boot_head_sha,
"current_head_sha": self.current_head_sha,
"inventory_complete": self.inventory_complete,
"incomplete_reasons": list(self.incomplete_reasons),
"mode": self.mode,
"mutation_hold": self.mutation_hold,
"overall_status": self.overall_status,
"items": [i.as_dict() for i in self.items],
"proposed_follow_ups": [f.as_dict() for f in self.proposed_follow_ups],
"resolved_count": self.resolved_count,
"unresolved_count": self.unresolved_count,
"skipped_count": self.skipped_count,
"note": self.note,
"links": {
"umbrella": 655,
"vision": 652,
"roadmap": 653,
"issue": 662,
"checkpoint_schema": 660,
"drain_proof": 661,
},
}
def _item(
dimension: str,
status: str,
summary: str,
*,
details: dict[str, Any] | None = None,
follow_up: bool = False,
) -> ReconcileItem:
return ReconcileItem(
dimension=dimension,
status=status,
summary=summary,
details=dict(details or {}),
follow_up_required=follow_up,
)
def _lease_freshness(lease: Mapping[str, Any]) -> str:
fr = lease.get("freshness")
if isinstance(fr, Mapping):
return str(fr.get("freshness") or fr.get("status") or "unknown")
if isinstance(fr, str):
return fr
# Fall back to pure classifier when raw lease rows are supplied.
try:
return str(lease_lifecycle.classify_lease_freshness(dict(lease)).get("freshness") or "unknown")
except Exception: # noqa: BLE001 - pure path must not raise on bad rows
return "unknown"
def _is_live_freshness(freshness: str) -> bool:
return freshness in {"active", "live", "fresh"}
def _is_mutating_phase(phase: str | None) -> bool:
p = (phase or "").strip().lower()
if not p:
return False
if p in MUTATING_PHASES:
return True
# Soft match for compound phases like "author_implementing".
return any(token in p for token in MUTATING_PHASES)
def _detect_interrupted_mutations(
leases: Sequence[Mapping[str, Any]],
pending_mutations: Sequence[Mapping[str, Any]],
) -> list[dict[str, Any]]:
"""Return interrupted-mutation evidence (never auto-resumes writes)."""
found: list[dict[str, Any]] = []
for raw in pending_mutations or ():
if not isinstance(raw, Mapping):
continue
found.append(
{
"source": "pending_mutation_inventory",
"status": "unresolved",
"phase": raw.get("phase"),
"session_id": raw.get("session_id"),
"work_kind": raw.get("work_kind") or raw.get("kind"),
"work_number": raw.get("work_number") or raw.get("number"),
"reason": raw.get("reason")
or "pending mutation recorded across process restart",
"resume_allowed": False,
}
)
for lease in leases or ():
if not isinstance(lease, Mapping):
continue
phase = lease.get("phase")
freshness = _lease_freshness(lease)
if not _is_mutating_phase(str(phase) if phase is not None else None):
continue
# A mutating phase whose owner is not live is interrupted.
if _is_live_freshness(freshness):
# Still live after restart is itself surprising — flag for review.
found.append(
{
"source": "lease_mutating_phase",
"status": "unresolved",
"phase": phase,
"freshness": freshness,
"lease_id": lease.get("lease_id"),
"session_id": lease.get("session_id"),
"work_kind": lease.get("work_kind"),
"work_number": lease.get("work_number"),
"worktree_path": lease.get("worktree_path"),
"reason": (
"mutating lease phase still classified live after restart; "
"do not auto-resume writes"
),
"resume_allowed": False,
}
)
else:
found.append(
{
"source": "lease_mutating_phase",
"status": "unresolved",
"phase": phase,
"freshness": freshness,
"lease_id": lease.get("lease_id"),
"session_id": lease.get("session_id"),
"work_kind": lease.get("work_kind"),
"work_number": lease.get("work_number"),
"worktree_path": lease.get("worktree_path"),
"reason": (
f"mutating lease phase '{phase}' with non-live freshness "
f"'{freshness}' — interrupted by restart"
),
"resume_allowed": False,
}
)
return found
def _detect_duplicate_work(
leases: Sequence[Mapping[str, Any]],
) -> list[dict[str, Any]]:
"""Surface duplicate live claims on the same work item."""
by_work: dict[tuple[Any, Any], list[Mapping[str, Any]]] = {}
for lease in leases or ():
if not isinstance(lease, Mapping):
continue
if not _is_live_freshness(_lease_freshness(lease)):
continue
key = (lease.get("work_kind"), lease.get("work_number"))
if key[0] is None or key[1] is None:
continue
by_work.setdefault(key, []).append(lease)
dups: list[dict[str, Any]] = []
for (kind, number), rows in sorted(by_work.items(), key=lambda kv: str(kv[0])):
if len(rows) < 2:
continue
dups.append(
{
"work_kind": kind,
"work_number": number,
"claim_count": len(rows),
"session_ids": [r.get("session_id") for r in rows],
"lease_ids": [r.get("lease_id") for r in rows],
}
)
return dups
def _follow_up_for_item(item: ReconcileItem) -> FollowUpIssue | None:
if not item.follow_up_required:
return None
title = f"[post-restart] unresolved {item.dimension} after MCP restart"
body = (
f"## Post-restart reconcile follow-up (#662)\n\n"
f"**Dimension:** `{item.dimension}`\n"
f"**Status:** `{item.status}`\n"
f"**Summary:** {item.summary}\n\n"
f"```json\n{item.details!r}\n```\n\n"
f"Parent umbrella: #655 · Vision: #652 · Roadmap: #653 · Reconcile: #662\n"
f"Do **not** auto-resume write mutations; reconcile evidence first.\n"
)
return FollowUpIssue(
title=title,
body=body,
dimension=item.dimension,
severity="high" if item.dimension == DIM_MUTATIONS else "medium",
)
def reconcile_after_restart(
inventory: Mapping[str, Any],
*,
now: datetime | None = None,
mode: str = MODE_LOG_ONLY,
reconcile_id: str | None = None,
) -> RestartCompletionProof:
"""Classify a post-restart inventory into a completion proof (#662).
Parameters
----------
inventory:
Gathered facts. Expected keys (all optional except completeness):
* ``inventory_complete`` (bool) fail closed when false
* ``incomplete_reasons`` (list[str])
* ``service_health`` (dict with ``healthy`` bool)
* ``clients`` (list) connected client descriptors
* ``sessions`` (list)
* ``leases`` (list, optionally with ``freshness``)
* ``checkpoints`` (list | None) durable session checkpoints (#660)
* ``checkpoints_available`` (bool) False when #660 schema absent
* ``worktree_bindings`` (list)
* ``pending_mutations`` (list) explicit interrupted-mutation evidence
* ``capabilities`` (dict with optional ``stale`` / heads)
* ``boot_head_sha`` / ``current_head_sha``
* ``queue_state`` (dict)
mode:
``log_only`` (default) or ``enforce`` (sets mutation_hold on unresolved).
"""
started = now or _utc_now()
mode_norm = (mode or MODE_LOG_ONLY).strip().lower()
if mode_norm not in {MODE_LOG_ONLY, MODE_ENFORCE}:
mode_norm = MODE_LOG_ONLY
inventory_complete = bool(inventory.get("inventory_complete", False))
incomplete_reasons = tuple(
str(r) for r in (inventory.get("incomplete_reasons") or []) if str(r).strip()
)
items: list[ReconcileItem] = []
# --- service health -------------------------------------------------
health = inventory.get("service_health") or {}
if not isinstance(health, Mapping):
health = {}
if not inventory_complete and "service_health" not in inventory:
items.append(
_item(
DIM_SERVICE_HEALTH,
ITEM_UNRESOLVED,
"service health unknown because inventory is incomplete",
details={"inventory_complete": False},
follow_up=True,
)
)
elif health.get("healthy") is True:
items.append(
_item(
DIM_SERVICE_HEALTH,
ITEM_RESOLVED,
"service health verified",
details=dict(health),
)
)
elif health.get("healthy") is False:
items.append(
_item(
DIM_SERVICE_HEALTH,
ITEM_UNRESOLVED,
"service health check failed",
details=dict(health),
follow_up=True,
)
)
else:
items.append(
_item(
DIM_SERVICE_HEALTH,
ITEM_DEGRADED,
"service health not reported; treating as degraded",
details=dict(health),
follow_up=True,
)
)
# --- clients --------------------------------------------------------
clients = list(inventory.get("clients") or [])
disconnected = [
c
for c in clients
if isinstance(c, Mapping) and c.get("connected") is False
]
if "clients" not in inventory:
items.append(
_item(
DIM_CLIENTS,
ITEM_SKIPPED,
"client inventory not supplied",
details={},
)
)
elif disconnected:
items.append(
_item(
DIM_CLIENTS,
ITEM_UNRESOLVED,
f"{len(disconnected)} disconnected client(s) need reconnect",
details={"disconnected": disconnected, "total": len(clients)},
follow_up=True,
)
)
else:
items.append(
_item(
DIM_CLIENTS,
ITEM_RESOLVED,
f"{len(clients)} client(s) accounted for",
details={"total": len(clients)},
)
)
# --- sessions -------------------------------------------------------
sessions = [s for s in (inventory.get("sessions") or []) if isinstance(s, Mapping)]
orphan_sessions = [
s
for s in sessions
if str(s.get("status") or "").lower() == "active"
and s.get("pid") is not None
and not lease_lifecycle.is_process_alive(s.get("pid"))
]
if orphan_sessions:
items.append(
_item(
DIM_SESSIONS,
ITEM_UNRESOLVED,
f"{len(orphan_sessions)} active session row(s) with dead owner pid",
details={
"orphan_session_ids": [s.get("session_id") for s in orphan_sessions],
"total_sessions": len(sessions),
},
follow_up=True,
)
)
else:
items.append(
_item(
DIM_SESSIONS,
ITEM_RESOLVED,
f"{len(sessions)} session row(s) reconciled (no dead-pid orphans)",
details={"total_sessions": len(sessions)},
)
)
# --- checkpoints (#660 soft) ----------------------------------------
checkpoints_available = inventory.get("checkpoints_available")
checkpoints = inventory.get("checkpoints")
if checkpoints_available is False or (
checkpoints is None and "checkpoints" not in inventory
):
items.append(
_item(
DIM_CHECKPOINTS,
ITEM_SKIPPED,
"durable session checkpoint schema not available yet (#660)",
details={"depends_on": 660},
)
)
else:
cp_list = [c for c in (checkpoints or []) if isinstance(c, Mapping)]
stale_cp = [c for c in cp_list if c.get("stale") or c.get("invalid")]
if stale_cp:
items.append(
_item(
DIM_CHECKPOINTS,
ITEM_UNRESOLVED,
f"{len(stale_cp)} checkpoint(s) invalid or stale vs live state",
details={"stale_count": len(stale_cp), "total": len(cp_list)},
follow_up=True,
)
)
else:
items.append(
_item(
DIM_CHECKPOINTS,
ITEM_RESOLVED,
f"{len(cp_list)} checkpoint(s) consistent with live state",
details={"total": len(cp_list)},
)
)
# --- leases / locks -------------------------------------------------
leases = [L for L in (inventory.get("leases") or []) if isinstance(L, Mapping)]
live_leases = [L for L in leases if _is_live_freshness(_lease_freshness(L))]
items.append(
_item(
DIM_LEASES,
ITEM_RESOLVED if inventory_complete else ITEM_DEGRADED,
f"{len(live_leases)} live lease(s) of {len(leases)} inventoried",
details={
"live_count": len(live_leases),
"total": len(leases),
"live_lease_ids": [L.get("lease_id") for L in live_leases],
},
follow_up=not inventory_complete,
)
)
# --- capabilities / stale runtime -----------------------------------
caps = inventory.get("capabilities") or {}
if not isinstance(caps, Mapping):
caps = {}
if caps.get("stale") is True:
items.append(
_item(
DIM_CAPABILITIES,
ITEM_UNRESOLVED,
"runtime code is stale vs on-disk master; restart did not reach parity",
details=dict(caps),
follow_up=True,
)
)
else:
items.append(
_item(
DIM_CAPABILITIES,
ITEM_RESOLVED,
"capability/runtime parity acceptable",
details=dict(caps) if caps else {"stale": False},
)
)
# --- worktrees ------------------------------------------------------
bindings = [
b for b in (inventory.get("worktree_bindings") or []) if isinstance(b, Mapping)
]
missing_wt = [
b
for b in bindings
if b.get("missing") is True or b.get("exists") is False
]
if "worktree_bindings" not in inventory:
items.append(
_item(
DIM_WORKTREES,
ITEM_SKIPPED,
"worktree binding inventory not supplied",
)
)
elif missing_wt:
items.append(
_item(
DIM_WORKTREES,
ITEM_UNRESOLVED,
f"{len(missing_wt)} worktree binding(s) missing on disk",
details={"missing": missing_wt, "total": len(bindings)},
follow_up=True,
)
)
else:
items.append(
_item(
DIM_WORKTREES,
ITEM_RESOLVED,
f"{len(bindings)} worktree binding(s) present",
details={"total": len(bindings)},
)
)
# --- interrupted mutations (AC4) ------------------------------------
pending = [
m
for m in (inventory.get("pending_mutations") or [])
if isinstance(m, Mapping)
]
interrupted = _detect_interrupted_mutations(leases, pending)
if interrupted:
items.append(
_item(
DIM_MUTATIONS,
ITEM_UNRESOLVED,
f"{len(interrupted)} interrupted mutation(s); write resume forbidden",
details={"interrupted": interrupted},
follow_up=True,
)
)
else:
items.append(
_item(
DIM_MUTATIONS,
ITEM_RESOLVED,
"no interrupted mutations detected",
details={"interrupted": []},
)
)
# --- duplicates -----------------------------------------------------
dups = _detect_duplicate_work(leases)
if dups:
items.append(
_item(
DIM_DUPLICATES,
ITEM_UNRESOLVED,
f"{len(dups)} work item(s) have multiple live claims",
details={"duplicates": dups},
follow_up=True,
)
)
else:
items.append(
_item(
DIM_DUPLICATES,
ITEM_RESOLVED,
"no duplicate live claims detected",
details={"duplicates": []},
)
)
# --- queue ----------------------------------------------------------
queue = inventory.get("queue_state")
if queue is None:
items.append(
_item(
DIM_QUEUE,
ITEM_SKIPPED,
"allocator queue state not supplied",
)
)
elif isinstance(queue, Mapping) and queue.get("safe_to_resume") is False:
items.append(
_item(
DIM_QUEUE,
ITEM_UNRESOLVED,
"allocator queue not safe to resume",
details=dict(queue),
follow_up=True,
)
)
else:
items.append(
_item(
DIM_QUEUE,
ITEM_RESOLVED,
"allocator queue state acceptable",
details=dict(queue) if isinstance(queue, Mapping) else {},
)
)
# Incomplete inventory always degrades the whole proof.
if not inventory_complete:
# Ensure at least one follow-up names the incomplete inventory.
items.append(
_item(
"inventory",
ITEM_UNRESOLVED,
"control-plane inventory incomplete; reconcile cannot claim success",
details={"reasons": list(incomplete_reasons)},
follow_up=True,
)
)
resolved = sum(1 for i in items if i.status == ITEM_RESOLVED)
unresolved = sum(1 for i in items if i.status in {ITEM_UNRESOLVED, ITEM_DEGRADED})
skipped = sum(1 for i in items if i.status == ITEM_SKIPPED)
if not inventory_complete or any(i.status == ITEM_UNRESOLVED for i in items):
if any(i.status == ITEM_UNRESOLVED for i in items) and inventory_complete:
overall = STATUS_DEGRADED
elif not inventory_complete:
overall = STATUS_FAILED
else:
overall = STATUS_DEGRADED
elif any(i.status == ITEM_DEGRADED for i in items):
overall = STATUS_DEGRADED
else:
overall = STATUS_COMPLETE
# Enforce mode holds mutations whenever anything is unresolved/failed.
mutation_hold = False
if mode_norm == MODE_ENFORCE and overall in {STATUS_DEGRADED, STATUS_FAILED}:
mutation_hold = True
if mode_norm == MODE_ENFORCE and any(
i.dimension == DIM_MUTATIONS and i.status == ITEM_UNRESOLVED for i in items
):
mutation_hold = True
follow_ups = tuple(
fu for i in items if (fu := _follow_up_for_item(i)) is not None
)
finished = _utc_now() if now is None else now
note = (
"Read-only completion proof. Never auto-resumes write mutations. "
"Unresolved items require durable follow-up before claiming clean restart. "
f"Mode={mode_norm}."
)
return RestartCompletionProof(
schema_version=SCHEMA_VERSION,
reconcile_version=RECONCILE_VERSION,
reconcile_id=(reconcile_id or f"reconcile-{uuid4().hex[:12]}"),
started_at=_ts(started),
finished_at=_ts(finished),
boot_head_sha=(
str(inventory.get("boot_head_sha")).strip()
if inventory.get("boot_head_sha")
else None
),
current_head_sha=(
str(inventory.get("current_head_sha")).strip()
if inventory.get("current_head_sha")
else None
),
inventory_complete=inventory_complete,
incomplete_reasons=incomplete_reasons,
mode=mode_norm,
mutation_hold=mutation_hold,
overall_status=overall,
items=tuple(items),
proposed_follow_ups=follow_ups,
resolved_count=resolved,
unresolved_count=unresolved,
skipped_count=skipped,
note=note,
)
def mutations_allowed(proof: RestartCompletionProof | Mapping[str, Any] | None) -> bool:
"""Return whether write mutations may proceed under the given proof."""
if proof is None:
return True # no proof yet → caller decides; enforce path sets hold
if isinstance(proof, RestartCompletionProof):
return not proof.mutation_hold
if isinstance(proof, Mapping):
return not bool(proof.get("mutation_hold"))
return True
-3
View File
@@ -1,3 +0,0 @@
[pytest]
testpaths = tests
norecursedirs = branches .git venv __pycache__ graphify-out
-583
View File
@@ -1,583 +0,0 @@
"""Scoped MCP recovery playbook (#669).
Operational recovery must prefer the *narrowest* action that can fix the
symptom. Full MCP / host restarts are last-resort rungs on a documented
ladder; the coordinator refuses those rungs unless a prior attempt log
shows narrower recoveries already failed (or break-glass is authorized).
This module is pure classification and recommendation:
* No network, filesystem, or process I/O.
* Never restarts anything.
* Narrow recovery *execution* is delegated to existing tools/docs (linked
per rung) the playbook records which rung to try next and whether
escalation to a broad restart is allowed.
Design lineage: umbrella #655, class matrix #663, coordinator #658,
auto-reconnect #584, stale-runtime #610, contamination #630, audit #665.
Vision #652 / roadmap #653.
"""
from __future__ import annotations
from dataclasses import dataclass, field
from datetime import datetime, timezone
from enum import Enum
from typing import Any, Mapping, Sequence
PLAYBOOK_VERSION = "1.0.0-issue-669"
# Attempt outcomes that count as "tried and insufficient" for escalation.
INSUFFICIENT_OUTCOMES = frozenset(
{
"failed",
"insufficient",
"denied",
"unresolved",
"timeout",
"error",
}
)
# Break-glass / operator override still records that the ladder was skipped.
OUTCOME_BREAK_GLASS = "break_glass"
OUTCOME_SUCCESS = "success"
OUTCOME_SKIPPED = "skipped"
class RecoveryAction(str, Enum):
"""Ordered recovery ladder (narrow → broad)."""
CLIENT_RECONNECT = "client_reconnect"
CAPABILITY_REFRESH = "capability_refresh"
SESSION_RECONNECT = "session_reconnect"
CONFIGURATION_RELOAD = "configuration_reload"
LEASE_RECOVERY = "lease_recovery"
WORKER_RESTART = "worker_restart"
ROLE_RUNTIME_RESTART = "role_runtime_restart"
CONNECTOR_RESTART = "connector_restart"
ROLLING_MCP_RESTART = "rolling_mcp_restart"
FULL_MCP_RESTART = "full_mcp_restart"
HOST_RESTART = "host_restart"
# Classes that require a prior narrow-attempt log (unless break-glass).
BROAD_RESTART_ACTIONS: frozenset[RecoveryAction] = frozenset(
{
RecoveryAction.ROLLING_MCP_RESTART,
RecoveryAction.FULL_MCP_RESTART,
RecoveryAction.HOST_RESTART,
}
)
# Map #663 restart_class strings onto playbook actions.
RESTART_CLASS_TO_ACTION: dict[str, RecoveryAction] = {
"client_reconnect": RecoveryAction.CLIENT_RECONNECT,
"session_reconnect": RecoveryAction.SESSION_RECONNECT,
"configuration_reload": RecoveryAction.CONFIGURATION_RELOAD,
"worker_restart": RecoveryAction.WORKER_RESTART,
"role_runtime_restart": RecoveryAction.ROLE_RUNTIME_RESTART,
"connector_restart": RecoveryAction.CONNECTOR_RESTART,
"rolling_mcp_restart": RecoveryAction.ROLLING_MCP_RESTART,
"full_mcp_restart": RecoveryAction.FULL_MCP_RESTART,
"host_restart": RecoveryAction.HOST_RESTART,
}
@dataclass(frozen=True)
class RecoveryRung:
"""One rung on the recovery ladder."""
action: RecoveryAction
rank: int
summary: str
# Existing implementation or explicit delegation target.
implementation: str
issue_links: tuple[str, ...]
self_service: bool
# Restart-class permission when this rung is requested via coordinator.
restart_class: str | None = None
def as_dict(self) -> dict[str, Any]:
return {
"action": self.action.value,
"rank": self.rank,
"summary": self.summary,
"implementation": self.implementation,
"issue_links": list(self.issue_links),
"self_service": self.self_service,
"restart_class": self.restart_class,
}
# Canonical ladder. Rank 0 is narrowest.
RECOVERY_LADDER: tuple[RecoveryRung, ...] = (
RecoveryRung(
RecoveryAction.CLIENT_RECONNECT,
0,
"Reconnect the IDE/client MCP transport (EOF / transport flap).",
"Host auto-reconnect or explicit client reconnect; "
"docs/mcp-namespace-eof-recovery.md",
("#584", "#655"),
True,
"client_reconnect",
),
RecoveryRung(
RecoveryAction.CAPABILITY_REFRESH,
1,
"Re-resolve task capability and clear stale permission context.",
"Delegated: gitea_resolve_task_capability + gitea_whoami "
"(no process change).",
("#610", "#685", "#655"),
True,
None,
),
RecoveryRung(
RecoveryAction.SESSION_RECONNECT,
2,
"Rebind identity, workspace, and namespace for one session.",
"Delegated: gitea_get_runtime_context + explicit worktree_path "
"rebind (#618); docs/mcp-namespace-health.md",
("#543", "#618", "#655"),
True,
"session_reconnect",
),
RecoveryRung(
RecoveryAction.CONFIGURATION_RELOAD,
3,
"Gracefully reload configuration without replacing the daemon.",
"restart_coordinator class configuration_reload; console "
"system.reload_namespace (#642).",
("#642", "#663", "#655"),
False,
"configuration_reload",
),
RecoveryRung(
RecoveryAction.LEASE_RECOVERY,
4,
"Recover or rebind stale leases/locks without a process restart.",
"Delegated: issue lock recovery / lease lifecycle paths "
"(#702, #753, #790).",
("#702", "#753", "#790", "#655"),
False,
None,
),
RecoveryRung(
RecoveryAction.WORKER_RESTART,
5,
"Restart one worker after its own lease and mutation scope drains.",
"restart_coordinator class worker_restart (#663).",
("#663", "#655"),
False,
"worker_restart",
),
RecoveryRung(
RecoveryAction.ROLE_RUNTIME_RESTART,
6,
"Restart one role runtime and re-probe that namespace only.",
"restart_coordinator class role_runtime_restart; console "
"system.restart_namespace (#642).",
("#642", "#663", "#655"),
False,
"role_runtime_restart",
),
RecoveryRung(
RecoveryAction.CONNECTOR_RESTART,
7,
"Restart one connector while unrelated runtimes stay available.",
"restart_coordinator class connector_restart (#663).",
("#663", "#655"),
False,
"connector_restart",
),
RecoveryRung(
RecoveryAction.ROLLING_MCP_RESTART,
8,
"Drain/restart/verify one instance at a time (HA path).",
"restart_coordinator class rolling_mcp_restart; design #668.",
("#668", "#663", "#655"),
False,
"rolling_mcp_restart",
),
RecoveryRung(
RecoveryAction.FULL_MCP_RESTART,
9,
"Full stable-control MCP process restart after verified full drain.",
"restart_coordinator class full_mcp_restart; requires attempt log "
"unless break-glass (#669).",
("#658", "#661", "#663", "#669", "#655"),
False,
"full_mcp_restart",
),
RecoveryRung(
RecoveryAction.HOST_RESTART,
10,
"Host/infrastructure restart — broadest last-resort action.",
"restart_coordinator class host_restart; operator-owned.",
("#663", "#669", "#655"),
False,
"host_restart",
),
)
_LADDER_BY_ACTION: dict[RecoveryAction, RecoveryRung] = {
rung.action: rung for rung in RECOVERY_LADDER
}
# Symptom tokens → preferred first rung (decision tree, #663 lineage).
SYMPTOM_TO_FIRST_ACTION: dict[str, RecoveryAction] = {
"transport_eof": RecoveryAction.CLIENT_RECONNECT,
"client_closing_eof": RecoveryAction.CLIENT_RECONNECT,
"transport_flap": RecoveryAction.CLIENT_RECONNECT,
"namespace_disconnected": RecoveryAction.CLIENT_RECONNECT,
"stale_capability": RecoveryAction.CAPABILITY_REFRESH,
"permission_stale": RecoveryAction.CAPABILITY_REFRESH,
"runtime_reconnect_required": RecoveryAction.CAPABILITY_REFRESH,
"stale_runtime": RecoveryAction.SESSION_RECONNECT,
"worktree_unbound": RecoveryAction.SESSION_RECONNECT,
"namespace_unhealthy": RecoveryAction.SESSION_RECONNECT,
"config_drift": RecoveryAction.CONFIGURATION_RELOAD,
"profile_misbound": RecoveryAction.CONFIGURATION_RELOAD,
"stale_lease": RecoveryAction.LEASE_RECOVERY,
"dead_pid_lock": RecoveryAction.LEASE_RECOVERY,
"orphan_worktree": RecoveryAction.LEASE_RECOVERY,
"single_worker_stuck": RecoveryAction.WORKER_RESTART,
"role_runtime_dead": RecoveryAction.ROLE_RUNTIME_RESTART,
"connector_dead": RecoveryAction.CONNECTOR_RESTART,
"ha_instance_unhealthy": RecoveryAction.ROLLING_MCP_RESTART,
"daemon_corrupt": RecoveryAction.FULL_MCP_RESTART,
"full_process_deadlock": RecoveryAction.FULL_MCP_RESTART,
"host_unresponsive": RecoveryAction.HOST_RESTART,
}
def _utc_now() -> datetime:
return datetime.now(timezone.utc)
def resolve_action(value: RecoveryAction | str) -> RecoveryAction:
"""Resolve a recovery action or fail closed for unknown values."""
if isinstance(value, RecoveryAction):
return value
text = str(value or "").strip()
# Accept #663 restart_class aliases.
if text in RESTART_CLASS_TO_ACTION:
return RESTART_CLASS_TO_ACTION[text]
try:
return RecoveryAction(text)
except ValueError as exc:
raise ValueError(
f"unknown recovery action {value!r}; deny (fail closed, #669)"
) from exc
def ladder_rank(action: RecoveryAction | str) -> int:
resolved = resolve_action(action)
return _LADDER_BY_ACTION[resolved].rank
def rung_for(action: RecoveryAction | str) -> RecoveryRung:
return _LADDER_BY_ACTION[resolve_action(action)]
def normalize_attempt(raw: Mapping[str, Any]) -> dict[str, Any] | None:
"""Normalize one prior-recovery attempt record; return None if unusable."""
if not isinstance(raw, Mapping):
return None
action_raw = raw.get("action") or raw.get("recovery_action") or raw.get(
"restart_class"
)
if not action_raw:
return None
try:
action = resolve_action(str(action_raw))
except ValueError:
return None
outcome = str(
raw.get("outcome") or raw.get("status") or raw.get("result") or ""
).strip().lower()
if not outcome:
return None
recorded_at = raw.get("recorded_at") or raw.get("at") or raw.get("timestamp")
reason = str(raw.get("reason") or raw.get("detail") or "").strip()
actor = str(raw.get("actor") or raw.get("session_id") or "").strip()
return {
"action": action.value,
"outcome": outcome,
"reason": reason,
"actor": actor,
"recorded_at": recorded_at,
"rank": ladder_rank(action),
"raw": dict(raw),
}
def normalize_attempt_log(
attempts: Sequence[Mapping[str, Any]] | None,
) -> list[dict[str, Any]]:
"""Return usable attempt records in ladder order."""
out: list[dict[str, Any]] = []
for raw in attempts or ():
norm = normalize_attempt(raw)
if norm is not None:
out.append(norm)
out.sort(key=lambda a: (a["rank"], str(a.get("recorded_at") or "")))
return out
def narrower_insufficient_attempts(
attempts: Sequence[Mapping[str, Any]] | None,
*,
requested: RecoveryAction | str,
) -> list[dict[str, Any]]:
"""Return prior attempts narrower than *requested* that were insufficient."""
target_rank = ladder_rank(requested)
usable = []
for attempt in normalize_attempt_log(attempts):
if attempt["rank"] >= target_rank:
continue
if attempt["outcome"] in INSUFFICIENT_OUTCOMES:
usable.append(attempt)
return usable
@dataclass(frozen=True)
class EscalationAssessment:
"""Whether a requested broad recovery may proceed given the attempt log."""
requested_action: str
allowed: bool
require_attempt_log: bool
break_glass: bool
reasons: list[str] = field(default_factory=list)
qualifying_attempts: list[dict[str, Any]] = field(default_factory=list)
recommended_next: list[dict[str, Any]] = field(default_factory=list)
playbook_version: str = PLAYBOOK_VERSION
def as_dict(self) -> dict[str, Any]:
return {
"playbook_version": self.playbook_version,
"requested_action": self.requested_action,
"allowed": self.allowed,
"require_attempt_log": self.require_attempt_log,
"break_glass": self.break_glass,
"reasons": list(self.reasons),
"qualifying_attempts": list(self.qualifying_attempts),
"recommended_next": list(self.recommended_next),
}
def assess_escalation(
requested: RecoveryAction | str,
*,
prior_recovery_attempts: Sequence[Mapping[str, Any]] | None = None,
break_glass: bool = False,
) -> EscalationAssessment:
"""Gate broad restarts on a prior narrow-attempt log (#669 AC3).
Narrow / mid-ladder actions do not require a prior attempt log.
``full_mcp_restart``, ``host_restart``, and ``rolling_mcp_restart``
require at least one *insufficient* narrower attempt unless
``break_glass`` is true.
"""
action = resolve_action(requested)
require_log = action in BROAD_RESTART_ACTIONS
reasons: list[str] = []
qualifying = narrower_insufficient_attempts(
prior_recovery_attempts, requested=action
)
if not require_log:
return EscalationAssessment(
requested_action=action.value,
allowed=True,
require_attempt_log=False,
break_glass=bool(break_glass),
reasons=["narrow recovery; attempt log not required"],
qualifying_attempts=qualifying,
recommended_next=[],
)
if break_glass:
return EscalationAssessment(
requested_action=action.value,
allowed=True,
require_attempt_log=True,
break_glass=True,
reasons=[
"break-glass authorized; broad restart permitted without "
"narrow-attempt log (#669)"
],
qualifying_attempts=qualifying,
recommended_next=[],
)
if qualifying:
return EscalationAssessment(
requested_action=action.value,
allowed=True,
require_attempt_log=True,
break_glass=False,
reasons=[
f"{len(qualifying)} narrower recovery attempt(s) recorded as "
"insufficient; escalation permitted"
],
qualifying_attempts=qualifying,
recommended_next=[],
)
# Deny: recommend the next untried narrow rung(s).
recommended = recommend_actions(
symptoms=(),
prior_recovery_attempts=prior_recovery_attempts,
max_actions=3,
)
reasons.append(
f"{action.value} requires a prior attempt log of insufficient "
"narrower recoveries (or break-glass); none found — deny (fail "
"closed, #669)"
)
return EscalationAssessment(
requested_action=action.value,
allowed=False,
require_attempt_log=True,
break_glass=False,
reasons=reasons,
qualifying_attempts=[],
recommended_next=recommended.get("recommended_actions") or [],
)
def recommend_actions(
*,
symptoms: Sequence[str] = (),
prior_recovery_attempts: Sequence[Mapping[str, Any]] | None = None,
max_actions: int = 5,
) -> dict[str, Any]:
"""Return ordered recommended recovery actions for the given symptoms.
Soft mode (rollout): recommendations only callers decide whether to
hard-gate. Hard mode for broad restarts is :func:`assess_escalation`.
"""
attempted_success = {
a["action"]
for a in normalize_attempt_log(prior_recovery_attempts)
if a["outcome"] == OUTCOME_SUCCESS
}
attempted_any = {
a["action"] for a in normalize_attempt_log(prior_recovery_attempts)
}
first_actions: list[RecoveryAction] = []
for symptom in symptoms:
key = str(symptom or "").strip().lower().replace(" ", "_").replace("-", "_")
mapped = SYMPTOM_TO_FIRST_ACTION.get(key)
if mapped is not None and mapped not in first_actions:
first_actions.append(mapped)
# Default entry: client reconnect then walk the ladder.
if not first_actions:
first_actions = [RecoveryAction.CLIENT_RECONNECT]
recommended: list[dict[str, Any]] = []
seen: set[str] = set()
min_rank = min(ladder_rank(a) for a in first_actions)
for rung in RECOVERY_LADDER:
if rung.rank < min_rank:
continue
if rung.action.value in attempted_success:
continue
if rung.action.value in seen:
continue
# Prefer rungs not yet attempted; still list previously-failed ones
# only if nothing else remains.
entry = rung.as_dict()
entry["already_attempted"] = rung.action.value in attempted_any
recommended.append(entry)
seen.add(rung.action.value)
if len(recommended) >= max(1, int(max_actions)):
break
return {
"playbook_version": PLAYBOOK_VERSION,
"symptoms": [str(s) for s in symptoms],
"recommended_actions": recommended,
"ladder": [r.as_dict() for r in RECOVERY_LADDER],
"read_only": True,
"hard_gate_note": (
"Broad restarts (rolling/full/host) still require "
"assess_escalation / coordinator attempt-log enforcement."
),
}
def build_attempt_record(
action: RecoveryAction | str,
*,
outcome: str,
reason: str = "",
actor: str = "",
recorded_at: str | None = None,
extra: Mapping[str, Any] | None = None,
) -> dict[str, Any]:
"""Build a durable-shaped attempt log entry for inventory/audit (#665)."""
resolved = resolve_action(action)
record = {
"action": resolved.value,
"outcome": str(outcome or "").strip().lower(),
"reason": str(reason or "").strip(),
"actor": str(actor or "").strip(),
"recorded_at": recorded_at or _utc_now().isoformat(),
"rank": ladder_rank(resolved),
"playbook_version": PLAYBOOK_VERSION,
}
if extra:
record["extra"] = dict(extra)
return record
def recovery_metrics(
attempts: Sequence[Mapping[str, Any]] | None,
) -> dict[str, Any]:
"""Compute the fraction of recoveries that avoided full/host restart.
A recovery *episode* is approximated as one attempt with
``outcome=success``. Successes on non-broad rungs count as avoided full
restart; successes on full/host count as full-restart recoveries.
"""
norms = normalize_attempt_log(attempts)
successes = [a for a in norms if a["outcome"] == OUTCOME_SUCCESS]
broad_success = [
a
for a in successes
if resolve_action(a["action"])
in {RecoveryAction.FULL_MCP_RESTART, RecoveryAction.HOST_RESTART}
]
avoided = [a for a in successes if a not in broad_success]
total = len(successes)
fraction_avoided = (len(avoided) / total) if total else None
return {
"playbook_version": PLAYBOOK_VERSION,
"attempts_total": len(norms),
"successes_total": total,
"successes_avoided_full_restart": len(avoided),
"successes_full_or_host_restart": len(broad_success),
"fraction_avoided_full_restart": fraction_avoided,
"insufficient_attempts": sum(
1 for a in norms if a["outcome"] in INSUFFICIENT_OUTCOMES
),
}
def ladder_document() -> dict[str, Any]:
"""Machine-readable ladder for docs/tools inventory."""
return {
"playbook_version": PLAYBOOK_VERSION,
"parent_issues": ["#655", "#652", "#653"],
"enforcement_issue": "#669",
"ladder": [r.as_dict() for r in RECOVERY_LADDER],
"broad_restart_actions": [a.value for a in sorted(BROAD_RESTART_ACTIONS, key=lambda x: x.value)],
"insufficient_outcomes": sorted(INSUFFICIENT_OUTCOMES),
"symptom_map": {k: v.value for k, v in sorted(SYMPTOM_TO_FIRST_ACTION.items())},
}
-2
View File
@@ -9,8 +9,6 @@ cryptography==49.0.0
h11==0.16.0
httpcore==1.0.9
httpx==0.28.1
# Starlette 1.3.x TestClient prefers httpx2; plain httpx remains for MCP/runtime (#682).
httpx2==2.9.1
httpx-sse==0.4.3
idna==3.18
iniconfig==2.3.0
-851
View File
@@ -1,851 +0,0 @@
"""MCP restart coordinator and impact analysis (#658 / #669).
Before any sanctioned MCP restart, a central coordinator must evaluate the
live control-plane state active sessions, leases/locks, in-flight issue/PR
work, mutations, worktrees, and recovery history and produce an *impact
preview* so operators (and the web console, #642/#652) can see the blast
radius **before** concurrent LLM work is disrupted.
Design rules (mirrors the read-only posture of ``workflow_dashboard`` /
``lease_lifecycle``):
* **Pure classification.** :func:`evaluate_restart_impact` takes an already
gathered inventory and returns a structured report. It never touches the
network, the filesystem, or a live process, so multi-session fixtures can
drive every branch in unit tests. The coordinator *never restarts anything*;
a mutative apply path is a later child gated by a drain proof (non-goal here).
* **Fail closed.** If the inventory is not explicitly complete, the verdict is
``unsafe`` / deny an incomplete evaluation must never green-light a restart.
* **Narrow-first (#669).** Broad classes (rolling / full / host) require a
prior attempt log of insufficient narrower recoveries unless break-glass is
authorized. See :mod:`recovery_playbook`.
* **No secrets.** Session ids, pids, and profiles are operational metadata, not
credentials; nothing secret flows through this module.
The single sanctioned entry point post-#657 is the MCP tool
``gitea_request_mcp_restart`` (dry-run by default), which gathers the inventory
from the #613 control-plane DB and calls :func:`evaluate_restart_impact`.
"""
from __future__ import annotations
from dataclasses import dataclass, field
from datetime import datetime, timezone
from enum import Enum
from typing import Any, Mapping, Sequence
import lease_lifecycle
import recovery_playbook
COORDINATOR_VERSION = "1.2.0-issue-669"
# Restart verdicts. Exactly the three the acceptance criteria name.
VERDICT_SAFE = "safe"
VERDICT_UNSAFE = "unsafe"
VERDICT_OVERRIDE = "override"
# Blast-radius severity bands.
BLAST_NONE = "none"
BLAST_LOW = "low"
BLAST_MEDIUM = "medium"
BLAST_HIGH = "high"
# A live lease with a live owner process is treated as active in-flight work.
LEASE_FRESHNESS_LIVE = "active"
# Default staleness window for a session heartbeat (seconds). A session whose
# last heartbeat is older than this is not counted as live even if its row is
# still marked ``active`` — it is assumed dead/detached.
DEFAULT_SESSION_HEARTBEAT_STALE_SECONDS = 900
class RestartClass(str, Enum):
"""The only restart/recovery classes accepted by the coordinator."""
CLIENT_RECONNECT = "client_reconnect"
SESSION_RECONNECT = "session_reconnect"
WORKER_RESTART = "worker_restart"
ROLE_RUNTIME_RESTART = "role_runtime_restart"
CONNECTOR_RESTART = "connector_restart"
CONFIGURATION_RELOAD = "configuration_reload"
ROLLING_MCP_RESTART = "rolling_mcp_restart"
FULL_MCP_RESTART = "full_mcp_restart"
HOST_RESTART = "host_restart"
@dataclass(frozen=True)
class RestartClassPolicy:
"""Least-privilege policy for one :class:`RestartClass`."""
restart_class: RestartClass
required_permission: str
expected_blast_radius: str
drain_requirement: str
full_drain_required: bool
approval_requirement: str
audit_requirement: str
recovery_behavior: str
request_roles: tuple[str, ...]
execution_roles: tuple[str, ...]
def as_dict(self) -> dict[str, Any]:
return {
"restart_class": self.restart_class.value,
"required_permission": self.required_permission,
"expected_blast_radius": self.expected_blast_radius,
"drain_requirement": self.drain_requirement,
"full_drain_required": self.full_drain_required,
"approval_requirement": self.approval_requirement,
"audit_requirement": self.audit_requirement,
"recovery_behavior": self.recovery_behavior,
"request_roles": list(self.request_roles),
"execution_roles": list(self.execution_roles),
}
WORKER_ROLES = ("author", "reviewer", "merger", "reconciler")
CONTROL_ROLES = ("controller", "operator", "admin")
ALL_REQUEST_ROLES = WORKER_ROLES + CONTROL_ROLES
RESTART_CLASS_POLICIES: dict[RestartClass, RestartClassPolicy] = {
RestartClass.CLIENT_RECONNECT: RestartClassPolicy(
RestartClass.CLIENT_RECONNECT,
"mcp.reconnect.client",
BLAST_NONE,
"none",
False,
"self_service",
"record class, actor, client namespace, reason, and outcome",
"Reconnect only the caller's client transport; no daemon or peer session changes.",
ALL_REQUEST_ROLES,
ALL_REQUEST_ROLES,
),
RestartClass.SESSION_RECONNECT: RestartClassPolicy(
RestartClass.SESSION_RECONNECT,
"mcp.reconnect.session",
BLAST_LOW,
"requesting_session_safe_point",
False,
"self_service",
"record class, actor, session id, reason, and outcome",
"Rebind identity, capability, and workspace state for one session.",
ALL_REQUEST_ROLES,
ALL_REQUEST_ROLES,
),
RestartClass.WORKER_RESTART: RestartClassPolicy(
RestartClass.WORKER_RESTART,
"mcp.restart.worker.request",
BLAST_LOW,
"target_worker",
False,
"controller_approval_and_automated_gates",
"record class, actor, target worker, approval, drain proof, and outcome",
"Restart one worker after its own lease and mutation scope is drained.",
ALL_REQUEST_ROLES,
("operator", "admin"),
),
RestartClass.ROLE_RUNTIME_RESTART: RestartClassPolicy(
RestartClass.ROLE_RUNTIME_RESTART,
"mcp.restart.role_runtime.request",
BLAST_MEDIUM,
"target_role_runtime",
False,
"controller_approval_and_automated_gates",
"record class, actor, role namespace, approval, drain proof, and outcome",
"Restart only the selected role runtime and then re-probe that namespace.",
ALL_REQUEST_ROLES,
("operator", "admin"),
),
RestartClass.CONNECTOR_RESTART: RestartClassPolicy(
RestartClass.CONNECTOR_RESTART,
"mcp.restart.connector.request",
BLAST_MEDIUM,
"target_connector",
False,
"controller_approval_and_automated_gates",
"record class, actor, connector id, approval, drain proof, and outcome",
"Restart one connector while unrelated role runtimes remain available.",
ALL_REQUEST_ROLES,
("operator", "admin"),
),
RestartClass.CONFIGURATION_RELOAD: RestartClassPolicy(
RestartClass.CONFIGURATION_RELOAD,
"mcp.reload.configuration.request",
BLAST_LOW,
"mutation_quiesce",
False,
"controller_approval_and_automated_gates",
"record class, actor, configuration revision, approval, and outcome",
"Gracefully reload configuration without replacing the daemon process.",
ALL_REQUEST_ROLES,
("operator", "admin"),
),
RestartClass.ROLLING_MCP_RESTART: RestartClassPolicy(
RestartClass.ROLLING_MCP_RESTART,
"mcp.restart.rolling.request",
BLAST_MEDIUM,
"one_instance_at_a_time",
False,
"controller_approval_and_automated_gates",
"record class, actor, instance order, approval, per-instance drains, and outcome",
"Drain, restart, verify, and restore one instance before advancing to the next.",
CONTROL_ROLES,
("operator", "admin"),
),
RestartClass.FULL_MCP_RESTART: RestartClassPolicy(
RestartClass.FULL_MCP_RESTART,
"mcp.restart.full.request",
BLAST_HIGH,
"all_sessions_and_mutations",
True,
"controller_approval_and_automated_gates",
"record class, actor, full impact report, approval, drain proof, and outcome",
"Stop and restore the complete MCP runtime only after a verified full drain.",
CONTROL_ROLES,
("operator", "admin"),
),
RestartClass.HOST_RESTART: RestartClassPolicy(
RestartClass.HOST_RESTART,
"mcp.restart.host.request",
BLAST_HIGH,
"all_host_work",
True,
"controller_approval_plus_infrastructure_operator",
"record class, actor, host, incident or change id, approval, drain proof, and outcome",
"Hand off to infrastructure ownership; reconcile every runtime after the host returns.",
("controller", "operator", "admin"),
("operator", "admin"),
),
}
def resolve_restart_class(value: RestartClass | str) -> RestartClass:
"""Resolve a restart class or fail closed for an unknown value."""
if isinstance(value, RestartClass):
return value
try:
return RestartClass(str(value).strip())
except ValueError as exc:
raise ValueError(f"unknown restart class {value!r}; deny (fail closed)") from exc
def restart_class_policy(value: RestartClass | str) -> RestartClassPolicy:
"""Return the canonical policy for *value*."""
return RESTART_CLASS_POLICIES[resolve_restart_class(value)]
def permissions_for_role(role: str | None) -> tuple[str, ...]:
"""Return request permissions granted to a workflow role by this policy."""
normalized = str(role or "").strip().lower()
return tuple(
policy.required_permission
for policy in RESTART_CLASS_POLICIES.values()
if normalized in policy.request_roles
)
def _utc_now() -> datetime:
return datetime.now(timezone.utc)
def _parse_ts(value: str | None) -> datetime | None:
return lease_lifecycle._parse_ts(value)
@dataclass(frozen=True)
class SessionImpact:
"""One MCP session a restart would terminate."""
session_id: str
role: str | None
profile: str | None
pid: int | None
status: str | None
alive: bool | None
heartbeat_stale: bool
is_requester: bool
live: bool
connector: str | None = None
def as_dict(self) -> dict[str, Any]:
return {
"session_id": self.session_id,
"role": self.role,
"profile": self.profile,
"pid": self.pid,
"status": self.status,
"alive": self.alive,
"heartbeat_stale": self.heartbeat_stale,
"is_requester": self.is_requester,
"live": self.live,
"connector": self.connector,
}
@dataclass(frozen=True)
class LeaseImpact:
"""One control-plane lease a restart would disrupt."""
lease_id: str | None
session_id: str | None
role: str | None
phase: str | None
freshness: str | None
work_kind: str | None
work_number: int | None
worktree_path: str | None
disruptive: bool
is_mutation: bool
is_critical_section: bool
connector: str | None = None
def as_dict(self) -> dict[str, Any]:
return {
"lease_id": self.lease_id,
"session_id": self.session_id,
"role": self.role,
"phase": self.phase,
"freshness": self.freshness,
"work_kind": self.work_kind,
"work_number": self.work_number,
"worktree_path": self.worktree_path,
"disruptive": self.disruptive,
"is_mutation": self.is_mutation,
"is_critical_section": self.is_critical_section,
"connector": self.connector,
}
@dataclass(frozen=True)
class RestartImpactReport:
"""Impact preview DTO returned to the console / operator (#642/#652)."""
coordinator_version: str
restart_class: str
restart_policy: dict[str, Any]
policy_enforced: bool
permission_authorized: bool
role_authorized: bool
approval_satisfied: bool
authorization_reasons: list[str]
evaluated_at: str
dry_run: bool
restart_performed: bool
inventory_complete: bool
verdict: str
allow_restart: bool
override_would_allow: bool
operator_override: bool
blast_radius: str
reasons: list[str]
affected_sessions: list[SessionImpact]
affected_leases: list[LeaseImpact]
critical_sections: list[LeaseImpact]
affected_issues: list[int]
affected_prs: list[int]
mutations: list[LeaseImpact]
terminal_lock: dict[str, Any] | None
ack_state: dict[str, str]
prior_recovery_attempts: list[dict[str, Any]]
counts: dict[str, int]
audit_record: dict[str, Any]
incomplete_reasons: list[str] = field(default_factory=list)
# #669 playbook escalation gate (attempt-log enforcement).
playbook_escalation: dict[str, Any] = field(default_factory=dict)
attempt_log_satisfied: bool = True
break_glass: bool = False
def as_dict(self) -> dict[str, Any]:
return {
"coordinator_version": self.coordinator_version,
"restart_class": self.restart_class,
"restart_policy": dict(self.restart_policy),
"policy_enforced": self.policy_enforced,
"permission_authorized": self.permission_authorized,
"role_authorized": self.role_authorized,
"approval_satisfied": self.approval_satisfied,
"authorization_reasons": list(self.authorization_reasons),
"evaluated_at": self.evaluated_at,
"dry_run": self.dry_run,
"restart_performed": self.restart_performed,
"inventory_complete": self.inventory_complete,
"incomplete_reasons": list(self.incomplete_reasons),
"verdict": self.verdict,
"allow_restart": self.allow_restart,
"override_would_allow": self.override_would_allow,
"operator_override": self.operator_override,
"blast_radius": self.blast_radius,
"reasons": list(self.reasons),
"affected_sessions": [s.as_dict() for s in self.affected_sessions],
"affected_leases": [l.as_dict() for l in self.affected_leases],
"critical_sections": [l.as_dict() for l in self.critical_sections],
"affected_issues": list(self.affected_issues),
"affected_prs": list(self.affected_prs),
"mutations": [l.as_dict() for l in self.mutations],
"terminal_lock": self.terminal_lock,
"ack_state": dict(self.ack_state),
"prior_recovery_attempts": list(self.prior_recovery_attempts),
"counts": dict(self.counts),
"audit_record": dict(self.audit_record),
"playbook_escalation": dict(self.playbook_escalation),
"attempt_log_satisfied": self.attempt_log_satisfied,
"break_glass": self.break_glass,
}
def _classify_session(
row: Mapping[str, Any],
*,
now: datetime,
requesting_session_id: str | None,
heartbeat_stale_seconds: int,
) -> SessionImpact:
session_id = str(row.get("session_id") or "")
pid = row.get("pid")
status = (row.get("status") or "").strip().lower() or None
alive = lease_lifecycle.is_process_alive(pid) if pid is not None else None
hb = _parse_ts(row.get("last_heartbeat_at"))
heartbeat_stale = bool(
hb is not None and (now - hb).total_seconds() > heartbeat_stale_seconds
)
live = bool(status == "active" and alive is not False and not heartbeat_stale)
return SessionImpact(
session_id=session_id,
role=row.get("role"),
profile=row.get("profile"),
pid=pid,
status=status,
alive=alive,
heartbeat_stale=heartbeat_stale,
is_requester=bool(
requesting_session_id and session_id == requesting_session_id
),
live=live,
connector=(str(row.get("connector") or "").strip() or None),
)
# Lease phases that represent an active mutation in flight (as opposed to a
# mere allocation/claim with no work committed yet). An active lease in any of
# these phases is a critical section a restart must not sever.
_MUTATING_PHASES = frozenset(
{
"implementing",
"publishing",
"pushing",
"committing",
"reviewing",
"merging",
"reconciling",
"conflict_fix",
}
)
def _classify_lease(row: Mapping[str, Any]) -> LeaseImpact:
freshness_obj = row.get("freshness")
if isinstance(freshness_obj, Mapping):
freshness = str(freshness_obj.get("freshness") or "").strip().lower() or None
else:
freshness = str(freshness_obj or "").strip().lower() or None
phase = (row.get("phase") or "").strip().lower() or None
worktree = row.get("worktree_path")
disruptive = freshness == LEASE_FRESHNESS_LIVE
# A live lease is a mutation-in-flight if it carries an author worktree or
# its phase names a mutating step. All disruptive leases are critical
# sections a restart would sever regardless.
is_mutation = bool(
disruptive and (bool(worktree) or (phase in _MUTATING_PHASES))
)
number = row.get("work_number")
try:
number = int(number) if number is not None else None
except (TypeError, ValueError):
number = None
return LeaseImpact(
lease_id=row.get("lease_id"),
session_id=row.get("session_id"),
role=row.get("role"),
phase=phase,
freshness=freshness,
work_kind=(str(row.get("work_kind") or "").strip().lower() or None),
work_number=number,
worktree_path=worktree,
disruptive=disruptive,
is_mutation=is_mutation,
is_critical_section=disruptive,
connector=(str(row.get("connector") or "").strip() or None),
)
def _blast_radius(*, session_count: int, work_count: int, mutation_count: int) -> str:
if mutation_count > 0 or work_count >= 3 or session_count >= 3:
return BLAST_HIGH
if work_count > 0 or session_count == 2:
return BLAST_MEDIUM
if session_count == 1:
return BLAST_LOW
return BLAST_NONE
def evaluate_restart_impact(
inventory: Mapping[str, Any],
*,
now: datetime | None = None,
operator_override: bool = False,
requesting_session_id: str | None = None,
dry_run: bool = True,
session_heartbeat_stale_seconds: int = DEFAULT_SESSION_HEARTBEAT_STALE_SECONDS,
restart_class: RestartClass | str | None = None,
requester_role: str | None = None,
requester_permissions: Sequence[str] | None = None,
controller_approved: bool = False,
operator_authorized: bool = False,
target_session_id: str | None = None,
target_role: str | None = None,
target_connector: str | None = None,
break_glass: bool = False,
) -> RestartImpactReport:
"""Evaluate a proposed MCP restart and return an impact preview.
``inventory`` is a mapping with:
* ``sessions`` session rows (session_id, role, profile, pid, status,
last_heartbeat_at).
* ``leases`` control-plane lease rows, each ideally carrying an enriched
``freshness`` dict (as :func:`lease_lifecycle.list_active_leases` returns);
a bare string freshness is also accepted.
* ``terminal_lock`` the active terminal (merge) lock, if any.
* ``prior_recovery_attempts`` narrower recovery attempts already tried
(e.g. sanctioned reconnects) so the operator sees escalation history.
* ``inventory_complete`` bool. **Must** be explicitly True; a missing or
falsy value forces a deny (fail closed).
* ``incomplete_reasons`` optional reasons the inventory is incomplete.
The coordinator never restarts anything: ``restart_performed`` is always
False and the mutative apply path is a later drain-gated child.
"""
moment = now or _utc_now()
reasons: list[str] = []
authorization_reasons: list[str] = []
# ``None`` preserves the pre-#663 impact-only API for callers that have not
# yet been migrated. All MCP requests pass an explicit class and therefore
# take the fail-closed policy path.
policy_enforced = restart_class is not None
try:
resolved_class = resolve_restart_class(
restart_class or RestartClass.FULL_MCP_RESTART
)
policy = RESTART_CLASS_POLICIES[resolved_class]
unknown_class = False
except ValueError as exc:
resolved_class = None
policy = None
unknown_class = True
authorization_reasons.append(str(exc))
normalized_role = str(requester_role or "").strip().lower()
granted = {str(p).strip() for p in (requester_permissions or ())}
if policy_enforced and policy is not None:
permission_authorized = policy.required_permission in granted
role_authorized = normalized_role in policy.request_roles
if not permission_authorized:
authorization_reasons.append(
f"missing required permission {policy.required_permission!r}"
)
if not role_authorized:
authorization_reasons.append(
f"role {normalized_role or 'unknown'!r} may not request "
f"{policy.restart_class.value}"
)
elif unknown_class:
permission_authorized = False
role_authorized = False
else:
permission_authorized = True
role_authorized = True
if policy_enforced and policy is not None:
approval = policy.approval_requirement
if approval == "self_service":
approval_satisfied = True
elif approval == "controller_approval_plus_infrastructure_operator":
approval_satisfied = bool(controller_approved and operator_authorized)
else:
approval_satisfied = bool(controller_approved)
if not approval_satisfied:
authorization_reasons.append(
f"approval requirement not satisfied: {approval}"
)
elif unknown_class:
approval_satisfied = False
else:
approval_satisfied = True
inventory_complete = bool(inventory.get("inventory_complete", False))
incomplete_reasons = [str(r) for r in (inventory.get("incomplete_reasons") or [])]
sessions_raw: Sequence[Mapping[str, Any]] = inventory.get("sessions") or []
leases_raw: Sequence[Mapping[str, Any]] = inventory.get("leases") or []
terminal_lock = inventory.get("terminal_lock") or None
prior_recovery_attempts = [
dict(a) for a in (inventory.get("prior_recovery_attempts") or [])
]
# #669: broad restarts require a prior narrow-attempt log unless break-glass.
playbook_escalation: dict[str, Any] = {}
attempt_log_satisfied = True
if policy_enforced and resolved_class is not None:
try:
escalation = recovery_playbook.assess_escalation(
resolved_class.value,
prior_recovery_attempts=prior_recovery_attempts,
break_glass=bool(break_glass),
)
playbook_escalation = escalation.as_dict()
attempt_log_satisfied = bool(escalation.allowed)
if not attempt_log_satisfied:
authorization_reasons.extend(list(escalation.reasons))
except ValueError as exc:
# Unknown mapping should never happen for enum values; fail closed.
attempt_log_satisfied = False
playbook_escalation = {
"allowed": False,
"reasons": [str(exc)],
"playbook_version": recovery_playbook.PLAYBOOK_VERSION,
}
authorization_reasons.append(str(exc))
session_impacts = [
_classify_session(
s,
now=moment,
requesting_session_id=requesting_session_id,
heartbeat_stale_seconds=session_heartbeat_stale_seconds,
)
for s in sessions_raw
]
lease_impacts = [_classify_lease(l) for l in leases_raw]
# Route impact through the selected class. Narrow classes never inherit a
# full-runtime drain merely because unrelated work exists.
target_complete = True
if resolved_class in {
RestartClass.CLIENT_RECONNECT,
RestartClass.SESSION_RECONNECT,
RestartClass.CONFIGURATION_RELOAD,
}:
scoped_sessions: list[SessionImpact] = []
scoped_leases: list[LeaseImpact] = []
elif resolved_class == RestartClass.WORKER_RESTART:
selected_session = (target_session_id or "").strip()
target_complete = bool(selected_session)
scoped_sessions = [
s for s in session_impacts if s.session_id == selected_session
]
scoped_leases = [
l for l in lease_impacts if l.session_id == selected_session
]
elif resolved_class == RestartClass.ROLE_RUNTIME_RESTART:
selected_role = (target_role or "").strip().lower()
target_complete = bool(selected_role)
scoped_sessions = [
s for s in session_impacts if str(s.role or "").lower() == selected_role
]
scoped_leases = [
l for l in lease_impacts if str(l.role or "").lower() == selected_role
]
elif resolved_class == RestartClass.CONNECTOR_RESTART:
selected_connector = (target_connector or "").strip()
target_complete = bool(selected_connector)
scoped_sessions = [
s for s in session_impacts if s.connector == selected_connector
]
scoped_leases = [
l for l in lease_impacts if l.connector == selected_connector
]
else:
scoped_sessions = list(session_impacts)
scoped_leases = list(lease_impacts)
if policy_enforced and not target_complete:
authorization_reasons.append(
f"target required for {resolved_class.value if resolved_class else 'unknown class'}"
)
other_live_sessions = [
s for s in scoped_sessions if s.live and not s.is_requester
]
disruptive_leases = [l for l in scoped_leases if l.disruptive]
critical_sections = [l for l in scoped_leases if l.is_critical_section]
mutations = [l for l in scoped_leases if l.is_mutation]
terminal_lock_in_scope = (
terminal_lock
if resolved_class
not in {
RestartClass.CLIENT_RECONNECT,
RestartClass.SESSION_RECONNECT,
}
else None
)
affected_issues = sorted(
{
l.work_number
for l in disruptive_leases
if l.work_kind == "issue" and l.work_number is not None
}
)
affected_prs = sorted(
{
l.work_number
for l in disruptive_leases
if l.work_kind == "pr" and l.work_number is not None
}
)
disruptive = bool(
disruptive_leases or other_live_sessions or terminal_lock_in_scope
)
authorization_ok = bool(
not unknown_class
and permission_authorized
and role_authorized
and approval_satisfied
and target_complete
and attempt_log_satisfied
)
if policy_enforced and not authorization_ok:
verdict = VERDICT_UNSAFE
allow_restart = False
reasons.append("restart class authorization denied (fail closed)")
reasons.extend(authorization_reasons)
elif not inventory_complete:
verdict = VERDICT_UNSAFE
allow_restart = False
reasons.append(
"inventory incomplete: restart evaluation cannot confirm blast "
"radius — deny (fail closed, #658)"
)
reasons.extend(incomplete_reasons)
elif not disruptive:
verdict = VERDICT_SAFE
allow_restart = True
reasons.append("no other live sessions, live leases, or terminal lock")
elif operator_override:
verdict = VERDICT_OVERRIDE
allow_restart = True
reasons.append(
"live work present; operator override accepts the blast radius"
)
else:
verdict = VERDICT_UNSAFE
allow_restart = False
reasons.append(
"live work would be disrupted; restart denied without operator "
"override"
)
if critical_sections and inventory_complete:
reasons.append(
f"{len(critical_sections)} critical section(s) in flight "
"(active lease with a live owner)"
)
if terminal_lock_in_scope:
reasons.append("active terminal (merge) lock present")
override_would_allow = bool(inventory_complete and disruptive)
blast_radius = _blast_radius(
session_count=len(other_live_sessions),
work_count=len(affected_issues) + len(affected_prs),
mutation_count=len(mutations),
)
# Acknowledgement is a later child (drain protocol); expose per-session
# placeholders so the console can render the ack column now.
ack_state = {s.session_id: "pending" for s in other_live_sessions}
counts = {
"sessions_total": len(session_impacts),
"sessions_live_other": len(other_live_sessions),
"leases_total": len(lease_impacts),
"leases_disruptive": len(disruptive_leases),
"critical_sections": len(critical_sections),
"mutations": len(mutations),
"affected_issues": len(affected_issues),
"affected_prs": len(affected_prs),
"prior_recovery_attempts": len(prior_recovery_attempts),
"attempt_log_satisfied": attempt_log_satisfied,
}
audit_record = {
"event": "restart_impact_evaluated",
"coordinator_version": COORDINATOR_VERSION,
"restart_class": (
resolved_class.value if resolved_class else str(restart_class or "")
),
"required_permission": (
policy.required_permission if policy is not None else None
),
"evaluated_at": moment.isoformat(),
"dry_run": dry_run,
"operator_override": bool(operator_override),
"requesting_session_id": requesting_session_id,
"inventory_complete": inventory_complete,
"verdict": verdict,
"allow_restart": allow_restart,
"blast_radius": blast_radius,
"counts": counts,
"attempt_log_satisfied": attempt_log_satisfied,
"break_glass": bool(break_glass),
"playbook_version": recovery_playbook.PLAYBOOK_VERSION,
}
return RestartImpactReport(
coordinator_version=COORDINATOR_VERSION,
restart_class=(
resolved_class.value if resolved_class else str(restart_class or "")
),
restart_policy=policy.as_dict() if policy is not None else {},
policy_enforced=policy_enforced,
permission_authorized=permission_authorized,
role_authorized=role_authorized,
approval_satisfied=approval_satisfied,
authorization_reasons=authorization_reasons,
evaluated_at=moment.isoformat(),
dry_run=dry_run,
restart_performed=False,
inventory_complete=inventory_complete,
verdict=verdict,
allow_restart=allow_restart,
override_would_allow=override_would_allow,
operator_override=bool(operator_override),
blast_radius=blast_radius,
reasons=reasons,
affected_sessions=session_impacts,
affected_leases=lease_impacts,
critical_sections=critical_sections,
affected_issues=affected_issues,
affected_prs=affected_prs,
mutations=mutations,
terminal_lock=(
dict(terminal_lock_in_scope)
if isinstance(terminal_lock_in_scope, Mapping)
else terminal_lock_in_scope
),
ack_state=ack_state,
prior_recovery_attempts=prior_recovery_attempts,
counts=counts,
audit_record=audit_record,
incomplete_reasons=incomplete_reasons,
playbook_escalation=playbook_escalation,
attempt_log_satisfied=attempt_log_satisfied,
break_glass=bool(break_glass),
)
-37
View File
@@ -74,43 +74,6 @@ def check_author_mutation_namespace(
return True, []
def check_author_role_kind(
mutation_task: str,
profile: dict,
) -> tuple[bool, list[str]]:
"""Author-exclusive wall for durable-lock mutations (#953 F1).
``check_author_mutation_namespace`` walls off reviewer-bound sessions, which
is the whole gate for tasks whose required permission is itself author-only
(``gitea.pr.create``, ``gitea.repo.commit``). It is *not* sufficient for a
task gated on ``gitea.issue.comment``, which every configured role holds: a
merger, controller, or reconciler session would clear both the namespace
check and the permission gate and still reach the durable write.
Opt-in per call site and additive. It refuses any active profile whose
derived role kind is not exactly ``author`` for a task the router declares
author-required, and grants nothing to anyone a ``mixed`` profile is
refused rather than admitted.
"""
required_role = role_session_router.required_role_for_task(mutation_task)
if required_role != "author":
return True, []
allowed = profile.get("allowed_operations") or []
forbidden = profile.get("forbidden_operations") or []
active_role = derive_role_kind(allowed, forbidden)
if active_role == "author":
return True, []
profile_name = profile.get("profile_name") or ""
namespace = infer_mcp_namespace(profile_name)
return False, [
f"author mutation '{mutation_task}' blocked: active session role kind is "
f"'{active_role}', not 'author' ({profile_name} / {namespace}); this "
"operation writes a durable author issue lock and is author-exclusive",
]
def mutation_audit_context(mutation_task: str, profile: dict, *,
remote=None, repository=None) -> dict:
"""Structured mutation metadata for audit records (#209)."""
-12
View File
@@ -73,12 +73,6 @@ AUTHOR_TASKS = frozenset({
"claim_issue",
"create_branch",
"push_branch",
"bootstrap_author_issue_worktree",
"gitea_bootstrap_author_issue_worktree",
# #953: recovery of an incomplete bootstrap lock is an author-only durable
# state mutation and belongs to the same class as bootstrap itself.
"recover_incomplete_bootstrap_lock",
"gitea_recover_incomplete_bootstrap_lock",
"create_pr",
"comment_pr",
"address_pr_change_requests",
@@ -116,12 +110,6 @@ TASK_REQUIRED_ROLE = {
"claim_issue": "author",
"create_branch": "author",
"push_branch": "author",
# #953: without this entry ``required_role_for_task`` returns None and
# ``role_namespace_gate.check_author_mutation_namespace`` short-circuits to
# "allowed" — the namespace wall on the recovery tool would be inert. The
# capability map already records the same role; both tables must agree.
"recover_incomplete_bootstrap_lock": "author",
"gitea_recover_incomplete_bootstrap_lock": "author",
"create_pr": "author",
"comment_pr": "author",
"address_pr_change_requests": "author",
+9 -146
View File
@@ -1,16 +1,9 @@
"""Root checkout guard (#475).
The project root checkout is the stable control checkout on its integration
branch. Author/reviewer/merge flows must fail closed when the control checkout
is contaminated (wrong branch, detached HEAD, dirty, or HEAD behind/ahead of the
tracking integration ref). Isolated ``branches/...`` worktrees remain allowed.
The tracking ref is *derived per repository* (#983) rather than assumed to be
``prgs/master``: a namespace bound to another repository say remote ``MDCPS``
on branch ``dev`` is gated against ``refs/remotes/MDCPS/dev``. Derivation is
delegated to :mod:`canonical_repository_root`, the authoritative
repository/context resolver, so gating and parity reporting share one resolved
target instead of maintaining two disagreeing hardcoded defaults.
The project root checkout is the stable control checkout on master/prgs/master.
Author/reviewer/merge flows must fail closed when the control checkout is
contaminated (wrong branch, detached HEAD, dirty, or HEAD behind/ahead of
prgs/master). Isolated ``branches/...`` worktrees remain allowed.
"""
from __future__ import annotations
@@ -18,7 +11,6 @@ from __future__ import annotations
import os
import subprocess
import canonical_repository_root
from author_mutation_worktree import is_path_under_branches
from reviewer_worktree import parse_dirty_tracked_files
@@ -28,65 +20,19 @@ REMEDIATION = (
)
BASE_BRANCHES = frozenset({"master", "main", "dev"})
# Legacy PRGS-specific probe order. Retained only for callers that pass an
# explicit ``remote_refs`` override; it is no longer the silent default, because
# inheriting it in a non-PRGS checkout compared that checkout against a ref it
# can never have (#983).
REMOTE_MASTER_REFS = ("prgs/master", "refs/remotes/prgs/master")
def _derive_probe_refs(root: str, explicit_remote: str | None) -> dict:
"""Derive the ordered tracking refs to probe for *root*.
Returns the derivation payload from :mod:`canonical_repository_root` plus a
``refs`` tuple, which is empty when the target is not provable.
"""
derived = canonical_repository_root.resolve_target_base_ref(
root, explicit_remote=explicit_remote
)
return {
"refs": tuple(derived.get("tracking_refs") or ()),
"remote": derived.get("remote"),
"branch": derived.get("branch"),
"source": derived.get("source"),
"proven": bool(derived.get("proven")),
"reason_code": derived.get("reason_code"),
"reasons": list(derived.get("reasons") or []),
"cached_remote_head_branch": derived.get("cached_remote_head_branch"),
"cached_remote_head_conflicts": bool(derived.get("cached_remote_head_conflicts")),
}
def resolve_remote_master_sha(
canonical_repo_root: str,
*,
remote_refs: tuple[str, ...] | None = None,
explicit_remote: str | None = None,
) -> str | None:
"""Return the commit SHA for the tracking integration ref when available.
This remains the single place that turns a ref into a SHA. Returns None when
the target cannot be derived or resolved exactly what this function already
returned when ``rev-parse`` failed. Callers that must fail closed on missing
evidence (the #749/#757 bootstrap path) surface that None as *missing
evidence*, never as "no constraint".
*explicit_remote* is a caller-supplied disambiguation and is named that way
deliberately: an internally inferred remote handed back in would suppress the
resolver's ambiguity gate (#983 B2). No production caller supplies it.
"""
"""Return the commit SHA for the tracking master ref when available."""
root = (canonical_repo_root or "").strip()
if not root:
return None
if remote_refs:
probe: tuple[str, ...] = tuple(remote_refs)
else:
derived = _derive_probe_refs(root, explicit_remote)
if not derived["proven"]:
return None
probe = derived["refs"]
for ref in probe:
for ref in remote_refs or REMOTE_MASTER_REFS:
res = subprocess.run(
["git", "-C", root, "rev-parse", "--verify", ref],
capture_output=True,
@@ -100,84 +46,6 @@ def resolve_remote_master_sha(
return None
def resolve_remote_master_ref_state(
canonical_repo_root: str,
*,
remote_refs: tuple[str, ...] | None = None,
explicit_remote: str | None = None,
) -> dict:
"""Resolve the tracking integration ref together with its commit SHA.
Returns ``sha``, the ``ref`` it came from, the derived ``remote`` /
``branch``, a machine-checkable ``reason_code``, and ``reasons``. ``sha`` is
None whenever the target cannot be resolved never a fallback to some other
repository's commit.
The SHA itself is obtained through :func:`resolve_remote_master_sha` rather
than by re-probing here, so one public function stays authoritative for
ref-to-SHA resolution. An explicit *remote_refs* keeps the historical
behaviour exactly: those refs are probed in order and no derivation happens.
"""
root = (canonical_repo_root or "").strip()
state: dict = {
"sha": None,
"ref": None,
"remote": None,
"branch": None,
"source": None,
"reason_code": None,
"reasons": [],
"cached_remote_head_branch": None,
"cached_remote_head_conflicts": False,
}
if not root:
state["reasons"].append("no canonical repository root supplied (fail closed)")
return state
if remote_refs:
probe: tuple[str, ...] = tuple(remote_refs)
state["source"] = "explicit_remote_refs"
else:
derived = _derive_probe_refs(root, explicit_remote)
state["remote"] = derived["remote"]
state["branch"] = derived["branch"]
state["source"] = derived["source"]
state["cached_remote_head_branch"] = derived["cached_remote_head_branch"]
state["cached_remote_head_conflicts"] = derived["cached_remote_head_conflicts"]
if not derived["proven"]:
state["reason_code"] = derived["reason_code"]
state["reasons"] = derived["reasons"]
return state
probe = derived["refs"]
sha = resolve_remote_master_sha(root, remote_refs=probe, explicit_remote=explicit_remote)
if not sha:
state["reasons"].append(
"tracking integration ref "
f"{' / '.join(probe) if probe else '(none derived)'} does not resolve in "
f"'{root}' (fail closed)"
)
return state
state["sha"] = sha
# Name the ref that actually carries this commit. When the SHA comes from a
# test double no probe will match, so fall back to the first derived ref,
# which is the one the guard is conceptually comparing against.
for ref in probe:
res = subprocess.run(
["git", "-C", root, "rev-parse", "--verify", ref],
capture_output=True,
text=True,
check=False,
)
if res.returncode == 0 and (res.stdout or "").strip() == sha:
state["ref"] = ref
break
else:
state["ref"] = probe[0] if probe else None
return state
resolve_tracking_master_sha = resolve_remote_master_sha
@@ -191,9 +59,8 @@ def assess_root_checkout_guard(
remote_master_sha: str | None,
resolved_role: str | None = None,
actual_role: str | None = None,
remote_master_ref: str | None = None,
) -> dict:
"""Fail closed when the control checkout is not clean on its integration ref.
"""Fail closed when the control checkout is not clean master/prgs/master.
``resolved_role`` is the preflight-resolved *task* role and ``actual_role``
is the *active profile* role (#540). The reconciler exemption honours either
@@ -235,13 +102,9 @@ def assess_root_checkout_guard(
)
if remote_master_sha and head_sha and head_sha != remote_master_sha:
# #983: name the ref that was actually compared. Reporting a literal
# 'prgs/master' in a checkout gated against refs/remotes/MDCPS/dev sends
# the operator to inspect a ref that repository does not have.
ref_label = (remote_master_ref or "").strip() or "the tracking integration ref"
reasons.append(
f"control checkout HEAD does not match {ref_label} "
f"(HEAD {head_sha[:12]}, {ref_label} {remote_master_sha[:12]})"
"control checkout HEAD does not match prgs/master "
f"(HEAD {head_sha[:12]}, prgs/master {remote_master_sha[:12]})"
)
proven = not reasons
+8 -10
View File
@@ -43,21 +43,19 @@ repo_root="$(cd "$script_dir/.." && pwd)"
# Enforce issue-linked, traceable branch names (issue → branch → worktree → PR).
if [[ "$allow_unlinked" -eq 0 ]]; then
if [[ "$dry_run" -eq 0 ]] && [[ ! "$branch" =~ ^review/pr-[0-9]+-.+ ]]; then
locked_branch=$(python3 -c "
locked_branch=$(python3 -c "
import sys
sys.path.insert(0, '$repo_root')
import issue_lock_store
print(issue_lock_store.resolve_locked_branch_for_session('$branch'))
")
if [[ -z "$locked_branch" ]]; then
echo "Error: No session issue lock is bound. Call gitea_lock_issue before branch creation (fail closed)." >&2
exit 2
fi
if [[ "$branch" != "$locked_branch" ]]; then
echo "Error: Requested branch '$branch' does not match locked branch '$locked_branch' (fail closed)." >&2
exit 2
fi
if [[ -z "$locked_branch" ]]; then
echo "Error: No session issue lock is bound. Call gitea_lock_issue before branch creation (fail closed)." >&2
exit 2
fi
if [[ "$branch" != "$locked_branch" ]]; then
echo "Error: Requested branch '$branch' does not match locked branch '$locked_branch' (fail closed)." >&2
exit 2
fi
if [[ "$branch" =~ ^(fix|feat|docs|chore)/issue-[0-9]+-.+ ]] \
-22
View File
@@ -204,28 +204,6 @@ proposed command before running it; `gitea_audit_runtime_recovery_contamination`
to inspect or (reconciler-only) clear the marker. Full contrast in
`docs/mcp-namespace-eof-recovery.md`.
## Connected is not attached (#708)
A host/CLI MCP inventory showing **Connected** is not proof the tools are usable.
The active session can expose **none** of a Connected server's tool namespaces —
a *session attachment* failure, distinct from config drift (#672),
transport-closed (#584), and resolver EOF (#685).
Required preflight proof is **live tool visibility plus `gitea_whoami` on the role
namespace**, never host Connected status alone. Call
`gitea_assess_mcp_namespace_attachment` with the Connected set and the namespaces
actually attached to the session; it returns the typed condition
**mcp_connected_namespaces_missing** with per-namespace connected-vs-attached proof,
and records the verdict so review and merge fail closed while a required namespace
is unattached.
Only sanctioned recovery: the client attach/reconnect path
(`gitea_request_mcp_reconnect`), then full preflight
(`gitea_whoami``gitea_resolve_task_capability` → task). Never recover by direct
module import, CLI/raw API mutation, profile hopping, session-state overrides,
process kills, or `.env`/mtime edits. A final report must not claim a healthy
session without attachment proof.
## Shell Spawn Hard-Stop Rule
`exit_code: -1` with empty stdout/stderr means the shell failed to spawn — not a
-5
View File
@@ -63,11 +63,6 @@ CONTAMINATION_GATED_TASKS = frozenset({
"merge_pr",
"delete_branch",
"complete_issue",
# Web console recovery playbooks that write (#644). These mutate runtime
# binding and process state, so a live contamination marker must block them
# exactly as it blocks the Gitea-side mutations above. The reconciler
# cleanup playbook is the designated remedy and is exempted by its caller.
"console_recovery_apply",
})
CONTAMINATION_KIND = "stable_branch_push"
-154
View File
@@ -32,56 +32,6 @@ TASK_CAPABILITY_MAP: dict[str, dict[str, str]] = {
"permission": "gitea.issue.comment",
"role": "author",
},
# #790 Slice A: prove an owned author lease is still active. Strictly
# narrower than lock_issue — it can only slide a lease this exact session
# already owns, never acquire, take over, or revive one — so it gates on the
# same authority rather than introducing an operation name that every
# already-configured author profile would be missing.
"heartbeat_issue_lock": {
"permission": "gitea.issue.comment",
"role": "author",
},
# #953: target-specific upgrade of an incomplete bootstrap lock (explicit
# operation, never a widening of lock_issue). Author-only, and the tool
# additionally proves exact-owner claimant match before writing.
"recover_incomplete_bootstrap_lock": {
"permission": "gitea.issue.comment",
"role": "author",
},
"gitea_recover_incomplete_bootstrap_lock": {
"permission": "gitea.issue.comment",
"role": "author",
},
# #953: read-only lock contract inspection. Read permission only — it must
# never be able to mutate.
"inspect_issue_lock_contract": {
"permission": "gitea.read",
"role": "author",
},
"gitea_inspect_issue_lock_contract": {
"permission": "gitea.read",
"role": "author",
},
# #860: dirty orphaned same-claimant worktree recovery (explicit operation).
"recover_dirty_orphaned_issue_worktree": {
"permission": "gitea.issue.comment",
"role": "author",
},
"gitea_recover_dirty_orphaned_issue_worktree": {
"permission": "gitea.issue.comment",
"role": "author",
},
# #864: dirty-preserving same-claimant author-session rebind (dead owner PID).
# Author MCP tool path. Reconciler execute is gated inside the tool via
# authorize_reconciler_execute + role_kind checks (not this map entry).
"rebind_dirty_same_claimant_author_session": {
"permission": "gitea.issue.comment",
"role": "author",
},
"gitea_rebind_dirty_same_claimant_author_session": {
"permission": "gitea.issue.comment",
"role": "author",
},
"set_issue_labels": {
"permission": "gitea.issue.comment",
"role": "author",
@@ -108,14 +58,6 @@ TASK_CAPABILITY_MAP: dict[str, dict[str, str]] = {
"permission": "gitea.branch.create",
"role": "author",
},
"bootstrap_author_issue_worktree": {
"permission": "gitea.branch.create",
"role": "author",
},
"gitea_bootstrap_author_issue_worktree": {
"permission": "gitea.branch.create",
"role": "author",
},
"push_branch": {
"permission": "gitea.branch.push",
"role": "author",
@@ -153,45 +95,6 @@ TASK_CAPABILITY_MAP: dict[str, dict[str, str]] = {
"permission": "gitea.branch.push",
"role": "author",
},
# #662: post-restart reconcile is read-only inventory + pure classification.
# Durable follow-up issue creation is a separate apply path (not this task).
"reconcile_after_restart": {
"permission": "gitea.read",
"role": "author",
},
"gitea_reconcile_after_restart": {
"permission": "gitea.read",
"role": "author",
},
# #978: instance-level fleet identity/health snapshot. gitea.read is the
# operation gate; the tool additionally restricts role_kind to
# controller|reconciler so author/reviewer/merger cannot use it as a
# mutation surface and no unrelated write permission is introduced.
"snapshot_instance_fleet": {
"permission": "gitea.read",
"role": "controller",
},
"gitea_snapshot_instance_fleet": {
"permission": "gitea.read",
"role": "controller",
},
# #644: Phase 2 Web Console recovery tasks.
"clear_stale_binding": {
"permission": "gitea.read",
"role": "author",
},
"rebind_session_worktree": {
"permission": "gitea.read",
"role": "author",
},
# The console playbook orchestrates gitea_reconcile_merged_cleanups, whose
# own gate is gitea.read (matching the existing reconcile_merged_cleanups
# entry). Declaring a stricter permission here stated a second, conflicting
# authority for one operation.
"reconcile_cleanups": {
"permission": "gitea.read",
"role": "reconciler",
},
# PR synchronization lifecycle: assess is read-only (any role with gitea.read);
# update-by-merge is author-only and mutates the PR head via Gitea API.
"assess_pr_sync_status": {
@@ -398,25 +301,6 @@ TASK_CAPABILITY_MAP: dict[str, dict[str, str]] = {
"permission": "gitea.branch.delete",
"role": "reconciler",
},
# #970: auditing missing worktree bindings is read-only; retiring one is a
# control-plane cleanup mutation and carries the same reconciler-only
# authority as any other reconciliation cleanup (review 644 B2).
"audit_missing_worktree_bindings": {
"permission": "gitea.read",
"role": "reconciler",
},
"gitea_audit_missing_worktree_bindings": {
"permission": "gitea.read",
"role": "reconciler",
},
"reconcile_missing_worktree_bindings": {
"permission": "gitea.branch.delete",
"role": "reconciler",
},
"gitea_reconcile_missing_worktree_bindings": {
"permission": "gitea.branch.delete",
"role": "reconciler",
},
"work_issue": {
"permission": "gitea.pr.create",
"role": "author",
@@ -451,21 +335,6 @@ TASK_CAPABILITY_MAP: dict[str, dict[str, str]] = {
"role": "controller",
},
# #642: sanctioned host-daemon lifecycle controls. Deliberately *not* a
# ``gitea.*`` operation — restarting an MCP namespace is a host action, not
# a Gitea API call, and no configured Gitea profile should be able to
# satisfy it by accident. Authority comes from the console RBAC model plus
# out-of-band operator authorization (#630); these entries exist so the
# console cannot invent an authority the capability layer never declared.
"restart_namespace": {
"permission": "runtime.restart_namespace",
"role": "controller",
},
"reload_namespace": {
"permission": "runtime.reload_namespace",
"role": "controller",
},
# #601 first-class lease lifecycle — inspect/list need read; mutations gate on
# ownership in the control-plane DB (not a separate Gitea write permission).
"list_workflow_leases": {
@@ -598,15 +467,6 @@ TASK_CAPABILITY_MAP: dict[str, dict[str, str]] = {
"permission": "gitea.issue.comment",
"role": "author",
},
# #651 console analytics ingest — control-plane DB write, not a Gitea API
# call. Authority comes from console RBAC (operator+) plus phase gating;
# permission string is a non-Gitea runtime capability so no Gitea profile
# can satisfy it by accident.
"record_analytics_usage": {
"permission": "runtime.record_analytics_usage",
"role": "author",
},
}
@@ -617,15 +477,6 @@ TASK_CAPABILITY_MAP: dict[str, dict[str, str]] = {
# merger lease (#763).
_PREFLIGHT_TASK_TRANSITIONS = frozenset({
("review_pr", "acquire_reviewer_pr_lease"),
# #850: native author issue worktree bootstrap
("work_issue", "bootstrap_author_issue_worktree"),
("bootstrap_author_issue_worktree", "lock_issue"),
# #860: dirty-orphan recovery and related work_issue transitions (master)
("work_issue", "lock_issue"),
("work_issue", "recover_dirty_orphaned_issue_worktree"),
("work_issue", "gitea_recover_dirty_orphaned_issue_worktree"),
("work_issue", "commit_files"),
("work_issue", "gitea_commit_files"),
})
@@ -672,8 +523,6 @@ ROLE_EXCLUSIVE_TASKS: frozenset[str] = frozenset(
"gitea_release_merger_pr_lease",
"create_branch",
"push_branch",
"bootstrap_author_issue_worktree",
"gitea_bootstrap_author_issue_worktree",
"publish_unpublished_branch",
"create_pr",
"commit_files",
@@ -684,8 +533,6 @@ ROLE_EXCLUSIVE_TASKS: frozenset[str] = frozenset(
"delete_branch",
"cleanup_merged_pr_branch",
"reconciliation_cleanup",
"reconcile_missing_worktree_bindings",
"gitea_reconcile_missing_worktree_bindings",
"work_issue",
"work-issue",
}
@@ -701,7 +548,6 @@ ISSUE_MUTATION_TOOL_TASKS: dict[str, str] = {
"gitea_set_issue_labels": "set_issue_labels",
"gitea_cleanup_terminal_pr_labels": "cleanup_terminal_pr_labels",
"gitea_create_label": "create_label",
"gitea_bootstrap_author_issue_worktree": "bootstrap_author_issue_worktree",
"gitea_commit_files": "commit_files",
}
-2
View File
@@ -44,8 +44,6 @@ def _reset_mutation_authority(monkeypatch):
]:
monkeypatch.delenv(env_key, raising=False)
monkeypatch.setenv("GITEA_CLIENT_MANAGED", "1")
# Isolate durable session-state files so tests never share host cache (#559).
import tempfile
@@ -1,243 +0,0 @@
"""Allocator epic / child-only container pre-rank exclusion (#844).
Covers:
* Issue #631-shaped child-only epic is excluded before ranking.
* Implementable child issues remain eligible and can be selected.
* Ordinary issues that merely mention "epic" in title/body are not excluded.
* Excluded containers never receive assignments or workflow leases.
* Structured skip reason ``epic_or_child_only_container`` is reported.
"""
from __future__ import annotations
import os
import tempfile
import unittest
from allocator_service import (
OUTCOME_ASSIGNED,
OUTCOME_PREVIEW,
SKIP_EPIC_OR_CHILD_ONLY_CONTAINER,
WorkCandidate,
allocate_next_work,
classify_epic_or_child_only_container,
)
from control_plane_db import ControlPlaneDB
REMOTE = "prgs"
ORG = "Scaled-Tech-Consulting"
REPO = "Gitea-Tools"
# Minimal body mirroring issue #631 authoritative scope language.
_EPIC_631_BODY = """
## Scope (umbrella)
This epic owns the **product roadmap and linkage** for the Web Console.
Implementation is delivered via child issues only.
## Explicit non-goals
* Do not implement product features in this epic issue itself.
* No product feature implementation is claimed complete solely on this epic.
"""
_CHILD_BODY = """
## Problem
Operators need a workflow-event timeline model for Phase 1.
## Acceptance criteria
- [ ] Timeline model API exists
"""
def _issue(
number: int,
*,
title: str = "",
body: str = "",
labels: tuple[str, ...] = ("status:ready", "type:feature"),
priority: int = 20,
) -> WorkCandidate:
return WorkCandidate(
kind="issue",
number=number,
state="open",
labels=labels,
title=title or f"issue {number}",
body=body,
priority=priority,
)
class ClassifyEpicContainerTest(unittest.TestCase):
def test_631_shaped_body_and_title_is_container(self) -> None:
c = _issue(
631,
title="Epic: MCP Control Plane Web Console",
body=_EPIC_631_BODY,
)
is_c, detail = classify_epic_or_child_only_container(c)
self.assertTrue(is_c)
self.assertIsNotNone(detail)
self.assertIn("body_marker", detail or "")
def test_body_markers_without_epic_title(self) -> None:
c = _issue(
900,
title="Control plane roadmap tracker",
body="Implementation is delivered via child issues only.",
)
is_c, _ = classify_epic_or_child_only_container(c)
self.assertTrue(is_c)
def test_epic_label_alone_is_container(self) -> None:
c = _issue(
901,
title="Roadmap linkage",
body="Track children.",
labels=("status:ready", "type:epic"),
)
is_c, detail = classify_epic_or_child_only_container(c)
self.assertTrue(is_c)
self.assertIn("type:epic", detail or "")
def test_title_epic_prefix_alone_not_container(self) -> None:
"""Title-only 'Epic:' without body scope evidence stays eligible (#844)."""
c = _issue(
902,
title="Epic: something mentioned only in title",
body="Implement a concrete fix for the allocator skip list.",
)
is_c, detail = classify_epic_or_child_only_container(c)
self.assertFalse(is_c)
self.assertIsNone(detail)
def test_incidental_epic_word_not_container(self) -> None:
c = _issue(
903,
title="Document epic handoff conventions",
body=(
"Update the docs so implementable issues that mention an epic "
"remain independently executable."
),
)
is_c, _ = classify_epic_or_child_only_container(c)
self.assertFalse(is_c)
def test_prs_never_classified(self) -> None:
pr = WorkCandidate(
kind="pr",
number=10,
state="open",
title="Epic: fake",
body="Implementation is delivered via child issues only.",
head_sha="a" * 40,
priority=5,
)
is_c, _ = classify_epic_or_child_only_container(pr)
self.assertFalse(is_c)
class AllocateEpicContainerExclusionTest(unittest.TestCase):
def setUp(self) -> None:
self._tmp = tempfile.TemporaryDirectory()
self.addCleanup(self._tmp.cleanup)
self.db = ControlPlaneDB(os.path.join(self._tmp.name, "cp.sqlite3"))
def _alloc(self, candidates, **kwargs):
defaults = dict(
session_id="sess-844",
role="author",
remote=REMOTE,
org=ORG,
repo=REPO,
profile_name="prgs-author",
username="jcwalker3",
claims={},
apply=False,
)
defaults.update(kwargs)
return allocate_next_work(self.db, candidates=candidates, **defaults)
def test_631_shaped_epic_excluded_child_selected(self) -> None:
epic = _issue(
631,
title="Epic: MCP Control Plane Web Console",
body=_EPIC_631_BODY,
)
child = _issue(
637,
title="Web Console: Workflow-event timeline model (Phase 1)",
body=_CHILD_BODY,
)
res = self._alloc([epic, child], apply=False)
self.assertTrue(res["success"], res)
self.assertEqual(res["outcome"], OUTCOME_PREVIEW)
self.assertEqual(res["selected"]["number"], 637)
skipped = {s["number"]: s for s in res["skipped"]}
self.assertIn(631, skipped)
self.assertEqual(
skipped[631]["reason_code"], SKIP_EPIC_OR_CHILD_ONLY_CONTAINER
)
self.assertIn(SKIP_EPIC_OR_CHILD_ONLY_CONTAINER, skipped[631]["reason"])
def test_container_cannot_receive_assignment_or_lease(self) -> None:
epic = _issue(
631,
title="Epic: MCP Control Plane Web Console",
body=_EPIC_631_BODY,
)
res = self._alloc([epic], apply=True)
self.assertTrue(res["success"], res)
# Only container present → no safe work; never assigned_work.
self.assertNotEqual(res["outcome"], OUTCOME_ASSIGNED)
self.assertIsNone(res.get("assignment"))
self.assertIsNone(res.get("selected"))
skipped = {s["number"]: s for s in res["skipped"]}
self.assertEqual(
skipped[631]["reason_code"], SKIP_EPIC_OR_CHILD_ONLY_CONTAINER
)
# No lease row for the epic.
leases = self.db.list_active_leases(
remote=REMOTE, org=ORG, repo=REPO
) if hasattr(self.db, "list_active_leases") else []
# Prefer generic inventory if available.
if not leases and hasattr(self.db, "list_leases"):
leases = self.db.list_leases(remote=REMOTE, org=ORG, repo=REPO)
for lease in leases or []:
work_number = lease.get("work_number") if isinstance(lease, dict) else None
self.assertNotEqual(work_number, 631)
def test_incidental_epic_title_remains_eligible(self) -> None:
ordinary = _issue(
700,
title="Document epic handoff conventions",
body="Write runbook text about epic vs child issues.",
)
res = self._alloc([ordinary], apply=False)
self.assertTrue(res["success"], res)
self.assertEqual(res["selected"]["number"], 700)
self.assertEqual(res["skipped"], [])
def test_apply_selects_child_not_epic(self) -> None:
epic = _issue(
631,
title="Epic: MCP Control Plane Web Console",
body=_EPIC_631_BODY,
)
child = _issue(
637,
title="Web Console: Workflow-event timeline model (Phase 1)",
body=_CHILD_BODY,
)
res = self._alloc([epic, child], apply=True)
self.assertTrue(res["success"], res)
self.assertEqual(res["outcome"], OUTCOME_ASSIGNED)
self.assertEqual(res["selected"]["number"], 637)
self.assertEqual(res["assignment"]["work_number"], 637)
if __name__ == "__main__":
unittest.main()
-158
View File
@@ -7,7 +7,6 @@ import tempfile
import threading
import unittest
from concurrent.futures import ThreadPoolExecutor, as_completed
from datetime import datetime, timezone
from allocator_service import (
OUTCOME_ASSIGNED,
@@ -16,7 +15,6 @@ from allocator_service import (
OUTCOME_PREVIEW,
OUTCOME_WAIT,
WorkCandidate,
_drop_expired_claims,
allocate_next_work,
candidate_from_dict,
classify_skip,
@@ -364,161 +362,5 @@ class AllocatorServiceTest(unittest.TestCase):
self.assertIn("unavailable", res["reasons"][0].lower())
class SideEffectFreeAllocationTest(unittest.TestCase):
"""``side_effect_free`` dry runs write nothing to the control plane (#643).
A plain ``apply=False`` still called ``upsert_session`` and
``expire_stale_leases`` before the apply branch was consulted, so a caller
advertising a read-only preview mutated on every call one unreferenced
session row per preview, plus a global lease sweep.
"""
def setUp(self) -> None:
self._tmp = tempfile.TemporaryDirectory()
self.db = ControlPlaneDB(os.path.join(self._tmp.name, "cp.sqlite3"))
def tearDown(self) -> None:
self._tmp.cleanup()
def _alloc(self, **kwargs):
defaults = dict(
db=self.db,
session_id="s-preview",
role="author",
remote="prgs",
org="org",
repo="repo",
candidates=[
WorkCandidate(kind="issue", number=643, labels=("status:ready",))
],
apply=False,
profile_name="prgs-author",
username="jcwalker3",
)
defaults.update(kwargs)
return allocate_next_work(**defaults)
def _session_ids(self) -> set[str]:
return {str(r.get("session_id")) for r in self.db.list_sessions()}
def test_side_effect_free_preview_writes_no_session_row(self):
before = self._session_ids()
result = self._alloc(side_effect_free=True)
self.assertEqual(result["outcome"], OUTCOME_PREVIEW)
self.assertEqual(self._session_ids(), before)
self.assertNotIn("s-preview", self._session_ids())
def test_plain_dry_run_still_registers_a_session(self):
# The default is unchanged for every existing caller.
self._alloc()
self.assertIn("s-preview", self._session_ids())
def test_repeated_previews_do_not_accumulate_rows(self):
for index in range(5):
self._alloc(side_effect_free=True, session_id=f"s-{index}")
self.assertEqual(self._session_ids(), set())
def test_side_effect_free_does_not_sweep_stale_leases(self):
self.db.upsert_session(session_id="owner", role="author", pid=1)
assigned = self.db.assign_and_lease(
session_id="owner",
role="author",
remote="prgs",
org="org",
repo="repo",
kind="issue",
number=999,
lease_ttl_seconds=-60, # already expired
)
self.assertEqual(assigned.outcome, "assigned")
self._alloc(side_effect_free=True)
# The expired row is still 'active' in the DB: nothing swept it.
statuses = {
r["lease_id"]: r["status"]
for r in self.db.list_leases(
remote="prgs", org="org", repo="repo",
statuses=("active", "expired"),
)
}
self.assertEqual(statuses.get(assigned.lease_id), "active")
def test_expired_claims_are_filtered_in_memory_so_work_stays_selectable(self):
"""The read-only mirror of the sweep: expired claims must not block."""
self.db.upsert_session(session_id="owner", role="author", pid=1)
self.db.assign_and_lease(
session_id="owner",
role="author",
remote="prgs",
org="org",
repo="repo",
kind="issue",
number=643,
lease_ttl_seconds=-60, # expired: must not withhold #643
)
result = self._alloc(side_effect_free=True)
self.assertEqual(result["outcome"], OUTCOME_PREVIEW)
self.assertEqual(result["selected"]["number"], 643)
def test_a_live_claim_still_withholds_the_work(self):
self.db.upsert_session(session_id="owner", role="author", pid=1)
self.db.assign_and_lease(
session_id="owner",
role="author",
remote="prgs",
org="org",
repo="repo",
kind="issue",
number=643,
lease_ttl_seconds=3600,
)
result = self._alloc(side_effect_free=True)
self.assertNotEqual(result["outcome"], OUTCOME_ASSIGNED)
self.assertNotEqual((result.get("selected") or {}).get("number"), 643)
def test_side_effect_free_with_apply_fails_closed(self):
result = self._alloc(side_effect_free=True, apply=True)
self.assertFalse(result["success"])
self.assertEqual(result["outcome"], OUTCOME_NO_SAFE)
self.assertIsNone(result["assignment"])
self.assertIn("incompatible with apply", result["reasons"][0])
# And it reserved nothing.
self.assertEqual(
self.db.list_leases(remote="prgs", org="org", repo="repo"), []
)
class DropExpiredClaimsTest(unittest.TestCase):
"""The in-memory expiry filter behind side-effect-free previews (#643)."""
def test_unparseable_expiry_is_kept_rather_than_assumed_free(self):
claims = {
("issue", 1): {"lease_id": "l1", "expires_at": "not-a-date"},
("issue", 2): {"lease_id": "l2"},
("issue", 3): {"lease_id": "l3", "expires_at": None},
}
self.assertEqual(_drop_expired_claims(claims), claims)
def test_expired_dropped_and_future_kept(self):
now = datetime(2026, 7, 25, 12, 0, tzinfo=timezone.utc)
claims = {
("issue", 1): {"expires_at": "2026-07-25T11:59:59+00:00"},
("issue", 2): {"expires_at": "2026-07-25T12:00:01+00:00"},
("issue", 3): {"expires_at": "2026-07-25T12:00:00+00:00"}, # boundary
}
kept = _drop_expired_claims(claims, now=now)
self.assertEqual(set(kept), {("issue", 2)})
def test_naive_and_zulu_timestamps_are_treated_as_utc(self):
now = datetime(2026, 7, 25, 12, 0, tzinfo=timezone.utc)
claims = {
("issue", 1): {"expires_at": "2026-07-25T11:00:00"}, # naive, past
("issue", 2): {"expires_at": "2026-07-25T13:00:00Z"}, # zulu, future
}
kept = _drop_expired_claims(claims, now=now)
self.assertEqual(set(kept), {("issue", 2)})
if __name__ == "__main__":
unittest.main()
-601
View File
@@ -1,601 +0,0 @@
"""Regression test suite for native author issue worktree bootstrap (#850)."""
from __future__ import annotations
import json
import os
import shutil
import subprocess
import tempfile
import unittest
from unittest import mock
import author_issue_bootstrap
import task_capability_map
def _concurrent_bootstrap_worker(args: tuple[str, int, str, str, str, str]) -> dict:
repo_dir, issue_num, key, lock_dir, journal_dir, master_sha = args
os.environ["GITEA_BOOTSTRAP_JOURNAL_DIR"] = journal_dir
return author_issue_bootstrap.bootstrap_author_issue_worktree(
issue_number=issue_num,
canonical_repo_root=repo_dir,
expected_base_sha=master_sha,
idempotency_key=key,
lock_dir=lock_dir,
owner_session="session-concurrent-test",
active_identity="jcwalker3",
active_profile="prgs-author",
)
class TestAuthorIssueBootstrap(unittest.TestCase):
"""Test suite covering AC1-AC10 and comment #14959 specification."""
def setUp(self):
self.tmp_dir = tempfile.mkdtemp(prefix="test_bootstrap_")
self.repo_dir = os.path.join(self.tmp_dir, "repo")
os.makedirs(self.repo_dir)
# Initialize synthetic git repo
subprocess.run(["git", "init", "-b", "master"], cwd=self.repo_dir, check=True, capture_output=True)
subprocess.run(["git", "config", "user.name", "Test User"], cwd=self.repo_dir, check=True)
subprocess.run(["git", "config", "user.email", "[email protected]"], cwd=self.repo_dir, check=True)
readme = os.path.join(self.repo_dir, "README.md")
with open(readme, "w", encoding="utf-8") as f:
f.write("# Test Repo\n")
subprocess.run(["git", "add", "README.md"], cwd=self.repo_dir, check=True, capture_output=True)
subprocess.run(["git", "commit", "-m", "initial commit"], cwd=self.repo_dir, check=True, capture_output=True)
rev_res = subprocess.run(["git", "rev-parse", "HEAD"], cwd=self.repo_dir, capture_output=True, text=True, check=True)
self.master_sha = rev_res.stdout.strip()
self.branches_dir = os.path.join(self.repo_dir, "branches")
os.makedirs(self.branches_dir, exist_ok=True)
self.lock_dir = os.path.join(self.tmp_dir, "locks")
os.makedirs(self.lock_dir, exist_ok=True)
self.journal_dir = os.path.join(self.tmp_dir, "journals")
os.makedirs(self.journal_dir, exist_ok=True)
os.environ["GITEA_BOOTSTRAP_JOURNAL_DIR"] = self.journal_dir
def tearDown(self):
os.environ.pop("GITEA_BOOTSTRAP_JOURNAL_DIR", None)
shutil.rmtree(self.tmp_dir, ignore_errors=True)
def test_bootstrap_success_path(self):
"""AC1/AC3/AC8: Successful bootstrap creates branch, worktree, registration, and lock proof."""
key = "test_key_success_1"
res = author_issue_bootstrap.bootstrap_author_issue_worktree(
issue_number=850,
canonical_repo_root=self.repo_dir,
# assignment_id/lease_id omitted: optional unless verified live.
expected_base_sha=self.master_sha,
idempotency_key=key,
remote="prgs",
lock_dir=self.lock_dir,
owner_session="session-test-1234",
)
self.assertTrue(res.get("success"), f"Bootstrap failed: {res}")
self.assertFalse(res.get("replayed"))
self.assertEqual(res.get("issue_number"), 850)
self.assertEqual(res.get("base_sha"), self.master_sha)
self.assertIn("branches/fix-issue-850-native-mcp-bootstrap", res.get("worktree_path"))
# Verify worktree directory exists and is registered
worktree_path = res["worktree_path"]
self.assertTrue(os.path.isdir(worktree_path))
wt_list = subprocess.run(["git", "-C", self.repo_dir, "worktree", "list"], capture_output=True, text=True, check=True)
self.assertIn(worktree_path, wt_list.stdout)
# Verify phase journal written
journal = author_issue_bootstrap.load_phase_journal(key, journal_dir=self.lock_dir)
self.assertIsNotNone(journal)
self.assertTrue(journal.get("completed"))
self.assertEqual(journal.get("current_phase"), author_issue_bootstrap.PHASE_7_TRANSITION_COMPLETED)
def test_idempotent_replay(self):
"""Item 2: Replaying with identical key returns cached transition without duplicate creation."""
key = "test_key_idempotent_1"
res1 = author_issue_bootstrap.bootstrap_author_issue_worktree(
issue_number=850,
canonical_repo_root=self.repo_dir,
idempotency_key=key,
lock_dir=self.lock_dir,
owner_session="session-test-1234",
)
self.assertTrue(res1["success"], f"res1 failed: {res1}")
self.assertFalse(res1.get("replayed"))
# Second call
res2 = author_issue_bootstrap.bootstrap_author_issue_worktree(
issue_number=850,
canonical_repo_root=self.repo_dir,
idempotency_key=key,
lock_dir=self.lock_dir,
owner_session="session-test-1234",
)
self.assertTrue(res2["success"], f"res2 failed: {res2}")
self.assertTrue(res2.get("replayed"))
self.assertEqual(res1["worktree_path"], res2["worktree_path"])
def test_stale_concurrency_pin_refusal(self):
"""Item 3: Mismatched expected base SHA fails closed without silent rebasing."""
stale_sha = "0000000000000000000000000000000000000000"
res = author_issue_bootstrap.bootstrap_author_issue_worktree(
issue_number=850,
canonical_repo_root=self.repo_dir,
expected_base_sha=stale_sha,
lock_dir=self.lock_dir,
owner_session="session-test-1234",
)
self.assertFalse(res["success"])
self.assertEqual(res.get("reason_code"), "stale_concurrency_pin")
self.assertIn("exact_next_action", res)
def test_path_outside_branches_root_refusal(self):
"""Item 6: Worktree path outside branches/ root is refused."""
outside_path = os.path.join(self.tmp_dir, "outside_worktree")
res = author_issue_bootstrap.bootstrap_author_issue_worktree(
issue_number=850,
canonical_repo_root=self.repo_dir,
worktree_path=outside_path,
lock_dir=self.lock_dir,
owner_session="session-test-1234",
)
self.assertFalse(res["success"])
self.assertEqual(res.get("reason_code"), "path_outside_canonical_branches_root")
def test_preexisting_dirty_worktree_preservation(self):
"""Item 6: Preexisting dirty worktree fails closed and is NOT modified or cleaned."""
branch = "fix/issue-850-dirty-test"
wt_path = os.path.join(self.branches_dir, "fix-issue-850-dirty-test")
subprocess.run(["git", "-C", self.repo_dir, "worktree", "add", "-b", branch, wt_path], check=True, capture_output=True)
# Create dirty untracked file
dirty_file = os.path.join(wt_path, "dirty.txt")
with open(dirty_file, "w") as f:
f.write("dirty edits\n")
res = author_issue_bootstrap.bootstrap_author_issue_worktree(
issue_number=850,
canonical_repo_root=self.repo_dir,
branch_name=branch,
worktree_path=wt_path,
lock_dir=self.lock_dir,
owner_session="session-test-1234",
)
self.assertFalse(res["success"])
self.assertEqual(res.get("reason_code"), "preexisting_dirty_worktree")
# Prove dirty file is preserved byte-for-byte
self.assertTrue(os.path.exists(dirty_file))
with open(dirty_file, "r") as f:
self.assertEqual(f.read(), "dirty edits\n")
def test_compensating_recovery_on_failed_phase(self):
"""AC4/Item 4: Failure during transition rolls back ONLY newly created artifacts."""
key = "test_key_recovery_1"
# Simulate partial progress in journal
journal = {
"idempotency_key": key,
"issue_number": 850,
"branch_name": "fix/issue-850-recovery-test",
"worktree_path": os.path.join(self.branches_dir, "fix-issue-850-recovery-test"),
"artifacts_created": {
"branch_created": True,
"worktree_dir_created": True,
"worktree_registered": True,
"lock_created": False,
},
"failure_reason": "simulated lock failure",
"current_phase": author_issue_bootstrap.PHASE_5_REGISTRATION_VERIFIED,
"completed": False,
}
# Create the branch and worktree manually to simulate partial state
subprocess.run(["git", "-C", self.repo_dir, "branch", journal["branch_name"]], check=True, capture_output=True)
subprocess.run(["git", "-C", self.repo_dir, "worktree", "add", journal["worktree_path"], journal["branch_name"]], check=True, capture_output=True)
# Run compensating recovery
rec = author_issue_bootstrap.run_compensating_recovery(journal, self.repo_dir)
self.assertTrue(rec["executed"])
self.assertIn(f"worktree_path:{journal['worktree_path']}", rec["rolled_back"])
self.assertIn(f"branch:{journal['branch_name']}", rec["rolled_back"])
# Prove worktree directory and branch were rolled back
self.assertFalse(os.path.exists(journal["worktree_path"]))
branch_check = subprocess.run(["git", "-C", self.repo_dir, "rev-parse", "--verify", journal["branch_name"]], capture_output=True, text=True, check=False)
self.assertNotEqual(branch_check.returncode, 0)
def test_cross_process_concurrency(self):
"""Review #525 Finding 1: Genuine cross-process concurrency locking prevents corruption."""
import concurrent.futures
key = "test_concurrent_key_850"
args = (self.repo_dir, 850, key, self.lock_dir, self.journal_dir, self.master_sha)
with concurrent.futures.ProcessPoolExecutor(max_workers=2) as executor:
fut1 = executor.submit(_concurrent_bootstrap_worker, args)
fut2 = executor.submit(_concurrent_bootstrap_worker, args)
res1 = fut1.result(timeout=10)
res2 = fut2.result(timeout=10)
self.assertTrue(res1["success"], f"res1 failed: {res1}")
self.assertTrue(res2["success"], f"res2 failed: {res2}")
# One process performs creation, the other process receives idempotent replay
replayed_count = sum(1 for r in (res1, res2) if r.get("replayed"))
created_count = sum(1 for r in (res1, res2) if not r.get("replayed"))
self.assertEqual(replayed_count, 1)
self.assertEqual(created_count, 1)
self.assertEqual(res1["worktree_path"], res2["worktree_path"])
def test_interrupted_replay_preserves_artifacts_created_provenance(self):
"""Review #525 Finding 2: Replaying incomplete journal preserves creation provenance monotonically."""
key = "test_key_interrupted_replay_1"
branch = "fix/issue-850-interrupted-replay"
wt_path = os.path.join(self.branches_dir, "fix-issue-850-interrupted-replay")
# Simulate Phase 2/3 completion where branch and worktree directory were created by this transition
journal = {
"idempotency_key": key,
"issue_number": 850,
"branch_name": branch,
"worktree_path": wt_path,
"active_identity": "jcwalker3",
"active_profile": "prgs-author",
"remote": "prgs",
"org": "Scaled-Tech-Consulting",
"repo": "Gitea-Tools",
"phases": {
author_issue_bootstrap.PHASE_1_REQUEST_ACCEPTED: {"status": "completed"},
author_issue_bootstrap.PHASE_2_BRANCH_CONFIRMED: {"status": "completed", "created": True},
},
"artifacts_created": {
"branch_created": True,
"worktree_dir_created": True,
"worktree_registered": True,
"lock_created": False,
},
"current_phase": author_issue_bootstrap.PHASE_3_PATH_RESERVED,
"completed": False,
}
# Pre-create the branch and worktree on disk to simulate partial state after crash
subprocess.run(["git", "-C", self.repo_dir, "branch", branch, self.master_sha], check=True, capture_output=True)
subprocess.run(["git", "-C", self.repo_dir, "worktree", "add", wt_path, branch], check=True, capture_output=True)
author_issue_bootstrap.save_phase_journal(journal, journal_dir=self.lock_dir)
# Now resume/replay the transition but simulate lock binding failure during Phase 6
with mock.patch("issue_lock_store.bind_session_lock", side_effect=RuntimeError("Lock failure test")):
res = author_issue_bootstrap.bootstrap_author_issue_worktree(
issue_number=850,
canonical_repo_root=self.repo_dir,
branch_name=branch,
worktree_path=wt_path,
idempotency_key=key,
lock_dir=self.lock_dir,
owner_session="session-test-1234",
)
self.assertFalse(res["success"])
self.assertEqual(res.get("reason_code"), "issue_lock_acquisition_failed")
# Verify that compensating recovery correctly deleted transition-created branch & worktree
# because creation provenance was preserved across replay (NOT downgraded to False!)
self.assertFalse(os.path.exists(wt_path))
branch_check = subprocess.run(["git", "-C", self.repo_dir, "rev-parse", "--verify", branch], capture_output=True, text=True, check=False)
self.assertNotEqual(branch_check.returncode, 0)
def test_transition_created_only_compensation(self):
"""Review #525 Finding 4: Preexisting branch is NOT deleted by compensation when only worktree was transition-created."""
key = "test_key_preexisting_branch_compensation"
preexisting_branch = "fix/issue-850-preexisting"
wt_path = os.path.join(self.branches_dir, "fix-issue-850-preexisting")
# Create branch BEFORE bootstrap (preexisting branch)
subprocess.run(["git", "-C", self.repo_dir, "branch", preexisting_branch, self.master_sha], check=True, capture_output=True)
# Call bootstrap with simulated failure during Phase 6 (lock binding)
with mock.patch("issue_lock_store.bind_session_lock", side_effect=RuntimeError("Simulated lock failure")):
res = author_issue_bootstrap.bootstrap_author_issue_worktree(
issue_number=850,
canonical_repo_root=self.repo_dir,
branch_name=preexisting_branch,
worktree_path=wt_path,
idempotency_key=key,
lock_dir=self.lock_dir,
owner_session="session-test-1234",
)
self.assertFalse(res["success"])
# Worktree dir was created by transition -> removed by compensation
self.assertFalse(os.path.exists(wt_path))
# Preexisting branch was NOT created by transition -> MUST BE PRESERVED!
branch_check = subprocess.run(["git", "-C", self.repo_dir, "rev-parse", "--verify", preexisting_branch], capture_output=True, text=True, check=False)
self.assertEqual(branch_check.returncode, 0, "Preexisting branch was deleted by mistake!")
def test_incompatible_idempotency_replay_refusal(self):
"""Review #525 Finding 4: Replaying key with incompatible parameters returns refusal."""
key = "test_key_incompatible_replay"
res1 = author_issue_bootstrap.bootstrap_author_issue_worktree(
issue_number=850,
canonical_repo_root=self.repo_dir,
branch_name="fix/issue-850-param-a",
idempotency_key=key,
lock_dir=self.lock_dir,
owner_session="session-test-1234",
)
self.assertTrue(res1["success"])
# Second call with different branch_name
res2 = author_issue_bootstrap.bootstrap_author_issue_worktree(
issue_number=850,
canonical_repo_root=self.repo_dir,
branch_name="fix/issue-850-param-b",
idempotency_key=key,
lock_dir=self.lock_dir,
owner_session="session-test-1234",
)
self.assertFalse(res2["success"])
self.assertEqual(res2.get("reason_code"), "incompatible_idempotency_replay")
def test_exact_next_action_satisfiable_via_mcp(self):
"""Review #525 Finding 4: exact_next_action provides satisfiable MCP actions, not shell commands."""
key = "test_key_next_action_mcp"
stale_sha = "0000000000000000000000000000000000000000"
res = author_issue_bootstrap.bootstrap_author_issue_worktree(
issue_number=850,
canonical_repo_root=self.repo_dir,
expected_base_sha=stale_sha,
lock_dir=self.lock_dir,
owner_session="session-test-1234",
)
next_action = res.get("exact_next_action", "")
self.assertNotIn("scripts/worktree-start", next_action)
self.assertNotIn("git worktree add", next_action)
self.assertNotIn("bash", next_action.lower())
def test_missing_owner_session_refusal(self):
"""Finding D: Missing owner_session context fails closed with typed refusal and zero mutation."""
res = author_issue_bootstrap.bootstrap_author_issue_worktree(
issue_number=850,
canonical_repo_root=self.repo_dir,
owner_session=None,
lock_dir=self.lock_dir,
)
self.assertFalse(res["success"])
self.assertEqual(res.get("reason_code"), "missing_owner_session")
self.assertIn("exact_next_action", res)
def test_symlink_lock_file_refusal(self):
"""Finding C: BootstrapTransitionLock refuses to follow symlinks."""
key = "test_symlink_lock_key"
safe_key = "".join(c if c.isalnum() or c in ("-", "_", ".") else "_" for c in key)
lock_path = os.path.join(self.lock_dir, f"{safe_key}.lock")
target_file = os.path.join(self.tmp_dir, "fake_target")
with open(target_file, "w") as f:
f.write("target")
os.symlink(target_file, lock_path)
with self.assertRaises(RuntimeError) as ctx:
with author_issue_bootstrap.BootstrapTransitionLock(key, journal_dir=self.lock_dir):
pass
self.assertIn("symlink", str(ctx.exception).lower())
def test_lock_directory_escape_refusal(self):
"""Finding C: BootstrapTransitionLock refuses keys that escape lock directory."""
with mock.patch("os.path.abspath", return_value="/tmp/outside/evil_key.lock"):
with self.assertRaises(RuntimeError) as ctx:
author_issue_bootstrap.BootstrapTransitionLock("key", journal_dir=self.lock_dir)
self.assertIn("escapes", str(ctx.exception).lower())
def test_missing_active_identity_refusal(self):
"""F-5: Missing active_identity parameter fails closed."""
res = author_issue_bootstrap.bootstrap_author_issue_worktree(
issue_number=850,
canonical_repo_root=self.repo_dir,
owner_session="session-test-1234",
active_identity=None,
active_profile="prgs-author",
lock_dir=self.lock_dir,
)
self.assertFalse(res["success"])
self.assertEqual(res.get("reason_code"), "missing_active_identity")
def test_missing_active_profile_refusal(self):
"""F-5: Missing active_profile parameter fails closed."""
res = author_issue_bootstrap.bootstrap_author_issue_worktree(
issue_number=850,
canonical_repo_root=self.repo_dir,
owner_session="session-test-1234",
active_identity="jcwalker3",
active_profile=None,
lock_dir=self.lock_dir,
)
self.assertFalse(res["success"])
self.assertEqual(res.get("reason_code"), "missing_active_profile")
def test_dirty_worktree_preserved_during_recovery(self):
"""F-4: Compensating recovery does not delete dirty worktree."""
branch = "fix/issue-850-rec-dirty"
wt_path = os.path.join(self.branches_dir, "fix-issue-850-rec-dirty")
subprocess.run(["git", "-C", self.repo_dir, "worktree", "add", "-b", branch, wt_path], check=True, capture_output=True)
dirty_file = os.path.join(wt_path, "dirty.txt")
with open(dirty_file, "w") as f:
f.write("uncommitted work")
journal = {
"idempotency_key": "test_dirty_rec",
"issue_number": 850,
"branch_name": branch,
"worktree_path": wt_path,
"artifacts_created": {
"worktree_dir_created": True,
"worktree_registered": True,
},
"failure_reason": "test dirty recovery",
}
rec = author_issue_bootstrap.run_compensating_recovery(journal, self.repo_dir, journal_dir=self.lock_dir)
self.assertTrue(os.path.exists(wt_path))
self.assertIn(f"worktree_path_preserved_dirty:{wt_path}", rec["rolled_back"])
def test_branch_with_commits_preserved_during_recovery(self):
"""F-4: Compensating recovery does not delete branch with author commits."""
branch = "fix/issue-850-rec-commits"
subprocess.run(["git", "-C", self.repo_dir, "branch", branch, self.master_sha], check=True, capture_output=True)
# Add a commit on the branch
wt_path = os.path.join(self.branches_dir, "fix-issue-850-rec-commits")
subprocess.run(["git", "-C", self.repo_dir, "worktree", "add", wt_path, branch], check=True, capture_output=True)
cfile = os.path.join(wt_path, "commit.txt")
with open(cfile, "w") as f:
f.write("author commit")
subprocess.run(["git", "-C", wt_path, "add", "commit.txt"], check=True, capture_output=True)
subprocess.run(["git", "-C", wt_path, "commit", "-m", "author commit"], check=True, capture_output=True)
subprocess.run(["git", "-C", self.repo_dir, "worktree", "remove", "--force", wt_path], check=True, capture_output=True)
journal = {
"idempotency_key": "test_commits_rec",
"issue_number": 850,
"branch_name": branch,
"resolved_base_sha": self.master_sha,
"artifacts_created": {
"branch_created": True,
},
"failure_reason": "test commit branch recovery",
}
rec = author_issue_bootstrap.run_compensating_recovery(journal, self.repo_dir, journal_dir=self.lock_dir)
branch_check = subprocess.run(["git", "-C", self.repo_dir, "rev-parse", "--verify", branch], capture_output=True, text=True, check=False)
self.assertEqual(branch_check.returncode, 0, "Branch with commits was deleted!")
self.assertIn(f"branch_preserved_commits:{branch}", rec["rolled_back"])
def test_task_capability_map_integration(self):
"""Verify task_capability_map has bootstrap_author_issue_worktree configured correctly."""
self.assertEqual(task_capability_map.required_role("bootstrap_author_issue_worktree"), "author")
self.assertEqual(task_capability_map.required_permission("bootstrap_author_issue_worktree"), "gitea.branch.create")
self.assertTrue(task_capability_map.preflight_task_matches("work_issue", "bootstrap_author_issue_worktree"))
self.assertTrue(task_capability_map.preflight_task_matches("bootstrap_author_issue_worktree", "lock_issue"))
def test_unverified_assignment_lease_ids_fail_closed(self):
"""Review #531 Finding 4: fabricated assignment/lease IDs are refused."""
res = author_issue_bootstrap.bootstrap_author_issue_worktree(
issue_number=850,
canonical_repo_root=self.repo_dir,
assignment_id="asn-fabricated",
lease_id="lease-fabricated",
expected_base_sha=self.master_sha,
lock_dir=self.lock_dir,
owner_session="session-test-1234",
)
self.assertFalse(res["success"])
self.assertIn(
res.get("reason_code"),
{
"unknown_lease_id",
"assignment_lease_lookup_failed",
"incomplete_assignment_lease_ids",
},
)
def test_partial_assignment_lease_ids_fail_closed(self):
"""Review #531 Finding 4: one of assignment_id/lease_id alone is incomplete."""
res = author_issue_bootstrap.bootstrap_author_issue_worktree(
issue_number=850,
canonical_repo_root=self.repo_dir,
assignment_id="asn-only",
lock_dir=self.lock_dir,
owner_session="session-test-1234",
)
self.assertFalse(res["success"])
self.assertEqual(res.get("reason_code"), "incomplete_assignment_lease_ids")
def test_stale_diverged_branch_is_not_accepted_via_merge_base(self):
"""Review #531 Finding 3: any common ancestor is not enough; require master ⊆ branch."""
branch = "fix/issue-850-stale-divergent"
# Create branch from current master, then advance master so branch lacks tip.
subprocess.run(
["git", "-C", self.repo_dir, "branch", branch, self.master_sha],
check=True,
capture_output=True,
)
# Make a new commit on master (orphan path so branch does not contain it).
marker = os.path.join(self.repo_dir, "master-advance.txt")
with open(marker, "w") as f:
f.write("advance master\n")
subprocess.run(["git", "-C", self.repo_dir, "add", "master-advance.txt"], check=True, capture_output=True)
subprocess.run(
["git", "-C", self.repo_dir, "commit", "-m", "advance master past branch"],
check=True,
capture_output=True,
)
new_master = subprocess.run(
["git", "-C", self.repo_dir, "rev-parse", "HEAD"],
capture_output=True,
text=True,
check=True,
).stdout.strip()
res = author_issue_bootstrap.bootstrap_author_issue_worktree(
issue_number=850,
canonical_repo_root=self.repo_dir,
branch_name=branch,
expected_base_sha=new_master,
lock_dir=self.lock_dir,
owner_session="session-test-1234",
)
self.assertFalse(res["success"])
self.assertEqual(res.get("reason_code"), "incompatible_existing_branch")
def test_compensating_recovery_attempts_lease_release(self):
"""Review #531 Finding 5: recovery invokes lease release when lease_id is present."""
from unittest import mock
journal = {
"idempotency_key": "test_lease_rec",
"issue_number": 850,
"owner_session": "session-test-1234",
"lease_id": "lease-abc",
"branch_name": "fix/issue-850-lease-rec",
"artifacts_created": {"lock_created": True},
"failure_reason": "simulated",
"completed": False,
}
with mock.patch.object(
author_issue_bootstrap.lease_lifecycle,
"release_lease",
return_value={"success": True},
) as rel, mock.patch.object(
author_issue_bootstrap.control_plane_db,
"ControlPlaneDB",
return_value=mock.Mock(),
):
rec = author_issue_bootstrap.run_compensating_recovery(
journal, self.repo_dir, journal_dir=self.lock_dir
)
self.assertTrue(rec["executed"])
rel.assert_called_once()
self.assertIn("lease:lease-abc", rec["rolled_back"])
class TestCanonicalRootNoStringSplit(unittest.TestCase):
def test_fallback_uses_commonpath_not_substring_split(self):
"""Review #531 Finding 2: no norm.split('/branches/') fallback."""
import inspect
import author_mutation_worktree as amw
src = inspect.getsource(amw.resolve_canonical_repo_root)
self.assertNotIn('split("/branches/")', src)
self.assertNotIn("split('/branches/')", src)
# Fallback recovers repo root from a nested branches worktree path.
with tempfile.TemporaryDirectory() as tmp:
repo = os.path.join(tmp, "repo")
wt = os.path.join(repo, "branches", "fix-issue-850-x")
os.makedirs(wt)
# git unavailable path: pass missing workspace so fallback is used.
root = amw.resolve_canonical_repo_root("/missing/path", wt)
self.assertEqual(root, os.path.realpath(repo))
if __name__ == "__main__":
unittest.main()
-22
View File
@@ -28,21 +28,6 @@ class TestPathUnderBranches(unittest.TestCase):
amw.is_path_under_branches("/repo/other-checkout", self.ROOT)
)
def test_unrelated_branches_dir_fails(self):
self.assertFalse(
amw.is_path_under_branches("/tmp/branches/evil", self.ROOT)
)
def test_prefix_confusion_fails(self):
self.assertFalse(
amw.is_path_under_branches(f"{self.ROOT}/branches-other/foo", self.ROOT)
)
def test_traversal_fails(self):
self.assertFalse(
amw.is_path_under_branches(f"{self.ROOT}/branches/../evil", self.ROOT)
)
class TestAssessAuthorMutationWorktree(unittest.TestCase):
ROOT = "/repo/Gitea-Tools"
@@ -86,13 +71,6 @@ class TestAssessAuthorMutationWorktree(unittest.TestCase):
self.assertTrue(result["proven"])
self.assertFalse(result["block"])
def test_path_shaped_branches_ancestry_without_isdir(self):
"""Review #551: commonpath recovery must not require on-disk isdir."""
fake_wt = "/repo/Gitea-Tools/branches/issue-274"
root = amw.resolve_canonical_repo_root(fake_wt, fake_wt)
self.assertEqual(root, "/repo/Gitea-Tools")
self.assertNotIn('split("/branches/")', open(amw.__file__).read())
class TestPreflightIntegration(unittest.TestCase):
def test_verify_preflight_blocks_control_checkout_with_test_porcelain(self):
-724
View File
@@ -1266,730 +1266,6 @@ class TestSecondRemediationIntegration(unittest.TestCase):
self.assertIn("delete_acknowledged", delete_actions[0])
self.assertTrue(delete_actions[0].get("verified_absent"))
def test_issue_851_worktree_removed_when_remote_blocked_only_by_worktree_binding(self):
"""#851: remote blocked by worktree_binding must not skip safe worktree removal.
Lifecycle: remove clean owned worktree reassess ownership delete
remote only if independently safe. Unrelated entries stay untouched.
"""
from mcp_server import gitea_reconcile_merged_cleanups
target_branch = "fix/issue-844-exclude-epic-containers"
foreign_branch = "fix/issue-999-unrelated-active"
worktree_path = "/tmp/branches/fix-issue-844-exclude-epic-containers"
ownership_calls = []
remove_calls = []
delete_api_calls = []
def fake_collect(**kwargs):
ownership_calls.append(dict(kwargs))
# Ownership is reassessed *after* independent worktree removal (#851).
# Target worktree is already gone → no worktree_binding remains.
# Foreign branch keeps an active author lease → remote delete blocked.
if kwargs.get("branch") == foreign_branch:
# Match session-bound org/repo + host used by the tool resolve path.
return {
"records": [
{
"category": guard.OWNERSHIP_CATEGORY_AUTHOR_LEASE,
"status": "active",
"remote": kwargs.get("remote") or "prgs",
"host": kwargs.get("host") or "gitea.example.com",
"org": kwargs.get("org") or "Scaled-Tech-Consulting",
"repo": kwargs.get("repo") or "Gitea-Tools",
"branch": foreign_branch,
"reclaim_allowed": False,
}
],
"inventory_error": False,
}
return {"records": [], "inventory_error": False}
def fake_remove(project_root, branch, worktree_path=None):
remove_calls.append(
{"branch": branch, "worktree_path": worktree_path}
)
return {
"success": True,
"performed": True,
"message": f"removed worktree {worktree_path}",
"worktree_path": worktree_path,
}
def fake_probe(h, o, r, auth, br):
return guard.classify_branch_readback_http_status(
404, not_found_scope=guard.NOT_FOUND_SCOPE_BRANCH
)
def fake_api(method, url, auth, **kwargs):
if method == "DELETE":
delete_api_calls.append(url)
return {}
report = {
"entries": [
{
"pr_number": 848,
"head_branch": target_branch,
"remote_branch": {"safe_to_delete_remote": True},
"local_worktree": {
"safe_to_remove_worktree": True,
"worktree_path": worktree_path,
},
},
{
"pr_number": 999,
"head_branch": foreign_branch,
"remote_branch": {"safe_to_delete_remote": True},
"local_worktree": {
"safe_to_remove_worktree": False,
"worktree_path": None,
},
},
],
"reviewer_scratch_entries": [],
}
patch(
"mcp_server.get_profile",
return_value={
"profile_name": "prgs-reconciler",
"role": "reconciler",
"allowed_operations": [
"gitea.read",
"gitea.branch.delete",
"gitea.pr.close",
],
"forbidden_operations": [],
},
).start()
patch("mcp_server.api_get_all", return_value=[]).start()
patch(
"mcp_server.merged_cleanup_reconcile.build_reconciliation_report",
return_value=report,
).start()
patch(
"mcp_server.merged_cleanup_reconcile.discover_reviewer_scratch_worktrees",
return_value=[],
).start()
patch(
"mcp_server.audit_reconciliation_mode.check_cleanup_execution_allowed",
return_value=(True, []),
).start()
patch("mcp_server.verify_preflight_purity", return_value=None).start()
patch(
"mcp_server._collect_branch_ownership_records",
side_effect=fake_collect,
).start()
patch("mcp_server._probe_remote_branch", side_effect=fake_probe).start()
patch(
"mcp_server.merged_cleanup_reconcile.remove_local_worktree",
side_effect=fake_remove,
).start()
self.mock_api.side_effect = fake_api
res = gitea_reconcile_merged_cleanups(
dry_run=False,
execute_confirmed=True,
remote="prgs",
)
self.assertTrue(res.get("performed") or res.get("executed"))
actions = res.get("actions") or []
remove_actions = [
a for a in actions if a.get("action") == "remove_local_worktree"
]
self.assertEqual(len(remove_actions), 1, actions)
self.assertTrue(remove_actions[0].get("success"))
self.assertEqual(remove_calls[0]["branch"], target_branch)
self.assertEqual(remove_calls[0]["worktree_path"], worktree_path)
# Target remote delete succeeds after worktree removal + reassessment.
target_deletes = [
a
for a in actions
if a.get("action") == "delete_remote_branch"
and a.get("branch") == target_branch
]
self.assertEqual(len(target_deletes), 1, actions)
self.assertTrue(target_deletes[0].get("success"))
self.assertTrue(target_deletes[0].get("after_worktree_removal"))
self.assertTrue(target_deletes[0].get("ownership_reassessed"))
self.assertTrue(target_deletes[0].get("verified_absent"))
# Foreign branch remains protected (author lease) and is not deleted.
foreign_deletes = [
a
for a in actions
if a.get("action") == "delete_remote_branch"
and a.get("branch") == foreign_branch
]
self.assertEqual(len(foreign_deletes), 1, actions)
self.assertFalse(foreign_deletes[0].get("success"))
self.assertEqual(
foreign_deletes[0].get("blocker_kind"), "active_branch_ownership"
)
self.assertIn(
guard.OWNERSHIP_CATEGORY_AUTHOR_LEASE,
foreign_deletes[0].get("blocking_categories") or [],
)
# Only the target branch should hit the DELETE API.
self.assertEqual(len(delete_api_calls), 1)
# Ownership collected for target (post-removal) and foreign; worktree
# removal happened before target remote delete in the action log.
target_idx = next(
i
for i, a in enumerate(actions)
if a.get("action") == "remove_local_worktree"
)
delete_idx = next(
i
for i, a in enumerate(actions)
if a.get("action") == "delete_remote_branch"
and a.get("branch") == target_branch
and a.get("success")
)
self.assertLess(target_idx, delete_idx)
def test_issue_851_dirty_worktree_not_removed_and_remote_stays_protected(self):
"""#851: dirty/foreign worktrees remain protected; no unsafe cleanup."""
from mcp_server import gitea_reconcile_merged_cleanups
branch = "fix/issue-851-dirty"
remove_calls = []
def fake_collect(**kwargs):
return {
"records": [
{
"category": guard.OWNERSHIP_CATEGORY_WORKTREE_BINDING,
"status": "active",
"remote": kwargs.get("remote") or "prgs",
"host": kwargs.get("host") or "gitea.example.com",
"org": kwargs.get("org") or "Scaled-Tech-Consulting",
"repo": kwargs.get("repo") or "Gitea-Tools",
"branch": branch,
"reclaim_allowed": False,
}
],
"inventory_error": False,
}
report = {
"entries": [
{
"pr_number": 851,
"head_branch": branch,
"remote_branch": {"safe_to_delete_remote": True},
"local_worktree": {
"safe_to_remove_worktree": False,
"worktree_path": "/tmp/dirty-wt",
},
}
],
"reviewer_scratch_entries": [],
}
patch(
"mcp_server.get_profile",
return_value={
"profile_name": "prgs-reconciler",
"role": "reconciler",
"allowed_operations": [
"gitea.read",
"gitea.branch.delete",
],
"forbidden_operations": [],
},
).start()
patch("mcp_server.api_get_all", return_value=[]).start()
patch(
"mcp_server.merged_cleanup_reconcile.build_reconciliation_report",
return_value=report,
).start()
patch(
"mcp_server.merged_cleanup_reconcile.discover_reviewer_scratch_worktrees",
return_value=[],
).start()
patch(
"mcp_server.audit_reconciliation_mode.check_cleanup_execution_allowed",
return_value=(True, []),
).start()
patch("mcp_server.verify_preflight_purity", return_value=None).start()
patch(
"mcp_server._collect_branch_ownership_records",
side_effect=fake_collect,
).start()
patch(
"mcp_server.merged_cleanup_reconcile.remove_local_worktree",
side_effect=lambda *a, **k: remove_calls.append(k) or {
"success": True,
"performed": True,
},
).start()
self.mock_api.side_effect = lambda *a, **k: {}
res = gitea_reconcile_merged_cleanups(
dry_run=False,
execute_confirmed=True,
remote="prgs",
)
actions = res.get("actions") or []
self.assertEqual(remove_calls, [])
self.assertFalse(
any(a.get("action") == "remove_local_worktree" for a in actions)
)
deletes = [
a for a in actions if a.get("action") == "delete_remote_branch"
]
self.assertEqual(len(deletes), 1)
self.assertFalse(deletes[0].get("success"))
self.assertEqual(deletes[0].get("blocker_kind"), "active_branch_ownership")
self.assertIn(
guard.OWNERSHIP_CATEGORY_WORKTREE_BINDING,
deletes[0].get("blocking_categories") or [],
)
def test_issue_851_idempotent_resume_when_worktree_already_absent(self):
"""#851: partial failures remain resumable and idempotent."""
from mcp_server import gitea_reconcile_merged_cleanups
branch = "fix/issue-851-resume"
ownership_calls = []
def fake_collect(**kwargs):
ownership_calls.append(kwargs)
return {"records": [], "inventory_error": False}
def fake_remove(project_root, branch, worktree_path=None):
return {
"success": False,
"performed": False,
"message": f"worktree not found: {worktree_path}",
}
def fake_probe(h, o, r, auth, br):
return guard.classify_branch_readback_http_status(
404, not_found_scope=guard.NOT_FOUND_SCOPE_BRANCH
)
report = {
"entries": [
{
"pr_number": 851,
"head_branch": branch,
"remote_branch": {"safe_to_delete_remote": True},
"local_worktree": {
"safe_to_remove_worktree": True,
"worktree_path": "/tmp/already-gone",
},
}
],
"reviewer_scratch_entries": [],
}
patch(
"mcp_server.get_profile",
return_value={
"profile_name": "prgs-reconciler",
"role": "reconciler",
"allowed_operations": [
"gitea.read",
"gitea.branch.delete",
],
"forbidden_operations": [],
},
).start()
patch("mcp_server.api_get_all", return_value=[]).start()
patch(
"mcp_server.merged_cleanup_reconcile.build_reconciliation_report",
return_value=report,
).start()
patch(
"mcp_server.merged_cleanup_reconcile.discover_reviewer_scratch_worktrees",
return_value=[],
).start()
patch(
"mcp_server.audit_reconciliation_mode.check_cleanup_execution_allowed",
return_value=(True, []),
).start()
patch("mcp_server.verify_preflight_purity", return_value=None).start()
patch(
"mcp_server._collect_branch_ownership_records",
side_effect=fake_collect,
).start()
patch("mcp_server._probe_remote_branch", side_effect=fake_probe).start()
patch(
"mcp_server.merged_cleanup_reconcile.remove_local_worktree",
side_effect=fake_remove,
).start()
self.mock_api.side_effect = lambda *a, **k: {}
res = gitea_reconcile_merged_cleanups(
dry_run=False,
execute_confirmed=True,
remote="prgs",
)
actions = res.get("actions") or []
removes = [a for a in actions if a.get("action") == "remove_local_worktree"]
deletes = [a for a in actions if a.get("action") == "delete_remote_branch"]
self.assertEqual(len(removes), 1)
self.assertFalse(removes[0].get("success"))
self.assertEqual(len(deletes), 1)
self.assertTrue(deletes[0].get("success"))
self.assertTrue(deletes[0].get("after_worktree_removal"))
self.assertTrue(ownership_calls)
class TestIssue855ExactPrSelector(unittest.TestCase):
"""#855: exact pr_number pin for reconcile_merged_cleanups (#851 lifecycle)."""
def setUp(self):
self._remotes = patch.dict(
mcp_server.REMOTES,
{
"prgs": {
"host": "gitea.example.com",
"org": "Scaled-Tech-Consulting",
"repo": "Gitea-Tools",
}
},
)
self._remotes.start()
patch("gitea_audit.audit_enabled", return_value=False).start()
self.mock_api = patch("mcp_server.api_request").start()
self.mock_all = patch("mcp_server.api_get_all", return_value=[]).start()
patch("mcp_server.get_auth_header", return_value=FAKE_AUTH).start()
patch(
"mcp_server.merged_cleanup_reconcile.is_head_ancestor_of_ref",
return_value=True,
).start()
patch(
"mcp_server.get_profile",
return_value=dict(RECONCILER_WITH_DELETE),
).start()
patch(
"mcp_server._profile_operation_gate",
return_value=[],
).start()
patch(
"mcp_server._collect_branch_ownership_records",
return_value={"records": [], "inventory_error": False},
).start()
patch(
"mcp_server.merged_cleanup_reconcile.discover_reviewer_scratch_worktrees",
return_value=[],
).start()
patch("mcp_server.verify_preflight_purity", return_value=None).start()
patch(
"mcp_server.audit_reconciliation_mode.check_cleanup_execution_allowed",
return_value=(True, []),
).start()
def tearDown(self):
patch.stopall()
def _merged_pr(self, number, branch, sha="c" * 40):
return {
"number": number,
"title": f"PR {number}",
"body": f"Closes #{number - 4}",
"merged": True,
"merged_at": "2026-07-23T12:00:00Z",
"merge_commit_sha": "f" * 40,
"state": "closed",
"head": {"ref": branch, "sha": sha},
"base": {"ref": "master"},
}
def test_exact_pr_848_ignores_newer_852_in_batch_queue(self):
"""pr_number=848 selects only #848 even when #852 is newer/first."""
from mcp_server import gitea_reconcile_merged_cleanups
pr_848 = self._merged_pr(
848, "fix/issue-844-exclude-epic-containers", sha="c3f282ba" + "0" * 32
)
# Closed list would rank #852 first in batch mode; exact pin must ignore it.
closed_batch = [
self._merged_pr(852, "fix/issue-851-cleanup-worktree-before-remote-delete"),
pr_848,
self._merged_pr(849, "fix/issue-849-other"),
self._merged_pr(846, "fix/issue-846-other"),
self._merged_pr(845, "fix/issue-845-other"),
]
batch_fetch_calls = []
def fake_api(method, url, *args, **kwargs):
if method == "GET" and url.rstrip("/").endswith("/pulls/848"):
return dict(pr_848)
if method == "GET" and "/pulls/" in url:
raise AssertionError(f"unexpected PR fetch: {url}")
if method == "GET" and "/branches/" in url:
return {"name": "present"}
return {}
def fake_all(url, auth, limit=None):
batch_fetch_calls.append((url, limit))
if "state=open" in url:
return []
if "state=closed" in url:
# Exact mode must not use the closed batch list.
raise AssertionError(
"exact pr_number mode must not page closed PRs: " + url
)
return []
self.mock_api.side_effect = fake_api
self.mock_all.side_effect = fake_all
patch(
"mcp_server._remote_branch_exists",
return_value=True,
).start()
patch(
"mcp_server.merged_cleanup_reconcile.build_reconciliation_report",
side_effect=lambda **kwargs: {
"entries": [
{
"pr_number": int(pr["number"]),
"head_branch": (pr.get("head") or {}).get("ref"),
"issue_number": 844,
"remote_branch": {
"safe_to_delete_remote": True,
"head_branch": (pr.get("head") or {}).get("ref"),
},
"local_worktree": {
"safe_to_remove_worktree": True,
"worktree_path": (
"/tmp/branches/fix-issue-844-exclude-epic-containers"
),
},
"planned_execution_order": (
mcp_server.merged_cleanup_reconcile.plan_cleanup_execution_order(
remote_assessment={"safe_to_delete_remote": True},
local_assessment={"safe_to_remove_worktree": True},
)
),
}
for pr in kwargs.get("closed_prs") or []
if pr.get("merged_at") or pr.get("merged")
],
"reviewer_scratch_entries": [],
"merged_pr_count": len(kwargs.get("closed_prs") or []),
},
).start()
res = gitea_reconcile_merged_cleanups(
dry_run=True,
pr_number=848,
remote="prgs",
org="Scaled-Tech-Consulting",
repo="Gitea-Tools",
)
self.assertTrue(res.get("success"))
self.assertFalse(res.get("performed"))
self.assertEqual(res.get("selection_mode"), "exact_pr")
self.assertEqual(res.get("selected_pr_number"), 848)
entries = res.get("entries") or []
self.assertEqual(len(entries), 1, entries)
self.assertEqual(entries[0].get("pr_number"), 848)
self.assertEqual(
entries[0].get("head_branch"),
"fix/issue-844-exclude-epic-containers",
)
# No other PR appears in plan.
self.assertEqual(list((res.get("planned_execution_orders") or {}).keys()), ["848"])
plan = (res.get("planned_execution_orders") or {}).get("848") or []
actions = [s.get("action") for s in plan]
self.assertEqual(
actions,
[
"remove_local_worktree",
"reassess_branch_ownership",
"delete_remote_branch",
],
)
# Prove we never scanned the multi-PR closed batch.
self.assertFalse(any("state=closed" in (u or "") for u, _ in batch_fetch_calls))
# closed_batch fixture must remain unused (sanity).
self.assertEqual(closed_batch[0]["number"], 852)
def test_exact_pr_execute_only_mutates_selected_pr(self):
"""Execute with pr_number must never touch #845/#846/#849/#852."""
from mcp_server import gitea_reconcile_merged_cleanups
pr_848 = self._merged_pr(848, "fix/issue-844-exclude-epic-containers")
worktree_path = "/tmp/branches/fix-issue-844-exclude-epic-containers"
remove_calls = []
delete_api_calls = []
ownership_branches = []
def fake_api(method, url, *args, **kwargs):
if method == "GET" and url.rstrip("/").endswith("/pulls/848"):
return dict(pr_848)
if method == "DELETE":
delete_api_calls.append(url)
# Forbid foreign PR branch deletion by URL content.
for forbidden in ("845", "846", "849", "852"):
self.assertNotIn(forbidden, url)
return {}
def fake_remove(project_root, branch, worktree_path=None):
remove_calls.append({"branch": branch, "worktree_path": worktree_path})
return {
"success": True,
"performed": True,
"message": f"removed {worktree_path}",
"worktree_path": worktree_path,
}
def fake_collect(**kwargs):
ownership_branches.append(kwargs.get("branch"))
return {"records": [], "inventory_error": False}
def fake_probe(h, o, r, auth, br):
return guard.classify_branch_readback_http_status(
404, not_found_scope=guard.NOT_FOUND_SCOPE_BRANCH
)
self.mock_api.side_effect = fake_api
self.mock_all.side_effect = lambda url, auth, limit=None: []
patch("mcp_server._remote_branch_exists", return_value=True).start()
patch(
"mcp_server.merged_cleanup_reconcile.build_reconciliation_report",
return_value={
"entries": [
{
"pr_number": 848,
"head_branch": "fix/issue-844-exclude-epic-containers",
"remote_branch": {"safe_to_delete_remote": True},
"local_worktree": {
"safe_to_remove_worktree": True,
"worktree_path": worktree_path,
},
"planned_execution_order": [
{"action": "remove_local_worktree", "phase": 1},
{"action": "reassess_branch_ownership", "phase": 2},
{"action": "delete_remote_branch", "phase": 3},
],
}
],
"reviewer_scratch_entries": [
# Foreign scratch must be filtered before report execute loop;
# if present here it would still be a test failure if acted on.
],
"merged_pr_count": 1,
},
).start()
patch(
"mcp_server.merged_cleanup_reconcile.remove_local_worktree",
side_effect=fake_remove,
).start()
patch(
"mcp_server._collect_branch_ownership_records",
side_effect=fake_collect,
).start()
patch("mcp_server._probe_remote_branch", side_effect=fake_probe).start()
res = gitea_reconcile_merged_cleanups(
dry_run=False,
execute_confirmed=True,
pr_number=848,
remote="prgs",
org="Scaled-Tech-Consulting",
repo="Gitea-Tools",
)
self.assertTrue(res.get("performed") or res.get("executed"))
self.assertEqual(res.get("selection_mode"), "exact_pr")
self.assertEqual(res.get("selected_pr_number"), 848)
actions = res.get("actions") or []
pr_numbers_touched = {
a.get("pr_number") for a in actions if a.get("pr_number") is not None
}
self.assertTrue(pr_numbers_touched.issubset({None, 848}) or not pr_numbers_touched)
removes = [a for a in actions if a.get("action") == "remove_local_worktree"]
deletes = [a for a in actions if a.get("action") == "delete_remote_branch"]
self.assertEqual(len(removes), 1)
self.assertEqual(remove_calls[0]["branch"], "fix/issue-844-exclude-epic-containers")
self.assertEqual(len(deletes), 1)
self.assertTrue(deletes[0].get("success"))
self.assertTrue(deletes[0].get("after_worktree_removal"))
self.assertEqual(len(delete_api_calls), 1)
self.assertEqual(
ownership_branches, ["fix/issue-844-exclude-epic-containers"]
)
def test_exact_pr_unknown_fails_closed_without_mutation(self):
from mcp_server import gitea_reconcile_merged_cleanups
def fake_api(method, url, *args, **kwargs):
if method == "GET" and "/pulls/99999" in url:
raise RuntimeError("HTTP 404 Not Found")
raise AssertionError(f"unexpected API call {method} {url}")
self.mock_api.side_effect = fake_api
res = gitea_reconcile_merged_cleanups(
dry_run=True,
pr_number=99999,
remote="prgs",
)
self.assertFalse(res.get("success"))
self.assertFalse(res.get("performed"))
self.assertEqual(res.get("blocker_kind"), "pr_unresolvable")
self.assertIn("99999", " ".join(res.get("reasons") or []))
def test_exact_pr_not_merged_fails_closed(self):
from mcp_server import gitea_reconcile_merged_cleanups
def fake_api(method, url, *args, **kwargs):
if method == "GET" and url.rstrip("/").endswith("/pulls/900"):
return {
"number": 900,
"merged": False,
"merged_at": None,
"state": "open",
"head": {"ref": "feat/x", "sha": "a" * 40},
}
raise AssertionError(f"unexpected {method} {url}")
self.mock_api.side_effect = fake_api
res = gitea_reconcile_merged_cleanups(
dry_run=False,
execute_confirmed=True,
pr_number=900,
remote="prgs",
)
self.assertFalse(res.get("success"))
self.assertFalse(res.get("performed"))
self.assertEqual(res.get("blocker_kind"), "pr_not_merged")
def test_exact_pr_invalid_number_fails_closed(self):
from mcp_server import gitea_reconcile_merged_cleanups
res = gitea_reconcile_merged_cleanups(
dry_run=True,
pr_number=0,
remote="prgs",
)
self.assertFalse(res.get("success"))
self.assertEqual(res.get("blocker_kind"), "invalid_pr_number")
self.mock_api.assert_not_called()
def test_batch_mode_still_works_without_pr_number(self):
"""Unfiltered batch path remains backward compatible."""
from mcp_server import gitea_reconcile_merged_cleanups
self.mock_all.side_effect = lambda url, auth, limit=None: []
self.mock_api.side_effect = lambda *a, **k: {}
patch(
"mcp_server.merged_cleanup_reconcile.build_reconciliation_report",
return_value={
"entries": [],
"reviewer_scratch_entries": [],
"merged_pr_count": 0,
},
).start()
res = gitea_reconcile_merged_cleanups(dry_run=True, remote="prgs", limit=10)
self.assertTrue(res.get("success"))
self.assertEqual(res.get("selection_mode"), "batch")
self.assertIsNone(res.get("selected_pr_number"))
if __name__ == "__main__":
+2 -2
View File
@@ -35,7 +35,7 @@ CONFIG = {
],
"forbidden_operations": [],
"execution_profile": "full-author",
"allowed_repositories": ["Scaled-Tech-Consulting/Gitea-Tools", "Example-Org/Example-Repo"],
"allowed_repositories": ["Example-Org/Example-Repo"],
},
"reviewer-no-commit": {
"enabled": True,
@@ -50,7 +50,7 @@ CONFIG = {
"gitea.repo.commit", "gitea.pr.create", "gitea.branch.push"
],
"execution_profile": "reviewer-no-commit",
"allowed_repositories": ["Scaled-Tech-Consulting/Gitea-Tools", "Example-Org/Example-Repo"],
"allowed_repositories": ["Example-Org/Example-Repo"],
},
},
"rules": {"allow_runtime_switching": False},
+1 -14
View File
@@ -175,21 +175,8 @@ class TestLauncherSnippets(unittest.TestCase):
def test_only_safe_keys_no_secrets(self):
entry = gitea_config.launcher_entry("prgs", "/cfg/profiles.json")["gitea-tools"]
self.assertEqual(set(entry), {"command", "args", "env"})
# #978 B1: production launcher also injects trusted client identity.
required = {
"GITEA_MCP_CONFIG",
"GITEA_MCP_PROFILE",
"GITEA_CLIENT_MANAGED",
"GITEA_MCP_CLIENT",
"GITEA_MCP_CLIENT_INSTANCE",
"GITEA_MCP_INSTANCE_PROVENANCE",
}
self.assertTrue(required.issubset(set(entry["env"])), entry["env"])
self.assertEqual(set(entry["env"]), {"GITEA_MCP_CONFIG", "GITEA_MCP_PROFILE"})
self.assertEqual(entry["env"]["GITEA_MCP_PROFILE"], "prgs")
self.assertTrue(
entry["env"]["GITEA_MCP_CLIENT_INSTANCE"].startswith("inst-"),
entry["env"]["GITEA_MCP_CLIENT_INSTANCE"],
)
blob = json.dumps(entry).lower()
for word in ("token", "password", "secret"):
self.assertNotIn(word, blob)
+1 -226
View File
@@ -11,7 +11,6 @@ from datetime import timedelta
from control_plane_db import (
ControlPlaneDB,
ControlPlaneError,
InvalidWorkKindError,
LeaseRequiredError,
WORK_KINDS,
@@ -37,7 +36,7 @@ class ControlPlaneDBTest(unittest.TestCase):
rows = dict(conn.execute("SELECT key, value FROM schema_meta").fetchall())
finally:
conn.close()
self.assertEqual(rows["schema_version"], "5")
self.assertEqual(rows["schema_version"], "4")
self.assertIn("DB coordinates", rows["architecture"])
self.assertIn("bridge", rows["architecture"].lower())
@@ -821,229 +820,5 @@ class ControlPlaneDBTest(unittest.TestCase):
self.assertEqual(n, 1)
class SessionCheckpointTest(unittest.TestCase):
"""Durable MCP session checkpoint schema, redaction, and reconcile (#660)."""
def setUp(self) -> None:
self._tmp = tempfile.TemporaryDirectory()
self.db_path = os.path.join(self._tmp.name, "cp.sqlite3")
self.db = ControlPlaneDB(self.db_path)
def tearDown(self) -> None:
self._tmp.cleanup()
def _write(self, **overrides):
base = dict(
remote="prgs",
org="Scaled-Tech-Consulting",
repo="Gitea-Tools",
session_id="prgs-author-1-abc",
role="author",
work_kind="issue",
work_number=660,
branch="feat/issue-660-session-checkpoint-schema",
head_sha="deadbeef",
lease_id="lease-1",
workflow_stage="implementing",
last_completed_action="wrote schema",
next_valid_action="write tests",
recovery_instructions="re-lock #660 then continue tests",
)
base.update(overrides)
return self.db.write_session_checkpoint(**base)
# AC1 — schema documented and versioned.
def test_table_exists_and_row_carries_schema_version(self) -> None:
import sqlite3
conn = sqlite3.connect(self.db_path)
try:
names = {
r[0]
for r in conn.execute(
"SELECT name FROM sqlite_master WHERE type='table'"
).fetchall()
}
finally:
conn.close()
self.assertIn("session_checkpoints", names)
record = self._write()
self.assertEqual(record["checkpoint_schema_version"], 5)
# AC2 — checkpoints written for multi-role session fixtures.
def test_multi_role_fixtures_each_get_a_row(self) -> None:
roles = [
("prgs-author-1", "author", "issue", 660),
("prgs-reviewer-2", "reviewer", "pr", 795),
("prgs-merger-3", "merger", "pr", 862),
("prgs-controller-4", "controller", "issue", 653),
]
for session_id, role, kind, number in roles:
self._write(
session_id=session_id,
role=role,
work_kind=kind,
work_number=number,
lease_id=f"lease-{session_id}",
)
rows = self.db.list_session_checkpoints(remote="prgs")
self.assertEqual(len(rows), 4)
self.assertEqual(
{r["role"] for r in rows},
{"author", "reviewer", "merger", "controller"},
)
def test_upsert_is_current_state_and_audits_stage_change(self) -> None:
first = self._write(workflow_stage="implementing")
second = self._write(workflow_stage="testing")
self.assertEqual(first["checkpoint_id"], second["checkpoint_id"])
rows = self.db.list_session_checkpoints(
remote="prgs", session_id="prgs-author-1-abc"
)
self.assertEqual(len(rows), 1)
self.assertEqual(rows[0]["workflow_stage"], "testing")
import sqlite3
conn = sqlite3.connect(self.db_path)
try:
n = conn.execute(
"SELECT COUNT(*) FROM events "
"WHERE event_type = 'session_checkpoint_stage_change'"
).fetchone()[0]
finally:
conn.close()
self.assertEqual(n, 1)
def test_get_and_roundtrip_json_fields(self) -> None:
self._write(
capabilities=["gitea.repo.commit", "gitea.pr.create"],
evidence={"tests": "4 passing"},
pending_mutation={"op": "commit_files", "files": ["control_plane_db.py"]},
)
got = self.db.get_session_checkpoint(
remote="prgs",
org="Scaled-Tech-Consulting",
repo="Gitea-Tools",
session_id="prgs-author-1-abc",
work_kind="issue",
work_number=660,
)
self.assertIsNotNone(got)
self.assertEqual(got["capabilities"], ["gitea.repo.commit", "gitea.pr.create"])
self.assertEqual(got["evidence"], {"tests": "4 passing"})
self.assertEqual(got["pending_mutation"]["op"], "commit_files")
# AC4 — no secrets in stored records.
def test_secrets_are_redacted_before_storage(self) -> None:
self._write(
recovery_instructions=(
"resume with Authorization: Bearer sk-supersecrettoken then retry"
),
evidence={"authorization": "Bearer sk-anothersecret"},
pending_mutation={"url": "https://user:[email protected]/repo.git"},
)
import sqlite3
conn = sqlite3.connect(self.db_path)
try:
row = conn.execute(
"SELECT recovery_instructions, evidence, pending_mutation "
"FROM session_checkpoints"
).fetchone()
finally:
conn.close()
blob = " ".join(str(v) for v in row)
self.assertNotIn("sk-supersecrettoken", blob)
self.assertNotIn("sk-anothersecret", blob)
self.assertNotIn("password", blob)
self.assertIn("REDACTED", blob)
# AC3 — reconcile detects stale head / lease mismatch.
def test_reconcile_flags_stale_head(self) -> None:
record = self._write(head_sha="aaaa1111")
result = self.db.reconcile_session_checkpoint(
record, live_head_sha="bbbb2222", live_lease_active=True,
live_lease_id="lease-1",
)
self.assertTrue(result["stale"])
self.assertTrue(result["head_mismatch"])
self.assertFalse(result["lease_mismatch"])
self.assertEqual(result["reconcile_action"], "reconcile_required")
def test_reconcile_flags_dead_lease(self) -> None:
record = self._write(lease_id="lease-1", head_sha="aaaa1111")
result = self.db.reconcile_session_checkpoint(
record, live_head_sha="aaaa1111", live_lease_active=False,
)
self.assertTrue(result["stale"])
self.assertFalse(result["head_mismatch"])
self.assertTrue(result["lease_mismatch"])
def test_reconcile_reassigned_lease_is_stale(self) -> None:
record = self._write(lease_id="lease-1")
result = self.db.reconcile_session_checkpoint(
record, live_lease_active=True, live_lease_id="lease-999",
)
self.assertTrue(result["lease_mismatch"])
def test_reconcile_clean_state_is_safe_to_resume(self) -> None:
record = self._write(head_sha="aaaa1111", lease_id="lease-1")
result = self.db.reconcile_session_checkpoint(
record, live_head_sha="aaaa1111", live_lease_active=True,
live_lease_id="lease-1",
)
self.assertFalse(result["stale"])
self.assertEqual(result["reconcile_action"], "safe_to_resume")
def test_unknown_live_state_never_flags_mismatch(self) -> None:
record = self._write(head_sha="aaaa1111", lease_id="lease-1")
result = self.db.reconcile_session_checkpoint(record)
self.assertFalse(result["stale"])
# Drain gate — fail closed when a checkpoint is incomplete.
def test_drain_requires_complete_checkpoint(self) -> None:
with self.assertRaises(ControlPlaneError):
self.db.write_session_checkpoint(
remote="prgs",
org="Scaled-Tech-Consulting",
repo="Gitea-Tools",
session_id="prgs-author-1-abc",
role="author",
workflow_stage="draining",
# next_valid_action + recovery_instructions intentionally absent
require_complete=True,
)
# Nothing was written.
rows = self.db.list_session_checkpoints(remote="prgs")
self.assertEqual(rows, [])
def test_drain_write_succeeds_when_complete(self) -> None:
record = self._write(require_complete=True)
self.assertEqual(record["status"], "active")
self.assertEqual(self.db.checkpoint_completeness(record), [])
def test_missing_session_id_fails_closed(self) -> None:
with self.assertRaises(ControlPlaneError):
self.db.write_session_checkpoint(
remote="prgs", org="o", repo="r", session_id="",
)
def test_session_level_checkpoint_uses_sentinel_key(self) -> None:
# No work unit -> ('', 0) sentinel; a second session-level write upserts.
self.db.write_session_checkpoint(
remote="prgs", org="o", repo="r", session_id="s-sess",
workflow_stage="idle",
)
self.db.write_session_checkpoint(
remote="prgs", org="o", repo="r", session_id="s-sess",
workflow_stage="booting",
)
rows = self.db.list_session_checkpoints(remote="prgs", session_id="s-sess")
self.assertEqual(len(rows), 1)
self.assertEqual(rows[0]["work_kind"], "")
self.assertEqual(rows[0]["work_number"], 0)
if __name__ == "__main__":
unittest.main()
@@ -313,7 +313,6 @@ class TestCanonicalRootGuardBinding(_ServerHarness):
process_project_root=self.install_root,
remote="prgs",
require_binding=True,
mode="derivation",
)
self.assertFalse(got.get("block"), got.get("reasons"))
self.assertEqual(got["resolved_slug"], TARGET_SLUG)
@@ -1,483 +0,0 @@
"""Synthetic regression coverage for dirty orphaned worktree recovery (#860).
Modeled on the #850 / #855 shape without mutating their real state.
"""
from __future__ import annotations
import json
import os
import shutil
import tempfile
import unittest
from unittest import mock
import dirty_orphan_worktree_recovery as dorec
import issue_lock_store
DEAD_PID = 999_999_999
LIVE_PID = os.getpid()
BRANCH = "fix/issue-901-dirty-orphan"
SOURCE_WT = "/repo/branches/issue-901-dirty-orphan"
RECOVERY_WT_NAME = "recovery-issue-901-dirty-orphan"
LOCAL_HEAD = "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa"
REMOTE_HEAD = "bbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbb"
OTHER_HEAD = "cccccccccccccccccccccccccccccccccccccccc"
FP_A = dorec.sha256_bytes(b"dirty-a")
FP_B = dorec.sha256_bytes(b"dirty-b")
FP_C = dorec.sha256_bytes(b"dirty-c-conflict")
def durable_lock(**overrides):
"""#850-shaped PID-less malformed same-claimant lock."""
lock = {
"issue_number": 901,
"branch_name": BRANCH,
"worktree_path": SOURCE_WT,
"remote": "prgs",
"org": "Example-Org",
"repo": "Example-Repo",
# intentionally no pid / session_pid / work_lease expiry
"claimant": {"username": "author-user", "profile": "prgs-author"},
}
lock.update(overrides)
return lock
def base_kwargs(**overrides):
kwargs = {
"issue_number": 901,
"branch_name": BRANCH,
"source_worktree_path": SOURCE_WT,
"remote": "prgs",
"org": "Example-Org",
"repo": "Example-Repo",
"identity": "author-user",
"profile": "prgs-author",
"expected_local_head": LOCAL_HEAD,
"expected_remote_head": REMOTE_HEAD,
"expected_dirty_fingerprints": {"a.py": FP_A, "b.py": FP_B},
"current_branch": BRANCH,
"porcelain_status": " M a.py\n M b.py\n",
"observed_local_head": LOCAL_HEAD,
"observed_remote_head": REMOTE_HEAD,
"observed_dirty_fingerprints": {"a.py": FP_A, "b.py": FP_B},
"competing_live_locks": [],
"competing_live_sessions": [],
"workflow_lease_active": False,
"workflow_lease_expired": True,
"canonical_repo_root": "/repo",
"worktree_registered": True,
"current_pid": LIVE_PID,
}
kwargs.update(overrides)
return kwargs
def assess(lock=None, **overrides):
return dorec.assess_dirty_orphan_recovery(
durable_lock() if lock is None else lock, **base_kwargs(**overrides)
)
class FreshnessPidLess(unittest.TestCase):
def test_pid_less_lock_is_not_live(self):
freshness = issue_lock_store.assess_lock_freshness(durable_lock())
self.assertFalse(freshness["live"])
self.assertTrue(freshness.get("pid_missing"))
self.assertEqual(freshness["status"], "malformed")
def test_pid_less_with_far_future_expiry_still_not_live(self):
lock = durable_lock(
work_lease={
"operation_type": "author_issue_work",
"expires_at": "2999-01-01T00:00:00Z",
"last_heartbeat_at": "2999-01-01T00:00:00Z",
}
)
freshness = issue_lock_store.assess_lock_freshness(lock)
self.assertFalse(freshness["live"])
self.assertTrue(freshness.get("pid_missing"))
class EligibilityGranted(unittest.TestCase):
def test_dead_same_claimant_pid_less_dirty(self):
result = assess()
self.assertEqual(result["outcome"], dorec.ELIGIBLE)
self.assertTrue(result["eligible"])
def test_expired_workflow_lease_corroboration(self):
result = assess(workflow_lease_active=False, workflow_lease_expired=True)
self.assertTrue(result["eligible"])
def test_older_local_newer_remote_heads(self):
result = assess()
self.assertTrue(result["evidence"].get("heads_diverged"))
self.assertTrue(result["eligible"])
class EligibilityRefused(unittest.TestCase):
def test_active_owner_with_pid(self):
lock = durable_lock(pid=LIVE_PID, session_pid=LIVE_PID)
result = assess(lock=lock, owner_process_alive_override=True)
self.assertEqual(result["outcome"], dorec.REFUSED)
self.assertFalse(result["eligible"])
self.assertTrue(any("alive" in r for r in result["reasons"]))
def test_foreign_claimant(self):
result = assess(identity="other-user")
self.assertEqual(result["outcome"], dorec.REFUSED)
self.assertTrue(any("foreign claimant identity" in r for r in result["reasons"]))
def test_foreign_profile(self):
result = assess(profile="prgs-reviewer")
self.assertEqual(result["outcome"], dorec.REFUSED)
def test_fingerprint_mismatch(self):
result = assess(observed_dirty_fingerprints={"a.py": "0" * 64, "b.py": FP_B})
self.assertEqual(result["outcome"], dorec.REFUSED)
self.assertTrue(any("fingerprint mismatch" in r for r in result["reasons"]))
def test_head_mismatch(self):
result = assess(observed_local_head=OTHER_HEAD)
self.assertEqual(result["outcome"], dorec.REFUSED)
def test_remote_head_mismatch(self):
result = assess(observed_remote_head=OTHER_HEAD)
self.assertEqual(result["outcome"], dorec.REFUSED)
def test_path_not_under_branches(self):
result = assess(
source_worktree_path="/tmp/branches/evil",
# lock path also changed so worktree agreement holds
lock=durable_lock(worktree_path="/tmp/branches/evil"),
)
self.assertEqual(result["outcome"], dorec.REFUSED)
self.assertTrue(any("canonical branches" in r for r in result["reasons"]))
def test_unregistered_worktree(self):
result = assess(worktree_registered=False)
self.assertEqual(result["outcome"], dorec.REFUSED)
def test_active_workflow_lease(self):
result = assess(workflow_lease_active=True, workflow_lease_expired=False)
self.assertEqual(result["outcome"], dorec.REFUSED)
def test_unsafe_dirty_path_pin(self):
result = assess(
expected_dirty_fingerprints={"../etc/passwd": FP_A},
observed_dirty_fingerprints={"../etc/passwd": FP_A},
)
self.assertEqual(result["outcome"], dorec.REFUSED)
def test_symlink_escape_rejected_by_ancestry(self):
ok, reasons = dorec.is_path_under_canonical_branches(
"/tmp/branches/evil", canonical_repo_root="/repo"
)
self.assertFalse(ok)
self.assertTrue(reasons)
class ConflictDetection(unittest.TestCase):
def test_overlapping_upstream_change(self):
conflicts = dorec.detect_path_conflicts(
dirty_paths=["c.py"],
local_head_contents={"c.py": b"local-base"},
remote_head_contents={"c.py": b"remote-changed"},
dirty_contents={"c.py": b"dirty-c-conflict"},
)
self.assertEqual(len(conflicts), 1)
self.assertEqual(conflicts[0]["path"], "c.py")
def test_unchanged_upstream_no_conflict(self):
conflicts = dorec.detect_path_conflicts(
dirty_paths=["a.py"],
local_head_contents={"a.py": b"same"},
remote_head_contents={"a.py": b"same"},
dirty_contents={"a.py": b"dirty-a"},
)
self.assertEqual(conflicts, [])
class CrashSafeRecovery(unittest.TestCase):
def setUp(self):
self.tmp = tempfile.mkdtemp(prefix="dirty-orphan-")
self.repo = os.path.join(self.tmp, "repo")
self.branches = os.path.join(self.repo, "branches")
self.source = os.path.join(self.branches, "issue-901-dirty-orphan")
self.recovery = os.path.join(self.branches, RECOVERY_WT_NAME)
os.makedirs(self.source, exist_ok=True)
os.makedirs(self.branches, exist_ok=True)
# seed dirty files in source
with open(os.path.join(self.source, "a.py"), "wb") as fh:
fh.write(b"dirty-a")
with open(os.path.join(self.source, "b.py"), "wb") as fh:
fh.write(b"dirty-b")
self.journal_dir = os.path.join(self.tmp, "journals")
self.lock = durable_lock(worktree_path=self.source)
self.assessment = dorec.assess_dirty_orphan_recovery(
self.lock,
**base_kwargs(
source_worktree_path=self.source,
canonical_repo_root=self.repo,
),
)
class FakeGit(dorec.GitOps):
def __init__(self, recovery_path, head):
self.recovery_path = recovery_path
self.head = head
self.calls = []
def run(self, args, *, cwd):
self.calls.append((args, cwd))
if args[:3] == ["git", "worktree", "add"]:
os.makedirs(self.recovery_path, exist_ok=True)
return mock.Mock(returncode=0, stdout="", stderr="")
if args[:2] == ["git", "checkout"]:
return mock.Mock(returncode=0, stdout="", stderr="")
if args[:2] == ["git", "rev-parse"]:
return mock.Mock(returncode=0, stdout=self.head + "\n", stderr="")
return mock.Mock(returncode=0, stdout="", stderr="")
self.git = FakeGit(self.recovery, REMOTE_HEAD)
self.written_locks = []
def lock_writer(record):
self.written_locks.append(record)
self.lock_writer = lock_writer
def tearDown(self):
shutil.rmtree(self.tmp, ignore_errors=True)
def _run(self, **overrides):
kwargs = {
"assessment": self.assessment,
"existing_lock": self.lock,
"issue_number": 901,
"branch_name": BRANCH,
"source_worktree_path": self.source,
"recovery_worktree_path": self.recovery,
"remote": "prgs",
"org": "Example-Org",
"repo": "Example-Repo",
"identity": "author-user",
"profile": "prgs-author",
"expected_local_head": LOCAL_HEAD,
"expected_remote_head": REMOTE_HEAD,
"expected_dirty_fingerprints": {"a.py": FP_A, "b.py": FP_B},
"dirty_contents": {"a.py": b"dirty-a", "b.py": b"dirty-b"},
"local_head_contents": {"a.py": b"base-a", "b.py": b"base-b"},
"remote_head_contents": {"a.py": b"base-a", "b.py": b"base-b"},
"canonical_repo_root": self.repo,
"bind_lock": True,
"lock_writer": self.lock_writer,
"git_ops": self.git,
"journal_dir": self.journal_dir,
"session_pid": LIVE_PID,
}
kwargs.update(overrides)
return dorec.run_dirty_orphan_recovery(**kwargs)
def test_success_preserves_dirty_bytes_and_source(self):
result = self._run()
self.assertTrue(result["success"])
self.assertEqual(result["outcome"], dorec.RECOVERY_COMPLETED)
self.assertTrue(os.path.isdir(self.source))
with open(os.path.join(self.source, "a.py"), "rb") as fh:
self.assertEqual(fh.read(), b"dirty-a")
with open(os.path.join(self.recovery, "a.py"), "rb") as fh:
self.assertEqual(fh.read(), b"dirty-a")
with open(os.path.join(self.recovery, "b.py"), "rb") as fh:
self.assertEqual(fh.read(), b"dirty-b")
self.assertEqual(len(self.written_locks), 1)
rec = self.written_locks[0]
self.assertEqual(rec["session_pid"], LIVE_PID)
self.assertTrue(rec["dirty_orphan_recovery"]["recovered"])
self.assertTrue(rec["dirty_orphan_recovery"]["source_frozen"])
def test_conflict_leaves_governed_state(self):
result = self._run(
expected_dirty_fingerprints={"c.py": FP_C},
dirty_contents={"c.py": b"dirty-c-conflict"},
local_head_contents={"c.py": b"local-base"},
remote_head_contents={"c.py": b"remote-changed"},
)
# #860 F4: session binding is NOT finalized while conflicts remain
self.assertFalse(result["success"])
self.assertEqual(result["outcome"], dorec.CONFLICTS_PRESENT)
sidecar = os.path.join(self.recovery, "c.py.recovered-dirty")
self.assertTrue(os.path.isfile(sidecar))
state = os.path.join(
self.recovery, dorec.CONFLICT_STATE_DIR, dorec.CONFLICT_STATE_FILE
)
self.assertTrue(os.path.isfile(state))
with open(state, "r", encoding="utf-8") as fh:
payload = json.load(fh)
self.assertEqual(payload["resolution"], "author_edit_required")
def test_interrupt_before_journal_no_artifacts(self):
result = self._run(interrupt_after_phase=dorec.PHASE_ELIGIBILITY)
self.assertFalse(result["success"])
self.assertEqual(result["outcome"], "INTERRUPTED")
self.assertFalse(os.path.isdir(self.recovery))
def test_interrupt_after_journal_then_retry_idempotent(self):
first = self._run(interrupt_after_phase=dorec.PHASE_JOURNAL_PERSISTED)
self.assertEqual(first["outcome"], "INTERRUPTED")
self.assertTrue(first["journal"]["artifacts_created"]["journal"])
second = self._run()
self.assertTrue(second["success"])
# source still recoverable
with open(os.path.join(self.source, "a.py"), "rb") as fh:
self.assertEqual(fh.read(), b"dirty-a")
def test_interrupt_after_worktree_then_retry(self):
first = self._run(interrupt_after_phase=dorec.PHASE_RECOVERY_WORKTREE)
self.assertEqual(first["outcome"], "INTERRUPTED")
self.assertTrue(os.path.isdir(self.recovery))
second = self._run()
self.assertTrue(second["success"])
def test_interrupt_after_binding_then_retry_complete(self):
first = self._run(interrupt_after_phase=dorec.PHASE_BINDING)
self.assertEqual(first["outcome"], "INTERRUPTED")
second = self._run()
self.assertTrue(second["success"])
# completed journal makes further retries no-ops
third = self._run()
self.assertEqual(third["outcome"], dorec.RECOVERY_RESUMED)
def test_source_worktree_never_deleted(self):
self._run()
self.assertTrue(os.path.isdir(self.source))
self.assertTrue(os.path.isfile(os.path.join(self.source, "a.py")))
def test_fingerprint_drift_refuses_without_mutation(self):
result = self._run(dirty_contents={"a.py": b"CHANGED", "b.py": b"dirty-b"})
self.assertFalse(result["success"])
self.assertFalse(os.path.isdir(self.recovery))
class SessionBindingPreflight(unittest.TestCase):
def test_canonical_session_binding_recognized(self):
lock = {
"worktree_path": "/repo/branches/recovery",
"session_pid": LIVE_PID,
"dirty_orphan_recovery": {
"recovered": True,
"conflicts": [],
"recovery_worktree_path": "/repo/branches/recovery",
"source_worktree_path": SOURCE_WT,
"accepted_head": REMOTE_HEAD,
},
}
result = dorec.preflight_recognizes_recovered_provenance(lock)
self.assertTrue(result["recognized"])
def test_conflicts_block_commit_preflight(self):
lock = {
"worktree_path": "/repo/branches/recovery",
"session_pid": LIVE_PID,
"dirty_orphan_recovery": {
"recovered": True,
"conflicts": [{"path": "c.py"}],
},
}
result = dorec.preflight_recognizes_recovered_provenance(lock)
self.assertFalse(result["recognized"])
def test_active_foreign_does_not_mutate(self):
# assess-only path: foreign refused before run
result = assess(identity="intruder")
self.assertFalse(result["eligible"])
class JournalSymlinkRefusal(unittest.TestCase):
def test_symlink_journal_path_refused_on_load(self):
tmp = tempfile.mkdtemp()
try:
real = os.path.join(tmp, "real.json")
with open(real, "w", encoding="utf-8") as fh:
fh.write("{}")
link = os.path.join(tmp, "link.json")
os.symlink(real, link)
key = "symlink-test"
jdir = tmp
path = dorec._journal_path(key, journal_dir=jdir)
with open(path, "w", encoding="utf-8") as fh:
json.dump({"idempotency_key": key}, fh)
os.remove(path)
os.symlink(real, path)
with self.assertRaises(ValueError):
dorec.load_journal(key, journal_dir=jdir)
finally:
shutil.rmtree(tmp, ignore_errors=True)
class RealGitMultiWorktreeIntegration(unittest.TestCase):
def setUp(self):
import subprocess
self.tmp = tempfile.mkdtemp(prefix="git-integration-")
self.repo = os.path.join(self.tmp, "repo")
os.makedirs(self.repo, exist_ok=True)
subprocess.run(["git", "init"], cwd=self.repo, check=True, capture_output=True)
subprocess.run(["git", "config", "user.name", "Test User"], cwd=self.repo, check=True)
subprocess.run(["git", "config", "user.email", "[email protected]"], cwd=self.repo, check=True)
with open(os.path.join(self.repo, "init.txt"), "w") as fh:
fh.write("init")
subprocess.run(["git", "add", "."], cwd=self.repo, check=True)
subprocess.run(["git", "commit", "-m", "init"], cwd=self.repo, check=True)
branch = "fix/issue-999-test"
subprocess.run(["git", "branch", branch], cwd=self.repo, check=True)
self.branches = os.path.join(self.repo, "branches")
self.source = os.path.join(self.branches, "issue-999-test")
subprocess.run(["git", "worktree", "add", self.source, branch], cwd=self.repo, check=True)
self.dirty_path = os.path.join(self.source, "dirty.txt")
with open(self.dirty_path, "w") as fh:
fh.write("dirty-data")
def tearDown(self):
shutil.rmtree(self.tmp, ignore_errors=True)
def test_prepare_recovery_worktree_detached_no_exit_128(self):
import subprocess
head_sha = subprocess.check_output(["git", "rev-parse", "HEAD"], cwd=self.repo, text=True).strip()
rec_wt = os.path.join(self.branches, "recovery-issue-999-test")
res = dorec.prepare_recovery_worktree(
canonical_repo_root=self.repo,
recovery_worktree_path=rec_wt,
branch_name="fix/issue-999-test",
remote_head=head_sha,
)
self.assertTrue(res["success"], res.get("reasons"))
self.assertTrue(os.path.isdir(rec_wt))
def test_real_lock_rebind_recovery_sanctioned(self):
lock_dir = os.path.join(self.tmp, "locks")
rec_wt = os.path.join(self.branches, "recovery-issue-999-test")
os.makedirs(rec_wt, exist_ok=True)
record = {
"remote": "prgs",
"org": "Example-Org",
"repo": "Example-Repo",
"issue_number": 999,
"branch_name": "fix/issue-999-test",
"worktree_path": rec_wt,
"claimant": {"username": "author-user", "profile": "prgs-author"},
}
record_src = dict(record)
record_src["worktree_path"] = self.source
issue_lock_store.bind_session_lock(record_src, lock_dir=lock_dir)
path = issue_lock_store.bind_session_lock(
record,
lock_dir=lock_dir,
recovery_sanctioned=True,
)
self.assertTrue(os.path.isfile(path))
if __name__ == "__main__":
unittest.main()
File diff suppressed because it is too large Load Diff
-907
View File
@@ -1,907 +0,0 @@
"""Tests for the pre-restart drain proof and hard gate (#661).
Covers the acceptance criteria:
1. Restart apply without a proof fails closed.
2. A successful drain produces a verifiable proof.
3. An open unsafe mutation makes the proof fail (multi-session fixture).
4. Pass / fail / expired verification paths.
Plus the security posture: forged/tampered proofs are rejected, break-glass is
the only bypass and is never silent, a stale blast-radius fingerprint rejects a
proof, and no per-process secret ever leaks into a serialized artifact.
"""
from __future__ import annotations
import os
import unittest
from datetime import datetime, timedelta, timezone
import drain_proof as dp
import restart_coordinator as rc
NOW = datetime(2026, 7, 24, 6, 0, 0, tzinfo=timezone.utc)
SECRET = b"unit-test-drain-proof-secret-0123456789abcdef"
def _live_pid() -> int:
return os.getpid()
def _clean_drain_state() -> dict:
"""Every drain action succeeded, no sessions outstanding."""
return {
"assignments_stopped": True,
"checkpoints_complete": True,
"handoffs_verified": True,
"leases_handled": True,
"acks": {}, # no other live sessions to acknowledge
"ack_timeout_policy_applied": False,
}
def _safe_report() -> dict:
"""Impact report with no other live work: a restart here is safe."""
report = rc.evaluate_restart_impact(
{"sessions": [], "leases": [], "inventory_complete": True},
now=NOW,
requesting_session_id="prgs-controller-1-req",
)
return report.as_dict()
def _unsafe_mutation_report() -> dict:
"""Multi-session report: a second session holds a live author mutation."""
sessions = [
{
"session_id": "prgs-controller-1-req",
"role": "controller",
"profile": "prgs-controller",
"pid": _live_pid(),
"status": "active",
"last_heartbeat_at": NOW.isoformat(),
},
{
"session_id": "prgs-author-99",
"role": "author",
"profile": "prgs-author",
"pid": _live_pid(),
"status": "active",
"last_heartbeat_at": NOW.isoformat(),
},
]
leases = [
{
"lease_id": "lease-mut",
"session_id": "prgs-author-99",
"role": "author",
"phase": "implementing",
"work_kind": "issue",
"work_number": 661,
"worktree_path": "branches/issue-661",
"freshness": {"freshness": "active"},
}
]
report = rc.evaluate_restart_impact(
{"sessions": sessions, "leases": leases, "inventory_complete": True},
now=NOW,
requesting_session_id="prgs-controller-1-req",
)
return report.as_dict()
class BuildDrainProofTests(unittest.TestCase):
def test_clean_drain_produces_verifiable_clean_proof(self):
"""AC#2: a successful drain produces a verifiable proof."""
proof = dp.build_drain_proof(
impact_report=_safe_report(),
drain_state=_clean_drain_state(),
requesting_session_id="prgs-controller-1-req",
now=NOW,
secret=SECRET,
)
self.assertTrue(proof.clean)
self.assertEqual(proof.failed_checks, [])
self.assertEqual(
{c.name for c in proof.checks}, set(dp.REQUIRED_CHECKS)
)
result = dp.verify_drain_proof(
proof.as_dict(), now=NOW, secret=SECRET
)
self.assertTrue(result.valid, result.reasons)
self.assertFalse(result.expired)
self.assertFalse(result.tampered)
def test_open_mutation_makes_proof_unclean(self):
"""AC#3: an unsafe mutation still in flight fails the proof."""
proof = dp.build_drain_proof(
impact_report=_unsafe_mutation_report(),
drain_state=_clean_drain_state(),
now=NOW,
secret=SECRET,
)
self.assertFalse(proof.clean)
self.assertIn(dp.CHECK_NO_INFLIGHT_MUTATIONS, proof.failed_checks)
# Leases-handled also fails: the report still shows a disruptive lease.
self.assertIn(dp.CHECK_LEASES_HANDLED, proof.failed_checks)
result = dp.verify_drain_proof(proof.as_dict(), now=NOW, secret=SECRET)
self.assertFalse(result.valid)
def test_incomplete_inventory_fails_no_mutations_check(self):
proof = dp.build_drain_proof(
impact_report={"inventory_complete": False},
drain_state=_clean_drain_state(),
now=NOW,
secret=SECRET,
)
self.assertFalse(proof.clean)
self.assertIn(dp.CHECK_NO_INFLIGHT_MUTATIONS, proof.failed_checks)
def test_missing_checkpoint_flag_fails_closed(self):
state = _clean_drain_state()
del state["checkpoints_complete"]
proof = dp.build_drain_proof(
impact_report=_safe_report(), drain_state=state, now=NOW, secret=SECRET
)
self.assertFalse(proof.clean)
self.assertIn(dp.CHECK_CHECKPOINTS_COMPLETE, proof.failed_checks)
def test_non_true_flags_fail_closed(self):
"""A truthy-but-not-True value (e.g. the string 'yes') must not pass."""
state = _clean_drain_state()
state["assignments_stopped"] = "yes"
proof = dp.build_drain_proof(
impact_report=_safe_report(), drain_state=state, now=NOW, secret=SECRET
)
self.assertIn(dp.CHECK_ASSIGNMENTS_STOPPED, proof.failed_checks)
def test_ack_timeout_policy_satisfies_ack_check(self):
state = _clean_drain_state()
state["acks"] = {"prgs-author-99": "pending"}
state["ack_timeout_policy_applied"] = True
proof = dp.build_drain_proof(
impact_report=_safe_report(), drain_state=state, now=NOW, secret=SECRET
)
names = {c.name: c.passed for c in proof.checks}
self.assertTrue(names[dp.CHECK_ACKS_OR_TIMEOUT])
def test_outstanding_acks_without_timeout_fail(self):
state = _clean_drain_state()
state["acks"] = {"prgs-author-99": "pending"}
state["ack_timeout_policy_applied"] = False
proof = dp.build_drain_proof(
impact_report=_safe_report(), drain_state=state, now=NOW, secret=SECRET
)
self.assertIn(dp.CHECK_ACKS_OR_TIMEOUT, proof.failed_checks)
def test_all_acked_satisfies_ack_check(self):
state = _clean_drain_state()
state["acks"] = {"prgs-author-99": "acked", "prgs-author-2": "acknowledged"}
proof = dp.build_drain_proof(
impact_report=_safe_report(), drain_state=state, now=NOW, secret=SECRET
)
names = {c.name: c.passed for c in proof.checks}
self.assertTrue(names[dp.CHECK_ACKS_OR_TIMEOUT])
class VerifyDrainProofTests(unittest.TestCase):
def _clean_proof_dict(self) -> dict:
return dp.build_drain_proof(
impact_report=_safe_report(),
drain_state=_clean_drain_state(),
now=NOW,
secret=SECRET,
).as_dict()
def test_missing_proof_is_invalid(self):
result = dp.verify_drain_proof(None, now=NOW, secret=SECRET)
self.assertFalse(result.valid)
self.assertIsNone(result.proof_id)
def test_expired_proof_is_invalid(self):
"""AC#4: an expired proof fails verification."""
proof = self._clean_proof_dict()
later = NOW + timedelta(seconds=dp.DEFAULT_PROOF_TTL_SECONDS + 1)
result = dp.verify_drain_proof(proof, now=later, secret=SECRET)
self.assertFalse(result.valid)
self.assertTrue(result.expired)
def test_proof_valid_just_before_expiry(self):
proof = self._clean_proof_dict()
almost = NOW + timedelta(seconds=dp.DEFAULT_PROOF_TTL_SECONDS - 1)
result = dp.verify_drain_proof(proof, now=almost, secret=SECRET)
self.assertTrue(result.valid, result.reasons)
def test_wrong_secret_rejected(self):
"""A proof minted in a prior process (different secret) will not verify."""
proof = self._clean_proof_dict()
result = dp.verify_drain_proof(proof, now=NOW, secret=b"other-secret")
self.assertFalse(result.valid)
self.assertTrue(result.tampered)
def test_flipping_clean_flag_is_detected(self):
"""Forging clean=True on an unclean proof breaks the signature."""
unclean = dp.build_drain_proof(
impact_report=_unsafe_mutation_report(),
drain_state=_clean_drain_state(),
now=NOW,
secret=SECRET,
).as_dict()
self.assertFalse(unclean["clean"])
unclean["clean"] = True # forge
result = dp.verify_drain_proof(unclean, now=NOW, secret=SECRET)
self.assertFalse(result.valid)
self.assertTrue(result.tampered)
def test_tampering_a_check_is_detected(self):
unclean = dp.build_drain_proof(
impact_report=_unsafe_mutation_report(),
drain_state=_clean_drain_state(),
now=NOW,
secret=SECRET,
).as_dict()
for c in unclean["checks"]:
if c["name"] == dp.CHECK_NO_INFLIGHT_MUTATIONS:
c["passed"] = True # forge the failing check to pass
result = dp.verify_drain_proof(unclean, now=NOW, secret=SECRET)
self.assertFalse(result.valid)
self.assertTrue(result.tampered)
def test_missing_required_check_rejected(self):
proof = self._clean_proof_dict()
proof["checks"] = [
c for c in proof["checks"] if c["name"] != dp.CHECK_HANDOFFS_OK
]
result = dp.verify_drain_proof(proof, now=NOW, secret=SECRET)
self.assertFalse(result.valid)
def test_stale_fingerprint_rejected(self):
proof = self._clean_proof_dict()
result = dp.verify_drain_proof(
proof,
now=NOW,
secret=SECRET,
expected_impact_fingerprint="deadbeef",
)
self.assertFalse(result.valid)
def test_matching_fingerprint_accepted(self):
report = _safe_report()
proof = dp.build_drain_proof(
impact_report=report,
drain_state=_clean_drain_state(),
now=NOW,
secret=SECRET,
).as_dict()
fp = dp.impact_fingerprint(report)
result = dp.verify_drain_proof(
proof, now=NOW, secret=SECRET, expected_impact_fingerprint=fp
)
self.assertTrue(result.valid, result.reasons)
class GateApplyRestartTests(unittest.TestCase):
def _clean_proof_dict(self) -> dict:
return dp.build_drain_proof(
impact_report=_safe_report(),
drain_state=_clean_drain_state(),
now=NOW,
secret=SECRET,
).as_dict()
def test_apply_without_proof_denied(self):
"""AC#1: restart apply without a proof fails closed + raises incident."""
decision = dp.gate_apply_restart(proof=None, now=NOW, secret=SECRET)
self.assertFalse(decision.allow)
self.assertEqual(decision.verdict, dp.GATE_DENY)
self.assertIsNotNone(decision.incident)
self.assertEqual(
decision.incident["kind"], "restart_drain_gate_denied"
)
def test_apply_with_valid_proof_allowed(self):
decision = dp.gate_apply_restart(
proof=self._clean_proof_dict(), now=NOW, secret=SECRET
)
self.assertTrue(decision.allow)
self.assertEqual(decision.verdict, dp.GATE_ALLOW)
self.assertIsNone(decision.incident)
def test_apply_with_expired_proof_denied_with_incident(self):
later = NOW + timedelta(seconds=dp.DEFAULT_PROOF_TTL_SECONDS + 5)
decision = dp.gate_apply_restart(
proof=self._clean_proof_dict(), now=later, secret=SECRET
)
self.assertFalse(decision.allow)
self.assertIsNotNone(decision.incident)
def test_apply_with_unclean_proof_denied(self):
"""AC#3 at the gate: an unsafe-mutation proof is denied."""
unclean = dp.build_drain_proof(
impact_report=_unsafe_mutation_report(),
drain_state=_clean_drain_state(),
now=NOW,
secret=SECRET,
).as_dict()
decision = dp.gate_apply_restart(proof=unclean, now=NOW, secret=SECRET)
self.assertFalse(decision.allow)
self.assertIsNotNone(decision.incident)
def test_break_glass_allows_without_proof_but_records_bypass(self):
decision = dp.gate_apply_restart(
proof=None, now=NOW, secret=SECRET, break_glass=True
)
self.assertTrue(decision.allow)
self.assertEqual(decision.verdict, dp.GATE_BREAK_GLASS)
self.assertTrue(decision.break_glass)
self.assertIsNone(decision.incident)
self.assertTrue(decision.audit_record["break_glass"])
def test_denied_gate_carries_stale_fingerprint_reason(self):
decision = dp.gate_apply_restart(
proof=self._clean_proof_dict(),
now=NOW,
secret=SECRET,
expected_impact_fingerprint="not-the-fingerprint",
)
self.assertFalse(decision.allow)
class SecretHygieneTests(unittest.TestCase):
def test_secret_never_serialized(self):
proof = dp.build_drain_proof(
impact_report=_safe_report(),
drain_state=_clean_drain_state(),
now=NOW,
secret=SECRET,
)
blob = dp._canonical(proof.as_dict())
self.assertNotIn(SECRET.decode(), blob)
# The signature is a hex digest, not the raw secret.
self.assertNotIn(SECRET.hex(), blob)
def test_incident_descriptor_has_no_secret(self):
decision = dp.gate_apply_restart(proof=None, now=NOW, secret=SECRET)
blob = dp._canonical(decision.incident)
self.assertNotIn(SECRET.decode(), blob)
def _drained_report_with_live_sessions(count: int) -> dict:
"""Report with ``count`` other live sessions but nothing in flight.
Every other checklist item passes against this report, so a failure
isolates the acknowledgement check rather than tripping on mutations.
"""
sessions = [
{
"session_id": "prgs-controller-1-req",
"role": "controller",
"profile": "prgs-controller",
"pid": _live_pid(),
"status": "active",
"last_heartbeat_at": NOW.isoformat(),
}
]
for index in range(count):
sessions.append(
{
"session_id": f"prgs-author-{index}",
"role": "author",
"profile": "prgs-author",
"pid": _live_pid(),
"status": "active",
"last_heartbeat_at": NOW.isoformat(),
}
)
report = rc.evaluate_restart_impact(
{"sessions": sessions, "leases": [], "inventory_complete": True},
now=NOW,
requesting_session_id="prgs-controller-1-req",
)
return report.as_dict()
class AcknowledgementFailClosedTests(unittest.TestCase):
"""Acknowledgement evidence must fail closed unless explicitly verified.
Regression cover for the reviewed fail-open on PR #882: an absent ``acks``
key collapsed to ``{}`` and was read as "no other live sessions required to
acknowledge", so a proof minted clean and the restart gate allowed while the
impact report still showed other live sessions.
"""
def _state(self, **overrides) -> dict:
state = _clean_drain_state()
state.pop("acks", None)
state["ack_timeout_policy_applied"] = False
state.update(overrides)
return state
def _acks_check(self, proof) -> dp.DrainCheck:
return next(c for c in proof.checks if c.name == dp.CHECK_ACKS_OR_TIMEOUT)
def _build(self, report: dict, state: dict):
return dp.build_drain_proof(
impact_report=report, drain_state=state, now=NOW, secret=SECRET
)
def assertAcksFailClosed(self, report: dict, state: dict) -> None:
proof = self._build(report, state)
self.assertFalse(self._acks_check(proof).passed)
self.assertIn(dp.CHECK_ACKS_OR_TIMEOUT, proof.failed_checks)
self.assertFalse(proof.clean)
# --- missing / null / empty / malformed ------------------------------
def test_missing_acks_key_with_live_sessions_fails_closed(self):
"""The exact reviewed defect: absent key, three other live sessions."""
report = _drained_report_with_live_sessions(3)
self.assertEqual(report["counts"]["sessions_live_other"], 3)
state = self._state()
self.assertNotIn("acks", state)
proof = self._build(report, state)
check = self._acks_check(proof)
self.assertFalse(check.passed)
self.assertNotIn("no other live sessions", check.detail)
self.assertIn("fail closed", check.detail)
self.assertFalse(proof.clean)
self.assertEqual(proof.failed_checks, [dp.CHECK_ACKS_OR_TIMEOUT])
def test_none_acks_with_live_sessions_fails_closed(self):
self.assertAcksFailClosed(
_drained_report_with_live_sessions(2), self._state(acks=None)
)
def test_empty_acks_with_live_sessions_fails_closed(self):
self.assertAcksFailClosed(
_drained_report_with_live_sessions(1), self._state(acks={})
)
def test_malformed_acks_fail_closed(self):
for malformed in ([], "ack", 7, ("ack",), True):
with self.subTest(malformed=malformed):
self.assertAcksFailClosed(
_drained_report_with_live_sessions(1),
self._state(acks=malformed),
)
# --- stale / unproven values -----------------------------------------
def test_stale_or_unproven_ack_values_fail_closed(self):
for value in ("pending", "stale", "unknown", "", None, True, 1, NOW):
with self.subTest(value=value):
self.assertAcksFailClosed(
_drained_report_with_live_sessions(1),
self._state(acks={"prgs-author-0": value}),
)
def test_partial_coverage_fails_closed(self):
"""Fewer acknowledgements than the report's live-session count."""
self.assertAcksFailClosed(
_drained_report_with_live_sessions(3),
self._state(acks={"prgs-author-0": "ack"}),
)
def test_one_unacked_entry_among_many_fails_closed(self):
self.assertAcksFailClosed(
_drained_report_with_live_sessions(2),
self._state(acks={"prgs-author-0": "ack", "prgs-author-1": "pending"}),
)
def test_unproven_live_session_count_fails_closed(self):
"""A missing or malformed count cannot prove nobody had to acknowledge."""
malformed_counts = (
None,
{},
{"sessions_live_other": None},
{"sessions_live_other": "3"},
{"sessions_live_other": -1},
{"sessions_live_other": True},
)
for counts in malformed_counts:
with self.subTest(counts=counts):
report = _drained_report_with_live_sessions(0)
if counts is None:
report.pop("counts", None)
else:
report["counts"] = counts
self.assertAcksFailClosed(report, self._state())
# --- valid evidence still passes -------------------------------------
def test_complete_valid_acks_pass(self):
report = _drained_report_with_live_sessions(2)
state = self._state(
acks={"prgs-author-0": "ack", "prgs-author-1": "acknowledged"}
)
proof = self._build(report, state)
self.assertTrue(self._acks_check(proof).passed)
self.assertTrue(proof.clean)
self.assertEqual(proof.failed_checks, [])
def test_no_other_live_sessions_still_passes(self):
"""Intended behavior retained: zero live sessions needs no acks."""
report = _drained_report_with_live_sessions(0)
self.assertEqual(report["counts"]["sessions_live_other"], 0)
proof = self._build(report, self._state())
check = self._acks_check(proof)
self.assertTrue(check.passed)
self.assertIn("sessions_live_other=0", check.detail)
self.assertTrue(proof.clean)
# --- timeout policy cannot become a second fail-open ------------------
def test_unproven_timeout_policy_cannot_open_the_gate(self):
for value in (None, "true", "yes", 1, "True", [], {}):
with self.subTest(value=value):
self.assertAcksFailClosed(
_drained_report_with_live_sessions(2),
self._state(ack_timeout_policy_applied=value),
)
def test_explicit_timeout_policy_permits(self):
proof = self._build(
_drained_report_with_live_sessions(2),
self._state(ack_timeout_policy_applied=True),
)
check = self._acks_check(proof)
self.assertTrue(check.passed)
self.assertIn("timeout policy", check.detail)
self.assertTrue(proof.clean)
# --- the gate itself must deny ---------------------------------------
def test_failed_ack_check_denies_the_restart_gate(self):
report = _drained_report_with_live_sessions(3)
proof = self._build(report, self._state())
self.assertFalse(proof.clean)
decision = dp.gate_apply_restart(
proof=proof.as_dict(),
now=NOW,
secret=SECRET,
expected_impact_fingerprint=dp.impact_fingerprint(report),
)
self.assertFalse(decision.allow)
self.assertEqual(decision.verdict, dp.GATE_DENY)
self.assertIsNotNone(decision.incident)
def test_unclean_ack_proof_fails_verification(self):
report = _drained_report_with_live_sessions(3)
proof = self._build(report, self._state())
result = dp.verify_drain_proof(
proof.as_dict(),
now=NOW,
secret=SECRET,
expected_impact_fingerprint=dp.impact_fingerprint(report),
)
self.assertFalse(result.valid)
self.assertFalse(result.clean)
def _identity_report(*, requester: str, others: tuple[str, ...]) -> dict:
"""Report with explicitly named requester and other live sessions.
Unlike :func:`_drained_report_with_live_sessions`, the session ids are
chosen by the caller so a test can supply acknowledgements for the *wrong*
identities while keeping the count correct.
"""
sessions = [
{
"session_id": requester,
"role": "controller",
"profile": "prgs-controller",
"pid": _live_pid(),
"status": "active",
"last_heartbeat_at": NOW.isoformat(),
}
]
for session_id in others:
sessions.append(
{
"session_id": session_id,
"role": "author",
"profile": "prgs-author",
"pid": _live_pid(),
"status": "active",
"last_heartbeat_at": NOW.isoformat(),
}
)
report = rc.evaluate_restart_impact(
{"sessions": sessions, "leases": [], "inventory_complete": True},
now=NOW,
requesting_session_id=requester,
)
return report.as_dict()
class AcknowledgementIdentityBindingTests(unittest.TestCase):
"""Acknowledgement coverage must be bound to session identity, not counted.
Regression cover for the second reviewed fail-open on PR #882 (review 582,
blocker B1): coverage compared ``acked_count`` against
``counts.sessions_live_other``, so acknowledgements supplied for the
requesting session and for ids that do not exist satisfied the obligations
of the live sessions that never answered. The required identities are
carried by the report itself ``ack_state`` keys and ``affected_sessions``
filtered on ``live and not is_requester`` and only an acknowledgement
keyed by one of those ids may count for it.
"""
def _state(self, **overrides) -> dict:
state = _clean_drain_state()
state.pop("acks", None)
state["ack_timeout_policy_applied"] = False
state.update(overrides)
return state
def _acks_check(self, proof) -> dp.DrainCheck:
return next(c for c in proof.checks if c.name == dp.CHECK_ACKS_OR_TIMEOUT)
def _build(self, report: dict, state: dict):
return dp.build_drain_proof(
impact_report=report, drain_state=state, now=NOW, secret=SECRET
)
def assertAcksFailClosed(self, report: dict, state: dict) -> dp.DrainCheck:
"""Failure must propagate through the check, the proof, and the gate."""
proof = self._build(report, state)
check = self._acks_check(proof)
self.assertFalse(check.passed)
self.assertFalse(proof.clean)
self.assertIn(dp.CHECK_ACKS_OR_TIMEOUT, proof.failed_checks)
decision = dp.gate_apply_restart(
proof=proof.as_dict(),
now=NOW,
secret=SECRET,
expected_impact_fingerprint=dp.impact_fingerprint(report),
)
self.assertEqual(decision.verdict, dp.GATE_DENY)
self.assertFalse(decision.allow)
return check
# --- the reviewer's exact reproduction --------------------------------
def test_requester_plus_unknown_id_cannot_satisfy_two_live_sessions(self):
"""Review 582 B1 verbatim: requester + a nonexistent session.
``sessions_live_other=2`` with ``ack_state`` naming ``other-0`` and
``other-1``; the drain state supplies an acknowledgement from the
requesting session itself and from a session that does not exist. The
count matches, the identities do not.
"""
report = _identity_report(requester="req", others=("other-0", "other-1"))
self.assertEqual(report["counts"]["sessions_live_other"], 2)
self.assertEqual(
report["ack_state"], {"other-0": "pending", "other-1": "pending"}
)
state = self._state(acks={"req": "ack", "totally-bogus-session": "ack"})
check = self.assertAcksFailClosed(report, state)
self.assertIn("other-0", check.detail)
self.assertIn("other-1", check.detail)
self.assertIn("fail closed", check.detail)
# --- wrong / unknown / requester identities ---------------------------
def test_sufficient_count_of_wrong_ids_fails_closed(self):
"""Right cardinality, wrong identities: two acks, neither required."""
report = _identity_report(requester="req", others=("other-0", "other-1"))
state = self._state(acks={"ghost-a": "ack", "ghost-b": "ack"})
check = self.assertAcksFailClosed(report, state)
self.assertIn("do not count", check.detail)
def test_more_acks_than_required_still_fails_on_wrong_ids(self):
"""Coverage cannot be bought with volume: five acks, none required."""
report = _identity_report(requester="req", others=("other-0", "other-1"))
state = self._state(acks={f"ghost-{i}": "acknowledged" for i in range(5)})
self.assertAcksFailClosed(report, state)
def test_partial_identity_match_fails_closed(self):
"""One required id acknowledged, the rest padded with unknown ids."""
report = _identity_report(
requester="req", others=("other-0", "other-1", "other-2")
)
state = self._state(
acks={"other-0": "ack", "ghost-1": "ack", "ghost-2": "ack"}
)
check = self.assertAcksFailClosed(report, state)
self.assertIn("other-1", check.detail)
self.assertIn("other-2", check.detail)
def test_requester_ack_never_satisfies_another_sessions_obligation(self):
"""The requester is excluded from the required set and stays excluded."""
report = _identity_report(requester="req", others=("other-0",))
requester_rows = [s for s in report["affected_sessions"] if s["is_requester"]]
self.assertEqual([s["session_id"] for s in requester_rows], ["req"])
self.assertNotIn("req", report["ack_state"])
check = self.assertAcksFailClosed(report, self._state(acks={"req": "ack"}))
self.assertIn("other-0", check.detail)
def test_fabricated_ids_do_not_count_toward_coverage(self):
report = _identity_report(requester="req", others=("other-0",))
for bogus in ("", " ", "other-0 extra", "OTHER-0", "other-01", "0"):
with self.subTest(bogus=bogus):
self.assertAcksFailClosed(report, self._state(acks={bogus: "ack"}))
# --- per-session state must be explicitly valid ------------------------
def test_unproven_per_session_states_fail_closed(self):
"""A required id present but not explicitly acknowledged fails closed."""
report = _identity_report(requester="req", others=("other-0", "other-1"))
for value in ("pending", "stale", "unknown", "", None, True, 1, NOW):
with self.subTest(value=value):
self.assertAcksFailClosed(
report,
self._state(acks={"other-0": "ack", "other-1": value}),
)
def test_report_ack_state_placeholder_is_never_read_as_an_ack(self):
"""``ack_state`` values are the report's own placeholders, not evidence."""
report = _identity_report(requester="req", others=("other-0",))
report["ack_state"] = {"other-0": "ack"}
self.assertAcksFailClosed(report, self._state())
# --- missing / malformed / contradictory identity evidence -------------
def test_missing_identity_evidence_fails_closed(self):
report = _identity_report(requester="req", others=("other-0",))
report.pop("ack_state", None)
report.pop("affected_sessions", None)
check = self.assertAcksFailClosed(report, self._state(acks={"other-0": "ack"}))
self.assertIn("no session-identity evidence", check.detail)
def test_malformed_ack_state_fails_closed(self):
for malformed in ([], "other-0", 7, None, ("other-0",)):
with self.subTest(malformed=malformed):
report = _identity_report(requester="req", others=("other-0",))
report["ack_state"] = malformed
self.assertAcksFailClosed(
report, self._state(acks={"other-0": "ack"})
)
def test_non_string_ack_state_key_fails_closed(self):
report = _identity_report(requester="req", others=("other-0",))
report["ack_state"] = {7: "pending"}
self.assertAcksFailClosed(report, self._state(acks={"other-0": "ack"}))
def test_malformed_affected_sessions_fails_closed(self):
for malformed in ("sessions", 7, {"session_id": "other-0"}, [None], [7]):
with self.subTest(malformed=malformed):
report = _identity_report(requester="req", others=("other-0",))
report.pop("ack_state", None)
report["affected_sessions"] = malformed
self.assertAcksFailClosed(
report, self._state(acks={"other-0": "ack"})
)
def test_affected_sessions_without_explicit_booleans_fails_closed(self):
"""``live``/``is_requester`` must be real booleans, never inferred."""
report = _identity_report(requester="req", others=("other-0",))
report.pop("ack_state", None)
for row in report["affected_sessions"]:
if row["session_id"] == "other-0":
row["is_requester"] = "false"
self.assertAcksFailClosed(report, self._state(acks={"other-0": "ack"}))
def test_affected_sessions_missing_live_flag_fails_closed(self):
report = _identity_report(requester="req", others=("other-0",))
report.pop("ack_state", None)
for row in report["affected_sessions"]:
row.pop("live", None)
self.assertAcksFailClosed(report, self._state(acks={"other-0": "ack"}))
def test_contradictory_ack_state_and_affected_sessions_fails_closed(self):
"""Both views present and disagreeing is unresolvable, not a tie-break."""
report = _identity_report(requester="req", others=("other-0", "other-1"))
report["ack_state"] = {"other-0": "pending", "other-9": "pending"}
check = self.assertAcksFailClosed(
report, self._state(acks={"other-0": "ack", "other-9": "ack"})
)
self.assertIn("contradicts itself", check.detail)
def test_identity_count_mismatch_fails_closed(self):
"""Identity evidence that cannot be reconciled with the count denies."""
report = _identity_report(requester="req", others=("other-0", "other-1"))
report["counts"] = dict(report["counts"], sessions_live_other=1)
check = self.assertAcksFailClosed(
report, self._state(acks={"other-0": "ack", "other-1": "ack"})
)
self.assertIn("cannot be reconciled", check.detail)
def test_broken_identity_evidence_outranks_timeout_policy(self):
"""The sanctioned timeout path cannot paper over an unreadable report."""
report = _identity_report(requester="req", others=("other-0",))
report["ack_state"] = "not-a-mapping"
self.assertAcksFailClosed(report, self._state(ack_timeout_policy_applied=True))
# --- legitimate success is preserved -----------------------------------
def test_every_required_session_acknowledged_passes(self):
report = _identity_report(
requester="req", others=("other-0", "other-1", "other-2")
)
state = self._state(
acks={
"other-0": "ack",
"other-1": "acked",
"other-2": "acknowledged",
}
)
proof = self._build(report, state)
check = self._acks_check(proof)
self.assertTrue(check.passed)
self.assertTrue(proof.clean)
self.assertEqual(proof.failed_checks, [])
self.assertIn("acknowledged by identity", check.detail)
decision = dp.gate_apply_restart(
proof=proof.as_dict(),
now=NOW,
secret=SECRET,
expected_impact_fingerprint=dp.impact_fingerprint(report),
)
self.assertEqual(decision.verdict, dp.GATE_ALLOW)
self.assertTrue(decision.allow)
def test_required_session_ack_tolerates_surrounding_whitespace(self):
report = _identity_report(requester="req", others=("other-0",))
proof = self._build(report, self._state(acks={" other-0 ": " ACK "}))
self.assertTrue(self._acks_check(proof).passed)
self.assertTrue(proof.clean)
def test_no_other_live_sessions_still_passes_with_identity_evidence(self):
report = _identity_report(requester="req", others=())
self.assertEqual(report["counts"]["sessions_live_other"], 0)
self.assertEqual(report["ack_state"], {})
proof = self._build(report, self._state())
check = self._acks_check(proof)
self.assertTrue(check.passed)
self.assertIn("sessions_live_other=0", check.detail)
self.assertTrue(proof.clean)
def test_explicit_timeout_policy_retains_intended_behavior(self):
"""Valid, correctly typed timeout evidence still permits the check."""
report = _identity_report(requester="req", others=("other-0", "other-1"))
proof = self._build(report, self._state(ack_timeout_policy_applied=True))
check = self._acks_check(proof)
self.assertTrue(check.passed)
self.assertIn("timeout policy", check.detail)
self.assertTrue(proof.clean)
def test_timeout_policy_still_strictly_typed_under_identity_binding(self):
report = _identity_report(requester="req", others=("other-0",))
for value in (None, "true", "True", 1, [], {}):
with self.subTest(value=value):
self.assertAcksFailClosed(
report, self._state(ack_timeout_policy_applied=value)
)
if __name__ == "__main__":
unittest.main()
-199
View File
@@ -1,199 +0,0 @@
"""Regression tests for #628 building blocks (child scope only).
Honest scope: unit coverage of pre-existing APIs used by umbrella #628.
This module does **not** implement or verify all 21 umbrella acceptance
criteria, automatic handoff store/retrieve, multi-worker product wiring,
or end-to-end orchestration.
Covered building blocks:
- CTH format / parse / assess (`format_cth_body`, `parse_cth_comment`,
`assess_cth_comment`)
- Exclusive-ownership skip classification (`classify_skip` with
OWNERSHIP_FOREIGN vs OWNERSHIP_OWN)
- Durable dependency edges (`upsert_dependency_edge` / list) and skip
when dependency_unmet
- Edge state transition UNMET -> MET
Parent umbrella remains #628; this slice is a scoped child issue only.
"""
import unittest
from unittest.mock import MagicMock, patch
import os
import json
import tempfile
from canonical_thread_handoff import (
format_cth_body,
parse_cth_comment,
assess_cth_comment,
)
import dependency_graph
from control_plane_db import ControlPlaneDB
from allocator_service import (
WorkCandidate,
classify_skip,
ROLE_AUTHOR,
ROLE_REVIEWER,
ROLE_MERGER,
ROLE_RECONCILER,
OWNERSHIP_OWN,
OWNERSHIP_FOREIGN,
)
class TestIssue628Orchestration(unittest.TestCase):
def setUp(self):
self._tmp = tempfile.TemporaryDirectory()
self.db_path = os.path.join(self._tmp.name, "cp.sqlite3")
self.db = ControlPlaneDB(self.db_path)
def tearDown(self):
self._tmp.cleanup()
def test_canonical_handoff_serialization_and_retrieval(self):
"""Building block: format/parse/assess a CTH body (not full AC1/AC2 product path)."""
handoff = format_cth_body(
cth_type="Author Handoff",
status="completed",
next_owner="reviewer",
current_blocker="none",
decision="Implementation complete, tests passing",
proof="pytest tests/test_issue_628_orchestration.py passed",
next_action="Review PR and run reviewer pre-flight",
ready_to_paste_prompt="Review PR for child issue #878 (parent #628)",
)
self.assertIn("CTH: Author Handoff", handoff)
parsed = parse_cth_comment(handoff)
self.assertIsNotNone(parsed)
self.assertEqual(parsed["cth_type"], "Author Handoff")
assessment = assess_cth_comment(handoff)
self.assertFalse(assessment["block"])
def test_exclusive_task_unit_single_owner(self):
"""Building block: classify_skip foreign vs own ownership."""
candidate = WorkCandidate(
kind="issue",
number=878,
title="Child #878 ownership classify candidate",
state="open",
labels=["status:in-progress"],
blocked=False,
dependency_unmet=False,
)
# Foreign ownership MUST be skipped
skip_foreign = classify_skip(
c=candidate,
role=ROLE_AUTHOR,
terminal_pr=None,
claim_ownership=OWNERSHIP_FOREIGN,
)
self.assertIsNotNone(skip_foreign)
self.assertIn("active lease", skip_foreign)
# Own/Self claim remains selectable for session resumption
skip_self = classify_skip(
c=candidate,
role=ROLE_AUTHOR,
terminal_pr=None,
claim_ownership=OWNERSHIP_OWN,
)
self.assertIsNone(skip_self)
def test_durable_dependency_graph_blocking(self):
"""Building block: unmet dependency edge + classify_skip on dependency_unmet."""
# Upsert a blocking dependency edge between issue 878 and blocker 601
self.db.upsert_dependency_edge(
remote="prgs",
org="Scaled-Tech-Consulting",
repo="Gitea-Tools",
source_kind="issue",
source_number=878,
target_kind="issue",
target_number=601,
edge_type=dependency_graph.EDGE_ISSUE_BLOCKED_BY_ISSUE,
state=dependency_graph.STATE_UNMET,
blocking_condition="Target issue #601 is not closed",
completion_condition="Target issue #601 is closed",
evidence={"source": "unit_test"},
)
edges = self.db.list_dependency_edges(
remote="prgs",
org="Scaled-Tech-Consulting",
repo="Gitea-Tools",
source_kind="issue",
source_number=878,
)
self.assertEqual(len(edges), 1)
self.assertEqual(edges[0]["state"], "unmet")
self.assertEqual(edges[0]["target_number"], 601)
# When dependency is unmet, candidate is blocked from selection
candidate = WorkCandidate(
kind="issue",
number=878,
title="Blocked candidate",
state="open",
labels=[],
blocked=False,
dependency_unmet=True,
dependency_reason="issue#878 is blocked by unmet dependency issue#601",
)
skip_reason = classify_skip(
c=candidate,
role=ROLE_AUTHOR,
terminal_pr=None,
claim_ownership=OWNERSHIP_OWN,
)
self.assertIsNotNone(skip_reason)
self.assertIn("issue#601", skip_reason)
def test_dependency_completion_reevaluation(self):
"""Building block: dependency edge state can transition UNMET -> MET."""
self.db.upsert_dependency_edge(
remote="prgs",
org="Scaled-Tech-Consulting",
repo="Gitea-Tools",
source_kind="issue",
source_number=878,
target_kind="issue",
target_number=601,
edge_type=dependency_graph.EDGE_ISSUE_BLOCKED_BY_ISSUE,
state=dependency_graph.STATE_UNMET,
blocking_condition="Target issue #601 is open",
completion_condition="Target issue #601 is closed",
evidence={"source": "unit_test"},
)
# Mark edge as met upon target issue closure
self.db.upsert_dependency_edge(
remote="prgs",
org="Scaled-Tech-Consulting",
repo="Gitea-Tools",
source_kind="issue",
source_number=878,
target_kind="issue",
target_number=601,
edge_type=dependency_graph.EDGE_ISSUE_BLOCKED_BY_ISSUE,
state=dependency_graph.STATE_MET,
blocking_condition="Target issue #601 is open",
completion_condition="Target issue #601 is closed",
evidence={"source": "target_closed_event"},
)
edges = self.db.list_dependency_edges(
remote="prgs",
org="Scaled-Tech-Consulting",
repo="Gitea-Tools",
source_kind="issue",
source_number=878,
)
self.assertEqual(len(edges), 1)
self.assertEqual(edges[0]["state"], "met")
if __name__ == "__main__":
unittest.main()
@@ -1,249 +0,0 @@
"""Tests for post-restart MCP reconciliation and completion proof (#662)."""
from __future__ import annotations
import json
import os
import unittest
from datetime import datetime, timezone
import post_restart_reconcile as prr
NOW = datetime(2026, 7, 24, 12, 0, 0, tzinfo=timezone.utc)
def _base_inventory(**overrides):
inv = {
"inventory_complete": True,
"incomplete_reasons": [],
"service_health": {"healthy": True},
"clients": [{"session_id": "c1", "connected": True}],
"sessions": [
{
"session_id": "s-live",
"status": "active",
"pid": os.getpid(),
"role": "author",
}
],
"leases": [],
"checkpoints_available": False,
"worktree_bindings": [{"path": "/tmp/wt", "exists": True}],
"pending_mutations": [],
"capabilities": {"stale": False},
"boot_head_sha": "a" * 40,
"current_head_sha": "a" * 40,
"queue_state": {"safe_to_resume": True},
}
inv.update(overrides)
return inv
class IncompleteInventoryTest(unittest.TestCase):
def test_incomplete_inventory_fails_closed(self) -> None:
proof = prr.reconcile_after_restart(
{
"inventory_complete": False,
"incomplete_reasons": ["control-plane DB unavailable"],
},
now=NOW,
mode=prr.MODE_ENFORCE,
reconcile_id="test-incomplete",
)
self.assertEqual(proof.overall_status, prr.STATUS_FAILED)
self.assertTrue(proof.mutation_hold)
self.assertFalse(proof.inventory_complete)
self.assertTrue(proof.proposed_follow_ups)
self.assertIn("control-plane DB unavailable", proof.incomplete_reasons)
class HappyPathTest(unittest.TestCase):
def test_clean_restart_is_complete_without_mutation_hold(self) -> None:
proof = prr.reconcile_after_restart(
_base_inventory(),
now=NOW,
mode=prr.MODE_ENFORCE,
reconcile_id="test-clean",
)
self.assertEqual(proof.overall_status, prr.STATUS_COMPLETE)
self.assertFalse(proof.mutation_hold)
self.assertEqual(proof.unresolved_count, 0)
cp = next(i for i in proof.items if i.dimension == prr.DIM_CHECKPOINTS)
self.assertEqual(cp.status, prr.ITEM_SKIPPED)
links = proof.as_dict()["links"]
self.assertEqual(links["umbrella"], 655)
self.assertEqual(links["issue"], 662)
self.assertEqual(links["vision"], 652)
self.assertEqual(links["roadmap"], 653)
class InterruptedMutationTest(unittest.TestCase):
def test_mutating_lease_with_dead_owner_is_unresolved(self) -> None:
proof = prr.reconcile_after_restart(
_base_inventory(
leases=[
{
"lease_id": "lease-mut",
"session_id": "s-dead",
"phase": "implementing",
"work_kind": "issue",
"work_number": 662,
"worktree_path": "/tmp/wt-662",
"freshness": {"freshness": "stale_dead_process"},
}
]
),
now=NOW,
mode=prr.MODE_ENFORCE,
)
mut = next(i for i in proof.items if i.dimension == prr.DIM_MUTATIONS)
self.assertEqual(mut.status, prr.ITEM_UNRESOLVED)
self.assertTrue(mut.follow_up_required)
interrupted = mut.details["interrupted"]
self.assertEqual(len(interrupted), 1)
self.assertFalse(interrupted[0]["resume_allowed"])
self.assertTrue(proof.mutation_hold)
self.assertTrue(
any(f.dimension == prr.DIM_MUTATIONS for f in proof.proposed_follow_ups)
)
def test_explicit_pending_mutation_inventory(self) -> None:
proof = prr.reconcile_after_restart(
_base_inventory(
pending_mutations=[
{
"session_id": "s1",
"phase": "publishing",
"work_kind": "pr",
"work_number": 856,
"reason": "push interrupted mid-flight",
}
]
),
now=NOW,
mode=prr.MODE_LOG_ONLY,
)
mut = next(i for i in proof.items if i.dimension == prr.DIM_MUTATIONS)
self.assertEqual(mut.status, prr.ITEM_UNRESOLVED)
# log_only never holds mutations even when unresolved
self.assertFalse(proof.mutation_hold)
self.assertEqual(proof.overall_status, prr.STATUS_DEGRADED)
class DuplicateClaimsTest(unittest.TestCase):
def test_duplicate_live_claims_flagged(self) -> None:
proof = prr.reconcile_after_restart(
_base_inventory(
leases=[
{
"lease_id": "l1",
"session_id": "s1",
"phase": "allocated",
"work_kind": "issue",
"work_number": 100,
"freshness": {"freshness": "active"},
},
{
"lease_id": "l2",
"session_id": "s2",
"phase": "allocated",
"work_kind": "issue",
"work_number": 100,
"freshness": {"freshness": "active"},
},
]
),
now=NOW,
mode=prr.MODE_ENFORCE,
)
dups = next(i for i in proof.items if i.dimension == prr.DIM_DUPLICATES)
self.assertEqual(dups.status, prr.ITEM_UNRESOLVED)
self.assertEqual(dups.details["duplicates"][0]["claim_count"], 2)
self.assertTrue(proof.mutation_hold)
class OrphanSessionTest(unittest.TestCase):
def test_active_session_dead_pid_is_unresolved(self) -> None:
proof = prr.reconcile_after_restart(
_base_inventory(
sessions=[
{
"session_id": "ghost",
"status": "active",
"pid": 2_000_000_000,
"role": "author",
}
]
),
now=NOW,
mode=prr.MODE_ENFORCE,
)
sess = next(i for i in proof.items if i.dimension == prr.DIM_SESSIONS)
self.assertEqual(sess.status, prr.ITEM_UNRESOLVED)
self.assertIn("ghost", sess.details["orphan_session_ids"])
class CapabilityStaleTest(unittest.TestCase):
def test_stale_runtime_unresolved(self) -> None:
proof = prr.reconcile_after_restart(
_base_inventory(capabilities={"stale": True, "startup_head": "aaa"}),
now=NOW,
mode=prr.MODE_ENFORCE,
)
caps = next(i for i in proof.items if i.dimension == prr.DIM_CAPABILITIES)
self.assertEqual(caps.status, prr.ITEM_UNRESOLVED)
self.assertTrue(proof.mutation_hold)
class MutationsAllowedHelperTest(unittest.TestCase):
def test_mutations_allowed_respects_hold(self) -> None:
held = prr.reconcile_after_restart(
_base_inventory(
pending_mutations=[{"phase": "merging", "session_id": "x"}]
),
now=NOW,
mode=prr.MODE_ENFORCE,
)
self.assertFalse(prr.mutations_allowed(held))
self.assertFalse(prr.mutations_allowed(held.as_dict()))
clean = prr.reconcile_after_restart(
_base_inventory(), now=NOW, mode=prr.MODE_ENFORCE
)
self.assertTrue(prr.mutations_allowed(clean))
class CheckpointSoftDependencyTest(unittest.TestCase):
def test_checkpoints_when_schema_present(self) -> None:
proof = prr.reconcile_after_restart(
_base_inventory(
checkpoints_available=True,
checkpoints=[{"session_id": "s1", "stale": False}],
),
now=NOW,
)
cp = next(i for i in proof.items if i.dimension == prr.DIM_CHECKPOINTS)
self.assertEqual(cp.status, prr.ITEM_RESOLVED)
def test_stale_checkpoints_unresolved(self) -> None:
proof = prr.reconcile_after_restart(
_base_inventory(
checkpoints_available=True,
checkpoints=[{"session_id": "s1", "stale": True}],
),
now=NOW,
mode=prr.MODE_ENFORCE,
)
cp = next(i for i in proof.items if i.dimension == prr.DIM_CHECKPOINTS)
self.assertEqual(cp.status, prr.ITEM_UNRESOLVED)
class ProofSerializationTest(unittest.TestCase):
def test_as_dict_is_json_friendly(self) -> None:
proof = prr.reconcile_after_restart(_base_inventory(), now=NOW)
blob = json.dumps(proof.as_dict())
self.assertIn("reconcile_id", blob)
self.assertIn("proposed_follow_ups", blob)
if __name__ == "__main__":
unittest.main()
@@ -1,323 +0,0 @@
"""Tests for sanctioned Codex MCP reconnect request surface (#678)."""
from __future__ import annotations
import os
import unittest
from unittest import mock
import mcp_client_reconnect as mcr
class NormalizeReasonTests(unittest.TestCase):
def test_stale_runtime_aliases(self):
self.assertEqual(mcr.normalize_reason("stale-runtime"), mcr.REASON_STALE_RUNTIME)
self.assertEqual(mcr.normalize_reason("stale_runtime"), mcr.REASON_STALE_RUNTIME)
self.assertEqual(mcr.normalize_reason("STALE"), mcr.REASON_STALE_RUNTIME)
def test_transport_eof_aliases(self):
self.assertEqual(mcr.normalize_reason("transport_eof"), mcr.REASON_TRANSPORT_EOF)
self.assertEqual(mcr.normalize_reason("EOF"), mcr.REASON_TRANSPORT_EOF)
self.assertEqual(
mcr.normalize_reason("client_is_closing"), mcr.REASON_TRANSPORT_EOF
)
def test_missing_namespace(self):
self.assertEqual(
mcr.normalize_reason("missing_namespace"), mcr.REASON_MISSING_NAMESPACE
)
def test_empty_is_unspecified(self):
self.assertEqual(mcr.normalize_reason(None), mcr.REASON_UNSPECIFIED)
self.assertEqual(mcr.normalize_reason(""), mcr.REASON_UNSPECIFIED)
class BoundaryClassificationTests(unittest.TestCase):
def test_clean_when_shas_match(self):
self.assertEqual(
mcr.classify_boundary_status(
startup_sha="abc", current_master_sha="abc"
),
mcr.BOUNDARY_CLEAN,
)
def test_mismatch_when_shas_differ(self):
self.assertEqual(
mcr.classify_boundary_status(
startup_sha="aaa", current_master_sha="bbb"
),
mcr.BOUNDARY_MISMATCH,
)
def test_stale_when_live_stale(self):
self.assertEqual(
mcr.classify_boundary_status(
startup_sha="aaa",
current_master_sha="aaa",
live_stale=True,
),
mcr.BOUNDARY_STALE,
)
class BuildReconnectRequestTests(unittest.TestCase):
def test_stale_runtime_returns_typed_blocker_with_codex_steps(self):
result = mcr.build_reconnect_request(
namespace="gitea-author",
profile="prgs-author",
pid=1234,
session_id="sess-1",
startup_sha="aaa111",
current_master_sha="bbb222",
reason="stale-runtime",
client="codex",
restart_required=True,
stop_required=True,
)
self.assertTrue(result["success"])
self.assertTrue(result["read_only"])
self.assertFalse(result["reconnect_performed"])
self.assertFalse(result["mutation_performed"])
self.assertTrue(result["reconnect_needed"])
self.assertEqual(result["namespace"], "gitea-author")
self.assertEqual(result["profile"], "prgs-author")
self.assertEqual(result["pid"], 1234)
self.assertEqual(result["session_id"], "sess-1")
self.assertEqual(result["startup_sha"], "aaa111")
self.assertEqual(result["current_master_sha"], "bbb222")
self.assertEqual(result["boundary_status"], mcr.BOUNDARY_MISMATCH)
self.assertEqual(result["blocker_kind"], mcr.BLOCKER_OPERATOR_RECONNECT)
self.assertIsNotNone(result["typed_blocker"])
blocker = result["typed_blocker"]
self.assertEqual(blocker["namespaces"], ["gitea-author"])
self.assertEqual(blocker["why_reconnect_required"], mcr.REASON_STALE_RUNTIME)
self.assertTrue(any("Codex" in s or "Reload" in s for s in blocker["operator_ui_steps"]))
self.assertIn("pkill", " ".join(result["forbidden_recovery_paths"]).lower())
self.assertTrue(
mcr.reasons_never_suggest_forbidden(result["exact_safe_next_action"] or "")
)
# Must not recommend forbidden recovery.
for step in blocker["operator_ui_steps"]:
self.assertTrue(mcr.reasons_never_suggest_forbidden(step), step)
def test_transport_eof_typed_blocker(self):
result = mcr.build_reconnect_request(
namespace="gitea-reviewer",
reason="transport_eof",
client="claude_code",
)
self.assertTrue(result["reconnect_needed"])
self.assertEqual(result["reason"], mcr.REASON_TRANSPORT_EOF)
self.assertEqual(result["client"], "claude_code")
steps = " ".join(result["operator_ui_steps"]).lower()
self.assertIn("/mcp", steps)
def test_missing_namespace_typed_blocker(self):
result = mcr.build_reconnect_request(
namespace="gitea-merger",
reason="missing_namespace",
client="codex",
)
self.assertTrue(result["reconnect_needed"])
self.assertEqual(result["reason"], mcr.REASON_MISSING_NAMESPACE)
self.assertEqual(
result["typed_blocker"]["blocker_kind"], mcr.BLOCKER_OPERATOR_RECONNECT
)
def test_healthy_not_required(self):
result = mcr.build_reconnect_request(
namespace="gitea-tools",
startup_sha="deadbeef",
current_master_sha="deadbeef",
reason="not_required",
client="codex",
in_parity=True,
restart_required=False,
stop_required=False,
)
self.assertFalse(result["reconnect_needed"])
self.assertEqual(result["blocker_kind"], mcr.BLOCKER_NONE)
self.assertIsNone(result["typed_blocker"])
self.assertFalse(result["stop_required"])
self.assertFalse(result["restart_required"])
self.assertIn("not required", (result["exact_safe_next_action"] or "").lower())
def test_successful_reconnect_report_fields_present(self):
"""AC2: reconnect result reports required fields (even when needed)."""
result = mcr.build_reconnect_request(
namespace="gitea-controller",
profile="prgs-controller",
pid=99,
session_id="sid",
startup_sha="s" * 40,
current_master_sha="c" * 40,
reason="stale-runtime",
)
for key in (
"namespace",
"profile",
"pid",
"session_id",
"startup_sha",
"current_master_sha",
"boundary_status",
):
self.assertIn(key, result)
self.assertIsNotNone(result[key], key)
class ToolSurfaceTests(unittest.TestCase):
"""Exercise gitea_request_mcp_reconnect with a stubbed server context."""
def test_tool_is_registered_and_side_effect_free(self):
import gitea_mcp_server as srv
self.assertTrue(hasattr(srv, "gitea_request_mcp_reconnect"))
with mock.patch.object(srv, "_profile_operation_gate", return_value=None):
with mock.patch.object(
srv,
"get_profile",
return_value={
"profile_name": "prgs-author",
"role_kind": "author",
"role": "author",
},
):
with mock.patch.object(
srv,
"_current_master_parity",
return_value={
"startup_head": "a" * 40,
"current_head": "a" * 40,
"daemon_start_head": "a" * 40,
"local_head": "a" * 40,
"in_parity": True,
"stale": False,
"restart_required": False,
"determinable": True,
"live_stale": False,
"live_known": True,
"reasons": [],
},
):
with mock.patch.object(
srv.master_parity_gate,
"format_parity",
return_value="in parity",
):
with mock.patch.object(
srv.role_namespace_gate,
"infer_mcp_namespace",
return_value="gitea-author",
):
with mock.patch.object(
srv.session_ctx,
"mutation_context_audit_fields",
return_value={"session_profile": "prgs-author"},
):
result = srv.gitea_request_mcp_reconnect(
namespace="gitea-author",
reason="not_required",
client="codex",
remote="prgs",
)
self.assertTrue(result.get("success"))
self.assertFalse(result.get("reconnect_performed"))
self.assertFalse(result.get("mutation_performed"))
self.assertEqual(result.get("namespace"), "gitea-author")
self.assertEqual(result.get("profile"), "prgs-author")
self.assertEqual(result.get("pid"), os.getpid())
self.assertIn("startup_sha", result)
self.assertIn("current_master_sha", result)
self.assertIn("boundary_status", result)
self.assertTrue(
mcr.reasons_never_suggest_forbidden(
result.get("exact_safe_next_action") or ""
)
)
def test_tool_stale_returns_typed_blocker(self):
import gitea_mcp_server as srv
with mock.patch.object(srv, "_profile_operation_gate", return_value=None):
with mock.patch.object(
srv,
"get_profile",
return_value={
"profile_name": "prgs-reconciler",
"role_kind": "reconciler",
"role": "reconciler",
},
):
with mock.patch.object(
srv,
"_current_master_parity",
return_value={
"startup_head": "a" * 40,
"current_head": "b" * 40,
"daemon_start_head": "a" * 40,
"local_head": "b" * 40,
"in_parity": False,
"stale": True,
"restart_required": True,
"determinable": True,
"live_stale": True,
"live_known": True,
"reasons": ["stale"],
},
):
with mock.patch.object(
srv.master_parity_gate,
"format_parity",
return_value="stale",
):
with mock.patch.object(
srv.role_namespace_gate,
"infer_mcp_namespace",
return_value="gitea-reconciler",
):
with mock.patch.object(
srv.session_ctx,
"mutation_context_audit_fields",
return_value={},
):
result = srv.gitea_request_mcp_reconnect(
reason="stale-runtime",
client="codex",
)
self.assertTrue(result["reconnect_needed"])
self.assertEqual(
result["blocker_kind"], mcr.BLOCKER_OPERATOR_RECONNECT
)
self.assertIsNotNone(result["typed_blocker"])
self.assertIn("gitea-reconciler", result["typed_blocker"]["namespaces"])
self.assertTrue(result["stop_required"])
self.assertTrue(result["restart_required"])
self.assertTrue(
mcr.reasons_never_suggest_forbidden(
result.get("exact_safe_next_action") or ""
)
)
class InventoryRegistrationTests(unittest.TestCase):
def test_reconnect_path_in_restart_inventory(self):
import mcp_restart_paths as mrp
ids = {p.path_id for p in mrp.iter_restart_paths()}
self.assertIn("codex_client_reconnect_request", ids)
self.assertIn("ide_client_reconnect", ids)
def test_tool_name_in_documented_inventory(self):
import mcp_tool_inventory as inv
doc_path = os.path.join(
os.path.dirname(os.path.dirname(__file__)), inv.INVENTORY_DOC_PATH
)
with open(doc_path, encoding="utf-8") as handle:
documented = inv.parse_documented_inventory(handle.read())
self.assertIn("gitea_request_mcp_reconnect", documented)
if __name__ == "__main__":
unittest.main()
-72
View File
@@ -1,72 +0,0 @@
"""Starlette TestClient uses httpx2 without deprecation warning (#682)."""
from __future__ import annotations
import importlib
import sys
import unittest
import warnings
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
class TestHttpx2Installed(unittest.TestCase):
def test_httpx2_importable(self) -> None:
httpx2 = importlib.import_module("httpx2")
self.assertTrue(hasattr(httpx2, "Client") or hasattr(httpx2, "AsyncClient"))
class TestNoStarletteTestClientDeprecation(unittest.TestCase):
def test_starlette_testclient_import_does_not_warn(self) -> None:
# Force a fresh import path so a previous httpx-fallback warning cannot
# hide a regression after httpx2 is present. Restore sys.modules afterwards.
saved = {
name: sys.modules[name]
for name in list(sys.modules)
if name == "starlette.testclient" or name.startswith("starlette.testclient.")
}
try:
for name in list(saved):
del sys.modules[name]
with warnings.catch_warnings(record=True) as caught:
warnings.simplefilter("always")
import starlette.testclient as testclient # noqa: F401
dep = [
w
for w in caught
if "httpx" in str(w.message).lower()
and "deprecated" in str(w.message).lower()
]
self.assertEqual(
dep,
[],
f"expected no Starlette httpx deprecation warning; got: "
f"{[str(w.message) for w in dep]}",
)
finally:
sys.modules.update(saved)
def test_canonical_webui_testclient_import(self) -> None:
with warnings.catch_warnings(record=True) as caught:
warnings.simplefilter("always")
from tests.webui_testclient import TestClient
from webui.app import create_app
client = TestClient(create_app())
response = client.get("/")
self.assertIn(response.status_code, {200, 302, 404})
dep = [
w
for w in caught
if "httpx" in str(w.message).lower()
and "deprecated" in str(w.message).lower()
]
self.assertEqual(dep, [], [str(w.message) for w in dep])
if __name__ == "__main__":
unittest.main()

Some files were not shown because too many files have changed in this diff Show More