Compare commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
0088ecaf00 | ||
|
|
89657a06c4 | ||
|
|
ba3ea3012c | ||
|
|
ef14622ba0 | ||
|
|
a3f8f67c93 | ||
|
|
6d0015cabc | ||
|
|
dc99c15ffa | ||
|
|
9f759150b8 | ||
|
|
347464a057 | ||
|
|
1301a57de4 |
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,223 @@
|
||||
# ADR: MCP restart governance and authorization policy
|
||||
|
||||
- **Status:** Accepted (policy effective immediately for LLM and operator sessions; enforcement tooling may lag)
|
||||
- **Date:** 2026-07-23
|
||||
- **Tracking issue:** [#656](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/656)
|
||||
- **Policy version:** `restart-governance/v1`
|
||||
- **Related:**
|
||||
- Umbrella: [#655](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/655) — governed MCP restart coordination and zero-disruption recovery
|
||||
- Vision: [#652](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/652) — MCP Control Plane Web Console product vision (§A system health and process control)
|
||||
- Roadmap: [#653](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/653) — Control Plane Web Console phased delivery (Phase 2 restart controls)
|
||||
- Contamination guard: [#630](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/630) — blocks manual process-kill recovery
|
||||
- Console restart UX: [#642](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/642) — sanctioned restart and graceful reload
|
||||
- Existing restart / reconnect paths to inventory: [#591](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/591) — auto-restart on master advance (closed); [#584](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/584) — host auto-reconnect on transport flap
|
||||
- Stable-control runtime split: `docs/architecture/mcp-stable-control-runtime-policy-adr.md` (#615)
|
||||
- Client-namespace health: `docs/mcp-namespace-health.md` (#543)
|
||||
- Reconnect-only EOF recovery: `docs/mcp-namespace-eof-recovery.md`
|
||||
|
||||
## 1. Context
|
||||
|
||||
The Gitea MCP server is the **control plane** for real issue and PR mutations
|
||||
(create, comment, lock, review, merge, reconcile). The same process serves every
|
||||
role namespace (`gitea-author`, `gitea-reviewer`, `gitea-merger`,
|
||||
`gitea-reconciler`, `gitea-controller`) and holds the in-memory capability-gate
|
||||
code loaded at startup.
|
||||
|
||||
Restarting that process is destructive to concurrent work:
|
||||
|
||||
- It resets every session's identity, preflight, and capability-lease binding.
|
||||
- It can interrupt a mutation mid-critical-section (a lock acquire, a review
|
||||
submit, a merge), leaving durable state half-written.
|
||||
- Relaunching from the wrong checkout or worktree silently changes which code
|
||||
the control plane runs, defeating master-parity gates (#420 / #615).
|
||||
|
||||
Today there is **no durable written policy** stating who may restart MCP, under
|
||||
what conditions, that restart is a last resort, and how controller approval,
|
||||
automated safety gates, and break-glass interact. Operators and LLM sessions
|
||||
therefore invent restart behavior ad hoc, which makes concurrent multi-role work
|
||||
unsafe. #630 and #642 need this policy as their backbone.
|
||||
|
||||
This ADR defines that policy. It does **not** implement coordinator code or HA
|
||||
multi-instance restart (those are later children of #655).
|
||||
|
||||
## 2. Decision
|
||||
|
||||
### 2.1 v1 decision (recorded)
|
||||
|
||||
**Restart authority in v1 is `controller approval + automated safety gates`.**
|
||||
|
||||
A restart of the stable control runtime is authorized only when **both** hold:
|
||||
|
||||
1. A **controller** role explicitly approves the restart, recording an audit
|
||||
entry (who, why, scope, affected sessions), **and**
|
||||
2. The **automated safety gates** pass: a completed drain acknowledgement (no
|
||||
affected session is mid-critical-section) or a declared break-glass incident
|
||||
(§2.5).
|
||||
|
||||
Quorum among multiple controllers is **not** required day-one. It is deferred
|
||||
unless a later investigation (tracked under #653) proves single-controller
|
||||
approval is insufficient. This ADR records the v1 decision so enforcement code
|
||||
(#630) has a fixed target; changing it requires a superseding ADR.
|
||||
|
||||
### 2.2 Restart is a last resort — the recovery ladder
|
||||
|
||||
Restart is the **last** rung. Before any restart, exhaust the narrower
|
||||
recoveries, in order:
|
||||
|
||||
1. **Reconnect** the IDE/client MCP namespace (transport EOF, `client is
|
||||
closing: EOF`, transient `#584` flap). No process change. See
|
||||
`docs/mcp-namespace-eof-recovery.md`.
|
||||
2. **Refresh / rebind** the session workspace: re-run `gitea_whoami`,
|
||||
`gitea_resolve_task_capability`, and pass an explicit validated
|
||||
`worktree_path`. Fixes stale session context without touching the process.
|
||||
3. **Scoped restart** of a single misbehaving namespace/service (where the
|
||||
deployment supports per-service restart) rather than the whole control plane.
|
||||
4. **Full restart** of the stable control runtime process — operator-owned,
|
||||
controller-approved, drained.
|
||||
5. **Host / infrastructure restart** — the broadest action; same authorization
|
||||
as a full restart plus infrastructure ownership.
|
||||
|
||||
A session **must** try rungs 1–2 and record why they were insufficient before
|
||||
requesting a restart at rung 3 or above. Skipping straight to restart is a
|
||||
policy violation.
|
||||
|
||||
### 2.3 Authorization matrix
|
||||
|
||||
| Role | Reconnect (1) | Refresh/rebind (2) | Scoped restart (3) | Full restart (4) | Host restart (5) |
|
||||
|---|---|---|---|---|---|
|
||||
| **author** | self | self | request only | **forbidden** | forbidden |
|
||||
| **reviewer** | self | self | request only | **forbidden** | forbidden |
|
||||
| **merger** | self | self | request only | **forbidden** | forbidden |
|
||||
| **reconciler** | self | self | request only | **forbidden** | forbidden |
|
||||
| **controller** | self | self | **approve** (+gates) | **approve** (+gates) | request to operator |
|
||||
| **operator** | self | self | execute (controller-approved) | execute (controller-approved) | execute (controller-approved) |
|
||||
| **admin** | self | self | execute | execute | execute (break-glass) |
|
||||
|
||||
Legend: *self* = may perform for its own client session; *request only* = may
|
||||
raise a restart request but not authorize or execute it; *approve* = may
|
||||
authorize under §2.1 gates; *execute* = may perform the process action after the
|
||||
authorization is recorded.
|
||||
|
||||
Key invariants:
|
||||
|
||||
- **No LLM worker role (author/reviewer/merger/reconciler) may perform or
|
||||
authorize a full or host restart.** They may only reconnect/rebind their own
|
||||
client and file a restart request.
|
||||
- **Controller approval authorizes; operator/admin executes.** The approving
|
||||
controller and the executing operator may be the same human, but both the
|
||||
approval and the execution are audited.
|
||||
- Privileged process actions (full restart, host restart) are reserved to
|
||||
**operator/admin**, never to an automated worker.
|
||||
|
||||
### 2.4 Approved conditions
|
||||
|
||||
A restart at rung 3+ is approved only under one of these recorded conditions:
|
||||
|
||||
- **No affected sessions:** the control plane has no live session that would be
|
||||
interrupted (verified, not assumed).
|
||||
- **Full drain acknowledged:** every affected session has drained
|
||||
(no open critical section — no held mutation lease mid-write) and the drain is
|
||||
acknowledged in the audit record.
|
||||
- **Controller + gates:** controller approval plus passing automated safety
|
||||
gates (§2.1), the standard v1 path.
|
||||
- **Quorum:** not required in v1; reserved for a future superseding ADR.
|
||||
- **Break-glass:** an incident-backed emergency exception (§2.5).
|
||||
|
||||
Restart **never** bypasses mutation gates mid-critical-section. Drain before
|
||||
restart is mandatory except under break-glass with a declared incident.
|
||||
|
||||
### 2.5 Break-glass
|
||||
|
||||
Break-glass is a **separate, narrower** authorization path for emergencies where
|
||||
the normal drain-and-approve path cannot complete (e.g. the control plane is
|
||||
wedged and cannot drain).
|
||||
|
||||
Break-glass conditions:
|
||||
|
||||
- A declared incident record exists (id, timestamp, declarer) **before** the
|
||||
action.
|
||||
- The action is taken by **operator or admin** authority only — never by an LLM
|
||||
worker role, and never unilaterally by an operator with active peers when a
|
||||
controller is reachable.
|
||||
- The scope is the minimum necessary rung of the ladder.
|
||||
- A **mandatory post-hoc audit** entry is filed: what was restarted, why the
|
||||
normal path was impossible, which sessions were affected, and the incident id.
|
||||
|
||||
Break-glass suspends the drain requirement, not the audit requirement.
|
||||
|
||||
### 2.6 Explicit prohibitions
|
||||
|
||||
- **A unilateral LLM or operator full restart while active peer sessions
|
||||
exist is forbidden.** An LLM worker role must not kill, restart, or relaunch
|
||||
the MCP process; a lone operator must not full-restart over live peer work
|
||||
without controller approval or a break-glass incident.
|
||||
- Process-kill recovery is forbidden as a routine tool (#630). This ADR does not
|
||||
introduce a kill path.
|
||||
- Ambiguous policy state **denies** restart (§4).
|
||||
|
||||
## 3. Security requirements
|
||||
|
||||
- Full restart and host restart are **privileged**; only operator/admin execute
|
||||
them, only after a controller approval or break-glass incident is recorded.
|
||||
- Break-glass is a distinct authorization path with its own audit mandate; it is
|
||||
never the default and never silent.
|
||||
- **Every approval and every restart action is audited** (who approved, who
|
||||
executed, scope, affected sessions, condition, policy version). No restart is
|
||||
authorized without a durable audit entry.
|
||||
|
||||
## 4. Failure behavior
|
||||
|
||||
**Ambiguous policy → deny restart.** If it cannot be established that a
|
||||
restart is authorized under §2 — unknown affected-session state, missing
|
||||
controller approval, absent break-glass incident, or an unclassifiable request —
|
||||
the safe action is to **refuse** the restart and stop with a recovery report,
|
||||
never to restart on assumption.
|
||||
|
||||
## 5. Policy IDs (for enforcement code)
|
||||
|
||||
Enforcement code — the restart coordinator (a later child of #655), the #630
|
||||
contamination guard, and the #642 console restart UX — binds to these stable
|
||||
policy identifiers rather than to prose:
|
||||
|
||||
| Policy ID | Statement |
|
||||
|---|---|
|
||||
| `RG-01` | Restart is last resort; rungs 1–2 must be tried and recorded first (§2.2). |
|
||||
| `RG-02` | v1 authority = controller approval + automated safety gates (§2.1). |
|
||||
| `RG-03` | No LLM worker role performs or authorizes full/host restart (§2.3). |
|
||||
| `RG-04` | Full/host restart executed by operator/admin only, post approval (§2.3). |
|
||||
| `RG-05` | Drain before restart is mandatory except break-glass with incident (§2.4). |
|
||||
| `RG-06` | Break-glass requires a pre-declared incident and post-hoc audit (§2.5). |
|
||||
| `RG-07` | Unilateral LLM/operator full restart with active peers is forbidden (§2.6). |
|
||||
| `RG-08` | Ambiguous policy state denies restart (§4). |
|
||||
|
||||
The `restart-governance/v1` **policy version** field is emitted on future
|
||||
restart audit events so approvals can be reconciled against the policy revision
|
||||
in force.
|
||||
|
||||
## 6. Dogfooding
|
||||
|
||||
Gitea-Tools governs its own MCP control plane by this policy. Author, reviewer,
|
||||
merger, and reconciler sessions operating on this repository use the recovery
|
||||
ladder (§2.2) — reconnect and rebind, never self-restart — and any real restart
|
||||
of the Gitea-Tools stable control runtime follows the controller-approval +
|
||||
drain path defined here.
|
||||
|
||||
## 7. Acceptance and cross-links
|
||||
|
||||
This ADR is the authoritative restart-governance policy. It **must** stay
|
||||
cross-linked from the safety model and the web-console deployment boundary:
|
||||
|
||||
- `docs/safety-model.md` § Process restart governance references this ADR.
|
||||
- `docs/webui-deployment.md` references this ADR for restart/reload disposition.
|
||||
|
||||
It is linked to its issue lineage — umbrella **#655**, vision **#652**, roadmap
|
||||
**#653**, contamination guard **#630**, and console restart UX **#642** — in
|
||||
§ Related above.
|
||||
|
||||
## 8. Non-goals
|
||||
|
||||
- Implementing the restart coordinator or approval state machine (#630, later
|
||||
children of #655).
|
||||
- Implementing HA multi-instance restart or quorum machinery.
|
||||
- Introducing any process-kill or auto-restart tool; existing auto-restart
|
||||
behavior must be inventoried before any new restart tool is enabled.
|
||||
@@ -46,3 +46,17 @@ If shell helpers are unavailable and MCP commit cannot run, stop with a recovery
|
||||
report (restart session, clear hung terminals, use MCP-native commit). See
|
||||
[`llm-workflow-runbooks.md`](llm-workflow-runbooks.md) § MCP-native commit path
|
||||
(#260) and agent temp artifact cleanup (#261).
|
||||
|
||||
## 7. Process restart governance
|
||||
|
||||
Restarting the MCP control-plane process is destructive to concurrent multi-role
|
||||
work and is governed by a dedicated policy. Restart is a **last resort** behind
|
||||
narrower recoveries (reconnect, rebind), full/host restart is reserved to
|
||||
operator/admin under **controller approval + automated safety gates**, a
|
||||
unilateral LLM or operator full restart with active peers is **forbidden**, and
|
||||
ambiguous policy state **denies** restart. Break-glass is a separate,
|
||||
incident-backed path with a mandatory audit.
|
||||
|
||||
See [`architecture/mcp-restart-governance.md`](architecture/mcp-restart-governance.md)
|
||||
(#656) for the authorization matrix, the recovery ladder, break-glass
|
||||
conditions, and the `RG-01`–`RG-08` policy IDs.
|
||||
|
||||
@@ -55,6 +55,15 @@ shipped to the browser.
|
||||
assumption paths, and the client-secret policy. Use it to verify an instance is
|
||||
configured for internal-only operation.
|
||||
|
||||
## Process restart / reload disposition
|
||||
|
||||
The console never exposes a restart or reload control; process restart of the
|
||||
MCP control-plane runtime is governed separately. Restart is a last resort behind
|
||||
reconnect/rebind, full restart is operator/admin-only under controller approval
|
||||
plus safety gates, and break-glass is an incident-backed path. See
|
||||
[`architecture/mcp-restart-governance.md`](architecture/mcp-restart-governance.md)
|
||||
(#656).
|
||||
|
||||
## Non-goals (MVP)
|
||||
|
||||
- Full SSO or session login in the UI
|
||||
|
||||
+278
-1
@@ -440,13 +440,32 @@ def _session_author_lock_worktree() -> str | None:
|
||||
|
||||
Used to derive the author mutation workspace when no explicit
|
||||
``worktree_path`` or env binding is provided. Never invents a path.
|
||||
|
||||
#864: a session pointer whose owner PID is dead and is not this process
|
||||
must not force workspace binding for other issues — rebind is required for
|
||||
that issue, and a stale dead-owner pointer must not poison unrelated work.
|
||||
"""
|
||||
try:
|
||||
lock = issue_lock_store.read_session_issue_lock() or {}
|
||||
except Exception:
|
||||
return None
|
||||
path = (lock.get("worktree_path") or "").strip()
|
||||
return path or None
|
||||
if not path:
|
||||
return None
|
||||
pid = lock.get("session_pid")
|
||||
if pid is None:
|
||||
pid = lock.get("pid")
|
||||
try:
|
||||
pid_i = int(pid) if pid is not None else None
|
||||
except (TypeError, ValueError):
|
||||
pid_i = None
|
||||
if (
|
||||
pid_i is not None
|
||||
and pid_i != os.getpid()
|
||||
and not issue_lock_store.is_process_alive(pid_i)
|
||||
):
|
||||
return None
|
||||
return path
|
||||
|
||||
|
||||
def _resolve_preflight_workspace_path(worktree_path: str | None = None) -> str:
|
||||
@@ -2031,6 +2050,7 @@ import issue_lock_store # noqa: E402
|
||||
import issue_lock_adoption # noqa: E402
|
||||
import issue_lock_recovery # noqa: E402
|
||||
import issue_lock_renewal # noqa: E402
|
||||
import dirty_same_claimant_session_rebind # noqa: E402 # #864
|
||||
import stacked_pr_support # noqa: E402
|
||||
import merge_approval_gate # noqa: E402
|
||||
import review_quarantine # noqa: E402 # #695 contaminated formal-review quarantine
|
||||
@@ -4342,6 +4362,263 @@ def gitea_lock_issue(
|
||||
return result
|
||||
|
||||
|
||||
@mcp.tool()
|
||||
def gitea_rebind_dirty_same_claimant_author_session(
|
||||
issue_number: int,
|
||||
branch_name: str,
|
||||
worktree_path: str,
|
||||
old_pid: int,
|
||||
expected_local_head: str,
|
||||
expected_remote_head: str,
|
||||
expected_dirty_paths: list[str],
|
||||
expected_fingerprints: dict,
|
||||
remote: str = "dadeschools",
|
||||
host: str | None = None,
|
||||
org: str | None = None,
|
||||
repo: str | None = None,
|
||||
dry_run: bool = False,
|
||||
authorize_reconciler_execute: bool = False,
|
||||
) -> dict:
|
||||
"""Rebind a dirty registered issue worktree to this session (#864).
|
||||
|
||||
Sanctioned only when every pin agrees: same claimant, dead old_pid matching
|
||||
the durable lock, matching local/remote heads, exact dirty path set, and
|
||||
per-path sha256 fingerprints. Preserves every tracked/untracked byte.
|
||||
Does not sync remote, create recovery worktrees, clean, reset, or move heads.
|
||||
|
||||
Role gate:
|
||||
* author — must match the lock claimant identity/profile
|
||||
* reconciler — execute only when ``authorize_reconciler_execute=True``
|
||||
* reviewer/merger — always refuse
|
||||
|
||||
``gitea.issue.comment`` (author map entry) is required for mutation; dry_run
|
||||
still assesses fully but writes nothing. Permission alone is never ownership
|
||||
proof — every pin is re-checked server-side.
|
||||
|
||||
Args:
|
||||
issue_number: Tracking issue number on the durable lock.
|
||||
branch_name: Exact locked branch name.
|
||||
worktree_path: Registered dirty worktree path (must be under branches/).
|
||||
old_pid: Dead owner PID recorded on the lock (must match session_pid/pid).
|
||||
expected_local_head: Full local HEAD sha the caller observed.
|
||||
expected_remote_head: Full remote-tracking HEAD sha the caller observed.
|
||||
expected_dirty_paths: Exact set of dirty relative paths (tracked+untracked).
|
||||
expected_fingerprints: Map of relative path -> sha256 hex of file bytes.
|
||||
remote: Known instance — 'dadeschools' or 'prgs'.
|
||||
host/org/repo: Optional target overrides (validated against binding).
|
||||
dry_run: When true, assess only (no lock/session writes).
|
||||
authorize_reconciler_execute: Reconciler-only execute gate.
|
||||
"""
|
||||
role = _profile_role_kind(get_profile())
|
||||
role_norm = (role or "").strip().lower()
|
||||
|
||||
# Permission: authors need comment; dry_run assess is reachable under read
|
||||
# for diagnosis, but execute always needs comment. Reconciler execute also
|
||||
# needs comment when authorized.
|
||||
if dry_run:
|
||||
read_block = _profile_operation_gate("gitea.read")
|
||||
if read_block:
|
||||
return {
|
||||
"success": False,
|
||||
"dry_run": True,
|
||||
"reasons": read_block,
|
||||
"permission_report": _permission_block_report("gitea.read"),
|
||||
}
|
||||
else:
|
||||
blocked = _profile_permission_block(
|
||||
task_capability_map.required_permission(
|
||||
"rebind_dirty_same_claimant_author_session"
|
||||
),
|
||||
issue_number=issue_number,
|
||||
remote=remote,
|
||||
host=host,
|
||||
org=org,
|
||||
repo=repo,
|
||||
org_explicit=org is not None,
|
||||
repo_explicit=repo is not None,
|
||||
)
|
||||
if blocked:
|
||||
return blocked
|
||||
|
||||
if role_norm in {"reviewer", "merger"}:
|
||||
return {
|
||||
"success": False,
|
||||
"dry_run": bool(dry_run),
|
||||
"outcome": dirty_same_claimant_session_rebind.REFUSED,
|
||||
"reasons": [
|
||||
f"role '{role_norm}' cannot rebind dirty same-claimant author "
|
||||
"sessions (fail closed)"
|
||||
],
|
||||
}
|
||||
if role_norm == "reconciler" and not authorize_reconciler_execute and not dry_run:
|
||||
return {
|
||||
"success": False,
|
||||
"dry_run": False,
|
||||
"outcome": dirty_same_claimant_session_rebind.REFUSED,
|
||||
"reasons": [
|
||||
"reconciler role requires authorize_reconciler_execute=True "
|
||||
"to execute dirty same-claimant rebind (fail closed)"
|
||||
],
|
||||
}
|
||||
|
||||
h, o, r = _resolve(remote, host, org, repo)
|
||||
try:
|
||||
identity = _authenticated_username(h)
|
||||
except Exception:
|
||||
identity = None
|
||||
profile = get_profile()
|
||||
profile_name = profile.get("profile_name")
|
||||
|
||||
existing = _load_existing_issue_lock(
|
||||
remote=remote, org=o, repo=r, issue_number=issue_number
|
||||
)
|
||||
resolved_wt = os.path.realpath(os.path.abspath((worktree_path or "").strip()))
|
||||
inv = dirty_same_claimant_session_rebind.collect_dirty_inventory(resolved_wt)
|
||||
|
||||
branch_res = subprocess.run(
|
||||
["git", "-C", resolved_wt, "branch", "--show-current"],
|
||||
capture_output=True,
|
||||
text=True,
|
||||
check=False,
|
||||
)
|
||||
current_branch = (branch_res.stdout or "").strip() or None
|
||||
head_res = subprocess.run(
|
||||
["git", "-C", resolved_wt, "rev-parse", "HEAD"],
|
||||
capture_output=True,
|
||||
text=True,
|
||||
check=False,
|
||||
)
|
||||
local_head = (head_res.stdout or "").strip() if head_res.returncode == 0 else None
|
||||
|
||||
# Observe remote-tracking head without network when possible.
|
||||
remote_head = None
|
||||
for ref in (
|
||||
f"refs/remotes/origin/{branch_name}",
|
||||
f"origin/{branch_name}",
|
||||
f"refs/remotes/{remote}/{branch_name}",
|
||||
f"{remote}/{branch_name}",
|
||||
):
|
||||
rh = subprocess.run(
|
||||
["git", "-C", resolved_wt, "rev-parse", "--verify", "--quiet", ref],
|
||||
capture_output=True,
|
||||
text=True,
|
||||
check=False,
|
||||
)
|
||||
if rh.returncode == 0 and (rh.stdout or "").strip():
|
||||
remote_head = (rh.stdout or "").strip()
|
||||
break
|
||||
if remote_head is None:
|
||||
# Fall back to caller's pin only for observation absence — assessment
|
||||
# still requires pin==observed, so missing observation fails closed.
|
||||
remote_head = None
|
||||
|
||||
# Competing live locks (other issues / other worktrees).
|
||||
competing_live = []
|
||||
for entry in issue_lock_store.list_live_locks():
|
||||
competing_live.append(entry)
|
||||
|
||||
# Session pointers that claim this issue lock.
|
||||
competing_sessions = []
|
||||
lock_dir = issue_lock_store.default_lock_dir()
|
||||
lock_path = issue_lock_store.lock_file_path(
|
||||
remote=remote, org=o, repo=r, issue_number=issue_number, lock_dir=lock_dir
|
||||
)
|
||||
try:
|
||||
for name in os.listdir(lock_dir):
|
||||
if not name.startswith("session-") or not name.endswith(".json"):
|
||||
continue
|
||||
ptr = issue_lock_store.read_lock_file(os.path.join(lock_dir, name))
|
||||
if not ptr:
|
||||
continue
|
||||
ptr_lock = str(ptr.get("lock_file_path") or "").strip()
|
||||
if not ptr_lock:
|
||||
continue
|
||||
try:
|
||||
same = os.path.realpath(ptr_lock) == os.path.realpath(lock_path)
|
||||
except OSError:
|
||||
same = ptr_lock == lock_path
|
||||
if not same:
|
||||
continue
|
||||
try:
|
||||
sess_pid = int(str(name)[len("session-") : -len(".json")])
|
||||
except ValueError:
|
||||
sess_pid = ptr.get("pid")
|
||||
competing_sessions.append(
|
||||
{
|
||||
"pid": sess_pid,
|
||||
"lock_file_path": ptr_lock,
|
||||
"live": issue_lock_store.is_process_alive(sess_pid),
|
||||
}
|
||||
)
|
||||
except OSError:
|
||||
pass
|
||||
|
||||
# Best-effort workflow-lease scan: any live lock file whose work_lease is a
|
||||
# non-author workflow lease on this issue/branch counts as active.
|
||||
workflow_lease_active = False
|
||||
for path in issue_lock_store.iter_lock_files(lock_dir):
|
||||
rec = issue_lock_store.read_lock_file(path)
|
||||
if not rec:
|
||||
continue
|
||||
lease = rec.get("work_lease") if isinstance(rec.get("work_lease"), dict) else {}
|
||||
op = str(lease.get("operation_type") or "")
|
||||
if op and op != issue_lock_store.AUTHOR_ISSUE_WORK_LEASE:
|
||||
if rec.get("issue_number") == issue_number or str(
|
||||
rec.get("branch_name") or ""
|
||||
) == branch_name:
|
||||
if issue_lock_store.is_lease_live(rec):
|
||||
workflow_lease_active = True
|
||||
break
|
||||
|
||||
repo_root = _canonical_local_git_root()
|
||||
# permission_allowed reflects profile gate only — never ownership proof.
|
||||
permission_allowed = True
|
||||
|
||||
result = dirty_same_claimant_session_rebind.apply_dirty_same_claimant_session_rebind(
|
||||
remote=remote,
|
||||
org=o,
|
||||
repo=r,
|
||||
issue_number=issue_number,
|
||||
branch_name=branch_name,
|
||||
worktree_path=resolved_wt,
|
||||
claimant_identity=identity,
|
||||
claimant_profile=profile_name,
|
||||
old_pid=old_pid,
|
||||
expected_local_head=expected_local_head,
|
||||
expected_remote_head=expected_remote_head,
|
||||
expected_dirty_paths=list(expected_dirty_paths or []),
|
||||
expected_fingerprints=dict(expected_fingerprints or {}),
|
||||
existing_lock=existing,
|
||||
current_identity=identity,
|
||||
current_profile=profile_name,
|
||||
role_kind=role_norm or role,
|
||||
current_pid=os.getpid(),
|
||||
current_branch=current_branch,
|
||||
local_head=local_head,
|
||||
remote_head=remote_head,
|
||||
dirty_inventory=inv,
|
||||
competing_live_locks=competing_live,
|
||||
competing_sessions=competing_sessions,
|
||||
workflow_lease_active=workflow_lease_active,
|
||||
authorize_reconciler_execute=bool(authorize_reconciler_execute),
|
||||
permission_allowed=permission_allowed,
|
||||
repo_root=repo_root,
|
||||
dry_run=bool(dry_run),
|
||||
lock_dir=lock_dir,
|
||||
)
|
||||
result["observed"] = {
|
||||
"local_head": local_head,
|
||||
"remote_head": remote_head,
|
||||
"current_branch": current_branch,
|
||||
"dirty_paths": inv.get("dirty_paths"),
|
||||
"fingerprints": inv.get("fingerprints"),
|
||||
"identity": identity,
|
||||
"profile": profile_name,
|
||||
"role_kind": role_norm,
|
||||
}
|
||||
return result
|
||||
|
||||
|
||||
@mcp.tool()
|
||||
def gitea_assess_work_issue_duplicate(
|
||||
issue_number: int,
|
||||
|
||||
@@ -16,11 +16,16 @@ ISSUE_LOCK_FILE = os.environ.get("GITEA_ISSUE_LOCK_FILE", "/tmp/gitea_issue_lock
|
||||
SOURCE_LOCK_ISSUE = "gitea_lock_issue"
|
||||
SOURCE_LOCK_ADOPTION = "gitea_lock_issue_adoption"
|
||||
SOURCE_OPERATOR_OVERRIDE = "operator_override"
|
||||
# #864: dirty-preserving same-claimant author-session rebind (dead owner PID).
|
||||
SOURCE_DIRTY_SAME_CLAIMANT_REBIND = (
|
||||
"gitea_rebind_dirty_same_claimant_author_session"
|
||||
)
|
||||
|
||||
SANCTIONED_LOCK_SOURCES = frozenset({
|
||||
SOURCE_LOCK_ISSUE,
|
||||
SOURCE_LOCK_ADOPTION,
|
||||
SOURCE_OPERATOR_OVERRIDE,
|
||||
SOURCE_DIRTY_SAME_CLAIMANT_REBIND,
|
||||
})
|
||||
|
||||
_OPERATOR_OVERRIDE_ENV = "GITEA_ISSUE_LOCK_OPERATOR_OVERRIDE"
|
||||
|
||||
@@ -32,6 +32,17 @@ TASK_CAPABILITY_MAP: dict[str, dict[str, str]] = {
|
||||
"permission": "gitea.issue.comment",
|
||||
"role": "author",
|
||||
},
|
||||
# #864: dirty-preserving same-claimant author-session rebind (dead owner PID).
|
||||
# Author MCP tool path. Reconciler execute is gated inside the tool via
|
||||
# authorize_reconciler_execute + role_kind checks (not this map entry).
|
||||
"rebind_dirty_same_claimant_author_session": {
|
||||
"permission": "gitea.issue.comment",
|
||||
"role": "author",
|
||||
},
|
||||
"gitea_rebind_dirty_same_claimant_author_session": {
|
||||
"permission": "gitea.issue.comment",
|
||||
"role": "author",
|
||||
},
|
||||
"set_issue_labels": {
|
||||
"permission": "gitea.issue.comment",
|
||||
"role": "author",
|
||||
|
||||
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,107 @@
|
||||
"""Documentation acceptance for the MCP restart governance ADR (#656).
|
||||
|
||||
Enforces issue #656 acceptance criteria:
|
||||
|
||||
* AC1 — policy document exists with an authorization matrix and the recorded
|
||||
v1 decision (controller approval + automated safety gates).
|
||||
* AC2 — restart is stated as a last resort with enumerated narrower recoveries.
|
||||
* AC3 — a unilateral LLM full restart with affected sessions is forbidden.
|
||||
* AC4 — break-glass conditions are listed.
|
||||
* AC5 — the ADR is linked to #655, #652, #653, #630, #642, and is cross-linked
|
||||
from the safety model and the web-console deployment boundary docs.
|
||||
"""
|
||||
from pathlib import Path
|
||||
|
||||
REPO_ROOT = Path(__file__).resolve().parent.parent
|
||||
ADR = REPO_ROOT / "docs" / "architecture" / "mcp-restart-governance.md"
|
||||
ADR_BASENAME = "mcp-restart-governance.md"
|
||||
|
||||
CROSS_LINK_DOCS = (
|
||||
REPO_ROOT / "docs" / "safety-model.md",
|
||||
REPO_ROOT / "docs" / "webui-deployment.md",
|
||||
)
|
||||
|
||||
LINKED_ISSUES = ("#655", "#652", "#653", "#630", "#642")
|
||||
POLICY_IDS = ("RG-01", "RG-02", "RG-03", "RG-04", "RG-05", "RG-06", "RG-07", "RG-08")
|
||||
|
||||
|
||||
def _read(path: Path) -> str:
|
||||
assert path.is_file(), f"missing {path.relative_to(REPO_ROOT)}"
|
||||
return path.read_text(encoding="utf-8")
|
||||
|
||||
|
||||
def test_ac1_adr_exists_with_matrix_and_v1_decision():
|
||||
text = _read(ADR)
|
||||
lower = text.lower()
|
||||
assert text.lstrip().startswith("#"), "ADR lacks a title"
|
||||
assert "#656" in text
|
||||
assert "authorization matrix" in lower
|
||||
# The matrix is a real table with the worker and privileged roles.
|
||||
for role in ("author", "reviewer", "merger", "reconciler", "controller",
|
||||
"operator", "admin"):
|
||||
assert role in lower, f"authorization matrix missing role {role!r}"
|
||||
# Recorded v1 decision.
|
||||
assert "restart-governance/v1" in text
|
||||
assert "controller approval" in lower and "automated safety gates" in lower
|
||||
|
||||
|
||||
def test_ac2_restart_is_last_resort_with_narrower_recoveries():
|
||||
text = _read(ADR)
|
||||
lower = text.lower()
|
||||
assert "last resort" in lower
|
||||
# Enumerated narrower recoveries precede full restart on the ladder.
|
||||
for rung in ("reconnect", "rebind", "scoped restart", "full restart",
|
||||
"host"):
|
||||
assert rung in lower, f"recovery ladder missing rung {rung!r}"
|
||||
|
||||
|
||||
def test_ac3_forbids_unilateral_llm_full_restart_with_affected_sessions():
|
||||
text = _read(ADR)
|
||||
lower = text.lower()
|
||||
assert "forbidden" in lower
|
||||
assert "llm" in lower and "restart" in lower
|
||||
assert "unilateral" in lower
|
||||
# A worker role must not perform or authorize full/host restart.
|
||||
assert "must not" in lower
|
||||
|
||||
|
||||
def test_ac4_break_glass_conditions_listed():
|
||||
text = _read(ADR)
|
||||
lower = text.lower()
|
||||
assert "break-glass" in lower
|
||||
assert "incident" in lower
|
||||
assert "audit" in lower
|
||||
|
||||
|
||||
def test_ac5_adr_links_issue_lineage():
|
||||
text = _read(ADR)
|
||||
for issue in LINKED_ISSUES:
|
||||
assert issue in text, f"ADR must link issue {issue}"
|
||||
|
||||
|
||||
def test_ac5_safety_model_and_deployment_cross_link_adr():
|
||||
for path in CROSS_LINK_DOCS:
|
||||
text = _read(path)
|
||||
assert ADR_BASENAME in text, (
|
||||
f"{path.relative_to(REPO_ROOT)} must cross-link {ADR_BASENAME} "
|
||||
f"(issue #656 acceptance criterion 5)"
|
||||
)
|
||||
|
||||
|
||||
def test_policy_ids_present_for_enforcement_code():
|
||||
text = _read(ADR)
|
||||
for pid in POLICY_IDS:
|
||||
assert pid in text, f"policy id {pid} missing from ADR"
|
||||
|
||||
|
||||
def test_failure_behavior_denies_on_ambiguity():
|
||||
text = _read(ADR)
|
||||
lower = text.lower()
|
||||
assert "ambiguous" in lower and "deny" in lower
|
||||
|
||||
|
||||
def test_cross_links_do_not_embed_secrets():
|
||||
for path in (ADR,) + CROSS_LINK_DOCS:
|
||||
text = _read(path)
|
||||
for marker in ("ghp_", "BEGIN PRIVATE KEY", "Authorization: Bearer"):
|
||||
assert marker not in text, f"{path} contains {marker!r}"
|
||||
Reference in New Issue
Block a user