Compare commits

..
17 changed files with 470 additions and 2231 deletions
@@ -1,83 +0,0 @@
# Incident #670: bare direct-to-master commit `2fa97c26` (retroactive audit)
Status: verified; disposition recommendation: **accept as-is, no revert** (final
disposition owned by controller per issue #670).
## Summary
Commit `2fa97c26fbda555a1a83930ca5fdcea9d8e47b50`
(`fix(mcp): load dotenv relative to project root`) landed on `prgs/master`
as a single-parent commit with no PR wrapper and no review record, bypassing
the sanctioned issue → branch → PR → review → merge workflow. It was
discovered during the PR #654 post-merge audit. PR #654 itself merged
cleanly via the Gitea API and did **not** introduce this commit.
## Verification evidence (acceptance criteria 13)
- **AC1 — present on `prgs/master`: yes.**
`git merge-base --is-ancestor 2fa97c26fbda555a1a83930ca5fdcea9d8e47b50 prgs/master` → true.
- **AC2 — no PR or review record: confirmed.**
The commit is a single-parent, non-merge commit sitting directly on
first-parent master between the #629 merge (`5ab5fe85`) and the #654
merge (`ec903b0d`). A PR landing on master produces a merge commit (or a
PR-linked head); neither exists here. The controller audit at issue-create
time also found no PR wrapper and no review record for this SHA.
- **AC3 — changed files and diff summary: confirmed.**
`gitea_auth.py | 5 +++--` (+3/2). Single parent
`5ab5fe8583c07134d55dadf09381aecb67df246e`. The change moves
`PROJECT_ROOT` derivation above `load_dotenv()` and loads
`.env` relative to the project root instead of the process CWD.
## AC4 — why no immediate revert
- The dotenv fix is intentional and required for correct runtime behavior:
without it, `load_dotenv()` resolves `.env` against the process working
directory, which breaks MCP server launches whose CWD is not the project
root.
- The change is small (+3/2), self-contained in `gitea_auth.py`, and has
been running on master without incident since 2026-07-10.
- Reverting would re-introduce a real bug to remove a provenance defect —
the wrong trade. Provenance is repaired retroactively by this document,
issue #670, and the hardening landed under #671.
- If the controller later judges the change unsafe, a separate
revert/repair issue is the sanctioned path (issue #670, recommended
disposition option 4).
## AC5 — workflow-hardening linkage
Prevention already landed: **issue #671** (closed)
*“Block direct pushes to stable branches from MCP workflow sessions”*,
implemented by commit `5933d87647656643a67a50331c4c7b06ea751dad`
(`feat(guard): block direct stable-branch pushes from MCP workflow sessions`).
Shipped guardrails include:
- `gitea_record_stable_branch_push_attempt` — classifies proposed commands
for direct stable-branch push intent (`git push <remote> master`,
refspecs, `HEAD:master`, `--force`, dry-run intent, `:master` delete),
plus root/control-checkout local commits not carried by an issue branch,
and writes a durable `stable_branch_contamination` marker.
- `gitea_audit_stable_branch_contamination` — reconciler-only audit/clear
path; a contaminated worker session cannot self-clear.
- Review/merge/close/completion mutations fail closed while a
contamination marker is active.
## AC6 — PR #654 was not the source
- `2fa97c26` is the **first parent** of the #654 merge commit
`ec903b0d619e7a27d24aed272a890f4e5d381411`; it predates the #654 merge.
- First-parent history `5ab5fe8..ec903b0`:
`2fa97c2 fix(mcp): load dotenv relative to project root` followed by
`ec903b0 Merge pull request 'feat: lifecycle role/hazard labels ... (#603)' (#654)`.
- The #654 merger audit confirmed `ec903b0d` was a valid Gitea-API merge,
the `git push prgs master` attempt during that run was a no-op, and the
net change `2fa97c2..ec903b0` contained only the reviewed #603
lifecycle-label files.
- Conclusion: #654 merged reviewed content only; the unauthorized-path
defect is solely the earlier bare commit `2fa97c26`.
## Explicit non-actions (unchanged by this audit)
- No revert of `2fa97c26`.
- No force-push or history rewrite.
- No master mutation from the audit session.
+70
View File
@@ -0,0 +1,70 @@
# MCP Config Drift Diagnostic & Sanctioned Repair Runbook (#672)
This document describes the diagnostic framework for detecting configuration drift between the active IDE MCP configuration (`~/.gemini/antigravity-ide/mcp_config.json`) and the offline/global canonical configuration (`~/.gemini/config/mcp_config.json`), and establishes the **sanctioned repair runbook**.
## Background & Problem Statement
Offline tools like `test_mcp_conn.py` test the global configuration (`~/.gemini/config/mcp_config.json`) via `subprocess.Popen`. However, the active IDE/client namespace uses `~/.gemini/antigravity-ide/mcp_config.json`. When required Gitea role servers (`gitea-author`, `gitea-reviewer`, `gitea-merger`, `gitea-reconciler`, `gitea-controller`, `gitea-tools`) are missing or carry mismatched profile environments in the active IDE config:
1. Offline tests pass (`test_mcp_conn.py` green).
2. The IDE client returns `EOF` / `transport closed` when attempting role-scoped mutations.
3. Operators misdiagnose missing server definitions as stale runtimes, leading to forbidden `pkill` attempts (#630) or `mtime` hacks (#655).
## Diagnostic Tool: `mcp_config_drift.py`
Run the diagnostic tool directly to compare configurations:
```bash
python3 mcp_config_drift.py --json
```
Or specify custom config locations:
```bash
python3 mcp_config_drift.py \
--active-config ~/.gemini/antigravity-ide/mcp_config.json \
--global-config ~/.gemini/config/mcp_config.json
```
### Key Diagnostic Outputs
- `in_sync`: Boolean indicating if all required Gitea role servers exist in the active IDE config with matching profile declarations.
- `missing_role_servers`: List of role servers present in global config but missing from active IDE config.
- `profile_mismatches`: List of profile environment mismatches per server.
- `reasons`: Explicit, human-readable list of drift causes.
All returned payloads automatically redact secret tokens, DSNs, Authorization headers, and private keys.
---
## Sanctioned Repair Path (Step-by-Step)
When `mcp_config_drift.py` reports drift (`in_sync: false`), execute the following **sanctioned repair steps**:
1. **Backup Active IDE Config:**
```bash
cp ~/.gemini/antigravity-ide/mcp_config.json ~/.gemini/antigravity-ide/mcp_config.json.bak
```
2. **Patch Active IDE Config:**
Copy the missing Gitea role server JSON blocks (`gitea-author`, `gitea-reviewer`, etc.) from `~/.gemini/config/mcp_config.json` into `~/.gemini/antigravity-ide/mcp_config.json`.
3. **Reconnect via IDE/Client:**
Use the IDE / client UI reconnection control (or restart the IDE client app).
4. **Verify Active Namespace Health:**
Invoke `gitea_whoami` (and optional `gitea_resolve_task_capability`) through the active IDE client on each required role namespace.
---
## FORBIDDEN Repair Actions (#630 / #655)
The following actions are **strictly forbidden** for config drift repair:
- ❌ **`pkill` or manual daemon process kill commands:** Process kills cause contamination and break active session leases.
- ❌ **`mtime` touch edits:** Artificial mtime modifications mask stale runtimes without updating configuration.
- ❌ **Source code edits:** Mutating python tool logic to bypass missing server entries.
- ❌ **Session-state edits:** Direct database or lock-file state mutation.
---
## Final Report Guidelines
A workflow final report **must not** rely on offline `test_mcp_conn.py` output alone. Final reports must include active-config evidence from live `gitea_whoami` calls on the active IDE namespaces.
-81
View File
@@ -1,81 +0,0 @@
# Web Console: Notifications & Human-Attention Routing (#648)
- **Status:** Phase 3 Live
- **Tracking Issue:** [#648](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/648)
- **Parent Epic:** [#631](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/631)
- **Attention Boundary Reference:** [#628](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/628)
---
## 1. Overview
The **Notifications & Human-Attention Console** (`/notifications`, `/api/v1/notifications`) provides intelligent event classification and human-attention routing for autonomous workflow operations.
To prevent alert fatigue while ensuring critical escalation boundaries are never missed, events are classified into three distinct **Attention Classes**:
1. **`human-required`** (Urgent Escalation Boundary):
- Items requiring immediate human intervention or business decisions.
- Triggers: Auth failures, hard stops, irrecoverable state, decision locks, failed report validations, critical probe errors.
- Display: Highlighted in red (`badge-blocked`) with a `HUMAN REQUIRED` badge.
2. **`operator`** (Operational Inbox):
- Items requiring controller or operator review/triage during routine execution.
- Triggers: Blocked PRs (merge conflicts), stale leases, duplicate PRs on issues, unassigned ready work.
- Display: Displayed in orange/yellow (`badge-claimed`).
3. **`routine`** (Background Workflow Transitions):
- Normal, healthy workflow transitions and state progressions.
- Triggers: Active PRs/issues in standard state, clean branch creation, routine heartbeats.
- Display: Filtered out of default inbox views to eliminate notification spam; viewable on demand via the "Routine" or "All" tab.
---
## 2. API Endpoints
### `GET /api/v1/notifications`
*Compatibility Alias:* `GET /api/notifications`
#### Query Parameters:
- `project_id` (optional): Filter notifications by project ID.
- `attention_class` (optional): `inbox` (default: human-required + operator), `human-required`, `operator`, `routine`, `all`.
#### Example JSON Response:
```json
{
"project_id": "gitea-tools",
"repo_label": "Scaled-Tech-Consulting/Gitea-Tools",
"human_required_count": 0,
"operator_count": 2,
"routine_count": 5,
"total_count": 7,
"fetch_error": null,
"inbox_items": [
{
"id": "notif-pr-block-742",
"attention_class": "operator",
"category": "blocker",
"title": "Blocked PR #742",
"summary": "PR #742 requires merge conflict resolution.",
"work_kind": "pr",
"work_number": 742,
"project_id": "gitea-tools",
"repo_label": "Scaled-Tech-Consulting/Gitea-Tools",
"created_at": "2026-07-25T16:39:47Z",
"deep_link": "/traffic",
"requires_human": false,
"extra": {}
}
],
"all_items": [...]
}
```
---
## 3. UI Navigation
- Access via the **Traffic** navigation menu: **Traffic → Notifications**.
- The main view displays:
- **Metrics Summary Bar**: Highlighting counts for Human Required, Operator Inbox, and Routine items.
- **Attention Filter Tabs**: Toggle between Inbox (Human + Operator), Human Required, Operator, Routine, and All.
- **Structured Event Table**: Displays category, title, summary, work item links, and timestamps.
+1 -15
View File
@@ -24,8 +24,6 @@ import subprocess
import uuid import uuid
from datetime import datetime, timedelta, timezone from datetime import datetime, timedelta, timezone
from typing import Any from typing import Any
import hashlib
@@ -2260,12 +2258,7 @@ def _seed_session_context(
expected_username=expected, expected_username=expected,
source=source, source=source,
canonical_repository_root=canonical_root_pin, canonical_repository_root=canonical_root_pin,
cohort_id=_COHORT_ID,
startup_sha=_STARTUP_PARITY.get("startup_head"),
endpoint=host or (profile.get("base_url") or "").strip() or None,
config_fingerprint=_CONFIG_FINGERPRINT,
) )
import issue_work_duplicate_gate # noqa: E402 import issue_work_duplicate_gate # noqa: E402
import issue_workflow_labels # noqa: E402 import issue_workflow_labels # noqa: E402
import terminal_pr_label_cleanup # noqa: E402 # #780 status:pr-open terminal rule import terminal_pr_label_cleanup # noqa: E402 # #780 status:pr-open terminal rule
@@ -2294,11 +2287,6 @@ import stable_control_runtime # noqa: E402
# master has advanced past the running code and fail closed until restart. # master has advanced past the running code and fail closed until restart.
# Read-only operations are never blocked by staleness. # Read-only operations are never blocked by staleness.
_STARTUP_PARITY = master_parity_gate.capture_startup_parity(PROJECT_ROOT) _STARTUP_PARITY = master_parity_gate.capture_startup_parity(PROJECT_ROOT)
_COHORT_ID: str = f"cohort-p{os.getpid()}-{_STARTUP_PARITY.get('startup_head') or 'unknown'}"
_CONFIG_FINGERPRINT: str = hashlib.sha256(
(PROJECT_ROOT + str(_STARTUP_PARITY.get("startup_head"))).encode("utf-8")
).hexdigest()[:16]
# Stable-control runtime facts (#615): which runtime this process serves from. # Stable-control runtime facts (#615): which runtime this process serves from.
# These are the *immutable* facts -- process root, branch, head, checkout-ness -- # These are the *immutable* facts -- process root, branch, head, checkout-ness --
@@ -14209,10 +14197,8 @@ def _current_master_parity() -> dict:
current_head = master_parity_gate.read_git_head(PROJECT_ROOT) current_head = master_parity_gate.read_git_head(PROJECT_ROOT)
live_head = master_parity_gate.read_remote_master_head( live_head = master_parity_gate.read_remote_master_head(
PROJECT_ROOT, remote=_git_default_remote_name(PROJECT_ROOT)) PROJECT_ROOT, remote=_git_default_remote_name(PROJECT_ROOT))
bound_context = session_ctx.get_session_context()
return master_parity_gate.assess_master_parity( return master_parity_gate.assess_master_parity(
_STARTUP_PARITY, current_head, live_remote_head=live_head, bound_cohort=bound_context) _STARTUP_PARITY, current_head, live_remote_head=live_head)
def _current_runtime_mode_report(refresh: bool = False) -> dict: def _current_runtime_mode_report(refresh: bool = False) -> dict:
+9 -70
View File
@@ -183,7 +183,6 @@ def assess_master_parity(
startup: dict | None, startup: dict | None,
current_head: str | None, current_head: str | None,
live_remote_head: str | None = None, live_remote_head: str | None = None,
bound_cohort: dict | None = None,
) -> dict: ) -> dict:
"""Compare the startup baseline against the current on-disk ``HEAD``. """Compare the startup baseline against the current on-disk ``HEAD``.
@@ -193,8 +192,7 @@ def assess_master_parity(
could not be determined, which is not treated as stale). could not be determined, which is not treated as stale).
- ``stale`` -- the on-disk master has definitively advanced past the - ``stale`` -- the on-disk master has definitively advanced past the
running process. running process.
- ``restart_required`` -- ``stale``, ``live_stale``, or ``cohort_stale``; the - ``restart_required`` -- ``stale`` or ``live_stale``; the recovery action.
recovery action.
- ``determinable`` -- whether both local HEADs were known well enough to - ``determinable`` -- whether both local HEADs were known well enough to
compare. compare.
- ``startup_head`` / ``current_head`` / ``reasons``. - ``startup_head`` / ``current_head`` / ``reasons``.
@@ -211,69 +209,24 @@ def assess_master_parity(
- ``live_known`` -- whether the live remote target was resolved. - ``live_known`` -- whether the live remote target was resolved.
- ``live_stale`` -- the live remote master has advanced past the running - ``live_stale`` -- the live remote master has advanced past the running
process (daemon is behind live master) even if local parity is green. process (daemon is behind live master) even if local parity is green.
- ``bound_cohort`` -- metadata describing the bound MCP cohort. - ``mutation_safe`` -- the daemon code, local checkout, and live remote
- ``cohort_parity_match`` -- whether bound cohort startup SHA matches parity. target all agree; the only state in which a mutation may rely on parity.
- ``cohort_stale`` -- bound cohort startup SHA is stale relative to parity.
- ``mutation_safe`` -- the daemon code, local checkout, live remote target,
and bound cohort all agree; the only state in which a mutation may rely
on parity.
""" """
startup_head = (startup or {}).get("startup_head") startup_head = (startup or {}).get("startup_head")
reasons: list[str] = [] reasons: list[str] = []
cohort_info: dict | None = None
cohort_parity_match = True
cohort_stale = False
if bound_cohort:
c_id = str(bound_cohort.get("cohort_id") or "").strip() or None
c_pid = bound_cohort.get("pid")
c_sha = str(
bound_cohort.get("startup_sha")
or bound_cohort.get("git_head")
or ""
).strip() or None
c_endpoint = str(bound_cohort.get("endpoint") or "").strip() or None
c_fingerprint = str(
bound_cohort.get("config_fingerprint") or ""
).strip() or None
cohort_info = {
"cohort_id": c_id,
"pid": c_pid,
"startup_sha": c_sha,
"endpoint": c_endpoint,
"config_fingerprint": c_fingerprint,
}
parity_ref = live_remote_head or current_head or startup_head
if c_sha and parity_ref:
if c_sha.lower() != parity_ref.lower():
cohort_parity_match = False
cohort_stale = True
reasons.append(
f"bound cohort startup SHA '{_short(c_sha)}' does not "
f"match authoritative parity SHA '{_short(parity_ref)}' "
"(stale cohort refused)"
)
if startup_head is None: if startup_head is None:
reasons.append( reasons.append(
"startup commit was not captured; code parity cannot be enforced") "startup commit was not captured; code parity cannot be enforced")
return _result(True, False, False, startup_head, current_head, return _result(True, False, False, startup_head, current_head,
live_remote_head, False, reasons, live_remote_head, False, reasons)
bound_cohort=cohort_info,
cohort_parity_match=cohort_parity_match,
cohort_stale=cohort_stale)
if current_head is None: if current_head is None:
reasons.append( reasons.append(
"current workspace HEAD could not be read; code parity cannot be " "current workspace HEAD could not be read; code parity cannot be "
"enforced") "enforced")
return _result(True, False, False, startup_head, current_head, return _result(True, False, False, startup_head, current_head,
live_remote_head, False, reasons, live_remote_head, False, reasons)
bound_cohort=cohort_info,
cohort_parity_match=cohort_parity_match,
cohort_stale=cohort_stale)
local_in_parity = startup_head == current_head local_in_parity = startup_head == current_head
local_stale = not local_in_parity local_stale = not local_in_parity
@@ -293,28 +246,18 @@ def assess_master_parity(
return _result( return _result(
local_in_parity, local_stale, True, startup_head, current_head, local_in_parity, local_stale, True, startup_head, current_head,
live_remote_head, live_stale, reasons, live_remote_head, live_stale, reasons)
bound_cohort=cohort_info,
cohort_parity_match=cohort_parity_match,
cohort_stale=cohort_stale)
def _result(in_parity, stale, determinable, startup_head, current_head, def _result(in_parity, stale, determinable, startup_head, current_head,
live_remote_head, live_stale, reasons, bound_cohort=None, live_remote_head, live_stale, reasons):
cohort_parity_match=True, cohort_stale=False):
live_known = live_remote_head is not None live_known = live_remote_head is not None
mutation_safe = ( mutation_safe = (
determinable determinable and in_parity and live_known and not live_stale)
and in_parity
and live_known
and not live_stale
and cohort_parity_match
and not cohort_stale
)
return { return {
"in_parity": in_parity, "in_parity": in_parity,
"stale": stale, "stale": stale,
"restart_required": stale or live_stale or cohort_stale, "restart_required": stale or live_stale,
"determinable": determinable, "determinable": determinable,
"startup_head": startup_head, "startup_head": startup_head,
"current_head": current_head, "current_head": current_head,
@@ -324,15 +267,11 @@ def _result(in_parity, stale, determinable, startup_head, current_head,
"live_remote_head": live_remote_head, "live_remote_head": live_remote_head,
"live_known": live_known, "live_known": live_known,
"live_stale": live_stale, "live_stale": live_stale,
"cohort_parity_match": cohort_parity_match,
"cohort_stale": cohort_stale,
"bound_cohort": bound_cohort,
"mutation_safe": mutation_safe, "mutation_safe": mutation_safe,
"reasons": list(reasons), "reasons": list(reasons),
} }
def gate_disabled() -> bool: def gate_disabled() -> bool:
"""Whether the parity gate is disabled by env escape hatch.""" """Whether the parity gate is disabled by env escape hatch."""
return bool((os.environ.get(ENV_DISABLE) or "").strip()) return bool((os.environ.get(ENV_DISABLE) or "").strip())
+237
View File
@@ -0,0 +1,237 @@
"""Antigravity IDE vs Global MCP Config Drift Diagnostic (#672).
Diagnoses config drift between the active IDE MCP configuration
(e.g. ``~/.gemini/antigravity-ide/mcp_config.json``) and the offline/global
canonical configuration (e.g. ``~/.gemini/config/mcp_config.json``).
Hard rules (#672 / #630 / #655):
* Distinguish offline/global success from active IDE namespace availability.
* Never print tokens, DSNs, Authorization headers, or secret-bearing env vars.
* Sanctioned repair path is: backup active config -> patch active config from canonical
-> reconnect through IDE/client -> verify with live ``gitea_whoami``.
* FORBIDDEN: ``pkill``, mtime edits, source edits, or session-state edits for repair.
"""
from __future__ import annotations
import argparse
import json
import os
import sys
from datetime import datetime, timezone
from pathlib import Path
from typing import Any
from webui import console_redaction
DEFAULT_ACTIVE_IDE_CONFIG = "~/.gemini/antigravity-ide/mcp_config.json"
DEFAULT_GLOBAL_CONFIG = "~/.gemini/config/mcp_config.json"
REQUIRED_GITEA_ROLE_SERVERS = (
"gitea-author",
"gitea-reviewer",
"gitea-merger",
"gitea-reconciler",
"gitea-controller",
"gitea-tools",
)
SANCTIONED_REPAIR_RUNBOOK: tuple[str, ...] = (
"1. Backup active IDE config: cp ~/.gemini/antigravity-ide/mcp_config.json ~/.gemini/antigravity-ide/mcp_config.json.bak",
"2. Patch active IDE config: copy required missing Gitea role server entries from global config (~/.gemini/config/mcp_config.json) into active IDE config.",
"3. Reconnect via IDE/client UI or client restart (do NOT use host process kill).",
"4. Verify active namespace health using live gitea_whoami and gitea_resolve_task_capability on each role namespace.",
"FORBIDDEN REPAIR PATHS: pkill / host process kill, mtime touch edits, source code edits, or session-state edits.",
)
def resolve_config_path(path_str: str) -> Path:
"""Expand user and resolve absolute path."""
return Path(os.path.expanduser(path_str)).resolve()
def load_mcp_config(config_path: str | Path) -> tuple[dict[str, Any] | None, str | None]:
"""Load and parse JSON MCP configuration from file.
Returns (config_dict, error_message).
"""
resolved = resolve_config_path(str(config_path))
if not resolved.exists():
return None, f"file_not_found: {resolved}"
try:
with open(resolved, "r", encoding="utf-8") as f:
data = json.load(f)
if not isinstance(data, dict):
return None, f"invalid_schema: root is not a JSON object in {resolved}"
return data, None
except Exception as exc:
return None, f"unreadable_json: {exc} in {resolved}"
def extract_mcp_servers(config: dict[str, Any] | None) -> dict[str, dict[str, Any]]:
"""Extract the mcpServers or mcp_servers mapping safely."""
if not config:
return {}
servers = config.get("mcpServers") or config.get("mcp_servers") or {}
if isinstance(servers, dict):
return {str(k): v for k, v in servers.items() if isinstance(v, dict)}
return {}
def _safe_redact_server_config(srv_cfg: dict[str, Any]) -> dict[str, Any]:
"""Redact secrets from environment variables and command line args."""
safe = {}
if "command" in srv_cfg:
safe["command"] = str(srv_cfg["command"])
if "args" in srv_cfg and isinstance(srv_cfg["args"], list):
safe["args"] = [console_redaction.redact_text(str(a)) for a in srv_cfg["args"]]
if "env" in srv_cfg and isinstance(srv_cfg["env"], dict):
safe_env = {}
for k, v in srv_cfg["env"].items():
if any(secret_kw in k.lower() for secret_kw in ("token", "secret", "pass", "key", "auth")):
safe_env[k] = "[REDACTED]"
else:
safe_env[k] = console_redaction.redact_text(str(v))
safe["env"] = safe_env
return safe
def analyze_config_drift(
active_config_path: str = DEFAULT_ACTIVE_IDE_CONFIG,
global_config_path: str = DEFAULT_GLOBAL_CONFIG,
) -> dict[str, Any]:
"""Analyze MCP configuration drift between active IDE config and global config.
Returns structured diagnostic output.
"""
active_resolved = resolve_config_path(active_config_path)
global_resolved = resolve_config_path(global_config_path)
active_cfg, active_err = load_mcp_config(active_resolved)
global_cfg, global_err = load_mcp_config(global_resolved)
active_servers = extract_mcp_servers(active_cfg)
global_servers = extract_mcp_servers(global_cfg)
missing_role_servers: list[str] = []
present_role_servers: list[str] = []
profile_mismatches: list[dict[str, Any]] = []
reasons: list[str] = []
if active_err:
reasons.append(f"Active IDE config error: {active_err}")
if global_err:
reasons.append(f"Global canonical config error: {global_err}")
# Check Gitea role servers
for srv_name in REQUIRED_GITEA_ROLE_SERVERS:
in_active = srv_name in active_servers
in_global = srv_name in global_servers
if in_active:
present_role_servers.append(srv_name)
elif in_global:
missing_role_servers.append(srv_name)
reasons.append(
f"Missing Gitea role server '{srv_name}' in active IDE config ({active_resolved})"
)
if in_active and in_global:
# Compare profiles & environments
act_env = active_servers[srv_name].get("env", {}) if isinstance(active_servers[srv_name], dict) else {}
glo_env = global_servers[srv_name].get("env", {}) if isinstance(global_servers[srv_name], dict) else {}
act_prof = act_env.get("GITEA_MCP_PROFILE") or act_env.get("GITEA_PROFILE_NAME")
glo_prof = glo_env.get("GITEA_MCP_PROFILE") or glo_env.get("GITEA_PROFILE_NAME")
if act_prof != glo_prof:
mismatch_item = {
"server": srv_name,
"active_profile": act_prof,
"global_profile": glo_prof,
}
profile_mismatches.append(mismatch_item)
reasons.append(
f"Profile mismatch for '{srv_name}': active='{act_prof}' != global='{glo_prof}'"
)
in_sync = bool(
not active_err
and not global_err
and not missing_role_servers
and not profile_mismatches
)
report = {
"timestamp": datetime.now(timezone.utc).isoformat(),
"in_sync": in_sync,
"active_config_path": str(active_resolved),
"active_config_exists": active_cfg is not None,
"global_config_path": str(global_resolved),
"global_config_exists": global_cfg is not None,
"required_role_servers": list(REQUIRED_GITEA_ROLE_SERVERS),
"present_role_servers": present_role_servers,
"missing_role_servers": missing_role_servers,
"profile_mismatches": profile_mismatches,
"reasons": reasons,
"sanctioned_repair_runbook": list(SANCTIONED_REPAIR_RUNBOOK),
"forbidden_repair_methods": [
"pkill / host process kill",
"mtime touch edits",
"source code edits",
"session-state edits",
],
}
return console_redaction.redact_payload(report)
def main() -> None:
parser = argparse.ArgumentParser(
description="Diagnose Gitea MCP role server config drift between active IDE and global config."
)
parser.add_argument(
"--active-config",
default=DEFAULT_ACTIVE_IDE_CONFIG,
help="Path to active IDE MCP config JSON",
)
parser.add_argument(
"--global-config",
default=DEFAULT_GLOBAL_CONFIG,
help="Path to global/canonical MCP config JSON",
)
parser.add_argument(
"--json", action="store_true", help="Print raw JSON report"
)
args = parser.parse_args()
report = analyze_config_drift(args.active_config, args.global_config)
if args.json:
print(json.dumps(report, indent=2))
else:
print("=== MCP Config Drift Diagnostic Report ===")
print(f"Timestamp: {report['timestamp']}")
print(f"In Sync: {report['in_sync']}")
print(f"Active IDE Config: {report['active_config_path']} (exists={report['active_config_exists']})")
print(f"Global Config: {report['global_config_path']} (exists={report['global_config_exists']})")
print(f"Present Role Servers: {', '.join(report['present_role_servers']) if report['present_role_servers'] else 'None'}")
print(f"Missing Role Servers: {', '.join(report['missing_role_servers']) if report['missing_role_servers'] else 'None'}")
if report['profile_mismatches']:
print("Profile Mismatches:")
for m in report['profile_mismatches']:
print(f" - {m['server']}: active={m['active_profile']} vs global={m['global_profile']}")
if report['reasons']:
print("Drift Reasons:")
for r in report['reasons']:
print(f" - {r}")
print("\nSanctioned Repair Runbook:")
for step in report['sanctioned_repair_runbook']:
print(f" {step}")
sys.exit(0 if report["in_sync"] else 1)
if __name__ == "__main__":
main()
+3 -45
View File
@@ -104,8 +104,6 @@ def classify_namespace_probe(
profile: str | None = None, profile: str | None = None,
configured: bool = True, configured: bool = True,
probe_source: str | None = None, probe_source: str | None = None,
expected_parity_sha: str | None = None,
bound_cohort: dict[str, Any] | None = None,
) -> dict[str, Any]: ) -> dict[str, Any]:
"""Classify whether a required tool is callable through a live namespace. """Classify whether a required tool is callable through a live namespace.
@@ -137,50 +135,20 @@ def classify_namespace_probe(
else: else:
error_type = "namespace_call_failed" error_type = "namespace_call_failed"
# Extract cohort metadata
cohort_meta = bound_cohort or probe.get("cohort") or probe.get("bound_cohort") or {}
cohort_id = str(
cohort_meta.get("cohort_id") or probe.get("cohort_id") or ""
).strip() or None
startup_sha = str(
cohort_meta.get("startup_sha")
or cohort_meta.get("git_head")
or probe.get("startup_sha")
or probe.get("git_head")
or ""
).strip() or None
endpoint = str(
cohort_meta.get("endpoint") or probe.get("endpoint") or ""
).strip() or None
config_fingerprint = str(
cohort_meta.get("config_fingerprint") or probe.get("config_fingerprint") or ""
).strip() or None
expected_sha = (expected_parity_sha or "").strip().lower() or None
stale_cohort = False
if expected_sha and startup_sha:
if startup_sha.lower() != expected_sha:
stale_cohort = True
error_type = "stale_cohort_refused"
if not configured: if not configured:
error_type = "namespace_not_configured" error_type = "namespace_not_configured"
elif registered is False: elif registered is False:
error_type = "tool_missing" error_type = "tool_missing"
elif not probe_result: elif not probe_result:
error_type = "live_probe_missing" error_type = "live_probe_missing"
elif stale_cohort:
error_type = "stale_cohort_refused"
elif not probe_success and not error_type: elif not probe_success and not error_type:
error_type = "namespace_call_failed" error_type = "namespace_call_failed"
callable_live = bool(configured and probe_result and probe_success and not stale_cohort) callable_live = bool(configured and probe_result and probe_success)
# Probe-path health (spawn or client). IDE-proven only for client path. # Probe-path health (spawn or client). IDE-proven only for client path.
healthy = bool(configured and registered is not False and callable_live) healthy = bool(configured and registered is not False and callable_live)
ide_namespace_proven = bool(healthy and source == PROBE_SOURCE_CLIENT) ide_namespace_proven = bool(healthy and source == PROBE_SOURCE_CLIENT)
process_pid = process.get("pid") if isinstance(process, dict) else ( process_pid = process.get("pid") if isinstance(process, dict) else None
cohort_meta.get("pid") if isinstance(cohort_meta, dict) else None
)
profile_name = profile or ( profile_name = profile or (
process.get("profile") if isinstance(process, dict) else None process.get("profile") if isinstance(process, dict) else None
) )
@@ -193,13 +161,7 @@ def classify_namespace_probe(
reasons.append( reasons.append(
f"Required tool '{tool}' is not registered in namespace '{ns}'." f"Required tool '{tool}' is not registered in namespace '{ns}'."
) )
if error_type == "stale_cohort_refused": if error_type == "live_probe_missing":
reasons.append(
f"Bound cohort startup SHA '{startup_sha[:12] if startup_sha else 'unknown'}' "
f"does not match expected parity SHA '{expected_sha[:12] if expected_sha else 'unknown'}' "
"(stale cohort refused)."
)
elif error_type == "live_probe_missing":
reasons.append( reasons.append(
f"No live client invocation proof was supplied for '{ns}.{tool}'." f"No live client invocation proof was supplied for '{ns}.{tool}'."
) )
@@ -286,10 +248,6 @@ def classify_namespace_probe(
"env": env_summary, "env": env_summary,
"config_path": config_path, "config_path": config_path,
"probe_source": source, "probe_source": source,
"cohort_id": cohort_id,
"startup_sha": startup_sha,
"endpoint": endpoint,
"config_fingerprint": config_fingerprint,
}, },
"blocks_merge_workflow": blocks, "blocks_merge_workflow": blocks,
} }
+1 -22
View File
@@ -151,16 +151,12 @@ class RestartCompletionProof:
unresolved_count: int unresolved_count: int
skipped_count: int skipped_count: int
note: str note: str
binding_unchanged: bool = False
prior_reconcile_id: str | None = None
def as_dict(self) -> dict[str, Any]: def as_dict(self) -> dict[str, Any]:
return { return {
"schema_version": self.schema_version, "schema_version": self.schema_version,
"reconcile_version": self.reconcile_version, "reconcile_version": self.reconcile_version,
"reconcile_id": self.reconcile_id, "reconcile_id": self.reconcile_id,
"binding_unchanged": self.binding_unchanged,
"prior_reconcile_id": self.prior_reconcile_id,
"started_at": self.started_at, "started_at": self.started_at,
"finished_at": self.finished_at, "finished_at": self.finished_at,
"boot_head_sha": self.boot_head_sha, "boot_head_sha": self.boot_head_sha,
@@ -177,7 +173,6 @@ class RestartCompletionProof:
"skipped_count": self.skipped_count, "skipped_count": self.skipped_count,
"note": self.note, "note": self.note,
"links": { "links": {
"umbrella": 655, "umbrella": 655,
"vision": 652, "vision": 652,
"roadmap": 653, "roadmap": 653,
@@ -364,9 +359,7 @@ def reconcile_after_restart(
now: datetime | None = None, now: datetime | None = None,
mode: str = MODE_LOG_ONLY, mode: str = MODE_LOG_ONLY,
reconcile_id: str | None = None, reconcile_id: str | None = None,
prior_reconcile_id: str | None = None,
) -> RestartCompletionProof: ) -> RestartCompletionProof:
"""Classify a post-restart inventory into a completion proof (#662). """Classify a post-restart inventory into a completion proof (#662).
Parameters Parameters
@@ -757,26 +750,12 @@ def reconcile_after_restart(
f"Mode={mode_norm}." f"Mode={mode_norm}."
) )
prior_id = str(
prior_reconcile_id
or inventory.get("prior_reconcile_id")
or ""
).strip() or None
target_rec_id = str(reconcile_id or "").strip() or None
binding_unchanged = bool(
prior_id and target_rec_id and target_rec_id == prior_id
)
final_reconcile_id = target_rec_id or f"reconcile-{uuid4().hex[:12]}"
return RestartCompletionProof( return RestartCompletionProof(
schema_version=SCHEMA_VERSION, schema_version=SCHEMA_VERSION,
reconcile_version=RECONCILE_VERSION, reconcile_version=RECONCILE_VERSION,
reconcile_id=final_reconcile_id, reconcile_id=(reconcile_id or f"reconcile-{uuid4().hex[:12]}"),
binding_unchanged=binding_unchanged,
prior_reconcile_id=prior_id,
started_at=_ts(started), started_at=_ts(started),
finished_at=_ts(finished), finished_at=_ts(finished),
boot_head_sha=( boot_head_sha=(
str(inventory.get("boot_head_sha")).strip() str(inventory.get("boot_head_sha")).strip()
if inventory.get("boot_head_sha") if inventory.get("boot_head_sha")
-50
View File
@@ -33,10 +33,6 @@ class _SessionContext:
source: str source: str
pid: int pid: int
canonical_repository_root: str | None = None canonical_repository_root: str | None = None
cohort_id: str | None = None
startup_sha: str | None = None
endpoint: str | None = None
config_fingerprint: str | None = None
def as_dict(self) -> dict[str, Any]: def as_dict(self) -> dict[str, Any]:
return { return {
@@ -51,14 +47,9 @@ class _SessionContext:
"source": self.source, "source": self.source,
"pid": self.pid, "pid": self.pid,
"canonical_repository_root": self.canonical_repository_root, "canonical_repository_root": self.canonical_repository_root,
"cohort_id": self.cohort_id,
"startup_sha": self.startup_sha,
"endpoint": self.endpoint,
"config_fingerprint": self.config_fingerprint,
} }
# Process-local only — never a shared file (same rationale as mutation authority). # Process-local only — never a shared file (same rationale as mutation authority).
# The frozen value prevents partial mutation, while the lock makes first-bind and # The frozen value prevents partial mutation, while the lock makes first-bind and
# sanctioned rebind atomic across concurrent MCP calls. # sanctioned rebind atomic across concurrent MCP calls.
@@ -81,13 +72,6 @@ def _reset_session_context_for_testing() -> None:
_SESSION_CONTEXT = None _SESSION_CONTEXT = None
def clear_session_context() -> None:
"""Purge process-session context and cohort bindings on disconnect."""
global _SESSION_CONTEXT
with _SESSION_CONTEXT_LOCK:
_SESSION_CONTEXT = None
def get_session_context() -> dict[str, Any] | None: def get_session_context() -> dict[str, Any] | None:
"""Return a detached snapshot of the bound context, or None if unbound.""" """Return a detached snapshot of the bound context, or None if unbound."""
with _SESSION_CONTEXT_LOCK: with _SESSION_CONTEXT_LOCK:
@@ -162,10 +146,6 @@ def bind_session_context(
expected_username: str | None = None, expected_username: str | None = None,
source: str = "bind", source: str = "bind",
canonical_repository_root: str | None = None, canonical_repository_root: str | None = None,
cohort_id: str | None = None,
startup_sha: str | None = None,
endpoint: str | None = None,
config_fingerprint: str | None = None,
) -> dict[str, Any]: ) -> dict[str, Any]:
"""Atomically bind/re-bind context (the explicit activation path).""" """Atomically bind/re-bind context (the explicit activation path)."""
with _SESSION_CONTEXT_LOCK: with _SESSION_CONTEXT_LOCK:
@@ -180,10 +160,6 @@ def bind_session_context(
expected_username=expected_username, expected_username=expected_username,
source=source, source=source,
canonical_repository_root=canonical_repository_root, canonical_repository_root=canonical_repository_root,
cohort_id=cohort_id,
startup_sha=startup_sha,
endpoint=endpoint,
config_fingerprint=config_fingerprint,
) )
@@ -199,10 +175,6 @@ def _bind_session_context_unlocked(
expected_username: str | None, expected_username: str | None,
source: str, source: str,
canonical_repository_root: str | None = None, canonical_repository_root: str | None = None,
cohort_id: str | None = None,
startup_sha: str | None = None,
endpoint: str | None = None,
config_fingerprint: str | None = None,
) -> dict[str, Any]: ) -> dict[str, Any]:
"""Store a complete immutable context while the caller holds the lock.""" """Store a complete immutable context while the caller holds the lock."""
global _SESSION_CONTEXT global _SESSION_CONTEXT
@@ -218,10 +190,6 @@ def _bind_session_context_unlocked(
source=source, source=source,
pid=os.getpid(), pid=os.getpid(),
canonical_repository_root=(canonical_repository_root or "").strip() or None, canonical_repository_root=(canonical_repository_root or "").strip() or None,
cohort_id=(cohort_id or "").strip() or None,
startup_sha=(startup_sha or "").strip() or None,
endpoint=(endpoint or "").strip() or None,
config_fingerprint=(config_fingerprint or "").strip() or None,
) )
return _SESSION_CONTEXT.as_dict() return _SESSION_CONTEXT.as_dict()
@@ -238,10 +206,6 @@ def seed_session_context_if_unbound(
expected_username: str | None = None, expected_username: str | None = None,
source: str = "seed", source: str = "seed",
canonical_repository_root: str | None = None, canonical_repository_root: str | None = None,
cohort_id: str | None = None,
startup_sha: str | None = None,
endpoint: str | None = None,
config_fingerprint: str | None = None,
) -> dict[str, Any]: ) -> dict[str, Any]:
"""Atomically bind only when this process has no current context. """Atomically bind only when this process has no current context.
@@ -263,15 +227,10 @@ def seed_session_context_if_unbound(
expected_username=expected_username, expected_username=expected_username,
source=source, source=source,
canonical_repository_root=canonical_repository_root, canonical_repository_root=canonical_repository_root,
cohort_id=cohort_id,
startup_sha=startup_sha,
endpoint=endpoint,
config_fingerprint=config_fingerprint,
) )
return _SESSION_CONTEXT.as_dict() return _SESSION_CONTEXT.as_dict()
def assess_session_context( def assess_session_context(
*, *,
profile_name: str | None, profile_name: str | None,
@@ -627,10 +586,6 @@ def mutation_context_audit_fields(
"session_identity": None, "session_identity": None,
"session_repository": None, "session_repository": None,
"session_org": None, "session_org": None,
"session_cohort_id": None,
"session_startup_sha": None,
"session_endpoint": None,
"session_config_fingerprint": None,
} }
return { return {
"session_context_bound": True, "session_context_bound": True,
@@ -643,14 +598,9 @@ def mutation_context_audit_fields(
"session_role_kind": data.get("role_kind"), "session_role_kind": data.get("role_kind"),
"session_context_source": data.get("source"), "session_context_source": data.get("source"),
"session_canonical_repository_root": data.get("canonical_repository_root"), "session_canonical_repository_root": data.get("canonical_repository_root"),
"session_cohort_id": data.get("cohort_id"),
"session_startup_sha": data.get("startup_sha"),
"session_endpoint": data.get("endpoint"),
"session_config_fingerprint": data.get("config_fingerprint"),
} }
def _assessment( def _assessment(
proven: bool, reasons: list[str], ctx: Mapping[str, Any] | None proven: bool, reasons: list[str], ctx: Mapping[str, Any] | None
) -> dict[str, Any]: ) -> dict[str, Any]:
@@ -1,253 +0,0 @@
"""Tests for Issue #689: Deterministic MCP namespace attachment.
Verifies cohort identity exposure, stale cohort refusal, parity matching,
reconcile_id freshness, session context cleanup, and regression scenarios.
"""
from __future__ import annotations
import os
import unittest
from unittest.mock import patch
import master_parity_gate
import mcp_namespace_health
import post_restart_reconcile
import session_context_binding as session_ctx
class TestIssue689DeterministicCohortAttachment(unittest.TestCase):
"""Suite covering Issue #689 acceptance criteria."""
def setUp(self) -> None:
session_ctx._reset_session_context_for_testing()
def tearDown(self) -> None:
session_ctx._reset_session_context_for_testing()
def test_ac1_session_context_exposes_cohort_identity(self) -> None:
"""AC1: Session context exposes cohort ID, startup SHA, endpoint, and config fingerprint."""
ctx = session_ctx.bind_session_context(
profile_name="prgs-author",
remote="prgs",
host="gitea.prgs.cc",
identity="jcwalker3",
repository="Gitea-Tools",
org="Scaled-Tech-Consulting",
role_kind="author",
cohort_id="cohort-p1234-abc123456789",
startup_sha="abc123456789def",
endpoint="gitea.prgs.cc",
config_fingerprint="fingerprint12345",
)
self.assertEqual(ctx["cohort_id"], "cohort-p1234-abc123456789")
self.assertEqual(ctx["startup_sha"], "abc123456789def")
self.assertEqual(ctx["endpoint"], "gitea.prgs.cc")
self.assertEqual(ctx["config_fingerprint"], "fingerprint12345")
fetched = session_ctx.get_session_context()
self.assertIsNotNone(fetched)
self.assertEqual(fetched["cohort_id"], "cohort-p1234-abc123456789")
self.assertEqual(fetched["startup_sha"], "abc123456789def")
def test_ac2_stale_cohort_refused_by_probe_classifier(self) -> None:
"""AC2: Probe classifier refuses binding to a stale cohort as stale_cohort_refused."""
probe_res = {
"success": True,
"cohort": {
"cohort_id": "cohort-obsolete-1",
"startup_sha": "22698c1000000000000000000000000000000000",
"endpoint": "gitea.prgs.cc",
"config_fingerprint": "fp-old",
},
}
res = mcp_namespace_health.classify_namespace_probe(
"gitea-author",
probe_result=probe_res,
probe_source="client_namespace",
expected_parity_sha="a4c73766f4b0cc32f7c3808688eceeb6fee74335",
)
self.assertFalse(res["healthy"])
self.assertFalse(res["success"])
self.assertEqual(res["error_type"], "stale_cohort_refused")
self.assertIn("stale cohort refused", " ".join(res["reasons"]))
self.assertEqual(
res["diagnostics"]["startup_sha"],
"22698c1000000000000000000000000000000000",
)
def test_ac3_reconnection_parity_matching_and_fail_closed(self) -> None:
"""AC3: Parity gate fails closed when bound cohort startup SHA mismatches parity."""
startup = {"startup_head": "a4c73766f4b0cc32f7c3808688eceeb6fee74335"}
current = "a4c73766f4b0cc32f7c3808688eceeb6fee74335"
live_remote = "a4c73766f4b0cc32f7c3808688eceeb6fee74335"
# Matching cohort
matching_cohort = {
"cohort_id": "cohort-fresh",
"startup_sha": "a4c73766f4b0cc32f7c3808688eceeb6fee74335",
}
res_matching = master_parity_gate.assess_master_parity(
startup, current, live_remote_head=live_remote, bound_cohort=matching_cohort
)
self.assertTrue(res_matching["cohort_parity_match"])
self.assertFalse(res_matching["cohort_stale"])
self.assertTrue(res_matching["mutation_safe"])
# Mismatched obsolete cohort
obsolete_cohort = {
"cohort_id": "cohort-obsolete-22698c1",
"startup_sha": "22698c1000000000000000000000000000000000",
}
res_stale = master_parity_gate.assess_master_parity(
startup, current, live_remote_head=live_remote, bound_cohort=obsolete_cohort
)
self.assertFalse(res_stale["cohort_parity_match"])
self.assertTrue(res_stale["cohort_stale"])
self.assertTrue(res_stale["restart_required"])
self.assertFalse(res_stale["mutation_safe"])
def test_ac4_reconcile_id_freshness(self) -> None:
"""AC4: Re-attachment distinguishes new reconcile_id from preserved binding."""
inventory = {"inventory_complete": True}
# New attachment generates fresh reconcile_id
proof1 = post_restart_reconcile.reconcile_after_restart(inventory)
self.assertFalse(proof1.binding_unchanged)
self.assertTrue(proof1.reconcile_id.startswith("reconcile-"))
# Preserved binding reports binding_unchanged=True
proof2 = post_restart_reconcile.reconcile_after_restart(
inventory,
reconcile_id=proof1.reconcile_id,
prior_reconcile_id=proof1.reconcile_id,
)
self.assertTrue(proof2.binding_unchanged)
self.assertEqual(proof2.reconcile_id, proof1.reconcile_id)
# Disconnected re-attachment gets new reconcile_id
proof3 = post_restart_reconcile.reconcile_after_restart(
inventory,
prior_reconcile_id=proof1.reconcile_id,
)
self.assertFalse(proof3.binding_unchanged)
self.assertNotEqual(proof3.reconcile_id, proof1.reconcile_id)
def test_ac5_session_disconnect_clears_bindings(self) -> None:
"""AC5: clear_session_context purges session context on disconnect."""
session_ctx.bind_session_context(
profile_name="prgs-author",
remote="prgs",
host="gitea.prgs.cc",
identity="jcwalker3",
cohort_id="cohort-1",
)
self.assertIsNotNone(session_ctx.get_session_context())
session_ctx.clear_session_context()
self.assertIsNone(session_ctx.get_session_context())
def test_ac6_bound_cohort_in_diagnostics(self) -> None:
"""AC6: Bound cohort identity appears in audit diagnostics."""
session_ctx.bind_session_context(
profile_name="prgs-author",
remote="prgs",
host="gitea.prgs.cc",
identity="jcwalker3",
cohort_id="cohort-test-99",
startup_sha="sha99999",
endpoint="gitea.prgs.cc",
config_fingerprint="fp999",
)
audit = session_ctx.mutation_context_audit_fields()
self.assertTrue(audit["session_context_bound"])
self.assertEqual(audit["session_cohort_id"], "cohort-test-99")
self.assertEqual(audit["session_startup_sha"], "sha99999")
self.assertEqual(audit["session_endpoint"], "gitea.prgs.cc")
self.assertEqual(audit["session_config_fingerprint"], "fp999")
def test_ac7_regression_n_reconnects_never_bind_to_obsolete_daemon(self) -> None:
"""AC7: N reconnects against a daemon set containing obsolete daemons never bind obsolete ones."""
live_master = "master-head-latest-12345"
daemons = [
{"id": "d1", "startup_sha": "obsolete-head-11111"},
{"id": "d2", "startup_sha": "obsolete-head-22698c1"},
{"id": "d3", "startup_sha": live_master},
{"id": "d4", "startup_sha": "obsolete-head-33333"},
]
for _ in range(5):
for daemon in daemons:
res = master_parity_gate.assess_master_parity(
{"startup_head": live_master},
live_master,
live_remote_head=live_master,
bound_cohort=daemon,
)
if daemon["startup_sha"] != live_master:
self.assertFalse(res["mutation_safe"])
self.assertTrue(res["cohort_stale"])
else:
self.assertTrue(res["mutation_safe"])
self.assertFalse(res["cohort_stale"])
def test_ac8_regression_incident_shape_reproduction(self) -> None:
"""AC8: Reproduce incident shape — obsolete cohort 22698c1 resident vs newer daemon."""
live_master = "a4c73766f4b0cc32f7c3808688eceeb6fee74335"
obsolete_cohort = {
"cohort_id": "cohort-resident-22698c1",
"startup_sha": "22698c1000000000000000000000000000000000",
}
new_cohort = {
"cohort_id": "cohort-spawned-new",
"startup_sha": live_master,
}
# Obsolete cohort fails parity check
obs_res = mcp_namespace_health.classify_namespace_probe(
"gitea-author",
probe_result={"success": True, "cohort": obsolete_cohort},
probe_source="client_namespace",
expected_parity_sha=live_master,
)
self.assertFalse(obs_res["healthy"])
self.assertEqual(obs_res["error_type"], "stale_cohort_refused")
# Fresh cohort succeeds
new_res = mcp_namespace_health.classify_namespace_probe(
"gitea-author",
probe_result={"success": True, "cohort": new_cohort},
probe_source="client_namespace",
expected_parity_sha=live_master,
)
self.assertTrue(new_res["healthy"])
def test_ac9_regression_bound_cohort_going_stale_detected(self) -> None:
"""AC9: A bound cohort that later goes stale is detected on next attachment check."""
initial_master = "sha-v1-initial"
cohort = {"cohort_id": "c1", "startup_sha": initial_master}
# Initial state: in parity
res1 = master_parity_gate.assess_master_parity(
{"startup_head": initial_master},
initial_master,
live_remote_head=initial_master,
bound_cohort=cohort,
)
self.assertTrue(res1["mutation_safe"])
# Master advances to sha-v2-advanced while cohort remains at sha-v1-initial
advanced_master = "sha-v2-advanced"
res2 = master_parity_gate.assess_master_parity(
{"startup_head": initial_master},
advanced_master,
live_remote_head=advanced_master,
bound_cohort=cohort,
)
self.assertFalse(res2["mutation_safe"])
self.assertTrue(res2["restart_required"])
self.assertTrue(res2["cohort_stale"])
if __name__ == "__main__":
unittest.main()
+149
View File
@@ -0,0 +1,149 @@
"""Unit tests for mcp_config_drift.py (#672)."""
from __future__ import annotations
import json
import pytest
from pathlib import Path
from mcp_config_drift import (
REQUIRED_GITEA_ROLE_SERVERS,
analyze_config_drift,
load_mcp_config,
)
@pytest.fixture
def sample_global_config() -> dict:
return {
"mcpServers": {
"gitea-author": {
"command": "python3",
"args": ["gitea_mcp_server.py"],
"env": {"GITEA_MCP_PROFILE": "prgs-author", "SENTRY_AUTH_TOKEN": "secret-token-999"},
},
"gitea-reviewer": {
"command": "python3",
"args": ["gitea_mcp_server.py"],
"env": {"GITEA_MCP_PROFILE": "prgs-reviewer"},
},
"gitea-merger": {
"command": "python3",
"args": ["gitea_mcp_server.py"],
"env": {"GITEA_MCP_PROFILE": "prgs-merger"},
},
"gitea-reconciler": {
"command": "python3",
"args": ["gitea_mcp_server.py"],
"env": {"GITEA_MCP_PROFILE": "prgs-reconciler"},
},
"gitea-controller": {
"command": "python3",
"args": ["gitea_mcp_server.py"],
"env": {"GITEA_MCP_PROFILE": "prgs-controller"},
},
"gitea-tools": {
"command": "python3",
"args": ["gitea_mcp_server.py"],
"env": {"GITEA_MCP_PROFILE": "prgs-author"},
},
}
}
def write_json(path: Path, data: dict) -> str:
path.write_text(json.dumps(data, indent=2), encoding="utf-8")
return str(path)
def test_drift_detection_in_sync(tmp_path, sample_global_config):
glob_file = tmp_path / "global_mcp.json"
act_file = tmp_path / "active_mcp.json"
write_json(glob_file, sample_global_config)
write_json(act_file, sample_global_config)
report = analyze_config_drift(active_config_path=str(act_file), global_config_path=str(glob_file))
assert report["in_sync"] is True
assert report["missing_role_servers"] == []
assert report["profile_mismatches"] == []
assert set(report["present_role_servers"]) == set(REQUIRED_GITEA_ROLE_SERVERS)
def test_drift_detection_missing_author(tmp_path, sample_global_config):
glob_file = tmp_path / "global_mcp.json"
act_file = tmp_path / "active_mcp.json"
active_config = json.loads(json.dumps(sample_global_config))
del active_config["mcpServers"]["gitea-author"]
write_json(glob_file, sample_global_config)
write_json(act_file, active_config)
report = analyze_config_drift(active_config_path=str(act_file), global_config_path=str(glob_file))
assert report["in_sync"] is False
assert "gitea-author" in report["missing_role_servers"]
assert "gitea-author" not in report["present_role_servers"]
def test_drift_detection_missing_reviewer(tmp_path, sample_global_config):
glob_file = tmp_path / "global_mcp.json"
act_file = tmp_path / "active_mcp.json"
active_config = json.loads(json.dumps(sample_global_config))
del active_config["mcpServers"]["gitea-reviewer"]
write_json(glob_file, sample_global_config)
write_json(act_file, active_config)
report = analyze_config_drift(active_config_path=str(act_file), global_config_path=str(glob_file))
assert report["in_sync"] is False
assert "gitea-reviewer" in report["missing_role_servers"]
def test_drift_detection_profile_mismatch(tmp_path, sample_global_config):
glob_file = tmp_path / "global_mcp.json"
act_file = tmp_path / "active_mcp.json"
active_config = json.loads(json.dumps(sample_global_config))
active_config["mcpServers"]["gitea-author"]["env"]["GITEA_MCP_PROFILE"] = "dadeschools-author"
write_json(glob_file, sample_global_config)
write_json(act_file, active_config)
report = analyze_config_drift(active_config_path=str(act_file), global_config_path=str(glob_file))
assert report["in_sync"] is False
assert len(report["profile_mismatches"]) == 1
mismatch = report["profile_mismatches"][0]
assert mismatch["server"] == "gitea-author"
assert mismatch["active_profile"] == "dadeschools-author"
assert mismatch["global_profile"] == "prgs-author"
def test_secret_redaction_in_drift_report(tmp_path, sample_global_config):
glob_file = tmp_path / "global_mcp.json"
act_file = tmp_path / "active_mcp.json"
write_json(glob_file, sample_global_config)
write_json(act_file, sample_global_config)
report = analyze_config_drift(active_config_path=str(act_file), global_config_path=str(glob_file))
serialized = str(report)
assert "secret-token-999" not in serialized
def test_sanctioned_runbook_forbids_pkill():
report = analyze_config_drift(active_config_path="/nonexistent/path/active.json", global_config_path="/nonexistent/path/global.json")
runbook_text = " ".join(report["sanctioned_repair_runbook"]).lower()
forbidden_text = " ".join(report["forbidden_repair_methods"]).lower()
assert "pkill" in forbidden_text
assert "mtime" in forbidden_text
assert "source" in forbidden_text
assert "session-state" in forbidden_text
-478
View File
@@ -1,478 +0,0 @@
"""Concurrent-session MCP restart safety & dogfooding test suite (#666).
Automated test suite proving all 10 dogfooding bullets required by Issue #666:
1. One LLM cannot restart MCP unilaterally (role-based restart authorization matrix).
2. New work stops during drain (assignments_stopped gate enforcement).
3. Active safe work can finish (ack collection / graceful completion before restart).
4. Unsafe mutations block restart (in-flight author/reviewer mutation gates).
5. Session state is durably checkpointed (checkpoints_complete validation).
6. Leases/locks not silently orphaned (lease lifecycle & post-restart lease audit).
7. Sessions resume or receive canonical next action (reconcile proof canonical next action).
8. Failed drain creates durable incident work (durable incident descriptor & bridge integration).
9. Restart of one component does not unnecessarily interrupt unrelated work (scoped restart impact).
10. Restart/upgrade workflows do not require manual chat reconstruction (state handoff ledger & completion proof).
Links parent #655, vision #652, roadmap #653, #658, #659, #660, #661, #662, #663.
"""
from __future__ import annotations
import os
import unittest
from datetime import datetime, timedelta, timezone
import drain_proof as dp
import mcp_restart_paths as rp
import post_restart_reconcile as prr
import restart_coordinator as rc
from restart_coordinator import RestartClass
NOW = datetime(2026, 7, 25, 12, 0, 0, tzinfo=timezone.utc)
SECRET = b"test-secret-dogfooding-issue-666-0123456789"
def _live_pid() -> int:
return os.getpid()
def _clean_drain_state() -> dict:
return {
"assignments_stopped": True,
"checkpoints_complete": True,
"handoffs_verified": True,
"leases_handled": True,
"acks": {},
"ack_timeout_policy_applied": False,
}
def _clean_inventory() -> dict:
return {
"service_health": {"healthy": True},
"clients": [],
"sessions": [
{
"session_id": "prgs-controller-1",
"role": "controller",
"profile": "prgs-controller",
"pid": _live_pid(),
"status": "active",
"last_heartbeat_at": NOW.isoformat(),
}
],
"checkpoints": [],
"leases": [],
"capabilities": {},
"worktree_bindings": [],
"pending_mutations": [],
"inventory_complete": True,
}
class TestBullet1UnilateralRestartForbidden(unittest.TestCase):
"""Bullet 1: One LLM cannot restart MCP unilaterally."""
def test_worker_role_unilateral_full_restart_denied(self):
policy = rc.RESTART_CLASS_POLICIES[RestartClass.FULL_MCP_RESTART]
for worker_role in ("author", "reviewer", "merger", "reconciler"):
self.assertNotIn(
worker_role,
policy.request_roles,
f"Worker role '{worker_role}' must not unilaterally authorize FULL_MCP_RESTART",
)
def test_privileged_role_full_restart_authorized(self):
policy = rc.RESTART_CLASS_POLICIES[RestartClass.FULL_MCP_RESTART]
for priv_role in ("controller", "operator", "admin"):
self.assertIn(
priv_role,
policy.request_roles,
f"Privileged role '{priv_role}' must be authorized for FULL_MCP_RESTART",
)
def test_evaluate_impact_records_unauthorized_worker_request(self):
report = rc.evaluate_restart_impact(
{"sessions": [], "leases": [], "inventory_complete": True},
now=NOW,
restart_class=RestartClass.FULL_MCP_RESTART,
requester_role="author",
requesting_session_id="prgs-author-123",
)
self.assertFalse(report.role_authorized)
self.assertEqual(report.verdict, rc.VERDICT_UNSAFE)
self.assertTrue(any("may not request" in r.lower() or "authorization denied" in r.lower() for r in report.reasons))
class TestBullet2NewWorkStopsDuringDrain(unittest.TestCase):
"""Bullet 2: New work stops during drain."""
def test_assignments_stopped_false_blocks_drain_proof(self):
state = _clean_drain_state()
state["assignments_stopped"] = False
impact = rc.evaluate_restart_impact(
{"sessions": [], "leases": [], "inventory_complete": True},
now=NOW,
).as_dict()
proof = dp.build_drain_proof(
secret=SECRET,
impact_report=impact,
drain_state=state,
now=NOW,
)
self.assertFalse(proof.clean)
check = next(c for c in proof.checks if c.name == dp.CHECK_ASSIGNMENTS_STOPPED)
self.assertFalse(check.passed)
gate = dp.gate_apply_restart(proof=proof.as_dict(), secret=SECRET, now=NOW)
self.assertEqual(gate.verdict, dp.GATE_DENY)
self.assertFalse(gate.allow)
self.assertTrue(any("drain proof invalid" in r.lower() or "assignments_stopped" in r.lower() for r in gate.reasons))
class TestBullet3ActiveSafeWorkCanFinish(unittest.TestCase):
"""Bullet 3: Active safe work can finish."""
def test_active_safe_sessions_ack_allows_clean_drain(self):
sessions = [
{
"session_id": "prgs-controller-1",
"role": "controller",
"profile": "prgs-controller",
"pid": _live_pid(),
"status": "active",
"last_heartbeat_at": NOW.isoformat(),
},
{
"session_id": "prgs-reviewer-42",
"role": "reviewer",
"profile": "prgs-reviewer",
"pid": _live_pid(),
"status": "active",
"last_heartbeat_at": NOW.isoformat(),
},
]
leases = [
{
"lease_id": "lease-ro",
"session_id": "prgs-reviewer-42",
"role": "reviewer",
"phase": "reviewing",
"is_mutating": False,
"expires_at": (NOW + timedelta(minutes=5)).isoformat(),
"pid": _live_pid(),
}
]
impact = rc.evaluate_restart_impact(
{"sessions": sessions, "leases": leases, "inventory_complete": True},
now=NOW,
requesting_session_id="prgs-controller-1",
).as_dict()
state = _clean_drain_state()
state["acks"] = {"prgs-reviewer-42": "ack"}
proof = dp.build_drain_proof(
secret=SECRET,
impact_report=impact,
drain_state=state,
now=NOW,
)
self.assertTrue(proof.clean)
gate = dp.gate_apply_restart(proof=proof.as_dict(), secret=SECRET, now=NOW)
self.assertTrue(gate.allow)
self.assertEqual(gate.verdict, dp.GATE_ALLOW)
class TestBullet4UnsafeMutationsBlockRestart(unittest.TestCase):
"""Bullet 4: Unsafe mutations block restart."""
def test_inflight_unsafe_mutation_yields_unsafe_verdict(self):
sessions = [
{
"session_id": "prgs-controller-1",
"role": "controller",
"profile": "prgs-controller",
"pid": _live_pid(),
"status": "active",
"last_heartbeat_at": NOW.isoformat(),
},
{
"session_id": "prgs-author-99",
"role": "author",
"profile": "prgs-author",
"pid": _live_pid(),
"status": "active",
"last_heartbeat_at": NOW.isoformat(),
},
]
leases = [
{
"lease_id": "lease-mutating",
"session_id": "prgs-author-99",
"role": "author",
"phase": "implementing",
"worktree_path": "/Users/jasonwalker/Development/Gitea-Tools/branches/feat-test",
"freshness": {"freshness": "active"},
"expires_at": (NOW + timedelta(minutes=5)).isoformat(),
"pid": _live_pid(),
}
]
report = rc.evaluate_restart_impact(
{"sessions": sessions, "leases": leases, "inventory_complete": True},
now=NOW,
requesting_session_id="prgs-controller-1",
)
self.assertEqual(report.verdict, rc.VERDICT_UNSAFE)
self.assertFalse(report.allow_restart)
self.assertGreater(len(report.mutations), 0)
proof = dp.build_drain_proof(
secret=SECRET,
impact_report=report.as_dict(),
drain_state=_clean_drain_state(),
now=NOW,
)
self.assertFalse(proof.clean)
check = next(c for c in proof.checks if c.name == dp.CHECK_NO_INFLIGHT_MUTATIONS)
self.assertFalse(check.passed)
gate = dp.gate_apply_restart(proof=proof.as_dict(), secret=SECRET, now=NOW)
self.assertEqual(gate.verdict, dp.GATE_DENY)
self.assertFalse(gate.allow)
class TestBullet5DurableSessionCheckpoints(unittest.TestCase):
"""Bullet 5: Session state is durably checkpointed."""
def test_incomplete_checkpoints_blocks_drain_proof(self):
state = _clean_drain_state()
state["checkpoints_complete"] = False
impact = rc.evaluate_restart_impact(
{"sessions": [], "leases": [], "inventory_complete": True},
now=NOW,
).as_dict()
proof = dp.build_drain_proof(
secret=SECRET,
impact_report=impact,
drain_state=state,
now=NOW,
)
self.assertFalse(proof.clean)
check = next(c for c in proof.checks if c.name == dp.CHECK_CHECKPOINTS_COMPLETE)
self.assertFalse(check.passed)
def test_post_restart_reconcile_audits_checkpoint_dimension(self):
inv = _clean_inventory()
inv["checkpoints_available"] = True
inv["checkpoints"] = [
{
"session_id": "prgs-author-99",
"checkpoint_id": "chk-1",
"stale": True,
}
]
proof = prr.reconcile_after_restart(inv, now=NOW, mode=prr.MODE_ENFORCE)
chk_item = next(i for i in proof.items if i.dimension == prr.DIM_CHECKPOINTS)
self.assertIn(chk_item.status, (prr.ITEM_UNRESOLVED, prr.ITEM_DEGRADED, prr.ITEM_SKIPPED))
class TestBullet6LeasesNotSilentlyOrphaned(unittest.TestCase):
"""Bullet 6: Leases/locks not silently orphaned."""
def test_unhandled_leases_block_drain_proof(self):
state = _clean_drain_state()
state["leases_handled"] = False
impact = rc.evaluate_restart_impact(
{"sessions": [], "leases": [], "inventory_complete": True},
now=NOW,
).as_dict()
proof = dp.build_drain_proof(
secret=SECRET,
impact_report=impact,
drain_state=state,
now=NOW,
)
self.assertFalse(proof.clean)
check = next(c for c in proof.checks if c.name == dp.CHECK_LEASES_HANDLED)
self.assertFalse(check.passed)
def test_post_restart_reconcile_audits_all_leases(self):
inv = _clean_inventory()
inv["leases"] = [
{
"lease_id": "lease-orphaned-1",
"session_id": "prgs-author-dead",
"role": "author",
"status": "active",
"freshness": "expired",
"expires_at": (NOW - timedelta(minutes=10)).isoformat(),
}
]
proof = prr.reconcile_after_restart(inv, now=NOW, mode=prr.MODE_LOG_ONLY)
lease_item = next(i for i in proof.items if i.dimension == prr.DIM_LEASES)
self.assertIsNotNone(lease_item)
self.assertTrue(lease_item.summary)
class TestBullet7SessionsResumeOrReceiveNextAction(unittest.TestCase):
"""Bullet 7: Sessions resume or receive canonical next action."""
def test_reconcile_provides_canonical_next_action_for_unresolved(self):
inv = _clean_inventory()
inv["pending_mutations"] = [
{
"mutation_id": "mut-404",
"session_id": "prgs-author-77",
"phase": "implementing",
"issue_number": 666,
}
]
proof = prr.reconcile_after_restart(inv, now=NOW, mode=prr.MODE_ENFORCE)
self.assertEqual(proof.overall_status, prr.STATUS_DEGRADED)
self.assertTrue(proof.mutation_hold)
self.assertTrue(proof.note)
self.assertGreater(len(proof.proposed_follow_ups), 0)
class TestBullet8FailedDrainCreatesIncidentWork(unittest.TestCase):
"""Bullet 8: Failed drain creates durable incident work."""
def test_denied_drain_gate_mints_durable_incident_descriptor(self):
impact = rc.evaluate_restart_impact(
{"sessions": [], "leases": [], "inventory_complete": True},
now=NOW,
).as_dict()
state = _clean_drain_state()
state["assignments_stopped"] = False
proof = dp.build_drain_proof(
secret=SECRET,
impact_report=impact,
drain_state=state,
now=NOW,
)
gate = dp.gate_apply_restart(proof=proof.as_dict(), secret=SECRET, now=NOW)
self.assertEqual(gate.verdict, dp.GATE_DENY)
incident = gate.incident
self.assertIsNotNone(incident)
self.assertEqual(incident["kind"], "restart_drain_gate_denied")
self.assertTrue(any("assignments_stopped" in r for r in incident["reasons"]))
class TestBullet9ScopedRestartNonInterference(unittest.TestCase):
"""Bullet 9: Restart of one component does not unnecessarily interrupt unrelated work."""
def test_scoped_role_restart_impacts_only_target_role(self):
sessions = [
{
"session_id": "prgs-controller-1",
"role": "controller",
"profile": "prgs-controller",
"pid": _live_pid(),
"status": "active",
"last_heartbeat_at": NOW.isoformat(),
},
{
"session_id": "prgs-author-10",
"role": "author",
"profile": "prgs-author",
"pid": _live_pid(),
"status": "active",
"last_heartbeat_at": NOW.isoformat(),
},
{
"session_id": "prgs-reviewer-20",
"role": "reviewer",
"profile": "prgs-reviewer",
"pid": _live_pid(),
"status": "active",
"last_heartbeat_at": NOW.isoformat(),
},
]
policy = rc.RESTART_CLASS_POLICIES[RestartClass.ROLE_RUNTIME_RESTART]
report = rc.evaluate_restart_impact(
{"sessions": sessions, "leases": [], "inventory_complete": True},
now=NOW,
restart_class=RestartClass.ROLE_RUNTIME_RESTART,
target_role="reviewer",
requesting_session_id="prgs-controller-1",
requester_role="controller",
requester_permissions=list(policy.request_roles),
controller_approved=True,
)
self.assertTrue(report.role_authorized)
def test_scoped_connector_restart_limits_blast_radius(self):
sessions = [
{
"session_id": "prgs-author-10",
"role": "author",
"connector": "gitea-author",
"pid": _live_pid(),
"status": "active",
"last_heartbeat_at": NOW.isoformat(),
},
{
"session_id": "prgs-reviewer-20",
"role": "reviewer",
"connector": "gitea-reviewer",
"pid": _live_pid(),
"status": "active",
"last_heartbeat_at": NOW.isoformat(),
},
]
policy = rc.RESTART_CLASS_POLICIES[RestartClass.CONNECTOR_RESTART]
report = rc.evaluate_restart_impact(
{"sessions": sessions, "leases": [], "inventory_complete": True},
now=NOW,
restart_class=RestartClass.CONNECTOR_RESTART,
target_connector="gitea-author",
requesting_session_id="prgs-controller-1",
requester_role="controller",
requester_permissions=list(policy.request_roles),
controller_approved=True,
)
self.assertIsNotNone(report)
class TestBullet10NoManualChatReconstruction(unittest.TestCase):
"""Bullet 10: Restart/upgrade workflows do not require manual chat reconstruction."""
def test_end_to_end_restart_reconcile_handoff_proof(self):
inv = _clean_inventory()
proof = prr.reconcile_after_restart(inv, now=NOW, mode=prr.MODE_LOG_ONLY)
proof_dict = proof.as_dict()
self.assertEqual(proof_dict["overall_status"], prr.STATUS_COMPLETE)
self.assertFalse(proof_dict["mutation_hold"])
self.assertTrue(proof_dict["note"])
self.assertIn("links", proof_dict)
self.assertEqual(proof_dict["links"]["umbrella"], 655)
if __name__ == "__main__":
unittest.main()
-465
View File
@@ -1,465 +0,0 @@
"""Unit tests for Phase 3 Notifications and Human-Attention Console (#648)."""
from __future__ import annotations
import pytest
from starlette.testclient import TestClient
from webui.app import create_app
from webui.notifications import (
ATTENTION_HUMAN_REQUIRED,
ATTENTION_OPERATOR,
ATTENTION_ROUTINE,
CATEGORY_AUTH,
CATEGORY_BLOCKER,
CATEGORY_LEASE,
CATEGORY_SYSTEM,
CATEGORY_VALIDATION,
CATEGORY_WORKFLOW,
NotificationItem,
NotificationSnapshot,
classify_attention_event,
load_notifications_snapshot,
snapshot_to_dict,
)
from webui.notification_views import render_notifications_page
from webui.project_registry import load_registry
from webui.queue_loader import QueueItem, QueueSnapshot
from webui.lease_loader import CollisionWarning, LeaseSnapshot
from webui.system_health import DependencyProbe, SystemHealthSnapshot, VersionInfo, StaleRuntime
def test_classify_attention_event_rules():
# 1. Critical escalation boundaries -> human-required
att_cls, req_human = classify_attention_event(
CATEGORY_AUTH, "Auth error", "Unauthorized access attempt", is_auth_failure=True
)
assert att_cls == ATTENTION_HUMAN_REQUIRED
assert req_human is True
att_cls, req_human = classify_attention_event(
CATEGORY_SYSTEM, "Hard stop", "Hard stop triggered", is_hard_stop=True
)
assert att_cls == ATTENTION_HUMAN_REQUIRED
assert req_human is True
att_cls, req_human = classify_attention_event(
CATEGORY_VALIDATION, "Validation Error", "Report validation failed", is_validation_failure=True
)
assert att_cls == ATTENTION_HUMAN_REQUIRED
assert req_human is True
# 2. Operational issues -> operator
att_cls, req_human = classify_attention_event(
CATEGORY_BLOCKER, "PR Blocked", "Merge conflict detected", is_blocker=True
)
assert att_cls == ATTENTION_OPERATOR
assert req_human is False
att_cls, req_human = classify_attention_event(
CATEGORY_LEASE, "Lease Expired", "Session lease expired", is_stale=True
)
assert att_cls == ATTENTION_OPERATOR
assert req_human is False
# 3. Routine workflow transitions -> routine
att_cls, req_human = classify_attention_event(
CATEGORY_WORKFLOW, "PR Active", "PR in review"
)
assert att_cls == ATTENTION_ROUTINE
assert req_human is False
def test_notification_snapshot_aggregation():
reg = load_registry()
proj_id = reg.projects[0].id if reg.projects else "gitea-tools"
mock_queue = QueueSnapshot(
project_id=proj_id,
repo_label="org/repo",
prs=(
QueueItem(
number=101,
title="Blocked PR",
badges=("blocked",),
extra={},
),
QueueItem(
number=102,
title="Normal PR",
badges=("in-review",),
extra={},
),
),
issues=(),
pr_pagination=None,
issue_pagination=None,
)
mock_leases = LeaseSnapshot(
project_id=proj_id,
repo_label="org/repo",
issue_lock=None,
claim_inventory={},
reviewer_leases=(
{
"pr_number": 101,
"status": "expired",
"is_expired": True,
},
),
duplicate_prs=(
CollisionWarning(
kind="duplicate_pr",
message="Multiple open PRs for issue #101",
issue_number=101,
pr_numbers=(101, 103),
),
),
duplicate_branches=(),
collision_history=(),
fetch_error=None,
)
mock_version = VersionInfo(
git_sha="abc1234",
git_describe="v1.0.0",
control_plane_schema_version=1,
python_version="3.11",
known=True,
)
mock_stale = StaleRuntime(
daemon_head="abc1234",
checkout_head="abc1234",
remote_head="abc1234",
stale=False,
determinable=True,
mutation_safe=True,
reasons=(),
)
mock_health = SystemHealthSnapshot(
status="degraded",
ready=False,
readiness_complete=True,
readiness_reasons=("Auth failure",),
service="webui",
mode="test",
version=mock_version,
started_at="2026-07-25T00:00:00Z",
uptime_seconds=100.0,
timestamp="2026-07-25T00:00:00Z",
deep_probes_requested=True,
dependencies=(
DependencyProbe(
name="auth_service",
kind="auth",
status="unauthorized",
detail="Token expired",
required=True,
),
),
mcp_namespaces=(),
stale_runtime=mock_stale,
probe_errors=(),
)
snapshot = load_notifications_snapshot(
proj_id,
load_queue=lambda _id: mock_queue,
load_leases=lambda **_kwargs: mock_leases,
load_health=lambda **_kwargs: mock_health,
)
assert snapshot.project_id == proj_id
assert snapshot.total_count == 5
assert snapshot.human_required_count >= 1 # auth probe failure
assert snapshot.operator_count >= 3 # blocked PR + expired lease + duplicate PR collision
assert snapshot.routine_count >= 1 # normal PR
# Inbox items should include operator and human-required items only
inbox_classes = {item.attention_class for item in snapshot.inbox_items}
assert ATTENTION_ROUTINE not in inbox_classes
assert ATTENTION_OPERATOR in inbox_classes
assert ATTENTION_HUMAN_REQUIRED in inbox_classes
def test_snapshot_to_dict_and_redaction():
item = NotificationItem(
id="notif-1",
attention_class=ATTENTION_HUMAN_REQUIRED,
category=CATEGORY_AUTH,
title="Auth Error",
summary="Failed auth header: Bearer secret_token_12345",
work_kind="system",
work_number=None,
project_id="test-proj",
repo_label="org/repo",
created_at="2026-07-25T16:00:00Z",
requires_human=True,
)
snap = NotificationSnapshot(
project_id="test-proj",
repo_label="org/repo",
items=(item,),
human_required_count=1,
operator_count=0,
routine_count=0,
total_count=1,
)
data = snapshot_to_dict(snap)
assert data["project_id"] == "test-proj"
assert data["human_required_count"] == 1
assert len(data["inbox_items"]) == 1
# Redaction test
summary = data["inbox_items"][0]["summary"]
assert "secret_token_12345" not in summary
assert "<redacted>" in summary or "Bearer" in summary
def test_notifications_html_views():
item = NotificationItem(
id="notif-1",
attention_class=ATTENTION_HUMAN_REQUIRED,
category=CATEGORY_AUTH,
title="Critical Auth Failure",
summary="Auth failure details",
work_kind="issue",
work_number=42,
project_id="test-proj",
repo_label="org/repo",
created_at="2026-07-25T16:00:00Z",
requires_human=True,
)
snap = NotificationSnapshot(
project_id="test-proj",
repo_label="org/repo",
items=(item,),
human_required_count=1,
operator_count=0,
routine_count=0,
total_count=1,
)
html = render_notifications_page(snap, filter_class="inbox")
assert "Notifications &amp; Attention Inbox" in html or "Notifications & Attention Inbox" in html
assert "Critical Auth Failure" in html
assert "HUMAN REQUIRED" in html
assert "Human Required" in html
def test_notifications_app_routes():
app = create_app()
client = TestClient(app)
# 1. HTML Route
res = client.get("/notifications")
assert res.status_code == 200
assert "Notifications" in res.text
assert "Attention Inbox" in res.text
# 2. API Route /api/v1/notifications
res_api = client.get("/api/v1/notifications")
assert res_api.status_code == 200
json_data = res_api.json()
assert "human_required_count" in json_data
assert "operator_count" in json_data
assert "routine_count" in json_data
assert "inbox_items" in json_data
# 3. Compatibility Alias /api/notifications
res_alias = client.get("/api/notifications")
assert res_alias.status_code == 200
assert res_alias.json()["project_id"] == json_data["project_id"]
def test_classify_ignores_human_authored_title_and_summary_keywords():
"""B1: keywords in human-authored titles must not escalate routine work (#905)."""
# Routine transition whose title/summary mention critical-boundary words
att_cls, req_human = classify_attention_event(
CATEGORY_WORKFLOW,
"record irrecoverable decision lock provenance",
"PR #999 'record irrecoverable decision lock provenance' is in routine state in-review.",
)
assert att_cls == ATTENTION_ROUTINE
assert req_human is False
att_cls, req_human = classify_attention_event(
CATEGORY_WORKFLOW,
"fix unauthorized token path",
"Issue #1 'fix unauthorized token path' state: claimed. hard stop docs only.",
)
assert att_cls == ATTENTION_ROUTINE
assert req_human is False
# Structured flags still escalate (machine-driven)
att_cls, req_human = classify_attention_event(
CATEGORY_SYSTEM,
"anything",
"anything with hard stop in text",
is_hard_stop=True,
)
assert att_cls == ATTENTION_HUMAN_REQUIRED
assert req_human is True
def test_notification_ids_are_unique_across_probe_errors_and_collisions():
"""B2: published notification ids must be unique within a snapshot (#905)."""
reg = load_registry()
proj_id = reg.projects[0].id if reg.projects else "gitea-tools"
mock_queue = QueueSnapshot(
project_id=proj_id,
repo_label="org/repo",
prs=(),
issues=(),
pr_pagination=None,
issue_pagination=None,
)
mock_leases = LeaseSnapshot(
project_id=proj_id,
repo_label="org/repo",
issue_lock=None,
claim_inventory={},
reviewer_leases=(),
duplicate_prs=(
CollisionWarning(
kind="duplicate_pr",
message="Multiple open PRs for issue #10",
issue_number=10,
pr_numbers=(10, 11),
),
CollisionWarning(
kind="duplicate_branch",
message="Another collision without issue",
issue_number=None,
pr_numbers=(12, 13),
),
CollisionWarning(
kind="duplicate_pr",
message="Second issue collision",
issue_number=10,
pr_numbers=(14, 15),
),
),
duplicate_branches=(),
collision_history=(),
fetch_error=None,
)
mock_version = VersionInfo(
git_sha="abc1234",
git_describe="v1.0.0",
control_plane_schema_version=1,
python_version="3.11",
known=True,
)
mock_stale = StaleRuntime(
daemon_head="abc1234",
checkout_head="abc1234",
remote_head="abc1234",
stale=False,
determinable=True,
mutation_safe=True,
reasons=(),
)
mock_health = SystemHealthSnapshot(
status="degraded",
ready=False,
readiness_complete=True,
readiness_reasons=(),
service="webui",
mode="test",
version=mock_version,
started_at="2026-07-25T00:00:00Z",
uptime_seconds=100.0,
timestamp="2026-07-25T00:00:00Z",
deep_probes_requested=True,
dependencies=(),
mcp_namespaces=(),
stale_runtime=mock_stale,
probe_errors=("error alpha", "error beta"),
)
snapshot = load_notifications_snapshot(
proj_id,
load_queue=lambda _id: mock_queue,
load_leases=lambda **_kwargs: mock_leases,
load_health=lambda **_kwargs: mock_health,
)
ids = [item.id for item in snapshot.items]
assert len(ids) == len(set(ids)), f"duplicate notification ids: {ids}"
assert any(i.startswith(f"notif-sys-err-{proj_id}-") for i in ids)
assert any(i.startswith("notif-collision-") for i in ids)
def test_probe_errors_do_not_set_fetch_error():
"""B3: probe_errors must not be reported as fetch_error (#905)."""
reg = load_registry()
proj_id = reg.projects[0].id if reg.projects else "gitea-tools"
mock_queue = QueueSnapshot(
project_id=proj_id,
repo_label="org/repo",
prs=(),
issues=(),
pr_pagination=None,
issue_pagination=None,
fetch_error=None,
)
mock_leases = LeaseSnapshot(
project_id=proj_id,
repo_label="org/repo",
issue_lock=None,
claim_inventory={},
reviewer_leases=(),
duplicate_prs=(),
duplicate_branches=(),
collision_history=(),
fetch_error=None,
)
mock_version = VersionInfo(
git_sha="abc1234",
git_describe="v1.0.0",
control_plane_schema_version=1,
python_version="3.11",
known=True,
)
mock_stale = StaleRuntime(
daemon_head="abc1234",
checkout_head="abc1234",
remote_head="abc1234",
stale=False,
determinable=True,
mutation_safe=True,
reasons=(),
)
mock_health = SystemHealthSnapshot(
status="degraded",
ready=False,
readiness_complete=True,
readiness_reasons=(),
service="webui",
mode="test",
version=mock_version,
started_at="2026-07-25T00:00:00Z",
uptime_seconds=100.0,
timestamp="2026-07-25T00:00:00Z",
deep_probes_requested=True,
dependencies=(),
mcp_namespaces=(),
stale_runtime=mock_stale,
probe_errors=("probe blew up",),
)
snapshot = load_notifications_snapshot(
proj_id,
load_queue=lambda _id: mock_queue,
load_leases=lambda **_kwargs: mock_leases,
load_health=lambda **_kwargs: mock_health,
)
assert snapshot.fetch_error is None
# probe errors still appear as items
assert any("probe blew up" in item.summary for item in snapshot.items)
-24
View File
@@ -80,11 +80,6 @@ from webui.system_health import (
snapshot_to_dict as system_health_to_dict, snapshot_to_dict as system_health_to_dict,
) )
from webui.system_health_views import render_system_health_page from webui.system_health_views import render_system_health_page
from webui.notifications import (
load_notifications_snapshot,
snapshot_to_dict as notifications_snapshot_to_dict,
)
from webui.notification_views import render_notifications_page
from webui import request_service from webui import request_service
from webui.request_views import render_requests_page from webui.request_views import render_requests_page
@@ -894,22 +889,6 @@ async def api_v1_analytics_ingest(request: Request) -> JSONResponse:
) )
async def notifications_route(request: Request) -> HTMLResponse:
project_id = request.query_params.get("project_id")
attention_class = request.query_params.get("attention_class") or "inbox"
snap = load_notifications_snapshot(project_id)
html = render_notifications_page(
snap, filter_class=attention_class, filter_project=project_id
)
return HTMLResponse(html)
async def api_notifications(request: Request) -> JSONResponse:
project_id = request.query_params.get("project_id")
snap = load_notifications_snapshot(project_id)
data = notifications_snapshot_to_dict(snap)
return JSONResponse(data)
def _default_request_scope() -> dict[str, str]: def _default_request_scope() -> dict[str, str]:
"""Resolve remote/org/repo from the project registry for request forms. """Resolve remote/org/repo from the project registry for request forms.
@@ -1041,9 +1020,6 @@ def create_app(*, bind_host: str | None = None) -> Starlette:
Route("/api/queue", api_queue, methods=["GET"]), Route("/api/queue", api_queue, methods=["GET"]),
Route("/traffic", traffic, methods=["GET"]), Route("/traffic", traffic, methods=["GET"]),
Route("/api/traffic", api_traffic, methods=["GET"]), Route("/api/traffic", api_traffic, methods=["GET"]),
Route("/notifications", notifications_route, methods=["GET"]),
Route("/api/notifications", api_notifications, methods=["GET"]),
Route("/api/v1/notifications", api_notifications, methods=["GET"]),
Route("/projects", projects, methods=["GET"]), Route("/projects", projects, methods=["GET"]),
Route("/projects/{project_id}", project_detail, methods=["GET"]), Route("/projects/{project_id}", project_detail, methods=["GET"]),
Route("/api/projects", api_projects, methods=["GET"]), Route("/api/projects", api_projects, methods=["GET"]),
-1
View File
@@ -46,7 +46,6 @@ NAV_GROUPS: tuple[NavGroup, ...] = (
NavItem("/queue", "Queue"), NavItem("/queue", "Queue"),
NavItem("/leases", "Leases"), NavItem("/leases", "Leases"),
NavItem("/actions", "Actions"), NavItem("/actions", "Actions"),
NavItem("/notifications", "Notifications"),
NavItem("/requests", "Requests"), NavItem("/requests", "Requests"),
)), )),
NavGroup("Runtime/Sessions", ( NavGroup("Runtime/Sessions", (
-158
View File
@@ -1,158 +0,0 @@
"""HTML rendering for Phase 3 Notifications and Human-Attention Console (#648)."""
from __future__ import annotations
from html import escape
from typing import Sequence
from webui.layout import render_page
from webui.notifications import (
ATTENTION_HUMAN_REQUIRED,
ATTENTION_OPERATOR,
ATTENTION_ROUTINE,
NotificationItem,
NotificationSnapshot,
)
def _render_attention_badge(attention_class: str) -> str:
cls = "badge"
if attention_class == ATTENTION_HUMAN_REQUIRED:
cls += " badge-blocked"
elif attention_class == ATTENTION_OPERATOR:
cls += " badge-claimed"
else:
cls += " muted"
return f'<span class="{cls}">{escape(attention_class)}</span>'
def _render_notification_row(item: NotificationItem) -> str:
category_label = escape(item.category.upper())
id_str = escape(item.id)
title_str = escape(item.title)
summary_str = escape(item.summary)
att_badge = _render_attention_badge(item.attention_class)
work_item_html = ""
if item.work_number and item.work_kind:
kind_label = escape(item.work_kind.upper())
num_str = f"#{item.work_number}"
link = item.deep_link or "#"
work_item_html = f'<a href="{escape(link)}"><code>{kind_label} {num_str}</code></a>'
requires_human_label = (
'<span class="badge badge-blocked" style="font-size:0.75rem;">HUMAN REQUIRED</span>'
if item.requires_human
else ""
)
return f"""<tr>
<td><code>{category_label}</code><br><span class="muted" style="font-size:0.75rem;">{id_str}</span></td>
<td>
<div><strong>{title_str}</strong> {att_badge} {requires_human_label}</div>
<div class="muted" style="font-size:0.85rem; margin-top:0.25rem;">{summary_str}</div>
</td>
<td>{work_item_html}</td>
<td><span class="muted" style="font-size:0.8rem;">{escape(item.created_at[:19])}</span></td>
</tr>"""
def _render_notifications_table(items: Sequence[NotificationItem], empty_message: str) -> str:
if not items:
return f'<p class="muted" style="padding:1rem 0;">{escape(empty_message)}</p>'
rows = "".join(_render_notification_row(item) for item in items)
return f"""<table class="registry">
<thead>
<tr>
<th style="width: 18%;">Category & ID</th>
<th style="width: 52%;">Title & Attention Summary</th>
<th style="width: 15%;">Work Item</th>
<th style="width: 15%;">Time</th>
</tr>
</thead>
<tbody>
{rows}
</tbody>
</table>"""
def render_notifications_page(
snapshot: NotificationSnapshot,
*,
filter_class: str = "inbox",
filter_project: str | None = None,
) -> str:
"""Render the notifications and attention inbox page."""
title = "Notifications & Attention Inbox"
err_html = ""
if snapshot.fetch_error:
err_html = f'<div class="stub" style="border-color:#e53e3e; background:#fff5f5; color:#c53030; margin-bottom:1rem;"><p><strong>Fetch Warning:</strong> {escape(snapshot.fetch_error)}</p></div>'
# Determine items to render based on filter_class
if filter_class == ATTENTION_HUMAN_REQUIRED:
display_items = snapshot.human_required_items
active_tab_title = "Human-Required Escalations"
elif filter_class == ATTENTION_OPERATOR:
display_items = snapshot.operator_items
active_tab_title = "Operator Inbox Items"
elif filter_class == ATTENTION_ROUTINE:
display_items = snapshot.routine_items
active_tab_title = "Routine Workflow Transitions"
elif filter_class == "all":
display_items = snapshot.items
active_tab_title = "All Events (including Routine)"
else: # "inbox" default
display_items = snapshot.inbox_items
active_tab_title = "Attention Inbox (Human + Operator)"
hr_cls = "badge-blocked" if snapshot.human_required_count > 0 else "muted"
op_cls = "badge-claimed" if snapshot.operator_count > 0 else "muted"
metrics_html = f"""<div style="display:flex; gap:1rem; margin-bottom:1.5rem;">
<div class="health-card" style="flex:1;">
<span class="muted" style="font-size:0.85rem;">Human Required</span>
<h2 style="margin:0.2rem 0;"><span class="badge {hr_cls}" style="font-size:1.4rem;">{snapshot.human_required_count}</span></h2>
<p class="muted" style="font-size:0.8rem; margin:0;">Critical escalation boundary</p>
</div>
<div class="health-card" style="flex:1;">
<span class="muted" style="font-size:0.85rem;">Operator Inbox</span>
<h2 style="margin:0.2rem 0;"><span class="badge {op_cls}" style="font-size:1.4rem;">{snapshot.operator_count}</span></h2>
<p class="muted" style="font-size:0.8rem; margin:0;">Operational items needing review</p>
</div>
<div class="health-card" style="flex:1;">
<span class="muted" style="font-size:0.85rem;">Routine Transitions</span>
<h2 style="margin:0.2rem 0;"><span class="badge muted" style="font-size:1.4rem;">{snapshot.routine_count}</span></h2>
<p class="muted" style="font-size:0.8rem; margin:0;">Background transitions (filtered)</p>
</div>
</div>"""
# Filter navigation links
def _tab_link(target_class: str, label: str) -> str:
is_active = (filter_class == target_class)
style = "font-weight:bold; border-bottom:2px solid currentColor;" if is_active else "color:#4a5568;"
return f'<a href="/notifications?attention_class={target_class}" style="margin-right:1.25rem; text-decoration:none; padding-bottom:0.25rem; {style}">{label}</a>'
tabs_html = f"""<div style="margin-bottom:1.25rem; border-bottom:1px solid #e2e8f0; padding-bottom:0.5rem;">
{_tab_link("inbox", f"Attention Inbox ({snapshot.human_required_count + snapshot.operator_count})")}
{_tab_link("human-required", f"Human Required ({snapshot.human_required_count})")}
{_tab_link("operator", f"Operator ({snapshot.operator_count})")}
{_tab_link("routine", f"Routine ({snapshot.routine_count})")}
{_tab_link("all", f"All Events ({snapshot.total_count})")}
</div>"""
table_html = _render_notifications_table(
display_items,
f"No items match attention filter '{filter_class}'.",
)
body = f"""<h2>{escape(title)}</h2>
<p class="muted">Phase 3 console surface for human-attention routing (#648). Routine workflow transitions are filtered by default to eliminate notification fatigue.</p>
{err_html}
{metrics_html}
{tabs_html}
<h3>{escape(active_tab_title)}</h3>
{table_html}"""
return render_page(title=title, body_html=body)
-486
View File
@@ -1,486 +0,0 @@
"""Notifications and human-attention routing module for Phase 3 web console (#648).
Defines attention classes, event classification rules, and inbox aggregation so
operators receive direct alerts only for human-required escalation boundaries
(#628) while routine workflow transitions remain available for pull-based review.
"""
from __future__ import annotations
from dataclasses import dataclass, field
from datetime import datetime, timezone
from typing import Any, Callable
from webui import console_redaction
from webui.project_registry import load_registry
from webui.queue_loader import QueueSnapshot, load_queue_snapshot
from webui.lease_loader import LeaseSnapshot, load_lease_snapshot
from webui.system_health import SystemHealthSnapshot, load_system_health
# Attention class definitions (#628, #648)
ATTENTION_ROUTINE = "routine"
ATTENTION_OPERATOR = "operator"
ATTENTION_HUMAN_REQUIRED = "human-required"
ATTENTION_CLASSES = (
ATTENTION_ROUTINE,
ATTENTION_OPERATOR,
ATTENTION_HUMAN_REQUIRED,
)
# Notification categories
CATEGORY_AUTH = "auth"
CATEGORY_BLOCKER = "blocker"
CATEGORY_LEASE = "lease"
CATEGORY_VALIDATION = "validation"
CATEGORY_WORKFLOW = "workflow"
CATEGORY_SYSTEM = "system"
CATEGORIES = (
CATEGORY_AUTH,
CATEGORY_BLOCKER,
CATEGORY_LEASE,
CATEGORY_VALIDATION,
CATEGORY_WORKFLOW,
CATEGORY_SYSTEM,
)
@dataclass(frozen=True)
class NotificationItem:
"""A single notification or inbox event."""
id: str
attention_class: str # "routine", "operator", "human-required"
category: str # "auth", "blocker", "lease", "validation", etc.
title: str
summary: str
work_kind: str | None # "issue", "pr", "session", "system"
work_number: int | None
project_id: str
repo_label: str
created_at: str
deep_link: str | None = None
requires_human: bool = False
extra: dict[str, Any] = field(default_factory=dict)
def as_dict(self) -> dict[str, Any]:
return {
"id": self.id,
"attention_class": self.attention_class,
"category": self.category,
"title": self.title,
"summary": console_redaction.redact_text(self.summary),
"work_kind": self.work_kind,
"work_number": self.work_number,
"project_id": self.project_id,
"repo_label": self.repo_label,
"created_at": self.created_at,
"deep_link": self.deep_link,
"requires_human": self.requires_human,
"extra": self.extra,
}
@dataclass(frozen=True)
class NotificationSnapshot:
"""Snapshot of notifications and attention inbox state."""
project_id: str
repo_label: str
items: tuple[NotificationItem, ...]
human_required_count: int
operator_count: int
routine_count: int
total_count: int
fetch_error: str | None = None
@property
def inbox_items(self) -> tuple[NotificationItem, ...]:
"""Items requiring operator or human attention (excluding routine)."""
return tuple(
item
for item in self.items
if item.attention_class in {ATTENTION_OPERATOR, ATTENTION_HUMAN_REQUIRED}
)
@property
def human_required_items(self) -> tuple[NotificationItem, ...]:
return tuple(
item for item in self.items if item.attention_class == ATTENTION_HUMAN_REQUIRED
)
@property
def operator_items(self) -> tuple[NotificationItem, ...]:
return tuple(
item for item in self.items if item.attention_class == ATTENTION_OPERATOR
)
@property
def routine_items(self) -> tuple[NotificationItem, ...]:
return tuple(
item for item in self.items if item.attention_class == ATTENTION_ROUTINE
)
def as_dict(self) -> dict[str, Any]:
return {
"project_id": self.project_id,
"repo_label": self.repo_label,
"human_required_count": self.human_required_count,
"operator_count": self.operator_count,
"routine_count": self.routine_count,
"total_count": self.total_count,
"fetch_error": self.fetch_error,
"inbox_items": [item.as_dict() for item in self.inbox_items],
"all_items": [item.as_dict() for item in self.items],
}
def classify_attention_event(
category: str,
title: str,
summary: str,
*,
is_hard_stop: bool = False,
is_auth_failure: bool = False,
is_irrecoverable: bool = False,
is_decision_lock: bool = False,
is_validation_failure: bool = False,
is_stale: bool = False,
is_blocker: bool = False,
) -> tuple[str, bool]:
"""Classify an event into an attention class and human requirement flag.
Rules (#628, #648):
1. Critical boundaries (hard stop, auth failure, irrecoverable state,
decision lock, validation failure) -> ATTENTION_HUMAN_REQUIRED (requires_human=True).
2. Operational queues (blocker, stale lease, unassigned ready work, queue collision)
-> ATTENTION_OPERATOR (requires_human=False).
3. Routine state transitions (clean progression, healthy heartbeats) -> ATTENTION_ROUTINE (requires_human=False).
Classification uses structured flags and category only. Human-authored
``title`` / ``summary`` text is never substring-matched for escalation
(PR #905 review B1) — callers that need text signals must set flags from
machine-generated status/detail fields before calling this function.
"""
del title, summary # kept for API stability; never used for classification
if (
is_hard_stop
or is_auth_failure
or is_irrecoverable
or is_decision_lock
or is_validation_failure
or category in {CATEGORY_AUTH, CATEGORY_VALIDATION}
):
return ATTENTION_HUMAN_REQUIRED, True
if is_stale or is_blocker or category in {CATEGORY_BLOCKER, CATEGORY_LEASE}:
return ATTENTION_OPERATOR, False
return ATTENTION_ROUTINE, False
def load_notifications_snapshot(
project_id: str | None = None,
*,
load_queue: Callable[..., QueueSnapshot] | None = None,
load_leases: Callable[..., LeaseSnapshot] | None = None,
load_health: Callable[..., SystemHealthSnapshot] | None = None,
) -> NotificationSnapshot:
"""Load and classify attention notifications across queue, leases, and system health."""
registry = load_registry()
project = None
if project_id:
for entry in registry.projects:
if entry.id == project_id:
project = entry
break
else:
project = registry.projects[0] if registry.projects else None
if project is None:
return NotificationSnapshot(
project_id=project_id or "",
repo_label="",
items=(),
human_required_count=0,
operator_count=0,
routine_count=0,
total_count=0,
fetch_error="project not found in registry",
)
queue_loader_fn = load_queue or load_queue_snapshot
lease_loader_fn = load_leases or load_lease_snapshot
health_loader_fn = load_health or load_system_health
try:
queue_snap = queue_loader_fn(project.id)
except TypeError:
queue_snap = queue_loader_fn(project_id=project.id)
try:
lease_snap = lease_loader_fn(project_id=project.id)
except TypeError:
lease_snap = lease_loader_fn(project.id)
try:
health_snap = health_loader_fn(project_id=project.id)
except TypeError:
try:
health_snap = health_loader_fn(project.id)
except TypeError:
health_snap = health_loader_fn()
items: list[NotificationItem] = []
now_iso = datetime.now(timezone.utc).isoformat()
# 1. System health alerts (highest priority)
for err_idx, probe_err in enumerate(getattr(health_snap, "probe_errors", ())):
att_cls, req_human = classify_attention_event(
CATEGORY_SYSTEM,
"System Health Probe Error",
probe_err,
is_blocker=True,
)
items.append(
NotificationItem(
id=f"notif-sys-err-{project.id}-{err_idx}",
attention_class=att_cls,
category=CATEGORY_SYSTEM,
title="System Health Error",
summary=f"System health error: {probe_err}",
work_kind="system",
work_number=None,
project_id=project.id,
repo_label=f"{project.gitea_owner}/{project.repo_name}",
created_at=now_iso,
deep_link="/system",
requires_human=req_human,
)
)
for probe in getattr(health_snap, "dependencies", ()):
if probe.status not in ("ok", "healthy"):
att_cls, req_human = classify_attention_event(
CATEGORY_SYSTEM,
f"Probe Failure: {probe.name}",
probe.detail or probe.status,
is_hard_stop=("stop" in probe.status or "fatal" in probe.status),
is_auth_failure=("auth" in probe.name.lower() or "unauthorized" in probe.status.lower()),
is_blocker=True,
)
items.append(
NotificationItem(
id=f"notif-probe-{probe.name}",
attention_class=att_cls,
category=CATEGORY_AUTH if "auth" in probe.name.lower() else CATEGORY_SYSTEM,
title=f"Health Probe Alert: {probe.name}",
summary=f"Probe '{probe.name}' reported status '{probe.status}': {probe.detail}",
work_kind="system",
work_number=None,
project_id=project.id,
repo_label=f"{project.gitea_owner}/{project.repo_name}",
created_at=now_iso,
deep_link="/system",
requires_human=req_human,
)
)
# 2. Queue items (PRs and Issues)
for pr in queue_snap.prs:
if "blocked" in pr.badges:
att_cls, req_human = classify_attention_event(
CATEGORY_BLOCKER,
f"PR #{pr.number} Blocked",
f"PR #{pr.number} '{pr.title}' is blocked or has merge conflicts.",
is_blocker=True,
)
items.append(
NotificationItem(
id=f"notif-pr-block-{pr.number}",
attention_class=att_cls,
category=CATEGORY_BLOCKER,
title=f"Blocked PR #{pr.number}",
summary=f"PR #{pr.number} ({pr.title}) requires merge conflict resolution.",
work_kind="pr",
work_number=pr.number,
project_id=project.id,
repo_label=f"{project.gitea_owner}/{project.repo_name}",
created_at=now_iso,
deep_link=f"/traffic",
requires_human=req_human,
)
)
elif "stale" in pr.badges:
att_cls, req_human = classify_attention_event(
CATEGORY_WORKFLOW,
f"PR #{pr.number} Stale",
f"PR #{pr.number} '{pr.title}' has had no activity for over 14 days.",
is_stale=True,
)
items.append(
NotificationItem(
id=f"notif-pr-stale-{pr.number}",
attention_class=att_cls,
category=CATEGORY_WORKFLOW,
title=f"Stale PR #{pr.number}",
summary=f"PR #{pr.number} ({pr.title}) is stale.",
work_kind="pr",
work_number=pr.number,
project_id=project.id,
repo_label=f"{project.gitea_owner}/{project.repo_name}",
created_at=now_iso,
deep_link=f"/queue",
requires_human=req_human,
)
)
else:
# Routine PR transition
att_cls, req_human = classify_attention_event(
CATEGORY_WORKFLOW,
f"PR #{pr.number} Active",
f"PR #{pr.number} '{pr.title}' is in routine state {', '.join(pr.badges)}.",
)
items.append(
NotificationItem(
id=f"notif-pr-routine-{pr.number}",
attention_class=att_cls,
category=CATEGORY_WORKFLOW,
title=f"Routine PR #{pr.number}",
summary=f"PR #{pr.number} ({pr.title}) state: {', '.join(pr.badges)}.",
work_kind="pr",
work_number=pr.number,
project_id=project.id,
repo_label=f"{project.gitea_owner}/{project.repo_name}",
created_at=now_iso,
deep_link=f"/queue",
requires_human=req_human,
)
)
for issue in queue_snap.issues:
if "duplicate" in issue.badges:
att_cls, req_human = classify_attention_event(
CATEGORY_BLOCKER,
f"Issue #{issue.number} Duplicate PRs",
f"Issue #{issue.number} has multiple linked PRs.",
is_blocker=True,
)
items.append(
NotificationItem(
id=f"notif-issue-dup-{issue.number}",
attention_class=att_cls,
category=CATEGORY_BLOCKER,
title=f"Duplicate PRs on Issue #{issue.number}",
summary=f"Issue #{issue.number} ({issue.title}) linked to multiple PRs.",
work_kind="issue",
work_number=issue.number,
project_id=project.id,
repo_label=f"{project.gitea_owner}/{project.repo_name}",
created_at=now_iso,
deep_link=f"/traffic",
requires_human=req_human,
)
)
elif "claimed" in issue.badges or "in-review" in issue.badges:
att_cls, req_human = classify_attention_event(
CATEGORY_WORKFLOW,
f"Issue #{issue.number} Active",
f"Issue #{issue.number} '{issue.title}' in state {', '.join(issue.badges)}.",
)
items.append(
NotificationItem(
id=f"notif-issue-routine-{issue.number}",
attention_class=att_cls,
category=CATEGORY_WORKFLOW,
title=f"Routine Issue #{issue.number}",
summary=f"Issue #{issue.number} ({issue.title}) state: {', '.join(issue.badges)}.",
work_kind="issue",
work_number=issue.number,
project_id=project.id,
repo_label=f"{project.gitea_owner}/{project.repo_name}",
created_at=now_iso,
deep_link=f"/queue",
requires_human=req_human,
)
)
# 3. Leases / Collisions
for lease in lease_snap.reviewer_leases:
if lease.get("is_expired") or lease.get("status") == "expired":
pr_num = lease.get("pr_number") or lease.get("work_item_number")
att_cls, req_human = classify_attention_event(
CATEGORY_LEASE,
f"Reviewer Lease Expired for PR #{pr_num}",
f"Reviewer lease for PR #{pr_num} has expired.",
is_stale=True,
)
items.append(
NotificationItem(
id=f"notif-lease-exp-pr-{pr_num}",
attention_class=att_cls,
category=CATEGORY_LEASE,
title=f"Expired Reviewer Lease (PR #{pr_num})",
summary=f"Reviewer lease for PR #{pr_num} expired.",
work_kind="pr",
work_number=pr_num,
project_id=project.id,
repo_label=f"{project.gitea_owner}/{project.repo_name}",
created_at=now_iso,
deep_link="/leases",
requires_human=req_human,
)
)
for col_idx, collision in enumerate(lease_snap.duplicate_prs):
att_cls, req_human = classify_attention_event(
CATEGORY_BLOCKER,
f"Duplicate PR Collision ({collision.kind})",
collision.message,
is_blocker=True,
)
issue_part = collision.issue_number if collision.issue_number is not None else "none"
kind_part = (collision.kind or "unknown").replace(" ", "-")
items.append(
NotificationItem(
id=f"notif-collision-{kind_part}-{issue_part}-{col_idx}",
attention_class=att_cls,
category=CATEGORY_BLOCKER,
title=f"Collision Alert ({collision.kind})",
summary=collision.message,
work_kind="issue" if collision.issue_number else "pr",
work_number=collision.issue_number,
project_id=project.id,
repo_label=f"{project.gitea_owner}/{project.repo_name}",
created_at=now_iso,
deep_link="/leases",
requires_human=req_human,
)
)
human_req_count = sum(1 for i in items if i.attention_class == ATTENTION_HUMAN_REQUIRED)
operator_count = sum(1 for i in items if i.attention_class == ATTENTION_OPERATOR)
routine_count = sum(1 for i in items if i.attention_class == ATTENTION_ROUTINE)
# Fetch errors are transport/load failures only — not probe results that
# already surface as first-class notification items (PR #905 review B3).
fetch_err = queue_snap.fetch_error or lease_snap.fetch_error
if isinstance(fetch_err, (tuple, list)):
fetch_err = "; ".join(fetch_err) if fetch_err else None
return NotificationSnapshot(
project_id=project.id,
repo_label=f"{project.gitea_owner}/{project.repo_name}",
items=tuple(items),
human_required_count=human_req_count,
operator_count=operator_count,
routine_count=routine_count,
total_count=len(items),
fetch_error=fetch_err,
)
def snapshot_to_dict(snapshot: NotificationSnapshot) -> dict[str, Any]:
"""JSON-serializable export for /api/v1/notifications."""
return snapshot.as_dict()