Compare commits

...
Author SHA1 Message Date
sysadminandClaude Opus 4.8 b2f6e9a6dc feat(webui): versioned project registry API (Closes #635)
Evolve the MVP project registry (#427) into a versioned, fail-closed
project registry API for the console (Phase 1, read-only).

- Add schema version 2 with project `status`, per-step onboarding
  `state`/`required`, optional redacted `last_seen_health`, and
  `remote_name`. Version 1 files stay loadable and are normalized with
  explicit defaults.
- Serve `/api/v1/projects` and `/api/v1/projects/{project_id}` with API
  provenance (`api_version`, `schema_version`, `source`). `/api/projects`
  is retained as an unversioned Phase 1 alias.
- Replace bare `ValueError` with `RegistryError`, carrying an operator
  `remediation` and `field_path`; invalid registries fail closed as a
  500 JSON payload or a dedicated HTML error page instead of a traceback.
- Reject credential-shaped keys before any DTO is built, reusing
  `registry_safety.is_forbidden_key` as the single source of truth
  shared with the worker registry (#798).
- Render HTML views from `project_to_dict`, so the console and the JSON
  API cannot disagree about status or onboarding progress.
- Document the contract in docs/webui-project-registry-api.md.

Tests: registry load/validate (valid, missing project, schema
validation, credential rejection) and API route coverage.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-07-22 16:12:07 -04:00
sysadmin 9eb0f29cef Merge pull request 'feat(webui): worker registry and configuration schema (Closes #798)' (#810) from feat/issue-798-worker-registry-schema into master 2026-07-22 06:48:29 -05:00
sysadmin 0b29404031 Merge pull request 'docs(webui): MCP Control Plane Web Console architecture ADR (Closes #632)' (#796) from docs/issue-632-web-console-architecture into master 2026-07-22 06:22:11 -05:00
jcwalker3andClaude Opus 4.8 5463f58933 feat(webui): worker registry and configuration schema (Closes #798)
Add the declarative worker registry that epic #797 makes the source of
truth for the scheduled multi-LLM worker fleet.

Providers and configured workers are modelled as separate entities so a
provider can be listed with no worker configured, and so provider facts
are not copied into every worker record. A worker records provider,
model, project, role, namespace, profile, workflow, schedule, timeout,
enabled state, and scheduler metadata.

Validation fails closed: unknown fields are refused rather than ignored,
so a typo cannot silently disable a timeout; a worker naming an
undeclared provider is rejected; worker ids, provider ids, and
LaunchAgent labels must be unique.

Persistence is atomic (temp file in the same directory, fsync, replace).
Every superseded document is retained as a numbered revision, and
rollback republishes a chosen revision as a new head, so history stays
append-only and a rollback is itself reversible.

The credential-rejection guard is extracted to webui/registry_safety.py
so both registries share one implementation instead of two copies of a
security check; project_registry.py keeps identical behaviour.

Scope: data model, validation, persistence only. No routes, scheduler,
process control, or provider probing - those are #799/#800/#804/#805.
The workers array ships empty because populating it is #808.

Tests: tests/test_webui_worker_registry.py, 44 cases.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-07-22 06:10:34 -05:00
sysadmin 5032965e3a Merge pull request 'feat: make master-parity live-remote aware so a stale daemon fails closed (Closes #610)' (#788) from feat/issue-610-live-remote-parity into master 2026-07-22 05:51:56 -05:00
jcwalker3andClaude Opus 4.8 e33b8d3712 docs(webui): add MCP Control Plane Web Console architecture ADR (Closes #632)
Adds the durable architecture record epic #631 lacked: layer contract,
authority boundaries, the redaction boundary, /api/v1 versioning with a
compatibility rule for the existing unversioned MVP exports, a target page
map, per-child component ownership for all twenty children (#632-#651),
phase gates, and explicit forbidden paths.

Documentation only. No UI, API, or deployment change.

- docs/architecture/webui-control-plane-console-architecture-adr.md (new)
- docs/webui-local-dev.md: cross-link to the ADR
- tests/test_webui_architecture_docs.py: doc gate for the acceptance
  criteria and the cross-link (12 cases)

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-07-22 04:52:21 -05:00
13 changed files with 2773 additions and 117 deletions
@@ -0,0 +1,201 @@
# ADR: MCP Control Plane Web Console architecture and information architecture
- **Status:** Proposed (documentation only; blocks no code, gates every #631 child)
- **Date:** 2026-07-22
- **Tracking issue:** [#632](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/632) — architecture and information architecture (Phase 1)
- **Parent epic:** [#631](https://gitea.prgs.cc/Scaled-Tech-Consulting/Gitea-Tools/issues/631) — MCP Control Plane Web Console
- **Foundation (closed, extend — do not recreate):** #425 tracker and children #426 skeleton, #427 projects, #428 prompts, #429 queue, #430 runtime, #431 audit paste, #432 worktrees, #433 leases, #434 gated actions, #435 auth/deployment boundary, #436 tests/CI
- **Related:** `mcp-allocator-control-plane-observability-adr.md`, `mcp-stable-control-runtime-policy-adr.md`, `control-plane-db-substrate.md`, `../safety-model.md`, `../tool-boundaries.md`, `../credential-isolation.md`, `../webui-local-dev.md`, `../webui-deployment.md`
## 1. Context
The MVP web UI shipped under `webui/` as a read-only Starlette application with ten operator routes and a JSON export beside most of them. It is a working foundation, not the console product described by epic #631, and it carries no durable architecture record: no layer contract, no authority boundary, no API versioning rule, no page map, and no statement of which phase may open a write path.
Twenty children (#632#651) hang off #631. Without one architecture document each implementer re-derives boundaries, and the most likely failure is not a bad view — it is a privileged action wired into the browser before the authorization and audit model of #633 exists.
This ADR is the single retrievable design source for the console. It decides structure only. It implements no UI, no API, and no change to deployment topology.
## 2. Decision summary (core)
| Layer | Owns | Must not |
|-------|------|----------|
| **Browser UI** | Rendering, navigation, operator affordances | Hold tokens, call Gitea/providers directly, or execute an action the server did not gate |
| **HTTP route layer** (`webui/app.py`) | Versioned routing, authentication, authorization, redaction boundary, audit emission | Contain domain logic or reach past a loader to a raw credential |
| **Domain loaders** (`webui/*_loader.py`, `*_scanner.py`, `runtime_health.py`, `project_registry.py`) | Assembling read models from authoritative sources | Mutate anything, or emit unredacted secrets across the boundary |
| **Gitea** | Durable work record: issues, PRs, comments, reviews, labels, merges | Be the concurrency lock under multi-session load |
| **Control-plane DB** | Sessions, assignment, leases, heartbeats, events | Replace Gitea history |
| **MCP tools / capability gates** | Mutation authorization | Be re-implemented, mirrored, or bypassed by console code |
| **External providers** (Sentry/GlitchTip, AI providers) | Incident and usage data | Assign work or mutate Gitea outside the #612 bridge |
**One-liner:** **Gitea records. The DB coordinates. MCP tools authorize. The console projects state and executes only capability-checked, audited actions. Providers observe.**
## 3. Console surface today versus target
`webui/app.py` currently registers these routes (see `../webui-local-dev.md` for the operator-facing table): `/`, `/health`, `/queue`, `/projects`, `/projects/{id}`, `/prompts`, `/runtime`, `/audit`, `/worktrees`, `/leases`, `/actions`, and the unversioned exports `/api/queue`, `/api/projects`, `/api/prompts`, `/api/runtime`, `/api/audit`, `/api/worktrees`, `/api/leases`, `/api/actions`, `/api/actions/{id}/preview`, `/api/actions/{id}/attempt`.
Every one of these is **retained and evolved**. No child issue may recreate a route from scratch; each states in its PR which MVP surface it extends and what it changes.
## 4. Authority boundaries
### 4.1 Gitea (durable record)
Authoritative for issue and PR identity and state, comments, reviews and verdicts, labels, merges, and branch refs. When the console and Gitea disagree about durable state, Gitea wins and the console view is refreshed — never the reverse.
### 4.2 Control-plane DB (coordination)
Authoritative for live coordination: which session holds which assignment or lease, heartbeat freshness, expiry, and the allocation event log. The console reads it; only allocator and lease tools write it.
### 4.3 MCP capability gates (authorization)
`task_capability_map.py` and `gitea_resolve_task_capability` remain the only authority that decides whether a mutation may run. The console asks; it never answers. A console action that cannot name the MCP tool it delegates to is not an action — it is a defect.
### 4.4 Filesystem and git (local state)
Issue lock files, `branches/` worktrees, and registered git worktrees are read through existing scanners. The console never deletes, rebinds, or force-clears local state outside a Phase 2 gated action.
### 4.5 Providers (observe only)
Sentry/GlitchTip and AI providers are read surfaces. The #612 incident bridge is the only path that turns an observation into Gitea work.
## 5. Request flow and the redaction boundary
```text
browser ──HTTP──> route layer ──> domain loader ──> Gitea REST
│ ├──> control-plane DB
│ ├──> filesystem / git
│ └──> providers
[redaction boundary]
audit event
```
| Stage | May hold credentials | Emits |
|-------|----------------------|-------|
| Loader → route layer | yes (server-side, via `gitea_auth`) | domain objects |
| Route layer → browser | **no** | redacted DTOs, HTML |
Two invariants govern the boundary and are non-negotiable for every child:
1. **No secrets to the browser.** Tokens, keychain identifiers, Authorization headers, raw provider endpoints, and credential-bearing URLs are redacted by default, consistent with `../safety-model.md` §3 and `../credential-isolation.md`. Serializers redact; templates do not sanitize after the fact.
2. **No ungated mutations.** A write reaches an authoritative system only by delegating to an MCP tool that passed its own capability gate. HTML forms and JSON endpoints are transport, never authority.
## 6. API naming and versioning
**Decision:** all console APIs added from Phase 1 onward are served under `/api/v1/...`.
- Nouns are plural and hierarchical: `/api/v1/inventory/leases`, `/api/v1/system/health`.
- Read endpoints are `GET` and side-effect free.
- Phase 2 action endpoints are `POST /api/v1/actions/{action_id}/preview` and `POST /api/v1/actions/{action_id}/execute`; `preview` stays side-effect free and returns a mutation ledger.
- The existing unversioned MVP exports remain as **compatibility aliases** for the whole of Phase 1 so the current operator flow never breaks. They may be retired no earlier than Phase 2, and only after the replacing `v1` route ships and `../webui-local-dev.md` records the swap.
- A breaking change to a `v1` payload requires `/api/v2/...`, not an in-place edit.
- Every JSON payload carries enough provenance for an auditor to tell where the data came from — at minimum the source system and whether the inventory was complete, matching the pagination-proof habit the MVP queue export already established.
## 7. Page map
| Page | Purpose | Owning child | Evolves |
|------|---------|--------------|---------|
| `/` | Console shell, navigation, next-safe-action summary | #638 | MVP `/` (#426) |
| `/system` | System-health dashboard | #639 | new, backed by #634 |
| `/traffic` | Workflow traffic control, queues, blockers | #640 | MVP `/queue` (#429) |
| `/runtime` | Runtime and session view | #641 | MVP `/runtime` (#430) |
| `/projects`, `/projects/{id}` | Project registry and onboarding | #635 | MVP `/projects` (#427) |
| `/inventory` | Sessions, leases, locks, worktrees in one surface | #636 | MVP `/leases` (#433) + `/worktrees` (#432) |
| `/timeline` | Workflow events and conversation timeline | #637 | new |
| `/actions` | Gated action registry, preview, execution | #642, #643, #644 | MVP `/actions` (#434) |
| `/gitea` | Issue and PR linkage console | #645 | new |
| `/policy` | Guardrail visibility, then versioned editing | #646, #647 | new |
| `/notifications` | Human-attention routing | #648 | new |
| `/observability` | Sentry/GlitchTip correlation and durable issue creation | #649 | new |
| `/providers` | AI-provider connections and insights | #650 | new |
| `/analytics` | Usage, token cost, latency, workflow performance | #651 | new |
| `/audit` | Final-report validator preview and audit log | #431 foundation, extended by #633 | MVP `/audit` (#431) |
| `/prompts`, `/prompts/{id}` | Canonical prompt library | #638 | MVP `/prompts` (#428) |
| `/health` | Liveness and deployment metadata | #634 | MVP `/health` (#435) |
## 8. Component ownership for every epic child
Each #631 child maps to at least one architectural component defined above.
| Child | Capability area | Primary component | Phase |
|-------|-----------------|-------------------|-------|
| #632 | Architecture and information architecture | this ADR | 1 |
| #633 | Authorization, RBAC, secret redaction, audit and retention | route layer + redaction boundary (§5) | 1 |
| #634 | Read-only system-health API | `/api/v1/system/health` + health loader | 1 |
| #635 | Project registry API evolution | `/api/v1/projects` + `project_registry.py` | 1 |
| #636 | Session, lease, lock, worktree inventory API | `/api/v1/inventory/*` + `lease_loader.py`, `worktree_scanner.py` | 1 |
| #637 | Workflow-event and conversation timeline model | `/api/v1/events` + control-plane DB event log | 1 |
| #638 | Application shell evolution | browser UI layer + `layout.py` | 1 |
| #639 | System-health dashboard | `/system` page over #634 | 1 |
| #640 | Workflow traffic-control view | `/traffic` page over the queue loader | 1 |
| #641 | Runtime and session view | `/runtime` page over `runtime_health.py` | 1 |
| #642 | Sanctioned restart and graceful reload controls | gated action framework, restart class | 2 |
| #643 | Requests, intent preview, authorization, workflow initiation | `/api/v1/actions/*` execute path | 2 |
| #644 | Stale-runtime recovery, worktree rebinding, reconciliation controls | gated actions over filesystem/git authority | 2 |
| #645 | Gitea issue and PR linkage console | `/gitea` page over Gitea authority | 3 |
| #646 | Workflow policy and guardrail visibility | `/policy` read view over the capability map | 3 |
| #647 | Versioned policy editing, validation, simulation, approval, rollback | `/policy` write path, gated | 3 |
| #648 | Notifications and human-attention routing | notification component over the event model | 3 |
| #649 | Sentry/GlitchTip connections, correlation, durable issue creation | provider layer + #612 incident bridge | 4 |
| #650 | AI-provider connections and operational insights | provider layer | 4 |
| #651 | Model usage, token cost, latency, workflow analytics | analytics component over the event model | 4 |
Related but **outside** this epic: #667 (restart status, impact preview, and approval controls) belongs to the #655 restart-governance umbrella and must reuse the #642 action class rather than adding a second restart surface.
## 9. Phase gates
| Phase | May ship | Entry condition |
|-------|----------|-----------------|
| **1 — read-only visibility** | `GET` pages and `GET /api/v1/...` | this ADR accepted |
| **2 — controlled actions** | gated `POST` action execution | #633 authorization, RBAC, and audit model landed |
| **3 — orchestration and policy** | linkage, policy visibility, versioned policy editing | Phase 1 inventory plus the Phase 2 action framework |
| **4 — insights** | provider correlation, analytics | evidence-backed sources from Phases 13 |
Phase 1 must not open a mutation endpoint, and the read-only guard that returns `405 read-only-mvp` stays in force until the Phase 2 entry condition is met. A phase is not entered by exception; if a control is urgent, the entry condition is what gets prioritized.
## 10. Security and workflow safety
- **Fail closed** on unknown authentication, missing RBAC mapping, or ambiguous lease ownership. An unknown state renders as blocked, never as permitted.
- **Redact by default**, per §5.
- **Every privileged action** requires a resolved capability, an explicit operator confirmation, and a durable audit event naming actor, action, target, and outcome.
- **Contamination surfaces.** Session contamination — including a manually killed MCP daemon (#630) — must be shown and must block clean claims rather than being silently repaired.
- **Deployment boundary unchanged.** Loopback by default, with the existing refusal of public binds (#435). This ADR documents that target; it does not widen it.
## 11. Forbidden paths
These are rejected designs, not preferences:
1. **Raw provider incidents as work.** The allocator never receives an unclassified Sentry/GlitchTip incident; only the #612 bridge turns an observation into a Gitea issue.
2. **Browser-held tokens.** No credential, keychain identifier, or Authorization header is ever sent to the browser or embedded in a client bundle.
3. **Process-kill recovery.** The console must not expose `pkill`, process-identifier termination, or any host process kill as a recovery affordance (#630). Restart is the sanctioned, operator-owned path of #642 and the #655 umbrella.
4. **Ungated browser mutations.** No review, approval, merge, close, or comment may originate from the browser without passing an MCP capability gate.
5. **Policy invented in the console.** The console projects policy from the capability map and canonical workflows; it never encodes a second copy.
6. **Recreating MVP scope.** Re-implementing a #426#436 surface without an explicit evolve-or-extend statement is out of bounds.
## 12. Approval checklist (readable without chat history)
A controller can accept or reject this ADR against these six points alone:
1. Layers and their owners are defined (§2) and each authority is named (§4).
2. The redaction boundary and the two invariants are stated (§5).
3. API versioning is decided, including what happens to the existing unversioned routes (§6).
4. A page map exists and names an owning child for every page (§7).
5. Every #631 child maps to at least one component and one phase (§8).
6. Phase gates and forbidden paths are explicit (§9, §11).
## 13. Open questions and follow-ups
Unresolved choices are recorded here rather than settled by implication. Each needs its own durable issue before the phase that depends on it:
- **Authentication mechanism.** Whether the console authenticates via an access proxy (Cloudflare Access or equivalent) or an application-level session is deferred to #633. This ADR requires only that it fail closed.
- **Event model substrate.** Whether the #637 timeline reads the control-plane event log directly or through a projection is deferred to #637.
- **CI path filter coverage.** `webui/ci_paths.py` triggers the web UI suite on `webui/`, `tests/test_webui_*`, and `docs/webui*`. This ADR lives under `docs/architecture/`, so editing it alone does not trigger that gate; the accompanying `tests/test_webui_architecture_docs.py` does run in the full suite. Widening the filter is a small follow-up, deliberately not bundled into a documentation-only change.
- **Retention.** Audit-event retention duration is owned by #633.
## 14. Acceptance
Accepting this ADR means:
- Phase 1 children may proceed against the layers, page map, and API rules above.
- Phase 2 children may not open a write path until #633 lands.
- Any deviation is recorded as an amendment to this file with its own issue reference, not as an undocumented divergence in code.
+14 -2
View File
@@ -37,6 +37,16 @@ Optional environment variables:
See [webui-deployment.md](webui-deployment.md) for internal-only serving, See [webui-deployment.md](webui-deployment.md) for internal-only serving,
Cloudflare Access/WARP/VPN guidance, and unsafe bind overrides (#435). Cloudflare Access/WARP/VPN guidance, and unsafe bind overrides (#435).
See
[architecture/webui-control-plane-console-architecture-adr.md](architecture/webui-control-plane-console-architecture-adr.md)
for the console architecture: layer and authority boundaries, the redaction
boundary, `/api/v1/...` versioning, the target page map, and the phase gates
that govern when a write path may open (#632, epic #631).
See [webui-project-registry-api.md](webui-project-registry-api.md) for the
versioned project registry contract: registry schema versions 1 and 2, project
status, onboarding checklist state, and the fail-closed error payloads (#635).
## Routes (MVP) ## Routes (MVP)
| Path | Description | | Path | Description |
@@ -45,9 +55,11 @@ Cloudflare Access/WARP/VPN guidance, and unsafe bind overrides (#435).
| `/health` | JSON liveness (`status`, `service`, `mode`, `timestamp`) | | `/health` | JSON liveness (`status`, `service`, `mode`, `timestamp`) |
| `/queue` | Live PR and issue queue dashboard (#429) | | `/queue` | Live PR and issue queue dashboard (#429) |
| `/api/queue` | JSON queue export with pagination metadata | | `/api/queue` | JSON queue export with pagination metadata |
| `/projects` | Project registry list (#427) | | `/projects` | Project registry list with status and onboarding progress (#427, #635) |
| `/projects/{id}` | Project detail + onboarding checklist | | `/projects/{id}` | Project detail + onboarding checklist |
| `/api/projects` | JSON registry export | | `/api/v1/projects` | Versioned JSON registry export (#635) |
| `/api/v1/projects/{id}` | Versioned JSON project detail (#635) |
| `/api/projects` | JSON registry export — unversioned Phase 1 alias of `/api/v1/projects` |
| `/prompts` | Prompt library with per-prompt copy buttons (#428) | | `/prompts` | Prompt library with per-prompt copy buttons (#428) |
| `/api/prompts` | JSON prompt export with workflow hashes | | `/api/prompts` | JSON prompt export with workflow hashes |
| `/runtime` | MCP runtime health and stale detection (#430) | | `/runtime` | MCP runtime health and stale detection (#430) |
+213
View File
@@ -0,0 +1,213 @@
# Project registry API (#635)
Phase 1 of the [console architecture ADR](architecture/webui-control-plane-console-architecture-adr.md)
gives the project registry a versioned, read-only API. This document is the
field-by-field contract for that API and for the registry file behind it.
Everything here is **read-only**. The console never writes the registry; an
operator edits the JSON file, and an invalid file fails closed rather than
rendering a partial inventory.
## Routes
| Route | Method | Description |
|-------|--------|-------------|
| `/api/v1/projects` | GET | Versioned registry export: all projects, with provenance |
| `/api/v1/projects/{project_id}` | GET | Single project; `404` with `project_not_found` when unknown |
| `/api/projects` | GET | Unversioned MVP alias (#427), retained for all of Phase 1 |
| `/projects` | GET | HTML list — status and onboarding progress per project |
| `/projects/{project_id}` | GET | HTML detail — identity, profiles, paths, checklist |
Per ADR section 6 the unversioned alias may be retired no earlier than Phase 2,
and only after this document and `webui-local-dev.md` record the swap. The alias
returns the same payload as `/api/v1/projects`, including the legacy `version`
and `source_path` keys #427 consumers already read.
The HTML views render from the same DTO the JSON routes serialize
(`project_to_dict`), so the console and the API cannot disagree about a
project's status or onboarding progress.
## Registry file
Default location: `webui/data/projects.registry.json`. Override with the
`WEBUI_PROJECT_REGISTRY` environment variable.
Schema versions: **1** and **2** are accepted; **2** is current. A version 1
file loads unchanged and is normalized with the documented defaults, so an
existing operator registry keeps working without edits.
### Root
| Field | Type | Required | Notes |
|-------|------|----------|-------|
| `version` | int | yes | `1` or `2`. Anything else fails closed |
| `projects` | array | yes | Must be non-empty |
### Project
| Field | Type | Required | Default | Notes |
|-------|------|----------|---------|-------|
| `id` | string | yes | — | Stable registry id used in URLs |
| `repo_name` | string | yes | — | Gitea repository name |
| `gitea_owner` | string | yes | — | Owning org or user |
| `remote_host` | string | yes | — | Instance base URL, no credentials |
| `remote_name` | string | no | `null` | Logical remote label, e.g. `prgs` (v2) |
| `default_branch` | string | yes | — | Stable branch name |
| `local_checkout_path` | string | yes | — | Control checkout path |
| `status` | string | no | `active` | `active`, `onboarding`, `paused`, `archived` (v2) |
| `profiles` | object | yes | — | Must map `author`, `reviewer`, `reconciler` |
| `workflow_paths` | object | yes | — | Non-empty; label to repo-relative path |
| `schema_paths` | object | no | `{}` | Label to repo-relative path |
| `onboarding_checklist` | array | no | `[]` | See below |
| `last_seen_health` | object | no | `null` | Redacted health only (v2) |
### Onboarding step
| Field | Type | Required | Default | Notes |
|-------|------|----------|---------|-------|
| `id` | string | yes | — | Stable step id |
| `title` | string | yes | — | Short operator-facing label |
| `description` | string | yes | — | Self-contained; assumes no chat history |
| `state` | string | no | `pending` | `complete`, `pending`, `blocked`, `not_applicable` (v2) |
| `required` | bool | no | `true` | Optional steps never block readiness (v2) |
### Last-seen health
| Field | Type | Required | Notes |
|-------|------|----------|-------|
| `status` | string | no (default `unknown`) | `healthy`, `degraded`, `unreachable`, `unknown` |
| `checked_at` | string | no | ISO-8601 UTC timestamp, e.g. `2026-01-01T00:00:00Z` |
| `detail` | string | no | Short redacted note |
Health is recorded metadata, not a live probe: Phase 1 performs no outbound
health checks. Endpoints, tokens, and keychain identifiers must never appear
here.
## Response shape
`GET /api/v1/projects`:
```json
{
"api_version": "v1",
"schema_version": 2,
"version": 2,
"source_path": "/path/to/webui/data/projects.registry.json",
"source": {
"kind": "file",
"path": "/path/to/webui/data/projects.registry.json",
"inventory_complete": true
},
"project_count": 1,
"projects": [
{
"id": "example",
"repo_name": "Example",
"gitea_owner": "Org",
"repo_full_name": "Org/Example",
"remote_host": "https://gitea.example.invalid",
"remote_name": "example-remote",
"default_branch": "main",
"local_checkout_path": ".",
"status": "active",
"profiles": {"author": "...", "reviewer": "...", "reconciler": "..."},
"workflow_paths": {"skill": "skills/..."},
"schema_paths": {},
"onboarding_checklist": [
{
"id": "profiles",
"title": "Configure execution profiles",
"description": "...",
"state": "complete",
"required": true
}
],
"onboarding_summary": {
"total": 1,
"complete": 1,
"pending": 0,
"blocked": 0,
"not_applicable": 0,
"required_outstanding": 0,
"onboarding_complete": true
},
"last_seen_health": null
}
]
}
```
`GET /api/v1/projects/{project_id}` returns `api_version`, `schema_version`,
`source`, and a single `project` object with the same fields.
The `source` block satisfies the ADR section 6 provenance rule: every payload
states where the data came from and whether the inventory is complete. A
file-backed registry is always complete — there is no pagination to truncate it.
`onboarding_summary` is derived, never stored. `required_outstanding` counts
steps that are `required` **and** in state `pending` or `blocked`;
`onboarding_complete` is true when that count is zero.
## Fail-closed errors
Validation failures raise `RegistryError`, which routes render instead of a
traceback.
`404` — unknown project id on `/api/v1/projects/{project_id}`:
```json
{
"error": "project_not_found",
"project_id": "not-registered",
"known_project_ids": ["example"],
"remediation": "Request one of the known project ids, or add the project ...",
"source": {"kind": "file", "path": "...", "inventory_complete": true}
}
```
`500` — invalid registry, on both the versioned route and the alias:
```json
{
"error": "registry_invalid",
"detail": "unsupported registry version: 42",
"remediation": "Set 'version' to one of 1, 2 (current schema is 2) ...",
"field_path": "version",
"source_path": "/path/to/registry.json"
}
```
`field_path` points at the offending location (`projects[0].profiles.reconciler`,
`projects[0].onboarding_checklist[2].state`, and so on). The HTML routes render
the same detail, field, source, and remediation on a "Project registry
unavailable" page.
Conditions that fail closed:
* file missing or unreadable;
* invalid JSON (the remediation names line and column);
* root not an object, or `projects` missing/empty;
* unsupported `version`;
* a credential-shaped key anywhere in the file (`token`, `*_secret`, `auth_*`, and similar);
* a project missing a required field, or missing an `author`/`reviewer`/`reconciler` profile;
* an unknown `status`, onboarding `state`, or health `status`.
## Credential rule
The registry stores redacted metadata only. Credential-shaped keys are
rejected at load time, before any DTO is built, consistent with
[safety-model.md](safety-model.md) and
[credential-isolation.md](credential-isolation.md). Tokens live in the keychain
and are resolved server-side by `gitea_auth`.
## Migrating a version 1 registry
1. Set `"version": 2`.
2. Optionally add `"status"` per project (omitted means `active`).
3. Optionally add `"remote_name"` per project.
4. Optionally add `"state"` and `"required"` to each onboarding step (omitted
means `pending` and `true`).
5. Optionally add `"last_seen_health"`.
No step is mandatory: a version 1 file keeps loading. Bumping the version only
declares that the file may use the v2 fields.
+149
View File
@@ -0,0 +1,149 @@
"""Documentation acceptance for the web console architecture ADR (#632 / epic #631).
Enforces the acceptance criteria of issue #632:
* AC1 — the ADR exists and covers layers, authority, phases, API versioning,
and a page map.
* AC2 — every #631 child (#632#651) maps to at least one architectural
component.
* AC3 — the closed MVP (#425#436) is stated as foundation, not recreated.
* AC4 — forbidden paths are explicit: raw provider incidents as work,
browser-held tokens, process-kill recovery.
* AC5 — a controller can approve the document without reading chat history.
Plus the linkage requirement: ``docs/webui-local-dev.md`` cross-links the ADR.
"""
from pathlib import Path
REPO_ROOT = Path(__file__).resolve().parent.parent
ADR = (
REPO_ROOT
/ "docs"
/ "architecture"
/ "webui-control-plane-console-architecture-adr.md"
)
ADR_BASENAME = "webui-control-plane-console-architecture-adr.md"
LOCAL_DEV = REPO_ROOT / "docs" / "webui-local-dev.md"
# Epic #631 children, phases 1-4 (twenty capability areas).
EPIC_CHILDREN = tuple(f"#{number}" for number in range(632, 652))
def _read(path: Path) -> str:
assert path.is_file(), f"missing {path.relative_to(REPO_ROOT)}"
return path.read_text(encoding="utf-8")
def test_ac1_adr_exists_with_required_sections():
text = _read(ADR)
lower = text.lower()
assert text.lstrip().startswith("#"), "ADR lacks a title"
assert "#631" in text and "#632" in text
for heading in (
"## 2. Decision summary",
"## 4. Authority boundaries",
"## 5. Request flow and the redaction boundary",
"## 6. API naming and versioning",
"## 7. Page map",
"## 8. Component ownership",
"## 9. Phase gates",
"## 11. Forbidden paths",
):
assert heading in text, f"ADR must contain section {heading!r}"
assert "browser ui" in lower and "domain loader" in lower
assert "control-plane db" in lower and "capability gate" in lower
def test_ac1_api_versioning_is_decided_including_legacy_routes():
text = _read(ADR)
assert "/api/v1/" in text, "ADR must decide the versioned API prefix"
assert "/api/v2/" in text, "ADR must state how breaking changes are handled"
lower = text.lower()
assert "compatibility alias" in lower, (
"ADR must say what happens to the existing unversioned MVP exports"
)
def test_ac1_page_map_covers_mvp_routes():
text = _read(ADR)
for route in ("`/`", "`/health`", "`/projects`", "`/prompts`", "`/runtime`",
"`/audit`", "`/actions`"):
assert route in text, f"page map must account for MVP route {route}"
def test_ac2_every_epic_child_maps_to_a_component():
text = _read(ADR)
ownership = text.split("## 8. Component ownership", 1)[-1].split("## 9.", 1)[0]
missing = [child for child in EPIC_CHILDREN if child not in ownership]
assert not missing, (
f"epic #631 children without an architectural component: {missing}"
)
def test_ac2_every_child_row_declares_a_phase():
text = _read(ADR)
ownership = text.split("## 8. Component ownership", 1)[-1].split("## 9.", 1)[0]
for child in EPIC_CHILDREN:
row = next(
(line for line in ownership.splitlines() if line.startswith(f"| {child} ")),
None,
)
assert row is not None, f"no ownership row for {child}"
assert row.rstrip().endswith(("| 1 |", "| 2 |", "| 3 |", "| 4 |")), (
f"ownership row for {child} must end with its phase: {row!r}"
)
def test_ac3_mvp_is_foundation_not_recreated():
text = _read(ADR)
assert "#425" in text and "#436" in text
lower = text.lower()
assert "do not recreate" in lower or "recreating mvp scope" in lower
assert "retained and evolved" in lower
def test_ac4_forbidden_paths_are_explicit():
text = _read(ADR)
forbidden = text.split("## 11. Forbidden paths", 1)[-1].split("## 12.", 1)[0]
lower = forbidden.lower()
assert "raw provider incidents" in lower and "#612" in forbidden
assert "browser-held tokens" in lower
assert "process-kill recovery" in lower and "#630" in forbidden
assert "ungated browser mutations" in lower
def test_ac5_approval_checklist_is_self_contained():
text = _read(ADR)
assert "## 12. Approval checklist" in text
checklist = text.split("## 12. Approval checklist", 1)[-1].split("## 13.", 1)[0]
for marker in ("1.", "2.", "3.", "4.", "5.", "6."):
assert marker in checklist, f"approval checklist missing item {marker}"
def test_adr_states_the_two_boundary_invariants():
text = _read(ADR)
lower = text.lower()
assert "no secrets to the browser" in lower
assert "no ungated mutations" in lower
def test_open_questions_are_recorded_not_implied():
text = _read(ADR)
assert "## 13. Open questions and follow-ups" in text
section = text.split("## 13. Open questions and follow-ups", 1)[-1]
assert "#633" in section, "deferred authorization work must name its issue"
def test_local_dev_doc_cross_links_the_adr():
text = _read(LOCAL_DEV)
assert ADR_BASENAME in text, (
"docs/webui-local-dev.md must cross-link the console architecture ADR "
"(issue #632 scope)"
)
def test_docs_do_not_embed_secrets():
for path in (ADR, LOCAL_DEV):
text = _read(path)
for marker in ("ghp_", "BEGIN PRIVATE KEY", "Authorization: Bearer"):
assert marker not in text, f"{path.name} contains {marker!r}"
+336 -31
View File
@@ -1,4 +1,4 @@
"""Tests for web UI project registry (#427).""" """Tests for web UI project registry (#427) and its API evolution (#635)."""
import json import json
import sys import sys
import tempfile import tempfile
@@ -11,57 +11,246 @@ from starlette.testclient import TestClient
from webui.app import create_app from webui.app import create_app
from webui.project_registry import ( from webui.project_registry import (
CURRENT_SCHEMA_VERSION,
REGISTRY_API_VERSION,
SUPPORTED_SCHEMA_VERSIONS,
RegistryError,
default_registry_path, default_registry_path,
load_registry, load_registry,
onboarding_summary,
project_to_dict, project_to_dict,
) )
from webui.registry_safety import is_forbidden_key
_REPO_ROOT = Path(__file__).resolve().parent.parent
_API_DOC = _REPO_ROOT / "docs" / "webui-project-registry-api.md"
class TestProjectRegistryLoader(unittest.TestCase): def _valid_project(**overrides):
project = {
"id": "example",
"repo_name": "Example",
"gitea_owner": "Org",
"remote_host": "https://gitea.example.invalid",
"default_branch": "main",
"local_checkout_path": ".",
"profiles": {"author": "a", "reviewer": "r", "reconciler": "c"},
"workflow_paths": {"skill": "skills/x.md"},
}
project.update(overrides)
return project
def _write_registry(payload) -> Path:
with tempfile.NamedTemporaryFile("w", suffix=".json", delete=False) as handle:
json.dump(payload, handle)
return Path(handle.name)
class RegistryFileCase(unittest.TestCase):
"""Base class that cleans up temporary registry files."""
def setUp(self):
self._temp_paths: list[Path] = []
def tearDown(self):
for path in self._temp_paths:
path.unlink(missing_ok=True)
def write_registry(self, payload) -> Path:
path = _write_registry(payload)
self._temp_paths.append(path)
return path
class TestProjectRegistryLoader(RegistryFileCase):
def test_default_registry_loads_gitea_tools(self): def test_default_registry_loads_gitea_tools(self):
registry = load_registry() registry = load_registry()
self.assertEqual(registry.version, 1) self.assertEqual(registry.version, CURRENT_SCHEMA_VERSION)
self.assertEqual(registry.schema_version, CURRENT_SCHEMA_VERSION)
self.assertEqual(registry.api_version, REGISTRY_API_VERSION)
self.assertEqual(len(registry.projects), 1) self.assertEqual(len(registry.projects), 1)
project = registry.projects[0] project = registry.projects[0]
self.assertEqual(project.id, "gitea-tools") self.assertEqual(project.id, "gitea-tools")
self.assertEqual(project.repo_name, "Gitea-Tools") self.assertEqual(project.repo_name, "Gitea-Tools")
self.assertEqual(project.gitea_owner, "Scaled-Tech-Consulting") self.assertEqual(project.gitea_owner, "Scaled-Tech-Consulting")
self.assertEqual(project.repo_full_name, "Scaled-Tech-Consulting/Gitea-Tools")
self.assertEqual(project.remote_host, "https://gitea.prgs.cc") self.assertEqual(project.remote_host, "https://gitea.prgs.cc")
self.assertEqual(project.remote_name, "prgs")
self.assertEqual(project.status, "active")
self.assertEqual(project.profiles["author"], "prgs-author") self.assertEqual(project.profiles["author"], "prgs-author")
self.assertEqual(project.profiles["reviewer"], "prgs-reviewer") self.assertEqual(project.profiles["reviewer"], "prgs-reviewer")
self.assertEqual(project.profiles["reconciler"], "prgs-reconciler") self.assertEqual(project.profiles["reconciler"], "prgs-reconciler")
self.assertIn("skill", project.workflow_paths) self.assertIn("skill", project.workflow_paths)
self.assertGreaterEqual(len(project.onboarding_checklist), 4) self.assertGreaterEqual(len(project.onboarding_checklist), 4)
def test_registry_rejects_credential_keys(self): def test_default_registry_onboarding_summary_is_complete(self):
payload = { summary = onboarding_summary(load_registry().projects[0])
self.assertEqual(summary.total, summary.complete)
self.assertEqual(summary.required_outstanding, 0)
self.assertTrue(summary.onboarding_complete)
def test_version_1_registry_still_loads_with_defaults(self):
path = self.write_registry({
"version": 1, "version": 1,
"projects": [ "projects": [
{ _valid_project(
"id": "bad", onboarding_checklist=[
"repo_name": "Bad", {"id": "step", "title": "Step", "description": "Do it"}
"gitea_owner": "Org", ]
"remote_host": "https://gitea.example.invalid", )
"default_branch": "main",
"local_checkout_path": ".",
"profiles": {
"author": "a",
"reviewer": "r",
"reconciler": "c",
},
"workflow_paths": {"skill": "skills/x.md"},
"api_token": "secret",
}
], ],
} })
registry = load_registry(path)
self.assertEqual(registry.schema_version, 1)
self.assertIn(1, SUPPORTED_SCHEMA_VERSIONS)
project = registry.projects[0]
self.assertEqual(project.status, "active")
self.assertIsNone(project.remote_name)
self.assertIsNone(project.last_seen_health)
step = project.onboarding_checklist[0]
self.assertEqual(step.state, "pending")
self.assertTrue(step.required)
self.assertFalse(onboarding_summary(project).onboarding_complete)
def test_onboarding_summary_counts_states(self):
path = self.write_registry({
"version": 2,
"projects": [
_valid_project(
onboarding_checklist=[
{"id": "a", "title": "A", "description": "d", "state": "complete"},
{"id": "b", "title": "B", "description": "d", "state": "blocked"},
{
"id": "c",
"title": "C",
"description": "d",
"state": "pending",
"required": False,
},
{
"id": "d",
"title": "D",
"description": "d",
"state": "not_applicable",
},
]
)
],
})
summary = onboarding_summary(load_registry(path).projects[0])
self.assertEqual(summary.total, 4)
self.assertEqual(summary.complete, 1)
self.assertEqual(summary.blocked, 1)
self.assertEqual(summary.pending, 1)
self.assertEqual(summary.not_applicable, 1)
# Only the blocked step is both required and outstanding.
self.assertEqual(summary.required_outstanding, 1)
self.assertFalse(summary.onboarding_complete)
def test_last_seen_health_is_parsed_when_present(self):
path = self.write_registry({
"version": 2,
"projects": [
_valid_project(
last_seen_health={
"status": "degraded",
"checked_at": "2026-01-01T00:00:00Z",
"detail": "daemon restart pending",
}
)
],
})
health = load_registry(path).projects[0].last_seen_health
self.assertIsNotNone(health)
self.assertEqual(health.status, "degraded")
self.assertEqual(health.checked_at, "2026-01-01T00:00:00Z")
def test_registry_rejects_credential_keys(self):
path = self.write_registry({
"version": 1,
"projects": [_valid_project(id="bad", api_token="redacted-placeholder")],
})
with self.assertRaises(RegistryError) as ctx:
load_registry(path)
self.assertIn("credential", ctx.exception.remediation.lower())
self.assertEqual(ctx.exception.field_path, "projects[0].api_token")
def test_unsupported_version_fails_closed_with_remediation(self):
path = self.write_registry({"version": 99, "projects": [_valid_project()]})
with self.assertRaises(RegistryError) as ctx:
load_registry(path)
self.assertIn("unsupported registry version", ctx.exception.message)
self.assertIn(str(CURRENT_SCHEMA_VERSION), ctx.exception.remediation)
self.assertEqual(ctx.exception.field_path, "version")
def test_missing_required_field_fails_closed(self):
broken = _valid_project()
del broken["default_branch"]
path = self.write_registry({"version": 2, "projects": [broken]})
with self.assertRaises(RegistryError) as ctx:
load_registry(path)
self.assertIn("default_branch", ctx.exception.message)
self.assertEqual(ctx.exception.field_path, "projects[0]")
def test_unknown_status_fails_closed(self):
path = self.write_registry({
"version": 2,
"projects": [_valid_project(status="mystery")],
})
with self.assertRaises(RegistryError) as ctx:
load_registry(path)
self.assertEqual(ctx.exception.field_path, "projects[0].status")
self.assertIn("active", ctx.exception.remediation)
def test_unknown_onboarding_state_fails_closed(self):
path = self.write_registry({
"version": 2,
"projects": [
_valid_project(
onboarding_checklist=[
{"id": "a", "title": "A", "description": "d", "state": "almost"}
]
)
],
})
with self.assertRaises(RegistryError) as ctx:
load_registry(path)
self.assertEqual(
ctx.exception.field_path,
"projects[0].onboarding_checklist[0].state",
)
def test_missing_profile_role_fails_closed(self):
path = self.write_registry({
"version": 2,
"projects": [_valid_project(profiles={"author": "a", "reviewer": "r"})],
})
with self.assertRaises(RegistryError) as ctx:
load_registry(path)
self.assertEqual(ctx.exception.field_path, "projects[0].profiles.reconciler")
def test_empty_projects_fails_closed(self):
path = self.write_registry({"version": 2, "projects": []})
with self.assertRaises(RegistryError) as ctx:
load_registry(path)
self.assertEqual(ctx.exception.field_path, "projects")
def test_invalid_json_fails_closed_with_location(self):
with tempfile.NamedTemporaryFile("w", suffix=".json", delete=False) as handle: with tempfile.NamedTemporaryFile("w", suffix=".json", delete=False) as handle:
json.dump(payload, handle) handle.write("{not json")
path = Path(handle.name) path = Path(handle.name)
try: self._temp_paths.append(path)
with self.assertRaises(ValueError): with self.assertRaises(RegistryError) as ctx:
load_registry(path) load_registry(path)
finally: self.assertIn("not valid JSON", ctx.exception.message)
path.unlink(missing_ok=True) self.assertIn("line", ctx.exception.remediation)
def test_missing_file_fails_closed(self):
missing = Path(tempfile.gettempdir()) / "webui-registry-does-not-exist.json"
with self.assertRaises(RegistryError) as ctx:
load_registry(missing)
self.assertIn("could not be read", ctx.exception.message)
def test_default_registry_path_points_at_packaged_data(self): def test_default_registry_path_points_at_packaged_data(self):
path = default_registry_path() path = default_registry_path()
@@ -81,31 +270,147 @@ class TestProjectRegistryRoutes(unittest.TestCase):
self.assertIn("prgs-author", response.text) self.assertIn("prgs-author", response.text)
self.assertNotIn("child issue", response.text.lower()) self.assertNotIn("child issue", response.text.lower())
def test_projects_page_shows_status_and_progress(self):
response = self.client.get("/projects")
self.assertIn("Status", response.text)
self.assertIn("Onboarding", response.text)
self.assertIn("4/4 complete", response.text)
def test_project_detail_renders_checklist(self): def test_project_detail_renders_checklist(self):
response = self.client.get("/projects/gitea-tools") response = self.client.get("/projects/gitea-tools")
self.assertEqual(response.status_code, 200) self.assertEqual(response.status_code, 200)
self.assertIn("Onboarding checklist", response.text) self.assertIn("Onboarding checklist", response.text)
self.assertIn("Configure execution profiles", response.text) self.assertIn("Configure execution profiles", response.text)
self.assertIn("branches/", response.text) self.assertIn("branches/", response.text)
self.assertIn("Complete", response.text)
self.assertIn("required outstanding 0", response.text)
def test_project_detail_404(self): def test_project_detail_404(self):
response = self.client.get("/projects/unknown-repo") response = self.client.get("/projects/unknown-repo")
self.assertEqual(response.status_code, 404) self.assertEqual(response.status_code, 404)
def test_api_projects_json(self): def test_api_projects_alias_stays_compatible(self):
response = self.client.get("/api/projects") response = self.client.get("/api/projects")
self.assertEqual(response.status_code, 200) self.assertEqual(response.status_code, 200)
data = response.json() data = response.json()
self.assertEqual(data["version"], 1) # #427 consumers keep these keys.
self.assertEqual(data["version"], CURRENT_SCHEMA_VERSION)
self.assertIn("source_path", data)
self.assertEqual(len(data["projects"]), 1) self.assertEqual(len(data["projects"]), 1)
self.assertEqual(data["projects"][0]["id"], "gitea-tools") self.assertEqual(data["projects"][0]["id"], "gitea-tools")
self.assertIn("onboarding_checklist", data["projects"][0]) self.assertIn("onboarding_checklist", data["projects"][0])
def test_api_v1_projects_payload(self):
response = self.client.get("/api/v1/projects")
self.assertEqual(response.status_code, 200)
data = response.json()
self.assertEqual(data["api_version"], REGISTRY_API_VERSION)
self.assertEqual(data["schema_version"], CURRENT_SCHEMA_VERSION)
self.assertEqual(data["project_count"], 1)
self.assertEqual(data["source"]["kind"], "file")
self.assertTrue(data["source"]["inventory_complete"])
project = data["projects"][0]
self.assertEqual(project["status"], "active")
self.assertEqual(project["remote_name"], "prgs")
self.assertEqual(
project["repo_full_name"], "Scaled-Tech-Consulting/Gitea-Tools"
)
self.assertTrue(project["onboarding_summary"]["onboarding_complete"])
self.assertEqual(project["onboarding_checklist"][0]["state"], "complete")
self.assertIsNone(project["last_seen_health"])
def test_api_v1_project_detail(self):
response = self.client.get("/api/v1/projects/gitea-tools")
self.assertEqual(response.status_code, 200)
data = response.json()
self.assertEqual(data["api_version"], REGISTRY_API_VERSION)
self.assertEqual(data["project"]["id"], "gitea-tools")
self.assertEqual(data["source"]["kind"], "file")
def test_api_v1_project_detail_missing_fails_closed(self):
response = self.client.get("/api/v1/projects/not-registered")
self.assertEqual(response.status_code, 404)
data = response.json()
self.assertEqual(data["error"], "project_not_found")
self.assertEqual(data["project_id"], "not-registered")
self.assertIn("gitea-tools", data["known_project_ids"])
self.assertIn("remediation", data)
def test_api_v1_projects_is_read_only(self):
response = self.client.post("/api/v1/projects", json={})
self.assertEqual(response.status_code, 405)
self.assertEqual(response.json()["error"], "read-only-mvp")
def test_project_to_dict_is_json_safe(self): def test_project_to_dict_is_json_safe(self):
registry = load_registry() registry = load_registry()
encoded = json.dumps(project_to_dict(registry.projects[0])) dto = project_to_dict(registry.projects[0])
encoded = json.dumps(dto)
self.assertIn("gitea-tools", encoded) self.assertIn("gitea-tools", encoded)
# Prose may mention tokens; no serialized *key* may look like a secret.
for key in dto:
with self.subTest(key=key):
self.assertFalse(is_forbidden_key(key))
class TestInvalidRegistryFailsClosedOverHttp(RegistryFileCase):
def setUp(self):
super().setUp()
self.path = self.write_registry({"version": 42, "projects": []})
self.client = TestClient(create_app())
def _with_bad_registry(self, url: str):
import os
from unittest import mock
with mock.patch.dict(
os.environ, {"WEBUI_PROJECT_REGISTRY": str(self.path)}, clear=False
):
return self.client.get(url)
def test_api_v1_reports_actionable_error(self):
response = self._with_bad_registry("/api/v1/projects")
self.assertEqual(response.status_code, 500)
data = response.json()
self.assertEqual(data["error"], "registry_invalid")
self.assertIn("unsupported registry version", data["detail"])
self.assertTrue(data["remediation"])
self.assertEqual(data["field_path"], "version")
def test_unversioned_alias_reports_actionable_error(self):
response = self._with_bad_registry("/api/projects")
self.assertEqual(response.status_code, 500)
self.assertEqual(response.json()["error"], "registry_invalid")
def test_html_page_reports_actionable_error(self):
response = self._with_bad_registry("/projects")
self.assertEqual(response.status_code, 500)
self.assertIn("Project registry unavailable", response.text)
self.assertIn("Remediation", response.text)
class TestProjectRegistryApiDocs(unittest.TestCase):
def test_api_contract_is_documented(self):
self.assertTrue(_API_DOC.is_file(), f"missing {_API_DOC}")
text = _API_DOC.read_text(encoding="utf-8")
for token in (
"/api/v1/projects",
"/api/v1/projects/{project_id}",
"/api/projects",
"onboarding_summary",
"last_seen_health",
"registry_invalid",
"#635",
):
with self.subTest(token=token):
self.assertIn(token, text)
def test_route_table_lists_versioned_routes(self):
local_dev = (_REPO_ROOT / "docs" / "webui-local-dev.md").read_text(
encoding="utf-8"
)
self.assertIn("/api/v1/projects", local_dev)
self.assertIn("webui-project-registry-api.md", local_dev)
if __name__ == "__main__": if __name__ == "__main__":
unittest.main() unittest.main()
+458
View File
@@ -0,0 +1,458 @@
"""Tests for the worker registry and configuration schema (#798, epic #797)."""
import json
import sys
import tempfile
import unittest
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
from webui.worker_registry import (
ALLOWED_ROLES,
SCHEMA_VERSION,
RegistryValidationError,
WorkerRegistry,
default_registry_path,
find_provider,
find_worker,
history_dir,
list_revisions,
load_registry,
registry_to_dict,
registry_to_document,
rollback_to_revision,
save_registry,
validate_payload,
worker_to_dict,
workers_for_provider,
)
_EXPECTED_PROVIDER_IDS = ("claude", "grok", "codex", "agy", "kimi-k")
def _provider(provider_id: str = "claude", **overrides) -> dict:
payload = {
"id": provider_id,
"display_name": "Claude",
"vendor": "Anthropic",
"executable": "claude",
"available": True,
"models": ["claude-opus-4-8"],
"notes": "",
}
payload.update(overrides)
return payload
def _worker(worker_id: str = "claude-author", **overrides) -> dict:
payload = {
"id": worker_id,
"display_name": "Claude author",
"provider": "claude",
"model": "claude-opus-4-8",
"project": "gitea-tools",
"role": "author",
"namespace": "gitea-author",
"profile": "prgs-author",
"workflow": "skills/llm-project-workflow/workflows/work-issue.md",
"schedule": {"kind": "cron", "expression": "0 * * * *"},
"timeout_seconds": 3600,
"enabled": True,
"scheduler": {"kind": "launchd", "label": "cc.prgs.claude.author"},
"notes": "",
}
payload.update(overrides)
return payload
def _document(providers=None, workers=None, **overrides) -> dict:
payload = {
"version": SCHEMA_VERSION,
"revision": 1,
"updated_at": "2026-07-22T00:00:00Z",
"providers": providers if providers is not None else [_provider()],
"workers": workers if workers is not None else [_worker()],
}
payload.update(overrides)
return payload
class _TempRegistryCase(unittest.TestCase):
"""Base case giving each test an isolated registry file."""
def setUp(self):
self._tmp = tempfile.TemporaryDirectory()
self.addCleanup(self._tmp.cleanup)
self.path = Path(self._tmp.name) / "workers.registry.json"
def write(self, document: dict) -> Path:
self.path.write_text(json.dumps(document, indent=2) + "\n", encoding="utf-8")
return self.path
def parse(self, document: dict) -> WorkerRegistry:
return validate_payload(document, source_path=self.path)
class TestPackagedRegistry(unittest.TestCase):
"""AC: the declarative registry is the source of truth and ships with the app."""
def test_default_path_points_at_packaged_data(self):
path = default_registry_path()
self.assertEqual(path.name, "workers.registry.json")
self.assertEqual(path.parent.name, "data")
def test_packaged_registry_loads_and_validates(self):
registry = load_registry()
self.assertEqual(registry.version, SCHEMA_VERSION)
self.assertGreaterEqual(registry.revision, 1)
def test_packaged_registry_declares_all_five_providers(self):
registry = load_registry()
self.assertEqual(
tuple(provider.id for provider in registry.providers),
_EXPECTED_PROVIDER_IDS,
)
def test_packaged_registry_carries_no_credentials(self):
raw = default_registry_path().read_text(encoding="utf-8").lower()
for marker in ("token", "password", "secret", "api_key", "credential"):
self.assertNotIn(marker, raw)
class TestSeparateEntities(_TempRegistryCase):
"""AC: providers and configured workers are separate entities."""
def test_provider_may_exist_with_no_workers(self):
registry = self.parse(
_document(providers=[_provider("grok", display_name="Grok")], workers=[])
)
self.assertEqual(len(registry.providers), 1)
self.assertEqual(registry.workers, ())
self.assertEqual(workers_for_provider(registry, "grok"), ())
def test_many_workers_may_share_one_provider(self):
registry = self.parse(
_document(
workers=[
_worker("claude-author"),
_worker(
"claude-reviewer",
role="reviewer",
namespace="gitea-reviewer",
profile="prgs-reviewer",
scheduler={"kind": "launchd", "label": "cc.prgs.claude.reviewer"},
),
]
)
)
self.assertEqual(len(workers_for_provider(registry, "claude")), 2)
self.assertEqual(len(registry.providers), 1)
def test_worker_referencing_unknown_provider_is_refused(self):
with self.assertRaises(RegistryValidationError) as ctx:
self.parse(_document(workers=[_worker(provider="mystery")]))
self.assertIn("unknown provider", str(ctx.exception))
def test_lookup_helpers(self):
registry = self.parse(_document())
self.assertIsNotNone(find_worker(registry, "claude-author"))
self.assertIsNone(find_worker(registry, "absent"))
self.assertIsNotNone(find_provider(registry, "claude"))
self.assertIsNone(find_provider(registry, "absent"))
class TestRecordedFields(_TempRegistryCase):
"""AC: records provider, model, project, role, namespace/profile, workflow,
schedule, timeout, enabled state, and scheduler metadata."""
def test_every_required_field_is_recorded(self):
registry = self.parse(_document())
worker = registry.workers[0]
self.assertEqual(worker.provider, "claude")
self.assertEqual(worker.model, "claude-opus-4-8")
self.assertEqual(worker.project, "gitea-tools")
self.assertEqual(worker.role, "author")
self.assertEqual(worker.namespace, "gitea-author")
self.assertEqual(worker.profile, "prgs-author")
self.assertEqual(worker.workflow, "skills/llm-project-workflow/workflows/work-issue.md")
self.assertEqual(worker.schedule.kind, "cron")
self.assertEqual(worker.schedule.expression, "0 * * * *")
self.assertEqual(worker.timeout_seconds, 3600)
self.assertTrue(worker.enabled)
self.assertEqual(worker.scheduler.kind, "launchd")
self.assertEqual(worker.scheduler.label, "cc.prgs.claude.author")
def test_each_required_field_is_individually_required(self):
for field in (
"provider", "model", "project", "role", "namespace",
"profile", "workflow", "schedule", "timeout_seconds",
"enabled", "scheduler", "id", "display_name",
):
with self.subTest(field=field):
worker = _worker()
worker.pop(field)
with self.assertRaises(RegistryValidationError):
self.parse(_document(workers=[worker]))
def test_all_sanctioned_roles_are_accepted(self):
for role in ALLOWED_ROLES:
with self.subTest(role=role):
registry = self.parse(_document(workers=[_worker(role=role)]))
self.assertEqual(registry.workers[0].role, role)
def test_unsanctioned_role_is_refused(self):
with self.assertRaises(RegistryValidationError) as ctx:
self.parse(_document(workers=[_worker(role="admin")]))
self.assertIn("role must be one of", str(ctx.exception))
def test_worker_dict_round_trips_every_field(self):
registry = self.parse(_document())
encoded = worker_to_dict(registry.workers[0])
self.assertEqual(encoded, _worker())
json.dumps(encoded) # must stay JSON-safe for the #799 API
class TestSchemaValidation(_TempRegistryCase):
"""AC: supports schema validation — and fails closed."""
def test_unsupported_version_is_refused(self):
with self.assertRaises(RegistryValidationError):
self.parse(_document(version=2))
def test_root_must_be_an_object(self):
with self.assertRaises(RegistryValidationError):
validate_payload([], source_path=self.path)
def test_providers_must_be_non_empty(self):
with self.assertRaises(RegistryValidationError):
self.parse(_document(providers=[]))
def test_unknown_top_level_field_is_refused(self):
with self.assertRaises(RegistryValidationError) as ctx:
self.parse(_document(fleet=[]))
self.assertIn("unknown fields", str(ctx.exception))
def test_unknown_worker_field_is_refused_not_ignored(self):
# A typo'd field must not be silently dropped: "timeout_second" would
# otherwise read as "no timeout declared".
worker = _worker()
worker["timeout_second"] = 30
with self.assertRaises(RegistryValidationError) as ctx:
self.parse(_document(workers=[worker]))
self.assertIn("timeout_second", str(ctx.exception))
def test_credentials_are_refused_anywhere_in_the_document(self):
for label, mutate in (
("provider.api_token", lambda doc: doc["providers"][0].__setitem__("api_token", "x")),
("worker.password", lambda doc: doc["workers"][0].__setitem__("password", "x")),
("root.secret", lambda doc: doc.__setitem__("secret", "x")),
):
with self.subTest(field=label):
document = _document()
mutate(document)
with self.assertRaises(ValueError) as ctx:
self.parse(document)
self.assertIn("credential", str(ctx.exception).lower())
def test_duplicate_worker_id_is_refused(self):
workers = [_worker("dup"), _worker("dup", scheduler={"kind": "manual"})]
with self.assertRaises(RegistryValidationError) as ctx:
self.parse(_document(workers=workers))
self.assertIn("duplicate worker id", str(ctx.exception))
def test_duplicate_provider_id_is_refused(self):
with self.assertRaises(RegistryValidationError) as ctx:
self.parse(_document(providers=[_provider("claude"), _provider("claude")], workers=[]))
self.assertIn("duplicate provider id", str(ctx.exception))
def test_duplicate_launchagent_label_is_refused(self):
# Two workers sharing a label would silently overwrite each other's agent.
workers = [
_worker("a", scheduler={"kind": "launchd", "label": "cc.prgs.same"}),
_worker("b", scheduler={"kind": "launchd", "label": "cc.prgs.same"}),
]
with self.assertRaises(RegistryValidationError) as ctx:
self.parse(_document(workers=workers))
self.assertIn("duplicate scheduler label", str(ctx.exception))
def test_manual_scheduler_needs_no_label_and_many_may_coexist(self):
workers = [
_worker("a", scheduler={"kind": "manual"}),
_worker("b", scheduler={"kind": "manual"}),
]
registry = self.parse(_document(workers=workers))
self.assertEqual([w.scheduler.label for w in registry.workers], [None, None])
def test_launchd_scheduler_requires_a_label(self):
with self.assertRaises(RegistryValidationError) as ctx:
self.parse(_document(workers=[_worker(scheduler={"kind": "launchd"})]))
self.assertIn("label is required", str(ctx.exception))
def test_unknown_scheduler_kind_is_refused(self):
with self.assertRaises(RegistryValidationError):
self.parse(_document(workers=[_worker(scheduler={"kind": "systemd", "label": "x"})]))
def test_timeout_must_be_a_positive_bounded_integer(self):
for bad in (0, -1, "3600", 1.5, True, 86_401):
with self.subTest(timeout=bad):
with self.assertRaises(RegistryValidationError):
self.parse(_document(workers=[_worker(timeout_seconds=bad)]))
def test_enabled_must_be_a_real_boolean(self):
for bad in ("true", 1, None):
with self.subTest(enabled=bad):
with self.assertRaises(RegistryValidationError):
self.parse(_document(workers=[_worker(enabled=bad)]))
def test_identifier_shape_is_enforced(self):
for bad in ("Claude Author", "-leading", "UPPER", ""):
with self.subTest(worker_id=bad):
with self.assertRaises(RegistryValidationError):
self.parse(_document(workers=[_worker(bad)]))
class TestScheduleValidation(_TempRegistryCase):
"""Schedules are declarations; next-run computation belongs to #803."""
def test_interval_schedule_requires_positive_seconds(self):
registry = self.parse(
_document(workers=[_worker(schedule={"kind": "interval", "seconds": 900})])
)
self.assertEqual(registry.workers[0].schedule.seconds, 900)
with self.assertRaises(RegistryValidationError):
self.parse(_document(workers=[_worker(schedule={"kind": "interval"})]))
with self.assertRaises(RegistryValidationError):
self.parse(_document(workers=[_worker(schedule={"kind": "interval", "seconds": 0})]))
def test_cron_schedule_requires_five_fields(self):
with self.assertRaises(RegistryValidationError) as ctx:
self.parse(_document(workers=[_worker(schedule={"kind": "cron", "expression": "0 *"})]))
self.assertIn("five crontab fields", str(ctx.exception))
def test_manual_schedule_needs_no_timing(self):
registry = self.parse(_document(workers=[_worker(schedule={"kind": "manual"})]))
schedule = registry.workers[0].schedule
self.assertEqual(schedule.kind, "manual")
self.assertIsNone(schedule.seconds)
self.assertIsNone(schedule.expression)
def test_fields_from_the_wrong_kind_are_refused(self):
with self.assertRaises(RegistryValidationError) as ctx:
self.parse(_document(workers=[_worker(schedule={"kind": "manual", "seconds": 60})]))
self.assertIn("not valid for kind", str(ctx.exception))
def test_unknown_schedule_kind_is_refused(self):
with self.assertRaises(RegistryValidationError):
self.parse(_document(workers=[_worker(schedule={"kind": "hourly"})]))
class TestAtomicPersistence(_TempRegistryCase):
"""AC: atomic persistence."""
def test_save_then_load_round_trips(self):
registry = self.parse(_document())
save_registry(registry, self.path)
reloaded = load_registry(self.path)
self.assertEqual(
[worker_to_dict(w) for w in reloaded.workers],
[worker_to_dict(w) for w in registry.workers],
)
def test_save_leaves_no_temp_files_behind(self):
registry = self.parse(_document())
save_registry(registry, self.path)
save_registry(registry, self.path)
leftovers = [p.name for p in self.path.parent.iterdir() if p.name.startswith(".")]
self.assertEqual(leftovers, [])
def test_save_refuses_to_persist_an_invalid_document(self):
registry = self.parse(_document())
broken = WorkerRegistry(
version=registry.version,
revision=registry.revision,
updated_at=registry.updated_at,
providers=registry.providers,
# A worker whose provider is not declared in the registry.
workers=tuple(
type(worker)(**{**worker.__dict__, "provider": "vanished"})
for worker in registry.workers
),
source_path=self.path,
)
with self.assertRaises(RegistryValidationError):
save_registry(broken, self.path)
self.assertFalse(self.path.exists(), "invalid save must not create the file")
def test_document_shape_excludes_local_paths_but_api_shape_includes_it(self):
registry = self.parse(_document())
self.assertNotIn("source_path", registry_to_document(registry))
self.assertEqual(registry_to_dict(registry)["source_path"], str(self.path))
class TestVersioningAndRollback(_TempRegistryCase):
"""AC: versioning and rollback."""
def _seed(self) -> WorkerRegistry:
self.write(_document())
return load_registry(self.path)
def test_revision_increments_on_each_save(self):
registry = self._seed()
self.assertEqual(registry.revision, 1)
second = save_registry(registry, self.path)
self.assertEqual(second.revision, 2)
third = save_registry(second, self.path)
self.assertEqual(third.revision, 3)
def test_updated_at_is_refreshed_and_utc(self):
registry = self._seed()
saved = save_registry(registry, self.path)
self.assertRegex(saved.updated_at, r"^\d{4}-\d{2}-\d{2}T\d{2}:\d{2}:\d{2}Z$")
def test_superseded_revisions_are_retained(self):
registry = self._seed()
second = save_registry(registry, self.path)
save_registry(second, self.path)
self.assertEqual(list_revisions(self.path), (1, 2))
self.assertTrue(history_dir(self.path).is_dir())
def test_rollback_restores_prior_content_as_a_new_revision(self):
self.write(_document(workers=[_worker("original")]))
registry = load_registry(self.path)
changed = WorkerRegistry(
version=registry.version,
revision=registry.revision,
updated_at=registry.updated_at,
providers=registry.providers,
workers=(), # operator deletes every worker
source_path=self.path,
)
save_registry(changed, self.path)
self.assertEqual(load_registry(self.path).workers, ())
restored = rollback_to_revision(1, self.path)
self.assertEqual([w.id for w in restored.workers], ["original"])
# Append-only: the rollback publishes a new head rather than rewinding.
self.assertGreater(restored.revision, 2)
self.assertEqual([w.id for w in load_registry(self.path).workers], ["original"])
def test_rollback_to_unknown_revision_fails_closed(self):
self._seed()
with self.assertRaises(RegistryValidationError) as ctx:
rollback_to_revision(99, self.path)
self.assertIn("not retained", str(ctx.exception))
def test_revision_must_be_a_positive_integer(self):
for bad in (0, -1, "1", None):
with self.subTest(revision=bad):
with self.assertRaises(RegistryValidationError):
self.parse(_document(revision=bad))
def test_history_is_empty_before_any_save(self):
self.write(_document())
self.assertEqual(list_revisions(self.path), ())
if __name__ == "__main__":
unittest.main()
+72 -5
View File
@@ -11,8 +11,20 @@ from starlette.routing import Route
from webui.deployment_boundary import deployment_snapshot from webui.deployment_boundary import deployment_snapshot
from webui.layout import render_page from webui.layout import render_page
from webui.project_registry import find_project, load_registry, registry_to_dict from webui.project_registry import (
from webui.project_views import render_project_detail, render_projects_list ProjectRegistry,
RegistryError,
find_project,
known_project_ids,
load_registry,
project_detail_to_dict,
registry_to_dict,
)
from webui.project_views import (
render_project_detail,
render_projects_list,
render_registry_error,
)
from webui.prompt_library import find_prompt, library_to_dict from webui.prompt_library import find_prompt, library_to_dict
from webui.prompt_views import render_prompt_detail, render_prompts_page from webui.prompt_views import render_prompt_detail, render_prompts_page
from final_report_validator import FINAL_REPORT_TASK_KINDS from final_report_validator import FINAL_REPORT_TASK_KINDS
@@ -81,14 +93,26 @@ async def api_queue(_request: Request) -> JSONResponse:
return JSONResponse(queue_snapshot_to_dict(load_queue_snapshot())) return JSONResponse(queue_snapshot_to_dict(load_queue_snapshot()))
def _load_project_registry() -> tuple[ProjectRegistry | None, RegistryError | None]:
"""Load the registry, converting validation failure into a fail-closed pair."""
try:
return load_registry(), None
except RegistryError as exc:
return None, exc
async def projects(_request: Request) -> HTMLResponse: async def projects(_request: Request) -> HTMLResponse:
registry = load_registry() registry, error = _load_project_registry()
if error is not None:
return HTMLResponse(render_registry_error(error), status_code=500)
return HTMLResponse(render_projects_list(registry)) return HTMLResponse(render_projects_list(registry))
async def project_detail(request: Request) -> HTMLResponse: async def project_detail(request: Request) -> HTMLResponse:
project_id = request.path_params["project_id"] project_id = request.path_params["project_id"]
registry = load_registry() registry, error = _load_project_registry()
if error is not None:
return HTMLResponse(render_registry_error(error), status_code=500)
project = find_project(registry, project_id) project = find_project(registry, project_id)
if project is None: if project is None:
return HTMLResponse( return HTMLResponse(
@@ -106,10 +130,47 @@ async def project_detail(request: Request) -> HTMLResponse:
async def api_projects(_request: Request) -> JSONResponse: async def api_projects(_request: Request) -> JSONResponse:
registry = load_registry() """Unversioned MVP alias, retained through Phase 1 (#632 section 6)."""
registry, error = _load_project_registry()
if error is not None:
return JSONResponse(error.to_dict(), status_code=500)
return JSONResponse(registry_to_dict(registry)) return JSONResponse(registry_to_dict(registry))
async def api_v1_projects(_request: Request) -> JSONResponse:
registry, error = _load_project_registry()
if error is not None:
return JSONResponse(error.to_dict(), status_code=500)
return JSONResponse(registry_to_dict(registry))
async def api_v1_project_detail(request: Request) -> JSONResponse:
project_id = request.path_params["project_id"]
registry, error = _load_project_registry()
if error is not None:
return JSONResponse(error.to_dict(), status_code=500)
project = find_project(registry, project_id)
if project is None:
return JSONResponse(
{
"error": "project_not_found",
"project_id": project_id,
"known_project_ids": known_project_ids(registry),
"remediation": (
"Request one of the known project ids, or add the project to the "
"registry file named in 'source'."
),
"source": {
"kind": "file",
"path": str(registry.source_path),
"inventory_complete": True,
},
},
status_code=404,
)
return JSONResponse(project_detail_to_dict(registry, project))
async def prompts(_request: Request) -> HTMLResponse: async def prompts(_request: Request) -> HTMLResponse:
return HTMLResponse(render_prompts_page()) return HTMLResponse(render_prompts_page())
@@ -268,6 +329,12 @@ def create_app(*, bind_host: str | None = None) -> Starlette:
Route("/projects", projects, methods=["GET"]), Route("/projects", projects, methods=["GET"]),
Route("/projects/{project_id}", project_detail, methods=["GET"]), Route("/projects/{project_id}", project_detail, methods=["GET"]),
Route("/api/projects", api_projects, methods=["GET"]), Route("/api/projects", api_projects, methods=["GET"]),
Route("/api/v1/projects", api_v1_projects, methods=["GET"]),
Route(
"/api/v1/projects/{project_id}",
api_v1_project_detail,
methods=["GET"],
),
Route("/prompts", prompts, methods=["GET"]), Route("/prompts", prompts, methods=["GET"]),
Route("/prompts/{prompt_id}", prompt_detail, methods=["GET"]), Route("/prompts/{prompt_id}", prompt_detail, methods=["GET"]),
Route("/api/prompts", api_prompts, methods=["GET"]), Route("/api/prompts", api_prompts, methods=["GET"]),
+16 -6
View File
@@ -1,13 +1,15 @@
{ {
"version": 1, "version": 2,
"projects": [ "projects": [
{ {
"id": "gitea-tools", "id": "gitea-tools",
"repo_name": "Gitea-Tools", "repo_name": "Gitea-Tools",
"gitea_owner": "Scaled-Tech-Consulting", "gitea_owner": "Scaled-Tech-Consulting",
"remote_name": "prgs",
"remote_host": "https://gitea.prgs.cc", "remote_host": "https://gitea.prgs.cc",
"default_branch": "master", "default_branch": "master",
"local_checkout_path": ".", "local_checkout_path": ".",
"status": "active",
"profiles": { "profiles": {
"author": "prgs-author", "author": "prgs-author",
"reviewer": "prgs-reviewer", "reviewer": "prgs-reviewer",
@@ -26,24 +28,32 @@
{ {
"id": "profiles", "id": "profiles",
"title": "Configure execution profiles", "title": "Configure execution profiles",
"description": "Install author, reviewer, and reconciler MCP profiles (prgs-author, prgs-reviewer, prgs-reconciler) in separate namespaces. Tokens stay in keychain — never in this registry." "description": "Install author, reviewer, and reconciler MCP profiles (prgs-author, prgs-reviewer, prgs-reconciler) in separate namespaces. Tokens stay in keychain — never in this registry.",
"state": "complete",
"required": true
}, },
{ {
"id": "mcp_config", "id": "mcp_config",
"title": "Wire MCP v2 contexts", "title": "Wire MCP v2 contexts",
"description": "Copy and customize gitea-mcp.v2-contexts.example.json for your machine. Map this repo path under projects with default_owner Scaled-Tech-Consulting and default_repo Gitea-Tools." "description": "Copy and customize gitea-mcp.v2-contexts.example.json for your machine. Map this repo path under projects with default_owner Scaled-Tech-Consulting and default_repo Gitea-Tools.",
"state": "complete",
"required": true
}, },
{ {
"id": "wiki_gate", "id": "wiki_gate",
"title": "Wiki publication readiness", "title": "Wiki publication readiness",
"description": "For wiki-tracked work, satisfy the live Gitea Wiki proof gate (#224) before closing issues. See docs/wiki/Safety-and-Gates.md." "description": "For wiki-tracked work, satisfy the live Gitea Wiki proof gate (#224) before closing issues. See docs/wiki/Safety-and-Gates.md.",
"state": "complete",
"required": true
}, },
{ {
"id": "branches_layout", "id": "branches_layout",
"title": "Isolate work under branches/", "title": "Isolate work under branches/",
"description": "All LLM task edits happen in worktrees under branches/. Main checkout stays clean; use skills/llm-project-workflow templates for start-issue and review flows." "description": "All LLM task edits happen in worktrees under branches/. Main checkout stays clean; use skills/llm-project-workflow templates for start-issue and review flows.",
"state": "complete",
"required": true
} }
] ]
} }
] ]
} }
+57
View File
@@ -0,0 +1,57 @@
{
"version": 1,
"revision": 1,
"updated_at": "2026-07-22T00:00:00Z",
"providers": [
{
"id": "claude",
"display_name": "Claude",
"vendor": "Anthropic",
"executable": "claude",
"available": true,
"models": [
"claude-opus-4-8",
"claude-sonnet-5",
"claude-haiku-4-5-20251001"
],
"notes": "Model list is a declaration. Live enumeration and version inspection belong to the provider adapter framework (#800)."
},
{
"id": "grok",
"display_name": "Grok",
"vendor": "xAI",
"executable": "grok",
"available": true,
"models": [],
"notes": "Models enumerated by the provider adapter (#800); not declared here."
},
{
"id": "codex",
"display_name": "Codex",
"vendor": "OpenAI",
"executable": "codex",
"available": true,
"models": [],
"notes": "Models enumerated by the provider adapter (#800); not declared here."
},
{
"id": "agy",
"display_name": "AGY",
"vendor": "Antigravity",
"executable": "agy",
"available": true,
"models": [],
"notes": "MCP allowlist gating applies to this provider; confirm server-side allowlist before configuring a worker."
},
{
"id": "kimi-k",
"display_name": "Kimi K",
"vendor": "Moonshot AI",
"executable": "kimi",
"available": true,
"models": [],
"notes": "Provider id is kimi-k; the executable on PATH is kimi. Models enumerated by the provider adapter (#800)."
}
],
"workers": []
}
+451 -48
View File
@@ -1,4 +1,16 @@
"""Load and validate the web UI project registry (#427).""" """Load and validate the web UI project registry (#427, evolved for #635).
Phase 1 of the console architecture ADR keeps this loader read-only. It owns
the versioned project registry contract served at ``/api/v1/projects``:
* the on-disk file carries a ``version`` (schema version 1 or 2);
* version 1 files stay loadable and are normalized with explicit defaults, so
an operator registry written for #427 keeps working;
* every validation failure raises :class:`RegistryError`, which carries an
actionable ``remediation`` string instead of leaking a traceback;
* serialization never emits credentials — credential-shaped keys are rejected
at load time, before any DTO is built.
"""
from __future__ import annotations from __future__ import annotations
@@ -8,27 +20,59 @@ from dataclasses import dataclass
from pathlib import Path from pathlib import Path
from typing import Any from typing import Any
_FORBIDDEN_EXACT_KEYS = frozenset({ from webui.registry_safety import is_forbidden_key
"token",
"password",
"secret",
"credential",
"auth",
"api_key",
"api-key",
})
_FORBIDDEN_KEY_PREFIXES = ("auth_", "api_key_", "api-key_")
_FORBIDDEN_KEY_SUFFIXES = ("_token", "_secret", "_password", "_credential", "_auth")
#: Version of the JSON contract served under ``/api/v1/...``.
REGISTRY_API_VERSION = "v1"
#: Schema version written by this repository's packaged registry.
CURRENT_SCHEMA_VERSION = 2
#: Schema versions this loader accepts. Version 1 is normalized on load.
SUPPORTED_SCHEMA_VERSIONS = (1, 2)
#: Lifecycle state of a registered project.
PROJECT_STATUSES = ("active", "onboarding", "paused", "archived")
_DEFAULT_PROJECT_STATUS = "active"
#: Completion state of a single onboarding step.
ONBOARDING_STATES = ("complete", "pending", "blocked", "not_applicable")
_DEFAULT_ONBOARDING_STATE = "pending"
#: Redacted, last-seen health of a project's control plane.
HEALTH_STATUSES = ("healthy", "degraded", "unreachable", "unknown")
class RegistryError(ValueError):
"""A registry file could not be loaded or failed validation.
Carries an operator-facing ``remediation`` so routes can fail closed with
an actionable message rather than a stack trace.
"""
def __init__(
self,
message: str,
*,
remediation: str,
source_path: Path | None = None,
field_path: str | None = None,
) -> None:
super().__init__(message)
self.message = message
self.remediation = remediation
self.source_path = source_path
self.field_path = field_path
def to_dict(self) -> dict[str, Any]:
"""Serialize for a fail-closed JSON error response."""
return {
"error": "registry_invalid",
"detail": self.message,
"remediation": self.remediation,
"field_path": self.field_path,
"source_path": str(self.source_path) if self.source_path else None,
}
def _is_forbidden_key(key: str) -> bool:
lowered = key.lower()
if lowered in _FORBIDDEN_EXACT_KEYS:
return True
return (
lowered.startswith(_FORBIDDEN_KEY_PREFIXES)
or lowered.endswith(_FORBIDDEN_KEY_SUFFIXES)
)
_REQUIRED_PROJECT_FIELDS = ( _REQUIRED_PROJECT_FIELDS = (
"id", "id",
@@ -49,6 +93,30 @@ class OnboardingStep:
id: str id: str
title: str title: str
description: str description: str
state: str = _DEFAULT_ONBOARDING_STATE
required: bool = True
@dataclass(frozen=True)
class OnboardingSummary:
"""Aggregate onboarding progress for a single project."""
total: int
complete: int
pending: int
blocked: int
not_applicable: int
required_outstanding: int
onboarding_complete: bool
@dataclass(frozen=True)
class ProjectHealth:
"""Redacted last-seen health. Never carries endpoints or credentials."""
status: str
checked_at: str | None
detail: str | None
@dataclass(frozen=True) @dataclass(frozen=True)
@@ -63,6 +131,13 @@ class ProjectRecord:
workflow_paths: dict[str, str] workflow_paths: dict[str, str]
schema_paths: dict[str, str] schema_paths: dict[str, str]
onboarding_checklist: tuple[OnboardingStep, ...] onboarding_checklist: tuple[OnboardingStep, ...]
status: str = _DEFAULT_PROJECT_STATUS
remote_name: str | None = None
last_seen_health: ProjectHealth | None = None
@property
def repo_full_name(self) -> str:
return f"{self.gitea_owner}/{self.repo_name}"
@dataclass(frozen=True) @dataclass(frozen=True)
@@ -71,6 +146,15 @@ class ProjectRegistry:
projects: tuple[ProjectRecord, ...] projects: tuple[ProjectRecord, ...]
source_path: Path source_path: Path
@property
def schema_version(self) -> int:
"""Alias of :attr:`version` — the schema version read from disk."""
return self.version
@property
def api_version(self) -> str:
return REGISTRY_API_VERSION
def default_registry_path() -> Path: def default_registry_path() -> Path:
override = os.environ.get("WEBUI_PROJECT_REGISTRY", "").strip() override = os.environ.get("WEBUI_PROJECT_REGISTRY", "").strip()
@@ -79,52 +163,221 @@ def default_registry_path() -> Path:
return (Path(__file__).resolve().parent / "data" / "projects.registry.json").resolve() return (Path(__file__).resolve().parent / "data" / "projects.registry.json").resolve()
def _reject_credential_keys(obj: Any, *, path: str = "") -> None: def _reject_credential_keys(obj: Any, *, path: str = "", source: Path | None = None) -> None:
"""Recursive credential-key guard that reports an actionable ``field_path``.
Key *shape* is decided by :func:`webui.registry_safety.is_forbidden_key`, the
single source of truth shared with the worker registry (#798).
"""
if isinstance(obj, dict): if isinstance(obj, dict):
for key, value in obj.items(): for key, value in obj.items():
key_path = f"{path}.{key}" if path else key key_path = f"{path}.{key}" if path else key
if _is_forbidden_key(key): if is_forbidden_key(key):
raise ValueError(f"registry must not store credentials ({key_path})") raise RegistryError(
_reject_credential_keys(value, path=key_path) f"registry must not store credentials ({key_path})",
remediation=(
f"Remove the credential-shaped key '{key_path}' from the registry. "
"Tokens live in the keychain and are resolved server-side by "
"gitea_auth; the registry is redacted metadata only."
),
source_path=source,
field_path=key_path,
)
_reject_credential_keys(value, path=key_path, source=source)
elif isinstance(obj, list): elif isinstance(obj, list):
for index, item in enumerate(obj): for index, item in enumerate(obj):
_reject_credential_keys(item, path=f"{path}[{index}]") _reject_credential_keys(item, path=f"{path}[{index}]", source=source)
def _parse_onboarding(raw: list[dict[str, Any]] | None) -> tuple[OnboardingStep, ...]: def _require_enum(
if not raw: value: Any,
*,
allowed: tuple[str, ...],
field_path: str,
source: Path | None,
) -> str:
text = str(value)
if text not in allowed:
raise RegistryError(
f"{field_path} must be one of {', '.join(allowed)} (got {text!r})",
remediation=(
f"Set {field_path} to one of: {', '.join(allowed)}. "
"Unknown values fail closed so the console never renders an "
"unverified state."
),
source_path=source,
field_path=field_path,
)
return text
def _parse_onboarding(
raw: Any,
*,
project_path: str,
source: Path | None,
) -> tuple[OnboardingStep, ...]:
if raw is None:
return () return ()
if not isinstance(raw, list):
raise RegistryError(
f"{project_path}.onboarding_checklist must be an array",
remediation=(
f"Rewrite {project_path}.onboarding_checklist as a JSON array of "
"steps with id, title, description, and optional state."
),
source_path=source,
field_path=f"{project_path}.onboarding_checklist",
)
steps: list[OnboardingStep] = [] steps: list[OnboardingStep] = []
for item in raw: for index, item in enumerate(raw):
step_path = f"{project_path}.onboarding_checklist[{index}]"
if not isinstance(item, dict):
raise RegistryError(
f"{step_path} must be an object",
remediation=f"Rewrite {step_path} as an object with id, title, description.",
source_path=source,
field_path=step_path,
)
missing = [field for field in ("id", "title", "description") if field not in item]
if missing:
raise RegistryError(
f"{step_path} missing required fields: {', '.join(missing)}",
remediation=(
f"Add {', '.join(missing)} to {step_path}. Every onboarding step "
"must be self-describing for an operator who has no chat history."
),
source_path=source,
field_path=step_path,
)
state = _require_enum(
item.get("state", _DEFAULT_ONBOARDING_STATE),
allowed=ONBOARDING_STATES,
field_path=f"{step_path}.state",
source=source,
)
steps.append( steps.append(
OnboardingStep( OnboardingStep(
id=str(item["id"]), id=str(item["id"]),
title=str(item["title"]), title=str(item["title"]),
description=str(item["description"]), description=str(item["description"]),
state=state,
required=bool(item.get("required", True)),
) )
) )
return tuple(steps) return tuple(steps)
def _parse_project(raw: dict[str, Any]) -> ProjectRecord: def _parse_health(
raw: Any,
*,
project_path: str,
source: Path | None,
) -> ProjectHealth | None:
if raw is None:
return None
if not isinstance(raw, dict):
raise RegistryError(
f"{project_path}.last_seen_health must be an object when present",
remediation=(
f"Rewrite {project_path}.last_seen_health as an object with status "
f"(one of {', '.join(HEALTH_STATUSES)}), optional checked_at and detail, "
"or remove it. Never store endpoints or credentials here."
),
source_path=source,
field_path=f"{project_path}.last_seen_health",
)
status = _require_enum(
raw.get("status", "unknown"),
allowed=HEALTH_STATUSES,
field_path=f"{project_path}.last_seen_health.status",
source=source,
)
checked_at = raw.get("checked_at")
detail = raw.get("detail")
return ProjectHealth(
status=status,
checked_at=str(checked_at) if checked_at is not None else None,
detail=str(detail) if detail is not None else None,
)
def _parse_project(raw: Any, *, index: int, source: Path | None) -> ProjectRecord:
project_path = f"projects[{index}]"
if not isinstance(raw, dict):
raise RegistryError(
f"{project_path} must be an object",
remediation=f"Rewrite {project_path} as a JSON object describing one project.",
source_path=source,
field_path=project_path,
)
missing = [field for field in _REQUIRED_PROJECT_FIELDS if field not in raw] missing = [field for field in _REQUIRED_PROJECT_FIELDS if field not in raw]
if missing: if missing:
raise ValueError(f"project missing required fields: {', '.join(missing)}") raise RegistryError(
f"{project_path} missing required fields: {', '.join(missing)}",
remediation=(
f"Add {', '.join(missing)} to {project_path}. See "
"docs/webui-project-registry-api.md for the field-by-field contract."
),
source_path=source,
field_path=project_path,
)
profiles = raw["profiles"] profiles = raw["profiles"]
if not isinstance(profiles, dict): if not isinstance(profiles, dict):
raise ValueError("profiles must be an object") raise RegistryError(
f"{project_path}.profiles must be an object",
remediation=(
f"Rewrite {project_path}.profiles as an object mapping "
f"{', '.join(_REQUIRED_PROFILE_ROLES)} to MCP profile names."
),
source_path=source,
field_path=f"{project_path}.profiles",
)
for role in _REQUIRED_PROFILE_ROLES: for role in _REQUIRED_PROFILE_ROLES:
if role not in profiles or not profiles[role]: if role not in profiles or not profiles[role]:
raise ValueError(f"profiles.{role} is required") raise RegistryError(
f"{project_path}.profiles.{role} is required",
remediation=(
f"Set {project_path}.profiles.{role} to the configured MCP profile "
"name for that role. Role separation is a workflow-safety invariant."
),
source_path=source,
field_path=f"{project_path}.profiles.{role}",
)
workflow_paths = raw["workflow_paths"] workflow_paths = raw["workflow_paths"]
if not isinstance(workflow_paths, dict) or not workflow_paths: if not isinstance(workflow_paths, dict) or not workflow_paths:
raise ValueError("workflow_paths must be a non-empty object") raise RegistryError(
f"{project_path}.workflow_paths must be a non-empty object",
remediation=(
f"Add at least a 'skill' entry to {project_path}.workflow_paths pointing "
"at the project's canonical workflow skill."
),
source_path=source,
field_path=f"{project_path}.workflow_paths",
)
schema_paths = raw.get("schema_paths") or {} schema_paths = raw.get("schema_paths") or {}
if not isinstance(schema_paths, dict): if not isinstance(schema_paths, dict):
raise ValueError("schema_paths must be an object when present") raise RegistryError(
f"{project_path}.schema_paths must be an object when present",
remediation=(
f"Rewrite {project_path}.schema_paths as an object of label to repo path, "
"or remove it."
),
source_path=source,
field_path=f"{project_path}.schema_paths",
)
status = _require_enum(
raw.get("status", _DEFAULT_PROJECT_STATUS),
allowed=PROJECT_STATUSES,
field_path=f"{project_path}.status",
source=source,
)
remote_name = raw.get("remote_name")
return ProjectRecord( return ProjectRecord(
id=str(raw["id"]), id=str(raw["id"]),
@@ -136,61 +389,211 @@ def _parse_project(raw: dict[str, Any]) -> ProjectRecord:
profiles={role: str(profiles[role]) for role in _REQUIRED_PROFILE_ROLES}, profiles={role: str(profiles[role]) for role in _REQUIRED_PROFILE_ROLES},
workflow_paths={key: str(value) for key, value in workflow_paths.items()}, workflow_paths={key: str(value) for key, value in workflow_paths.items()},
schema_paths={key: str(value) for key, value in schema_paths.items()}, schema_paths={key: str(value) for key, value in schema_paths.items()},
onboarding_checklist=_parse_onboarding(raw.get("onboarding_checklist")), onboarding_checklist=_parse_onboarding(
raw.get("onboarding_checklist"),
project_path=project_path,
source=source,
),
status=status,
remote_name=str(remote_name) if remote_name else None,
last_seen_health=_parse_health(
raw.get("last_seen_health"),
project_path=project_path,
source=source,
),
) )
def load_registry(path: Path | None = None) -> ProjectRegistry: def load_registry(path: Path | None = None) -> ProjectRegistry:
"""Load the versioned project registry from disk.""" """Load the versioned project registry from disk.
Raises:
RegistryError: whenever the file is unreadable, is not valid JSON, or
fails schema validation. The error carries an operator remediation.
"""
source = (path or default_registry_path()).resolve() source = (path or default_registry_path()).resolve()
raw_text = source.read_text(encoding="utf-8") try:
payload = json.loads(raw_text) raw_text = source.read_text(encoding="utf-8")
except OSError as exc:
raise RegistryError(
f"registry file could not be read: {exc.strerror or exc}",
remediation=(
f"Create a readable registry at {source}, or point "
"WEBUI_PROJECT_REGISTRY at an existing file."
),
source_path=source,
) from exc
try:
payload = json.loads(raw_text)
except json.JSONDecodeError as exc:
raise RegistryError(
f"registry is not valid JSON: {exc.msg} (line {exc.lineno}, column {exc.colno})",
remediation=(
f"Fix the JSON syntax in {source} at line {exc.lineno}, column {exc.colno}."
),
source_path=source,
) from exc
if not isinstance(payload, dict): if not isinstance(payload, dict):
raise ValueError("registry root must be an object") raise RegistryError(
"registry root must be an object",
remediation=(
"Wrap the registry in a JSON object with 'version' and 'projects' keys."
),
source_path=source,
)
version = payload.get("version") version = payload.get("version")
if version != 1: if version not in SUPPORTED_SCHEMA_VERSIONS:
raise ValueError(f"unsupported registry version: {version!r}") supported = ", ".join(str(item) for item in SUPPORTED_SCHEMA_VERSIONS)
raise RegistryError(
f"unsupported registry version: {version!r}",
remediation=(
f"Set 'version' to one of {supported} (current schema is "
f"{CURRENT_SCHEMA_VERSION}). Migration notes live in "
"docs/webui-project-registry-api.md."
),
source_path=source,
field_path="version",
)
_reject_credential_keys(payload) _reject_credential_keys(payload, source=source)
projects_raw = payload.get("projects") projects_raw = payload.get("projects")
if not isinstance(projects_raw, list) or not projects_raw: if not isinstance(projects_raw, list) or not projects_raw:
raise ValueError("projects must be a non-empty array") raise RegistryError(
"projects must be a non-empty array",
remediation=(
"Add at least one project object to 'projects'. An empty console "
"registry fails closed rather than rendering a blank inventory."
),
source_path=source,
field_path="projects",
)
projects = tuple(_parse_project(item) for item in projects_raw) projects = tuple(
return ProjectRegistry(version=version, projects=projects, source_path=source) _parse_project(item, index=index, source=source)
for index, item in enumerate(projects_raw)
)
return ProjectRegistry(version=int(version), projects=projects, source_path=source)
def onboarding_summary(project: ProjectRecord) -> OnboardingSummary:
"""Aggregate a project's onboarding checklist state."""
steps = project.onboarding_checklist
counts = {state: 0 for state in ONBOARDING_STATES}
for step in steps:
counts[step.state] += 1
required_outstanding = sum(
1
for step in steps
if step.required and step.state in ("pending", "blocked")
)
return OnboardingSummary(
total=len(steps),
complete=counts["complete"],
pending=counts["pending"],
blocked=counts["blocked"],
not_applicable=counts["not_applicable"],
required_outstanding=required_outstanding,
onboarding_complete=required_outstanding == 0,
)
def project_to_dict(project: ProjectRecord) -> dict[str, Any]: def project_to_dict(project: ProjectRecord) -> dict[str, Any]:
"""Serialize a project for JSON API responses.""" """Serialize a project for JSON API responses and HTML views.
The HTML views render from this same DTO, so the console and the API can
never disagree about a project's status or onboarding progress.
"""
summary = onboarding_summary(project)
health = project.last_seen_health
return { return {
"id": project.id, "id": project.id,
"repo_name": project.repo_name, "repo_name": project.repo_name,
"gitea_owner": project.gitea_owner, "gitea_owner": project.gitea_owner,
"repo_full_name": project.repo_full_name,
"remote_host": project.remote_host, "remote_host": project.remote_host,
"remote_name": project.remote_name,
"default_branch": project.default_branch, "default_branch": project.default_branch,
"local_checkout_path": project.local_checkout_path, "local_checkout_path": project.local_checkout_path,
"status": project.status,
"profiles": dict(project.profiles), "profiles": dict(project.profiles),
"workflow_paths": dict(project.workflow_paths), "workflow_paths": dict(project.workflow_paths),
"schema_paths": dict(project.schema_paths), "schema_paths": dict(project.schema_paths),
"onboarding_checklist": [ "onboarding_checklist": [
{"id": step.id, "title": step.title, "description": step.description} {
"id": step.id,
"title": step.title,
"description": step.description,
"state": step.state,
"required": step.required,
}
for step in project.onboarding_checklist for step in project.onboarding_checklist
], ],
"onboarding_summary": {
"total": summary.total,
"complete": summary.complete,
"pending": summary.pending,
"blocked": summary.blocked,
"not_applicable": summary.not_applicable,
"required_outstanding": summary.required_outstanding,
"onboarding_complete": summary.onboarding_complete,
},
"last_seen_health": (
None
if health is None
else {
"status": health.status,
"checked_at": health.checked_at,
"detail": health.detail,
}
),
} }
def registry_to_dict(registry: ProjectRegistry) -> dict[str, Any]: def registry_to_dict(registry: ProjectRegistry) -> dict[str, Any]:
"""Serialize the whole registry, including API provenance (#632 section 6)."""
return { return {
"api_version": registry.api_version,
"schema_version": registry.schema_version,
# Retained for the unversioned MVP alias consumers (#427).
"version": registry.version, "version": registry.version,
"source_path": str(registry.source_path), "source_path": str(registry.source_path),
"source": {
"kind": "file",
"path": str(registry.source_path),
"inventory_complete": True,
},
"project_count": len(registry.projects),
"projects": [project_to_dict(project) for project in registry.projects], "projects": [project_to_dict(project) for project in registry.projects],
} }
def project_detail_to_dict(
registry: ProjectRegistry,
project: ProjectRecord,
) -> dict[str, Any]:
"""Serialize a single project for ``/api/v1/projects/{project_id}``."""
return {
"api_version": registry.api_version,
"schema_version": registry.schema_version,
"source": {
"kind": "file",
"path": str(registry.source_path),
"inventory_complete": True,
},
"project": project_to_dict(project),
}
def find_project(registry: ProjectRegistry, project_id: str) -> ProjectRecord | None: def find_project(registry: ProjectRegistry, project_id: str) -> ProjectRecord | None:
for project in registry.projects: for project in registry.projects:
if project.id == project_id: if project.id == project_id:
return project return project
return None return None
def known_project_ids(registry: ProjectRegistry) -> list[str]:
return [project.id for project in registry.projects]
+107 -25
View File
@@ -1,34 +1,65 @@
"""HTML views for project registry pages (#427).""" """HTML views for project registry pages (#427, evolved for #635).
Every view renders from :func:`webui.project_registry.project_to_dict`, the
same DTO the ``/api/v1/projects`` JSON responses use, so the HTML console and
the API can never disagree about status or onboarding progress.
"""
from __future__ import annotations from __future__ import annotations
import html import html
from typing import Any
from webui.layout import render_page from webui.layout import render_page
from webui.project_registry import ProjectRecord, ProjectRegistry from webui.project_registry import (
ProjectRecord,
ProjectRegistry,
RegistryError,
project_to_dict,
)
_STATE_LABELS = {
"complete": "Complete",
"pending": "Pending",
"blocked": "Blocked",
"not_applicable": "Not applicable",
}
def _escape(text: str) -> str: def _escape(text: str) -> str:
return html.escape(text, quote=True) return html.escape(text, quote=True)
def _progress_label(summary: dict[str, Any]) -> str:
total = summary["total"]
if not total:
return "no steps"
label = f"{summary['complete']}/{total} complete"
if summary["blocked"]:
label += f", {summary['blocked']} blocked"
return label
def render_projects_list(registry: ProjectRegistry) -> str: def render_projects_list(registry: ProjectRegistry) -> str:
rows = [] rows = []
for project in registry.projects: for project in registry.projects:
dto = project_to_dict(project)
rows.append( rows.append(
"<tr>" "<tr>"
f"<td><a href=\"/projects/{_escape(project.id)}\">{_escape(project.repo_name)}</a></td>" f"<td><a href=\"/projects/{_escape(dto['id'])}\">{_escape(dto['repo_name'])}</a></td>"
f"<td>{_escape(project.gitea_owner)}</td>" f"<td>{_escape(dto['gitea_owner'])}</td>"
f"<td>{_escape(project.remote_host)}</td>" f"<td>{_escape(dto['remote_host'])}</td>"
f"<td>{_escape(project.default_branch)}</td>" f"<td>{_escape(dto['default_branch'])}</td>"
f"<td><code>{_escape(project.profiles['author'])}</code></td>" f"<td><code>{_escape(dto['status'])}</code></td>"
f"<td>{_escape(_progress_label(dto['onboarding_summary']))}</td>"
f"<td><code>{_escape(dto['profiles']['author'])}</code></td>"
"</tr>" "</tr>"
) )
table = ( table = (
"<table class=\"registry\">" "<table class=\"registry\">"
"<thead><tr>" "<thead><tr>"
"<th>Repository</th><th>Owner</th><th>Remote</th>" "<th>Repository</th><th>Owner</th><th>Remote</th>"
"<th>Branch</th><th>Author profile</th>" "<th>Branch</th><th>Status</th><th>Onboarding</th><th>Author profile</th>"
"</tr></thead>" "</tr></thead>"
f"<tbody>{''.join(rows)}</tbody></table>" f"<tbody>{''.join(rows)}</tbody></table>"
) )
@@ -36,32 +67,39 @@ def render_projects_list(registry: ProjectRegistry) -> str:
"<h2>Projects</h2>" "<h2>Projects</h2>"
"<p>Configured repositories managed by the MCP Control Plane.</p>" "<p>Configured repositories managed by the MCP Control Plane.</p>"
f"<p class=\"meta\">Registry: <code>{_escape(str(registry.source_path))}</code> " f"<p class=\"meta\">Registry: <code>{_escape(str(registry.source_path))}</code> "
f"(version {registry.version})</p>" f"(schema version {registry.schema_version}, "
f"API {_escape(registry.api_version)})</p>"
f"{table}" f"{table}"
"<p><a href=\"/api/projects\">JSON API</a></p>" "<p><a href=\"/api/v1/projects\">JSON API</a> "
"(<a href=\"/api/projects\">unversioned alias</a>)</p>"
) )
return render_page(title="Projects", body_html=body) return render_page(title="Projects", body_html=body)
def render_project_detail(project: ProjectRecord) -> str: def render_project_detail(project: ProjectRecord) -> str:
dto = project_to_dict(project)
profile_rows = "".join( profile_rows = "".join(
f"<tr><th>{_escape(role)}</th><td><code>{_escape(name)}</code></td></tr>" f"<tr><th>{_escape(role)}</th><td><code>{_escape(name)}</code></td></tr>"
for role, name in project.profiles.items() for role, name in dto["profiles"].items()
) )
workflow_rows = "".join( workflow_rows = "".join(
f"<tr><th>{_escape(key)}</th><td><code>{_escape(path)}</code></td></tr>" f"<tr><th>{_escape(key)}</th><td><code>{_escape(path)}</code></td></tr>"
for key, path in project.workflow_paths.items() for key, path in dto["workflow_paths"].items()
) )
schema_rows = "".join( schema_rows = "".join(
f"<tr><th>{_escape(key)}</th><td><code>{_escape(path)}</code></td></tr>" f"<tr><th>{_escape(key)}</th><td><code>{_escape(path)}</code></td></tr>"
for key, path in project.schema_paths.items() for key, path in dto["schema_paths"].items()
) )
checklist_items = [] checklist_items = []
for index, step in enumerate(project.onboarding_checklist, start=1): for index, step in enumerate(dto["onboarding_checklist"], start=1):
state_label = _STATE_LABELS.get(step["state"], step["state"])
requirement = "required" if step["required"] else "optional"
checklist_items.append( checklist_items.append(
"<li>" f"<li class=\"step-{_escape(step['state'])}\">"
f"<strong>{index}. {_escape(step.title)}</strong>" f"<strong>{index}. {_escape(step['title'])}</strong>"
f"<p>{_escape(step.description)}</p>" f" <span class=\"badge\">{_escape(state_label)}</span>"
f" <span class=\"meta\">({_escape(requirement)})</span>"
f"<p>{_escape(step['description'])}</p>"
"</li>" "</li>"
) )
checklist_html = ( checklist_html = (
@@ -69,16 +107,38 @@ def render_project_detail(project: ProjectRecord) -> str:
if checklist_items if checklist_items
else "<p>No onboarding steps defined.</p>" else "<p>No onboarding steps defined.</p>"
) )
summary = dto["onboarding_summary"]
summary_html = (
"<p class=\"meta\">Onboarding: "
f"{_escape(_progress_label(summary))}; required outstanding "
f"{summary['required_outstanding']}.</p>"
)
health = dto["last_seen_health"]
health_html = (
"<p class=\"meta\">No health probe recorded (Phase 1 is read-only).</p>"
if health is None
else (
"<table class=\"detail\">"
f"<tr><th>Status</th><td><code>{_escape(health['status'])}</code></td></tr>"
f"<tr><th>Checked at</th><td>{_escape(str(health['checked_at'] or 'unknown'))}</td></tr>"
f"<tr><th>Detail</th><td>{_escape(str(health['detail'] or ''))}</td></tr>"
"</table>"
)
)
remote_name = dto["remote_name"] or "unset"
body = ( body = (
f"<h2>{_escape(project.repo_name)}</h2>" f"<h2>{_escape(dto['repo_name'])}</h2>"
"<p><a href=\"/projects\">← All projects</a></p>" "<p><a href=\"/projects\">← All projects</a></p>"
"<h3>Identity</h3>" "<h3>Identity</h3>"
"<table class=\"detail\">" "<table class=\"detail\">"
f"<tr><th>Registry id</th><td><code>{_escape(project.id)}</code></td></tr>" f"<tr><th>Registry id</th><td><code>{_escape(dto['id'])}</code></td></tr>"
f"<tr><th>Gitea owner</th><td>{_escape(project.gitea_owner)}</td></tr>" f"<tr><th>Status</th><td><code>{_escape(dto['status'])}</code></td></tr>"
f"<tr><th>Remote host</th><td>{_escape(project.remote_host)}</td></tr>" f"<tr><th>Gitea owner</th><td>{_escape(dto['gitea_owner'])}</td></tr>"
f"<tr><th>Default branch</th><td><code>{_escape(project.default_branch)}</code></td></tr>" f"<tr><th>Repository</th><td><code>{_escape(dto['repo_full_name'])}</code></td></tr>"
f"<tr><th>Local checkout</th><td><code>{_escape(project.local_checkout_path)}</code></td></tr>" f"<tr><th>Remote name</th><td><code>{_escape(remote_name)}</code></td></tr>"
f"<tr><th>Remote host</th><td>{_escape(dto['remote_host'])}</td></tr>"
f"<tr><th>Default branch</th><td><code>{_escape(dto['default_branch'])}</code></td></tr>"
f"<tr><th>Local checkout</th><td><code>{_escape(dto['local_checkout_path'])}</code></td></tr>"
"</table>" "</table>"
"<h3>Profiles</h3>" "<h3>Profiles</h3>"
f"<table class=\"detail\">{profile_rows}</table>" f"<table class=\"detail\">{profile_rows}</table>"
@@ -86,8 +146,30 @@ def render_project_detail(project: ProjectRecord) -> str:
f"<table class=\"detail\">{workflow_rows}</table>" f"<table class=\"detail\">{workflow_rows}</table>"
"<h3>Schema paths</h3>" "<h3>Schema paths</h3>"
f"<table class=\"detail\">{schema_rows}</table>" f"<table class=\"detail\">{schema_rows}</table>"
"<h3>Last seen health</h3>"
f"{health_html}"
"<h3>Onboarding checklist</h3>" "<h3>Onboarding checklist</h3>"
"<p class=\"meta\">Read-only MVP — complete these steps outside the UI.</p>" "<p class=\"meta\">Read-only — complete these steps outside the UI.</p>"
f"{summary_html}"
f"{checklist_html}" f"{checklist_html}"
f"<p><a href=\"/api/v1/projects/{_escape(dto['id'])}\">JSON detail</a></p>"
) )
return render_page(title=project.repo_name, body_html=body) return render_page(title=dto["repo_name"], body_html=body)
def render_registry_error(error: RegistryError) -> str:
"""Render a fail-closed page for an invalid registry."""
source = str(error.source_path) if error.source_path else "unknown"
field = error.field_path or "n/a"
body = (
"<h2>Project registry unavailable</h2>"
"<p>The registry failed validation, so the console refuses to render a "
"partial inventory.</p>"
"<table class=\"detail\">"
f"<tr><th>Detail</th><td>{_escape(error.message)}</td></tr>"
f"<tr><th>Field</th><td><code>{_escape(field)}</code></td></tr>"
f"<tr><th>Source</th><td><code>{_escape(source)}</code></td></tr>"
f"<tr><th>Remediation</th><td>{_escape(error.remediation)}</td></tr>"
"</table>"
)
return render_page(title="Project registry unavailable", body_html=body)
+52
View File
@@ -0,0 +1,52 @@
"""Shared credential-rejection guard for web UI registries (#427, #798).
Registries are operator-editable declarative files that the web UI loads and,
for the worker registry, writes back. None of them may ever carry a secret:
credentials belong in the keychain and reach worker processes through
environment injection, never through a file the browser layer can read.
The check is structural rather than value-based on purpose. A value scanner has
to guess what a secret looks like; a key scanner refuses the *shape* of a
credential field, so an operator cannot introduce one by accident and a later
loader cannot silently pass one through.
"""
from __future__ import annotations
from typing import Any
_FORBIDDEN_EXACT_KEYS = frozenset({
"token",
"password",
"secret",
"credential",
"auth",
"api_key",
"api-key",
})
_FORBIDDEN_KEY_PREFIXES = ("auth_", "api_key_", "api-key_")
_FORBIDDEN_KEY_SUFFIXES = ("_token", "_secret", "_password", "_credential", "_auth")
def is_forbidden_key(key: str) -> bool:
"""Return True when *key* names a credential field."""
lowered = key.lower()
if lowered in _FORBIDDEN_EXACT_KEYS:
return True
return (
lowered.startswith(_FORBIDDEN_KEY_PREFIXES)
or lowered.endswith(_FORBIDDEN_KEY_SUFFIXES)
)
def reject_credential_keys(obj: Any, *, path: str = "", subject: str = "registry") -> None:
"""Raise ValueError when *obj* carries a credential-shaped key at any depth."""
if isinstance(obj, dict):
for key, value in obj.items():
key_path = f"{path}.{key}" if path else key
if is_forbidden_key(key):
raise ValueError(f"{subject} must not store credentials ({key_path})")
reject_credential_keys(value, path=key_path, subject=subject)
elif isinstance(obj, list):
for index, item in enumerate(obj):
reject_credential_keys(item, path=f"{path}[{index}]", subject=subject)
+647
View File
@@ -0,0 +1,647 @@
"""Declarative worker registry and configuration schema (#798, epic #797).
The registry is the single source of truth for the scheduled multi-LLM worker
fleet. It is a versioned JSON document holding two *separate* entity kinds:
* **Providers** — the LLM runtimes a worker can be built on (Claude, Grok,
Codex, AGY, Kimi K). A provider describes the runtime itself: vendor,
executable name, models it can serve, and whether it is available on this
machine. Providers exist whether or not any worker uses them.
* **Workers** — a configured *instance*: one provider, one model, one project,
one role, one MCP namespace/profile, one workflow, one schedule. Several
workers may share a provider; a worker naming an undeclared provider is
refused.
Keeping them separate is what lets #799 list all five providers even when a
provider currently has no configured worker, and it stops provider facts from
being copied into (and drifting across) every worker record.
Scope boundary. This module owns the data model, its validation, and its
persistence. It does **not** schedule anything, launch anything, probe provider
executables, or serve HTTP. Loading a registry never touches a process; the
live fields a dashboard wants (PID, elapsed time, next run) are derived
elsewhere (#799, #801, #803, #804) from these declarations.
Safety invariants:
* No credential may be stored (:mod:`webui.registry_safety`), so the registry
stays safe to render and to hand to a browser layer.
* Validation fails closed. Unknown fields are refused rather than ignored, so a
typo cannot silently disable a timeout or a role binding.
* Writes are atomic and every superseded document is retained as a numbered
revision, so a bad edit is recoverable by rollback rather than hand-repair.
"""
from __future__ import annotations
import json
import os
import re
import tempfile
from dataclasses import dataclass
from datetime import datetime, timezone
from pathlib import Path
from typing import Any
from webui.registry_safety import reject_credential_keys
SCHEMA_VERSION = 1
#: Roles a worker may hold. These mirror the sanctioned MCP role kinds; a
#: worker may not invent one, because the role selects the namespace/profile
#: whose capability gates constrain it.
ALLOWED_ROLES = ("author", "reviewer", "merger", "reconciler", "cleanup")
#: Scheduler backends the registry can describe. ``manual`` means the worker is
#: only ever started on request and has no recurring trigger.
ALLOWED_SCHEDULER_KINDS = ("launchd", "manual")
#: Schedule kinds. Next-run computation belongs to #803; this module only
#: guarantees the declaration is well formed.
ALLOWED_SCHEDULE_KINDS = ("interval", "cron", "manual")
_REQUIRED_PROVIDER_FIELDS = ("id", "display_name", "vendor", "executable", "available")
_OPTIONAL_PROVIDER_FIELDS = ("models", "notes")
_REQUIRED_WORKER_FIELDS = (
"id",
"display_name",
"provider",
"model",
"project",
"role",
"namespace",
"profile",
"workflow",
"schedule",
"timeout_seconds",
"enabled",
"scheduler",
)
_OPTIONAL_WORKER_FIELDS = ("notes",)
_ID_RE = re.compile(r"^[a-z0-9][a-z0-9._-]*$")
#: Guards against an operator writing a timeout that would let a worker hold a
#: lease effectively forever. 24h is far above any sanctioned cycle.
_MAX_TIMEOUT_SECONDS = 86_400
#: How many superseded revisions to retain beside the live file.
_HISTORY_LIMIT = 20
_TOP_LEVEL_FIELDS = frozenset({"version", "revision", "updated_at", "providers", "workers"})
class RegistryValidationError(ValueError):
"""Raised when a registry document violates the schema."""
@dataclass(frozen=True)
class ProviderRecord:
id: str
display_name: str
vendor: str
executable: str
available: bool
models: tuple[str, ...]
notes: str
@dataclass(frozen=True)
class ScheduleSpec:
kind: str
#: Set for ``interval`` schedules.
seconds: int | None
#: Set for ``cron`` schedules — a five-field crontab expression.
expression: str | None
@dataclass(frozen=True)
class SchedulerSpec:
kind: str
#: LaunchAgent label; required for ``launchd``, absent for ``manual``.
label: str | None
@dataclass(frozen=True)
class WorkerRecord:
id: str
display_name: str
provider: str
model: str
project: str
role: str
namespace: str
profile: str
workflow: str
schedule: ScheduleSpec
timeout_seconds: int
enabled: bool
scheduler: SchedulerSpec
notes: str
@dataclass(frozen=True)
class WorkerRegistry:
version: int
revision: int
updated_at: str
providers: tuple[ProviderRecord, ...]
workers: tuple[WorkerRecord, ...]
source_path: Path
# ── paths ────────────────────────────────────────────────────────────────────
def default_registry_path() -> Path:
"""Location of the packaged worker registry, overridable for tests/deploys."""
override = os.environ.get("WEBUI_WORKER_REGISTRY", "").strip()
if override:
return Path(override).expanduser().resolve()
return (Path(__file__).resolve().parent / "data" / "workers.registry.json").resolve()
def history_dir(path: Path | None = None) -> Path:
"""Directory holding superseded revisions of *path*."""
source = (path or default_registry_path()).resolve()
return source.parent / f"{source.name}.history"
# ── field helpers ────────────────────────────────────────────────────────────
def _require_exact_fields(
raw: Any,
*,
required: tuple[str, ...],
optional: tuple[str, ...],
subject: str,
) -> dict[str, Any]:
if not isinstance(raw, dict):
raise RegistryValidationError(f"{subject} must be an object")
missing = [field for field in required if field not in raw]
if missing:
raise RegistryValidationError(
f"{subject} missing required fields: {', '.join(sorted(missing))}"
)
unknown = sorted(set(raw) - set(required) - set(optional))
if unknown:
# Fail closed: silently dropping an unrecognized key is how a typo'd
# "timeout_second" ends up meaning "no timeout".
raise RegistryValidationError(f"{subject} has unknown fields: {', '.join(unknown)}")
return raw
def _require_identifier(value: Any, *, subject: str) -> str:
text = str(value).strip()
if not _ID_RE.match(text):
raise RegistryValidationError(
f"{subject} must be lowercase alphanumeric with '.', '_', or '-' (got {value!r})"
)
return text
def _require_text(value: Any, *, subject: str) -> str:
if not isinstance(value, str):
raise RegistryValidationError(f"{subject} must be a string (got {value!r})")
text = value.strip()
if not text:
raise RegistryValidationError(f"{subject} must be a non-empty string")
return text
def _require_bool(value: Any, *, subject: str) -> bool:
if not isinstance(value, bool):
raise RegistryValidationError(f"{subject} must be a boolean (got {value!r})")
return value
def _require_positive_int(value: Any, *, subject: str, maximum: int | None = None) -> int:
if isinstance(value, bool) or not isinstance(value, int):
raise RegistryValidationError(f"{subject} must be an integer (got {value!r})")
if value <= 0:
raise RegistryValidationError(f"{subject} must be greater than zero (got {value})")
if maximum is not None and value > maximum:
raise RegistryValidationError(f"{subject} must not exceed {maximum} (got {value})")
return value
# ── parsing ──────────────────────────────────────────────────────────────────
def _parse_provider(raw: Any) -> ProviderRecord:
data = _require_exact_fields(
raw,
required=_REQUIRED_PROVIDER_FIELDS,
optional=_OPTIONAL_PROVIDER_FIELDS,
subject="provider",
)
provider_id = _require_identifier(data["id"], subject="provider.id")
models_raw = data.get("models") or []
if not isinstance(models_raw, list):
raise RegistryValidationError(f"provider[{provider_id}].models must be an array")
models = tuple(
_require_text(item, subject=f"provider[{provider_id}].models[]") for item in models_raw
)
return ProviderRecord(
id=provider_id,
display_name=_require_text(
data["display_name"], subject=f"provider[{provider_id}].display_name"
),
vendor=_require_text(data["vendor"], subject=f"provider[{provider_id}].vendor"),
executable=_require_text(data["executable"], subject=f"provider[{provider_id}].executable"),
available=_require_bool(data["available"], subject=f"provider[{provider_id}].available"),
models=models,
notes=str(data.get("notes") or "").strip(),
)
def _parse_schedule(raw: Any, *, subject: str) -> ScheduleSpec:
if not isinstance(raw, dict):
raise RegistryValidationError(f"{subject} must be an object")
kind = _require_text(raw.get("kind"), subject=f"{subject}.kind")
if kind not in ALLOWED_SCHEDULE_KINDS:
raise RegistryValidationError(
f"{subject}.kind must be one of {', '.join(ALLOWED_SCHEDULE_KINDS)} (got {kind!r})"
)
seconds: int | None = None
expression: str | None = None
if kind == "interval":
if "seconds" not in raw:
raise RegistryValidationError(f"{subject}.seconds is required for interval schedules")
seconds = _require_positive_int(raw["seconds"], subject=f"{subject}.seconds")
elif kind == "cron":
if "expression" not in raw:
raise RegistryValidationError(f"{subject}.expression is required for cron schedules")
expression = _require_text(raw["expression"], subject=f"{subject}.expression")
if len(expression.split()) != 5:
raise RegistryValidationError(
f"{subject}.expression must have five crontab fields (got {expression!r})"
)
allowed = {"kind"}
if kind == "interval":
allowed.add("seconds")
elif kind == "cron":
allowed.add("expression")
unknown = sorted(set(raw) - allowed)
if unknown:
raise RegistryValidationError(
f"{subject} has fields not valid for kind {kind!r}: {', '.join(unknown)}"
)
return ScheduleSpec(kind=kind, seconds=seconds, expression=expression)
def _parse_scheduler(raw: Any, *, subject: str) -> SchedulerSpec:
if not isinstance(raw, dict):
raise RegistryValidationError(f"{subject} must be an object")
kind = _require_text(raw.get("kind"), subject=f"{subject}.kind")
if kind not in ALLOWED_SCHEDULER_KINDS:
raise RegistryValidationError(
f"{subject}.kind must be one of {', '.join(ALLOWED_SCHEDULER_KINDS)} (got {kind!r})"
)
label: str | None = None
if kind == "launchd":
if "label" not in raw:
raise RegistryValidationError(f"{subject}.label is required for launchd schedulers")
label = _require_text(raw["label"], subject=f"{subject}.label")
allowed = {"kind"}
if kind == "launchd":
allowed.add("label")
unknown = sorted(set(raw) - allowed)
if unknown:
raise RegistryValidationError(
f"{subject} has fields not valid for kind {kind!r}: {', '.join(unknown)}"
)
return SchedulerSpec(kind=kind, label=label)
def _parse_worker(raw: Any) -> WorkerRecord:
data = _require_exact_fields(
raw,
required=_REQUIRED_WORKER_FIELDS,
optional=_OPTIONAL_WORKER_FIELDS,
subject="worker",
)
worker_id = _require_identifier(data["id"], subject="worker.id")
role = _require_text(data["role"], subject=f"worker[{worker_id}].role")
if role not in ALLOWED_ROLES:
raise RegistryValidationError(
f"worker[{worker_id}].role must be one of {', '.join(ALLOWED_ROLES)} (got {role!r})"
)
return WorkerRecord(
id=worker_id,
display_name=_require_text(
data["display_name"], subject=f"worker[{worker_id}].display_name"
),
provider=_require_identifier(data["provider"], subject=f"worker[{worker_id}].provider"),
model=_require_text(data["model"], subject=f"worker[{worker_id}].model"),
project=_require_text(data["project"], subject=f"worker[{worker_id}].project"),
role=role,
namespace=_require_text(data["namespace"], subject=f"worker[{worker_id}].namespace"),
profile=_require_text(data["profile"], subject=f"worker[{worker_id}].profile"),
workflow=_require_text(data["workflow"], subject=f"worker[{worker_id}].workflow"),
schedule=_parse_schedule(data["schedule"], subject=f"worker[{worker_id}].schedule"),
timeout_seconds=_require_positive_int(
data["timeout_seconds"],
subject=f"worker[{worker_id}].timeout_seconds",
maximum=_MAX_TIMEOUT_SECONDS,
),
enabled=_require_bool(data["enabled"], subject=f"worker[{worker_id}].enabled"),
scheduler=_parse_scheduler(data["scheduler"], subject=f"worker[{worker_id}].scheduler"),
notes=str(data.get("notes") or "").strip(),
)
def _require_unique(values: list[str], *, subject: str) -> None:
seen: set[str] = set()
for value in values:
if value in seen:
raise RegistryValidationError(f"duplicate {subject}: {value}")
seen.add(value)
def validate_payload(payload: Any, *, source_path: Path) -> WorkerRegistry:
"""Validate a decoded registry document and return the typed registry.
Raises :class:`RegistryValidationError` on any violation; never partially
accepts a document.
"""
if not isinstance(payload, dict):
raise RegistryValidationError("registry root must be an object")
version = payload.get("version")
if version != SCHEMA_VERSION:
raise RegistryValidationError(f"unsupported registry version: {version!r}")
reject_credential_keys(payload, subject="worker registry")
unknown = sorted(set(payload) - _TOP_LEVEL_FIELDS)
if unknown:
raise RegistryValidationError(f"registry has unknown fields: {', '.join(unknown)}")
revision = _require_positive_int(payload.get("revision"), subject="revision")
updated_at = _require_text(payload.get("updated_at"), subject="updated_at")
providers_raw = payload.get("providers")
if not isinstance(providers_raw, list) or not providers_raw:
raise RegistryValidationError("providers must be a non-empty array")
providers = tuple(_parse_provider(item) for item in providers_raw)
_require_unique([provider.id for provider in providers], subject="provider id")
workers_raw = payload.get("workers")
if not isinstance(workers_raw, list):
raise RegistryValidationError("workers must be an array")
workers = tuple(_parse_worker(item) for item in workers_raw)
_require_unique([worker.id for worker in workers], subject="worker id")
# Referential integrity: a worker naming an undeclared provider would look
# configured while being unrunnable, which is exactly the ambiguous
# ownership the epic requires to fail closed.
known_providers = {provider.id for provider in providers}
for worker in workers:
if worker.provider not in known_providers:
raise RegistryValidationError(
f"worker[{worker.id}].provider references unknown provider {worker.provider!r}"
)
# A LaunchAgent label identifies a job to launchd; two workers sharing one
# would silently overwrite each other's agent.
_require_unique(
[worker.scheduler.label for worker in workers if worker.scheduler.label],
subject="scheduler label",
)
return WorkerRegistry(
version=version,
revision=revision,
updated_at=updated_at,
providers=providers,
workers=workers,
source_path=source_path,
)
def load_registry(path: Path | None = None) -> WorkerRegistry:
"""Load and validate the worker registry from disk."""
source = (path or default_registry_path()).resolve()
payload = json.loads(source.read_text(encoding="utf-8"))
return validate_payload(payload, source_path=source)
# ── serialization ────────────────────────────────────────────────────────────
def provider_to_dict(provider: ProviderRecord) -> dict[str, Any]:
return {
"id": provider.id,
"display_name": provider.display_name,
"vendor": provider.vendor,
"executable": provider.executable,
"available": provider.available,
"models": list(provider.models),
"notes": provider.notes,
}
def _schedule_to_dict(schedule: ScheduleSpec) -> dict[str, Any]:
payload: dict[str, Any] = {"kind": schedule.kind}
if schedule.kind == "interval":
payload["seconds"] = schedule.seconds
elif schedule.kind == "cron":
payload["expression"] = schedule.expression
return payload
def _scheduler_to_dict(scheduler: SchedulerSpec) -> dict[str, Any]:
payload: dict[str, Any] = {"kind": scheduler.kind}
if scheduler.kind == "launchd":
payload["label"] = scheduler.label
return payload
def worker_to_dict(worker: WorkerRecord) -> dict[str, Any]:
return {
"id": worker.id,
"display_name": worker.display_name,
"provider": worker.provider,
"model": worker.model,
"project": worker.project,
"role": worker.role,
"namespace": worker.namespace,
"profile": worker.profile,
"workflow": worker.workflow,
"schedule": _schedule_to_dict(worker.schedule),
"timeout_seconds": worker.timeout_seconds,
"enabled": worker.enabled,
"scheduler": _scheduler_to_dict(worker.scheduler),
"notes": worker.notes,
}
def registry_to_document(registry: WorkerRegistry) -> dict[str, Any]:
"""Serialize to the on-disk document shape (no local paths embedded)."""
return {
"version": registry.version,
"revision": registry.revision,
"updated_at": registry.updated_at,
"providers": [provider_to_dict(provider) for provider in registry.providers],
"workers": [worker_to_dict(worker) for worker in registry.workers],
}
def registry_to_dict(registry: WorkerRegistry) -> dict[str, Any]:
"""Serialize for JSON API responses (adds the resolved source path)."""
document = registry_to_document(registry)
document["source_path"] = str(registry.source_path)
return document
def find_worker(registry: WorkerRegistry, worker_id: str) -> WorkerRecord | None:
for worker in registry.workers:
if worker.id == worker_id:
return worker
return None
def find_provider(registry: WorkerRegistry, provider_id: str) -> ProviderRecord | None:
for provider in registry.providers:
if provider.id == provider_id:
return provider
return None
def workers_for_provider(registry: WorkerRegistry, provider_id: str) -> tuple[WorkerRecord, ...]:
return tuple(worker for worker in registry.workers if worker.provider == provider_id)
# ── persistence ──────────────────────────────────────────────────────────────
def _utc_now() -> str:
return datetime.now(timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ")
def _atomic_write(path: Path, payload: str) -> None:
"""Write *payload* to *path* atomically: temp file in the same dir, fsync, replace."""
parent = path.parent
parent.mkdir(parents=True, exist_ok=True)
fd, temp_path = tempfile.mkstemp(prefix=f".{path.name}-", suffix=".tmp", dir=parent)
try:
with os.fdopen(fd, "w", encoding="utf-8") as handle:
handle.write(payload)
handle.flush()
os.fsync(handle.fileno())
os.replace(temp_path, path)
finally:
if os.path.exists(temp_path):
try:
os.remove(temp_path)
except OSError:
pass
def _revision_path(directory: Path, revision: int) -> Path:
return directory / f"rev-{revision:06d}.json"
def _prune_history(path: Path) -> None:
directory = history_dir(path)
revisions = list_revisions(path)
excess = len(revisions) - _HISTORY_LIMIT
for revision in revisions[: max(0, excess)]:
_revision_path(directory, revision).unlink(missing_ok=True)
def _archive_current(path: Path) -> int | None:
"""Copy the live document into the history dir under its own revision number."""
if not path.exists():
return None
try:
existing = json.loads(path.read_text(encoding="utf-8"))
revision = int(existing.get("revision", 0))
except (json.JSONDecodeError, TypeError, ValueError, AttributeError):
# An unreadable live file has no trustworthy revision number to file it
# under, so it cannot join the history chain.
return None
if revision <= 0:
return None
_atomic_write(
_revision_path(history_dir(path), revision),
json.dumps(existing, indent=2, sort_keys=True) + "\n",
)
_prune_history(path)
return revision
def list_revisions(path: Path | None = None) -> tuple[int, ...]:
"""Revision numbers retained in history for *path*, oldest first."""
directory = history_dir(path)
if not directory.is_dir():
return ()
revisions: list[int] = []
for entry in directory.glob("rev-*.json"):
try:
revisions.append(int(entry.stem.split("-", 1)[1]))
except (IndexError, ValueError):
continue
return tuple(sorted(revisions))
def save_registry(
registry: WorkerRegistry,
path: Path | None = None,
*,
updated_at: str | None = None,
) -> WorkerRegistry:
"""Validate, archive the superseded revision, then atomically persist a new one.
The stored revision is always the previous revision plus one, so a reader
can tell two documents apart even when their content is otherwise equal.
Returns the registry exactly as persisted.
"""
target = (path or registry.source_path or default_registry_path()).resolve()
document = registry_to_document(registry)
# Re-validate before writing: a registry assembled in memory has not
# necessarily been through the loader.
validate_payload(document, source_path=target)
archived = _archive_current(target)
document["revision"] = (archived + 1) if archived is not None else registry.revision
document["updated_at"] = updated_at or _utc_now()
persisted = validate_payload(document, source_path=target)
_atomic_write(target, json.dumps(document, indent=2, sort_keys=True) + "\n")
return persisted
def rollback_to_revision(revision: int, path: Path | None = None) -> WorkerRegistry:
"""Restore a retained *revision* as a new head revision.
History is append-only: rolling back does not delete the revisions in
between, it republishes the chosen one under the next revision number, so a
rollback is itself reversible.
"""
target = (path or default_registry_path()).resolve()
snapshot_path = _revision_path(history_dir(target), revision)
if not snapshot_path.exists():
available = ", ".join(str(item) for item in list_revisions(target)) or "(none)"
raise RegistryValidationError(
f"revision {revision} is not retained for {target.name}; available: {available}"
)
payload = json.loads(snapshot_path.read_text(encoding="utf-8"))
restored = validate_payload(payload, source_path=target)
return save_registry(restored, target)