# Stale worker retirement (#980) The #978 instance-fleet snapshot made registry accuracy observable but deliberately read-only: a registry full of rows whose owning processes are long gone stays full. This document describes the sanctioned way to retire those rows — a dry-run-first, compare-and-swap-protected workflow available only to controller and reconciler namespaces. Related: [instance-fleet-identity.md](instance-fleet-identity.md) (#978), [post-restart-reconcile.md](post-restart-reconcile.md) (#662). ## Why a dedicated registry token `mcp_fleet_snapshot._consistency_token` seeds its digest with `snapshot_at`, formatted at second precision. Its output — surfaced as `registry_revision` and `consistency_token` on the snapshot — therefore changes on **every call**, even when no registry row changed. Any compare-and-swap gated on it can never pass: a dry-run/apply cycle spanning more than one second aborts unconditionally. That token remains useful as an observation stamp and is unchanged. #980 adds a separate, *stable* token instead: | Token | Module | Derived from | Stable across time? | | --- | --- | --- | --- | | `registry_revision` / `consistency_token` | `mcp_fleet_snapshot` | `snapshot_at` + a subset of row fields | **No** — moves every second | | `registry_fingerprint` | `mcp_fleet_retirement` | canonical retirement-relevant row content only | **Yes** | | `candidate_fingerprint` | `mcp_fleet_retirement` | canonical content of the selected candidate rows | **Yes** | `registry_fingerprint` guarantees: * identical canonical registry contents always produce the same token, whenever they are observed; * row iteration order never affects the token (serialized rows are sorted); * any retirement-relevant change moves it — row creation or deletion, identity change, heartbeat or TTL change, ownership change, registration-state change, PID change, or repository-binding change. The exact field set is `mcp_fleet_retirement.FINGERPRINT_FIELDS`. Deliberately excluded: `token_fingerprint` (credential-adjacent, never a retirement input), the four `*_revision` columns (revision drift is an independent restart concern and is not part of the eligibility conjunction), and the `retired_*` bookkeeping columns this feature adds. Numeric values are canonicalised, so a TTL that round-trips through SQLite as `900.0` hashes identically to `900`. ## Eligibility — the conjunction A registration is retired only when **every** one of these holds. Any missing or contradictory evidence preserves the row. | Requirement | Preserve reason code when it fails | | --- | --- | | `status` is `active` | `already_terminal_registration` | | Every field the conjunction reads is present (`REQUIRED_IDENTITY_FIELDS`) | `incomplete_registry_identity` | | `last_heartbeat_at` parses as a UTC stamp | `unparsable_heartbeat` | | Worker is not live | `worker_live` | | PID probe returns a definite answer | `pid_liveness_unknown` | | PID probe says the process is gone | `pid_alive` | | Heartbeat has expired under the canonical TTL | `heartbeat_not_expired` | | `ownership_state` is exactly `stale` | `ambiguous_ownership_state` | | Repository binding present and canonical | `repository_binding_ambiguous` | | No identity evidence shared with a live or unprobeable worker | `conflicting_identity_evidence` | | No other active row claims the same (instance, namespace) while one may be live | `client_instance_conflict` | | Not an active workflow-lease owner | `protected_active_workflow_owner` | | Instance identity is launcher-minted (`inst-…`) | `untrusted_identity_provenance` | | Row's `host_id` matches the host running retirement | `host_binding_unproven` | | Boot identity is known on both sides | `boot_identity_unknown` | | The pid number is not occupied by a different incarnation | `pid_reuse_detected` | Eligible rows carry `eligible_stale_orphan`. Three properties are worth stating explicitly: * **`pid_alive` can only withdraw liveness, never grant it** (`WorkerRegistry.is_live`, #948 AC7). A heartbeat-lapsed but still-running process therefore classifies as `stale` in the snapshot, yet #980's added `pid_alive is False` requirement preserves it. An unprobeable PID (`None`) also fails closed. * **Affirmative identity proof is required (review 657 B2).** An earlier revision required only that the pre-existing registry columns were non-null — which a legacy `legacy-pid-…` row satisfies trivially, so a row that proved nothing about *which* process it described was retireable. Retirement now needs both halves of a positive proof: * **Attribution** — a launcher-minted `inst-…` `client_instance_id`, so the row is known to belong to one specific application launch rather than having been inferred from pid proximity. * **Fencing** — `host_id`, `boot_id`, and `process_start_time`, which turn a bare pid into a statement about one process: which machine it ran on, which boot of that machine, and which incarnation of that pid number. **Consequence, stated plainly:** registrations written before these columns existed, and any row on a legacy instance identity, are preserved *permanently*. They are retired only after their worker re-registers under a trusted identity — never on weaker evidence. That the alternative would leave legacy rows outstanding indefinitely is not a reason to relax the proof. * **A live pid is an absolute block.** Even across a boot boundary, where the number provably cannot belong to the registered process, an occupied pid preserves the row rather than being argued away by the fencing proof. Multiple processes belonging to one legitimate worker cohort are not treated as multiple independent workers: the fleet model from #948/#978 is preserved unchanged, and sharing a role or profile is never a duplicate. ## Tools ### `gitea_plan_stale_worker_retirement` Read-only. Controller and reconciler only. | Parameter | Meaning | | --- | --- | | `remote` | `dadeschools` or `prgs` | | `host`, `org`, `repo` | Optional overrides (audit context) | | `canonical_repository` | Expected repository binding; defaults to the process root | Returns `registry_fingerprint`, `candidate_fingerprint`, `candidate_worker_identities`, per-worker `candidates` and `preserved` entries (each with `reason_code`, `detail`, and structured `evidence`), `preserved_reason_counts`, `assessed_count`, `candidate_count`, `preserved_count`, and `protected_active_workflow_owners`. `mutation_performed` is always `false` and `read_only` is always `true`. Planning is deterministic: the same authoritative registry contents produce the same plan and the same tokens regardless of when they are observed. ### `gitea_apply_stale_worker_retirement` Mutating. Controller and reconciler only. | Parameter | Meaning | | --- | --- | | `registry_fingerprint` | The exact stable token the plan returned | | `candidate_fingerprint` | The exact candidate-set token the plan returned | | `worker_identities` | The exact candidate identities (list, or JSON / comma-separated string) | | `remote`, `host`, `org`, `repo` | As above | | `canonical_repository` | Must match the value the plan used | Before touching the registry, apply fails closed on: profile permission, role kind, master parity (`mutation_safe`), stable-runtime mode, capability resolution refreshed immediately before mutation, worker-registry availability, workflow-lease enumeration failure, and daemon-cohort uniqueness (`classify_cohort`). A matching token is necessary but never sufficient. Inside one `BEGIN IMMEDIATE` transaction (`WorkerRegistry.retire_stale_workers`) the server: 1. re-reads the authoritative rows; 2. recomputes `registry_fingerprint` from *those* rows and compares — a mismatch returns `registry_revision_moved` with `retired_count: 0` and no write; 3. recomputes the eligibility plan from *those* rows — re-reading the active workflow leases rather than reusing the set captured before the transaction opened, so a lease acquired after planning still preserves its worker — and compares `candidate_fingerprint`. A mismatch returns `candidate_set_moved` with `retired_count: 0` and no write; a lease-enumeration failure raises and rolls the transaction back; 4. revalidates every requested identity against that fresh plan; 5. retires each survivor with a guarded `UPDATE` that additionally asserts `status`, `last_heartbeat_at`, `generation_id`, `session_id`, `fencing_epoch`, and `pid` are unchanged. A guard that matches no row preserves the worker with `row_changed_since_plan`. There is no window between a safety check and its matching write, so a worker that comes back to life, changes ownership, or is retired concurrently cannot be removed on the strength of a stale observation. Any exception — including a commit failure — rolls the whole transaction back and returns `transaction_failed` with `success: false`, `retired_count: 0`, and `mutation_performed: false`; a partial write can never be reported as success. ### Outcomes | `outcome` | Meaning | `mutation_performed` | | --- | --- | --- | | `planned` | Dry-run result | `false` | | `applied` | Transaction ran; see `retired` / `preserved` | `true` only if something was retired | | `registry_revision_moved` | Registry changed between plan and apply | `false` | | `candidate_set_moved` | Eligibility verdict changed between plan and apply | `false` | | `already_retired` | Every requested row is already retired (idempotent replay) | `false` | | `nothing_requested` | Empty target list | `false` | | `transaction_failed` | Rolled back; nothing retired | `false` | ## What retirement does to the fleet snapshot A retired row keeps its history: `status` moves to `retired` and `retired_at`, `retired_by`, `retirement_reason` are recorded. Nothing is deleted. Because `retired` is not `active`, the #978 snapshot counts the row as **historical**, not stale, so `stale_worker_count` falls and historical rows never make the live fleet unsafe by themselves. **Retirement does not repair untrusted live identity.** Live workers registered under legacy `pid-`/`proc-` instance identities are preserved untouched and keep their `legacy_incomplete_identity` blockers. Retiring every stale row can therefore legitimately produce: * `stale_worker_count: 0` * residual live `legacy_incomplete_identity` blockers * `live_fleet_safe: false` That is a truthful result, and the `post_apply` block reports the remaining blockers rather than claiming the fleet became safe. Trusted `client_instance_id` propagation through launchers is a separate enrollment problem. ## Permissions Plan and apply are authorized differently, and deliberately so (review 657 B1). | | Plan | Apply | | --- | --- | --- | | Capability | `gitea.read` | `gitea.worker_registry.retire` | | Nature | observational; opens no transaction, writes nothing | mutation | | Roles | `controller`, `reconciler` | `controller`, `reconciler` | An earlier revision authorized apply with `gitea.read` alone, reasoning that the mutation lands in the local control-plane registry rather than in Gitea. That the write is local makes it **no less a mutation**: sharing an observational permission class with plan meant any profile that could *look* could also *destroy*. Apply now requires its own capability. * **Denied:** author, reviewer, merger, and every ordinary read-only profile — they lack the capability, so they fail closed on the permission itself rather than on the role check alone. The role restriction remains as defence in depth: a profile mistakenly granted the capability still cannot reach apply from an author, reviewer, or merger role. * The capability is checked at entry **and** re-resolved immediately before the registry mutation, so a profile change mid-call cannot be outrun. * **No new Gitea write permission is introduced.** `gitea.worker_registry.retire` authorizes exactly one local control-plane transition (`worker_registrations.status -> retired`) and grants no branch, issue, PR, review, merge, or restart authority. No author permission is broadened. * The fleet snapshot remains observational: nothing here turns it into a gate on ordinary author work. ### Operator step No profile holds `gitea.worker_registry.retire` by default, so apply is inert until an operator adds it to the `allowed_operations` of the controller or reconciler profile in `profiles.json`. Removing it again immediately and completely revokes apply, while leaving plan and every other capability untouched. That grant is a configuration change and is outside the scope of the code that implements this feature. ## External-state fencing `BEGIN IMMEDIATE` locks the worker registry and nothing else, so two inputs the decision depends on sit outside the transaction's isolation domain: the workflow-lease table in a separate control-plane database, and OS process liveness. Re-reading them once during revalidation is not sufficient — the per-target loop runs afterwards, so a lease acquired (or a pid revived) after revalidation but before a given row's `UPDATE` would go unnoticed, and the registry-column guard cannot catch it because no registry column changed. Two mechanisms close that window, both applied per target immediately before its own write: * **`external_fence_fn`** — a version token over active leases (`external_state_fingerprint`), captured inside the transaction *before* the authoritative read and re-compared before every guarded `UPDATE`. Any movement raises, rolling back the whole transaction: once the world has changed, every remaining per-row decision was computed against a world that no longer exists. An unreadable lease store raises rather than returning a token, because "unreadable" must not silently compare equal to "unchanged". * **`liveness_fn`** — a re-probe of process liveness and fencing identity that must affirmatively re-establish that this exact process is gone. It compares `process_start_time`, so a pid number reused since the plan is refused rather than accepted. A caller that supplies no `liveness_fn` retires nothing (`liveness_reprobe_unavailable`) rather than proceeding unfenced. The guarded `UPDATE` additionally asserts `host_id`, `boot_id`, and `process_start_time` are unchanged, and all three participate in the CAS token, so fencing movement alone is enough to abort. ## Non-goals * Killing or restarting processes. * Editing session files or configuration. * Direct database cleanup outside the sanctioned transaction. * Retiring live workers. * Backfilling trusted identity for legacy workers. * Rewriting worker ownership. * Cleaning unrelated workflow-lease or issue-claim registries.