Files
Gitea-Tools/docs/instance-fleet-identity.md
T
jcwalker3andClaude Opus 4.8 7e19079b5c feat(launcher): trusted client-instance identity for project-scoped launches
Issue #985. The production launcher could mint one trusted
GITEA_MCP_CLIENT_INSTANCE per launch, but two gaps kept real launches on
untrusted legacy-pid-* identities:

1. build_application_mcp_servers required a profile for all five sanctioned
   namespaces, so a project-scoped configuration exposing only author,
   reviewer, and merger could not use it without inventing controller and
   reconciler profiles that must not exist.
2. The module produced configuration data but had no runnable entry point, so
   every real launch bypassed it entirely.

Changes:

- resolve_launch_namespaces() validates an explicit namespace subset as an
  allow-list; unknown, duplicate, and empty selections are refused rather than
  silently narrowing a launch. Omitting it preserves five-namespace behaviour.
- build_application_mcp_servers() accepts that subset, requires profiles only
  for the launched namespaces, starts only those workers, and reports
  excluded_namespaces / project_scoped.
- collect_instance_ids_from_mcp_servers() inspects the launch's own namespaces
  instead of an assumed five, and reports missing_servers, so a three-namespace
  launch can prove shared attribution without reading as two absent workers.
- Runnable entry point: python3 -m mcp_application_launcher mints one trusted
  identity, writes a per-launch 0600 mcpServers config, and execs the client.
  CLIENT_LAUNCH_SPECS is a data-driven registry so other supported clients use
  the same mint-once/propagate-to-all mechanism. --dry-run prints the plan.
- Provenance sealing now fails closed. The inst- format is public and
  reproducible, so format alone could previously let anyone who set one
  environment variable manufacture a trusted identity. Trust now additionally
  requires GITEA_MCP_INSTANCE_PROVENANCE=trusted_launcher, which only the
  launcher writes; a well-formed but unsealed value is classified
  unsealed_launcher and refused, while still being reported for diagnosis.

Deliberate behaviour change: tests/test_issue_978_instance_fleet_snapshot.py
test_client_hints_trusted_when_set previously asserted that a well-formed ID
alone was trusted. It now supplies the launcher seal, and a new companion test
asserts the unsealed case fails closed. This tightens the contract; no
assertion was weakened.

No static or persistent per-project instance IDs are introduced, duplicate
worker and cohort detection are untouched, and no fleet or mutation gate is
relaxed.

Tests: tests/test_issue_985_project_scoped_launcher.py, 47 passed, 3 subtests.
Full suite from a branches/ worktree: 28 failed, 6252 passed, 6 skipped against
a master baseline at 32ab8392 of 28 failed, 6204 passed, 6 skipped; the failing
sets are byte-identical, so zero regressions and zero masked failures.

Closes #985

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-07-31 03:15:53 -05:00

13 KiB

Instance-level fleet identity and health snapshots

Issue #978. Companion primitives: #948 (worker ownership), #975 (heartbeat lifecycle), #951 (restart receipts).

Why this exists

The multi-instance fleet gate (#963) needs a native answer to:

  • which application launches are live;
  • which of the five namespace workers belong to each launch;
  • whether two Codex (or Claude, Grok, …) launches are distinct;
  • whether any live collision is real rather than a shared client type.

Process tables, PID proximity, configuration files, and direct SQLite access are not production evidence. The sanctioned surface is the read-only MCP tool gitea_snapshot_instance_fleet on controller and reconciler namespaces.

Identity hierarchy

Identity Scope Who generates it Lifetime
client_type Application family (codex, claude_code, gemini, grok, …) Launcher sets GITEA_MCP_CLIENT Stable for the product
client_instance_id One running application launch Trusted launcher, once per launch, as GITEA_MCP_CLIENT_INSTANCE Fresh launch → new ID; reconnect of same launch → same ID; full app restart → new ID
fleet_run_id Operator-approved enrollment / canary cohort Operator / controller sets GITEA_MCP_FLEET_RUN_ID Duration of the approved rollout
namespace author | reviewer | merger | controller | reconciler Profile / MCP server binding Process lifetime
worker_id / worker_identity One namespace worker process Runtime registry at first registration New on worker restart; not reused
session_id Client session ownership Launcher GITEA_MCP_CLIENT_SESSION or runtime Session lifetime
generation_id One daemon launch Runtime at process boot New on process restart
process_identity / PID OS process Runtime Process lifetime

Rules

  1. Multiple active instances may share the same client_type.
  2. Every application launch receives a distinct client_instance_id.
  3. All namespace workers of one launch report the same client_instance_id. For a whole-fleet launch that is five workers; for a project-scoped launch it is exactly the namespaces that launch started (#985).
  4. Each namespace worker has a distinct worker_identity, process identity, generation, and PID.
  5. Instance identity is never inferred from PID proximity, timestamps, or client type alone.
  6. Live reuse of one client_instance_id with conflicting workers fails closed.
  7. Sharing only a profile or client_type is not a duplicate.

This deliberately replaces any permanent exactly_one_per_profile fleet model (#949 assumption) as the operating rule for multi-instance fleets.

Project-scoped launches and the runnable entry point (#985)

Not every project exposes all five namespaces. A project-scoped configuration such as Weekly Briefings intentionally has author, reviewer, and merger only, and has no controller or reconciler profile to supply. Requiring all five would force operators to invent profiles that must not exist, so the launcher accepts an explicitly validated subset.

Running a launch

python3 -m mcp_application_launcher \
    --client claude_code \
    --namespaces author,reviewer,merger \
    --profile author=prgs-author \
    --profile reviewer=prgs-reviewer \
    --profile merger=prgs-merger

That mints one trusted client_instance_id, writes a per-launch mcpServers config, and executes:

claude --mcp-config <per-launch.json> --strict-mcp-config

Add --dry-run to print the plan (namespaces, excluded namespaces, minted ID, argv, config path) as JSON and start nothing. Omit --namespaces to launch all five exactly as before. Arguments after the flags are forwarded to the client.

Other supported clients register one entry in mcp_application_launcher.CLIENT_LAUNCH_SPECS; because the argv builder only ever receives a config this module wrote, every client necessarily goes through the same mint-once/propagate-to-all mechanism.

Why the config is per-launch and never .mcp.json

Persisting a trusted ID into a shared, reused .mcp.json would give two concurrent sessions the same trusted identity — precisely the reuse case the duplicate gate must reject. The launcher therefore writes a fresh owner-readable-only (0600) config per launch. Do not commit one, and do not copy a minted GITEA_MCP_CLIENT_INSTANCE into any checked-in configuration.

Provenance sealing

The inst-… format is public and reproducible, so format alone cannot establish trust: anyone able to set one environment variable could otherwise hand-write a valid-looking ID and be believed. Trust therefore requires both the format and GITEA_MCP_INSTANCE_PROVENANCE=trusted_launcher, which only the launcher writes. A well-formed but unsealed identity is classified unsealed_launcher and fails closed — it is still reported for diagnosis, but never authorizes trusted attribution. Manually asserted trust is not possible.

Migration and coordinated relaunch

Existing configurations that predate this work set no instance key at all, so their workers register under legacy-pid-* (legacy_incomplete) and remain mutually indistinguishable when two launches share a profile.

Migrating is a coordinated relaunch, not an in-place edit — a running worker cannot acquire an identity it was not started with:

  1. Stop every LLM client currently running Gitea MCP workers. A single surviving legacy cohort keeps the fleet ambiguous.
  2. Relaunch each application through the command above.
  3. Repeat per application. Concurrent launches are expected and safe: each receives its own trusted ID.

Post-launch verification

  • gitea_get_runtime_contextprovenance_assessment.attachment should show an inst-… client_instance_id with instance_id_provenance: trusted_launcher, not legacy-pid-* / legacy_incomplete.
  • Every namespace of the same launch must report that same ID; two different launches must report different IDs.
  • gitea_resolve_task_capability should return exact_safe_next_action: "None; ready for operations." with no blocker_kind. A runtime_reconnect_required naming duplicate PIDs per profile means at least one cohort is still on legacy identity, or a second cohort is genuinely running.
  • mcp_application_launcher.collect_instance_ids_from_mcp_servers(servers, namespaces=[...]) returns shared_single_trusted_id for a built config, and reports missing_servers when an expected namespace entry is absent.

How five workers join one instance

  1. The host starts one application instance (for example one Codex session).

  2. The production application launcher (mcp_application_launcher.build_application_mcp_servers / gitea_config.multi_namespace_launcher_entries) mints exactly one client_instance_id via mcp_fleet_snapshot.generate_client_instance_id and injects the same env into every namespace worker:

    GITEA_MCP_CLIENT=<client_type>
    GITEA_MCP_CLIENT_INSTANCE=<client_instance_id>
    GITEA_MCP_INSTANCE_PROVENANCE=trusted_launcher
    GITEA_MCP_CLIENT_SESSION=<session_id>          # optional but recommended
    GITEA_MCP_FLEET_RUN_ID=<enrollment id>         # when on an approved canary
    GITEA_CLIENT_MANAGED=1
    GITEA_MCP_PROFILE=<role-profile>
    
  3. The MCP client starts the five namespace processes (gitea-author, gitea-reviewer, gitea-merger, gitea-controller, gitea-reconciler) from that generated mcpServers map — each entry carries the same instance ID.

  4. Each worker registers once into the worker registry with its own worker_identity, namespace, generation_id, and PID, but the shared client_instance_id. Workers never mint a trusted instance ID themselves.

Only IDs matching the launcher format inst-<client>-<timestamp>-<digest> are trusted. Missing, legacy-pid-* / pid-* placeholders, and ordinary user-supplied strings fail closed as untrusted and cannot authorize multi-instance fleet mutation safety.

Worker reconnect (same process, same registration) keeps the instance ID. Worker restart (new process) mints a new worker identity and generation but must still receive the same GITEA_MCP_CLIENT_INSTANCE from the parent application if it is the same launch. Full application restart mints a new client_instance_id (call the launcher without a prior ID).

Mutation gate (instance-aware)

The capability / runtime diagnostic gate (_check_mcp_runtimes_diagnostics) is instance-aware:

  • Two legitimate application instances that share a profile (same GITEA_MCP_PROFILE) are allowed when each has a distinct trusted client_instance_id.
  • More than one live worker for the same (client_instance_id, profile/namespace) fails closed as a duplicate namespace worker.
  • Multiple processes sharing a profile without trusted instance evidence still fail closed (indistinguishable from a duplicate).

Approved fleet manifest

An operator (or controller enrollment step) obtains an approved manifest as a list of expected instances, for example:

{
  "instances": [
    {
      "client_type": "codex",
      "client_instance_id": "inst-codex-20260730T120000Z-abc123def456",
      "fleet_run_id": "canary-963-2026-07-30",
      "namespaces": ["author", "reviewer", "merger", "controller", "reconciler"]
    },
    {
      "client_type": "codex",
      "client_instance_id": "inst-codex-20260730T120100Z-fed654cba321",
      "fleet_run_id": "canary-963-2026-07-30",
      "namespaces": ["author", "reviewer", "merger", "controller", "reconciler"]
    }
  ]
}

Pass that list as expected_manifest to gitea_snapshot_instance_fleet. When a manifest is supplied:

  • instances on the manifest but not live → missing_expected;
  • live instances not on the manifest → unmanifested;
  • both are active blockers for fleet-gate safety.

Without a manifest, the snapshot still enumerates the live fleet and classifies identity collisions; it does not invent enrollment policy.

Snapshot consistency

Each snapshot includes:

  • snapshot_at — UTC timestamp;
  • consistency_token / registry_revision — content digest over worker identity, instance id, status, heartbeat, fencing, and generation.

Two successive snapshots can prove heartbeat continuity via mcp_fleet_snapshot.compare_snapshot_heartbeats.

Classification (active vs historical)

Active blockers (make live_fleet_safe=false):

  • missing expected instance;
  • unmanifested extra instance;
  • duplicate namespace worker within one instance;
  • live client_instance_id collision;
  • reused worker / session / generation / process / PID / fencing identity;
  • orphaned or unowned workers;
  • unknown client;
  • foreign-repository workers;
  • old-revision workers;
  • stale workers still marked active;
  • legacy incomplete instance identity.

Historical dead rows (status released/superseded) are reported as historical_dead findings with active_blocker=false. They never automatically make the live fleet unsafe.

Fail-closed behaviour and recovery

Condition Mutation safety Diagnostic reads Recovery
Two instances, same client_type, distinct IDs Safe (if otherwise healthy) Available None needed
Live reuse of one client_instance_id Unsafe Available Stop the colliding launch or re-issue a distinct ID
Two workers, same namespace, one instance Unsafe Available Stop the extra worker
Legacy registration without trusted instance ID Unsafe for fleet mutation Available Relaunch with launcher-issued GITEA_MCP_CLIENT_INSTANCE
Historical dead row only Does not block alone Available No action required for fleet safety
Registry unavailable Snapshot fails closed N/A Repair registry path / reconnect namespaces

Incomplete or untrusted instance identity cannot authorize unsafe mutation. Diagnostic reads remain available where gitea.read allows.

Permissions

  • Allowed: controller, reconciler with gitea.read.
  • Denied: author, reviewer, merger (even with gitea.read for other tools).
  • No new mutation permissions are granted to any role.
  • mcp_worker_identity — registry, worker identity, heartbeats (#948, #975).
  • mcp_fleet_snapshot — pure snapshot + classification (#978).
  • gitea_snapshot_instance_fleet — sanctioned MCP tool (#978).
  • gitea_get_runtime_context — single-process view (not fleet-wide).

Non-goals

  • Starting the #963 canary.
  • Purging historical registry rows.
  • Preserving “exactly one process per profile” as the permanent model.
  • Using shell / process-table / SQLite inspection as production fleet evidence.