11 KiB
11 KiB
ADR: High-availability and rolling-restart architecture for Gitea MCP control plane
- Status: Proposed (Design ADR under #668)
- Date: 2026-07-25
- Tracking Issue: #668
- Policy Version:
mcp-ha-rolling-restart/v1 - Related:
- Parent: #655 — Governed MCP restart coordination and zero-disruption recovery
- Governance Policy: #656 /
docs/architecture/mcp-restart-governance.md - Control-Plane DB Substrate: #613 /
docs/architecture/control-plane-db-substrate.md - Runtime Policy: #615 /
docs/architecture/mcp-stable-control-runtime-policy-adr.md - Product Vision: #652 (Phase 5 Maturity)
- Delivery Roadmap: #653
1. Context & Problem Statement
The Gitea MCP server operates as the authoritative control plane for managing issues, Pull Requests, code mutations, formal reviews, and workflow reconciliations. Under single-process governance (#656), process restarts are strictly controlled using pre-flight checks, drain phases, and operator approvals.
However, a single-instance control plane inherently presents fundamental constraints:
- Downtime during updates: Even a perfectly executed single-process drain requires a window where incoming client requests must be paused or rejected while the server binary or python environment reloads.
- Single point of failure: Infrastructure issues, process crashes, or unhandled host-level terminations immediately disconnect active LLM sessions and leave transient workflows incomplete.
- Multi-agent concurrency bottlenecks: High volumes of concurrent multi-LLM tasks put all lock management, lease allocation, and Gitea API interactions through a single process event loop.
To achieve true zero-disruption operation and seamless rolling deployments without stopping active work, the system requires a high-availability (HA), multi-instance MCP architecture.
2. Architectural Principles & Non-Goals
2.1 Core Architectural Principles
- Gitea as Canonical Work SoT: Gitea remains the ultimate System of Record (SoT) for issue states, pull requests, labels, and audit comments. The MCP control plane does not duplicate domain entities.
- Control-Plane DB as Multi-Instance State Substrate: The control-plane SQLite/durable database (#613) acts as the single source of truth for workflow leases, session tokens, assignment records, and lock fences across all MCP nodes.
- Stateless Worker Nodes: MCP role server processes (
gitea-author,gitea-reviewer,gitea-merger,gitea-reconciler,gitea-controller) maintain no unique in-memory state; any node can handle any request given a valid session resume token. - Fail-Closed Split-Brain Defense: In any network partition or quorum loss scenario, nodes must fail closed rather than risk double-mutations or conflicting Gitea states.
2.2 Non-Goals
- Replacing Gitea: We do not replace Gitea issue/PR tracking with an independent database.
- Immediate Multi-Node Cluster Execution in v1: This ADR defines the target architecture and phased roadmap; immediate implementation occurs incrementally post-[#655] v1.
3. High-Availability & Rolling-Restart Architecture
3.1 Architecture Overview
+----------------------------+
| LLM Clients / IDE Sessions |
+--------------+-------------+
|
v
+----------------------------+
| HA Proxy / Router |
| (Health-based & Affinity) |
+------+--------------+------+
| |
+--------------+ +--------------+
v v
+--------------------+ +--------------------+
| MCP Instance Node A| | MCP Instance Node B|
| (Version N) | | (Version N+1) |
+---------+----------+ +---------+----------+
| |
+----------------------+----------------------+
|
v
+----------------------------+
| Control-Plane DB Substrate|
| (Shared Lease & Locks) |
+--------------+-------------+
|
v
+----------------------------+
| Gitea API |
+----------------------------+
3.2 Key System Components
A. Multiple MCP Instance Cohorts
- The control plane runs across
N \ge 2redundant process nodes. - Dual-namespace deployment allows running the old version (Node A) alongside a updated version (Node B) during rolling upgrades.
B. Shared Durable Session Storage & Resume Tokens
- Session context, preflight verification proofs, and capability resolution states are stored in the shared control-plane database.
- Client requests carry an explicit
session_idandresume_token. If an MCP instance restarts or a request routes to a different instance, the target node validates the token against the database without requiring full session re-initialization.
C. Shared Lease Authority & Fencing Counters
- Workflow leases (
gitea_allocate_next_work,gitea_adopt_workflow_lease) use monotonic fencing tokens (lease_generation_id). - When Node B acquires or renews a lease, it increments the generation counter. Any delayed or out-of-order write attempt from Node A using an older generation token is rejected by database constraints.
D. Leader Election & Coordinated Drain
- Node clusters elect a primary coordinator node for administrative background tasks (such as stale lease cleanup or incident Watchdogs).
- During a rolling deployment:
- Node B (new version) is launched and registers as healthy.
- Router directs new session creations to Node B.
- Node A enters
MAINTENANCE_DRAINstatus (#659), completing in-flight mutations while refusing new tasks. - Once all active sessions migrate or complete, Node A shuts down cleanly.
E. Idempotent Mutations & Failover Safety
- All state-changing tool executions (PR creation, review submission, merge operations, label changes) carry a deterministic
idempotency_key. - If a network connection flaps or a node fails mid-mutation, the re-issued request with the same
idempotency_keyis recognized by the control-plane substrate, returning the existing recorded result without repeating side effects on Gitea.
F. Schema Version Compatibility
- Database migrations follow non-breaking additive patterns.
- During rolling upgrades where Node A (Version
N) and Node B (VersionN+1) run concurrently, both versions operate against the shared schema without structural conflicts.
4. Split-Brain & Failure Behavior
4.1 Split-Brain Risk Scenarios & Mitigation
| Scenario | Risk | Mitigation Strategy |
|---|---|---|
| Network Partition between Nodes | Both Node A and Node B attempt to process operations for the same issue/PR. | Generation Fencing: Lease renewal requires updating the DB generation counter. The node isolated from the DB fails closed immediately. |
| Stale Node Recovery | Node A recovers after a long pause and executes a queued mutation. | Lease Expiry & TTL Fencing: Transactions verify that expires_at > NOW() within the atomic SQLite transaction boundaries. |
| Database Connection Loss | Node loses access to shared control-plane DB substrate. | Strict Fail-Closed: The node immediately marks all task capabilities as blocked and rejects mutation tools until DB connectivity is re-established. |
5. Phased Implementation Milestones
flowchart TD
M1[Milestone 1: Shared Control-Plane DB Schema & Resume Tokens] --> M2[Milestone 2: Idempotent Mutation Layer]
M2 --> M3[Milestone 3: Health Routing & Standby Failover]
M3 --> M4[Milestone 4: Active-Active Rolling Deployment & Auto-Drain]
Milestone 1: Shared Control-Plane DB Schema & Resume Tokens (Post-#655)
- Extend [#613] Control-Plane DB schema to store multi-instance node heartbeat records and session resume tokens.
- Enable session lookup across instances via
session_id.
Milestone 2: Idempotent Mutation Layer & Lease Fencing
- Add mandatory
idempotency_keytracking to all Gitea mutation tools. - Implement monotonic lease fencing counters in
gitea_allocate_next_workandgitea_adopt_workflow_lease.
Milestone 3: Health-Based Routing & Active-Passive Standby
- Introduce lightweight proxy/router capable of checking node health endpoints.
- Implement active-standby failover where standby node automatically assumes work if active node fails health checks.
Milestone 4: Active-Active Horizontal Deployment & Rolling Upgrade Automation
- Enable true active-active multi-instance execution.
- Integrate automated zero-downtime rolling upgrades coordinated with
gitea_request_mcp_restartmaintenance drain.
6. Observability & Audit Requirements
High-availability control plane operations must expose clear telemetry and audit trails:
- Node Registry Telemetry: Active nodes, version numbers, uptime, and heartbeat timestamps reported via
gitea_get_runtime_context. - Lease Fencing Metrics: Tracking lease acquire latency, fence rejection counts, and lease handoff durations.
- Failover & Re-route Audit Logs: Durable logging of session migrations between nodes, drain initiation, and process retirement events.
7. Tradeoffs & Accepted Risks
- Increased Architectural Complexity: Moving from a single process to a multi-instance control plane requires robust DB locking, proxy routing, and migration governance.
- Database Dependency: The control-plane database substrate becomes a critical shared dependency for multi-node deployments. High availability for the underlying SQLite file system / DB must be guaranteed.