Files
Gitea-Tools/docs/architecture/mcp-ha-rolling-restart.md
T

11 KiB

ADR: High-availability and rolling-restart architecture for Gitea MCP control plane

  • Status: Proposed (Design ADR under #668)
  • Date: 2026-07-25
  • Tracking Issue: #668
  • Policy Version: mcp-ha-rolling-restart/v1
  • Related:
    • Parent: #655 — Governed MCP restart coordination and zero-disruption recovery
    • Governance Policy: #656 / docs/architecture/mcp-restart-governance.md
    • Control-Plane DB Substrate: #613 / docs/architecture/control-plane-db-substrate.md
    • Runtime Policy: #615 / docs/architecture/mcp-stable-control-runtime-policy-adr.md
    • Product Vision: #652 (Phase 5 Maturity)
    • Delivery Roadmap: #653

1. Context & Problem Statement

The Gitea MCP server operates as the authoritative control plane for managing issues, Pull Requests, code mutations, formal reviews, and workflow reconciliations. Under single-process governance (#656), process restarts are strictly controlled using pre-flight checks, drain phases, and operator approvals.

However, a single-instance control plane inherently presents fundamental constraints:

  1. Downtime during updates: Even a perfectly executed single-process drain requires a window where incoming client requests must be paused or rejected while the server binary or python environment reloads.
  2. Single point of failure: Infrastructure issues, process crashes, or unhandled host-level terminations immediately disconnect active LLM sessions and leave transient workflows incomplete.
  3. Multi-agent concurrency bottlenecks: High volumes of concurrent multi-LLM tasks put all lock management, lease allocation, and Gitea API interactions through a single process event loop.

To achieve true zero-disruption operation and seamless rolling deployments without stopping active work, the system requires a high-availability (HA), multi-instance MCP architecture.


2. Architectural Principles & Non-Goals

2.1 Core Architectural Principles

  • Gitea as Canonical Work SoT: Gitea remains the ultimate System of Record (SoT) for issue states, pull requests, labels, and audit comments. The MCP control plane does not duplicate domain entities.
  • Control-Plane DB as Multi-Instance State Substrate: The control-plane SQLite/durable database (#613) acts as the single source of truth for workflow leases, session tokens, assignment records, and lock fences across all MCP nodes.
  • Stateless Worker Nodes: MCP role server processes (gitea-author, gitea-reviewer, gitea-merger, gitea-reconciler, gitea-controller) maintain no unique in-memory state; any node can handle any request given a valid session resume token.
  • Fail-Closed Split-Brain Defense: In any network partition or quorum loss scenario, nodes must fail closed rather than risk double-mutations or conflicting Gitea states.

2.2 Non-Goals

  • Replacing Gitea: We do not replace Gitea issue/PR tracking with an independent database.
  • Immediate Multi-Node Cluster Execution in v1: This ADR defines the target architecture and phased roadmap; immediate implementation occurs incrementally post-[#655] v1.

3. High-Availability & Rolling-Restart Architecture

3.1 Architecture Overview

                        +----------------------------+
                        | LLM Clients / IDE Sessions |
                        +--------------+-------------+
                                       |
                                       v
                        +----------------------------+
                        |    HA Proxy / Router       |
                        | (Health-based & Affinity)  |
                        +------+--------------+------+
                               |              |
                +--------------+              +--------------+
                v                                            v
     +--------------------+                        +--------------------+
     | MCP Instance Node A|                        | MCP Instance Node B|
     | (Version N)        |                        | (Version N+1)      |
     +---------+----------+                        +---------+----------+
               |                                             |
               +----------------------+----------------------+
                                      |
                                      v
                        +----------------------------+
                        |  Control-Plane DB Substrate|
                        |  (Shared Lease & Locks)    |
                        +--------------+-------------+
                                       |
                                       v
                        +----------------------------+
                        |        Gitea API           |
                        +----------------------------+

3.2 Key System Components

A. Multiple MCP Instance Cohorts

  • The control plane runs across N \ge 2 redundant process nodes.
  • Dual-namespace deployment allows running the old version (Node A) alongside a updated version (Node B) during rolling upgrades.

B. Shared Durable Session Storage & Resume Tokens

  • Session context, preflight verification proofs, and capability resolution states are stored in the shared control-plane database.
  • Client requests carry an explicit session_id and resume_token. If an MCP instance restarts or a request routes to a different instance, the target node validates the token against the database without requiring full session re-initialization.

C. Shared Lease Authority & Fencing Counters

  • Workflow leases (gitea_allocate_next_work, gitea_adopt_workflow_lease) use monotonic fencing tokens (lease_generation_id).
  • When Node B acquires or renews a lease, it increments the generation counter. Any delayed or out-of-order write attempt from Node A using an older generation token is rejected by database constraints.

D. Leader Election & Coordinated Drain

  • Node clusters elect a primary coordinator node for administrative background tasks (such as stale lease cleanup or incident Watchdogs).
  • During a rolling deployment:
    1. Node B (new version) is launched and registers as healthy.
    2. Router directs new session creations to Node B.
    3. Node A enters MAINTENANCE_DRAIN status (#659), completing in-flight mutations while refusing new tasks.
    4. Once all active sessions migrate or complete, Node A shuts down cleanly.

E. Idempotent Mutations & Failover Safety

  • All state-changing tool executions (PR creation, review submission, merge operations, label changes) carry a deterministic idempotency_key.
  • If a network connection flaps or a node fails mid-mutation, the re-issued request with the same idempotency_key is recognized by the control-plane substrate, returning the existing recorded result without repeating side effects on Gitea.

F. Schema Version Compatibility

  • Database migrations follow non-breaking additive patterns.
  • During rolling upgrades where Node A (Version N) and Node B (Version N+1) run concurrently, both versions operate against the shared schema without structural conflicts.

4. Split-Brain & Failure Behavior

4.1 Split-Brain Risk Scenarios & Mitigation

Scenario Risk Mitigation Strategy
Network Partition between Nodes Both Node A and Node B attempt to process operations for the same issue/PR. Generation Fencing: Lease renewal requires updating the DB generation counter. The node isolated from the DB fails closed immediately.
Stale Node Recovery Node A recovers after a long pause and executes a queued mutation. Lease Expiry & TTL Fencing: Transactions verify that expires_at > NOW() within the atomic SQLite transaction boundaries.
Database Connection Loss Node loses access to shared control-plane DB substrate. Strict Fail-Closed: The node immediately marks all task capabilities as blocked and rejects mutation tools until DB connectivity is re-established.

5. Phased Implementation Milestones

flowchart TD
    M1[Milestone 1: Shared Control-Plane DB Schema & Resume Tokens] --> M2[Milestone 2: Idempotent Mutation Layer]
    M2 --> M3[Milestone 3: Health Routing & Standby Failover]
    M3 --> M4[Milestone 4: Active-Active Rolling Deployment & Auto-Drain]

Milestone 1: Shared Control-Plane DB Schema & Resume Tokens (Post-#655)

  • Extend [#613] Control-Plane DB schema to store multi-instance node heartbeat records and session resume tokens.
  • Enable session lookup across instances via session_id.

Milestone 2: Idempotent Mutation Layer & Lease Fencing

  • Add mandatory idempotency_key tracking to all Gitea mutation tools.
  • Implement monotonic lease fencing counters in gitea_allocate_next_work and gitea_adopt_workflow_lease.

Milestone 3: Health-Based Routing & Active-Passive Standby

  • Introduce lightweight proxy/router capable of checking node health endpoints.
  • Implement active-standby failover where standby node automatically assumes work if active node fails health checks.

Milestone 4: Active-Active Horizontal Deployment & Rolling Upgrade Automation

  • Enable true active-active multi-instance execution.
  • Integrate automated zero-downtime rolling upgrades coordinated with gitea_request_mcp_restart maintenance drain.

6. Observability & Audit Requirements

High-availability control plane operations must expose clear telemetry and audit trails:

  • Node Registry Telemetry: Active nodes, version numbers, uptime, and heartbeat timestamps reported via gitea_get_runtime_context.
  • Lease Fencing Metrics: Tracking lease acquire latency, fence rejection counts, and lease handoff durations.
  • Failover & Re-route Audit Logs: Durable logging of session migrations between nodes, drain initiation, and process retirement events.

7. Tradeoffs & Accepted Risks

  • Increased Architectural Complexity: Moving from a single process to a multi-instance control plane requires robust DB locking, proxy routing, and migration governance.
  • Database Dependency: The control-plane database substrate becomes a critical shared dependency for multi-node deployments. High availability for the underlying SQLite file system / DB must be guaranteed.