Add gitea_request_mcp_reconnect as a report-only callable surface so Codex and other agent hosts can request host/IDE reconnect with typed blockers and exact UI steps. Never kills processes or edits config. Co-Authored-By: Grok 4.5 <[email protected]>
175 lines
8.6 KiB
Markdown
175 lines
8.6 KiB
Markdown
# Recovering from `client is closing: EOF` on a Gitea MCP namespace (#543)
|
|
|
|
## Symptom
|
|
|
|
A tool call through a Gitea MCP namespace — `gitea-author`, `gitea-reviewer`,
|
|
`gitea-merger`, or the shared `gitea-tools` namespace — fails immediately with:
|
|
|
|
```
|
|
client is closing: EOF
|
|
```
|
|
|
|
Every subsequent call to that same namespace returns the same error, including
|
|
cheap read tools such as `gitea_whoami` and `gitea_list_profiles`. Other MCP
|
|
servers registered with the same client (for example `context7`) keep working,
|
|
so this is **not** a global MCP-client outage.
|
|
|
|
## Why this is not a code defect
|
|
|
|
This failure is a **transport-level** condition in the IDE / MCP client manager,
|
|
not a missing or broken tool:
|
|
|
|
- The tool can be present and registered in the Python `FastMCP` tool manager.
|
|
- Direct Python inspection of the server confirms the tool exists.
|
|
- Running the server manually and sending JSON-RPC over stdio works fine
|
|
(offline spawn) — that path does **not** prove the IDE namespace is healthy.
|
|
|
|
The client manager entered a closed state after the backing subprocess for that
|
|
namespace terminated (or was killed) behind its back. Once closed, the client
|
|
does **not** re-spawn the child on the next tool call — it just replays
|
|
`client is closing: EOF`. The OS process may even still be alive if a parent
|
|
language-server process is holding the stdio pipes open.
|
|
|
|
This is the canonical "registered in FastMCP ≠ callable through the namespace"
|
|
false-ready state. It is distinct from the **stale-runtime** family in #531 /
|
|
#544, where the process is reachable but running behind `master`; that case is
|
|
detected by the `ps`-based `_check_mcp_runtimes_diagnostics` in
|
|
`gitea_mcp_server.py`. The EOF case is a dead/closed transport, not a stale one,
|
|
so the `ps` check alone will not surface it.
|
|
|
|
## Recovery path (canonical — client reconnect only)
|
|
|
|
Do the steps in order. Stop as soon as a live **client-namespace** call succeeds.
|
|
|
|
1. **Confirm the blast radius.** Call a cheap read tool on the failing namespace
|
|
(`gitea_whoami` or `gitea_list_profiles`). Then call the same tool on a
|
|
different MCP server (e.g. `context7`).
|
|
- Only the Gitea namespace fails → single-namespace transport close. Continue.
|
|
- Every server fails → restart the whole MCP client, not just one namespace.
|
|
|
|
2. **Request the sanctioned reconnect surface (#678), then reconnect through
|
|
the client — not the shell.** From a still-reachable Gitea MCP namespace
|
|
(or after host auto-reconnect), call:
|
|
|
|
```text
|
|
gitea_request_mcp_reconnect(
|
|
namespace="gitea-author", # or gitea-reviewer / gitea-merger / …
|
|
reason="transport_eof",
|
|
client="codex", # or claude_code / generic
|
|
)
|
|
```
|
|
|
|
The tool is **report-only**: it never restarts a process. It returns
|
|
namespace, profile, pid/session, startup SHA, current master SHA, boundary
|
|
status, and a **typed blocker** with exact operator UI steps for Codex
|
|
(Reload Developer Tools / per-server reconnect) or Claude Code (`/mcp`).
|
|
Then perform the host reconnect those steps describe so the client spawns a
|
|
fresh subprocess and re-opens the pipe. That clears the closed-client state
|
|
that a bare `kill`/respawn from a terminal does **not**.
|
|
|
|
3. **Do not "fix" it by importing the server or poking the process.** Reaching
|
|
for `python -c 'import gitea_mcp_server ...'`, raw JSON-RPC from a shell,
|
|
killing PIDs to force a respawn, or touching MCP config mtimes does **not**
|
|
restore the *client's* view of the namespace and violates the daemon-import
|
|
guard (#558, `docs/mcp-daemon-import-guard.md`). The only sanctioned repair
|
|
is a **client reconnect / relaunch** (or the typed operator path returned by
|
|
`gitea_request_mcp_reconnect`).
|
|
|
|
4. **Verify through the same path the workflow will use.** After reconnect, call
|
|
the specific tool the blocked workflow needs — not just any tool — through
|
|
the target namespace. For a merge that means calling the merger-authorized
|
|
adoption/merge tool through `gitea-merger`. A green `gitea_whoami` on one
|
|
namespace does **not** prove another namespace or another tool is callable.
|
|
Record success with:
|
|
|
|
```text
|
|
gitea_assess_mcp_namespace_health(..., probe_source="client_namespace")
|
|
```
|
|
|
|
5. **If reconnect does not clear it,** relaunch the client entirely, then repeat
|
|
step 4. If EOF persists after a full relaunch, the backing subprocess is
|
|
failing to start — inspect its stderr / launch config (command path, venv,
|
|
`*_MCP_CONFIG`, `*_MCP_PROFILE` env) rather than retrying the call. Still
|
|
do not use PID kill or config-touch as the primary recovery.
|
|
|
|
## Diagnostics to capture when reporting EOF
|
|
|
|
Include all of these so the failure is actionable and reproducible:
|
|
|
|
- **Namespace name** that returned EOF (`gitea-author` / `gitea-reviewer` /
|
|
`gitea-merger` / `gitea-tools`).
|
|
- **Tool** that was called and the **exact** error string.
|
|
- **PID** of the backing process (if any) and whether it was still alive
|
|
(informational only — not a recovery action).
|
|
- **Profile / env** for that namespace (execution profile, `*_MCP_PROFILE`,
|
|
worktree binding such as `GITEA_AUTHOR_WORKTREE`).
|
|
- **Config path** the client launched the server from.
|
|
- Result of the **cross-server control** call (did `context7` succeed?).
|
|
|
|
## Offline spawn probe (non-authoritative)
|
|
|
|
`test_mcp_conn.py` performs a full JSON-RPC handshake against a **fresh
|
|
subprocess** (`initialize` → `initialized` → `tools/list` → `tools/call`) and
|
|
classifies with `probe_source=offline_spawn`. That is useful for offline
|
|
launch/registration debugging. It is **not** proof the IDE-managed namespace is
|
|
healthy. See `docs/mcp-namespace-health.md`.
|
|
|
|
## Do-not list during EOF recovery
|
|
|
|
- Do **not** retry a blocked merge/adoption until the required tool is confirmed
|
|
callable through the merger-authorized **client** namespace (see #543).
|
|
- Do **not** clean, reset, or rebind a **foreign** worktree to work around the
|
|
error.
|
|
- Do **not** bypass the namespace with direct imports, raw API/curl, or
|
|
in-memory state restoration.
|
|
- Do **not** kill MCP PIDs or touch config mtimes as a substitute for client
|
|
reconnect.
|
|
|
|
## Sanctioned recovery vs forbidden process manipulation (#630)
|
|
|
|
Both restore a working namespace. Only one leaves the session trustworthy.
|
|
|
|
**Sanctioned — the runtime is repaired by whoever owns it:**
|
|
|
|
- IDE/host auto-reconnect, or an explicit client reconnect (`/mcp reconnect`).
|
|
- Relaunching the IDE/client so it respawns the daemons it started.
|
|
- An operator-owned restart performed outside the workflow session.
|
|
|
|
**Forbidden — the session manipulates the processes its own proof depends on:**
|
|
|
|
- `pkill -f mcp_server.py`, `pkill -f gitea_mcp_server`, broad `pkill -f mcp`.
|
|
- `killall` of a daemon, or `kill <pid>` of an MCP daemon pid.
|
|
- Any pattern broad enough to take unrelated namespaces with it
|
|
(`pkill -f python`), even when it never names MCP.
|
|
|
|
Read-only inspection (`ps aux | grep mcp_server`) is neither: it proves nothing
|
|
and breaks nothing. A `kill` of some unrelated pid is reported as *ambiguous*
|
|
rather than contaminating, so ordinary subprocess work is never false-blocked.
|
|
|
|
**What happens on a detected attempt.** `gitea_record_daemon_process_kill_attempt`
|
|
classifies a proposed command and, when it is a manual daemon kill, writes a
|
|
durable contamination marker for the active profile identity. While that marker
|
|
is live every review / merge / close / completion mutation fails closed;
|
|
`comment_issue` and `lock_issue` stay allowed so the contaminated worker can
|
|
still post its audit comment and hand off. The final report must surface the
|
|
contaminated recovery and must not claim a clean session.
|
|
|
|
Contamination is **not self-clearable**. Only
|
|
`gitea_audit_runtime_recovery_contamination` with `action=clear`, run under a
|
|
reconciler profile, removes it. The marker is recovery-critical, so it does not
|
|
expire into cleanliness when the session-state TTL lapses.
|
|
|
|
**Operator-authorized host maintenance stays permitted.** Authorization is read
|
|
from the `GITEA_OPERATOR_DAEMON_MAINTENANCE_AUTHORIZATION` environment variable
|
|
and from nowhere else — set outside the session by the operator who owns the
|
|
host, and recorded as an audit reference on the assessment. It is deliberately
|
|
not a tool argument: a session must never be able to authorize itself.
|
|
|
|
## Related
|
|
|
|
- #630 — manual daemon killing as contaminated recovery (this contrast, enforced).
|
|
- #531 / #544 — stale-runtime detection (`ps`-based); sibling failure mode.
|
|
- #558 / `docs/mcp-daemon-import-guard.md` — why shell imports are not a repair.
|
|
- `docs/mcp-client-registration.md` — per-server registration contract.
|
|
- `docs/mcp-namespace-health.md` — probe sources and mutation enforcement.
|