Supervise MCP daemon cohort lifecycle: one eligible cohort per profile, drain and reap superseded cohorts #900
Open
opened 2026-07-24 21:58:07 -05:00 by jcwalker3
·
0 comments
No Branch/Tag Specified
master
feat/issue-644-console-recovery
feat/issue-643-request-preview-initiate
fix/issue-897-permission-stale-runtime-classification
feat/issue-641-runtime-session-view
feat/issue-663-restart-classes
feat/issue-661-drain-proof-hard-gate
fix/issue-854-semantic-container-exclusion
issue-640
fix/issue-682-starlette-httpx2
v1.1.0
Labels
Clear labels
allocator
anti-stomp
architecture
bug
chore
codex
concurrency
contamination
control-plane
dashboard
database
design
documentation
enhancement
gitea
glitchtip
important
incident
incident-bridge
integration
jenkins
labels
leases
mcp
mcp-health
mcp-menu
multi-project
mutating
nice-to-have
observability
portability
preflight
protected-branch
queue
read-only
reconnect
recovery
refactor
release
reliability
resumable-review
reviewer
roadmap
safety
security
self-hosted
sentry
stale-runtime
status:blocked
status:in-progress
status:pr-open
status:ready
terminal-lock
testing
tracker
type:bug
type:feature
type:feature
type:guardrail
visibility
workflow
workflow-hardening
workflow-hardening
Controller-owned work allocator
Prevent concurrent LLM session stomping
Architecture / structural design
OpenAI Codex client / workflow session surface
Concurrent session safety
Workflow or session contamination incident
MCP control-plane coordination and allocation authority
MCP operational dashboard/queue view
Internal coordination storage (SQLite/Postgres)
Design / investigation, no implementation
Docs / runbooks
New feature or improvement
Gitea MCP workflow
GlitchTip integration
Operational or process incident requiring durable audit trail
Sentry-to-Gitea incident bridging
Integration testing
Jenkins integration
Label taxonomy management
Lease adopt/release/expire lifecycle
MCP server / tooling
MCP namespace and runtime health
MCP menu surface
Work spanning multiple monitoring projects or Gitea repos
Mutating action; requires gating
Observability, metrics, traces, error reporting
Cross-platform / portability
Shared preflight gates before mutation
Protected branch / stable-branch policy concern
Work queue visibility and allocation
Read-only, no mutation
MCP client reconnect/reload recovery path
Recovery paths for stale/foreign leases
Code refactor / restructure
Release / versioning
Reliability / failure handling
Persist and resume prepared review verdicts across sessions
Reviewer workflow tooling
Roadmap / umbrella issue
Safety rails and fail-closed mutation guards
Security / trust boundary
Self-hosted infrastructure integration
Sentry error monitoring integration
Stale backend daemon / runtime-vs-master parity failures
Issue is blocked
Issue is being worked on
Issue has an open pull request
Issue is ready for work
Terminal review lock (#332) path
Tests / test coverage
Issue tracker hygiene / meta
Bug or defect
Feature or enhancement
Feature or enhancement
Safety gate or guardrail
Workflow state visibility for LLMs/operators
Cross-tool workflow
LLM workflow coordination hardening
LLM workflow coordination hardening
Milestone
No items
No Milestone
Projects
Clear projects
No projects
Notifications
Due Date
No due date set.
Dependencies
No dependencies set.
Reference: Scaled-Tech-Consulting/Gitea-Tools#900
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Related: epic #887, #686, #655, #659, #669, #678
Problem
Nothing owns the lifetime of an MCP daemon cohort. Reconnects create new cohorts, superseded cohorts are never drained or reaped, and the surviving population is only ever inspected — never bounded.
Read-only process inventory on this host found roughly forty live
mcp_server.pyprocesses in at least five distinct cohorts, spawned across four separate days and still resident: cohorts atWed06AM,Thu03AM,5:45AM,10:32PM, and10:33PM. Every one of them holds production credentials and reads and writes the same session, lease, lock, and decision stores. A comparable inventory was already recorded on #686 during the 2026-07-12 incident, which found approximately fifty-six such processes. The population has never been bounded in the intervening time, which is the defect: the accumulation is structural, not incidental.The existing issues address neighbouring concerns and stop short of this one. #686 detects manually launched duplicates by provenance and fails closed on them. #659 drains work before a restart. #669 prefers narrow recovery over full reset. None of them establishes how many cohorts may exist, retires the ones that have been superseded, or bounds growth across repeated reconnects.
Coverage gap this issue closes
Assessed against the daemon lifecycle requirement set at master
a4c73766f4b0cc32f7c3808688eceeb6fee74335:Acceptance criteria
Non-goals
Duplicate verdict
NOT a duplicate. Verified by reading the bodies, comments, and acceptance criteria of #686, #655, #659, #669, and #678. #686 is the closest and remains detection-only: its criteria inventory and reject untrusted processes but never bound the sanctioned population or retire a superseded cohort.