Compare commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
2d95e0fcc6 | ||
|
|
8ba1c5b87c | ||
|
|
a20975688d | ||
|
|
25bc2a3291 |
@@ -292,6 +292,108 @@ health, workflow/schema SHA-256 hashes, and stale-runtime warnings when the
|
||||
checkout is behind merged safety-gate changes. Restart guidance links to #420;
|
||||
no tokens or MCP restart actions are exposed.
|
||||
|
||||
## Workflow-event timeline (#637)
|
||||
|
||||
`GET /api/v1/timeline` is a read-only, versioned aggregation of workflow
|
||||
events from every available source into one normalised, filterable stream. It
|
||||
is the model layer for the Phase 1 timeline console view (a later child issue
|
||||
of #631); this issue ships the schema, adapters, and read API only.
|
||||
|
||||
### Schema (versioned)
|
||||
|
||||
`webui/timeline.py` declares `TIMELINE_SCHEMA_VERSION` (currently `1`) and the
|
||||
frozen `WorkflowEvent` record. Every response carries `schema_version` so a
|
||||
consumer can branch on shape. One event:
|
||||
|
||||
```json
|
||||
{
|
||||
"source": "control_plane",
|
||||
"event_type": "lease.renew",
|
||||
"event_key": "cp:1421",
|
||||
"timestamp": "2026-07-23T02:00:00Z",
|
||||
"actor": null,
|
||||
"role": null,
|
||||
"issue_number": 637,
|
||||
"pr_number": null,
|
||||
"session_id": null,
|
||||
"tool_name": null,
|
||||
"decision": null,
|
||||
"message": "lease renewed",
|
||||
"correlation_id": "issue#637",
|
||||
"evidence_refs": [],
|
||||
"sensitive": true
|
||||
}
|
||||
```
|
||||
|
||||
`event_key` is stable and unique per source (`cp:<event_id>`,
|
||||
`cth:<kind>:<number>:<comment_id>`), so pagination and dedup are deterministic.
|
||||
|
||||
### Sources and field authority
|
||||
|
||||
| Source | Adapter | Authority |
|
||||
|---|---|---|
|
||||
| Control-plane `events` ⋈ `work_items` | `adapt_cp_events` | `event_type`, `message`, `timestamp`, issue/PR scope come from the CP database, read through a `mode=ro` URI (never creates the DB or runs migrations) |
|
||||
| Gitea Canonical Thread Handoff comments | `adapt_cth_comments` | `actor`, `role` (next owner), `decision`, `evidence_refs`, `timestamp` come from the parsed CTH comment body (`canonical_thread_handoff`) |
|
||||
|
||||
Handoff comments are thread-scoped: they are only read when the request filters
|
||||
by a single `issue` or `pr`. Otherwise the handoff source reports `not run`
|
||||
with a reason — it is never rendered as empty-and-healthy. Each source degrades
|
||||
independently: an unavailable control-plane DB or a failed comment fetch is a
|
||||
`sources[]` entry with `ok:false` and a `reason`, never a dropped timeline.
|
||||
|
||||
### Query parameters
|
||||
|
||||
`issue`, `pr`, `session` (conjunctive filters); `limit` (default 50, max 500)
|
||||
and `offset` for pagination; `remote`, `org`, `repo` to override the default
|
||||
registry-project scope. Events sort ascending by
|
||||
`(timestamp, source_rank, event_key)`; missing timestamps sort last.
|
||||
|
||||
### Filter authority, and refusing what cannot be answered
|
||||
|
||||
A filter dimension is only meaningful for a source whose records carry it.
|
||||
Each source declares its own support in `_SOURCE_FILTER_SUPPORT` and reports it
|
||||
per response as `supported_filters` / `unsupported_filters`:
|
||||
|
||||
| Source | issue | pr | session |
|
||||
|---|---|---|---|
|
||||
| `control_plane` | yes | yes | **no** — the `events` table is `(event_id, work_item_id, event_type, message, created_at)` and records no session |
|
||||
| `gitea_handoff` | yes | yes | yes — a CTH comment declares its own `Session:` field |
|
||||
|
||||
`session_id` is read only from that declared CTH field. It is never inferred
|
||||
from a work item, an actor, or message text, and a value that is
|
||||
redaction-altering or bare-secret-shaped is dropped rather than emitted.
|
||||
|
||||
When **no source that ran** can carry a requested dimension, the request is
|
||||
refused rather than answered: the response is `422` with `ok:false` and a
|
||||
structured `error` naming `unsupported_filters` and the per-source reason. A
|
||||
`200` with zero events would tell an operator that no such activity exists,
|
||||
which is a stronger — and false — claim than "this cannot be answered here".
|
||||
A source that *can* answer the dimension and simply matched nothing still
|
||||
returns `200` with `ok:true` and an empty page.
|
||||
|
||||
### Redaction
|
||||
|
||||
Every free-text field (event messages, decision/proof text, roles, actors) is
|
||||
passed through the console redaction policy (`webui.console_redaction`, backed
|
||||
by `gitea_audit.redact`) before it leaves the module, failing closed to the
|
||||
placeholder. No unredacted tool arguments or secrets are ever emitted, and a
|
||||
generation error never drops raw data to a caller or a log.
|
||||
|
||||
Redaction also runs *before* any structured value is derived from free text.
|
||||
`evidence_refs` are extracted from already-redacted proof/decision text, and a
|
||||
commit reference is recognised only where the text declares one (`commit`,
|
||||
`head`, `base`, `sha`, …). An undeclared 40-character hex run has the exact
|
||||
shape of a Gitea access token, so it is never lifted out of prose into a
|
||||
structured field. Every reference is then independently revalidated against an
|
||||
allowed shape and a second redaction pass immediately before serialization;
|
||||
anything unproven is dropped and the event is flagged `sensitive`.
|
||||
|
||||
### Tests
|
||||
|
||||
```bash
|
||||
pytest tests/test_webui_timeline.py -q
|
||||
```
|
||||
|
||||
## Tests
|
||||
|
||||
```bash
|
||||
|
||||
@@ -0,0 +1,618 @@
|
||||
"""Tests for the workflow-event timeline model and read API (#637).
|
||||
|
||||
Covers the acceptance criteria: versioned schema, adaptation of control-plane
|
||||
events and Gitea handoff comments, filter by issue/PR/session, redaction of
|
||||
secret-like payloads, and stable pagination.
|
||||
"""
|
||||
import json
|
||||
import os
|
||||
import sqlite3
|
||||
import sys
|
||||
import unittest
|
||||
from pathlib import Path
|
||||
|
||||
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
|
||||
|
||||
from starlette.testclient import TestClient
|
||||
|
||||
import control_plane_db
|
||||
from canonical_thread_handoff import format_cth_body
|
||||
from webui import timeline
|
||||
from webui.app import create_app
|
||||
|
||||
|
||||
def _seed_db(path: str) -> None:
|
||||
"""Create a control-plane DB and seed scoped work_items + events."""
|
||||
# Constructing ControlPlaneDB runs the schema migration once.
|
||||
control_plane_db.ControlPlaneDB(db_path=path)
|
||||
conn = sqlite3.connect(path)
|
||||
try:
|
||||
conn.execute(
|
||||
"INSERT INTO work_items(remote, org, repo, kind, number, state, updated_at) "
|
||||
"VALUES (?,?,?,?,?,?,?)",
|
||||
("prgs", "Scaled-Tech-Consulting", "Gitea-Tools", "issue", 637, "open", "2026-07-23T00:00:00Z"),
|
||||
)
|
||||
issue_wid = conn.execute("SELECT last_insert_rowid()").fetchone()[0]
|
||||
conn.execute(
|
||||
"INSERT INTO work_items(remote, org, repo, kind, number, state, updated_at) "
|
||||
"VALUES (?,?,?,?,?,?,?)",
|
||||
("prgs", "Scaled-Tech-Consulting", "Gitea-Tools", "pr", 813, "open", "2026-07-23T00:00:00Z"),
|
||||
)
|
||||
pr_wid = conn.execute("SELECT last_insert_rowid()").fetchone()[0]
|
||||
# A work item for a different repo — must never appear in prgs/Gitea-Tools scope.
|
||||
conn.execute(
|
||||
"INSERT INTO work_items(remote, org, repo, kind, number, state, updated_at) "
|
||||
"VALUES (?,?,?,?,?,?,?)",
|
||||
("dadeschools", "Other", "Elsewhere", "issue", 1, "open", "2026-07-23T00:00:00Z"),
|
||||
)
|
||||
other_wid = conn.execute("SELECT last_insert_rowid()").fetchone()[0]
|
||||
|
||||
rows = [
|
||||
(issue_wid, "allocation", "assigned author work", "2026-07-23T01:00:00Z"),
|
||||
(issue_wid, "lease.renew", "token=ghs_ABCDEF1234567890abcdef lease renewed", "2026-07-23T02:00:00Z"),
|
||||
(pr_wid, "pr.opened", "PR opened for review", "2026-07-23T03:00:00Z"),
|
||||
(other_wid, "allocation", "off-scope event", "2026-07-23T04:00:00Z"),
|
||||
]
|
||||
conn.executemany(
|
||||
"INSERT INTO events(work_item_id, event_type, message, created_at) VALUES (?,?,?,?)",
|
||||
rows,
|
||||
)
|
||||
conn.commit()
|
||||
finally:
|
||||
conn.close()
|
||||
|
||||
|
||||
class TestSchema(unittest.TestCase):
|
||||
def test_schema_is_versioned(self):
|
||||
self.assertIsInstance(timeline.TIMELINE_SCHEMA_VERSION, int)
|
||||
self.assertGreaterEqual(timeline.TIMELINE_SCHEMA_VERSION, 1)
|
||||
|
||||
def test_event_to_dict_shape(self):
|
||||
ev = timeline.WorkflowEvent(
|
||||
source=timeline.SOURCE_CONTROL_PLANE,
|
||||
event_type="allocation",
|
||||
event_key="cp:1",
|
||||
timestamp="2026-07-23T01:00:00Z",
|
||||
issue_number=637,
|
||||
)
|
||||
d = ev.to_dict()
|
||||
for key in (
|
||||
"source", "event_type", "event_key", "timestamp", "actor", "role",
|
||||
"issue_number", "pr_number", "session_id", "tool_name", "decision",
|
||||
"message", "correlation_id", "evidence_refs", "sensitive",
|
||||
):
|
||||
self.assertIn(key, d)
|
||||
self.assertEqual(d["evidence_refs"], [])
|
||||
|
||||
|
||||
class TestCpAdapter(unittest.TestCase):
|
||||
def test_issue_and_pr_mapping(self):
|
||||
rows = [
|
||||
{"event_id": 1, "event_type": "allocation", "message": "x", "created_at": "2026-07-23T01:00:00Z", "kind": "issue", "number": 637},
|
||||
{"event_id": 2, "event_type": "pr.opened", "message": "y", "created_at": "2026-07-23T02:00:00Z", "kind": "pr", "number": 813},
|
||||
]
|
||||
events = timeline.adapt_cp_events(rows)
|
||||
self.assertEqual(len(events), 2)
|
||||
self.assertEqual(events[0].issue_number, 637)
|
||||
self.assertIsNone(events[0].pr_number)
|
||||
self.assertEqual(events[0].correlation_id, "issue#637")
|
||||
self.assertIsNone(events[1].issue_number)
|
||||
self.assertEqual(events[1].pr_number, 813)
|
||||
|
||||
def test_malformed_rows_skipped(self):
|
||||
rows = [
|
||||
{"event_id": None, "event_type": "x", "kind": "issue", "number": 1},
|
||||
{"event_id": 5, "event_type": "", "kind": "issue", "number": 1},
|
||||
{"event_id": 6, "event_type": "ok", "message": "m", "created_at": None, "kind": "issue", "number": 1},
|
||||
]
|
||||
events = timeline.adapt_cp_events(rows)
|
||||
self.assertEqual(len(events), 1)
|
||||
self.assertIsNone(events[0].timestamp)
|
||||
|
||||
def test_sensitive_event_flagged(self):
|
||||
rows = [{"event_id": 1, "event_type": "lease.renew", "message": "m", "created_at": "2026-07-23T01:00:00Z", "kind": "issue", "number": 1}]
|
||||
events = timeline.adapt_cp_events(rows)
|
||||
self.assertTrue(events[0].sensitive)
|
||||
|
||||
|
||||
class TestCthAdapter(unittest.TestCase):
|
||||
def test_cth_comment_becomes_event(self):
|
||||
body = format_cth_body(
|
||||
cth_type="Author Handoff",
|
||||
status="ready",
|
||||
next_owner="reviewer",
|
||||
decision="implement timeline",
|
||||
proof="commit abc1234 closes #637",
|
||||
next_action="review PR",
|
||||
ready_to_paste_prompt="Review PR #900 as reviewer",
|
||||
)
|
||||
comments = [{"id": 42, "body": body, "created_at": "2026-07-23T05:00:00Z", "user": {"login": "jcwalker3"}}]
|
||||
events = timeline.adapt_cth_comments(comments, kind="issue", number=637)
|
||||
self.assertEqual(len(events), 1)
|
||||
ev = events[0]
|
||||
self.assertEqual(ev.source, timeline.SOURCE_GITEA_HANDOFF)
|
||||
self.assertEqual(ev.event_type, "handoff:Author Handoff")
|
||||
self.assertEqual(ev.actor, "jcwalker3")
|
||||
self.assertEqual(ev.issue_number, 637)
|
||||
self.assertEqual(ev.event_key, "cth:issue:637:42")
|
||||
self.assertIn("#637", ev.evidence_refs)
|
||||
self.assertIn("abc1234", ev.evidence_refs)
|
||||
|
||||
def test_non_cth_comment_ignored(self):
|
||||
comments = [{"id": 1, "body": "just a normal comment", "created_at": "2026-07-23T05:00:00Z", "user": {"login": "x"}}]
|
||||
self.assertEqual(timeline.adapt_cth_comments(comments, kind="issue", number=1), [])
|
||||
|
||||
|
||||
class TestRedaction(unittest.TestCase):
|
||||
def test_cp_message_redacted(self):
|
||||
rows = [{"event_id": 1, "event_type": "lease", "message": "token=ghs_ABCDEF1234567890abcdef here", "created_at": "2026-07-23T01:00:00Z", "kind": "issue", "number": 1}]
|
||||
events = timeline.adapt_cp_events(rows)
|
||||
self.assertNotIn("ghs_ABCDEF1234567890abcdef", events[0].message or "")
|
||||
|
||||
def test_handoff_decision_redacted(self):
|
||||
body = format_cth_body(
|
||||
cth_type="Blocker",
|
||||
status="blocked",
|
||||
next_owner="author",
|
||||
decision="password=SuperSecret123! must rotate",
|
||||
proof="none",
|
||||
next_action="rotate",
|
||||
ready_to_paste_prompt="Rotate the credential and retry",
|
||||
)
|
||||
comments = [{"id": 7, "body": body, "created_at": "2026-07-23T05:00:00Z", "user": {"login": "x"}}]
|
||||
events = timeline.adapt_cth_comments(comments, kind="issue", number=1)
|
||||
self.assertNotIn("SuperSecret123!", events[0].decision or "")
|
||||
|
||||
|
||||
class TestFilterSortPaginate(unittest.TestCase):
|
||||
def _events(self):
|
||||
return [
|
||||
timeline.WorkflowEvent(source="control_plane", event_type="a", event_key="cp:3", timestamp="2026-07-23T03:00:00Z", pr_number=813),
|
||||
timeline.WorkflowEvent(source="control_plane", event_type="b", event_key="cp:1", timestamp="2026-07-23T01:00:00Z", issue_number=637),
|
||||
timeline.WorkflowEvent(source="control_plane", event_type="c", event_key="cp:2", timestamp="2026-07-23T02:00:00Z", issue_number=637, session_id="sess-1"),
|
||||
]
|
||||
|
||||
def test_filter_by_issue(self):
|
||||
out = timeline.filter_events(self._events(), issue=637)
|
||||
self.assertEqual({e.event_key for e in out}, {"cp:1", "cp:2"})
|
||||
|
||||
def test_filter_by_pr(self):
|
||||
out = timeline.filter_events(self._events(), pr=813)
|
||||
self.assertEqual([e.event_key for e in out], ["cp:3"])
|
||||
|
||||
def test_filter_by_session(self):
|
||||
out = timeline.filter_events(self._events(), session="sess-1")
|
||||
self.assertEqual([e.event_key for e in out], ["cp:2"])
|
||||
|
||||
def test_stable_sort_ascending(self):
|
||||
out = timeline.sort_events(self._events())
|
||||
self.assertEqual([e.event_key for e in out], ["cp:1", "cp:2", "cp:3"])
|
||||
|
||||
def test_missing_timestamp_sorts_last(self):
|
||||
evs = self._events() + [
|
||||
timeline.WorkflowEvent(source="control_plane", event_type="z", event_key="cp:9", timestamp=None)
|
||||
]
|
||||
out = timeline.sort_events(evs)
|
||||
self.assertEqual(out[-1].event_key, "cp:9")
|
||||
|
||||
def test_pagination_windows_and_next_offset(self):
|
||||
evs = timeline.sort_events(self._events())
|
||||
page1 = timeline.paginate(evs, limit=2, offset=0)
|
||||
self.assertEqual(len(page1.events), 2)
|
||||
self.assertEqual(page1.total, 3)
|
||||
self.assertEqual(page1.next_offset, 2)
|
||||
page2 = timeline.paginate(evs, limit=2, offset=2)
|
||||
self.assertEqual(len(page2.events), 1)
|
||||
self.assertIsNone(page2.next_offset)
|
||||
|
||||
def test_pagination_bounds_coerced(self):
|
||||
evs = self._events()
|
||||
page = timeline.paginate(evs, limit=-5, offset=-3)
|
||||
self.assertGreaterEqual(page.limit, 1)
|
||||
self.assertEqual(page.offset, 0)
|
||||
|
||||
|
||||
class TestCpReader(unittest.TestCase):
|
||||
def test_reads_scoped_events_only(self):
|
||||
import tempfile
|
||||
|
||||
with tempfile.TemporaryDirectory() as tmp:
|
||||
db = os.path.join(tmp, "cp.sqlite3")
|
||||
_seed_db(db)
|
||||
events, status = timeline.read_cp_events(
|
||||
remote="prgs", org="Scaled-Tech-Consulting", repo="Gitea-Tools", db_path=db
|
||||
)
|
||||
self.assertTrue(status.ok)
|
||||
# 3 scoped events; the dadeschools/Other event is excluded.
|
||||
self.assertEqual(len(events), 3)
|
||||
self.assertTrue(all(e.source == "control_plane" for e in events))
|
||||
# Redaction applied to the token-bearing message.
|
||||
joined = " ".join(e.message or "" for e in events)
|
||||
self.assertNotIn("ghs_ABCDEF1234567890abcdef", joined)
|
||||
|
||||
def test_missing_db_degrades(self):
|
||||
events, status = timeline.read_cp_events(
|
||||
remote="prgs", org="Scaled-Tech-Consulting", repo="Gitea-Tools",
|
||||
db_path="/nonexistent/path/to/cp.sqlite3",
|
||||
)
|
||||
self.assertEqual(events, [])
|
||||
self.assertFalse(status.ok)
|
||||
self.assertIsNotNone(status.reason)
|
||||
|
||||
|
||||
class TestLoadTimeline(unittest.TestCase):
|
||||
def test_handoff_not_run_without_thread_filter(self):
|
||||
import tempfile
|
||||
|
||||
with tempfile.TemporaryDirectory() as tmp:
|
||||
db = os.path.join(tmp, "cp.sqlite3")
|
||||
_seed_db(db)
|
||||
snap = timeline.load_timeline(
|
||||
remote="prgs", org="Scaled-Tech-Consulting", repo="Gitea-Tools", db_path=db
|
||||
)
|
||||
d = snap.to_dict()
|
||||
handoff = [s for s in d["sources"] if s["name"] == "gitea_handoff"][0]
|
||||
self.assertFalse(handoff["ok"])
|
||||
self.assertIn("thread-scoped", handoff["reason"])
|
||||
self.assertEqual(d["schema_version"], timeline.TIMELINE_SCHEMA_VERSION)
|
||||
|
||||
def test_handoff_included_via_injected_source(self):
|
||||
import tempfile
|
||||
|
||||
body = format_cth_body(
|
||||
cth_type="Author Handoff", status="ready", next_owner="reviewer",
|
||||
decision="d", proof="#637", next_action="review", ready_to_paste_prompt="Review PR #1 now",
|
||||
)
|
||||
|
||||
def source(kind, number):
|
||||
return [{"id": 1, "body": body, "created_at": "2026-07-23T09:00:00Z", "user": {"login": "jcwalker3"}}]
|
||||
|
||||
with tempfile.TemporaryDirectory() as tmp:
|
||||
db = os.path.join(tmp, "cp.sqlite3")
|
||||
_seed_db(db)
|
||||
snap = timeline.load_timeline(
|
||||
remote="prgs", org="Scaled-Tech-Consulting", repo="Gitea-Tools",
|
||||
issue=637, db_path=db, comment_source=source,
|
||||
)
|
||||
d = snap.to_dict()
|
||||
handoff = [s for s in d["sources"] if s["name"] == "gitea_handoff"][0]
|
||||
self.assertTrue(handoff["ok"])
|
||||
self.assertEqual(handoff["count"], 1)
|
||||
# Both a CP event and the handoff event for issue 637 appear, sorted.
|
||||
kinds = {e["source"] for e in d["events"]}
|
||||
self.assertEqual(kinds, {"control_plane", "gitea_handoff"})
|
||||
|
||||
def test_failing_comment_source_degrades_only_handoff(self):
|
||||
import tempfile
|
||||
|
||||
def boom(kind, number):
|
||||
raise RuntimeError("network down")
|
||||
|
||||
with tempfile.TemporaryDirectory() as tmp:
|
||||
db = os.path.join(tmp, "cp.sqlite3")
|
||||
_seed_db(db)
|
||||
snap = timeline.load_timeline(
|
||||
remote="prgs", org="Scaled-Tech-Consulting", repo="Gitea-Tools",
|
||||
issue=637, db_path=db, comment_source=boom,
|
||||
)
|
||||
d = snap.to_dict()
|
||||
cp = [s for s in d["sources"] if s["name"] == "control_plane"][0]
|
||||
handoff = [s for s in d["sources"] if s["name"] == "gitea_handoff"][0]
|
||||
self.assertTrue(cp["ok"])
|
||||
self.assertFalse(handoff["ok"])
|
||||
self.assertIn("network down", handoff["reason"])
|
||||
|
||||
|
||||
# A fabricated 40-character lowercase hex value with the shape of a Gitea
|
||||
# personal access token. Never a real credential — its only job is to prove it
|
||||
# cannot reach any part of a serialized timeline payload.
|
||||
SYNTHETIC_SECRET_40_HEX = "a3f9c17be44d2058e6b17c9d0f5321ab77c4e9d1"
|
||||
|
||||
|
||||
def _cth_comment(comment_id, *, created_at, session=None, decision="d", proof="none", **kw):
|
||||
"""Build one real CTH comment record, optionally declaring a session."""
|
||||
extra = {"Session": session} if session is not None else None
|
||||
body = format_cth_body(
|
||||
cth_type=kw.pop("cth_type", "Author Handoff"),
|
||||
status=kw.pop("status", "ready"),
|
||||
next_owner=kw.pop("next_owner", "reviewer"),
|
||||
decision=decision,
|
||||
proof=proof,
|
||||
next_action=kw.pop("next_action", "review"),
|
||||
ready_to_paste_prompt=kw.pop("ready_to_paste_prompt", "Review PR #1 now"),
|
||||
extra_fields=extra,
|
||||
)
|
||||
return {
|
||||
"id": comment_id,
|
||||
"body": body,
|
||||
"created_at": created_at,
|
||||
"user": {"login": kw.pop("login", "jcwalker3")},
|
||||
}
|
||||
|
||||
|
||||
class TestSessionFilterThroughAdapter(unittest.TestCase):
|
||||
"""F1: the session dimension must be real, or explicitly refused.
|
||||
|
||||
These drive the filter through the CTH adapter and the composed
|
||||
``load_timeline``/API path, never through a hand-built ``WorkflowEvent``.
|
||||
"""
|
||||
|
||||
def _source(self, comments):
|
||||
return lambda kind, number: list(comments)
|
||||
|
||||
def test_adapter_populates_declared_session(self):
|
||||
events = timeline.adapt_cth_comments(
|
||||
[_cth_comment(1, created_at="2026-07-23T05:00:00Z", session="sess-alpha")],
|
||||
kind="issue",
|
||||
number=637,
|
||||
)
|
||||
self.assertEqual(len(events), 1)
|
||||
self.assertEqual(events[0].session_id, "sess-alpha")
|
||||
|
||||
def test_session_filter_matches_through_adapter(self):
|
||||
import tempfile
|
||||
|
||||
comments = [
|
||||
_cth_comment(1, created_at="2026-07-23T09:00:00Z", session="sess-alpha"),
|
||||
_cth_comment(2, created_at="2026-07-23T08:00:00Z", session="sess-alpha"),
|
||||
_cth_comment(3, created_at="2026-07-23T10:00:00Z", session="sess-beta"),
|
||||
]
|
||||
with tempfile.TemporaryDirectory() as tmp:
|
||||
db = os.path.join(tmp, "cp.sqlite3")
|
||||
_seed_db(db)
|
||||
snap = timeline.load_timeline(
|
||||
remote="prgs", org="Scaled-Tech-Consulting", repo="Gitea-Tools",
|
||||
issue=637, session="sess-alpha", db_path=db,
|
||||
comment_source=self._source(comments),
|
||||
)
|
||||
d = snap.to_dict()
|
||||
self.assertTrue(d["ok"])
|
||||
self.assertIsNone(d["error"])
|
||||
keys = [e["event_key"] for e in d["events"]]
|
||||
# Only the two sess-alpha events, still in ascending timestamp order.
|
||||
self.assertEqual(keys, ["cth:issue:637:2", "cth:issue:637:1"])
|
||||
self.assertTrue(all(e["session_id"] == "sess-alpha" for e in d["events"]))
|
||||
self.assertEqual(d["pagination"]["total"], 2)
|
||||
|
||||
def test_session_filter_paginates_and_keeps_order(self):
|
||||
import tempfile
|
||||
|
||||
comments = [
|
||||
_cth_comment(i, created_at=f"2026-07-23T0{i}:00:00Z", session="sess-alpha")
|
||||
for i in range(1, 4)
|
||||
]
|
||||
with tempfile.TemporaryDirectory() as tmp:
|
||||
db = os.path.join(tmp, "cp.sqlite3")
|
||||
_seed_db(db)
|
||||
kwargs = dict(
|
||||
remote="prgs", org="Scaled-Tech-Consulting", repo="Gitea-Tools",
|
||||
issue=637, session="sess-alpha", db_path=db,
|
||||
comment_source=self._source(comments),
|
||||
)
|
||||
page1 = timeline.load_timeline(limit=2, offset=0, **kwargs).to_dict()
|
||||
page2 = timeline.load_timeline(limit=2, offset=2, **kwargs).to_dict()
|
||||
self.assertEqual(
|
||||
[e["event_key"] for e in page1["events"]],
|
||||
["cth:issue:637:1", "cth:issue:637:2"],
|
||||
)
|
||||
self.assertEqual(page1["pagination"]["next_offset"], 2)
|
||||
self.assertEqual([e["event_key"] for e in page2["events"]], ["cth:issue:637:3"])
|
||||
self.assertIsNone(page2["pagination"]["next_offset"])
|
||||
|
||||
def test_unknown_session_is_honestly_empty_when_supported(self):
|
||||
"""A source that *can* answer the dimension may legitimately match nothing."""
|
||||
import tempfile
|
||||
|
||||
comments = [_cth_comment(1, created_at="2026-07-23T09:00:00Z", session="sess-alpha")]
|
||||
with tempfile.TemporaryDirectory() as tmp:
|
||||
db = os.path.join(tmp, "cp.sqlite3")
|
||||
_seed_db(db)
|
||||
d = timeline.load_timeline(
|
||||
remote="prgs", org="Scaled-Tech-Consulting", repo="Gitea-Tools",
|
||||
issue=637, session="sess-nope", db_path=db,
|
||||
comment_source=self._source(comments),
|
||||
).to_dict()
|
||||
self.assertTrue(d["ok"])
|
||||
self.assertEqual(d["events"], [])
|
||||
|
||||
def test_session_filter_refused_when_no_source_can_answer(self):
|
||||
"""The F1 defect: an empty-and-healthy page for an unanswerable filter."""
|
||||
import tempfile
|
||||
|
||||
with tempfile.TemporaryDirectory() as tmp:
|
||||
db = os.path.join(tmp, "cp.sqlite3")
|
||||
_seed_db(db)
|
||||
d = timeline.load_timeline(
|
||||
remote="prgs", org="Scaled-Tech-Consulting", repo="Gitea-Tools",
|
||||
issue=637, session="sess-alpha", db_path=db, comment_source=None,
|
||||
).to_dict()
|
||||
self.assertFalse(d["ok"])
|
||||
self.assertEqual(d["error"]["code"], "filter_not_supported")
|
||||
self.assertEqual(d["error"]["unsupported_filters"], ["session"])
|
||||
self.assertEqual(d["events"], [])
|
||||
self.assertEqual(d["pagination"]["total"], 0)
|
||||
# The refusal states which source could not answer, and why.
|
||||
explained = {r["source"] for r in d["error"]["sources"]}
|
||||
self.assertEqual(explained, {"control_plane", "gitea_handoff"})
|
||||
|
||||
def test_control_plane_declares_session_unsupported(self):
|
||||
import tempfile
|
||||
|
||||
with tempfile.TemporaryDirectory() as tmp:
|
||||
db = os.path.join(tmp, "cp.sqlite3")
|
||||
_seed_db(db)
|
||||
d = timeline.load_timeline(
|
||||
remote="prgs", org="Scaled-Tech-Consulting", repo="Gitea-Tools",
|
||||
issue=637, session="sess-alpha", db_path=db,
|
||||
comment_source=self._source([]),
|
||||
).to_dict()
|
||||
cp = [s for s in d["sources"] if s["name"] == "control_plane"][0]
|
||||
handoff = [s for s in d["sources"] if s["name"] == "gitea_handoff"][0]
|
||||
self.assertNotIn("session", cp["supported_filters"])
|
||||
self.assertEqual(cp["unsupported_filters"], ["session"])
|
||||
self.assertIn("session", handoff["supported_filters"])
|
||||
self.assertEqual(handoff["unsupported_filters"], [])
|
||||
|
||||
def test_secret_shaped_session_value_is_dropped(self):
|
||||
events = timeline.adapt_cth_comments(
|
||||
[_cth_comment(1, created_at="2026-07-23T05:00:00Z", session=SYNTHETIC_SECRET_40_HEX)],
|
||||
kind="issue",
|
||||
number=637,
|
||||
)
|
||||
self.assertIsNone(events[0].session_id)
|
||||
self.assertNotIn(SYNTHETIC_SECRET_40_HEX, json.dumps(events[0].to_dict()))
|
||||
|
||||
|
||||
class TestEvidenceRefRedaction(unittest.TestCase):
|
||||
"""F2: evidence_refs must not be a hole in the redaction boundary."""
|
||||
|
||||
def _refs_for(self, *, proof="none", decision="d"):
|
||||
events = timeline.adapt_cth_comments(
|
||||
[_cth_comment(1, created_at="2026-07-23T05:00:00Z", proof=proof, decision=decision)],
|
||||
kind="issue",
|
||||
number=637,
|
||||
)
|
||||
self.assertEqual(len(events), 1)
|
||||
return events[0]
|
||||
|
||||
def test_assigned_secret_never_reaches_evidence_refs(self):
|
||||
ev = self._refs_for(proof=f"authenticated with token={SYNTHETIC_SECRET_40_HEX}")
|
||||
self.assertNotIn(SYNTHETIC_SECRET_40_HEX, ev.evidence_refs)
|
||||
self.assertNotIn(SYNTHETIC_SECRET_40_HEX, json.dumps(ev.to_dict()))
|
||||
|
||||
def test_bare_secret_shaped_value_never_reaches_evidence_refs(self):
|
||||
# An undeclared hex run in proof text is not evidence of anything, and
|
||||
# proof itself is never serialized — so the value has no way out.
|
||||
ev = self._refs_for(proof=f"proof {SYNTHETIC_SECRET_40_HEX}", decision="rotate the credential")
|
||||
self.assertNotIn(SYNTHETIC_SECRET_40_HEX, ev.evidence_refs)
|
||||
self.assertNotIn(SYNTHETIC_SECRET_40_HEX, json.dumps(ev.to_dict()))
|
||||
|
||||
def test_secret_absent_from_complete_serialized_payload(self):
|
||||
import tempfile
|
||||
|
||||
comments = [
|
||||
_cth_comment(
|
||||
1,
|
||||
created_at="2026-07-23T09:00:00Z",
|
||||
proof=f"lease token={SYNTHETIC_SECRET_40_HEX} and bare {SYNTHETIC_SECRET_40_HEX}",
|
||||
decision=f"rotate api_key={SYNTHETIC_SECRET_40_HEX}",
|
||||
)
|
||||
]
|
||||
with tempfile.TemporaryDirectory() as tmp:
|
||||
db = os.path.join(tmp, "cp.sqlite3")
|
||||
_seed_db(db)
|
||||
snap = timeline.load_timeline(
|
||||
remote="prgs", org="Scaled-Tech-Consulting", repo="Gitea-Tools",
|
||||
issue=637, db_path=db, comment_source=lambda k, n: comments,
|
||||
)
|
||||
payload = json.dumps(snap.to_dict())
|
||||
self.assertNotIn(SYNTHETIC_SECRET_40_HEX, payload)
|
||||
# And the surface is clean by the redaction policy's own detectors.
|
||||
from webui import console_redaction
|
||||
|
||||
self.assertEqual(console_redaction.scan_for_secrets(snap.to_dict()), [])
|
||||
|
||||
def test_legitimate_references_still_usable(self):
|
||||
head = "4f3a464a1c455b6ecaaae9b6eef496c8f8ed451a"
|
||||
ev = self._refs_for(
|
||||
proof=f"closes #637, PR #849, commit abc1234, at head {head}",
|
||||
decision="none",
|
||||
)
|
||||
for token in ("#637", "#849", "abc1234", head):
|
||||
self.assertIn(token, ev.evidence_refs)
|
||||
|
||||
def test_undeclared_hex_words_are_not_references(self):
|
||||
ev = self._refs_for(proof="the record was defaced and the facade decayed")
|
||||
self.assertEqual(ev.evidence_refs, ())
|
||||
|
||||
def test_refs_revalidated_independently_before_serialization(self):
|
||||
"""Extraction is not trusted: the validator drops anything unproven."""
|
||||
safe, dropped = timeline._validated_evidence_refs(
|
||||
["#637", "abc1234", "not-a-ref", f"token={SYNTHETIC_SECRET_40_HEX}", ""]
|
||||
)
|
||||
self.assertEqual(safe, ("#637", "abc1234"))
|
||||
self.assertTrue(dropped)
|
||||
self.assertNotIn(SYNTHETIC_SECRET_40_HEX, json.dumps(list(safe)))
|
||||
|
||||
def test_clean_event_is_not_marked_sensitive(self):
|
||||
ev = self._refs_for(proof="closes #637")
|
||||
self.assertEqual(ev.evidence_refs, ("#637",))
|
||||
self.assertFalse(ev.sensitive)
|
||||
|
||||
|
||||
class TestTimelineApi(unittest.TestCase):
|
||||
def setUp(self):
|
||||
self._prev_db = os.environ.get(control_plane_db.DB_PATH_ENV)
|
||||
self._prev_offline = os.environ.get("WEBUI_TEST_OFFLINE")
|
||||
import tempfile
|
||||
|
||||
self._tmpdir = tempfile.TemporaryDirectory()
|
||||
self._db = os.path.join(self._tmpdir.name, "cp.sqlite3")
|
||||
_seed_db(self._db)
|
||||
os.environ[control_plane_db.DB_PATH_ENV] = self._db
|
||||
os.environ["WEBUI_TEST_OFFLINE"] = "1"
|
||||
self.client = TestClient(create_app())
|
||||
|
||||
def tearDown(self):
|
||||
if self._prev_db is None:
|
||||
os.environ.pop(control_plane_db.DB_PATH_ENV, None)
|
||||
else:
|
||||
os.environ[control_plane_db.DB_PATH_ENV] = self._prev_db
|
||||
if self._prev_offline is None:
|
||||
os.environ.pop("WEBUI_TEST_OFFLINE", None)
|
||||
else:
|
||||
os.environ["WEBUI_TEST_OFFLINE"] = self._prev_offline
|
||||
self._tmpdir.cleanup()
|
||||
|
||||
def test_api_returns_timeline(self):
|
||||
resp = self.client.get("/api/v1/timeline")
|
||||
self.assertEqual(resp.status_code, 200)
|
||||
body = resp.json()
|
||||
self.assertEqual(body["schema_version"], timeline.TIMELINE_SCHEMA_VERSION)
|
||||
self.assertIn("events", body)
|
||||
self.assertIn("pagination", body)
|
||||
self.assertGreaterEqual(body["pagination"]["total"], 1)
|
||||
|
||||
def test_api_filter_by_issue(self):
|
||||
resp = self.client.get("/api/v1/timeline?issue=637")
|
||||
self.assertEqual(resp.status_code, 200)
|
||||
events = resp.json()["events"]
|
||||
self.assertTrue(events)
|
||||
self.assertTrue(all(e["issue_number"] == 637 for e in events))
|
||||
|
||||
def test_api_pagination(self):
|
||||
resp = self.client.get("/api/v1/timeline?limit=1&offset=0")
|
||||
self.assertEqual(resp.status_code, 200)
|
||||
pg = resp.json()["pagination"]
|
||||
self.assertEqual(pg["limit"], 1)
|
||||
self.assertEqual(len(resp.json()["events"]), 1)
|
||||
if pg["total"] > 1:
|
||||
self.assertTrue(pg["has_more"])
|
||||
|
||||
def test_api_is_read_only(self):
|
||||
resp = self.client.post("/api/v1/timeline")
|
||||
self.assertIn(resp.status_code, (404, 405))
|
||||
|
||||
def test_api_no_secret_leak(self):
|
||||
resp = self.client.get("/api/v1/timeline?issue=637")
|
||||
self.assertNotIn("ghs_ABCDEF1234567890abcdef", resp.text)
|
||||
|
||||
def test_api_refuses_unanswerable_session_filter(self):
|
||||
"""No handoff source is configured offline, so nothing can carry a session."""
|
||||
resp = self.client.get("/api/v1/timeline?issue=637&session=sess-alpha")
|
||||
self.assertEqual(resp.status_code, 422)
|
||||
body = resp.json()
|
||||
self.assertFalse(body["ok"])
|
||||
self.assertEqual(body["error"]["code"], "filter_not_supported")
|
||||
self.assertEqual(body["error"]["unsupported_filters"], ["session"])
|
||||
self.assertEqual(body["events"], [])
|
||||
self.assertEqual(body["pagination"]["total"], 0)
|
||||
|
||||
def test_api_unfiltered_read_stays_ok(self):
|
||||
body = self.client.get("/api/v1/timeline?issue=637").json()
|
||||
self.assertTrue(body["ok"])
|
||||
self.assertIsNone(body["error"])
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
unittest.main()
|
||||
+100
@@ -45,6 +45,7 @@ from webui.worktree_scanner import load_hygiene_snapshot, snapshot_to_dict as wo
|
||||
from webui.worktree_views import render_worktrees_page
|
||||
from webui.runtime_health import load_runtime_snapshot, snapshot_to_dict as runtime_snapshot_to_dict
|
||||
from webui.runtime_views import render_runtime_page
|
||||
from webui.timeline import load_timeline, snapshot_to_dict as timeline_snapshot_to_dict
|
||||
from webui.system_health import (
|
||||
API_PATH as SYSTEM_HEALTH_API_PATH,
|
||||
load_system_health,
|
||||
@@ -410,6 +411,104 @@ async def api_console_security_model(_request: Request) -> JSONResponse:
|
||||
})
|
||||
|
||||
|
||||
def _query_int(request: Request, key: str) -> int | None:
|
||||
"""Parse an optional integer query parameter; None when absent/invalid."""
|
||||
raw = request.query_params.get(key)
|
||||
if raw is None or not str(raw).strip():
|
||||
return None
|
||||
try:
|
||||
return int(str(raw).strip())
|
||||
except (TypeError, ValueError):
|
||||
return None
|
||||
|
||||
|
||||
def _derive_remote(host: str) -> str:
|
||||
"""Map a Gitea host to its known short remote name (control-plane scope key)."""
|
||||
text = (host or "").lower()
|
||||
if "prgs" in text:
|
||||
return "prgs"
|
||||
if "dadeschools" in text:
|
||||
return "dadeschools"
|
||||
return text.split(".")[0] if text else ""
|
||||
|
||||
|
||||
def _timeline_comment_source(host: str, org: str, repo: str):
|
||||
"""Build a fail-soft CTH-comment fetcher for one repo, or None when offline.
|
||||
|
||||
Returns a callable ``(kind, number) -> list[comment]``. Credentials or
|
||||
network failures raise inside the callable so ``load_timeline`` degrades the
|
||||
handoff source rather than the whole timeline. Offline test mode yields no
|
||||
live source so the handoff section reports ``not run``.
|
||||
"""
|
||||
import os
|
||||
|
||||
from gitea_auth import api_fetch_page, get_auth_header, repo_api_url
|
||||
|
||||
offline = (os.environ.get("WEBUI_TEST_OFFLINE") or "").strip().lower() in {"1", "true", "yes"}
|
||||
if offline:
|
||||
return None
|
||||
auth = get_auth_header(host)
|
||||
if not auth:
|
||||
return None
|
||||
|
||||
def _fetch(kind: str, number: int) -> list:
|
||||
segment = "pulls" if kind == "pr" else "issues"
|
||||
url = f"{repo_api_url(host, org, repo)}/{segment}/{int(number)}/comments"
|
||||
comments: list = []
|
||||
page = 1
|
||||
while page <= 20:
|
||||
raw, meta = api_fetch_page(url, auth, page=page, limit=50)
|
||||
comments.extend(raw)
|
||||
if bool(meta["is_final_page"]):
|
||||
break
|
||||
page += 1
|
||||
return comments
|
||||
|
||||
return _fetch
|
||||
|
||||
|
||||
async def api_v1_timeline(request: Request) -> JSONResponse:
|
||||
"""Read-only workflow-event timeline (#637). Filter by issue/PR/session."""
|
||||
from webui.queue_loader import _host_from_url # host normalisation helper
|
||||
|
||||
registry, error = _load_project_registry()
|
||||
if error is not None:
|
||||
return JSONResponse(error.to_dict(), status_code=500)
|
||||
project = registry.projects[0] if registry.projects else None
|
||||
|
||||
org = request.query_params.get("org") or (project.gitea_owner if project else "")
|
||||
repo = request.query_params.get("repo") or (project.repo_name if project else "")
|
||||
host = _host_from_url(project.remote_host) if project else ""
|
||||
remote = request.query_params.get("remote") or _derive_remote(host)
|
||||
|
||||
if not (remote and org and repo):
|
||||
return JSONResponse(
|
||||
{
|
||||
"error": "timeline_scope_unresolved",
|
||||
"detail": "no project in registry and no remote/org/repo query params provided",
|
||||
},
|
||||
status_code=400,
|
||||
)
|
||||
|
||||
comment_source = _timeline_comment_source(host, org, repo) if (host and org and repo) else None
|
||||
|
||||
snapshot = load_timeline(
|
||||
remote=remote,
|
||||
org=org,
|
||||
repo=repo,
|
||||
issue=_query_int(request, "issue"),
|
||||
pr=_query_int(request, "pr"),
|
||||
session=(request.query_params.get("session") or None),
|
||||
limit=_query_int(request, "limit"),
|
||||
offset=_query_int(request, "offset"),
|
||||
comment_source=comment_source,
|
||||
)
|
||||
# A filter no surviving source can carry is refused, not answered empty:
|
||||
# a 200 with zero events would tell the operator no such activity exists.
|
||||
status_code = 200 if snapshot.ok else 422
|
||||
return JSONResponse(timeline_snapshot_to_dict(snapshot), status_code=status_code)
|
||||
|
||||
|
||||
async def method_not_allowed(request: Request, _exc: Exception) -> Response:
|
||||
path = request.url.path
|
||||
if path in _AUDIT_MUTATION_PATHS and request.method == "POST":
|
||||
@@ -449,6 +548,7 @@ def create_app(*, bind_host: str | None = None) -> Starlette:
|
||||
Route("/api/prompts", api_prompts, methods=["GET"]),
|
||||
Route("/runtime", runtime, methods=["GET"]),
|
||||
Route("/api/runtime", api_runtime, methods=["GET"]),
|
||||
Route("/api/v1/timeline", api_v1_timeline, methods=["GET"]),
|
||||
Route("/audit", audit, methods=["GET", "POST"]),
|
||||
Route("/api/audit", api_audit, methods=["GET", "POST"]),
|
||||
Route("/worktrees", worktrees, methods=["GET"]),
|
||||
|
||||
@@ -0,0 +1,784 @@
|
||||
"""Workflow-event and conversation timeline model (#637, Phase 1).
|
||||
|
||||
Operators cannot browse a unified timeline of workflow events, decisions,
|
||||
tool calls, and handoffs: the evidence is scattered across control-plane
|
||||
events, Gitea canonical handoff comments, and local logs. This module defines
|
||||
one durable, versioned event schema and per-source adapters that normalise
|
||||
those scattered records into a single ``WorkflowEvent`` stream, plus a
|
||||
read-only query layer (filter by issue / PR / session, stable ordering,
|
||||
pagination) that the ``/api/v1/timeline`` route serves.
|
||||
|
||||
Design rules honoured here:
|
||||
|
||||
- **Read-only.** Sources are read; nothing is mutated. The control-plane
|
||||
database is opened through a ``mode=ro`` URI so a missing or unwritable DB
|
||||
degrades to a reason instead of creating directories or running migrations.
|
||||
- **Fail-soft per source.** An unavailable source degrades to a status with a
|
||||
reason rather than raising, and a source that could not run is never
|
||||
rendered as an empty-and-healthy timeline.
|
||||
- **Answerable filters only.** Each source declares which filter dimensions it
|
||||
can actually answer. A filter dimension no source that ran can carry is
|
||||
refused with an explicit reason rather than silently matching nothing: an
|
||||
empty page from an unanswerable filter reads to an operator as "no such
|
||||
activity", which is a different — and false — statement.
|
||||
- **Redaction at the boundary, fail closed.** Every free-text field (event
|
||||
messages, redacted tool arguments, decision/proof text) is run through the
|
||||
console redaction policy before it leaves this module, and *before* any
|
||||
structured value is derived from it — evidence references are extracted from
|
||||
redacted text, then independently revalidated before serialization. An
|
||||
unredactable value becomes the placeholder, and a value that cannot be proven
|
||||
safe is dropped — an unredacted payload is never emitted, and a generation
|
||||
error never drops raw data to a caller or a log.
|
||||
- **Stable ordering.** Events sort by ``(timestamp, source_rank, event_key)``
|
||||
with a deterministic tiebreak, so pagination is stable across calls and
|
||||
events with equal or missing timestamps keep a fixed order.
|
||||
|
||||
Non-goals (from the issue): no full chat replay, no mutation of historical
|
||||
events, no unredacted tool-argument storage.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import re
|
||||
import sqlite3
|
||||
from dataclasses import dataclass, replace
|
||||
from datetime import datetime, timezone
|
||||
from typing import Any, Callable, Iterable
|
||||
|
||||
import control_plane_db
|
||||
from webui import console_redaction
|
||||
|
||||
# The schema is versioned so consumers can branch on shape. Bump on any
|
||||
# breaking change to WorkflowEvent's serialized form.
|
||||
TIMELINE_SCHEMA_VERSION = 1
|
||||
|
||||
# Known event sources and their deterministic ordering rank. When two events
|
||||
# carry the same timestamp, the source rank breaks the tie before the
|
||||
# per-source event key, so a control-plane event and a handoff comment minted
|
||||
# in the same second always sort in a fixed order.
|
||||
SOURCE_CONTROL_PLANE = "control_plane"
|
||||
SOURCE_GITEA_HANDOFF = "gitea_handoff"
|
||||
_SOURCE_RANK = {
|
||||
SOURCE_CONTROL_PLANE: 0,
|
||||
SOURCE_GITEA_HANDOFF: 1,
|
||||
}
|
||||
|
||||
# The filter dimensions the query layer accepts.
|
||||
FILTER_ISSUE = "issue"
|
||||
FILTER_PR = "pr"
|
||||
FILTER_SESSION = "session"
|
||||
|
||||
# Which dimensions each source can actually answer. This is a property of the
|
||||
# underlying records, not of the query code: the control-plane ``events`` table
|
||||
# is (event_id, work_item_id, event_type, message, created_at) and carries no
|
||||
# session identity at all, so no control-plane event can ever match a session
|
||||
# filter. A CTH handoff comment can declare its session as a field, so the
|
||||
# handoff source answers all three. Filtering on a dimension the surviving
|
||||
# sources cannot carry is refused in ``load_timeline`` rather than answered
|
||||
# with an empty page.
|
||||
_SOURCE_FILTER_SUPPORT: dict[str, tuple[str, ...]] = {
|
||||
SOURCE_CONTROL_PLANE: (FILTER_ISSUE, FILTER_PR),
|
||||
SOURCE_GITEA_HANDOFF: (FILTER_ISSUE, FILTER_PR, FILTER_SESSION),
|
||||
}
|
||||
|
||||
# Why a source cannot answer a dimension, for the refusal reason an operator reads.
|
||||
_SOURCE_FILTER_LIMITS: dict[tuple[str, str], str] = {
|
||||
(SOURCE_CONTROL_PLANE, FILTER_SESSION): (
|
||||
"control-plane events carry no session identity "
|
||||
"(the events table has no session column)"
|
||||
),
|
||||
}
|
||||
|
||||
# A timestamp far in the future so events with no parseable timestamp sort
|
||||
# last (after everything real) instead of first, without raising.
|
||||
_MISSING_TS_SORT = "9999-12-31T23:59:59Z"
|
||||
|
||||
|
||||
def _parse_ts(value: str | None) -> str | None:
|
||||
"""Normalise a timestamp to ``...Z`` UTC, or None when unparseable."""
|
||||
if not value:
|
||||
return None
|
||||
text = str(value).strip()
|
||||
if not text:
|
||||
return None
|
||||
candidate = text[:-1] + "+00:00" if text.endswith("Z") else text
|
||||
try:
|
||||
parsed = datetime.fromisoformat(candidate)
|
||||
except ValueError:
|
||||
return None
|
||||
if parsed.tzinfo is None:
|
||||
parsed = parsed.replace(tzinfo=timezone.utc)
|
||||
return parsed.astimezone(timezone.utc).replace(microsecond=0).isoformat().replace("+00:00", "Z")
|
||||
|
||||
|
||||
def _redact(value: Any) -> Any:
|
||||
"""Redact a single free-text field, failing closed to the placeholder."""
|
||||
if value is None:
|
||||
return None
|
||||
return console_redaction.redact_text(str(value))
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class WorkflowEvent:
|
||||
"""One normalised timeline event.
|
||||
|
||||
Every field is optional except ``source``/``event_type``/``event_key``
|
||||
because sources carry different subsets. The class is frozen so an adapted
|
||||
event is an immutable record; a consumer that needs a variant builds a new
|
||||
one rather than mutating history.
|
||||
"""
|
||||
|
||||
source: str
|
||||
event_type: str
|
||||
event_key: str
|
||||
timestamp: str | None = None
|
||||
actor: str | None = None
|
||||
role: str | None = None
|
||||
issue_number: int | None = None
|
||||
pr_number: int | None = None
|
||||
session_id: str | None = None
|
||||
tool_name: str | None = None
|
||||
decision: str | None = None
|
||||
message: str | None = None
|
||||
correlation_id: str | None = None
|
||||
evidence_refs: tuple[str, ...] = ()
|
||||
sensitive: bool = False
|
||||
|
||||
def sort_key(self) -> tuple[str, int, str]:
|
||||
return (
|
||||
self.timestamp or _MISSING_TS_SORT,
|
||||
_SOURCE_RANK.get(self.source, 99),
|
||||
self.event_key,
|
||||
)
|
||||
|
||||
def to_dict(self) -> dict[str, Any]:
|
||||
return {
|
||||
"source": self.source,
|
||||
"event_type": self.event_type,
|
||||
"event_key": self.event_key,
|
||||
"timestamp": self.timestamp,
|
||||
"actor": self.actor,
|
||||
"role": self.role,
|
||||
"issue_number": self.issue_number,
|
||||
"pr_number": self.pr_number,
|
||||
"session_id": self.session_id,
|
||||
"tool_name": self.tool_name,
|
||||
"decision": self.decision,
|
||||
"message": self.message,
|
||||
"correlation_id": self.correlation_id,
|
||||
"evidence_refs": list(self.evidence_refs),
|
||||
"sensitive": self.sensitive,
|
||||
}
|
||||
|
||||
|
||||
# --------------------------------------------------------------------------- #
|
||||
# Adapters — pure functions from a source's raw records to WorkflowEvents. #
|
||||
# Each is total: a malformed record is skipped, never raised on. #
|
||||
# --------------------------------------------------------------------------- #
|
||||
|
||||
# Event types whose payload is treated as sensitive and always redaction-hard
|
||||
# (they can carry lease/session provenance or tool arguments).
|
||||
_SENSITIVE_EVENT_HINTS = ("lease", "capability", "token", "auth", "secret")
|
||||
|
||||
# Reference tokens (issue/PR/comment ids) and SHAs parsed out of proof text.
|
||||
_EVIDENCE_REF_RE = re.compile(r"(?:#|PR\s*#?|issue\s*#?|comment\s*#?)(\d+)", re.IGNORECASE)
|
||||
|
||||
# A commit reference is only recognised when the text *declares* it as one.
|
||||
# A bare lowercase hex run is not evidence of anything: at 40 characters it is
|
||||
# exactly the shape of a Gitea personal access token, and at 7 it also matches
|
||||
# ordinary words such as "defaced". Requiring an anchoring keyword keeps real
|
||||
# references ("commit abc1234", "at head a209756...", "base caaae9b6") usable
|
||||
# while refusing to lift an undeclared secret-shaped run out of free text.
|
||||
_SHA_RE = re.compile(
|
||||
r"(?i:\b(?:commit|sha|head|base|parent|revision|rev|merge[- ]base)\b[\s:=@#]*)"
|
||||
r"([0-9a-f]{7,40})\b"
|
||||
)
|
||||
|
||||
# Shapes a serialized evidence reference is allowed to take. Anything else is
|
||||
# dropped rather than emitted.
|
||||
_REF_ISSUE_SHAPE = re.compile(r"^#[0-9]{1,9}$")
|
||||
_REF_SHA_SHAPE = re.compile(r"^[0-9a-f]{7,40}$")
|
||||
|
||||
# A long undelimited hex run with no declaring context is treated as credential
|
||||
# material wherever it appears, never as an identifier.
|
||||
_BARE_SECRET_SHAPE = re.compile(r"^[0-9a-f]{32,}$")
|
||||
|
||||
|
||||
def _kind_to_numbers(kind: str | None, number: int | None) -> tuple[int | None, int | None]:
|
||||
"""Map a control-plane work-item (kind, number) to (issue_no, pr_no)."""
|
||||
if number is None:
|
||||
return (None, None)
|
||||
if kind == "pr":
|
||||
return (None, int(number))
|
||||
if kind == "issue":
|
||||
return (int(number), None)
|
||||
return (None, None)
|
||||
|
||||
|
||||
def _correlation_for(kind: str | None, number: int | None) -> str | None:
|
||||
if number is None or kind not in ("issue", "pr"):
|
||||
return None
|
||||
return f"{kind}#{number}"
|
||||
|
||||
|
||||
def _extract_evidence_refs(*texts: str | None) -> tuple[str, ...]:
|
||||
"""Extract issue/PR and declared-commit references from **redacted** text.
|
||||
|
||||
Callers must pass text that has already been through :func:`_redact`; this
|
||||
function derives a structured field from its input, so extracting ahead of
|
||||
redaction would republish whatever redaction was about to remove. Every
|
||||
reference is revalidated by :func:`_validated_evidence_refs` before it is
|
||||
serialized.
|
||||
"""
|
||||
refs: list[str] = []
|
||||
for text in texts:
|
||||
if not text:
|
||||
continue
|
||||
for match in _EVIDENCE_REF_RE.finditer(text):
|
||||
token = f"#{match.group(1)}"
|
||||
if token not in refs:
|
||||
refs.append(token)
|
||||
for match in _SHA_RE.finditer(text):
|
||||
token = match.group(1)
|
||||
if token not in refs:
|
||||
refs.append(token)
|
||||
return tuple(refs)
|
||||
|
||||
|
||||
def _validated_evidence_refs(refs: Iterable[str]) -> tuple[tuple[str, ...], bool]:
|
||||
"""Independently revalidate references immediately before serialization.
|
||||
|
||||
Extraction is not trusted on its own. A reference survives only when it has
|
||||
a known reference shape and is unchanged by a second redaction pass — a
|
||||
value the redaction policy would alter is credential material that must not
|
||||
be emitted as a structured field. A full 40-character SHA stays usable
|
||||
because extraction only accepts a hex run the source text explicitly
|
||||
declared as a commit. Returns ``(safe_refs, dropped_any)``; ``dropped_any``
|
||||
marks the event sensitive so the drop is visible rather than silent.
|
||||
"""
|
||||
safe: list[str] = []
|
||||
dropped = False
|
||||
for ref in refs or ():
|
||||
try:
|
||||
token = str(ref).strip()
|
||||
if not token:
|
||||
continue
|
||||
recognised = bool(_REF_ISSUE_SHAPE.match(token) or _REF_SHA_SHAPE.match(token))
|
||||
if not recognised:
|
||||
dropped = True
|
||||
continue
|
||||
if _redact(token) != token:
|
||||
dropped = True
|
||||
continue
|
||||
if token not in safe:
|
||||
safe.append(token)
|
||||
except Exception:
|
||||
# Fail closed: a reference that cannot be proven safe is dropped.
|
||||
dropped = True
|
||||
continue
|
||||
return (tuple(safe), dropped)
|
||||
|
||||
|
||||
def _safe_session_id(value: Any) -> str | None:
|
||||
"""Return a session identifier only when it is safe to emit.
|
||||
|
||||
The value is authoritative source data — a session the record names for
|
||||
itself — but it is still free text. It is dropped when redaction alters it
|
||||
or when it is a bare secret-shaped hex run, so a credential parked in a
|
||||
session field can never reach the payload or be echoed back by a filter.
|
||||
"""
|
||||
if value is None:
|
||||
return None
|
||||
text = str(value).strip()
|
||||
if not text:
|
||||
return None
|
||||
if _BARE_SECRET_SHAPE.match(text):
|
||||
return None
|
||||
return text if _redact(text) == text else None
|
||||
|
||||
|
||||
def adapt_cp_events(rows: Iterable[dict[str, Any]]) -> list[WorkflowEvent]:
|
||||
"""Adapt control-plane ``events`` rows (joined to work_items) into events.
|
||||
|
||||
Each row is expected to carry ``event_id``, ``event_type``, ``message``,
|
||||
``created_at`` and the joined work-item ``kind``/``number``. Rows missing
|
||||
an id or type are skipped so a partially written table never raises.
|
||||
"""
|
||||
events: list[WorkflowEvent] = []
|
||||
for row in rows or []:
|
||||
try:
|
||||
event_id = row.get("event_id")
|
||||
event_type = (row.get("event_type") or "").strip()
|
||||
if event_id is None or not event_type:
|
||||
continue
|
||||
kind = row.get("kind")
|
||||
number = row.get("number")
|
||||
issue_no, pr_no = _kind_to_numbers(kind, number)
|
||||
sensitive = any(hint in event_type.lower() for hint in _SENSITIVE_EVENT_HINTS)
|
||||
events.append(
|
||||
WorkflowEvent(
|
||||
source=SOURCE_CONTROL_PLANE,
|
||||
event_type=event_type,
|
||||
event_key=f"cp:{event_id}",
|
||||
timestamp=_parse_ts(row.get("created_at")),
|
||||
issue_number=issue_no,
|
||||
pr_number=pr_no,
|
||||
# No session_id: the control-plane events table is
|
||||
# (event_id, work_item_id, event_type, message, created_at)
|
||||
# and records no session. Inventing one from the work item
|
||||
# or the message text would be a guess, so this source
|
||||
# declares the session dimension unsupported instead
|
||||
# (_SOURCE_FILTER_SUPPORT) and the query layer refuses a
|
||||
# session filter it cannot honestly answer.
|
||||
message=_redact(row.get("message")),
|
||||
correlation_id=_correlation_for(kind, number),
|
||||
sensitive=sensitive,
|
||||
)
|
||||
)
|
||||
except Exception:
|
||||
# A single malformed row must not sink the whole adaptation.
|
||||
continue
|
||||
return events
|
||||
|
||||
|
||||
def adapt_cth_comments(
|
||||
comments: Iterable[dict[str, Any]],
|
||||
*,
|
||||
kind: str,
|
||||
number: int,
|
||||
) -> list[WorkflowEvent]:
|
||||
"""Adapt Gitea Canonical Thread Handoff (CTH) comments into events.
|
||||
|
||||
Only comments that parse as a CTH (``canonical_thread_handoff.parse_cth_comment``)
|
||||
become events; ordinary comments are ignored. ``kind``/``number`` scope the
|
||||
events to the issue or PR the comments belong to.
|
||||
"""
|
||||
# Imported lazily so this module has no import-time dependency on the
|
||||
# handoff parser when only the control-plane adapter is used.
|
||||
from canonical_thread_handoff import parse_cth_comment
|
||||
|
||||
issue_no, pr_no = _kind_to_numbers(kind, number)
|
||||
correlation = _correlation_for(kind, number)
|
||||
events: list[WorkflowEvent] = []
|
||||
for comment in comments or []:
|
||||
try:
|
||||
body = comment.get("body") or ""
|
||||
parsed = parse_cth_comment(body)
|
||||
if not parsed:
|
||||
continue
|
||||
fields = parsed.get("fields") or {}
|
||||
cth_type = parsed.get("cth_type") or "handoff"
|
||||
comment_id = comment.get("id")
|
||||
# Redaction runs first, and every derived value is taken from the
|
||||
# redacted text — deriving evidence refs from the raw proof would
|
||||
# re-emit exactly what redaction was about to remove.
|
||||
decision = _redact(fields.get("decision"))
|
||||
proof = _redact(fields.get("proof"))
|
||||
next_action = _redact(fields.get("next action"))
|
||||
refs, refs_dropped = _validated_evidence_refs(
|
||||
_extract_evidence_refs(proof, decision)
|
||||
)
|
||||
events.append(
|
||||
WorkflowEvent(
|
||||
source=SOURCE_GITEA_HANDOFF,
|
||||
event_type=f"handoff:{cth_type}",
|
||||
event_key=f"cth:{kind}:{number}:{comment_id}",
|
||||
timestamp=_parse_ts(comment.get("created_at")),
|
||||
actor=_redact((comment.get("user") or {}).get("login")),
|
||||
role=_redact(fields.get("next owner")),
|
||||
issue_number=issue_no,
|
||||
pr_number=pr_no,
|
||||
# A CTH names its own session when the producer records one;
|
||||
# it is read from that declared field, never inferred from
|
||||
# unrelated text.
|
||||
session_id=_safe_session_id(fields.get("session")),
|
||||
decision=decision,
|
||||
message=next_action or _redact(fields.get("status")),
|
||||
correlation_id=correlation,
|
||||
evidence_refs=refs,
|
||||
sensitive=refs_dropped,
|
||||
)
|
||||
)
|
||||
except Exception:
|
||||
continue
|
||||
return events
|
||||
|
||||
|
||||
# --------------------------------------------------------------------------- #
|
||||
# Read-only control-plane event source. #
|
||||
# --------------------------------------------------------------------------- #
|
||||
|
||||
_CP_EVENTS_QUERY = """
|
||||
SELECT e.event_id AS event_id,
|
||||
e.event_type AS event_type,
|
||||
e.message AS message,
|
||||
e.created_at AS created_at,
|
||||
w.kind AS kind,
|
||||
w.number AS number
|
||||
FROM events e
|
||||
JOIN work_items w ON e.work_item_id = w.work_item_id
|
||||
WHERE w.remote = ? AND w.org = ? AND w.repo = ?
|
||||
"""
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class SourceStatus:
|
||||
"""Fail-soft status for one timeline source.
|
||||
|
||||
``supported_filters`` states which filter dimensions this source's records
|
||||
can carry; ``unsupported_filters`` names the requested dimensions it cannot,
|
||||
so an operator can see *why* a source contributed nothing rather than being
|
||||
left to read an empty list as an absence of activity.
|
||||
"""
|
||||
|
||||
name: str
|
||||
ok: bool
|
||||
reason: str | None = None
|
||||
count: int = 0
|
||||
supported_filters: tuple[str, ...] = ()
|
||||
unsupported_filters: tuple[str, ...] = ()
|
||||
|
||||
def to_dict(self) -> dict[str, Any]:
|
||||
return {
|
||||
"name": self.name,
|
||||
"ok": self.ok,
|
||||
"reason": self.reason,
|
||||
"count": self.count,
|
||||
"supported_filters": list(self.supported_filters),
|
||||
"unsupported_filters": list(self.unsupported_filters),
|
||||
}
|
||||
|
||||
|
||||
def _cp_status(*, ok: bool, reason: str | None = None, count: int = 0) -> SourceStatus:
|
||||
return SourceStatus(
|
||||
SOURCE_CONTROL_PLANE,
|
||||
ok=ok,
|
||||
reason=reason,
|
||||
count=count,
|
||||
supported_filters=_SOURCE_FILTER_SUPPORT[SOURCE_CONTROL_PLANE],
|
||||
)
|
||||
|
||||
|
||||
def _handoff_status(*, ok: bool, reason: str | None = None, count: int = 0) -> SourceStatus:
|
||||
return SourceStatus(
|
||||
SOURCE_GITEA_HANDOFF,
|
||||
ok=ok,
|
||||
reason=reason,
|
||||
count=count,
|
||||
supported_filters=_SOURCE_FILTER_SUPPORT[SOURCE_GITEA_HANDOFF],
|
||||
)
|
||||
|
||||
|
||||
def read_cp_events(
|
||||
*,
|
||||
remote: str,
|
||||
org: str,
|
||||
repo: str,
|
||||
db_path: str | None = None,
|
||||
) -> tuple[list[WorkflowEvent], SourceStatus]:
|
||||
"""Read scoped control-plane events read-only. Never creates the DB.
|
||||
|
||||
Opens the SQLite file through a ``mode=ro`` URI: a health/timeline read
|
||||
must never create directories or run the schema migration that
|
||||
``ControlPlaneDB()`` performs on construction. A missing or unreadable DB
|
||||
degrades to a status with a reason.
|
||||
"""
|
||||
path = (db_path or control_plane_db.default_db_path()).strip()
|
||||
conn: sqlite3.Connection | None = None
|
||||
try:
|
||||
conn = sqlite3.connect(f"file:{path}?mode=ro", uri=True)
|
||||
conn.row_factory = sqlite3.Row
|
||||
cursor = conn.execute(_CP_EVENTS_QUERY, (remote, org, repo))
|
||||
rows = [dict(r) for r in cursor.fetchall()]
|
||||
except sqlite3.OperationalError as exc:
|
||||
return ([], _cp_status(ok=False, reason=f"control-plane DB unavailable: {exc}"))
|
||||
except sqlite3.Error as exc:
|
||||
return ([], _cp_status(ok=False, reason=f"control-plane read failed: {exc}"))
|
||||
finally:
|
||||
if conn is not None:
|
||||
conn.close()
|
||||
events = adapt_cp_events(rows)
|
||||
return (events, _cp_status(ok=True, count=len(events)))
|
||||
|
||||
|
||||
# --------------------------------------------------------------------------- #
|
||||
# Filter, sort, paginate. #
|
||||
# --------------------------------------------------------------------------- #
|
||||
|
||||
|
||||
def filter_events(
|
||||
events: Iterable[WorkflowEvent],
|
||||
*,
|
||||
issue: int | None = None,
|
||||
pr: int | None = None,
|
||||
session: str | None = None,
|
||||
) -> list[WorkflowEvent]:
|
||||
"""Filter events by issue number, PR number, and/or session id.
|
||||
|
||||
Filters are conjunctive. A filter that names a dimension an event does not
|
||||
carry excludes that event (an issue filter excludes PR-only events).
|
||||
"""
|
||||
out: list[WorkflowEvent] = []
|
||||
for ev in events:
|
||||
if issue is not None and ev.issue_number != issue:
|
||||
continue
|
||||
if pr is not None and ev.pr_number != pr:
|
||||
continue
|
||||
if session is not None and ev.session_id != session:
|
||||
continue
|
||||
out.append(ev)
|
||||
return out
|
||||
|
||||
|
||||
def sort_events(events: Iterable[WorkflowEvent]) -> list[WorkflowEvent]:
|
||||
"""Return events in stable timeline order (ascending)."""
|
||||
return sorted(events, key=lambda ev: ev.sort_key())
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class TimelinePage:
|
||||
"""One page of the sorted, filtered timeline."""
|
||||
|
||||
events: tuple[WorkflowEvent, ...]
|
||||
total: int
|
||||
limit: int
|
||||
offset: int
|
||||
|
||||
@property
|
||||
def next_offset(self) -> int | None:
|
||||
nxt = self.offset + len(self.events)
|
||||
return nxt if nxt < self.total else None
|
||||
|
||||
def to_dict(self) -> dict[str, Any]:
|
||||
return {
|
||||
"events": [ev.to_dict() for ev in self.events],
|
||||
"pagination": {
|
||||
"total": self.total,
|
||||
"limit": self.limit,
|
||||
"offset": self.offset,
|
||||
"returned": len(self.events),
|
||||
"next_offset": self.next_offset,
|
||||
"has_more": self.next_offset is not None,
|
||||
},
|
||||
}
|
||||
|
||||
|
||||
_MAX_LIMIT = 500
|
||||
_DEFAULT_LIMIT = 50
|
||||
|
||||
|
||||
def _coerce_bounds(limit: int | None, offset: int | None) -> tuple[int, int]:
|
||||
try:
|
||||
lim = int(limit) if limit is not None else _DEFAULT_LIMIT
|
||||
except (TypeError, ValueError):
|
||||
lim = _DEFAULT_LIMIT
|
||||
try:
|
||||
off = int(offset) if offset is not None else 0
|
||||
except (TypeError, ValueError):
|
||||
off = 0
|
||||
lim = max(1, min(lim, _MAX_LIMIT))
|
||||
off = max(0, off)
|
||||
return (lim, off)
|
||||
|
||||
|
||||
def paginate(events: list[WorkflowEvent], *, limit: int | None, offset: int | None) -> TimelinePage:
|
||||
lim, off = _coerce_bounds(limit, offset)
|
||||
window = events[off : off + lim]
|
||||
return TimelinePage(events=tuple(window), total=len(events), limit=lim, offset=off)
|
||||
|
||||
|
||||
# --------------------------------------------------------------------------- #
|
||||
# Composition — load_timeline aggregates all sources, fail-soft. #
|
||||
# --------------------------------------------------------------------------- #
|
||||
|
||||
# A comment source is a callable that, given (kind, number), returns the raw
|
||||
# Gitea comment list for that issue/PR. The route supplies a live fail-soft
|
||||
# fetcher; tests supply a fixture. When None, the handoff source is reported as
|
||||
# not-run (never silently empty-and-healthy).
|
||||
CommentSource = Callable[[str, int], list[dict[str, Any]]]
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class TimelineSnapshot:
|
||||
"""One answered timeline query.
|
||||
|
||||
``ok`` is False when the query could not be answered as asked — currently
|
||||
when a requested filter dimension no surviving source can carry was
|
||||
supplied. The page is then empty *and* the snapshot says so, because an
|
||||
``ok`` empty page is a claim that no such activity exists.
|
||||
"""
|
||||
|
||||
schema_version: int
|
||||
remote: str
|
||||
org: str
|
||||
repo: str
|
||||
filters: dict[str, Any]
|
||||
page: TimelinePage
|
||||
sources: tuple[SourceStatus, ...]
|
||||
ok: bool = True
|
||||
error: dict[str, Any] | None = None
|
||||
|
||||
def to_dict(self) -> dict[str, Any]:
|
||||
return {
|
||||
"ok": self.ok,
|
||||
"error": self.error,
|
||||
"schema_version": self.schema_version,
|
||||
"scope": {"remote": self.remote, "org": self.org, "repo": self.repo},
|
||||
"filters": self.filters,
|
||||
"sources": [s.to_dict() for s in self.sources],
|
||||
**self.page.to_dict(),
|
||||
}
|
||||
|
||||
|
||||
def _unanswerable_reasons(
|
||||
statuses: Iterable[SourceStatus], unanswerable: Iterable[str]
|
||||
) -> list[dict[str, str]]:
|
||||
"""Explain, per source, why each unanswerable dimension went unanswered."""
|
||||
out: list[dict[str, str]] = []
|
||||
for status in statuses:
|
||||
for dim in unanswerable:
|
||||
if dim not in status.supported_filters:
|
||||
reason = _SOURCE_FILTER_LIMITS.get(
|
||||
(status.name, dim), f"this source's records carry no {dim} identity"
|
||||
)
|
||||
elif not status.ok:
|
||||
reason = (
|
||||
f"this source can carry {dim} but did not run: "
|
||||
f"{status.reason or 'unavailable'}"
|
||||
)
|
||||
else:
|
||||
continue
|
||||
out.append({"source": status.name, "filter": dim, "reason": reason})
|
||||
return out
|
||||
|
||||
|
||||
def load_timeline(
|
||||
*,
|
||||
remote: str,
|
||||
org: str,
|
||||
repo: str,
|
||||
issue: int | None = None,
|
||||
pr: int | None = None,
|
||||
session: str | None = None,
|
||||
limit: int | None = None,
|
||||
offset: int | None = None,
|
||||
db_path: str | None = None,
|
||||
comment_source: CommentSource | None = None,
|
||||
) -> TimelineSnapshot:
|
||||
"""Aggregate every timeline source into one filtered, paginated snapshot.
|
||||
|
||||
Sources are read independently and fail soft: an unavailable source
|
||||
contributes a ``SourceStatus`` with ``ok=False`` and a reason, and never
|
||||
collapses the whole timeline. The handoff source only runs when a specific
|
||||
issue or PR is requested (a handoff comment belongs to one thread) and a
|
||||
``comment_source`` is available; otherwise it is reported as ``not run``
|
||||
rather than as an empty-and-healthy source.
|
||||
|
||||
A filter dimension that no surviving source can carry — a ``session``
|
||||
filter when the only source that ran is the control plane, whose events
|
||||
record no session — is refused with ``ok=False`` and a structured error
|
||||
instead of being answered with an empty page.
|
||||
"""
|
||||
all_events: list[WorkflowEvent] = []
|
||||
statuses: list[SourceStatus] = []
|
||||
|
||||
cp_events, cp_status = read_cp_events(remote=remote, org=org, repo=repo, db_path=db_path)
|
||||
all_events.extend(cp_events)
|
||||
statuses.append(cp_status)
|
||||
|
||||
# Gitea handoff comments are thread-scoped: only fetch when the caller
|
||||
# narrowed to one issue or PR, and only when a source was provided.
|
||||
handoff_target: tuple[str, int] | None = None
|
||||
if pr is not None:
|
||||
handoff_target = ("pr", pr)
|
||||
elif issue is not None:
|
||||
handoff_target = ("issue", issue)
|
||||
|
||||
if handoff_target is None:
|
||||
statuses.append(
|
||||
_handoff_status(
|
||||
ok=False,
|
||||
reason="not run: handoff comments are thread-scoped; filter by issue or pr to include them",
|
||||
)
|
||||
)
|
||||
elif comment_source is None:
|
||||
statuses.append(
|
||||
_handoff_status(
|
||||
ok=False,
|
||||
reason="not run: no comment source configured for this timeline read",
|
||||
)
|
||||
)
|
||||
else:
|
||||
kind, number = handoff_target
|
||||
try:
|
||||
comments = comment_source(kind, number) or []
|
||||
handoff_events = adapt_cth_comments(comments, kind=kind, number=number)
|
||||
all_events.extend(handoff_events)
|
||||
statuses.append(_handoff_status(ok=True, count=len(handoff_events)))
|
||||
except Exception as exc: # fail soft: a fetch/parse error degrades this source only
|
||||
statuses.append(_handoff_status(ok=False, reason=f"handoff source failed: {exc}"))
|
||||
|
||||
requested = tuple(
|
||||
name
|
||||
for name, value in ((FILTER_ISSUE, issue), (FILTER_PR, pr), (FILTER_SESSION, session))
|
||||
if value is not None
|
||||
)
|
||||
statuses = [
|
||||
replace(
|
||||
status,
|
||||
unsupported_filters=tuple(
|
||||
dim for dim in requested if dim not in status.supported_filters
|
||||
),
|
||||
)
|
||||
for status in statuses
|
||||
]
|
||||
filters = {"issue": issue, "pr": pr, "session": session}
|
||||
|
||||
# A dimension is answerable only if a source that actually ran can carry it.
|
||||
# If none can, refuse: an empty page would assert "no such activity", which
|
||||
# is a claim this timeline is not in a position to make.
|
||||
answerable: set[str] = set()
|
||||
for status in statuses:
|
||||
if status.ok:
|
||||
answerable.update(status.supported_filters)
|
||||
unanswerable = tuple(dim for dim in requested if dim not in answerable)
|
||||
|
||||
if unanswerable:
|
||||
return TimelineSnapshot(
|
||||
schema_version=TIMELINE_SCHEMA_VERSION,
|
||||
remote=remote,
|
||||
org=org,
|
||||
repo=repo,
|
||||
filters=filters,
|
||||
page=paginate([], limit=limit, offset=offset),
|
||||
sources=tuple(statuses),
|
||||
ok=False,
|
||||
error={
|
||||
"code": "filter_not_supported",
|
||||
"unsupported_filters": list(unanswerable),
|
||||
"detail": (
|
||||
"no timeline source that ran can answer "
|
||||
+ ", ".join(f"'{dim}'" for dim in unanswerable)
|
||||
+ "; the result is refused rather than returned empty"
|
||||
),
|
||||
"sources": _unanswerable_reasons(statuses, unanswerable),
|
||||
},
|
||||
)
|
||||
|
||||
filtered = filter_events(all_events, issue=issue, pr=pr, session=session)
|
||||
ordered = sort_events(filtered)
|
||||
page = paginate(ordered, limit=limit, offset=offset)
|
||||
|
||||
return TimelineSnapshot(
|
||||
schema_version=TIMELINE_SCHEMA_VERSION,
|
||||
remote=remote,
|
||||
org=org,
|
||||
repo=repo,
|
||||
filters=filters,
|
||||
page=page,
|
||||
sources=tuple(statuses),
|
||||
)
|
||||
|
||||
|
||||
def snapshot_to_dict(snapshot: TimelineSnapshot) -> dict[str, Any]:
|
||||
return snapshot.to_dict()
|
||||
Reference in New Issue
Block a user