The continuation-override comment said "step-end/continuation listeners". With no
agent/step-end emit, the surviving step-boundary listener is the durable step/end
SESSION event, so spell it "step/end session-event/continuation listeners" to avoid
implying a removed agent/* mirror. Comment-only; no behavior change.
Second-round Codex review of the PR-A taxonomy change found four issues, all
verified against the code:
- The /goal regression guard asserted only that the steered content reached
requests[1], which passes even with the hasSteering override (loop.ts) disabled:
leftover steering is re-enqueued as a next-turn queued message and also lands in
requests[1], one turn later. The guard now asserts the same-turn shape — ONE
turn, TWO steps, a steering/message recorded before step 2 — which is the
mechanism the override drives. Proven to fail red with the override disabled.
- The event-domain-semantics RFC's consequence list still described the pre-fix
behavior (step marked open AFTER step/start, so no step/end owed). It now states
the shipped behavior: the loop marks the step open BEFORE the append, so a
throwing step/start listener gets a balancing step/end via closeStep().
- architecture.md's loop pseudocode said only continuation listeners force
continuation; step/end session-event listeners (the /goal pattern) do too.
- The agent/turn-end JSDoc listed a `rejected` TurnEndReason that does not exist on
this branch (it belongs to the later interception work). Removed it and
regenerated the cordis catalog; `interrupted` (a real variant) stays.
Codex review of PR-A found three blockers:
- A throwing step/start session-event listener left an unbalanced log
(turn/start → step/start → turn/end with no step/end), which the invariants
oracle rejects — masked because that rejection was itself contained as a
throwing turn/end listener. Fix the root cause in the loop: mark the step open
BEFORE appending step/start (Session.append pushes before notifying), so the
outer catch's closeStep() appends the balancing step/end. The test now asserts
the balanced outcome (stepEnd:1, step/end before turn/end); proven load-bearing
(revert the reorder → the test goes red with stepEnd:0).
- Reintroduce the /goal-pattern guard deleted in the prior commit, migrated to a
step/end session-event listener (the surviving step-boundary hook point), with
a no-tools first step so it exercises the hasSteering continuation override.
- Update packages/core/agent/README.md: step boundaries are no longer agent/*
emits.
Pin the three-domain rule (session = durable fact log, agent = live runtime
surface, tools = registry/exec): a durable replayable fact is a SessionEvent; a
live interception or transient/live-object signal is an agent/tools Cordis
event. A boundary that is both is mirrored as an agent/* emit ONLY where a live
consumer needs the Agent handle.
Apply it to the boundary twins: drop agent/step-start and agent/step-end (no
production consumer needs the live Agent at a step boundary — consumers read the
durable step/start/step/end session events). Keep agent/turn-start/turn-end (the
stdio UI labels output by agent.id). Tests that observed step boundaries via the
removed emits now observe the durable session events; the pinned behavior is
unchanged.
Conservative subset of the proposed "remove boundary mirror events"
simplification; foundation for the Hooks subsystem's canonical event surface.
The Events intro carried "The harness declares N events across M scopes."
and the Services intro "The N `ctx.<key>` services the harness provides."
Both embed counts the generator recomputes from source, so every branch
that adds an event or service rewrites that one line — a guaranteed merge
conflict against any sibling branch that also touched the catalog, for
prose that adds nothing a reader can't get by scanning the page.
Remove the count clauses from the generator's render() and regenerate the
catalog. The freshness gate (verify-cordis-catalog) stays green.
Add a "Run the CI gates locally BEFORE marking a PR ready" subsection: the
CI-equivalent local command line, and the rule that `pnpm run test:coverage`
(per-file 100%, CI-enforced) — not `pnpm run test` — is the gating test command,
alongside hygiene/snapshot/doc-sync. A green `test` run can still fail CI on an
uncovered line, which is usually dead code the gate is correctly flagging.
The registry's validateArgs rejects a bad `status` enum before execute runs, so
the in-body re-check (`status !== 'pending' && …` → throw) was unreachable dead
code — line 70 was uncovered, failing the per-file 100% coverage gate. Narrow
the registry-guaranteed value with `status as TodoItem['status']` instead of
re-validating it, mirroring tool-bash (which only checks what the DSL can't
express). The malformed-status test still passes — it exercises the registry's
rejection, the actual path. Coverage back to 100%.
Codex confirmation review: the trim-the-stored-content fix had no test that
would fail if it regressed (existing assertions use already-trimmed todos).
Add a focused test asserting " plan the work " appends content "plan the
work". Verified it fails red against the pre-fix code.
Record the `todo-plan` snapshot scenario: a real prompt drives the model to call
todo_write, and the golden captures the resulting `plan` sessionUpdate (three
entries, priority synthesized as medium, status 1:1) plus the persisted
todo/write event. Registered in SCENARIOS; replays deterministically keyless.
Add a with-key coding-agent e2e that verifies the WORLD — a real model call to
todo_write lands a todo/write event whose snapshot is a valid, one-in-progress
list — not the agent's self-report. Wire tool-todo into the e2e harness.
toTodoList dedupes and length-checks on the trimmed content but stored the raw
item.content, so a todo with leading/trailing whitespace was deduped by its
trimmed form yet persisted untrimmed — the stored value and the uniqueness key
could differ. Store the trimmed content so the persisted list matches what was
validated.
Add @deepseek-ai/dsh-tool-todo (a new packages/todo/ group): a model-facing
todo_write(todos: [{content, status}]) tool with whole-list-replace semantics.
Each call appends the full list as a todo/write event to the calling agent's
session log; the current list is the most recent such event (last-write-wins).
Single-owner — a non-agent caller is rejected. Beyond the schema's
type/required/enum checks, execute rejects empty/duplicate content and more than
one in_progress task, narrowing the loosely-typed args into a real TodoItem[].
Both UIs render off the existing session/event: the stdio UI prints a glyphed
checklist; the ACP bridge maps the list to a `plan` sessionUpdate (todosToPlan
synthesizes the priority ACP requires; status maps 1:1). Wired into the
coding-agent, acp-agent, and snapshot example configs with a system-prompt nudge.
Tests: unit (schema, validation, append/replace, no-agent rejection, presentCall,
HMR-safety, Loader export-shape guard), full-loop integration through the agent
loop, the ACP todosToPlan mapping + stream-update arm, the stdio render arm, and
a session/load replay that re-emits the plan. New-group TS wiring added to
tsconfig.base/json/build. RFC + a doc-inventory sweep (architecture, packages
README, AGENTS layout, cookbook group list, example READMEs) ship with it.
The todo-plan ACP snapshot scenario is recorded separately (needs an API key).
Codex Phase 1 review: the event JSDoc described Phase 2 consumers (the
todo_write tool, stdio printing, ACP plan mapping) as current state, and put an
@mode tag on a SessionEventMap member. @mode is for first-class Cordis
`interface Events` entries the catalog generator reads — this event rides the
existing session/event emit and has no catalog row, so the tag was wrong.
Trim the JSDoc to the event's own contract (snapshot data shape,
last-write-wins, not-a-surface-event) and drop @mode; phrase TodoItem in terms
of its own purpose rather than a not-yet-present tool.
Add the TodoItem type and a todo/write SessionEventMap variant carrying the
whole todo list as a snapshot (last-write-wins on replay). It is NOT a
SurfaceEventType: it produces no LLM message and never reaches
deriveMessages(), so it carries no surfaceOp and stays off the surface — it is
durable, replayable UI state that rides the existing session/event emit.
Tests cover the snapshot-clone-on-append contract, last-write-wins, the
not-on-surface guarantee, and a seeded replay round-trip. Docs: session.md
gains the TodoItem type-equiv block + the event member; core.md's variant count
goes to twelve; the type-equiv manifest gains TodoItem.
The per-file 100% coverage gate flagged surface.ts line 46 — the
branch where a surface-eligible event type carries no surfaceOp marker
(isSurfaceEvent returns false). Exercise both guards directly: the
type-only eligibility check, the positive narrowing path, a
non-eligible type, and the markerless-but-eligible branch.
Type-aware ESLint loads every package tsconfig through the project
service and peaks at ~3.4GB RSS. The default V8 old-space ceiling
(~2GB) OOMs it (FATAL ERROR: Ineffective mark-compacts near heap
limit, exit 134) on both node 24 and 26. Set NODE_OPTIONS with an
8GB ceiling for the Lint step, comfortably above the peak.
P1: both merge parents shipped SCHEMA_VERSION=3 for different layouts (surface
columns vs seed_length), so an on-disk 3 was ambiguous and wrongly accepted.
Bump to 4 (merged layout) so the version check rejects both sibling v3s.
P2: a surface-eligible event with no surfaceOp lands in the log but vanishes
from deriveMessages() (surface is the sole derivation path). The typed append
overload enforces the marker only when the type arg is a literal; it collapses
to optional when widened to the union (a caller iterating raw events). Guard at
runtime in both append() and the seed constructor — no backward-compat for
surface-less logs. Shared seed fixtures carry surfaceOp explicitly and the
appendLog helper forwards it verbatim (no synthesized default). Exports
isSurfaceEligibleType. Regression tests for all three, each verified to fail
on the unfixed code.
Gates: typecheck, test (1115), snapshot (14), doc-sync, lint, build, hygiene green.
Reconciles the session-surface work (surfaceOp/sourceEventSeqs provenance as
the sole derivation path) with master's worktree-subagent series (fork-seed
boundary + out-of-process subagent backends).
Semantic reconciliations beyond the textual auto-merge:
- SQLite SCHEMA_VERSION: both sides bumped 2->3. Merged to a single v3 carrying
BOTH column families — master's seed_length on `sessions` and surface's
source_event_seqs/surface_op on `events`. writeRow + both INSERT sites bind
the full set; the schema doc lists all three added columns as the v2->v3 gap.
- agent-loop runStep request: master's `sessionId: session.id` and surface's
per-append surfaceOp/sourceEventSeqs coexist (different regions).
- Fork seed + surface: a fork seeds the child from the parent's LIVE events,
which now carry surfaceOp, so the child's surface rebuilds correctly. Verified
end-to-end — the subagent-fork replay recalls the inherited "SAFFRON" codeword
through the seeded prefix.
- Subagent snapshot fixtures (recorded pre-surface) re-enriched via KEYLESS
deterministic replay: only surfaceOp/sourceEventSeqs added onto existing
recorded lines (matched by seq), no recorded value changed. Not re-recorded
against the live API.
Gates: typecheck, test (1112), test:snapshot (14), doc-sync, lint, build,
hygiene all green.