diff --git a/AGENTS.md b/AGENTS.md index 3fcdcceeef..d53750222e 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -73,6 +73,19 @@ packages/ Harness packages, grouped by role at packages///. bash/ abstract bash executor seam (ctx.bash) — interface only bash-local/ local-subprocess BashExecutor implementation tool-bash/ model-facing bash/bash_output/bash_kill tool schemas + compact/ compaction capability family + compact/ abstract compaction seam (ctx.compact); backend + tool deferred + subagent/ subagent capability family + subagent/ provider-registry seam (ctx.subagents) + subagent-inprocess/ shared in-process run driver (library, registers nothing) + subagent-spawn/ in-process fresh-child backend + subagent-fork/ in-process backend seeded from the parent's completed-turn prefix + subagent-acp/ out-of-process child over ACP + tool-subagent/ model-facing delegation tool over ctx.subagents + todo/ todo/planning capability family + tool-todo/ model-facing todo_write tool: writes the whole task list to + the session log (todo/write), rendered as a stdio checklist / + ACP plan session-persistence/ persistence capability family session-persistence/ durable persistence seam + write coordinator session-persistence-jsonl/ JSONL-sidecar backend @@ -91,20 +104,23 @@ packages/ Harness packages, grouped by role at packages///. feeds stdin lines to the agent (shared by the demos) llm-replay/ record/replay adapter: short-circuits llm/stream from a recorded session JSONL (keyless snapshot tests) + subagent-mock/ scripted SubagentProvider for deterministic seam/tool tests util/ low-level zero-dependency utilities shared across groups brand/ type-only Branded nominal-typing primitive (no runtime code, no harness deps; owns the brand for cross-boundary ids) examples/ Runnable demos (not workspaces; see examples/AGENTS.md). Each is a THIN leaf cordis.yml: it picks the swappable backends (an LLM adapter, - a bash executor) and loads ONE app package (dsh-stdio-agent or - dsh-acp-agent), which bundles the agent-core spine + front-door - cluster + boot glue (a bin). No start.ts. echo-agent = mock model + - echo tool on dsh-stdio-agent (pnpm run demo:echo, no key). - coding-agent = the real thing: DeepSeek V4 + bash tools on the same - app (pnpm run demo:coding, needs DEEPSEEK_API_KEY). acp-agent = the - coding agent as an ACP server on dsh-acp-agent (pnpm run demo:acp, - needs DEEPSEEK_API_KEY). cordis.snapshot.yml = the acp leaf with - llm-replay for keyless snapshot replay. + a bash executor), loads ONE app package (dsh-stdio-agent or + dsh-acp-agent), and may add optional product tools or demo-local + teaching plugins. The app package bundles the agent-core spine + + front-door cluster + boot glue (a bin). No start.ts. echo-agent = + mock model + echo tool on dsh-stdio-agent (pnpm run demo:echo, no + key). coding-agent = the real thing: DeepSeek V4 + bash tools + + subagent + todo_write on the same app (pnpm run demo:coding, needs + DEEPSEEK_API_KEY). acp-agent = the coding agent as an ACP server on + dsh-acp-agent (pnpm run demo:acp, needs DEEPSEEK_API_KEY). + cordis.snapshot.yml = the acp leaf with llm-replay for keyless + snapshot replay. docs/ architecture.md — the design doc. module-graph.md — generated inter-package dependency graph (Mermaid; `pnpm run gen-module-graph`). rfc/ — design decisions and proposals, one kind of doc grouped by @@ -172,6 +188,30 @@ pnpm run demo:acp # run examples/acp-agent — the coding agent as an ACP # drive it from Zed or another ACP client) ``` +### Run the CI gates locally BEFORE marking a PR ready + +CI is the backstop, not the first place a gate runs. Before you open a non-draft PR or move one from draft to ready, run the same gates CI runs, on your own tree, and confirm they pass — do not lean on CI (or a Codex pass) to discover a red gate you could have caught locally. The CI-equivalent local run is: + +```sh +set -euo pipefail +pnpm run typecheck +pnpm run lint +pnpm run test:coverage +pnpm run test:snapshot +pnpm run doc-sync +pnpm run verify-module-graph +pnpm run build +pnpm run hygiene +out=$(printf 'echo ci smoke\n' | pnpm run demo:echo 2>&1) +printf '%s\n' "$out" | grep -q '\[tool call\] echo({"text":"ci smoke"})' +printf '%s\n' "$out" | grep -q '\[tool result\] ECHO: CI SMOKE' +ls .sessions/_no-cwd/main-session-*.jsonl >/dev/null +rm -rf .sessions +pnpm exec vitest run --config vitest.e2e.config.ts packages/ui/stdio-agent/tests/built-bin.e2e.ts packages/ui/acp-agent/tests/built-bin.e2e.ts +``` + +**`pnpm run test:coverage`, NOT `pnpm run test`, is the gating test command.** `pnpm run test` runs `vitest run` with no coverage; CI's node job runs `test:coverage`, which enforces a **per-file 100%** threshold on `packages/*/*/src`. A suite that is green under `test` can still fail CI on an uncovered line — and that uncovered line is often *dead code* the 100% gate is correctly flagging for deletion (see [§ Defensive patterns](#defensive-patterns-hard-won) "Line coverage is not behavior coverage"), not a missing test to bolt on. `hygiene` (knip + publint + workspace constraints + NodeNext types) and `test:snapshot` (keyless ACP replay) are likewise CI gates that `test` alone does not cover. When you rely on a Codex convergence pass for sign-off, check WHICH commands it ran: a pass that ran `test` but not `test:coverage`/`hygiene`/`doc-sync` has not exercised those gates. + ## Secrets / .env Real-API e2e tests (`pnpm run test:e2e`) read `DEEPSEEK_API_KEY` (and optionally `DEEPSEEK_BASE_URL`) from the environment, or from a gitignored `.env` at the repo root loaded via Node's native `process.loadEnvFile()`: @@ -237,15 +277,15 @@ In the **core** packages (`packages/llm/llm`, `packages/core/tools`, `packages/c Verbose documentation is fine **as long as docs and code stay strictly in sync**. Out-of-sync docs are worse than no docs. **When you change code, update its docs in the SAME change** — grep the package README and the module/JSDoc comments for the old behavior (config keys, defaults, error codes, wire field names, event names) and fix every hit. CI runs `pnpm run doc-sync` (`doc-typecheck` + `verify-cordis-catalog` + `verify-md-wrap` + `verify-md-links` + `verify-doc-refs` + `verify-package-paths` + `verify-rfc-classification` + `verify-type-equiv`), which typechecks every fenced `ts` block in `README.md`, `docs/**/*.md`, and `packages/*/*.md`, regenerates the cordis events/services catalog from source and fails if the committed copy is stale, asserts no hard-wrapped prose paragraphs, checks that every relative Markdown cross-link resolves, checks that every `docs/*.md` path cited in a source comment resolves, checks that every `packages/` reference naming a real package resolves, checks that every RFC is filed under a valid class folder and listed in its index, and checks that every ` ```ts type-equiv ` doc block still matches its source type — across those files plus `AGENTS.md` / `packages/AGENTS.md` — but that scope does NOT catch prose drift in `AGENTS.md` / `packages/AGENTS.md` / `packages/README.md` (config keys, defaults, error codes), so keeping those in sync remains on the author. Every module has a module-level doc comment explaining its role. Every exported class, interface, type, function, and non-obvious method has a JSDoc that explains semantics (not just the name) — contracts (what events fire when), disposal behavior, error behavior, and extension intent. Internal helpers get docs only where non-obvious. Prefer one-liners when one line suffices. -**Tag every new event with `@mode`.** The cordis events/services catalog ([docs/cordis-catalog/events-and-services.md](docs/cordis-catalog/events-and-services.md)) is GENERATED from source by `scripts/gen-cordis-catalog.ts` — never hand-edit it; run `pnpm run gen-cordis-catalog` and commit the result. When you add an event to an `interface Events` block, its JSDoc MUST carry a `@mode emit|waterfall|parallel` tag (the generator hard-errors without it): use `waterfall` when the signature ends with a `next: () => …` parameter (the listener transforms or vetoes via `next()`), `parallel` when the loop awaits a fan-out with no veto (e.g. an awaited `Promise | void` checkpoint like `session/flush`), and `emit` for plain fire-and-forget notifications. The generator also cross-checks the tag against the signature where the shape is conclusive (a trailing `next` ⇒ waterfall) and hard-errors on a contradiction. Write the rest of the event's JSDoc to stand alone — it is the catalog entry's prose. +**Tag every new event with `@mode`.** The cordis events/services catalog ([docs/cordis-catalog/events-and-services.md](docs/cordis-catalog/events-and-services.md)) is GENERATED from source by `scripts/gen-cordis-catalog.ts` — never hand-edit it; run `pnpm run gen-cordis-catalog` and commit the result. When you add an event to an `interface Events` block, its JSDoc MUST carry a `@mode emit|waterfall|parallel|serial` tag (the generator hard-errors without it): use `waterfall` when the signature ends with a `next: () => …` parameter (the listener transforms or vetoes via `next()`), `parallel` when the loop awaits a fan-out and must run every listener (e.g. an awaited `Promise | void` checkpoint like `session/flush`), `serial` when the loop awaits listeners in registration order and should isolate side effects (e.g. an ordered surface-mutation checkpoint like `agent/pre-step`; Cordis stops early if a listener returns a bail value, so `void` serial listeners must not return a semantic veto), and `emit` for plain fire-and-forget notifications. The generator also cross-checks the tag against the signature where the shape is conclusive (a trailing `next` ⇒ waterfall) and hard-errors on a contradiction. Write the rest of the event's JSDoc to stand alone — it is the catalog entry's prose. **The core-data-structures catalog is a maintained surface, not a write-once artifact.** [docs/core-data-structures/](docs/core-data-structures/core.md) catalogs the spine vocabulary (core.md) and the per-seam types (sub-pages). When a change adds, removes, or reshapes a type the catalog documents — a new `…Map` variant, a new content-block or session-event type, a field on `GenerateOptions`/`Agent`/`ToolDefinition`/a bash type, or a whole new core/seam type — update the catalog in the SAME change: edit the prose, and for a pasted ` ```ts type-equiv ` block, re-copy it verbatim and keep `scripts/type-equiv.manifest.json` 1:1 with the blocks. The `verify-type-equiv` gate catches a *drifted paste* of an already-documented type, but it canNOT tell you a brand-new core type was never documented — that judgment is on the author and the reviewer. The definition of "core" (the spine-vs-seam line) is in [core.md § What counts as "core"](docs/core-data-structures/core.md#what-counts-as-core); a genuinely spine-level new type belongs in core.md, a new capability's vocabulary on a sub-page. See [development.md](docs/development.md#documenting-types-verbatim-ts-type-equiv) for the `ts type-equiv` mechanics. **Document the CURRENT state — the "what" and "why" — never the PROCESS or HISTORY of how it got there.** A comment, JSDoc, or doc paragraph describes what the code *is* and why it is that way, as if it had always been so. Do NOT narrate the change that produced it: no "previously X, now Y", "changed from", "used to", "this replaces", "the old map", "renamed", "moved here", "as of this PR", or "(was …)". **In particular, NEVER name the change unit a reader cannot see — the PR, commit, or stack position that introduced the code — in a comment, JSDoc, OR a test name/description.** A `// (PR D's per-agent teardown)` aside, a `* Tests for the cancel primitive (PR C).` module doc, or an `it('… identity no longer matters')` title that only makes sense relative to a prior design are all the same violation: the reader of the current tree has no "PR D" or "old design" to anchor against, and the reference rots the moment the stack merges. Name the *mechanism* (`the session's AgentHandle teardown`), not the PR. Such phrasing rots the instant the next change lands, and a reader of the current code does not need the diff narrated in prose — that belongs in the commit message, the PR description, or an RFC (the durable home for "why we moved away from X"). Write "the owner token lives on the task in the executor" — not "ownership *now* lives on the executor instead of a plugin-local map". When a contrast genuinely aids understanding (a non-obvious choice between live alternatives), frame it against the alternative as a standing fact ("stored on the executor, NOT the tool plugin, so it survives an HMR reload"), not against the codebase's past. The same rule governs review-fix commits: the *commit message* records what the review caught; the *code comment* it touches states only the resulting truth. RFCs (`docs/rfc/`, grouped into `proposed/` / `implemented/` / `rejected/`) record the *why* behind choices a future reader would otherwise re-litigate (the vendoring policy, event-sourcing, the schema DSL are the existing examples). A PR that introduces such a decision — a new third-party runtime dependency over the vendoring default, a cross-package contract, a security/isolation model, a deviation from a documented architecture rule — writes the RFC in `implemented/` **in the same PR**, and links it from the relevant code. A proposal for future work not yet built goes in `proposed/`. A PR whose changes are mechanical, self-evident, or already covered by an existing RFC needs none — do not manufacture an RFC for a routine change. When unsure, the test is: would a competent maintainer six months from now ask "why was it done this way?" and be unable to answer from the code alone? If yes, write it. See [docs/rfc/README.md](docs/rfc/README.md) for the naming scheme and [docs/AGENTS.md](docs/AGENTS.md) for the cross-link convention. -**Markdown is not hard-wrapped**: write one line per paragraph and let the editor soft-wrap. Hard line breaks mid-paragraph make docs harder to edit and diff — a one-word change reflows and re-diffs the whole paragraph. This applies to prose only: leave fenced code blocks, tables, and list structure intact (a wrapped list item folds to one line per bullet). Code comments / JSDoc are exempt — they stay under the linter's column limit. `pnpm run verify-md-wrap` (part of `doc-sync`) enforces this across `README.md`, `docs/**/*.md`, `packages/*/*.md`, and `AGENTS.md` / `packages/AGENTS.md`; `pnpm run verify-md-links` (also part of `doc-sync`) checks that every relative cross-link in those files resolves. +**Markdown is not hard-wrapped**: write one line per paragraph and let the editor soft-wrap. Hard line breaks mid-paragraph make docs harder to edit and diff — a one-word change reflows and re-diffs the whole paragraph. This applies to prose only: leave fenced code blocks, tables, and list structure intact (a wrapped list item folds to one line per bullet). Code comments / JSDoc are exempt — they stay under the linter's column limit. `pnpm run verify-md-wrap` (part of `doc-sync`) enforces this across `README.md`, `docs/**/*.md`, `packages/*/*.md`, and `AGENTS.md` / `packages/AGENTS.md`; `pnpm run verify-md-links` (also part of `doc-sync`) checks that every relative cross-link in those files plus `examples/**/*.md` and `.agents/skills/**/*.md` resolves. -**Editing these instructions**: `AGENTS.md` is the real file; `CLAUDE.md` is a symlink to it (at the repo root and in `packages/`). Always edit `AGENTS.md` — never write through the `CLAUDE.md` symlink or replace it with a regular file. +**Editing these instructions**: `AGENTS.md` is the real file; `CLAUDE.md` is a symlink to it (at the repo root and in `packages/` / `examples/`). Always edit `AGENTS.md` — never write through the `CLAUDE.md` symlink or replace it with a regular file. ## Vendoring Policy diff --git a/docs/architecture.md b/docs/architecture.md index 34bb0a1280..3539a0dae3 100644 --- a/docs/architecture.md +++ b/docs/architecture.md @@ -27,6 +27,7 @@ For a catalog of the **data structures** this architecture moves around — the │ @deepseek-ai/dsh-fs-local (filesystem impl) │ │ @deepseek-ai/dsh-file-context (filesystem policy gate) │ │ @deepseek-ai/dsh-tool-fs (filesystem tools+executor)│ +│ @deepseek-ai/dsh-subagent-* (subagent providers) │ │ @deepseek-ai/dsh-session-persistence-jsonl (persistence impl)│ ├─────────────────────────────────────────────────────────────┤ │ @deepseek-ai/dsh-agent (vocabulary + registry) │ @@ -37,6 +38,8 @@ For a catalog of the **data structures** this architecture moves around — the │ @deepseek-ai/dsh-llm (abstract model service) │ │ @deepseek-ai/dsh-bash (abstract bash executor) │ │ @deepseek-ai/dsh-fs (filesystem provider seam) │ +│ @deepseek-ai/dsh-compact (abstract compaction seam) │ +│ @deepseek-ai/dsh-subagent (provider registry seam) │ ├─────────────────────────────────────────────────────────────┤ │ vendor/: cordis, loader, include, group, timer, hmr, │ │ logger-console, cosmokit, schemastery │ @@ -60,6 +63,7 @@ Dependency rule: **extension** plugins depend on interface packages, never on `d | `ctx.bash` | `BashExecutor` (abstract) | dsh-bash | bash execution seam: foreground runs + background tasks | | `ctx.fs` | `FileSystem` (abstract) | dsh-fs | filesystem provider seam: path resolution, stat, text read/stream, atomic writes/edits (optional version guard); owns the `fs/*` policy events | | `ctx.compact` | `CompactService` (abstract) | dsh-compact | compaction seam: decide when history is too large, summarize an older range into a single surface node | +| `ctx.subagents` | `SubagentService` | dsh-subagent | named provider registry for delegating a task to child agents | All registrations (`registerAdapter`, `section`, `tools`, `register`, …) go through `ctx.effect()` and return disposers, so plugin hot-reload (vendored HMR) and fiber disposal clean up automatically. @@ -94,11 +98,11 @@ A `Session` is an append-only log of typed `SessionEvent`s — the single source - `user/message` → user message - `assistant/message` → assistant message (raw `assistant/chunk` events are replay/UI data and are skipped in derivation; an empty-content `assistant/message`, which exists only to host a max-tokens step's `usage`, is skipped too) - `tool/result` → user message carrying a `tool-result` block -- `context/message`, `steering/message` → user-role messages wrapped in a tagged envelope (``) at their chronological position — the "system-reminder" pattern; models distinguish them from real user prompts by the envelope. **TODO(review)**: the real adapters now exist (the original precondition); the envelope still wants a deliberate review against live model behavior (`TODO(review)` in dsh-session). +- `context/message`, `steering/message` → user-role messages wrapped in a tagged envelope (``) at their chronological position — the "system-reminder" pattern; models distinguish them from real user prompts by the envelope. Live-adapter review has validated the tagged-envelope rendering against current DeepSeek behavior; provider-specific mismatches belong in that adapter. Replay/fork = `ctx.sessions.create(id, { seed: seedEvents })`. Trace/telemetry = listen to `session/event`. -**Durability seam**: `session/event` is a synchronous notification; persistence plugins buffer (write-behind) and drain at the awaited `session/flush` checkpoint the loop fires at every turn end. The durable backend is a real **capability seam**: the abstract `SessionPersistence` service (`dsh-session-persistence`, `ctx.sessionPersistence`) defines create/append/load/list over the existing `SessionEvent` (no parallel persisted type), and `dsh-session-persistence-jsonl` is the first implementation — an append-only JSONL log per session with crash-safe atomic writes, crash recovery that PRESERVES an interrupted turn (closing it with a synthetic `turn/end {interrupted}` rather than truncating — a turn can be huge), and a read/replay path. Session metadata (format version, cwd, lineage) travels separately as `SessionHeader`, attached to a `Session` via `session.header`. Resuming a persisted session into a live agent is `ctx.agents.resume({ resumeSessionId })`. A second backend, `dsh-session-persistence-sqlite` (`node:sqlite`, one row per `SessionEvent` — the row shape `(session_id, seq, type, time, data)` maps 1:1 onto it), passes the same `runPersistenceContract` suite, proving the seam is genuinely backend-agnostic. +**Durability seam**: `session/event` is a synchronous notification; persistence plugins buffer (write-behind) and drain at the awaited `session/flush` checkpoint the loop fires at every turn end. The durable backend is a real **capability seam**: the abstract `SessionPersistence` service (`dsh-session-persistence`, `ctx.sessionPersistence`) defines create/append/load/list over the existing `SessionEvent` (no parallel persisted type), and `dsh-session-persistence-jsonl` is the first implementation — an append-only JSONL log per session with crash-safe atomic writes, crash recovery that PRESERVES an interrupted turn (closing it with a synthetic `turn/end {interrupted}` rather than truncating — a turn can be huge), and a read/replay path. Session metadata (format version, cwd, lineage, seed boundary) travels separately as `SessionHeader`, attached to a `Session` via `session.header`. Resuming a persisted session into a live agent is `ctx.agents.resume({ resumeSessionId })`. A second backend, `dsh-session-persistence-sqlite` (`node:sqlite`, one row per `SessionEvent` — the row shape `(session_id, seq, type, time, data, source_event_seqs, surface_op)` maps 1:1 onto it), passes the same `runPersistenceContract` suite, proving the seam is genuinely backend-agnostic. ## Prompt assembly (dsh-system-prompt) @@ -141,10 +145,11 @@ forever: drain queued → 'turn/start' → session('user/message'…) → emit agent/turn-start STEP loop: drain steering (late steering from previous step's listeners) - session('step/start'); emit agent/step-start assembly = ctx.systemPrompt.assemble() ⟵ waterfall system-prompt/assemble + await ctx.serial('agent/pre-step') ⟵ surface mutation (compaction) OUTSIDE the step + session('step/start'); emit agent/step-start req = {model, system, tools, messages: session.deriveMessages(), signal} - req = waterfall agent/request ⟵ hooks, compaction, model switch + req = waterfall agent/request ⟵ hooks, model switch stream ctx.llm.stream(req) ⟵ waterfall llm/stream (raw chunks) session('assistant/chunk'); emit agent/stream-chunk if assembler.finish is error/aborted: throw ⟵ adapter's in-band error path → @@ -154,6 +159,7 @@ forever: session('assistant/message' {content, usage?}) log records what tool dispatch uses each tool-call (sequential, abort-checked between calls): session('tool/call'); ctx.tools.execute() ⟵ waterfall tools/execute + tool execution may append tool-owned session events, e.g. `todo/write` session('tool/result') drain steering → session('steering/message'); emit agent/steering emit agent/step-end @@ -200,16 +206,16 @@ Every MVP feature (including the TODO-marked ones), with the mechanism that impl | `/loop` | on `agent/turn-end`, `send()` the next iteration; or force-continue | | Dynamic workflow | orchestrator plugin on `agent/turn-end` / `agent/step-end` driving `send`/`steer` (+ sub-agents later) | | Queued + steering messages | core `Agent.send()` / `Agent.steer()` | -| Context compaction (auto + manual) | the `ctx.compact` seam ([dsh-compact](../packages/compact/compact)): a backend summarizes an older surface range into a single `user/message` `replace` op, bracketed by log-only `compact/*` events; auto = check token pressure at turn boundaries, manual = a `/compact` tool. See the [compaction capability-seam RFC](rfc/proposed/feature/2026-06-18-compaction-capability-seam.md) | +| Context compaction (auto + manual) | the `dsh-compact` seam (`ctx.compact`) + a backend (`dsh-compact-basic`) on the serial `agent/pre-step` seam: a backend summarizes an older surface range into a single `user/message` `replace` op, bracketed by log-only `compact/*` events; auto = check token pressure before each step — runaway-turn survival, manual = a (deferred) `/compact` tool invoking the same `ctx.compact` routine. See the [compaction capability-seam RFC](rfc/implemented/feature/2026-06-18-compaction-capability-seam.md) | | System prompt configurability | `ctx.systemPrompt.section()` with ordering | | AGENTS.md (root) | a section provider reading the file | | AGENTS.md (subdir, on-touch) + file-change notices | `agent.inject()` from a watcher / tool-result listener | -| Built-in tools (Read/Write/Edit/Bash/…) | `ctx.tools.register()`; schemas flow into the assembly automatically. **Bash: implemented** — `dsh-bash` (seam) + `dsh-bash-local` (subprocesses) + `dsh-tool-bash` (`bash`/`bash_output`/`bash_kill`, incl. background tasks) | +| Built-in tools (Read/Write/Edit/Bash/…) | `ctx.tools.register()`; schemas flow into the assembly automatically. **Bash: implemented** — `dsh-bash` (seam) + `dsh-bash-local` (subprocesses) + `dsh-tool-bash` (`bash`/`bash_output`/`bash_kill`, incl. background tasks). **`todo_write`: implemented** — `dsh-tool-todo` writes the whole task list to the session log (`todo/write`), rendered as a stdio checklist / ACP `plan` | | ToolSearch / progressive disclosure | wrap `agent/request`, filter `req.tools` | | Tool sandbox (landlock / sandbox-exec) | wrap `tools/execute`, or implement a sandboxing `BashExecutor` (the dsh-bash seam) | | Permission system / AskUserQuestion | wrap `tools/execute` (veto or ask); register an ask tool | | Plan mode | wrap `tools/execute` (deny writes) + `agent/request` (inject mode prompt) | -| Sub-agents (spawn / fork / steer) | TODO seam on `AgentLoop.create()`; fork = seed Session with parent events; `steer()` on the child handle | +| Sub-agent delegation | Implemented as the `ctx.subagents` provider-registry seam: `dsh-subagent-spawn` starts a fresh in-process child, `dsh-subagent-fork` seeds a child from the parent's completed-turn prefix, `dsh-subagent-acp` drives an out-of-process child over ACP, and `dsh-tool-subagent` exposes one configured provider to the model | | MCP | one plugin per server: discover tools → `ctx.tools.register()` | | Skills | section + tool registration; `inject()` skill content on invocation | | Memory | section provider + tool | @@ -227,7 +233,7 @@ Code skeletons for the three plugin shapes (tool, hook/permission-gate, UI) and Tracked here deliberately — each is designed-for but not implemented: -- **Sub-agent spawn/fork semantics** (seam: `AgentLoop.create()`); inter-agent channels beyond `send`/`steer`/events. -- **Compaction implementation** (auto thresholds, summarization prompts) on the `agent/request` seam, with its session-event types added by declaration merging. +- **Inter-agent channels beyond delegation** (shared state, streaming child output, background/poll semantics) remain out of scope for the current `ctx.subagents` seam. +- **Compaction** — the `dsh-compact` seam (`ctx.compact`) and the `dsh-compact-basic` backend exist (auto thresholds, summarization on the serial `agent/pre-step` seam, `compact/*` session events via declaration merging). The model-facing `/compact` consumer tool is still deferred. See [the compaction capability-seam RFC](rfc/implemented/feature/2026-06-18-compaction-capability-seam.md). - **Parallel tool execution** (concurrency-safety hints on ToolDefinition). - **Session branching/tree** (pi-style entry tree) if needed beyond seed-based forking. diff --git a/docs/cookbook/adding-a-package.md b/docs/cookbook/adding-a-package.md index 3c9cdb07d5..ef5f48d624 100644 --- a/docs/cookbook/adding-a-package.md +++ b/docs/cookbook/adding-a-package.md @@ -16,7 +16,7 @@ packages/// README.md # service API, events, extension points, design notes ``` -Choose an existing group when one matches the package's role (`core`, `llm`, `bash`, `session-persistence`, `ui`, `util`, or `support`). A new group is allowed, but it is a pure container: no `package.json`, no source files, and packages still sit exactly one level below it. +Choose an existing group when one matches the package's role (`core`, `llm`, `bash`, `compact`, `subagent`, `todo`, `session-persistence`, `ui`, `util`, or `support`). A new group is allowed, but it is a pure container: no `package.json`, no source files, and packages still sit exactly one level below it. package.json invariants (enforced by `pnpm run constraints` / `scripts/check-workspace-constraints.ts`): `private: true`, `version: 0.0.1`, `type: module`, `main: "lib/index.js"`, `types: "lib/types/index.d.ts"`, `exports["."].types: "./lib/types/index.d.ts"`, `exports["."].default: "./lib/index.js"`, `cordis` in BOTH peerDependencies and devDependencies (same range). Mirror every dsh peer dependency in devDependencies. `schemastery` goes in `dependencies` (it is a runtime validator), matching agent-loop. The `files` list is precise: `lib/index.js`, `lib/types/**/*.d.ts`, `lib/types/**/*.d.ts.map`, and `src`; do not publish `lib/types` JS or JS-map intermediates or stale root declaration files. CLI app packages with a package `bin` include `lib/bin.js` immediately after `lib/index.js` in `files`. diff --git a/docs/cordis-catalog/events-and-services.md b/docs/cordis-catalog/events-and-services.md index e95687d78c..7c1701af59 100644 --- a/docs/cordis-catalog/events-and-services.md +++ b/docs/cordis-catalog/events-and-services.md @@ -11,7 +11,7 @@ The **harness tier** below (the `@deepseek-ai/dsh-*` packages) is the vocabulary ## Events -Dispatch modes: **emit** (fire-and-forget), **waterfall** (each listener gets `next()` and may transform or veto — see [waterfall semantics](../architecture.md#cordis-waterfall-semantics-important)), **parallel** (awaited fan-out, no veto). The harness declares 27 events across 7 scopes. +Dispatch modes: **emit** (fire-and-forget), **waterfall** (each listener gets `next()` and may transform or veto — see [waterfall semantics](../architecture.md#cordis-waterfall-semantics-important)), **parallel** (awaited fan-out; all listeners run), **serial** (awaited in registration order until one returns a bail value — anything other than `null`, `false`, or `undefined`). ### `agent/*` @@ -49,7 +49,21 @@ A step or turn errored. The loop reports a failure here (plus the logger) even w Types: [Agent](../core-data-structures/core.md) -Source: [`packages/core/agent/src/types.ts:220`](../../packages/core/agent/src/types.ts) +Source: [`packages/core/agent/src/types.ts:254`](../../packages/core/agent/src/types.ts) + +#### `agent/pre-step` — serial + +Awaited pre-step surface-mutation checkpoint, fired once per step AFTER `turn/start` (and after the prior step closed) but BEFORE this step's `step/start` — so anything a listener appends lands OUTSIDE the step, between `turn/start`/`step/end` and the upcoming `step/start`. `step` is the number of the step about to start. The loop awaits `ctx.serial('agent/pre-step', …)` after assembling the system prompt, then opens the step and derives the request history ONCE from whatever the surface now holds. This is where compaction belongs: it mutates the session surface in place (shadowing an older range with a summary node) with its log-only `compact/*` records cleanly outside any step, and the single subsequent derive reflects the mutation — so there is no double-derive and no listener can see (or be expected to act on) an assembled `messages` array that does not exist yet. + +Serial (awaited in registration order), not a waterfall: a listener mutates the surface as a side effect; there is nothing to transform, but the loop must wait for the mutation to complete before opening the step and deriving. Cordis `serial` bails early if a listener returns a bail value; this event is typed and documented as `void`, so listeners must not return a semantic veto value. `fullSystemPrompt` is the assembled prompt a listener needs to measure pressure (the system prompt counts toward the budget). `signal` cancels any in-flight work a listener starts (e.g. a summarization model call). + +```ts cordis-catalog +'agent/pre-step'(agent: Agent, turn: number, step: number, fullSystemPrompt: string, signal: AbortSignal): Promise | void +``` + +Types: [Agent](../core-data-structures/core.md) + +Source: [`packages/core/agent/src/types.ts:214`](../../packages/core/agent/src/types.ts) #### `agent/queued` — emit @@ -65,7 +79,7 @@ Source: [`packages/core/agent/src/types.ts:156`](../../packages/core/agent/src/t #### `agent/request` — waterfall -Waterfall: mutate the fully-assembled GenerateOptions before the model call (hooks, compaction, model switching, tool filtering, …). Call `next()` to delegate, or return without it to short-circuit. +Waterfall: mutate the fully-assembled GenerateOptions before the model call (hooks, model switching, tool filtering, …). Call `next()` to delegate, or return without it to short-circuit. For surface mutation that must precede history derivation (compaction), use agent/pre-step instead — by the time this fires, `options.messages` is already derived. ```ts cordis-catalog 'agent/request'(agent: Agent, turn: number, step: number, options: GenerateOptions, next: () => Promise): Promise @@ -73,7 +87,7 @@ Waterfall: mutate the fully-assembled GenerateOptions before the model call (hoo Types: [Agent](../core-data-structures/core.md) · [GenerateOptions](../core-data-structures/core.md) -Source: [`packages/core/agent/src/types.ts:189`](../../packages/core/agent/src/types.ts) +Source: [`packages/core/agent/src/types.ts:223`](../../packages/core/agent/src/types.ts) #### `agent/status` — emit @@ -97,7 +111,7 @@ Steering content was injected into a running turn. Types: [Agent](../core-data-structures/core.md) · [ContentBlock](../core-data-structures/core.md) · [MessageSource](../core-data-structures/core.md) -Source: [`packages/core/agent/src/types.ts:214`](../../packages/core/agent/src/types.ts) +Source: [`packages/core/agent/src/types.ts:248`](../../packages/core/agent/src/types.ts) #### `agent/step-end` — emit @@ -121,7 +135,7 @@ Waterfall: post-process the assembled assistant Message before tool dispatch (va Types: [Agent](../core-data-structures/core.md) · [Message](../core-data-structures/core.md) -Source: [`packages/core/agent/src/types.ts:195`](../../packages/core/agent/src/types.ts) +Source: [`packages/core/agent/src/types.ts:229`](../../packages/core/agent/src/types.ts) #### `agent/step-start` — emit @@ -145,7 +159,7 @@ A raw StreamChunk arrived from the model (token-level UI/log feed). Types: [Agent](../core-data-structures/core.md) · [StreamChunk](../core-data-structures/llm-streaming.md) -Source: [`packages/core/agent/src/types.ts:209`](../../packages/core/agent/src/types.ts) +Source: [`packages/core/agent/src/types.ts:243`](../../packages/core/agent/src/types.ts) #### `agent/turn-continuation` — waterfall @@ -157,7 +171,7 @@ Waterfall: override the turn-continuation decision. The default (computed by the Types: [Agent](../core-data-structures/core.md) -Source: [`packages/core/agent/src/types.ts:202`](../../packages/core/agent/src/types.ts) +Source: [`packages/core/agent/src/types.ts:236`](../../packages/core/agent/src/types.ts) #### `agent/turn-end` — emit @@ -245,7 +259,7 @@ A session was created in the store. 'session/created'(session: Session): void ``` -Source: [`packages/core/session/src/index.ts:33`](../../packages/core/session/src/index.ts) +Source: [`packages/core/session/src/index.ts:34`](../../packages/core/session/src/index.ts) #### `session/event` — emit @@ -257,7 +271,7 @@ An event was appended to a session log (sync, fire-and-forget). This is the per- Types: [SessionEvent](../core-data-structures/core.md) -Source: [`packages/core/session/src/index.ts:39`](../../packages/core/session/src/index.ts) +Source: [`packages/core/session/src/index.ts:40`](../../packages/core/session/src/index.ts) #### `session/flush` — parallel @@ -267,7 +281,7 @@ Awaited durability checkpoint. The agent loop awaits `ctx.parallel('session/flus 'session/flush'(session: Session): Promise | void ``` -Source: [`packages/core/session/src/index.ts:48`](../../packages/core/session/src/index.ts) +Source: [`packages/core/session/src/index.ts:49`](../../packages/core/session/src/index.ts) ### `subagent/*` @@ -339,7 +353,7 @@ Source: [`packages/core/tools/src/index.ts:43`](../../packages/core/tools/src/in ## Services -The 11 `ctx.` services the harness provides. An abstract seam (e.g. `ctx.bash`) is implemented by a separate package; the interface is what consumers code against. +The `ctx.` services the harness provides. An abstract seam (e.g. `ctx.bash`) is implemented by a separate package; the interface is what consumers code against. ### `ctx.agentLoop` — `AgentLoop` @@ -411,11 +425,11 @@ Implementations MUST honor: - **Blocking**: no compaction begins while another is in progress for the same session. The recommended mechanism is the log-recorded lock — append `compact/start` before the slow work and `compact/end` after (even on failure) — so the lock is visible to replay and crash recovery. ```ts cordis-catalog -abstract compactIfNeeded( session: Session, systemPrompt?: string, model?: string, signal?: AbortSignal, ): Promise -abstract compactRegion( session: Session, start: number, end: number, model: string, signal?: AbortSignal, ): Promise +abstract compactIfNeeded( agent: CompactAgentContext, turn: number, step: number, fullSystemPrompt: string, signal: AbortSignal, ): Promise +abstract compactRegion( session: Session, start: number, end: number, agent: CompactAgentContext, turn: number, step: number, signal?: AbortSignal, ): Promise ``` -Source: [`packages/compact/compact/src/index.ts:57`](../../packages/compact/compact/src/index.ts) +Source: [`packages/compact/compact/src/index.ts:63`](../../packages/compact/compact/src/index.ts) ### `ctx.fs` — `FileSystem` (abstract seam) @@ -493,7 +507,7 @@ get(id: SessionId): Session | undefined list(): Session[] ``` -Source: [`packages/core/session/src/index.ts:321`](../../packages/core/session/src/index.ts) +Source: [`packages/core/session/src/index.ts:322`](../../packages/core/session/src/index.ts) ### `ctx.subagents` — `SubagentService` @@ -560,7 +574,7 @@ The framework surface every plugin inherits, beyond the harness vocabulary above ### Inherited `ctx` members - `ctx.on / ctx.once` — Register an event listener (disposable). ([`vendor/cordis/src/events.ts:29`](../../vendor/cordis/src/events.ts)) -- `ctx.emit / ctx.parallel / ctx.serial / ctx.bail / ctx.waterfall` — Dispatch an event (sync / awaited / first-non-nullish / veto-chain). ([`vendor/cordis/src/events.ts:29`](../../vendor/cordis/src/events.ts)) +- `ctx.emit / ctx.parallel / ctx.serial / ctx.bail / ctx.waterfall` — Dispatch an event (sync / awaited / first-bail / veto-chain). ([`vendor/cordis/src/events.ts:29`](../../vendor/cordis/src/events.ts)) - `ctx.plugin / ctx.inject` — Load a plugin / declare required services. ([`vendor/cordis/src/registry.ts:144`](../../vendor/cordis/src/registry.ts)) - `ctx.effect` — Register a disposable side effect tied to the fiber. ([`vendor/cordis/src/fiber.ts:9`](../../vendor/cordis/src/fiber.ts)) - `ctx.get / ctx.set / ctx.provide / ctx.accessor / ctx.mixin` — Low-level service-store access and binding. ([`vendor/cordis/src/reflect.ts:7`](../../vendor/cordis/src/reflect.ts)) diff --git a/docs/core-data-structures/compaction.md b/docs/core-data-structures/compaction.md index 9bdb987f81..05d1ce6c2e 100644 --- a/docs/core-data-structures/compaction.md +++ b/docs/core-data-structures/compaction.md @@ -1,6 +1,6 @@ # Compaction -The compaction seam — a [capability seam](../rfc/implemented/architecture/2026-06-13-capability-seams.md) split like bash: interface ([dsh-compact](../../packages/compact/compact), `ctx.compact`), implementation (a backend such as `dsh-compact-basic`, deferred), and consumer (a `/compact` tool, deferred). Compaction is **one optional capability**, not part of the agent-loop spine — so its vocabulary lives here, not in [core.md](core.md). A tokenizer- or template-based backend is a sibling package implementing the same interface. Unlike bash, the interface necessarily depends on `dsh-session` and `dsh-llm`: its verbs are defined over a `Session` and its output is the `ContentBlock` vocabulary (see the [compaction capability-seam RFC](../rfc/proposed/feature/2026-06-18-compaction-capability-seam.md)). +The compaction seam — a [capability seam](../rfc/implemented/architecture/2026-06-13-capability-seams.md) split like bash: interface ([dsh-compact](../../packages/compact/compact), `ctx.compact`), implementation (a backend such as [dsh-compact-basic](../../packages/compact/compact-basic)), and consumer (a `/compact` tool, deferred). Compaction is **one optional capability**, not part of the agent-loop spine — so its vocabulary lives here, not in [core.md](core.md). A tokenizer- or template-based backend is a sibling package implementing the same interface. Unlike bash, the interface necessarily depends on `dsh-session` and `dsh-llm`: its verbs are defined over a `Session` and its output is the `ContentBlock` vocabulary (see the [compaction capability-seam RFC](../rfc/implemented/feature/2026-06-18-compaction-capability-seam.md)). Source: [`packages/compact/compact/src/types.ts`](../../packages/compact/compact/src/types.ts) @@ -11,7 +11,7 @@ Compaction extends [`SessionEventMap`](session.md) with three event types via de | Event | Payload | Role | |---|---|---| | `compact/start` | `{ turn }` | acquires the log-recorded lock | -| `compact/summary` | `{ summary, shadowedRange, shadowedSeqs, shadowedTokenCount }` | provenance: the summary blocks, the shadowed seq range, and the estimated token count | +| `compact/summary` | `{ summary, shadowedRange, shadowedSeqs, shadowedTokenCount }` | provenance: the summary blocks, the shadowed surface-boundary pair (`start`/`end` seqs — a position span, not a numeric interval), the shadowed seqs in surface order, and the estimated token count | | `compact/end` | `{ turn, error? }` | releases the lock (`error` set when summarization threw) | The lock brackets the **whole** operation: `compact/start` is appended first, then summarization, the `compact/summary` provenance record, and the `user/message` replacement all land, and only then `compact/end`. Releasing the lock last turns a crash mid-operation into a detectable orphaned lock (a `compact/start` with no matching `compact/end`) rather than a `compact/end` that falsely claims compaction finished. @@ -32,9 +32,16 @@ interface CompactionResult { endSeq: number /** The summary content blocks produced by the backend. */ summary: ContentBlock[] - /** The seq range that was shadowed [start, end] inclusive. */ + /** + * The surface-boundary pair that was shadowed: the seqs of the first + * (`start`) and last (`end`) surface nodes of the replaced range. A + * surface-POSITION span, not a numeric seq interval — after a prior replace + * lands a fresh high-seq summary node at an older range's position, `start` + * can be GREATER than `end`. {@link CompactionResult.shadowedSeqs} is the + * authoritative set of shadowed nodes, in surface order. + */ shadowedRange: { start: number; end: number } - /** The seq numbers of all shadowed surface nodes. */ + /** The seqs of all shadowed surface nodes, in surface order. */ shadowedSeqs: number[] /** Estimated token count of the shadowed content. */ shadowedTokenCount: number @@ -43,4 +50,6 @@ interface CompactionResult { ## The service -`CompactService` (`ctx.compact`, abstract — defined in [`packages/compact/compact/src/index.ts`](../../packages/compact/compact/src/index.ts)) declares two abstract methods: `compactIfNeeded(session, systemPrompt?, model?, signal?)` checks token pressure and compacts an older range if the history is too large (returning `null` when nothing needs it), and `compactRegion(session, start, end, model, signal?)` forcibly summarizes surface nodes `[start, end]` into a single replacement node. Both take an optional `signal: AbortSignal` that a backend summarizing via `ctx.llm.stream()` must forward into the call's `GenerateOptions.signal`, so an abort or dispose tears down the in-flight summarization. The entire strategy — token estimation, retention policy, event sequencing, summarization — is a HOW decision owned by the implementation. +`CompactService` (`ctx.compact`, abstract — defined in [`packages/compact/compact/src/index.ts`](../../packages/compact/compact/src/index.ts)) declares two abstract methods: `compactIfNeeded(agent, turn, step, fullSystemPrompt, signal)` checks token pressure and compacts an older range if the history is too large (returning `null` when nothing needs it), and `compactRegion(session, start, end, agent, turn, step, signal?)` forcibly summarizes surface nodes `[start, end]` into a single replacement node. `compactIfNeeded`'s parameters are all required — the loop's `agent/pre-step` checkpoint supplies the agent, lifecycle context, assembled `fullSystemPrompt`, and turn `signal`. A backend summarizing via `ctx.llm.stream()` must forward `signal` into the call's `GenerateOptions.signal`, so an abort or dispose tears down the in-flight summarization. The entire strategy — token estimation, retention policy, event sequencing, summarization — is a HOW decision owned by the implementation. + +Auto-compaction runs on the serial `agent/pre-step` loop seam (fired once per step, after `turn/start` and BEFORE the step opens and its request history is derived), not the `agent/request` waterfall: compaction mutates the session surface in place — with its log-only `compact/*` records landing cleanly outside any step — and the loop derives the request from the already-compacted surface. Retention is turn-agnostic — the only structural guard is tool-pairing balance (a compacted region's edges are balanced cuts on the surface, so it never splits a step's tool-calls from their results), so a single runaway turn that alone exceeds the window compacts its own early closed steps rather than being retained verbatim. The backend that ships this (`dsh-compact-basic`) documents the retention walk, summary shrink validation, bounded re-compaction, and the crash/recoverable failure taxonomy. diff --git a/docs/core-data-structures/core.md b/docs/core-data-structures/core.md index 176d4b32a1..187a4b19a8 100644 --- a/docs/core-data-structures/core.md +++ b/docs/core-data-structures/core.md @@ -206,7 +206,7 @@ type SessionEvent = { /** * Seq numbers of events that are provenance sources of this event * (e.g. the `assistant/chunk` seqs that built an `assistant/message`, - * or the surface nodes shadowed by a compaction marker). + * or the surface nodes shadowed by a compaction replace node). */ sourceEventSeqs?: number[] /** How this event entered the surface; absent for non-surface events. */ @@ -215,7 +215,7 @@ type SessionEvent = { }[T] ``` -The eleven event variants (`turn/start`, `turn/end`, `step/start`, `step/end`, `user/message`, `context/message`, `assistant/chunk`, `assistant/message`, `tool/call`, `tool/result`, `steering/message`), the `deriveMessages()` projection rules, the `TurnTrigger`/`TurnEndReason` reasons, and the turn-enclosure invariant are on **[session.md](session.md)**. How the log is made durable — the `SessionPersistence` seam, JSONL/SQLite backends, the `session/flush` checkpoint, crash recovery, and `SessionHeader` — is on **[persistence.md](persistence.md)**. +The twelve event variants (`turn/start`, `turn/end`, `step/start`, `step/end`, `user/message`, `context/message`, `assistant/chunk`, `assistant/message`, `tool/call`, `tool/result`, `steering/message`, `todo/write`), the `deriveMessages()` projection rules, the `TurnTrigger`/`TurnEndReason` reasons, and the turn-enclosure invariant are on **[session.md](session.md)**. How the log is made durable — the `SessionPersistence` seam, JSONL/SQLite backends, the `session/flush` checkpoint, crash recovery, and `SessionHeader` — is on **[persistence.md](persistence.md)**. ## The agent handle @@ -307,7 +307,7 @@ interface Agent { } ``` -`AgentStatus` is `'idle' | 'running' | 'disposed'`. `AgentId` is a branded string. `AgentOptions` (`model?`, `systemPrompt?`) is merge-extensible — plugins add creation options by declaration merging. The `agent/*` event taxonomy (lifecycle, turn/step boundaries, the `agent/request`/`agent/step-result`/`agent/turn-continuation` waterfalls) is in [architecture.md § Event taxonomy](../architecture.md#event-taxonomy). +`AgentStatus` is `'idle' | 'running' | 'disposed'`. `AgentId` is a branded string. `AgentOptions` (`model?`, `systemPrompt?`) is merge-extensible — plugins add creation options by declaration merging. The `agent/*` event taxonomy (lifecycle, turn/step boundaries, the serial `agent/pre-step` surface-mutation seam, and the `agent/request`/`agent/step-result`/`agent/turn-continuation` waterfalls) is in [architecture.md § Event taxonomy](../architecture.md#event-taxonomy). ## `ToolDefinition` diff --git a/docs/core-data-structures/persistence.md b/docs/core-data-structures/persistence.md index 8d8f032514..327162792a 100644 --- a/docs/core-data-structures/persistence.md +++ b/docs/core-data-structures/persistence.md @@ -14,7 +14,7 @@ A backend that reloads a log crashed mid-turn finds an open `turn/start` with no ## `SessionHeader` — metadata beside the log -Per-session metadata travels **separately** from the event log: format version, cwd, and lineage are storage concerns, not conversation events, so they stay out of `SessionEventMap` and never reach `deriveMessages()`. The header is attached to a `Session` via `session.header`. +Per-session metadata travels **separately** from the event log: format version, cwd, lineage, and the seed boundary are storage concerns, not conversation events, so they stay out of `SessionEventMap` and never reach `deriveMessages()`. The header is attached to a `Session` via `session.header`. Source: [`packages/core/session/src/types.ts`](../../packages/core/session/src/types.ts) @@ -78,6 +78,6 @@ Replay/fork is therefore `ctx.sessions.create(id, { seed: seedEvents })`; resumi Both implement the same abstract `SessionPersistence` (create/append/load/list over `SessionEvent`) and pass `runPersistenceContract`, proving the seam is genuinely backend-agnostic: - **[dsh-session-persistence-jsonl](../../packages/session-persistence/session-persistence-jsonl)** — an append-only JSONL log per session with crash-safe atomic writes, the interrupted-turn crash recovery above, and a read/replay path. -- **[dsh-session-persistence-sqlite](../../packages/session-persistence/session-persistence-sqlite)** — `node:sqlite`, one row per `SessionEvent`. The row shape `(session_id, seq, type, time, data)` maps 1:1 onto the event, so there is no parallel persisted schema to keep in sync. +- **[dsh-session-persistence-sqlite](../../packages/session-persistence/session-persistence-sqlite)** — `node:sqlite`, one row per `SessionEvent`. The row shape `(session_id, seq, type, time, data, source_event_seqs, surface_op)` maps 1:1 onto the event, including optional surface metadata, so there is no parallel persisted schema to keep in sync. Multiple backends sharing one on-disk session coordinate writes through the [shared persistence write-coordinator](../rfc/implemented/architecture/2026-06-18-shared-persistence-write-coordinator.md). diff --git a/docs/core-data-structures/session.md b/docs/core-data-structures/session.md index 8f8ef4a800..1a0e209e41 100644 --- a/docs/core-data-structures/session.md +++ b/docs/core-data-structures/session.md @@ -35,6 +35,31 @@ interface SessionEventMap { 'tool/result': { turn: number; step: number; callId: CallId; content: ContentBlock[]; isError: boolean; error?: { name: string; code: string } } /** Steering content injected between steps of a running turn. */ 'steering/message': { turn: number; content: ContentBlock[]; source: MessageSource } + /** + * The agent's whole todo list, carried as a full snapshot and replaced + * wholesale on each write — the current list is the most recent `todo/write` + * (last-write-wins on replay, no fold). Appended by an owning agent via + * `session.append('todo/write', { todos })`. + * + * NOT a {@link SurfaceEventType}: it produces no LLM message and never reaches + * `deriveMessages()`, so it carries no `surfaceOp` and stays off the surface — + * it is durable, replayable UI state, distinct from the conversation history. + * It is a `SessionEventMap` member riding the existing `session/event` emit, + * not a first-class Cordis `interface Events` notification, so it has no + * cordis-catalog row. + */ + 'todo/write': { todos: TodoItem[] } +} +``` + +### `TodoItem` — one todo-list entry + +The unit of the `todo/write` event's whole-list snapshot. Deliberately minimal — a `content` line and a three-state `status` (no id, priority, or `activeForm`): the list is replaced wholesale on every write, so entries need no stable identity, and the status triple is exactly the ACP `PlanEntryStatus`, so a UI bridge can map a todo list onto an ACP `plan` 1:1 (synthesizing the priority ACP additionally requires). See the [todo_write RFC](../rfc/implemented/feature/2026-06-29-todo-write-tool.md). + +```ts type-equiv +export interface TodoItem { + content: string + status: 'pending' | 'in_progress' | 'completed' } ``` @@ -55,7 +80,7 @@ type SessionEvent = { /** * Seq numbers of events that are provenance sources of this event * (e.g. the `assistant/chunk` seqs that built an `assistant/message`, - * or the surface nodes shadowed by a compaction marker). + * or the surface nodes shadowed by a compaction replace node). */ sourceEventSeqs?: number[] /** How this event entered the surface; absent for non-surface events. */ diff --git a/docs/development.md b/docs/development.md index 99f67e671d..431d7b4dac 100644 --- a/docs/development.md +++ b/docs/development.md @@ -61,7 +61,7 @@ lefthook is configured in `lefthook.yml` as an early local checkpoint before rev The vendor manifest guard checks that changes under `vendor/*/src` are staged with the matching `vendor/README.md` manifest update. See `vendor/README.md` before editing vendored code. -These hooks do not exactly mirror CI. Notably, `pre-push` runs unit tests without coverage, while CI runs `pnpm run test:coverage`; CI also runs an echo-agent smoke test and exercises the matrix on Node 24 and 26. +These hooks do not exactly mirror CI. Notably, `pre-push` runs unit tests without coverage, while CI runs `pnpm run test:coverage`; CI also runs echo-agent and built-bin smoke tests and exercises the matrix on Node 24 and 26. ## CI gates @@ -78,6 +78,7 @@ The GitHub workflow runs these gates on each pull request: - `pnpm run build` - `pnpm run hygiene` - an echo-agent smoke test that checks the demo's tool call, tool result, and JSONL output +- built-bin smoke tests that run the published `lib/bin.js` entrypoints under plain `node` `pnpm run hygiene` is the local shorthand for `pnpm run knip && pnpm run publint && pnpm run constraints && pnpm run verify-node-next-types`; CI also runs `pnpm run constraints` as an earlier fail-fast step, then runs the full hygiene script after `pnpm run build`. @@ -121,6 +122,12 @@ The coding-agent demo uses the real DeepSeek adapter and needs `DEEPSEEK_API_KEY pnpm run demo:coding ``` +The ACP server demo exposes the same coding agent over JSON-RPC stdio and also needs `DEEPSEEK_API_KEY`: + +```sh +pnpm run demo:acp +``` + ## TODO markers Use one of three comment tags to flag known issues in the code, ordered by urgency: diff --git a/docs/i18n/terminology.md b/docs/i18n/terminology.md new file mode 100644 index 0000000000..f9baa40adc --- /dev/null +++ b/docs/i18n/terminology.md @@ -0,0 +1,113 @@ +# Terminology + +本表约定本仓库的中英术语统一译法。 + +| English | 中文 | 备注 | +|---|---|---| +| ACP | ACP | 首次出现可写:ACP(Agent Client Protocol) | +| AI | AI | 首次出现可写:人工智能(AI) | +| API | API | | +| CLI | CLI | 首次出现可写:命令行界面(CLI) | +| Cordis | Cordis | 保留英文 | +| Function Calling | Function Calling | 首次出现可写:Function Calling(函数调用) | +| HMR | HMR | 首次出现可写:热模块替换(HMR) | +| JSON Schema | JSON Schema | | +| JSONL | JSONL | | +| lint | lint | | +| loader | loader | | +| LLM | LLM | 首次出现可写:大语言模型(LLM) | +| MCP | MCP | | +| RAG | RAG | 首次出现可写:检索增强生成(RAG) | +| SDK | SDK | | +| SSE | SSE | 首次出现可写:SSE(Server-Sent Events) | +| agent | agent | 首次出现可写:agent(智能体) | +| agent loop | agent loop | | +| fiber | fiber | 首次出现可写:fiber(插件运行时) | +| fixture | fixture | 指测试前置数据或环境 | +| fork | fork | 保留英文 | +| harness | harness | 保留英文 | +| manifest | manifest | 描述模块或工具元数据的文件 | +| schema DSL | schema DSL | | +| schema | schema | 保留英文 | +| seam | seam | 首次出现可写:seam(扩展点) | +| skill | skill | 首次出现可写:skill(技能) | +| spawn | spawn | 保留英文 | +| steering | steering | 首次出现可写:steering(中途引导) | +| subagent | subagent | 首次出现可写:subagent(子 agent) | +| transcript | transcript | 首次出现可写:transcript(文本记录);指会话渲染给用户或编辑器的完整文本,区别于事件日志(event log) | +| waterfall | waterfall | 首次出现可写:waterfall(瀑布式事件) | +| wire format | 协议格式 | 首次出现可写:协议格式(wire format) | +| adapter contract | 适配器契约 | 首次出现可写:适配器契约(adapter contract) | +| adapter | 适配器 | | +| append-only | 仅追加 | | +| artifact | 产物 | | +| block | 块 | | +| background task | 后台任务 | | +| backend | 后端 | | +| capability | 能力 | | +| cancel | 取消 | | +| checkpoint | 检查点 | | +| chunk | 分片 | | +| compaction | compaction | 首次出现可写:compaction(上下文压缩);正文优先保留英文 | +| consumer | 消费方 | | +| content block | 内容块 | | +| config | 配置 | | +| context | 上下文 | | +| context compaction | 上下文压缩 | 首次出现可写:上下文压缩(context compaction) | +| coverage | 覆盖率 | | +| crash recovery | 崩溃恢复 | | +| dispose | dispose | 首次出现可写:dispose(释放资源);正文优先保留英文 | +| durability | 持久性 | | +| event log | 事件日志 | | +| event | 事件 | | +| event stream | 事件流 | | +| executor | 执行器 | | +| extension | 扩展 | | +| finish reason | 结束原因 | | +| foreground run | 前台运行 | | +| hook | 钩子 | | +| implementation | 实现 | | +| inference | 推理(inference) | 每次提及时保留英文括注,避免与 reasoning 混淆 | +| injection | 注入 | | +| interface | 接口 | | +| integration | 集成 | | +| memory | memory / 记忆 / 内存 | 按上下文区分:agent memory 译为“记忆”;resource/memory usage 译为“内存” | +| message | 消息 | | +| mod | 模组 | 区别于 module(模块);plugin 译作「插件」 | +| model provider | 模型提供方 | | +| module | 模块 | | +| permission | 权限 | | +| persistence | 持久化 | | +| pipeline | 流水线 | | +| plugin | 插件 | mod 对应“模组” | +| prompt | 提示词 | | +| provider | 提供方 | | +| provider-neutral | 提供方无关 | | +| quality gate | 质量门禁 | | +| registry | 注册表 | | +| reasoning | 推理(reasoning) | 需要和 inference 区分时保留英文括注;`reasoning_content` 译为“思考内容” | +| replay | 回放 | | +| resume | 恢复 | | +| runtime | 运行时 | | +| sandbox | 沙箱 | | +| service | 服务 | | +| session | 会话 | | +| session event | 会话事件 | | +| snapshot | 快照 | | +| spine | 主干 | | +| step | 步骤 | | +| stream | 流 | | +| streaming | 流式输出 | | +| system prompt | 系统提示词 | | +| taxonomy | 分类体系 | | +| token usage | token 用量 | | +| thinking | thinking | API 字段保留;模型模式译为“思考” | +| tool | 工具 | | +| tool call | 工具调用 | | +| tool result | 工具结果 | | +| tool schema | 工具 schema | | +| toolkit | 工具包 | | +| turn | 轮次 | | +| typecheck | 类型检查 | | +| vocabulary | 词汇 | | +| workflow | 工作流 | | diff --git a/docs/module-graph.md b/docs/module-graph.md index a5011007df..3fe0d5aefe 100644 --- a/docs/module-graph.md +++ b/docs/module-graph.md @@ -27,6 +27,10 @@ graph TD llm-replay --> llm llm-replay --> session session-persistence --> session + compact-basic --> agent + compact-basic --> compact + compact-basic --> llm + compact-basic --> session invariants --> agent invariants --> llm invariants --> session @@ -62,6 +66,9 @@ graph TD tool-fs --> llm tool-fs --> system-prompt tool-fs --> tools + tool-todo --> agent + tool-todo --> session + tool-todo --> tools agent-core --> agent agent-core --> agent-loop agent-core --> invariants @@ -117,6 +124,7 @@ graph TD | `fs-local` | `fs` | | `llm-replay` | `llm`, `session` | | `session-persistence` | `session` | +| `compact-basic` | `agent`, `compact`, `llm`, `session` | | `invariants` | `agent`, `llm`, `session` | | `session-persistence-jsonl` | `session`, `session-persistence` | | `session-persistence-sqlite` | `session`, `session-persistence` | @@ -127,6 +135,7 @@ graph TD | `subagent` | `agent`, `llm`, `tools` | | `tool-bash` | `agent`, `bash`, `llm`, `tools` | | `tool-fs` | `fs`, `llm`, `system-prompt`, `tools` | +| `tool-todo` | `agent`, `session`, `tools` | | `agent-core` | `agent`, `agent-loop`, `invariants`, `llm`, `session`, `system-prompt`, `tool-bash`, `tools` | | `subagent-acp` | `agent`, `llm`, `subagent` | | `subagent-inprocess` | `agent`, `llm`, `session`, `subagent` | diff --git a/docs/rfc/README.md b/docs/rfc/README.md index 4c36503617..6d12cc9c3d 100644 --- a/docs/rfc/README.md +++ b/docs/rfc/README.md @@ -44,7 +44,6 @@ Do NOT write one for a mechanical or local choice (a variable name, a one-file r | [Agent Client Protocol (ACP) support for external editors](proposed/feature/2026-06-14-acp-agent-client-protocol.md) | 2026-06-14 | | [Multiplex concurrent ACP sessions over one connection](proposed/feature/2026-06-14-acp-multi-session.md) | 2026-06-14 | | [Optional Code Mode — model writes TypeScript against an SDK of all tools](proposed/feature/2026-06-15-optional-code-mode.md) | 2026-06-15 | -| [Compaction as a capability seam (abstract contract + basic backend)](proposed/feature/2026-06-18-compaction-capability-seam.md) | 2026-06-18 | ### Simplification @@ -84,8 +83,10 @@ Do NOT write one for a mechanical or local choice (a variable name, a one-file r |---|---| | [Filesystem tool schemas — model-facing read/write/edit shapes](implemented/feature/2026-06-17-filesystem-tool-schemas.md) | 2026-06-17 | | [Rich ACP bash rendering — the terminal card (`_meta`) and command classification](implemented/feature/2026-06-18-acp-terminal-and-tool-rendering.md) | 2026-06-18 | +| [Compaction as a capability seam (abstract contract + basic backend)](implemented/feature/2026-06-18-compaction-capability-seam.md) | 2026-06-18 | | [Subagent capability seam](implemented/feature/2026-06-21-subagent-capability-seam.md) | 2026-06-21 | | [ACP subagent backend (out-of-process delegation)](implemented/feature/2026-06-22-acp-subagent-backend.md) | 2026-06-22 | +| [The `todo_write` tool — model task list as event-sourced session state](implemented/feature/2026-06-29-todo-write-tool.md) | 2026-06-29 | ### Simplification diff --git a/docs/rfc/implemented/feature/2026-06-18-compaction-capability-seam.md b/docs/rfc/implemented/feature/2026-06-18-compaction-capability-seam.md new file mode 100644 index 0000000000..9e08df2fbd --- /dev/null +++ b/docs/rfc/implemented/feature/2026-06-18-compaction-capability-seam.md @@ -0,0 +1,120 @@ +# RFC: Compaction as a capability seam (abstract contract + basic backend) + +Status: implemented (2026-06-18; retention/seam reform 2026-06-26) + +## Context + +A long-running agent conversation grows without bound. As the event log accumulates turns, the derived message history eventually approaches the model's context window — the model then truncates mid-response (`max-tokens`) or degrades. **Compaction** is the mitigation: replace a run of older history with a concise summary, keeping recent context intact. + +The [session surface](../../implemented/architecture/2026-06-18-session-surface.md) was built as the foundation for exactly this — a linked list over the event log with a `surfaceOp: { op: 'replace', start, end }` operation purpose-built to shadow a range of nodes and insert a replacement, with `sourceEventSeqs` recording provenance so the decision replays deterministically. What remained was the plugin that *decides what to compact and produces the summary*. + +Two forces shape the design. First, compaction is **swappable**: token counting can be a char/4 heuristic or a real tokenizer, and summarization can be a model call, a template, or a remote service — these vary independently of *when* and *which range* to compact. Second, a later commit (`ce43c25`) closed `SurfaceEventType` to five event types (`user/message`, `assistant/message`, `tool/result`, `context/message`, `steering/message`); only those may carry `surfaceOp`. A bespoke `compaction/*` event therefore **cannot** itself appear on the surface — the compiler rejects `surfaceOp` on it and the invariants plugin rejects it at runtime. + +## Decision + +### Compaction is a capability seam, split interface / implementation + +Per the [capability-seams RFC](../../implemented/architecture/2026-06-13-capability-seams.md), compaction ships as separate packages so the contract, the algorithm, and (later) the consumer surface evolve independently: + +1. **Interface** — `@deepseek-ai/dsh-compact`: an abstract `CompactService` owning the `ctx.compact` key, the `CompactionResult` vocabulary, and the `compact/*` session events. It declares `compactIfNeeded()` and `compactRegion()` as **abstract** — the contract states *what* compaction does, not *how*. +2. **Implementation** — `@deepseek-ai/dsh-compact-basic`: a concrete `BasicCompactService` that owns the entire algorithm — token estimation (char/4 + per-block overhead), the tail→head retention walk, summarization via `ctx.llm.stream()`, the surface replacement, the lock, and the `agent/pre-step` auto-compaction listener. A tokenizer-based or template-based backend is a sibling package (or a subclass overriding the two protected estimation/summarization hooks). +3. **Consumer** — deferred. A `/compact` tool and slash command will `inject: ['compact']` and call the contract; they are intentionally out of scope here so the seam settles first. + +### The contract depends on `dsh-session` and `dsh-llm` — a deliberate deviation + +The capability-seams RFC states the interface package "depends only on cordis" (true of `dsh-bash`, whose vocabulary is self-contained). Compaction **cannot** honor that: its verbs are defined *over* a `Session` (`compactRegion(session, start, end)`) and its output *is* the content vocabulary (`CompactionResult.summary: ContentBlock[]`). There is no way to express the contract without naming `Session`/`SessionEvent` (from `dsh-session`) and `ContentBlock` (from `dsh-llm`). + +This is not a coupling smell — it is the contract's domain. The "only cordis" guidance was always shorthand for "the interface depends only on what the contract genuinely names, and never on an implementation." `dsh-session` and `dsh-llm` are themselves interface/vocabulary packages, not implementations; `dsh-compact` still imports no backend. The seam's real invariant — *consumers and implementations evolve independently behind an abstract service* — holds intact. + +### Abstract `compactIfNeeded` / `compactRegion`, algorithm in the backend + +An earlier draft put the full algorithm (the retention walk, token-summing, text extraction) as concrete methods on the interface, with only `estimateContentTokens()` and `summarize()` abstract. That recouples the contract to one strategy: a backend that wants a different retention policy or a different event-sequencing would have to fight inherited concrete code. Making both core methods abstract puts every *how* decision in the backend, where it belongs, and keeps the interface a pure statement of *what*. The backend remains internally factored — `estimateContentTokens()` and `summarize()` are `protected` hooks a sub-backend can override without reimplementing the walk — but that factoring is the backend's private concern, not the contract's. + +`compactIfNeeded(agent, turn, step, fullSystemPrompt, signal)` takes **required** parameters (not the original all-optional shape). The auto-compaction seam (below) always supplies the agent, lifecycle context, assembled system prompt (counted toward the estimate), and the turn's abort signal, so optionality would only invite a hidden default at the seam. The session being compacted comes from the agent context. `compactRegion(session, start, end, agent, turn, step, signal?)` keeps an optional signal (a manual caller may omit it). Passing lifecycle context rather than a concrete model keeps router agents honest: the backend's summarization request can run through `agent/request`, where model-routing plugins already choose the actual model. + +### Auto-compaction runs on `agent/pre-step`, a dedicated surface-mutation seam + +Compaction is a **surface mutation**, not a request transform — and that distinction is the seam it belongs on. The loop's request lifecycle, per step, is: assemble the system prompt → open the step → derive the message history from the surface → run the `agent/request` waterfall → call the model. An earlier cut wedged compaction into the `agent/request` waterfall, which forced two problems: (1) the loop had already derived `messages` from the *stale* surface, so the listener had to mutate the surface and then *re-derive* and overwrite `request.messages` — a double-derive whose only purpose was to undo the premature first derive; and (2) `agent/request` also carries downstream-injected context a listener might have added to `request.messages`, which compaction cannot act on (it can only compact the surface), inviting the confusion of measuring tokens compaction can't shed. + +The fix is a dedicated loop seam, **`agent/pre-step`** (`@mode serial`), fired by the loop *after* system assembly and *before* the step opens (`step/start`): + +``` +assembly = ctx.systemPrompt.assemble() +await ctx.serial('agent/pre-step', agent, turn, step, system, signal) ⟵ compaction mutates the surface here +session('step/start') ⟵ the step opens AFTER the seam +messages = session.deriveMessages() ⟵ single derive, reflects the compaction +request = waterfall agent/request ⟵ pure request transform (hooks, model switch) +``` + +This makes the layering correct *by construction*: compaction mutates the surface, the loop derives **once** from the result (no double-derive), and at `pre-step` the assembled `messages` do not yet exist — so a listener structurally *cannot* see or be expected to act on downstream-injected context. `agent/request` reverts to a pure request transformer. Firing the seam **before** `step/start` (not inside the open step) is load-bearing for crash-safety: compaction's log-only `compact/*` records and its replacement node land *outside* any step, so the honest log structure a crash leaves (a dangling `compact/start` sitting before the synthetic `turn/end` that turn-repair appends) holds without a half-open step to reconcile. The seam is `serial` (awaited, in registration order), not `parallel`: a listener mutates the surface as a side effect — there is nothing to transform or return — and serial isolates listeners from each other so two surface-mutating listeners can never interleave their `session.append`s. Cordis `serial` does bail early if a listener returns a bail value, so `agent/pre-step` listeners are typed/documented to return `void` and must not use that bail channel as a semantic veto surface. + +This **amends** the original RFC's claim of "NO changes to `dsh-agent-loop`; compaction is a pure plugin." That claim was load-bearing for a wrong design — reusing `agent/request` was the mistake. Per the pre-release "foundation over blast radius" stance, adding the correct seam (one event declaration in `dsh-agent`, one awaited emit in the loop) beats preserving a no-change boast that locked in the double-derive. + +### Retention is turn-agnostic; tool-pairing balance is the only structural guard + +Auto-compaction fires before **every** step, not once per turn. This is **load-bearing for runaway-turn survival**: a tool-heavy ReAct turn appends an `assistant/message` + a `tool/result` per step, so the surface grows *within* a turn. A single turn can grow past the window on its own (a "runaway turn") — and the only moment to rescue it before the next model call overflows is the next step's `pre-step` checkpoint. Gating compaction to a turn's first step (or, worse, retaining the whole in-flight turn verbatim) re-opens exactly the hole compaction exists to close: the harness would die when compaction is most needed. + +So retention does **not** protect the in-flight turn, and turn boundaries play no role in it. `compactIfNeeded` walks the surface nodes tail→head, summing per-node token estimates, and retains the smallest tail-run of **whole units** whose total reaches `retainTokens`; everything older is compacted (head-anchored — see below). A *unit* is either a whole closed step (its `assistant/message` plus its `tool/result`s) or a single no-step node (a pre-step `user/message`, inter-step `steering/message`, or injection `context/message`). The walk rounds toward retaining *more*: when the raw token cutoff lands mid-step, it extends the retained side head-ward until the cut before the retained node is **tool-pairing balanced**. The single structural guard is therefore **tool-pairing balance** — a region's edges are balanced cuts on the *surface* (no unanswered `tool-call` crosses either edge), so a compacted region never splits a step's tool-calls from their `tool/result`s (which would produce a transcript every provider rejects). The check is decided over the surface linked list, **not** the log's `step/*` markers: a compaction lands a replacement node at a high log seq whose surface position is the head, so a log-position scan mis-reads its neighbours — `dsh-session` exports `isToolPairingBalanced(nodes, events, beforeSeq)` for the surface-anchored check. `compactRegion` enforces it strictly, throwing on a boundary that would split a step. + +A runaway turn thus compacts exactly like any other history: its early *closed* steps get summarized while its recent steps stay verbatim. When the only compactable content left is an un-splittable open tail step (its tool-calls have no results yet), compaction declines (`null`) and retries once that step closes. + +**Single-unit overflow is out of scope, by design.** If a single retained unit — one closed step, or a large free node such as a pasted `user/message` — *alone* exceeds the budget, compaction cannot help and the next model call may go out over-budget. Bounding an individual unit's size is a separate concern (output truncation), handled elsewhere; compaction makes no promise about it, and the harness without such a mechanism can still break on a single oversized unit. This is named honestly rather than papered over. + +### Head-anchoring: one auto checkpoint, always at the head + +`compactIfNeeded` always anchors the compacted range at the surface **head** (`nodes[0]`). After a first compaction lands a summary node at the head, the *second* compaction's range starts at that summary node and re-summarizes it together with the steps accumulated since — so the surface holds **at most one** auto-generated checkpoint, always at the head, re-consolidated each cycle (the backend's checkpoint-merge prompt makes this a cheap incremental merge — see below). This is *why* `CompactionResult.shadowedRange` is a **surface-position span, not a numeric seq interval**: after a replace lands a fresh high-seq summary node at an older range's position, `start` can be numerically **greater** than `end`. The range is resolved positionally (index into the ordered node list and slice), and `shadowedSeqs` is the authoritative set in surface order. (Manual `compactRegion` may target any aligned mid-range and so *can* leave several checkpoints; the checkpoint framing does not claim everything after it is recent.) + +### Approximate convergence invariant + +`resolveConfig` validates numeric knobs but does NOT reject based on a pretend summary-length invariant. Convergence is dynamic: provider output caps can be spent on hidden or surfaced reasoning tokens, and the model may emit a summary of unpredictable size. `maxTokens` is only the provider-side generation cap for the summarization call; reasoning blocks are stripped before the checkpoint is stored. If a compacted surface is still over threshold, `compactIfNeeded()` re-compacts the head checkpoint up to `compactionRetries` extra times, but each committed summary must be smaller than the content it shadows. The sole residual is the single-unit-overflow case above (a backward-rounded oversized step can push the retained tail over budget) — which is exactly the out-of-scope concern, not a thrash bug. + +### Surface replacement: `compact/*` events are log-only; one `user/message` carries the summary + +Because `SurfaceEventType` is closed, the summary cannot ride on a `compact/*` event. The backend instead appends a **single `user/message`** with `surfaceOp: { op: 'replace', start, end }` whose `content` is the (framed) summary and whose `sourceEventSeqs` covers the shadowed nodes *and* the bookkeeping events. The `compact/*` events are pure log records (lock + provenance). The surface mutation sits **inside** the lock — `compact/end` is the last event appended: + +``` +compact/start → log-only. Acquires the lock. +[summarize older range via the backend] +compact/summary → log-only. Provenance: raw summary, range, shadowed seqs, token count. +user/message → surfaceOp { op:'replace', start, end }. THE surface mutation (framed summary). + deriveMessages() renders it as a user-role message. +compact/end → log-only. Releases the lock (carries `error` on a recoverable failure). +``` + +`deriveMessages()` then yields `[summary_as_user_message, ...retained_nodes]`. Reusing `user/message` is honest rather than a workaround: a summary genuinely *is* user-role context. + +### Checkpoint framing + incremental merge (backend-private) + +The landed `user/message` is not the raw summary: the backend wraps it in a checkpoint preamble (so a resuming model reads it as established background, not a fresh request) and `` tags. The tags make a prior checkpoint detectable on the next cycle, and the summarization prompt then instructs the model to *merge it in place* (preserve still-true facts, drop stale) rather than re-summarize verbatim — a cheap incremental merge that needs no extra log/event machinery. The raw, unframed summary stays on the `compact/summary` provenance event. This framing is entirely a **backend HOW decision** — the contract only promises "a single replace `user/message` carries the (possibly framed) summary; the raw summary lives on `compact/summary`." A template or remote backend may frame differently or not at all. + +### Blocking via a log-recorded lock, plus a crash/recoverable failure taxonomy + +The `compact/start … compact/end` bracket is justified, in order of what now does the work: + +1. **Crash-detectable orphan + provenance** (primary). Summarization is a slow model call persisted *after* `compact/start`. A crash mid-summarization leaves a `compact/start` with no matching `compact/end` — a detectable orphan. Releasing the lock last (rather than first) converts the crash window from *silent corruption* into that detectable orphan. +2. **Prevents concurrent compaction.** `compactRegion` refuses to start if the current turn holds an unmatched `compact/start`. (The loop is single-threaded across the awaited `pre-step`, so this is also a re-entry tripwire — a thrown "already in progress" signals a real bug.) + +Two failure paths, both documented: + +- **Crash** (the loop dies mid-summarization): a dangling `compact/start`, no closer. Because `compact/*` are **log-only**, the orphan is **inert** — the surface replacement never landed, so the full, uncompacted history derives correctly. Generic turn-repair (`interruptedTurnClosers`) closes the turn with a synthetic `turn/end`; the orphan sits *before* that `turn/end`, so the turn-scoped in-progress check never sees it and a crash can't wedge future compaction. Compaction simply re-attempts at the next `pre-step`. +- **Recoverable** (summarization throws but the loop survives): the backend appends `compact/end` with its **`error`** field set, leaving the surface untouched, and the model call proceeds with full history. + +`compact/end` keeps its `error?` field (mirroring `tool/result`'s self-contained error — one event tells success from failure without correlating a sibling). There is no separate `compact/error` event. + +**Core session repair stays compaction-agnostic — deliberately.** `interruptedTurnClosers` is never taught about `compact/*`. Teaching it would force every future `xxx/start … xxx/end` plugin pair to patch a core module — exactly the coupling the capability-seam architecture exists to avoid. Because the log-only orphan is inert, no special repair is needed: generic turn-repair plus the inertness of an un-landed surface mutation is sufficient. + +## Consequences + +- **New packages**: `packages/compact/compact` (interface) and a sibling `compact-basic` (backend) under `packages/compact/`, wired into the root tsconfigs. The consumer tier is deferred. +- **New loop seam**: `agent/pre-step` (`@mode serial`) declared in `dsh-agent` and emitted by `dsh-agent-loop` after system assembly and before `step/start`. This is a documented change to the loop — `docs/architecture.md` records it and the generated cordis catalog carries its signature. +- **`SessionEventMap`** gains `compact/start` / `compact/summary` / `compact/end` by declaration merging (merge-extensible); `SurfaceEventType` is **not** touched. These are session events, not cordis `Events`, so the event-taxonomy gate needs no entry. +- **`dsh-session`** gains the tool-pairing balance predicate (`isToolPairingBalanced`, in `tool-pairing.ts`, exported from the package index) that `compactRegion`/`compactIfNeeded` use to keep a collapsed region from splitting a step's tool-call/result pair. The surface `replace` op and the surface-metadata runtime guard already existed and are reused. +- **`dsh-invariants`** drops its `surface replace: start must be <= end` assertion: a head-anchored compaction lands a high-seq replacement node at an older range's *position*, so `start > end` numerically is normal and valid (the range is positional, validated by the surface's `indexOf` checks that remain). The turn-enclosure invariant is reused unchanged. +- **Wiring**: `dsh-compact-basic` is loaded in `examples/coding-agent`'s `cordis.yml`, so the seam ships in the real demo (it was previously loaded nowhere). + +## Testing + +- **Unit** (`dsh-compact-basic`): the whole-unit retention walk, the convergence-invariant throw, both failure paths (`compact/end` with/without `error`), head-anchoring producing a non-monotonic `shadowedRange`, decline-on-open-tail, crash-orphan inertness, and the **runaway-turn regression** — a single oversized open turn compacts its early closed steps (proven to fail on the layer-2 protection it replaced). Driven through the real `dsh-invariants` plugin and the real Loader/inject path. +- **Loop** (`dsh-agent-loop`): `agent/pre-step` fires once per step, after `turn/start` and before `step/start`, awaited; a surface mutation in a `pre-step` listener lands outside the step and is reflected in the single derived request. +- **With-key e2e** (`examples/coding-agent`): a real model + real bash session with a lowered `contextWindow`/`retainTokens` triggers compaction mid-session; the test verifies the WORLD (a `compact/start…end` pair landed, the surface shrank, the agent still completed the task after compaction). This is compaction's first real-world exercise and the runaway-survival net. +- **Snapshot (deferred, named gap)**: a full-transcript snapshot of a runaway-turn compaction is NOT yet possible — `dsh-llm-replay` derives one model call per `(turn, step)` from `assistant/chunk` events, but the summarization call records no `assistant/chunk`s and carries no `sessionId` (it binds to the anonymous cursor and claims a non-existent extra script). Covering it needs net-new replay infrastructure (record/replay an interleaved summarization call) and is scheduled as a follow-up rather than discovered mid-build. diff --git a/docs/rfc/implemented/feature/2026-06-29-todo-write-tool.md b/docs/rfc/implemented/feature/2026-06-29-todo-write-tool.md new file mode 100644 index 0000000000..a2287918f9 --- /dev/null +++ b/docs/rfc/implemented/feature/2026-06-29-todo-write-tool.md @@ -0,0 +1,58 @@ +# RFC: The `todo_write` tool — model task list as event-sourced session state + +Status: implemented + +## Problem + +The harness gives the model bash and subagent tools but no way to record a structured task list. A todo list serves two co-equal purposes: it steers the model to plan multi-step work and keep the active task unambiguous (at most one active, exactly one while work remains), and it gives the human a live progress checklist. The ACP protocol has a native `plan` sessionUpdate that editors (Zed) already render, but the bridge never emitted one. Every reference coding agent surveyed (claude-code, opencode, codex, oh-my-pi, pi) ships some form of this; the harness had nothing. + +## Decision + +Add a model-facing `todo_write(todos: [{ content, status }])` tool whose whole-list state lives on the event-sourced session log as a new `todo/write` `SessionEventMap` variant. Both the stdio UI and the ACP bridge render off the existing `session/event` — the ACP bridge maps the list to a `plan` sessionUpdate. + +### Whole-list replace, three-state status + +The model sends the ENTIRE list every call; the new list replaces the old (last-write-wins on replay). This is the shape claude-code V1, opencode, and codex `update_plan` all use, and the shape the model is most trained on — no per-item ids, no delta protocol. `status` is exactly `pending | in_progress | completed`: the same triple as codex `update_plan` and, crucially, **identical to the ACP `PlanEntryStatus`**, so the bridge maps it 1:1 with no lossy translation. + +### State on the session log, not a service + +The list is appended as a `todo/write` event carrying the full `{ todos }` snapshot. The harness is event-sourced — the LLM history, tool calls, and turn structure all live on the log — so the todo list lives there too. This buys durability, replay, and `session/load` reconstruction for free: a reopened session re-derives the current list (the last `todo/write`) and the ACP bridge re-emits the `plan` on load, with no separate persistence backend, no in-memory service to rehydrate, and no extra wiring. An in-memory `ctx.todos` service would have had to reinvent all of that. + +### NOT a surface event + +`todo/write` is deliberately excluded from `SurfaceEventType`. The surface is the projection that produces the LLM message history (`deriveMessages()`); a todo write produces no conversation message. So it carries no `surfaceOp`, never joins the surface linked list, and never reaches `deriveMessages()` — it is durable, replayable *UI* state that travels alongside the conversation without being part of it. (The dev-mode invariants still require it to sit inside an open turn, which it always does: it is appended mid-step during a tool call.) + +### Priority synthesized only at the ACP boundary + +ACP's `PlanEntry` requires `content` + `priority` + `status`, but a `TodoItem` has no priority — the model never reasons about it. Rather than burden the schema with a field the model must always supply, the bridge synthesizes a constant `priority: 'medium'` on every entry when it builds the `plan`. Priority is an ACP wire requirement, not a harness concept, so it lives at exactly the boundary that needs it. + +### Dropped vs claude-code V1: `activeForm`, id, priority + +claude-code V1's item is `{ content, status, activeForm }`; later (V2) it grew ids, dependencies, and ownership — but only to support agent *swarms* (disk-backed, lock-guarded, per-item mutation). This tool keeps the item at the minimum: `{ content, status }`. No `activeForm` (the present-continuous label) — the UI shows `content`; no id — whole-list replace needs no stable identity; no priority — see above. Each dropped field is one less thing the model must produce on every call. + +### Single owner — no swarm machinery (YAGNI) + +The list belongs to the ONE agent session that called the tool (`exec.agent.session`); a non-agent caller is rejected. There is deliberately no shared/multi-owner scope, no capability seam (interface/impl/consumer), no scope resolver, and no delta protocol. The harness does have subagents, and a shared cross-agent list is conceivable — but building that now means designing for a form the product does not yet have. The whole-list-replace + single-owner shape is what claude-code V1, opencode, and codex all ship; if a shared list is ever needed, the on-log representation would change to per-item deltas (so concurrent writers can't clobber each other) and a scope resolver would choose the target log. That is a future RFC, not speculative scaffolding today. + +### Validation: the cheap middle + +The schema enforces type/required/enum. Beyond that, `execute` rejects empty or duplicate `content` and more than one `in_progress` task. claude-code leaves single-in-progress to the prompt; oh-my-pi enforces it in code. We take the middle: enforce the cheap invariants that make a plan *coherent* (no blank tasks, no dupes, at most one active), but leave ordering and the discipline of keeping the list current to the model via the tool description. A rejected write returns an `isError` result so the model self-corrects. + +## Why no cordis-catalog entry / no `@mode` + +`todo/write` is a member of `SessionEventMap`, not a first-class cordis `interface Events` event. The catalog generator (`scripts/gen-cordis-catalog.ts`) scans `interface Events` declarations; a `SessionEventMap` variant rides the existing `session/event` emit and produces no new catalog row. So it carries no `@mode` tag (which the generator requires only on `interface Events` members) — adding one would be meaningless. + +## Testing + +Four tiers, designed up front: +- **Unit** — the session event (append/snapshot-clone/last-write-wins/not-on-surface); the tool (schema shape, arg validation via the real `ctx.tools.execute`, value validation, the event append + replacement, no-agent rejection, `presentCall`, HMR-safety); the ACP `todosToPlan` mapping; the stdio render arm. +- **Real-Loader path** — the plugin run through `Loader.unwrapExports`, asserting the namespace export shape survives (it HAS `inject`, so a stray default would crash at load — postmortem/0001). +- **Full-loop integration** — a scripted mock model calls `todo_write` through the real agent loop; the `todo/write` event lands and a second call replaces it. +- **`session/load` replay** — a persisted `todo/write` re-emits the `plan` update when a fresh ACP bridge loads the session. +- **With-key e2e + snapshot** — a real prompt induces a `todo_write`; the snapshot golden gains the `plan` notification and the log event. + +## Alternatives rejected + +- **In-memory `ctx.todos` service** — would reinvent durability, replay, and `session/load` reconstruction the log gives for free. +- **Per-item delta protocol** — only needed for a shared multi-owner list, which is out of scope; whole-list replace is simpler and matches the references. +- **Tool in `core/`** — `todo_write` is an extension tool registering on `ctx.tools`, not part of the spine; it lives in its own `packages/todo/` group like other tool families. diff --git a/docs/rfc/implemented/process/2026-06-20-generated-cordis-catalog.md b/docs/rfc/implemented/process/2026-06-20-generated-cordis-catalog.md index 2801ae209c..ae07f4b14c 100644 --- a/docs/rfc/implemented/process/2026-06-20-generated-cordis-catalog.md +++ b/docs/rfc/implemented/process/2026-06-20-generated-cordis-catalog.md @@ -20,7 +20,7 @@ Pure generation is correct here because the codebase is disciplined enough that Specific choices: -- **`@mode` tag, cross-checked.** Each harness event's JSDoc carries an explicit `@mode emit|waterfall|parallel` tag; the generator hard-errors on a missing tag. Where the signature shape is conclusive — a trailing `next: () => …` parameter is structurally a waterfall — it asserts the tag agrees and hard-errors on a contradiction. The emit-vs-parallel distinction is not structurally visible (`session/flush` returns `Promise | void` with no `next`), so it is trusted from the tag. The authoring rule lives in [AGENTS.md](../../../../AGENTS.md). +- **`@mode` tag, cross-checked.** Each harness event's JSDoc carries an explicit `@mode emit|waterfall|parallel|serial` tag; the generator hard-errors on a missing tag. Where the signature shape is conclusive — a trailing `next: () => …` parameter is structurally a waterfall — it asserts the tag agrees and hard-errors on a contradiction. The emit/parallel/serial distinction is not structurally visible (`session/flush` returns `Promise | void` with no `next`, as does the ordered `agent/pre-step` checkpoint), so it is trusted from the tag. The authoring rule lives in [AGENTS.md](../../../../AGENTS.md). - **Tiered scope.** The harness tier (the 8 `@deepseek-ai/dsh-*` services + their events) is rendered in full from source. The inherited tier (cordis-core `ctx.on/emit/effect/provide/…` + the `internal/*` events + loader/hmr/timer) is pinned vendor source a plugin also sees; it is rendered tersely (name + one-line + source pointer) from a curated table in the generator, NOT walked from the vendor AST — the cordis-core `Context` mixes true ctx members with non-service fields (`root`, `baseUrl`, `logger`), and the vendor surface changes only on a deliberate vendor sync. - **Cross-links to the data-structure catalog.** A type name in a signature (`GenerateOptions`, `StreamChunk`, `ToolDefinition`, …) links to the core-data-structures page that documents it. The map is a small hand-curated const in the generator — NOT `type-equiv.manifest.json`, which documents the `…Map` symbols while signatures reference the derived union names, and lists a few symbols on two pages. - **A dedicated fence.** Signature blocks use a ` ```ts cordis-catalog ` info string that `doc-typecheck` recognizes and skips (a bare signature fragment is not standalone-compilable), excluded from the opt-out ratio — the same treatment `type-equiv` blocks get. diff --git a/docs/rfc/proposed/feature/2026-06-18-compaction-capability-seam.md b/docs/rfc/proposed/feature/2026-06-18-compaction-capability-seam.md deleted file mode 100644 index 2d559fa65c..0000000000 --- a/docs/rfc/proposed/feature/2026-06-18-compaction-capability-seam.md +++ /dev/null @@ -1,59 +0,0 @@ -# RFC: Compaction as a capability seam (abstract contract + basic backend) - -Status: proposed (2026-06-18) - -## Context - -A long-running agent conversation grows without bound. As the event log accumulates turns, the derived message history eventually approaches the model's context window — the model then truncates mid-response (`max-tokens`) or degrades. **Compaction** is the mitigation: replace a run of older history with a concise summary, keeping recent context intact. - -The [session surface](../../implemented/architecture/2026-06-18-session-surface.md) was built as the foundation for exactly this — a linked list over the event log with a `surfaceOp: { op: 'replace', start, end }` operation purpose-built to shadow a range of nodes and insert a replacement, with `sourceEventSeqs` recording provenance so the decision replays deterministically. What remained was the plugin that *decides what to compact and produces the summary*. - -Two forces shape the design. First, compaction is **swappable**: token counting can be a char/4 heuristic or a real tokenizer, and summarization can be a model call, a template, or a remote service — these vary independently of *when* and *which range* to compact. Second, a later commit (`ce43c25`) closed `SurfaceEventType` to five event types (`user/message`, `assistant/message`, `tool/result`, `context/message`, `steering/message`); only those may carry `surfaceOp`. A bespoke `compaction/*` event therefore **cannot** itself appear on the surface — the compiler rejects `surfaceOp` on it and the invariants plugin rejects it at runtime. - -## Decision - -### Compaction is a capability seam, split interface / implementation - -Per the [capability-seams RFC](../../implemented/architecture/2026-06-13-capability-seams.md), compaction ships as separate packages so the contract, the algorithm, and (later) the consumer surface evolve independently: - -1. **Interface** — `@deepseek-ai/dsh-compact`: an abstract `CompactService` owning the `ctx.compact` key, the `CompactionResult` vocabulary, and the `compact/*` session events. It declares `compactIfNeeded()` and `compactRegion()` as **abstract** — the contract states *what* compaction does, not *how*. -2. **Implementation** — `@deepseek-ai/dsh-compact-basic`: a concrete `BasicCompactService` that owns the entire algorithm — token estimation (char/4 + per-block overhead), the tail→head retention walk, summarization via `ctx.llm.generate()`, the surface replacement, the lock, and the `agent/request` auto-compaction listener. A tokenizer-based or template-based backend is a sibling package (or a subclass overriding the two protected estimation/summarization hooks). -3. **Consumer** — deferred. A `/compact` tool and slash command will `inject: ['compact']` and call the contract; they are intentionally out of scope here so the seam settles first. - -### The contract depends on `dsh-session` and `dsh-llm` — a deliberate deviation - -The capability-seams RFC states the interface package "depends only on cordis" (true of `dsh-bash`, whose vocabulary is self-contained). Compaction **cannot** honor that: its verbs are defined *over* a `Session` (`compactRegion(session, start, end)`) and its output *is* the content vocabulary (`CompactionResult.summary: ContentBlock[]`). There is no way to express the contract without naming `Session`/`SessionEvent` (from `dsh-session`) and `ContentBlock` (from `dsh-llm`). - -This is not a coupling smell — it is the contract's domain. The "only cordis" guidance was always shorthand for "the interface depends only on what the contract genuinely names, and never on an implementation." `dsh-session` and `dsh-llm` are themselves interface/vocabulary packages, not implementations; `dsh-compact` still imports no backend. The seam's real invariant — *consumers and implementations evolve independently behind an abstract service* — holds intact. We record the deviation here so a future reader doesn't mistake it for an accident or "fix" it by smuggling `Session` behind an opaque handle. - -### Abstract `compactIfNeeded` / `compactRegion`, algorithm in the backend - -An earlier draft put the full algorithm (the retention walk, token-summing, text extraction) as concrete methods on the interface, with only `estimateContentTokens()` and `summarize()` abstract. That recouples the contract to one strategy: a backend that wants a different retention policy (e.g. turn-count instead of token-budget) or a different event-sequencing would have to fight inherited concrete code. Making both core methods abstract puts every *how* decision in the backend, where it belongs, and keeps the interface a pure statement of *what*. The backend remains internally factored — `estimateContentTokens()` and `summarize()` are `protected` hooks a sub-backend can override without reimplementing the walk — but that factoring is the backend's private concern, not the contract's. - -### Surface replacement: `compact/*` events are log-only; one `user/message` carries the summary - -Because `SurfaceEventType` is closed, the summary cannot ride on a `compact/*` event. The backend instead appends a **single `user/message`** with `surfaceOp: { op: 'replace', start, end }` whose `content` is the summary `ContentBlock[]` and whose `sourceEventSeqs` covers the shadowed nodes *and* the bookkeeping events. The `compact/*` events are pure log records (lock + provenance), never on the surface. The surface mutation sits **inside** the lock — `compact/end` is the last event appended: - -``` -compact/start → log-only. Acquires the lock. -[summarize older range via the backend] -compact/summary → log-only. Provenance: summary, range, shadowed seqs, token count. -user/message → surfaceOp { op:'replace', start, end }. THE surface mutation. - deriveMessages() renders it as a user-role message. -compact/end → log-only. Releases the lock. -``` - -Ordering the surface mutation **before** `compact/end` is deliberate: `session.append()` commits one event at a time, so there is no multi-event transaction to make the sequence atomic. Releasing the lock last converts the crash window from *silent corruption* (a `compact/end` that claims compaction finished while the surface was never shadowed) into a *detectable orphaned lock* (a `compact/start` with no matching `compact/end`), which a persistence backend already detects on reload. A `session/event` listener on `compact/end` likewise never sees the lock free before the replacement has landed. - -`deriveMessages()` then yields `[summary_as_user_message, ...retained_nodes]`. An alternative — extending `SurfaceEventType` to admit a `compact/*` type — was rejected: the closed union is a deliberate safety boundary (only message-producing events reach the model), and a summary genuinely *is* user-role context, so reusing `user/message` is honest rather than a workaround. - -### Blocking via a log-recorded lock, not a mutex - -Compaction must be serialized: no second compaction starts before the first finishes, and no ordinary events interleave the slow summarization. Rather than an in-memory mutex (invisible to replay, lost on crash), the lock **is** the log: `compactRegion` refuses to start if the last `compact/start` has no matching `compact/end` after it. `compact/start` is appended first (fast, synchronous), the slow model call runs, then the `compact/summary` and `user/message` replacement land, and only then is `compact/end` appended — in a `catch` that records the error, so a failed summarization can never wedge the lock. Because the backend runs compaction synchronously inside the `agent/request` waterfall, the loop is single-threaded for that window; the lock additionally gives observability and lets a persistence backend detect an orphaned `compact/start` on reload. - -## Consequences - -- **New packages**: `packages/compact/compact` (interface) and a sibling `compact-basic` (backend) under `packages/compact/`, wired into the three root tsconfigs. The consumer tier is deferred. -- **`SessionEventMap`** gains `compact/start` / `compact/summary` / `compact/end` by declaration merging (merge-extensible); `SurfaceEventType` is **not** touched. These are session events, not cordis `Events`, so the event-taxonomy gate needs no entry. -- **No changes** to `dsh-session`, `dsh-invariants`, or `dsh-agent-loop`: the surface replace op, the surface-metadata runtime guard, and the `agent/request` waterfall all already exist. Compaction is a pure plugin on documented seams. -- The capability-seams convention gains a second reference beyond bash, and a documented case where "interface depends only on cordis" relaxes to "depends only on interface/vocabulary packages the contract genuinely names." On acceptance, [AGENTS.md](../../../../AGENTS.md) § Conventions and [architecture.md](../../../architecture.md) § "Capability seams" should note this relaxation. diff --git a/examples/AGENTS.md b/examples/AGENTS.md index 67ea68ae6f..c5a797f5df 100644 --- a/examples/AGENTS.md +++ b/examples/AGENTS.md @@ -20,7 +20,7 @@ A keyless smoke that spawns the example from a temp cwd must set `TSX_TSCONFIG_P | Example | Keyless smoke | With-key smoke | |---|---|---| | `echo-agent` | `tests/echo.e2e.ts` — boots the real `cordis.yml`, drives the echo tool round-trip and the direct canned reply | **N/A — keyless by nature** (the `mock-echo` model has no real provider) | -| `coding-agent` | `tests/keyless-smoke.e2e.ts` — boots the full real tree (dummy key, no prompt → no model call), asserts banner + clean exit | `tests/{full-loop,coding-task,resume}.e2e.ts` — real model + real bash, world-verified | +| `coding-agent` | `tests/keyless-smoke.e2e.ts` — boots the full real tree (dummy key, no prompt → no model call), asserts banner + clean exit | `tests/{full-loop,coding-task,resume,compaction,todo-write}.e2e.ts` — real model + real bash + real todo_write, world-verified | | `acp-agent` | `pnpm run test:snapshot` — boots the real ACP subprocess and replays a recorded session keyless; `tests/acp.e2e.ts` also asserts stdout purity without a key | `tests/acp.e2e.ts` — real ACP prompt, verifies a file the agent wrote | See [the root AGENTS.md](../AGENTS.md) for repo-wide conventions and [docs/architecture.md](../docs/architecture.md) for the design. diff --git a/examples/README.md b/examples/README.md index 3fd9259e36..887bc18beb 100644 --- a/examples/README.md +++ b/examples/README.md @@ -1,6 +1,6 @@ # Examples -Runnable demos (not workspaces) that showcase how the harness is wired. Each example is now a **thin leaf**: a `cordis.yml` that picks the swappable backends (an LLM adapter, a bash executor) and loads ONE app package, plus any demo-only mocks. The composition — the spine, the front-door cluster, and the boot glue — lives in the app packages ([`@deepseek-ai/dsh-stdio-agent`](../packages/ui/stdio-agent), [`@deepseek-ai/dsh-acp-agent`](../packages/ui/acp-agent)) and the [`@deepseek-ai/dsh-agent-core`](../packages/core/agent-core) bundle they share. There is no `start.ts`; the `demo:*` scripts invoke each app package's `bin`. +Runnable demos (not workspaces) that showcase how the harness is wired. Each example is now a **thin leaf**: a `cordis.yml` that picks the swappable backends (an LLM adapter, a bash executor), loads ONE app package, and may add optional product tools or demo-only mocks. The composition — the spine, the front-door cluster, and the boot glue — lives in the app packages ([`@deepseek-ai/dsh-stdio-agent`](../packages/ui/stdio-agent), [`@deepseek-ai/dsh-acp-agent`](../packages/ui/acp-agent)) and the [`@deepseek-ai/dsh-agent-core`](../packages/core/agent-core) bundle they share. There is no `start.ts`; the `demo:*` scripts invoke each app package's `bin`. ## echo-agent @@ -15,7 +15,7 @@ Run with: `pnpm run demo:echo`. When prompted, type "echo " to trigge ## coding-agent -The real thing: DeepSeek V4 + the bash tool suite on the same `@deepseek-ai/dsh-stdio-agent` app. Where echo-agent proves the skeleton with mocks, this is a usable coding assistant. +The real thing: DeepSeek V4 + the bash tool suite, `subagent` delegation, and the `todo_write` task tracker on the same `@deepseek-ai/dsh-stdio-agent` app. Where echo-agent proves the skeleton with mocks, this is a usable coding assistant. Run with: `pnpm run demo:coding` (needs `DEEPSEEK_API_KEY` in the environment or a gitignored repo-root `.env`). See [coding-agent/README.md](coding-agent/README.md) for details. diff --git a/examples/acp-agent/README.md b/examples/acp-agent/README.md index 1df36715ec..58ba92cd0e 100644 --- a/examples/acp-agent/README.md +++ b/examples/acp-agent/README.md @@ -6,7 +6,7 @@ The DeepSeek Harness coding agent exposed as an **Agent Client Protocol (ACP)** pnpm run demo:acp # needs DEEPSEEK_API_KEY (repo-root .env or env) ``` -This example is just a leaf `cordis.yml`: it loads the [`@deepseek-ai/dsh-acp-agent`](../../packages/ui/acp-agent) app (which bundles the [`@deepseek-ai/dsh-agent-core`](../../packages/core/agent-core) spine, JSONL session persistence, and the `@deepseek-ai/dsh-acp` bridge — with **no pre-created agents**, since ACP `session/new` creates them on demand) plus the two swappable backends (`llm-deepseek`, `bash-local`). The app package bakes in the no-stdout-logger cluster, so a leaf has no logger entry to get wrong by default — keeping stdout pure for JSON-RPC. +This example is just a leaf `cordis.yml`: it loads the [`@deepseek-ai/dsh-acp-agent`](../../packages/ui/acp-agent) app (which bundles the [`@deepseek-ai/dsh-agent-core`](../../packages/core/agent-core) spine, JSONL session persistence, and the `@deepseek-ai/dsh-acp` bridge — with **no pre-created agents**, since ACP `session/new` creates them on demand), the swappable DeepSeek and bash backends, and the optional model-facing `subagent`/`subagent_fork`/`todo_write` tool entries. The app package bakes in the no-stdout-logger cluster, so a leaf has no logger entry to get wrong by default — keeping stdout pure for JSON-RPC. ## stdout is the protocol @@ -32,7 +32,7 @@ The editor sets each session's `cwd` to the project it opens; the agent's bash t ## Snapshot tests (record-once / replay-deterministic) -This example is the home of the harness's **snapshot tests** — they boot this server as a real subprocess, drive it with a deterministic input script, and diff its normalized output against committed golden files. The model is made deterministic by `@deepseek-ai/dsh-llm-replay`, a function/namespace plugin that installs an `llm/stream` waterfall listener and short-circuits it, serving model streams reconstructed from a recorded **session JSONL** fixture (`/session.jsonl`) — so replay needs no API key. The fixture IS the persisted session log: its `assistant/chunk` events carry every `StreamChunk`, so grouping them by `(turn, step)` reconstructs each `stream()` call (one model call per loop step). Recording is therefore "run the real agent once and harvest the `.jsonl`". The two failure modes not expressible as logged chunks — a pure throw before any chunk, and cancel/hang — use an optional `/replay.override.json` sidecar (a `ReplayEntry[]` that replaces the derived script). A scenario that needs the agent to operate on existing files ships an optional `/workspace/` directory — the harness copies its contents into the temp cwd before the run (see `workspace-edit`). See [docs/rfc/implemented/2026-06-19-acp-snapshot-tests.md](../../docs/rfc/implemented/2026-06-19-acp-snapshot-tests.md) for the full design. +This example is the home of the harness's **snapshot tests** — they boot this server as a real subprocess, drive it with a deterministic input script, and diff its normalized output against committed golden files. The model is made deterministic by `@deepseek-ai/dsh-llm-replay`, a function/namespace plugin that installs an `llm/stream` waterfall listener and short-circuits it, serving model streams reconstructed from a recorded **session JSONL** fixture (`/session.jsonl`) — so replay needs no API key. The fixture IS the persisted session log: its `assistant/chunk` events carry every `StreamChunk`, so grouping them by `(turn, step)` reconstructs each `stream()` call (one model call per loop step). Recording is therefore "run the real agent once and harvest the `.jsonl`". The two failure modes not expressible as logged chunks — a pure throw before any chunk, and cancel/hang — use an optional `/replay.override.json` sidecar (a `ReplayEntry[]` that replaces the derived script). A scenario that needs the agent to operate on existing files ships an optional `/workspace/` directory — the harness copies its contents into the temp cwd before the run (see `workspace-edit`). See [the ACP snapshot tests RFC](../../docs/rfc/implemented/testing/2026-06-19-acp-snapshot-tests.md) for the full design. ## MVP limitations diff --git a/examples/acp-agent/cordis.snapshot.yml b/examples/acp-agent/cordis.snapshot.yml index 5f36a11efa..f03cc49223 100644 --- a/examples/acp-agent/cordis.snapshot.yml +++ b/examples/acp-agent/cordis.snapshot.yml @@ -16,7 +16,9 @@ - id: llm-replay name: '@deepseek-ai/dsh-llm-replay' -# Local bash executor (the agent's only tool, via agent-core's tool-bash schema). +# Local bash executor for agent-core's tool-bash schema. +# FIXME(config-comments): keep this executor note from implying bash is the +# whole tool set; subagent and todo_write are loaded below. - id: bash name: '@deepseek-ai/dsh-bash-local' config: @@ -40,7 +42,15 @@ Use the subagent tool to delegate a focused, self-contained subtask to a fresh child agent (it works in its own context and returns only its - final result) — give it a complete, standalone instruction. + final result) — give it a complete, standalone instruction. Use + subagent_fork instead when the subtask needs THIS conversation's + context: the child inherits the log so far. + + For multi-step work, use the todo_write tool to track a task list: + send the WHOLE list each call (it replaces the previous one), keep at + most one task in_progress (exactly one while work remains), and mark a + task completed as soon as it is done. Skip it for trivial single-step + tasks. # The subagent seam + both in-process backends + two model-facing tools — # identical to cordis.yml's wiring (only the LLM backend differs above): spawn @@ -70,3 +80,8 @@ config: provider: fork toolName: subagent_fork + +# The model-facing todo_write tool — identical to cordis.yml's wiring, so a +# replayed todo_write tool call resolves to a real tool during snapshot replay. +- id: tool-todo + name: '@deepseek-ai/dsh-tool-todo' diff --git a/examples/acp-agent/cordis.yml b/examples/acp-agent/cordis.yml index a00d0e6036..31d4d5429a 100644 --- a/examples/acp-agent/cordis.yml +++ b/examples/acp-agent/cordis.yml @@ -1,9 +1,8 @@ # The acp-agent plugin tree: the ACP server. Also the snapshot RECORD config # (the dsh-acp-agent bin selects it for DSH_SNAPSHOT=record): a real llm-deepseek -# run whose persisted log the snapshot harness harvests. Just the two swappable -# backends — the DeepSeek adapter and the local bash executor — plus the ACP -# server app (@deepseek-ai/dsh-acp-agent), which bundles the agent-core spine, -# JSONL persistence, and the ACP bridge. +# run whose persisted log the snapshot harness harvests. The swappable DeepSeek +# adapter and local bash executor, the ACP server app (@deepseek-ai/dsh-acp-agent), +# and the optional model-facing subagent/todo tools loaded below. # # CRITICAL: this tree loads NO stdout logger and NO hmr — stdout is reserved for # the ACP JSON-RPC protocol (see packages/ui/acp). That guarantee is now a @@ -23,7 +22,9 @@ - deepseek-v4-flash - deepseek-v4-pro -# Local bash executor (the agent's only tool, via agent-core's tool-bash schema). +# Local bash executor for agent-core's tool-bash schema. +# FIXME(config-comments): keep this executor note from implying bash is the +# whole tool set; subagent and todo_write are loaded below. - id: bash name: '@deepseek-ai/dsh-bash-local' config: @@ -49,7 +50,15 @@ Use the subagent tool to delegate a focused, self-contained subtask to a fresh child agent (it works in its own context and returns only its - final result) — give it a complete, standalone instruction. + final result) — give it a complete, standalone instruction. Use + subagent_fork instead when the subtask needs THIS conversation's + context: the child inherits the log so far. + + For multi-step work, use the todo_write tool to track a task list: + send the WHOLE list each call (it replaces the previous one), keep at + most one task in_progress (exactly one while work remains), and mark a + task completed as soon as it is done. Skip it for trivial single-step + tasks. # The subagent seam + both in-process backends + two model-facing tools, as leaf # entries after the app (which provides ctx.agents/ctx.tools). spawn (a fresh @@ -81,3 +90,8 @@ config: provider: fork toolName: subagent_fork + +# The model-facing todo_write tool: whole-list task tracking written to the +# session log (todo/write), surfaced to the ACP client as a `plan` update. +- id: tool-todo + name: '@deepseek-ai/dsh-tool-todo' diff --git a/examples/acp-agent/tests/acp.snapshot.ts b/examples/acp-agent/tests/acp.snapshot.ts index 42e9107c04..94f18e64d3 100644 --- a/examples/acp-agent/tests/acp.snapshot.ts +++ b/examples/acp-agent/tests/acp.snapshot.ts @@ -52,6 +52,7 @@ const SCENARIOS: Scenario[] = [ { name: 'reject-extra-dirs', hasModelTurn: false, recorded: false }, { name: 'text-turn', hasModelTurn: true, recorded: true }, { name: 'tool-call-turn', hasModelTurn: true, recorded: true }, + { name: 'todo-plan', hasModelTurn: true, recorded: true }, { name: 'workspace-edit', hasModelTurn: true, recorded: true }, { name: 'multi-turn', hasModelTurn: true, recorded: true }, { name: 'error-finish', hasModelTurn: true, recorded: false }, diff --git a/examples/acp-agent/tests/snapshots/todo-plan/input.json b/examples/acp-agent/tests/snapshots/todo-plan/input.json new file mode 100644 index 0000000000..6cc82bdcae --- /dev/null +++ b/examples/acp-agent/tests/snapshots/todo-plan/input.json @@ -0,0 +1,7 @@ +{ + "steps": [ + { "op": "initialize" }, + { "op": "newSession" }, + { "op": "prompt", "text": "Use the todo_write tool to record a plan with exactly three todos: \"read the code\" (in_progress), \"write the fix\" (pending), \"run the tests\" (pending). Send all three in one todo_write call. Then reply with the single word DONE and stop." } + ] +} diff --git a/examples/acp-agent/tests/snapshots/todo-plan/session.jsonl b/examples/acp-agent/tests/snapshots/todo-plan/session.jsonl new file mode 100644 index 0000000000..6fc52b6a4b --- /dev/null +++ b/examples/acp-agent/tests/snapshots/todo-plan/session.jsonl @@ -0,0 +1,124 @@ +{"type":"session","version":0,"id":"259ed557-03cf-4f50-9592-fc7fdbece7f3","createdAt":1782701599718,"cwd":"/tmp/acp-snap-cwd-4xZzZ9"} +{"type":"turn/start","seq":0,"time":1782701599722,"data":{"turn":1,"trigger":{"kind":"message","source":{"kind":"user"}}}} +{"type":"user/message","seq":1,"time":1782701599722,"data":{"content":[{"type":"text","text":"Use the todo_write tool to record a plan with exactly three todos: \"read the code\" (in_progress), \"write the fix\" (pending), \"run the tests\" (pending). Send all three in one todo_write call. Then reply with the single word DONE and stop."}],"source":{"kind":"user"}},"surfaceOp":"append"} +{"type":"step/start","seq":2,"time":1782701599722,"data":{"turn":1,"step":1}} +{"type":"assistant/chunk","seq":3,"time":1782701600164,"data":{"turn":1,"step":1,"chunk":{"type":"block-start","index":0,"blockType":"reasoning"}}} +{"type":"assistant/chunk","seq":4,"time":1782701600164,"data":{"turn":1,"step":1,"chunk":{"type":"reasoning-delta","index":0,"text":"The"}}} +{"type":"assistant/chunk","seq":5,"time":1782701600271,"data":{"turn":1,"step":1,"chunk":{"type":"reasoning-delta","index":0,"text":" user"}}} +{"type":"assistant/chunk","seq":6,"time":1782701600298,"data":{"turn":1,"step":1,"chunk":{"type":"reasoning-delta","index":0,"text":" wants"}}} +{"type":"assistant/chunk","seq":7,"time":1782701600299,"data":{"turn":1,"step":1,"chunk":{"type":"reasoning-delta","index":0,"text":" me"}}} +{"type":"assistant/chunk","seq":8,"time":1782701600299,"data":{"turn":1,"step":1,"chunk":{"type":"reasoning-delta","index":0,"text":" to"}}} +{"type":"assistant/chunk","seq":9,"time":1782701600299,"data":{"turn":1,"step":1,"chunk":{"type":"reasoning-delta","index":0,"text":" use"}}} +{"type":"assistant/chunk","seq":10,"time":1782701600299,"data":{"turn":1,"step":1,"chunk":{"type":"reasoning-delta","index":0,"text":" todo"}}} +{"type":"assistant/chunk","seq":11,"time":1782701600325,"data":{"turn":1,"step":1,"chunk":{"type":"reasoning-delta","index":0,"text":"_write"}}} +{"type":"assistant/chunk","seq":12,"time":1782701600326,"data":{"turn":1,"step":1,"chunk":{"type":"reasoning-delta","index":0,"text":" to"}}} +{"type":"assistant/chunk","seq":13,"time":1782701600326,"data":{"turn":1,"step":1,"chunk":{"type":"reasoning-delta","index":0,"text":" record"}}} +{"type":"assistant/chunk","seq":14,"time":1782701600354,"data":{"turn":1,"step":1,"chunk":{"type":"reasoning-delta","index":0,"text":" a"}}} +{"type":"assistant/chunk","seq":15,"time":1782701600354,"data":{"turn":1,"step":1,"chunk":{"type":"reasoning-delta","index":0,"text":" plan"}}} +{"type":"assistant/chunk","seq":16,"time":1782701600355,"data":{"turn":1,"step":1,"chunk":{"type":"reasoning-delta","index":0,"text":" with"}}} +{"type":"assistant/chunk","seq":17,"time":1782701600355,"data":{"turn":1,"step":1,"chunk":{"type":"reasoning-delta","index":0,"text":" exactly"}}} +{"type":"assistant/chunk","seq":18,"time":1782701600355,"data":{"turn":1,"step":1,"chunk":{"type":"reasoning-delta","index":0,"text":" three"}}} +{"type":"assistant/chunk","seq":19,"time":1782701600383,"data":{"turn":1,"step":1,"chunk":{"type":"reasoning-delta","index":0,"text":" todos"}}} +{"type":"assistant/chunk","seq":20,"time":1782701600383,"data":{"turn":1,"step":1,"chunk":{"type":"reasoning-delta","index":0,"text":","}}} +{"type":"assistant/chunk","seq":21,"time":1782701600409,"data":{"turn":1,"step":1,"chunk":{"type":"reasoning-delta","index":0,"text":" then"}}} +{"type":"assistant/chunk","seq":22,"time":1782701600438,"data":{"turn":1,"step":1,"chunk":{"type":"reasoning-delta","index":0,"text":" reply"}}} +{"type":"assistant/chunk","seq":23,"time":1782701600438,"data":{"turn":1,"step":1,"chunk":{"type":"reasoning-delta","index":0,"text":" with"}}} +{"type":"assistant/chunk","seq":24,"time":1782701600438,"data":{"turn":1,"step":1,"chunk":{"type":"reasoning-delta","index":0,"text":" D"}}} +{"type":"assistant/chunk","seq":25,"time":1782701600465,"data":{"turn":1,"step":1,"chunk":{"type":"reasoning-delta","index":0,"text":"ONE"}}} +{"type":"assistant/chunk","seq":26,"time":1782701600466,"data":{"turn":1,"step":1,"chunk":{"type":"reasoning-delta","index":0,"text":"."}}} +{"type":"assistant/chunk","seq":27,"time":1782701600568,"data":{"turn":1,"step":1,"chunk":{"type":"block-start","index":1,"blockType":"tool-call"}}} +{"type":"assistant/chunk","seq":28,"time":1782701600568,"data":{"turn":1,"step":1,"chunk":{"type":"tool-call-delta","index":1,"id":"call_00_OK2ZF1DYrsKHQQtxfQlJ0810","name":"todo_write","argumentsDelta":""}}} +{"type":"assistant/chunk","seq":29,"time":1782701600568,"data":{"turn":1,"step":1,"chunk":{"type":"tool-call-delta","index":1,"id":"call_00_OK2ZF1DYrsKHQQtxfQlJ0810","name":"todo_write","argumentsDelta":"{"}}} +{"type":"assistant/chunk","seq":30,"time":1782701600568,"data":{"turn":1,"step":1,"chunk":{"type":"tool-call-delta","index":1,"id":"call_00_OK2ZF1DYrsKHQQtxfQlJ0810","name":"todo_write","argumentsDelta":"\""}}} +{"type":"assistant/chunk","seq":31,"time":1782701600576,"data":{"turn":1,"step":1,"chunk":{"type":"tool-call-delta","index":1,"id":"call_00_OK2ZF1DYrsKHQQtxfQlJ0810","name":"todo_write","argumentsDelta":"t"}}} +{"type":"assistant/chunk","seq":32,"time":1782701600577,"data":{"turn":1,"step":1,"chunk":{"type":"tool-call-delta","index":1,"id":"call_00_OK2ZF1DYrsKHQQtxfQlJ0810","name":"todo_write","argumentsDelta":"odos"}}} +{"type":"assistant/chunk","seq":33,"time":1782701600577,"data":{"turn":1,"step":1,"chunk":{"type":"tool-call-delta","index":1,"id":"call_00_OK2ZF1DYrsKHQQtxfQlJ0810","name":"todo_write","argumentsDelta":"\""}}} +{"type":"assistant/chunk","seq":34,"time":1782701600577,"data":{"turn":1,"step":1,"chunk":{"type":"tool-call-delta","index":1,"id":"call_00_OK2ZF1DYrsKHQQtxfQlJ0810","name":"todo_write","argumentsDelta":": "}}} +{"type":"assistant/chunk","seq":35,"time":1782701600604,"data":{"turn":1,"step":1,"chunk":{"type":"tool-call-delta","index":1,"id":"call_00_OK2ZF1DYrsKHQQtxfQlJ0810","name":"todo_write","argumentsDelta":"["}}} +{"type":"assistant/chunk","seq":36,"time":1782701600605,"data":{"turn":1,"step":1,"chunk":{"type":"tool-call-delta","index":1,"id":"call_00_OK2ZF1DYrsKHQQtxfQlJ0810","name":"todo_write","argumentsDelta":"{\""}}} +{"type":"assistant/chunk","seq":37,"time":1782701600605,"data":{"turn":1,"step":1,"chunk":{"type":"tool-call-delta","index":1,"id":"call_00_OK2ZF1DYrsKHQQtxfQlJ0810","name":"todo_write","argumentsDelta":"content"}}} +{"type":"assistant/chunk","seq":38,"time":1782701600605,"data":{"turn":1,"step":1,"chunk":{"type":"tool-call-delta","index":1,"id":"call_00_OK2ZF1DYrsKHQQtxfQlJ0810","name":"todo_write","argumentsDelta":"\":"}}} +{"type":"assistant/chunk","seq":39,"time":1782701600605,"data":{"turn":1,"step":1,"chunk":{"type":"tool-call-delta","index":1,"id":"call_00_OK2ZF1DYrsKHQQtxfQlJ0810","name":"todo_write","argumentsDelta":" \""}}} +{"type":"assistant/chunk","seq":40,"time":1782701600631,"data":{"turn":1,"step":1,"chunk":{"type":"tool-call-delta","index":1,"id":"call_00_OK2ZF1DYrsKHQQtxfQlJ0810","name":"todo_write","argumentsDelta":"read"}}} +{"type":"assistant/chunk","seq":41,"time":1782701600631,"data":{"turn":1,"step":1,"chunk":{"type":"tool-call-delta","index":1,"id":"call_00_OK2ZF1DYrsKHQQtxfQlJ0810","name":"todo_write","argumentsDelta":" the"}}} +{"type":"assistant/chunk","seq":42,"time":1782701600632,"data":{"turn":1,"step":1,"chunk":{"type":"tool-call-delta","index":1,"id":"call_00_OK2ZF1DYrsKHQQtxfQlJ0810","name":"todo_write","argumentsDelta":" code"}}} +{"type":"assistant/chunk","seq":43,"time":1782701600632,"data":{"turn":1,"step":1,"chunk":{"type":"tool-call-delta","index":1,"id":"call_00_OK2ZF1DYrsKHQQtxfQlJ0810","name":"todo_write","argumentsDelta":"\","}}} +{"type":"assistant/chunk","seq":44,"time":1782701600632,"data":{"turn":1,"step":1,"chunk":{"type":"tool-call-delta","index":1,"id":"call_00_OK2ZF1DYrsKHQQtxfQlJ0810","name":"todo_write","argumentsDelta":" \""}}} +{"type":"assistant/chunk","seq":45,"time":1782701600632,"data":{"turn":1,"step":1,"chunk":{"type":"tool-call-delta","index":1,"id":"call_00_OK2ZF1DYrsKHQQtxfQlJ0810","name":"todo_write","argumentsDelta":"status"}}} +{"type":"assistant/chunk","seq":46,"time":1782701600659,"data":{"turn":1,"step":1,"chunk":{"type":"tool-call-delta","index":1,"id":"call_00_OK2ZF1DYrsKHQQtxfQlJ0810","name":"todo_write","argumentsDelta":"\":"}}} +{"type":"assistant/chunk","seq":47,"time":1782701600660,"data":{"turn":1,"step":1,"chunk":{"type":"tool-call-delta","index":1,"id":"call_00_OK2ZF1DYrsKHQQtxfQlJ0810","name":"todo_write","argumentsDelta":" \""}}} +{"type":"assistant/chunk","seq":48,"time":1782701600660,"data":{"turn":1,"step":1,"chunk":{"type":"tool-call-delta","index":1,"id":"call_00_OK2ZF1DYrsKHQQtxfQlJ0810","name":"todo_write","argumentsDelta":"in"}}} +{"type":"assistant/chunk","seq":49,"time":1782701600660,"data":{"turn":1,"step":1,"chunk":{"type":"tool-call-delta","index":1,"id":"call_00_OK2ZF1DYrsKHQQtxfQlJ0810","name":"todo_write","argumentsDelta":"_pro"}}} +{"type":"assistant/chunk","seq":50,"time":1782701600660,"data":{"turn":1,"step":1,"chunk":{"type":"tool-call-delta","index":1,"id":"call_00_OK2ZF1DYrsKHQQtxfQlJ0810","name":"todo_write","argumentsDelta":"gress"}}} +{"type":"assistant/chunk","seq":51,"time":1782701600660,"data":{"turn":1,"step":1,"chunk":{"type":"tool-call-delta","index":1,"id":"call_00_OK2ZF1DYrsKHQQtxfQlJ0810","name":"todo_write","argumentsDelta":"\"},"}}} +{"type":"assistant/chunk","seq":52,"time":1782701600686,"data":{"turn":1,"step":1,"chunk":{"type":"tool-call-delta","index":1,"id":"call_00_OK2ZF1DYrsKHQQtxfQlJ0810","name":"todo_write","argumentsDelta":" {\""}}} +{"type":"assistant/chunk","seq":53,"time":1782701600687,"data":{"turn":1,"step":1,"chunk":{"type":"tool-call-delta","index":1,"id":"call_00_OK2ZF1DYrsKHQQtxfQlJ0810","name":"todo_write","argumentsDelta":"content"}}} +{"type":"assistant/chunk","seq":54,"time":1782701600687,"data":{"turn":1,"step":1,"chunk":{"type":"tool-call-delta","index":1,"id":"call_00_OK2ZF1DYrsKHQQtxfQlJ0810","name":"todo_write","argumentsDelta":"\":"}}} +{"type":"assistant/chunk","seq":55,"time":1782701600687,"data":{"turn":1,"step":1,"chunk":{"type":"tool-call-delta","index":1,"id":"call_00_OK2ZF1DYrsKHQQtxfQlJ0810","name":"todo_write","argumentsDelta":" \""}}} +{"type":"assistant/chunk","seq":56,"time":1782701600687,"data":{"turn":1,"step":1,"chunk":{"type":"tool-call-delta","index":1,"id":"call_00_OK2ZF1DYrsKHQQtxfQlJ0810","name":"todo_write","argumentsDelta":"write"}}} +{"type":"assistant/chunk","seq":57,"time":1782701600687,"data":{"turn":1,"step":1,"chunk":{"type":"tool-call-delta","index":1,"id":"call_00_OK2ZF1DYrsKHQQtxfQlJ0810","name":"todo_write","argumentsDelta":" the"}}} +{"type":"assistant/chunk","seq":58,"time":1782701600715,"data":{"turn":1,"step":1,"chunk":{"type":"tool-call-delta","index":1,"id":"call_00_OK2ZF1DYrsKHQQtxfQlJ0810","name":"todo_write","argumentsDelta":" fix"}}} +{"type":"assistant/chunk","seq":59,"time":1782701600715,"data":{"turn":1,"step":1,"chunk":{"type":"tool-call-delta","index":1,"id":"call_00_OK2ZF1DYrsKHQQtxfQlJ0810","name":"todo_write","argumentsDelta":"\","}}} +{"type":"assistant/chunk","seq":60,"time":1782701600715,"data":{"turn":1,"step":1,"chunk":{"type":"tool-call-delta","index":1,"id":"call_00_OK2ZF1DYrsKHQQtxfQlJ0810","name":"todo_write","argumentsDelta":" \""}}} +{"type":"assistant/chunk","seq":61,"time":1782701600715,"data":{"turn":1,"step":1,"chunk":{"type":"tool-call-delta","index":1,"id":"call_00_OK2ZF1DYrsKHQQtxfQlJ0810","name":"todo_write","argumentsDelta":"status"}}} +{"type":"assistant/chunk","seq":62,"time":1782701600716,"data":{"turn":1,"step":1,"chunk":{"type":"tool-call-delta","index":1,"id":"call_00_OK2ZF1DYrsKHQQtxfQlJ0810","name":"todo_write","argumentsDelta":"\":"}}} +{"type":"assistant/chunk","seq":63,"time":1782701600716,"data":{"turn":1,"step":1,"chunk":{"type":"tool-call-delta","index":1,"id":"call_00_OK2ZF1DYrsKHQQtxfQlJ0810","name":"todo_write","argumentsDelta":" \""}}} +{"type":"assistant/chunk","seq":64,"time":1782701600744,"data":{"turn":1,"step":1,"chunk":{"type":"tool-call-delta","index":1,"id":"call_00_OK2ZF1DYrsKHQQtxfQlJ0810","name":"todo_write","argumentsDelta":"pending"}}} +{"type":"assistant/chunk","seq":65,"time":1782701600744,"data":{"turn":1,"step":1,"chunk":{"type":"tool-call-delta","index":1,"id":"call_00_OK2ZF1DYrsKHQQtxfQlJ0810","name":"todo_write","argumentsDelta":"\"},"}}} +{"type":"assistant/chunk","seq":66,"time":1782701600744,"data":{"turn":1,"step":1,"chunk":{"type":"tool-call-delta","index":1,"id":"call_00_OK2ZF1DYrsKHQQtxfQlJ0810","name":"todo_write","argumentsDelta":" {\""}}} +{"type":"assistant/chunk","seq":67,"time":1782701600744,"data":{"turn":1,"step":1,"chunk":{"type":"tool-call-delta","index":1,"id":"call_00_OK2ZF1DYrsKHQQtxfQlJ0810","name":"todo_write","argumentsDelta":"content"}}} +{"type":"assistant/chunk","seq":68,"time":1782701600744,"data":{"turn":1,"step":1,"chunk":{"type":"tool-call-delta","index":1,"id":"call_00_OK2ZF1DYrsKHQQtxfQlJ0810","name":"todo_write","argumentsDelta":"\":"}}} +{"type":"assistant/chunk","seq":69,"time":1782701600744,"data":{"turn":1,"step":1,"chunk":{"type":"tool-call-delta","index":1,"id":"call_00_OK2ZF1DYrsKHQQtxfQlJ0810","name":"todo_write","argumentsDelta":" \""}}} +{"type":"assistant/chunk","seq":70,"time":1782701600770,"data":{"turn":1,"step":1,"chunk":{"type":"tool-call-delta","index":1,"id":"call_00_OK2ZF1DYrsKHQQtxfQlJ0810","name":"todo_write","argumentsDelta":"run"}}} +{"type":"assistant/chunk","seq":71,"time":1782701600770,"data":{"turn":1,"step":1,"chunk":{"type":"tool-call-delta","index":1,"id":"call_00_OK2ZF1DYrsKHQQtxfQlJ0810","name":"todo_write","argumentsDelta":" the"}}} +{"type":"assistant/chunk","seq":72,"time":1782701600770,"data":{"turn":1,"step":1,"chunk":{"type":"tool-call-delta","index":1,"id":"call_00_OK2ZF1DYrsKHQQtxfQlJ0810","name":"todo_write","argumentsDelta":" tests"}}} +{"type":"assistant/chunk","seq":73,"time":1782701600770,"data":{"turn":1,"step":1,"chunk":{"type":"tool-call-delta","index":1,"id":"call_00_OK2ZF1DYrsKHQQtxfQlJ0810","name":"todo_write","argumentsDelta":"\","}}} +{"type":"assistant/chunk","seq":74,"time":1782701600770,"data":{"turn":1,"step":1,"chunk":{"type":"tool-call-delta","index":1,"id":"call_00_OK2ZF1DYrsKHQQtxfQlJ0810","name":"todo_write","argumentsDelta":" \""}}} +{"type":"assistant/chunk","seq":75,"time":1782701600771,"data":{"turn":1,"step":1,"chunk":{"type":"tool-call-delta","index":1,"id":"call_00_OK2ZF1DYrsKHQQtxfQlJ0810","name":"todo_write","argumentsDelta":"status"}}} +{"type":"assistant/chunk","seq":76,"time":1782701600798,"data":{"turn":1,"step":1,"chunk":{"type":"tool-call-delta","index":1,"id":"call_00_OK2ZF1DYrsKHQQtxfQlJ0810","name":"todo_write","argumentsDelta":"\":"}}} +{"type":"assistant/chunk","seq":77,"time":1782701600799,"data":{"turn":1,"step":1,"chunk":{"type":"tool-call-delta","index":1,"id":"call_00_OK2ZF1DYrsKHQQtxfQlJ0810","name":"todo_write","argumentsDelta":" \""}}} +{"type":"assistant/chunk","seq":78,"time":1782701600799,"data":{"turn":1,"step":1,"chunk":{"type":"tool-call-delta","index":1,"id":"call_00_OK2ZF1DYrsKHQQtxfQlJ0810","name":"todo_write","argumentsDelta":"pending"}}} +{"type":"assistant/chunk","seq":79,"time":1782701600799,"data":{"turn":1,"step":1,"chunk":{"type":"tool-call-delta","index":1,"id":"call_00_OK2ZF1DYrsKHQQtxfQlJ0810","name":"todo_write","argumentsDelta":"\""}}} +{"type":"assistant/chunk","seq":80,"time":1782701600799,"data":{"turn":1,"step":1,"chunk":{"type":"tool-call-delta","index":1,"id":"call_00_OK2ZF1DYrsKHQQtxfQlJ0810","name":"todo_write","argumentsDelta":"}]"}}} +{"type":"assistant/chunk","seq":81,"time":1782701600826,"data":{"turn":1,"step":1,"chunk":{"type":"tool-call-delta","index":1,"id":"call_00_OK2ZF1DYrsKHQQtxfQlJ0810","name":"todo_write","argumentsDelta":"}"}}} +{"type":"assistant/chunk","seq":82,"time":1782701600884,"data":{"turn":1,"step":1,"chunk":{"type":"block-end","index":0,"block":{"type":"reasoning","text":"The user wants me to use todo_write to record a plan with exactly three todos, then reply with DONE."}}}} +{"type":"assistant/chunk","seq":83,"time":1782701600884,"data":{"turn":1,"step":1,"chunk":{"type":"block-end","index":1,"block":{"type":"tool-call","id":"call_00_OK2ZF1DYrsKHQQtxfQlJ0810","name":"todo_write","arguments":"{\"todos\": [{\"content\": \"read the code\", \"status\": \"in_progress\"}, {\"content\": \"write the fix\", \"status\": \"pending\"}, {\"content\": \"run the tests\", \"status\": \"pending\"}]}"}}}} +{"type":"assistant/chunk","seq":84,"time":1782701600884,"data":{"turn":1,"step":1,"chunk":{"type":"usage","usage":{"inputTokens":1706,"outputTokens":113,"cacheReadTokens":0,"reasoningTokens":23}}}} +{"type":"assistant/chunk","seq":85,"time":1782701600885,"data":{"turn":1,"step":1,"chunk":{"type":"finish","reason":{"kind":"tool-calls"}}}} +{"type":"assistant/message","seq":86,"time":1782701600886,"data":{"turn":1,"step":1,"content":[{"type":"reasoning","text":"The user wants me to use todo_write to record a plan with exactly three todos, then reply with DONE."},{"type":"tool-call","id":"call_00_OK2ZF1DYrsKHQQtxfQlJ0810","name":"todo_write","arguments":"{\"todos\": [{\"content\": \"read the code\", \"status\": \"in_progress\"}, {\"content\": \"write the fix\", \"status\": \"pending\"}, {\"content\": \"run the tests\", \"status\": \"pending\"}]}"}],"usage":{"inputTokens":1706,"outputTokens":113,"cacheReadTokens":0,"reasoningTokens":23}},"sourceEventSeqs":[3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,25,26,27,28,29,30,31,32,33,34,35,36,37,38,39,40,41,42,43,44,45,46,47,48,49,50,51,52,53,54,55,56,57,58,59,60,61,62,63,64,65,66,67,68,69,70,71,72,73,74,75,76,77,78,79,80,81,82,83,84,85],"surfaceOp":"append"} +{"type":"tool/call","seq":87,"time":1782701600887,"data":{"turn":1,"step":1,"callId":"call_00_OK2ZF1DYrsKHQQtxfQlJ0810","name":"todo_write","arguments":"{\"todos\": [{\"content\": \"read the code\", \"status\": \"in_progress\"}, {\"content\": \"write the fix\", \"status\": \"pending\"}, {\"content\": \"run the tests\", \"status\": \"pending\"}]}"}} +{"type":"todo/write","seq":88,"time":1782701600887,"data":{"todos":[{"content":"read the code","status":"in_progress"},{"content":"write the fix","status":"pending"},{"content":"run the tests","status":"pending"}]}} +{"type":"tool/result","seq":89,"time":1782701600887,"data":{"turn":1,"step":1,"callId":"call_00_OK2ZF1DYrsKHQQtxfQlJ0810","content":[{"type":"text","text":"Updated todo list: 2 pending, 1 in progress, 0 completed."}],"isError":false},"sourceEventSeqs":[87],"surfaceOp":"append"} +{"type":"step/end","seq":90,"time":1782701600888,"data":{"turn":1,"step":1}} +{"type":"step/start","seq":91,"time":1782701600888,"data":{"turn":1,"step":2}} +{"type":"assistant/chunk","seq":92,"time":1782701601276,"data":{"turn":1,"step":2,"chunk":{"type":"block-start","index":0,"blockType":"reasoning"}}} +{"type":"assistant/chunk","seq":93,"time":1782701601276,"data":{"turn":1,"step":2,"chunk":{"type":"reasoning-delta","index":0,"text":"The"}}} +{"type":"assistant/chunk","seq":94,"time":1782701601382,"data":{"turn":1,"step":2,"chunk":{"type":"reasoning-delta","index":0,"text":" todo"}}} +{"type":"assistant/chunk","seq":95,"time":1782701601410,"data":{"turn":1,"step":2,"chunk":{"type":"reasoning-delta","index":0,"text":" list"}}} +{"type":"assistant/chunk","seq":96,"time":1782701601437,"data":{"turn":1,"step":2,"chunk":{"type":"reasoning-delta","index":0,"text":" was"}}} +{"type":"assistant/chunk","seq":97,"time":1782701601438,"data":{"turn":1,"step":2,"chunk":{"type":"reasoning-delta","index":0,"text":" set"}}} +{"type":"assistant/chunk","seq":98,"time":1782701601438,"data":{"turn":1,"step":2,"chunk":{"type":"reasoning-delta","index":0,"text":" successfully"}}} +{"type":"assistant/chunk","seq":99,"time":1782701601466,"data":{"turn":1,"step":2,"chunk":{"type":"reasoning-delta","index":0,"text":"."}}} +{"type":"assistant/chunk","seq":100,"time":1782701601467,"data":{"turn":1,"step":2,"chunk":{"type":"reasoning-delta","index":0,"text":" Now"}}} +{"type":"assistant/chunk","seq":101,"time":1782701601467,"data":{"turn":1,"step":2,"chunk":{"type":"reasoning-delta","index":0,"text":" I"}}} +{"type":"assistant/chunk","seq":102,"time":1782701601467,"data":{"turn":1,"step":2,"chunk":{"type":"reasoning-delta","index":0,"text":" just"}}} +{"type":"assistant/chunk","seq":103,"time":1782701601494,"data":{"turn":1,"step":2,"chunk":{"type":"reasoning-delta","index":0,"text":" need"}}} +{"type":"assistant/chunk","seq":104,"time":1782701601494,"data":{"turn":1,"step":2,"chunk":{"type":"reasoning-delta","index":0,"text":" to"}}} +{"type":"assistant/chunk","seq":105,"time":1782701601495,"data":{"turn":1,"step":2,"chunk":{"type":"reasoning-delta","index":0,"text":" reply"}}} +{"type":"assistant/chunk","seq":106,"time":1782701601495,"data":{"turn":1,"step":2,"chunk":{"type":"reasoning-delta","index":0,"text":" with"}}} +{"type":"assistant/chunk","seq":107,"time":1782701601520,"data":{"turn":1,"step":2,"chunk":{"type":"reasoning-delta","index":0,"text":" the"}}} +{"type":"assistant/chunk","seq":108,"time":1782701601521,"data":{"turn":1,"step":2,"chunk":{"type":"reasoning-delta","index":0,"text":" single"}}} +{"type":"assistant/chunk","seq":109,"time":1782701601521,"data":{"turn":1,"step":2,"chunk":{"type":"reasoning-delta","index":0,"text":" word"}}} +{"type":"assistant/chunk","seq":110,"time":1782701601521,"data":{"turn":1,"step":2,"chunk":{"type":"reasoning-delta","index":0,"text":" D"}}} +{"type":"assistant/chunk","seq":111,"time":1782701601550,"data":{"turn":1,"step":2,"chunk":{"type":"reasoning-delta","index":0,"text":"ONE"}}} +{"type":"assistant/chunk","seq":112,"time":1782701601550,"data":{"turn":1,"step":2,"chunk":{"type":"reasoning-delta","index":0,"text":"."}}} +{"type":"assistant/chunk","seq":113,"time":1782701601550,"data":{"turn":1,"step":2,"chunk":{"type":"block-start","index":1,"blockType":"text"}}} +{"type":"assistant/chunk","seq":114,"time":1782701601551,"data":{"turn":1,"step":2,"chunk":{"type":"text-delta","index":1,"text":"D"}}} +{"type":"assistant/chunk","seq":115,"time":1782701601551,"data":{"turn":1,"step":2,"chunk":{"type":"text-delta","index":1,"text":"ONE"}}} +{"type":"assistant/chunk","seq":116,"time":1782701601551,"data":{"turn":1,"step":2,"chunk":{"type":"block-end","index":0,"block":{"type":"reasoning","text":"The todo list was set successfully. Now I just need to reply with the single word DONE."}}}} +{"type":"assistant/chunk","seq":117,"time":1782701601551,"data":{"turn":1,"step":2,"chunk":{"type":"block-end","index":1,"block":{"type":"text","text":"DONE"}}}} +{"type":"assistant/chunk","seq":118,"time":1782701601551,"data":{"turn":1,"step":2,"chunk":{"type":"usage","usage":{"inputTokens":174,"outputTokens":23,"cacheReadTokens":1664,"reasoningTokens":20}}}} +{"type":"assistant/chunk","seq":119,"time":1782701601551,"data":{"turn":1,"step":2,"chunk":{"type":"finish","reason":{"kind":"stop"}}}} +{"type":"assistant/message","seq":120,"time":1782701601551,"data":{"turn":1,"step":2,"content":[{"type":"reasoning","text":"The todo list was set successfully. Now I just need to reply with the single word DONE."},{"type":"text","text":"DONE"}],"usage":{"inputTokens":174,"outputTokens":23,"cacheReadTokens":1664,"reasoningTokens":20}},"sourceEventSeqs":[92,93,94,95,96,97,98,99,100,101,102,103,104,105,106,107,108,109,110,111,112,113,114,115,116,117,118,119],"surfaceOp":"append"} +{"type":"step/end","seq":121,"time":1782701601552,"data":{"turn":1,"step":2}} +{"type":"turn/end","seq":122,"time":1782701601552,"data":{"turn":1,"reason":{"kind":"completed"}}} diff --git a/examples/acp-agent/tests/snapshots/todo-plan/stdout.golden.jsonl b/examples/acp-agent/tests/snapshots/todo-plan/stdout.golden.jsonl new file mode 100644 index 0000000000..052b86fca0 --- /dev/null +++ b/examples/acp-agent/tests/snapshots/todo-plan/stdout.golden.jsonl @@ -0,0 +1,51 @@ +{"jsonrpc":"2.0","id":1,"result":{"protocolVersion":1,"agentInfo":{"name":"deepseek-harness-acp","version":"0.0.1"},"agentCapabilities":{"loadSession":true,"promptCapabilities":{"image":false,"audio":false,"embeddedContext":false}},"authMethods":[]}} +{"jsonrpc":"2.0","id":2,"result":{"sessionId":"{{sessionId}}"}} +{"jsonrpc":"2.0","method":"session/update","params":{"sessionId":"{{sessionId}}","update":{"sessionUpdate":"agent_thought_chunk","content":{"type":"text","text":"The"}}}} +{"jsonrpc":"2.0","method":"session/update","params":{"sessionId":"{{sessionId}}","update":{"sessionUpdate":"agent_thought_chunk","content":{"type":"text","text":" user"}}}} +{"jsonrpc":"2.0","method":"session/update","params":{"sessionId":"{{sessionId}}","update":{"sessionUpdate":"agent_thought_chunk","content":{"type":"text","text":" wants"}}}} +{"jsonrpc":"2.0","method":"session/update","params":{"sessionId":"{{sessionId}}","update":{"sessionUpdate":"agent_thought_chunk","content":{"type":"text","text":" me"}}}} +{"jsonrpc":"2.0","method":"session/update","params":{"sessionId":"{{sessionId}}","update":{"sessionUpdate":"agent_thought_chunk","content":{"type":"text","text":" to"}}}} +{"jsonrpc":"2.0","method":"session/update","params":{"sessionId":"{{sessionId}}","update":{"sessionUpdate":"agent_thought_chunk","content":{"type":"text","text":" use"}}}} +{"jsonrpc":"2.0","method":"session/update","params":{"sessionId":"{{sessionId}}","update":{"sessionUpdate":"agent_thought_chunk","content":{"type":"text","text":" todo"}}}} +{"jsonrpc":"2.0","method":"session/update","params":{"sessionId":"{{sessionId}}","update":{"sessionUpdate":"agent_thought_chunk","content":{"type":"text","text":"_write"}}}} +{"jsonrpc":"2.0","method":"session/update","params":{"sessionId":"{{sessionId}}","update":{"sessionUpdate":"agent_thought_chunk","content":{"type":"text","text":" to"}}}} +{"jsonrpc":"2.0","method":"session/update","params":{"sessionId":"{{sessionId}}","update":{"sessionUpdate":"agent_thought_chunk","content":{"type":"text","text":" record"}}}} +{"jsonrpc":"2.0","method":"session/update","params":{"sessionId":"{{sessionId}}","update":{"sessionUpdate":"agent_thought_chunk","content":{"type":"text","text":" a"}}}} +{"jsonrpc":"2.0","method":"session/update","params":{"sessionId":"{{sessionId}}","update":{"sessionUpdate":"agent_thought_chunk","content":{"type":"text","text":" plan"}}}} +{"jsonrpc":"2.0","method":"session/update","params":{"sessionId":"{{sessionId}}","update":{"sessionUpdate":"agent_thought_chunk","content":{"type":"text","text":" with"}}}} +{"jsonrpc":"2.0","method":"session/update","params":{"sessionId":"{{sessionId}}","update":{"sessionUpdate":"agent_thought_chunk","content":{"type":"text","text":" exactly"}}}} +{"jsonrpc":"2.0","method":"session/update","params":{"sessionId":"{{sessionId}}","update":{"sessionUpdate":"agent_thought_chunk","content":{"type":"text","text":" three"}}}} +{"jsonrpc":"2.0","method":"session/update","params":{"sessionId":"{{sessionId}}","update":{"sessionUpdate":"agent_thought_chunk","content":{"type":"text","text":" todos"}}}} +{"jsonrpc":"2.0","method":"session/update","params":{"sessionId":"{{sessionId}}","update":{"sessionUpdate":"agent_thought_chunk","content":{"type":"text","text":","}}}} +{"jsonrpc":"2.0","method":"session/update","params":{"sessionId":"{{sessionId}}","update":{"sessionUpdate":"agent_thought_chunk","content":{"type":"text","text":" then"}}}} +{"jsonrpc":"2.0","method":"session/update","params":{"sessionId":"{{sessionId}}","update":{"sessionUpdate":"agent_thought_chunk","content":{"type":"text","text":" reply"}}}} +{"jsonrpc":"2.0","method":"session/update","params":{"sessionId":"{{sessionId}}","update":{"sessionUpdate":"agent_thought_chunk","content":{"type":"text","text":" with"}}}} +{"jsonrpc":"2.0","method":"session/update","params":{"sessionId":"{{sessionId}}","update":{"sessionUpdate":"agent_thought_chunk","content":{"type":"text","text":" D"}}}} +{"jsonrpc":"2.0","method":"session/update","params":{"sessionId":"{{sessionId}}","update":{"sessionUpdate":"agent_thought_chunk","content":{"type":"text","text":"ONE"}}}} +{"jsonrpc":"2.0","method":"session/update","params":{"sessionId":"{{sessionId}}","update":{"sessionUpdate":"agent_thought_chunk","content":{"type":"text","text":"."}}}} +{"jsonrpc":"2.0","method":"session/update","params":{"sessionId":"{{sessionId}}","update":{"sessionUpdate":"tool_call","toolCallId":"call_00_OK2ZF1DYrsKHQQtxfQlJ0810","title":"Update todo list","kind":"other","status":"in_progress","rawInput":[{"content":"read the code","status":"in_progress"},{"content":"write the fix","status":"pending"},{"content":"run the tests","status":"pending"}]}}} +{"jsonrpc":"2.0","method":"session/update","params":{"sessionId":"{{sessionId}}","update":{"sessionUpdate":"plan","entries":[{"content":"read the code","priority":"medium","status":"in_progress"},{"content":"write the fix","priority":"medium","status":"pending"},{"content":"run the tests","priority":"medium","status":"pending"}]}}} +{"jsonrpc":"2.0","method":"session/update","params":{"sessionId":"{{sessionId}}","update":{"sessionUpdate":"tool_call_update","toolCallId":"call_00_OK2ZF1DYrsKHQQtxfQlJ0810","status":"completed","content":[{"type":"content","content":{"type":"text","text":"Updated todo list: 2 pending, 1 in progress, 0 completed."}}]}}} +{"jsonrpc":"2.0","method":"session/update","params":{"sessionId":"{{sessionId}}","update":{"sessionUpdate":"agent_thought_chunk","content":{"type":"text","text":"The"}}}} +{"jsonrpc":"2.0","method":"session/update","params":{"sessionId":"{{sessionId}}","update":{"sessionUpdate":"agent_thought_chunk","content":{"type":"text","text":" todo"}}}} +{"jsonrpc":"2.0","method":"session/update","params":{"sessionId":"{{sessionId}}","update":{"sessionUpdate":"agent_thought_chunk","content":{"type":"text","text":" list"}}}} +{"jsonrpc":"2.0","method":"session/update","params":{"sessionId":"{{sessionId}}","update":{"sessionUpdate":"agent_thought_chunk","content":{"type":"text","text":" was"}}}} +{"jsonrpc":"2.0","method":"session/update","params":{"sessionId":"{{sessionId}}","update":{"sessionUpdate":"agent_thought_chunk","content":{"type":"text","text":" set"}}}} +{"jsonrpc":"2.0","method":"session/update","params":{"sessionId":"{{sessionId}}","update":{"sessionUpdate":"agent_thought_chunk","content":{"type":"text","text":" successfully"}}}} +{"jsonrpc":"2.0","method":"session/update","params":{"sessionId":"{{sessionId}}","update":{"sessionUpdate":"agent_thought_chunk","content":{"type":"text","text":"."}}}} +{"jsonrpc":"2.0","method":"session/update","params":{"sessionId":"{{sessionId}}","update":{"sessionUpdate":"agent_thought_chunk","content":{"type":"text","text":" Now"}}}} +{"jsonrpc":"2.0","method":"session/update","params":{"sessionId":"{{sessionId}}","update":{"sessionUpdate":"agent_thought_chunk","content":{"type":"text","text":" I"}}}} +{"jsonrpc":"2.0","method":"session/update","params":{"sessionId":"{{sessionId}}","update":{"sessionUpdate":"agent_thought_chunk","content":{"type":"text","text":" just"}}}} +{"jsonrpc":"2.0","method":"session/update","params":{"sessionId":"{{sessionId}}","update":{"sessionUpdate":"agent_thought_chunk","content":{"type":"text","text":" need"}}}} +{"jsonrpc":"2.0","method":"session/update","params":{"sessionId":"{{sessionId}}","update":{"sessionUpdate":"agent_thought_chunk","content":{"type":"text","text":" to"}}}} +{"jsonrpc":"2.0","method":"session/update","params":{"sessionId":"{{sessionId}}","update":{"sessionUpdate":"agent_thought_chunk","content":{"type":"text","text":" reply"}}}} +{"jsonrpc":"2.0","method":"session/update","params":{"sessionId":"{{sessionId}}","update":{"sessionUpdate":"agent_thought_chunk","content":{"type":"text","text":" with"}}}} +{"jsonrpc":"2.0","method":"session/update","params":{"sessionId":"{{sessionId}}","update":{"sessionUpdate":"agent_thought_chunk","content":{"type":"text","text":" the"}}}} +{"jsonrpc":"2.0","method":"session/update","params":{"sessionId":"{{sessionId}}","update":{"sessionUpdate":"agent_thought_chunk","content":{"type":"text","text":" single"}}}} +{"jsonrpc":"2.0","method":"session/update","params":{"sessionId":"{{sessionId}}","update":{"sessionUpdate":"agent_thought_chunk","content":{"type":"text","text":" word"}}}} +{"jsonrpc":"2.0","method":"session/update","params":{"sessionId":"{{sessionId}}","update":{"sessionUpdate":"agent_thought_chunk","content":{"type":"text","text":" D"}}}} +{"jsonrpc":"2.0","method":"session/update","params":{"sessionId":"{{sessionId}}","update":{"sessionUpdate":"agent_thought_chunk","content":{"type":"text","text":"ONE"}}}} +{"jsonrpc":"2.0","method":"session/update","params":{"sessionId":"{{sessionId}}","update":{"sessionUpdate":"agent_thought_chunk","content":{"type":"text","text":"."}}}} +{"jsonrpc":"2.0","method":"session/update","params":{"sessionId":"{{sessionId}}","update":{"sessionUpdate":"agent_message_chunk","content":{"type":"text","text":"D"}}}} +{"jsonrpc":"2.0","method":"session/update","params":{"sessionId":"{{sessionId}}","update":{"sessionUpdate":"agent_message_chunk","content":{"type":"text","text":"ONE"}}}} +{"jsonrpc":"2.0","id":3,"result":{"stopReason":"end_turn"}} diff --git a/examples/coding-agent/README.md b/examples/coding-agent/README.md index 7585129382..6a156dcc48 100644 --- a/examples/coding-agent/README.md +++ b/examples/coding-agent/README.md @@ -1,7 +1,6 @@ # coding-agent -The first REAL agent wiring: DeepSeek V4 + the bash tool suite + stdio chat -+ JSONL persistence, loaded from `cordis.yml`. Where echo-agent proves the skeleton with mocks, this example is a usable coding assistant. +The real stdio coding-agent wiring: DeepSeek V4 + the bash tool suite + subagent delegation + `todo_write` + stdio chat + JSONL persistence, loaded from `cordis.yml`. Where echo-agent proves the skeleton with mocks, this example is a usable coding assistant. ## Run it @@ -12,7 +11,7 @@ The first REAL agent wiring: DeepSeek V4 + the bash tool suite + stdio chat pnpm run demo:coding ``` -Type a coding task. The agent's only tools are `bash` (+ `bash_output` / `bash_kill` for background tasks): file reads, writes, searches, and test runs all happen through shell commands, each in a fresh `bash -c` (the system prompt tells the model to pass `workdir` instead of `cd`). Reasoning streams dimmed; tool calls/results render inline. +Type a coding task. The agent works through `bash` (+ `bash_output` / `bash_kill` for background tasks): file reads, writes, searches, and test runs all happen through shell commands, each in a fresh `bash -c` (the system prompt tells the model to pass `workdir` instead of `cd`). It can also delegate with `subagent`/`subagent_fork` and track multi-step work with `todo_write` (a whole-list task tracker rendered as a checklist). Reasoning streams dimmed; tool calls/results render inline. ``` > fix the failing test in /path/to/project @@ -34,7 +33,7 @@ The id is wired through `cordis.yml` (`resumeSessionId: !!js process.env.RESUME_ ## What each leaf entry demonstrates -This example is a thin leaf `cordis.yml`: it picks the swappable backends and loads one app package. The spine (sessions, system-prompt, tools, agents, invariants, `agent-loop`) and the front-door cluster (console logger, JSONL persistence, readline UI, the pre-created `main` agent) all live inside the [`@deepseek-ai/dsh-stdio-agent`](../../packages/ui/stdio-agent) app and the [`@deepseek-ai/dsh-agent-core`](../../packages/core/agent-core) bundle it loads — so the leaf has only four entries: +This example is a thin leaf `cordis.yml`: it picks the swappable backends, loads one app package, and adds product tools that are intentionally outside the shared spine. The spine (sessions, system-prompt, tools, agents, invariants, `agent-loop`) and the front-door cluster (console logger, JSONL persistence, readline UI, the pre-created `main` agent) live inside the [`@deepseek-ai/dsh-stdio-agent`](../../packages/ui/stdio-agent) app and the [`@deepseek-ai/dsh-agent-core`](../../packages/core/agent-core) bundle it loads; the leaf wires the backends and model-facing optional tools: | Entry | Demonstrates | |---|---| @@ -42,11 +41,16 @@ This example is a thin leaf `cordis.yml`: it picks the swappable backends and lo | `llm-deepseek` | real `LlmAdapter` via config (`!!js process.env.…` secrets); swap one line to `@deepseek-ai/dsh-llm-pi-ai` for the library-backed twin | | `bash` (`dsh-bash-local`) | the executor implementation — the swappable half of the bash seam. The model-facing `bash`/`bash_output`/`bash_kill` tool schemas (`tool-bash`) come from `agent-core`, so only the executor is a leaf choice | | `stdio-agent` (`@deepseek-ai/dsh-stdio-agent`) | the app bundle: the agent-core spine + console logger + JSONL persistence + readline UI + a pre-created `main` agent. Its config carries the model, system prompt, `persistenceRoot` (`./.sessions`), and `resumeSessionId` — so persistence and the agent are configured here, not wired as separate leaf plugins | +| `subagent`, `subagent-spawn`, `subagent-fork` | the subagent provider registry plus the two in-process backends: a fresh child and a child seeded with the parent's completed-turn prefix | +| `tool-subagent`, `tool-subagent-fork` | two model-facing `dsh-tool-subagent` loads, each bound to a different provider and exposed under a distinct tool name (`subagent`, `subagent_fork`) | +| `tool-todo` | the model-facing `todo_write` tool; writes the whole task list to the session log and renders as a checklist in stdio | ## End-to-end tests (`pnpm run test:e2e`, key-gated) - `tests/full-loop.e2e.ts` — the canary: real model runs `echo e2e-ok` through the real bash tool; asserts `tool/call`/`tool/result` session events and the final answer. - `tests/coding-task.e2e.ts` — the swebench-style smoke: a temp dir holds `add.js` (with `a - b` where `a + b` belongs) and a failing `add.test.js`; the agent must fix the bug and verify. The test re-runs `node add.test.js` ITSELF and inspects the files — agent claims are not trusted. - `tests/resume.e2e.ts` — durable continuity across processes: run 1 tells the real model a secret code and persists the turn to a temp JSONL root, then the whole context is disposed; run 2 is a fresh context over the same root that RESUMES the session id and asks the model to recall the code. The recall can only come from the rehydrated log. +- `tests/compaction.e2e.ts` — the compaction smoke: a real multi-step bash task runs with a deliberately tiny context window so the auto-compaction listener fires MID-SESSION. Verifies the WORLD — a `compact/start…end` pair landed in the real log, the surface shrank (a replace node shadowed older nodes), and the agent still produced a correct final answer after compaction. +- `tests/todo-write.e2e.ts` — a real model drives the real `todo_write` tool and the test verifies the resulting `todo/write` session event. -Both self-skip without `DEEPSEEK_API_KEY`. +These self-skip without `DEEPSEEK_API_KEY`. The keyless boot smoke is `tests/keyless-smoke.e2e.ts` (boots the full real tree with a dummy key and no prompt, so no model call), which runs in the default e2e gate. diff --git a/examples/coding-agent/cordis.yml b/examples/coding-agent/cordis.yml index 355f8f102b..bc3c3f2ff7 100644 --- a/examples/coding-agent/cordis.yml +++ b/examples/coding-agent/cordis.yml @@ -28,7 +28,9 @@ - deepseek-v4-pro - deepseek-v4-flash -# Local bash executor (the model's only tool, via agent-core's tool-bash schema). +# Local bash executor for agent-core's tool-bash schema. +# FIXME(config-comments): keep this executor note from implying bash is the +# whole tool set; subagent and todo_write are loaded below. - id: bash name: '@deepseek-ai/dsh-bash-local' config: @@ -44,7 +46,7 @@ # under ./.sessions); unset starts a fresh session each run. resumeSessionId: !!js process.env.RESUME_SESSION_ID persistenceRoot: './.sessions' - welcome: 'coding-agent ready. Give it a coding task (its tools are bash and subagent).' + welcome: 'coding-agent ready. Give it a coding task (its tools are bash, subagent, and todo_write).' systemPrompt: | You are coding-agent, a CLI coding assistant. @@ -65,6 +67,26 @@ failures before moving on. Verify your work by running the code or tests. Keep answers brief and factual. + For multi-step work, use the todo_write tool to track a task list: + send the WHOLE list each call (it replaces the previous one), keep at + most one task in_progress (exactly one while work remains), and mark a + task completed as soon as it is done. Skip it for trivial single-step + tasks. + +# Automatic context compaction: when the derived history approaches the model's +# context window, summarize an older range into a checkpoint so a long-running +# or tool-heavy session keeps fitting. A leaf entry (needs ctx.llm + the +# agent-loop's `agent/pre-step` seam from the app above). +- id: compact-basic + name: '@deepseek-ai/dsh-compact-basic' + config: + contextWindow: 128000 + thresholdRatio: 0.8 + retainTokens: 20480 + summarizationModel: '' + maxTokens: 8192 + compactionRetries: 1 + # The subagent seam + BOTH in-process backends + two model-facing tools, as leaf # entries after the app (which provides ctx.agents/ctx.tools). spawn (a fresh # child) and fork (a child seeded with the parent's completed-turn prefix) are @@ -96,3 +118,8 @@ config: provider: fork toolName: subagent_fork + +# The model-facing todo_write tool: whole-list task tracking written to the +# session log (todo/write), rendered as a stdio checklist / ACP plan. +- id: tool-todo + name: '@deepseek-ai/dsh-tool-todo' diff --git a/examples/coding-agent/tests/compaction.e2e.ts b/examples/coding-agent/tests/compaction.e2e.ts new file mode 100644 index 0000000000..2b8f278be3 --- /dev/null +++ b/examples/coding-agent/tests/compaction.e2e.ts @@ -0,0 +1,108 @@ +import { mkdtemp, rm, writeFile } from 'node:fs/promises' +import { tmpdir } from 'node:os' +import { join } from 'node:path' +import { afterEach, describe, expect, it } from 'vitest' +import type { Context } from 'cordis' +import { AgentId } from '@deepseek-ai/dsh-agent' +import { codingHarness, finalText, SYSTEM_PROMPT, waitForIdle } from './harness.ts' + +/** + * The compaction smoke test: a real model runs a multi-step bash task with a + * deliberately tiny context window, so the auto-compaction listener fires + * MID-SESSION and summarizes the older history into a checkpoint. This is the + * first end-to-end exercise of the compaction seam (it is wired nowhere else), + * and the runaway-survival regression net — it proves a session that grows past + * the window keeps running rather than overflowing. Key-gated. + * + * Verifies the WORLD, not the agent's self-report: a compact/start…end pair + * landed in the real session log, the surface actually shrank (a replace node + * exists and shadowed older nodes), and the agent still produced a final answer + * after compaction (so the summarized history did not break the conversation). + * + * FIXME(compaction-snapshot): this key-gated e2e is the ONLY coverage of runaway + * compaction — there is no keyless full-transcript snapshot of it. dsh-llm-replay + * reconstructs one model call per (turn, step) from `assistant/chunk` events, but + * `summarize()` assembles its stream into a local BlockAssembler and appends no + * `assistant/chunk`, so the interleaved summarization call is unreplayable. A + * snapshot needs replay-harness work to serve that call; deferred as a follow-up. + */ + +let workdir: string | undefined +let ctx: Context | undefined + +afterEach(async () => { + await ctx?.fiber.dispose() + ctx = undefined + if (workdir !== undefined) await rm(workdir, { recursive: true, force: true }) + workdir = undefined +}) + +describe.skipIf(!process.env.DEEPSEEK_API_KEY)('compaction: a long session compacts mid-flight and keeps running', () => { + it('summarizes older history into a checkpoint without breaking the task', async () => { + workdir = await mkdtemp(join(tmpdir(), 'dsh-compaction-')) + // A handful of files for the model to read, so multiple bash steps + // accumulate surface nodes (tool calls + results) and grow the history past + // the (deliberately tiny) window. + for (let i = 1; i <= 6; i++) { + await writeFile(join(workdir, `file${i}.txt`), `This is file number ${i}. `.repeat(40)) + } + + // Tiny window so a couple of steps crosses the threshold. The generation + // cap is deliberately larger than the final checkpoint because + // reasoning-capable APIs count reasoning tokens against the provider output + // budget even though those blocks are stripped before the checkpoint is + // stored. + ctx = await codingHarness(workdir, { + compact: { + contextWindow: 2400, + thresholdRatio: 0.5, + retainTokens: 500, + summarizationModel: '', + maxTokens: 2048, + compactionRetries: 1, + }, + persistenceRoot: './.sessions', + }) + const agent = ctx.agentLoop.create(AgentId('e2e-compaction'), { + model: 'deepseek-v4-flash', + systemPrompt: SYSTEM_PROMPT, + }) + + agent.send([{ + type: 'text', + text: 'Read file1.txt, file2.txt, file3.txt, file4.txt, file5.txt, and file6.txt one at a ' + + 'time using cat (a separate bash command for each). After reading all six, tell me how ' + + 'many files you read and the number mentioned in file1.txt.', + }]) + await waitForIdle(ctx, agent) + + const events = [...agent.session.events] + + // A compaction ran: the start…end bracket landed in the real log. + const starts = events.filter(e => e.type === 'compact/start') + const ends = events.filter(e => e.type === 'compact/end') + expect(starts.length).toBeGreaterThan(0) + expect(ends.length).toBe(starts.length) // every start was released + + // It succeeded at least once: a compact/summary provenance event and a + // replace-op user/message (the surface mutation) both landed. + const summaries = events.filter(e => e.type === 'compact/summary') + expect(summaries.length).toBeGreaterThan(0) + const replaceNode = events.find((e) => { + const se = e as unknown as { type: string; surfaceOp?: unknown } + return se.type === 'user/message' && typeof se.surfaceOp === 'object' && se.surfaceOp !== null + }) + expect(replaceNode).toBeDefined() + + // The summary shadowed real older nodes (the surface shrank vs. the raw + // message-producing event count). + const summaryData = summaries[0]!.data as { shadowedSeqs: number[] } + expect(summaryData.shadowedSeqs.length).toBeGreaterThan(0) + + // The conversation survived compaction: the agent produced a final answer + // that reflects the work (it read six files). + const answer = finalText(events).toLowerCase() + expect(answer.length).toBeGreaterThan(0) + expect(answer).toMatch(/\b(6|six)\b/) + }, 240_000) +}) diff --git a/examples/coding-agent/tests/harness.ts b/examples/coding-agent/tests/harness.ts index fbe9b10db5..7ce24913cf 100644 --- a/examples/coding-agent/tests/harness.ts +++ b/examples/coding-agent/tests/harness.ts @@ -8,20 +8,42 @@ import AgentRegistry from '@deepseek-ai/dsh-agent' import AgentLoop, { ReactLoopAgent } from '@deepseek-ai/dsh-agent-loop' import { LocalBashExecutor } from '@deepseek-ai/dsh-bash-local' import * as ToolBash from '@deepseek-ai/dsh-tool-bash' +import * as ToolTodo from '@deepseek-ai/dsh-tool-todo' import * as LlmDeepSeek from '@deepseek-ai/dsh-llm-deepseek' import SessionPersistenceJsonl from '@deepseek-ai/dsh-session-persistence-jsonl' +import { BasicCompactService } from '@deepseek-ai/dsh-compact-basic' +import type { BasicCompactConfig } from '@deepseek-ai/dsh-compact-basic' /** * Shared harness for the coding-agent e2e suites: the full plugin stack - * with the real DeepSeek adapter and the real bash tool. Lives outside the - * *.e2e.ts pattern so importing it never re-registers another file's tests. + * with the real DeepSeek adapter and the real bash + todo_write tools. Lives + * outside the *.e2e.ts pattern so importing it never re-registers another + * file's tests. */ -export const SYSTEM_PROMPT = 'You are a coding agent. Your only tool is bash; ' - + 'do file operations with cat/grep/heredocs, check [exit code: N] markers, ' +export const SYSTEM_PROMPT = 'You are a coding agent. Use bash for file operations ' + + 'with cat/grep/heredocs; check [exit code: N] markers, ' + 'and report results briefly.' -export async function codingHarness(workdir: string, persistenceRoot?: string): Promise { +/** System prompt for the todo_write e2e: nudges the model to plan with the tool. */ +export const TODO_SYSTEM_PROMPT = 'You are a coding agent. For multi-step work, ' + + 'use the todo_write tool to track a task list: send the WHOLE list each call, ' + + 'keep at most one task in_progress (exactly one while work remains), and mark ' + + 'a task completed as soon as it is done.' + +/** Options for {@link codingHarness}. */ +export interface CodingHarnessOptions { + /** Durable JSONL persistence root (the resume suite needs it; others stay file-free). */ + persistenceRoot?: string + /** + * Load {@link BasicCompactService} with this config so the compaction e2e can + * trigger compaction at a small, controlled history size. Omitted ⇒ no + * compaction plugin (the default suites run without it). + */ + compact?: BasicCompactConfig +} + +export async function codingHarness(workdir: string, options: CodingHarnessOptions = {}): Promise { const ctx = new Context() await ctx.plugin(LlmService) await ctx.plugin(SessionStore) @@ -32,10 +54,14 @@ export async function codingHarness(workdir: string, persistenceRoot?: string): await ctx.plugin(LlmDeepSeek, { models: ['deepseek-v4-flash'] }) await ctx.plugin(LocalBashExecutor, { cwd: workdir, timeoutMs: 30_000 }) await ctx.plugin(ToolBash) + await ctx.plugin(ToolTodo) + // Compaction is opt-in: only the compaction e2e loads it, with a lowered + // contextWindow/retainTokens so a short real session crosses the threshold. + if (options.compact !== undefined) await ctx.plugin(BasicCompactService, options.compact) // Durable JSONL persistence is opt-in: only the resume e2e needs it, and the // other suites stay file-free. Loaded last so a resume's deferred // `ctx.inject(['sessionPersistence'])` resolves once this is present. - if (persistenceRoot !== undefined) await ctx.plugin(SessionPersistenceJsonl, { root: persistenceRoot }) + if (options.persistenceRoot !== undefined) await ctx.plugin(SessionPersistenceJsonl, { root: options.persistenceRoot }) return ctx } diff --git a/examples/coding-agent/tests/resume.e2e.ts b/examples/coding-agent/tests/resume.e2e.ts index 450938fc6d..4be11ed3ea 100644 --- a/examples/coding-agent/tests/resume.e2e.ts +++ b/examples/coding-agent/tests/resume.e2e.ts @@ -38,7 +38,7 @@ describe.skipIf(!process.env.DEEPSEEK_API_KEY)('resume: continue a persisted ses // Run 1: a fresh agent on a KNOWN session id learns a secret, then we // dispose the whole context (simulating process exit) so only the JSONL // log on disk survives. - ctx = await codingHarness(process.cwd(), root) + ctx = await codingHarness(process.cwd(), { persistenceRoot: root }) const first = ctx.agents.create({ agentId: AgentId('resume-1'), sessionId: SESSION_ID, @@ -52,7 +52,7 @@ describe.skipIf(!process.env.DEEPSEEK_API_KEY)('resume: continue a persisted ses // Run 2: a brand-new context over the SAME root resumes the persisted // session. The loaded event log seeds the live session, so the model sees // run 1's exchange as conversation history. - ctx = await codingHarness(process.cwd(), root) + ctx = await codingHarness(process.cwd(), { persistenceRoot: root }) const resumed = (await ctx.agents.resume({ agentId: AgentId('resume-2'), resumeSessionId: SESSION_ID, diff --git a/examples/coding-agent/tests/todo-write.e2e.ts b/examples/coding-agent/tests/todo-write.e2e.ts new file mode 100644 index 0000000000..33cac531cf --- /dev/null +++ b/examples/coding-agent/tests/todo-write.e2e.ts @@ -0,0 +1,49 @@ +import { afterEach, describe, expect, it } from 'vitest' +import type { Context } from 'cordis' +import { AgentId } from '@deepseek-ai/dsh-agent' +import { codingHarness, TODO_SYSTEM_PROMPT, waitForIdle } from './harness.ts' + +/** + * A REAL model drives the REAL todo_write tool: verify the WORLD (the session + * log gains a todo/write event whose snapshot the model actually produced), not + * the agent's self-report. Key-gated (see vitest.e2e.config.ts). + */ + +let ctx: Context | undefined + +afterEach(async () => { + await ctx?.fiber.dispose() + ctx = undefined +}) + +describe.skipIf(!process.env.DEEPSEEK_API_KEY)('todo_write: real model records a plan', () => { + it('appends a todo/write event with the model-produced task list', async () => { + ctx = await codingHarness(process.cwd()) + const agent = ctx.agentLoop.create(AgentId('e2e-todo'), { + model: 'deepseek-v4-flash', + systemPrompt: TODO_SYSTEM_PROMPT, + }) + + agent.send([{ type: 'text', text: + 'Use the todo_write tool to record a plan of exactly two steps: first ' + + '"inspect the failing test" (in_progress), then "apply the fix" (pending). ' + + 'Send both in one todo_write call, then reply with the single word DONE.' }]) + await waitForIdle(ctx, agent) + + const events = [...agent.session.events] + + // The model actually called the tool. + const calls = events.filter(event => event.type === 'tool/call') + expect(calls.some(event => event.data.name === 'todo_write')).toBe(true) + + // And the tool wrote a todo/write event to the log — verify the WORLD. + const todoEvents = events.filter(event => event.type === 'todo/write') + expect(todoEvents.length).toBeGreaterThan(0) + + const todos = (todoEvents.at(-1)!).data.todos + expect(todos).toEqual([ + { content: 'inspect the failing test', status: 'in_progress' }, + { content: 'apply the fix', status: 'pending' }, + ]) + }, 120_000) +}) diff --git a/packages/README.md b/packages/README.md index 32a1288601..1ecdb8c9eb 100644 --- a/packages/README.md +++ b/packages/README.md @@ -12,8 +12,9 @@ Packages are grouped by modular role at `packages///`. The group dir | [`llm/`](llm/README.md) | LLM capability family: the abstract service + provider adapters | Product — stable surface | | [`bash/`](bash/README.md) | Bash capability family: the executor seam, a local impl, and the model-facing tool | Product — stable surface | | [`fs/`](fs/README.md) | Filesystem capability family: the abstract seam, a local impl, and the model-facing file tools | Product — stable surface | -| [`compact/`](compact/README.md) | Compaction capability family: the abstract seam (backend + tool deferred) | Product — stable surface | +| [`compact/`](compact/README.md) | Compaction capability family: the abstract seam + a basic backend (tool deferred) | Product — stable surface | | [`subagent/`](subagent/README.md) | Subagent capability family: the provider-registry seam and the model-facing delegation tool | Product — stable surface | +| [`todo/`](todo/README.md) | Todo/planning family: the model-facing `todo_write` tool (whole-list task tracking on the session log) | Product — stable surface | | [`session-persistence/`](session-persistence/README.md) | Persistence capability family: the seam + JSONL/SQLite backends | Product — stable surface | | [`ui/`](ui/README.md) | Editor/client integration surfaces (the ACP bridge) | Product — stable surface | | [`support/`](support/README.md) | Dev/test/example infrastructure (invariants, stdio UI, replay adapter) | Support — lower compatibility expectations | @@ -30,7 +31,8 @@ dsh-bash ← dsh-brand (abstract executor seam; b dsh-session ← dsh-llm, dsh-brand dsh-system-prompt ← dsh-llm dsh-agent ← dsh-llm, dsh-session, dsh-brand -dsh-compact ← dsh-session, dsh-llm (abstract compaction seam; backend + tool deferred) +dsh-compact ← dsh-session, dsh-llm (abstract compaction seam; tool deferred) +dsh-compact-basic ← dsh-compact, dsh-session, dsh-llm, dsh-agent (char/4 + token-budget retention backend) dsh-tools ← dsh-llm, dsh-system-prompt, dsh-agent dsh-bash-local ← dsh-bash (BashExecutor impl) dsh-tool-bash ← dsh-bash, dsh-tools (bash tool schemas) @@ -40,17 +42,19 @@ dsh-file-context ← dsh-fs (observed-state + freshnes dsh-tool-fs ← dsh-fs, dsh-tools (file tools + executor) dsh-llm-deepseek ← dsh-llm (DeepSeek adapter) dsh-llm-pi-ai ← dsh-llm (pi-ai-backed adapter) -dsh-agent-loop ← dsh-llm, dsh-session, dsh-system-prompt, dsh-tools, dsh-agent +dsh-agent-loop ← dsh-llm, dsh-session, dsh-session-persistence, dsh-system-prompt, dsh-tools, dsh-agent dsh-invariants ← dsh-llm, dsh-session, dsh-agent (dev-mode contract checks) -dsh-acp ← dsh-agent, dsh-llm, dsh-session, dsh-session-persistence (ACP JSON-RPC bridge) +dsh-acp ← dsh-agent, dsh-llm, dsh-session, dsh-session-persistence, dsh-tools (ACP JSON-RPC bridge) dsh-ui-stdio ← dsh-agent, dsh-llm, dsh-session (stdio readline UI plugin) dsh-llm-replay ← dsh-llm, dsh-session (record/replay adapter for keyless snapshot tests) dsh-subagent ← dsh-agent, dsh-llm, dsh-tools (abstract subagent provider-registry seam) -dsh-subagent-mock ← dsh-subagent (scripted provider for tests) -dsh-subagent-spawn ← dsh-subagent, dsh-agent, dsh-session, dsh-llm (in-process fresh child + shared run driver) -dsh-subagent-fork ← dsh-subagent-spawn, dsh-agent, dsh-session (in-process child seeded from parent log) +dsh-subagent-inprocess ← dsh-subagent, dsh-agent, dsh-session, dsh-llm (shared in-process run driver) +dsh-subagent-mock ← dsh-subagent, dsh-agent, dsh-llm (scripted provider for tests) +dsh-subagent-spawn ← dsh-subagent, dsh-subagent-inprocess (in-process fresh child backend) +dsh-subagent-fork ← dsh-subagent, dsh-subagent-inprocess, dsh-agent, dsh-session (in-process child seeded from parent log) dsh-subagent-acp ← dsh-subagent, dsh-agent, dsh-llm, @agentclientprotocol/sdk (out-of-process child over ACP) -dsh-tool-subagent ← dsh-subagent, dsh-tools, dsh-agent (model-facing delegation tool) +dsh-tool-subagent ← dsh-subagent, dsh-tools, dsh-agent, dsh-llm (model-facing delegation tool) +dsh-tool-todo ← dsh-tools, dsh-agent, dsh-session (model-facing todo_write tool; whole list on the session log) dsh-agent-core ← timer, dsh-llm, dsh-session, dsh-system-prompt, dsh-tools, dsh-agent, dsh-invariants, dsh-tool-bash, dsh-agent-loop (the providerless spine, as one bundle plugin) dsh-stdio-agent ← dsh-agent-core, dsh-ui-stdio, dsh-session-persistence-jsonl, dsh-agent, dsh-session (stdio chat APP + bin) dsh-acp-agent ← dsh-agent-core, dsh-acp, dsh-session-persistence-jsonl (ACP server APP + bin) @@ -77,6 +81,7 @@ The rule: **extension** plugins depend on interfaces, never on the concrete loop | `file-context/` | `fs` | Policy gate plugin: observed-state + read-before-edit + version-guarded write/edit via the `fs/*` event gate | (no service — `fs/*` listeners) | | `tool-fs/` | `fs` | Model-facing `read`/`write`/`edit` tools + executor (reads via `ctx.fs`, owns read windowing, dispatches `fs/*`) | (registers on `ctx.tools`) | | `compact/` | `compact` | Abstract compaction seam + `compact/*` events + `CompactionResult` | `ctx.compact` | +| `compact-basic/` | `compact` | A backend: char/4 estimation + token-budget retention + `llm.stream()` summarization | (registers `ctx.compact`) | | `llm-deepseek/` | `llm` | DeepSeek API adapter (hand-rolled fetch/SSE) | (registers on `ctx.llm`) | | `llm-pi-ai/` | `llm` | DeepSeek adapter via `@earendil-works/pi-ai` (design twin) | (registers on `ctx.llm`) | | `session-persistence/` | `session-persistence` | Persistence seam + write coordinator | `ctx.sessionPersistence` | @@ -89,11 +94,13 @@ The rule: **extension** plugins depend on interfaces, never on the concrete loop | `ui-stdio/` | `support` | Minimal stdio (readline) UI plugin: renders `agent/*` events, feeds stdin lines to the agent | (drives `ctx.agents`) | | `llm-replay/` | `support` | Record/replay adapter: short-circuits `llm/stream` with chunks from a recorded session JSONL (keyless snapshot tests) | (listens on `llm/stream`) | | `subagent/` | `subagent` | Abstract subagent seam: named-provider registry for delegating to child agents | `ctx.subagents` | -| `subagent-spawn/` | `subagent` | In-process backend: a fresh child agent (+ the shared in-process run driver) | (registers on `ctx.subagents`) | +| `subagent-inprocess/` | `subagent` | Shared in-process subagent run driver used by spawn/fork; pure library, registers nothing | (none) | +| `subagent-spawn/` | `subagent` | In-process backend: a fresh child agent | (registers on `ctx.subagents`) | | `subagent-fork/` | `subagent` | In-process backend: a child agent seeded with the parent's completed-turn prefix | (registers on `ctx.subagents`) | | `subagent-acp/` | `subagent` | Out-of-process backend: a child agent in a spawned subprocess, driven over the Agent Client Protocol | (registers on `ctx.subagents`) | | `subagent-mock/` | `support` | Scripted `SubagentProvider` for testing the seam through the real load path | (registers on `ctx.subagents`) | | `tool-subagent/` | `subagent` | Model-facing `subagent` delegation tool over `ctx.subagents` | (registers on `ctx.tools`) | +| `tool-todo/` | `todo` | Model-facing `todo_write` tool; writes the whole task list to the session log (`todo/write`) | (registers on `ctx.tools`) | | `brand/` | `util` | Type-only `Branded` nominal-typing primitive (no runtime code, no harness deps) | (none — type-only) | Each package has its own `README.md` with purpose, service API, events, extension points, and deliberate non-goals (TODOs). diff --git a/packages/compact/README.md b/packages/compact/README.md index 0d63b3b8cd..10eaf1617a 100644 --- a/packages/compact/README.md +++ b/packages/compact/README.md @@ -1,11 +1,11 @@ # compact/ — compaction capability family -A three-package capability seam (see [capability seams](../../docs/rfc/implemented/architecture/2026-06-13-capability-seams.md)): an abstract compaction interface, a backend that summarizes, and the model-facing tool that consumes it. Only the interface tier exists today; the backend and consumer are deferred. All **product** packages. +A three-package capability seam (see [capability seams](../../docs/rfc/implemented/architecture/2026-06-13-capability-seams.md)): an abstract compaction interface, a backend that summarizes, and the model-facing tool that consumes it. The interface and a first backend (`compact-basic/`) exist; the consumer tool is deferred. All **product** packages. | Package | Role | ctx key | |---|---|---| | `compact/` | Abstract compaction seam (interface + `compact/*` events + `CompactionResult`) | `ctx.compact` | -| `compact-basic/` (deferred) | A backend: char/4 estimation + token-budget retention + `llm.stream()` summarization | (registers `ctx.compact`) | +| `compact-basic/` | A backend: char/4 estimation + token-budget retention + `llm.stream()` summarization | (registers `ctx.compact`) | | `tool-compact/` (deferred) | Model-facing `/compact` tool over `ctx.compact` | (registers on `ctx.tools`) | -The interface lives at `compact/compact/`. Unlike the bash seam, it depends on `dsh-session` and `dsh-llm` — its verbs are defined over a `Session` and its output is the `ContentBlock` vocabulary, so the contract cannot be expressed without naming them. That deviation from the "interface depends only on cordis" guidance is intentional and recorded in the [compaction capability-seam RFC](../../docs/rfc/proposed/feature/2026-06-18-compaction-capability-seam.md). A tokenizer- or template-based backend would replace `compact-basic` without touching the interface or the tool. +The interface lives at `compact/compact/`, the backend at `compact/compact-basic/`. Unlike the bash seam, it depends on `dsh-session` and `dsh-llm` — its verbs are defined over a `Session` and its output is the `ContentBlock` vocabulary, so the contract cannot be expressed without naming them. That deviation from the "interface depends only on cordis" guidance is intentional and recorded in the [compaction capability-seam RFC](../../docs/rfc/implemented/feature/2026-06-18-compaction-capability-seam.md). A tokenizer- or template-based backend would replace `compact-basic` without touching the interface or the tool. diff --git a/packages/compact/compact-basic/README.md b/packages/compact/compact-basic/README.md new file mode 100644 index 0000000000..c195ae42fb --- /dev/null +++ b/packages/compact/compact-basic/README.md @@ -0,0 +1,57 @@ +# @deepseek-ai/dsh-compact-basic + +The **basic compaction backend**: a `BasicCompactService` implementing the `@deepseek-ai/dsh-compact` seam with a char/4 token heuristic, token-budget retention, and summarization routed through the agent request pipeline. + +This is the implementation tier of the compaction capability — see the [interface package](../compact/README.md) for the seam and the [capability-seam RFC](../../../docs/rfc/implemented/feature/2026-06-18-compaction-capability-seam.md) for the design. + +## What it owns + +The abstract contract states only WHAT compaction does; this backend owns every HOW decision: + +- **Token estimation** — `estimateContentTokens()`: char/4 with per-block structural overhead (`text`/`reasoning` = `ceil(len/4) + 4`, `tool-call` from name + arguments, `tool-result` recursive, `image` = 85, unknown blocks via JSON length). +- **Retention policy** — `compactIfNeeded()` walks the surface nodes tail→head summing per-node token estimates, and retains the smallest tail-run of WHOLE units (a closed step, or a single no-step node such as a pre-step `user/message` or inter-step `steering/message`) whose total reaches `retainTokens`; everything older is compacted. Retention is **turn-agnostic** — turn boundaries play no role, so a single runaway turn that alone exceeds the window compacts its OWN early closed steps rather than being retained verbatim (the failure mode that motivated dropping turn-protection: a tool-heavy turn must stay compactable or the harness dies exactly when compaction is needed). The only structural guard is **tool-pairing balance**: the compacted region's edges are balanced cuts on the surface (no unanswered tool-call crosses either edge), so it never splits a step's `assistant/message` tool-calls from their `tool/result`s. When the only compactable content left is an un-splittable open tail step, it declines (returns `null`) and retries once an older step closes. **Single-unit overflow is out of scope, by design**: if one retained unit (a single closed step, or a large pasted `user/message`) ALONE exceeds the budget, compaction cannot help and the call may go out over-budget — bounding an individual unit's size is a separate concern. `compactRegion()` enforces tool-pairing balance strictly, throwing on a boundary that would split a step. `dsh-session` exports `isToolPairingBalanced` for the check. +- **Dynamic convergence** — no static summary-length config pretends to bound what the model will write. If framing/estimator/system overhead leaves the compacted surface above threshold, `compactIfNeeded()` re-compacts the head checkpoint up to `compactionRetries` extra times; if it still cannot get below threshold, it throws. A summary whose estimated stored size is not smaller than the shadowed content fails closed before it mutates the surface. +- **Summarization** — `summarize()`: a `GenerateOptions` request assembled via `BlockAssembler` with a fixed system prompt that asks for a structured checkpoint (Primary Request and Intent · Key Technical Concepts · Files and Code · Errors and Fixes · Pending Tasks · Current Work · Next Step · Critical Context), every section mandatory, exact paths/commands/identifiers preserved. The request runs through the `agent/request` waterfall before `ctx.llm.stream()`, so router agents that choose the concrete model there also route compaction summaries. `maxTokens` is the provider-side generation cap; only text blocks from the model's reply are kept before the checkpoint is stored (reasoning is dropped so private chain-of-thought never leaks into the durable summary, and a stray `tool-call` is dropped so the synthesized `user/message` summary cannot land an orphaned call with no matching `tool-result`). The compacted region is flattened to a plain-text transcript first: text and reasoning contribute their text, and every non-text block (image, tool-call, tool-result, plugin-added types) contributes a type-tagged placeholder (`[image]`, `[tool-call: name(args)]`, …) so the summarizer is told what existed rather than silently dropping it. +- **Checkpoint framing** — the raw summary is not landed directly. `compactRegion()` wraps it in a checkpoint preamble (so a resuming model reads it as a checkpoint, not a fresh user request, and builds on the captured context rather than restating it) plus `` tags. Because region compaction can be invoked manually, a surface may hold several checkpoints, so the framing does not claim everything after it is recent or verbatim. The tags make a prior checkpoint detectable in the transcript on the next compaction cycle: the summarization prompt then instructs the model to merge it in place (preserve still-true facts, drop stale ones) rather than re-summarize it verbatim — a cheap incremental merge that needs no extra log/event machinery. The unframed summary stays on the `compact/summary` provenance event. +- **Surface mutation** — `compactRegion()` appends the `compact/start` → `compact/summary` → `compact/end` log records and the single `user/message` replace node carrying the framed summary (see the interface README). +- **Auto-compaction** — an `agent/pre-step` listener delegates to `compactIfNeeded()` before every step (not just a turn's first — a tool-heavy turn grows the surface mid-turn, so a runaway turn still compacts, and per-step firing is the only moment to rescue it before overflow). `agent/pre-step` is a serial (awaited, in-order) surface-mutation checkpoint that fires after `turn/start` and BEFORE the step opens (`step/start`) and its request history is derived, so compaction mutates the surface — with its log-only `compact/*` records landing cleanly outside any step — and the loop derives once from the result: no double-derive, and the listener cannot see (or need to rewrite) an already-assembled `messages` array. The listener owns no threshold logic of its own (the single token-pressure check lives in `compactIfNeeded()`); because Cordis `serial` bails early on non-void return values, the listener returns `void` and does not use the dispatcher's bail channel as a veto surface. +- **Failure handling** — the `compact/start … compact/end` bracket is a log-recorded lock: it makes a crash mid-summarization a detectable orphan (a `compact/start` with no `compact/end`), records provenance, and prevents a concurrent compaction. Two failure paths: a **crash** (the loop dies mid-summarization) leaves a dangling `compact/start` that is inert — `compact/*` events are log-only, the surface replacement never landed, so the full history derives fine and generic turn-repair closes the turn; a **recoverable** failure (summarization throws but the loop survives) appends `compact/end` with its `error` field set, leaving the surface untouched so the call proceeds with full history. Core session repair stays compaction-agnostic by design — it never learns about `compact/*`. + +`estimateContentTokens()` and `summarize()` are overridable hooks: a tokenizer-based or template-based backend can subclass `BasicCompactService` and override just those, reusing the retention walk and surface plumbing. + +## Config (`BasicCompactConfig`) + +Every knob is **required** except `auto` — there is no concrete data yet to justify default thresholds/budgets, so a consumer states each value explicitly rather than inherit a guessed default. `auto` alone defaults to `true`. + +| Key | Required | Meaning | +|---|---|---| +| `contextWindow` | yes | Context window size in tokens. | +| `thresholdRatio` | yes | Compact when estimated usage exceeds this fraction of the window. | +| `retainTokens` | yes | Tokens of recent context to keep intact. | +| `summarizationModel` | yes | Model for summarization (`''` → use the agent's model). | +| `maxTokens` | yes | Provider generation cap for the summarization call; may include reasoning tokens. | +| `compactionRetries` | yes | Extra compaction attempts after the first if the compacted surface remains over threshold. | +| `auto` | no (default `true`) | Register the `agent/pre-step` auto-compaction listener. Set `false` for manual-only. | + +## Usage + +```ts +import type { Context } from 'cordis' +import { BasicCompactService } from '@deepseek-ai/dsh-compact-basic' + +export const name = 'compact-basic' +export const inject = ['llm'] + +export function apply(ctx: Context): void { + ctx.plugin(BasicCompactService, { + contextWindow: 128000, + thresholdRatio: 0.8, + retainTokens: 20480, + summarizationModel: '', + maxTokens: 8192, + compactionRetries: 1, + }) +} +``` + +Loading the plugin registers `ctx.compact`. With `auto: true` (the default) it compacts automatically under token pressure; a consumer (a future `/compact` tool) can also call `ctx.compact.compactIfNeeded(...)` or `ctx.compact.compactRegion(...)` directly. diff --git a/packages/compact/compact-basic/package.json b/packages/compact/compact-basic/package.json new file mode 100644 index 0000000000..c019796e0d --- /dev/null +++ b/packages/compact/compact-basic/package.json @@ -0,0 +1,42 @@ +{ + "name": "@deepseek-ai/dsh-compact-basic", + "description": "Basic compaction backend (char/4 token estimation + token-budget retention + llm.generate() summarization) for the DeepSeek Harness", + "version": "0.0.1", + "private": true, + "type": "module", + "main": "lib/index.js", + "types": "lib/types/index.d.ts", + "exports": { + ".": { + "types": "./lib/types/index.d.ts", + "default": "./lib/index.js" + }, + "./src/*": "./src/*", + "./package.json": "./package.json" + }, + "files": [ + "lib/index.js", + "lib/types/**/*.d.ts", + "lib/types/**/*.d.ts.map", + "src" + ], + "license": "BSD-3-Clause", + "peerDependencies": { + "@deepseek-ai/dsh-agent": "^0.0.1", + "@deepseek-ai/dsh-compact": "^0.0.1", + "@deepseek-ai/dsh-llm": "^0.0.1", + "@deepseek-ai/dsh-session": "^0.0.1", + "cordis": "^4.0.0-rc.6" + }, + "devDependencies": { + "@deepseek-ai/dsh-agent": "workspace:^", + "@deepseek-ai/dsh-agent-loop": "workspace:^", + "@deepseek-ai/dsh-compact": "workspace:^", + "@deepseek-ai/dsh-invariants": "workspace:^", + "@deepseek-ai/dsh-llm": "workspace:^", + "@deepseek-ai/dsh-session": "workspace:^", + "@deepseek-ai/dsh-system-prompt": "workspace:^", + "@deepseek-ai/dsh-tools": "workspace:^", + "cordis": "^4.0.0-rc.6" + } +} diff --git a/packages/compact/compact-basic/src/index.ts b/packages/compact/compact-basic/src/index.ts new file mode 100644 index 0000000000..f53dace461 --- /dev/null +++ b/packages/compact/compact-basic/src/index.ts @@ -0,0 +1,746 @@ +/** + * `BasicCompactService`: the first implementation of the + * `@deepseek-ai/dsh-compact` seam. It owns the entire compaction strategy: + * + * - **Token estimation** — char/4 heuristic with per-block structural overhead. + * - **Retention policy** — walk surface nodes tail→head, keep recent nodes up + * to a token budget, compact everything older. The cutoff is snapped forward + * to the next balanced tool-pairing boundary so a compacted region never + * splits a step's tool-call/result pair (an open tail step is never crossed — + * compaction declines and retries once it closes). + * - **Summarization** — `ctx.llm.stream()` assembled via `BlockAssembler` + * (the single model-call surface; same path the loop uses) with a fixed + * condense-the-history system prompt routed through `agent/request`. + * - **Surface mutation** — a single `user/message` replace node carries the + * summary; `compact/*` events are log-only lock + provenance records. + * - **Auto-compaction** — an `agent/pre-step` listener delegates to + * {@link BasicCompactService.compactIfNeeded} before EVERY step (so a + * tool-heavy turn that grows the surface mid-turn still compacts); it owns the + * sole token-pressure check. + * + * A different backend (real tokenizer, template summarizer, turn-count + * retention) either subclasses this and overrides the {@link + * BasicCompactService.estimateContentTokens} / {@link + * BasicCompactService.summarize} hooks, or implements the abstract + * {@link CompactService} from scratch. + * + * @module @deepseek-ai/dsh-compact-basic + */ + +import { Context } from 'cordis' +import { CompactService } from '@deepseek-ai/dsh-compact' +import type { CompactionResult } from '@deepseek-ai/dsh-compact' +import { BlockAssembler } from '@deepseek-ai/dsh-llm' +import type { ContentBlock, FinishReason, GenerateOptions, Message } from '@deepseek-ai/dsh-llm' +import type { Session, SessionEvent } from '@deepseek-ai/dsh-session' +import { isToolPairingBalanced } from '@deepseek-ai/dsh-session' +import type { Agent } from '@deepseek-ai/dsh-agent' +import type { BasicCompactConfig, ResolvedConfig } from './types.ts' +import { resolveConfig } from './types.ts' + +export type { BasicCompactConfig, ResolvedConfig } from './types.ts' +export { resolveConfig } from './types.ts' + +/** Per-block structural overhead for JSON framing / type tag. */ +const BLOCK_OVERHEAD = 4 + +/** Heuristic token count for an image block (~85 tokens for low-res URL). */ +const IMAGE_TOKEN_COST = 85 + +/** Role-field framing overhead added per message in {@link BasicCompactService.estimateTokens}. */ +const ROLE_OVERHEAD = 4 + +/** Tags wrapping the structured summary inside the landed checkpoint node. */ +const SUMMARY_OPEN_TAG = '' +const SUMMARY_CLOSE_TAG = '' + +/** + * The summarization system prompt: instructs the model to condense the + * conversation into a fixed, fully-populated structure rather than freeform + * bullets. The fixed structure guarantees coverage of the things a resuming + * model needs (original intent, pending work, the next step, critical context) + * and is stable across compaction cycles, so a prior checkpoint can be merged + * in place. The final rule keys off {@link SUMMARY_OPEN_TAG}: when the + * transcript already contains a prior checkpoint, the model consolidates rather + * than re-summarizing it verbatim (a cheap incremental-merge that needs no + * extra log/event machinery — the tag travels on the summary surface node). + */ +const SUMMARIZE_SYSTEM_PROMPT = [ + 'You are a compaction engine for an AI coding assistant. Condense the conversation transcript into a structured checkpoint that lets another model resume the work with no loss of essential context.', + '', + 'Output EXACTLY the Markdown structure below: keep every section, in order. Use terse bullets, not prose paragraphs. Write "(none)" for an empty section — never drop a section.', + '', + '## Primary Request and Intent', + "- [the user's original and evolving goals; quote verbatim where the exact wording matters]", + '', + '## Key Technical Concepts', + '- [technologies, frameworks, patterns, and conventions in play]', + '', + '## Files and Code', + '- [exact path: why it matters, key changes or snippets]', + '', + '## Errors and Fixes', + '- [error: how it was resolved, plus any related user feedback]', + '', + '## Pending Tasks', + '- [explicitly requested work not yet completed]', + '', + '## Current Work', + '- [precisely what was in progress at this checkpoint]', + '', + '## Next Step', + '- [the single next action, directly in line with the most recent request, or "(none)"]', + '', + '## Critical Context', + '- [decisions and their rationale, constraints, user preferences, open questions, data needed to continue]', + '', + 'Rules:', + '- Preserve exact file paths, commands, error strings, identifiers, and function signatures.', + '- Capture user feedback and explicit instructions faithfully, especially corrections.', + '- Do NOT mention this summarization process or that the context was compacted.', + `- If the transcript already contains a ${SUMMARY_OPEN_TAG} block, it is a PRIOR checkpoint. Do not copy it forward verbatim: preserve still-true facts, drop stale ones, and merge newer information into a single consolidated summary under the same structure.`, +].join('\n') + +/** + * Framing prepended to the landed summary so a resuming model reads it as a + * checkpoint rather than a fresh user request, and continues the task from it. + * It summarizes an earlier span of the conversation; the messages that follow + * are the continuation. Because region compaction can be invoked manually, a + * surface may hold several checkpoints, so the framing does NOT claim that + * everything after it is recent or verbatim — only that the captured context + * should be built on, not restated. + */ +const CHECKPOINT_PREAMBLE = + 'This is an automatically generated checkpoint condensing an earlier span of the conversation to free up context. Treat the captured context as established background and build on it without restating it. Continue the task directly from the messages that follow, without acknowledging this checkpoint.' + +/** + * Map a terminal `FinishReason` to the error a SUMMARIZATION must throw, or + * `undefined` for an acceptable finish. `FinishReason` is merge-extensible. + * + * Compaction fails CLOSED on a truncated summary: `error`, `aborted`, AND + * `max-tokens` all raise. Unlike an ordinary agent turn — where `max-tokens` is + * a normal "the model hit its budget" outcome the loop keeps — a summary cut off + * at the token cap is an INCOMPLETE checkpoint, and committing it would shadow + * (discard) the real history it summarizes. Raising here keeps the original + * surface intact (the caller appends `compact/end` with the error and the auto + * path proceeds with full history). `stop`/future kinds are accepted. + */ +function finishError(finish: FinishReason): Error | undefined { + switch (finish.kind) { + case 'error': { + const error = new Error(finish.message) as Error & { code?: string } + if (finish.code !== undefined) error.code = finish.code + return error + } + case 'aborted': { + const error = new Error('summarization stream aborted') as Error & { code?: string } + error.code = 'ABORTED' + return error + } + case 'max-tokens': { + const error = new Error('summarization truncated at the token cap (incomplete checkpoint)') as Error & { code?: string } + error.code = 'MAX_TOKENS' + return error + } + default: + return undefined + } +} + +/** + * Basic, dependency-light compaction backend. Defaults target a 128K context + * window, compacting at 80% utilization and retaining ~20K tokens of recent + * context. + */ +export class BasicCompactService extends CompactService { + static inject = ['llm'] + + /** Resolved configuration (`auto` defaulted). */ + readonly config: ResolvedConfig + + constructor(ctx: Context, config: BasicCompactConfig) { + super(ctx) + this.config = resolveConfig(config) + + if (this.config.auto) { + // Auto-compaction: delegate to compactIfNeeded before EVERY step. This is + // LOAD-BEARING for runaway-turn survival: a tool-heavy ReAct turn appends + // an assistant/message and a tool/result per step, so the surface (and the + // derived token count) grows WITHIN a turn. The only moment to rescue a + // turn that alone approaches the window is the next step's pre-step + // checkpoint; gating to a turn's first step would let a runaway turn + // overflow before the next turn's check. The listener owns NO threshold + // logic — compactIfNeeded is the single place that decides whether to + // compact, and its in-progress lock serializes concurrent attempts. + // + // It runs on `agent/pre-step` (a serial surface-mutation checkpoint fired + // AFTER turn/start but BEFORE step/start), NOT `agent/request`: compaction + // mutates the session surface, and the loop derives the request `messages` + // AFTER this fires — so a single derive already reflects the compaction, + // with no double-derive and no need to rewrite an already-assembled + // `messages` array. Firing pre-step (outside any open step) keeps the + // log-only `compact/*` records and the replacement node cleanly outside a + // step, so a crash mid-compaction leaves an inert orphan the turn-repair + // closes — never a half-open step. + ctx.on('agent/pre-step', async (agent: Agent, turn: number, step: number, fullSystemPrompt: string, signal: AbortSignal) => { + try { + const result = await this.compactIfNeeded(agent, turn, step, fullSystemPrompt, signal) + if (result) { + const after = this.estimateTokens(agent.session.deriveMessages(), fullSystemPrompt) + ctx.logger.info( + `compaction: shadowed ${result.shadowedSeqs.length} surface nodes ` + + `(seqs ${result.shadowedRange.start}-${result.shadowedRange.end}, ` + + `~${result.shadowedTokenCount} tokens) ` + + `→ ${after} estimated tokens after compaction`, + ) + } + } catch (error: unknown) { + // A failed compaction must not prevent the model call — the surface is + // untouched on failure, so the loop derives the full history and the + // call proceeds. + const msg = error instanceof Error ? error.message : String(error) + ctx.logger.warn(`compaction failed: ${msg}; proceeding with full history`) + } + }) + } + } + + // ---- Token estimation (overridable hooks) ---- + + // TODO: char/4 is a coarse heuristic. Replace with an exact count — a real + // tokenizer, or the provider's post-response `usage` (input tokens) fed back + // as a correction — so threshold decisions match the model's actual budget. + /** + * Estimate the token count of content blocks — char/4 with per-block + * overhead. Override in a subclass to plug in a real tokenizer. + */ + estimateContentTokens(blocks: readonly ContentBlock[]): number { + let tokens = 0 + for (const block of blocks) { + switch (block.type) { + case 'text': + case 'reasoning': + tokens += Math.ceil(block.text.length / 4) + BLOCK_OVERHEAD + break + case 'tool-call': + tokens += Math.ceil(block.name.length / 4) + + Math.ceil(block.arguments.length / 4) + + BLOCK_OVERHEAD + break + case 'tool-result': + tokens += this.estimateContentTokens(block.content) + BLOCK_OVERHEAD + break + case 'image': + tokens += IMAGE_TOKEN_COST + break + default: + // Unknown block types (merge-extensible ContentBlockMap): + // estimate conservatively via JSON stringify. + tokens += BLOCK_OVERHEAD + Math.ceil(JSON.stringify(block).length / 4) + } + } + return tokens + } + + /** + * Estimate token count for a single session event. Returns 0 for non-message + * event types (boundaries, chunks, usage, errors, compact markers). + */ + estimateEventTokens(event: SessionEvent): number { + switch (event.type) { + case 'user/message': + case 'assistant/message': + case 'context/message': + case 'steering/message': + case 'tool/result': + return this.estimateContentTokens(event.data.content) + default: + return 0 + } + } + + /** Estimate total tokens across a list of messages plus optional system prompt. */ + estimateTokens(messages: readonly Message[], systemPrompt?: string): number { + let total = 0 + for (const msg of messages) { + total += this.estimateContentTokens(msg.content) + total += ROLE_OVERHEAD + } + if (systemPrompt) total += Math.ceil(systemPrompt.length / 4) + return total + } + + /** + * Summarize conversation text into content blocks via `agent/request` plus + * `ctx.llm.stream()` assembled through a `BlockAssembler` (the single + * model-call surface). + * Override in a subclass for a template or remote summarizer. + * + * Honors the adapter failure contract: an adapter may report a model failure + * by throwing from `stream()` (propagated here) OR by ending the stream with + * a `finish {kind:'error'|'aborted'}` chunk — the latter is re-thrown so a + * provider error never yields an empty summary. + * + * Forwards `signal` into `GenerateOptions.signal` so an abort/dispose tears + * down the in-flight summarization rather than orphaning the model call. + */ + async summarize(text: string, agent: Agent, turn: number, step: number, signal?: AbortSignal): Promise { + const assembler = new BlockAssembler() + const options: GenerateOptions = { + model: this.config.summarizationModel || agent.options.model || '', + messages: [{ + role: 'user', + content: [{ type: 'text', text: `Summarize this conversation history:\n\n${text}\n\nSummary:` }], + }], + system: SUMMARIZE_SYSTEM_PROMPT, + maxTokens: this.config.maxTokens, + sessionId: agent.session.id, + } + // exactOptionalPropertyTypes: only set `signal` when present — assigning + // `undefined` to an optional `signal?: AbortSignal` is a type error. + if (signal) options.signal = signal + const request = await this.ctx.waterfall('agent/request', agent, turn, step, options, () => Promise.resolve(options)) + if (!request.model) { + throw new Error('no model available for summarization: set BasicCompactConfig.summarizationModel, AgentOptions.model, or supply one via the agent/request waterfall') + } + for await (const chunk of this.ctx.llm.stream(request)) { + assembler.push(chunk) + } + + const error = finishError(assembler.finish) + if (error) throw error + + const summary = this._textOnly(assembler.message().content) + if (!summary.some(block => block.type === 'text' && block.text.trim().length > 0)) { + throw new Error('summarization produced no text summary content') + } + + return summary + } + + // ---- Core API (implements the abstract contract) ---- + + /** + * The sole token-pressure gate: estimate the current surface-derived history, + * and if it exceeds the threshold (`contextWindow * thresholdRatio`), compact + * the oldest surface nodes outside the `retainTokens` budget. The auto- + * compaction listener delegates here rather than pre-checking, so this is the + * only place the decision lives. + * + * Retention is a UNIFORM tail→head walk over the whole surface — turn + * boundaries play NO role. Walking node-by-node from the tail and summing + * token estimates, once the retained total reaches `retainTokens` the cutoff + * is rounded to a balanced tool-pairing boundary: if the cut before the + * retained node is unbalanced (an unanswered tool-call sits before it — i.e. + * it is mid-step), the walk continues head-ward until the cut is balanced so + * the whole step is retained (never splitting a step's tool-calls from their + * results); if it stopped on a free node (a node belonging to no step), that + * cut is already balanced. This always rounds toward retaining MORE (retained + * ≥ `retainTokens`) and is boundary-safe by construction — no separate snap + * pass. + * + * The compacted range is always anchored at the surface HEAD (`nodes[0]`): + * auto-compaction re-consolidates any prior head checkpoint into one fresh + * checkpoint. Declines (`null`) when nothing is over threshold, when the whole + * surface fits the retain budget, or when no balanced cutoff exists in the + * compactable range (its only content is an open tail step — retry once it + * closes). + */ + override async compactIfNeeded( + agent: Agent, + turn: number, + step: number, + fullSystemPrompt: string, + signal: AbortSignal, + ): Promise { + const session = agent.session + const threshold = Math.floor(this.config.contextWindow * this.config.thresholdRatio) + let result: CompactionResult | null = null + for (let attempt = 0; attempt <= this.config.compactionRetries; attempt++) { + const totalTokens = this.estimateTokens(session.deriveMessages(), fullSystemPrompt) + if (totalTokens < threshold) return result + + const range = this._compactableRange(session) + if (range === null) { + /* v8 ignore else -- defensive for non-standard subclass mutations; the concrete replace keeps a compactable head checkpoint. */ + if (result === null) return null + /* v8 ignore next -- paired with the ignored defensive branch above. */ + break + } + + result = await this.compactRegion(session, range.start, range.end, agent, turn, step, signal) + } + + const totalTokens = this.estimateTokens(session.deriveMessages(), fullSystemPrompt) + if (totalTokens < threshold) return result + + throw new Error( + `compaction still above threshold after ${this.config.compactionRetries + 1} compaction attempts ` + + `(${totalTokens} estimated tokens >= threshold ${threshold})`, + ) + } + + override async compactRegion( + session: Session, + start: number, + end: number, + agent: Agent, + turn: number, + step: number, + signal?: AbortSignal, + ): Promise { + // Resolve the range by surface POSITION, not numeric seq interval. A prior + // replace lands a fresh high-seq summary node AT the shadowed range's + // position, so the surface order (head→tail) no longer tracks seq order — + // `[newSummarySeq, olderRetainedSeq, …]` is normal. Indexing into the + // ordered node list and slicing it is the only correct way to read a range; + // a `node.seq >= start && node.seq <= end` interval test would mis-collect + // nodes (and `start > end` would falsely reject) once that happens. + const nodes = session.surface.nodes + const startIdx = nodes.findIndex(n => n.seq === start) + const endIdx = nodes.findIndex(n => n.seq === end) + if (startIdx === -1) throw new Error(`compactRegion: start seq ${start} not found in surface`) + if (endIdx === -1) throw new Error(`compactRegion: end seq ${end} not found in surface`) + if (startIdx > endIdx) { + throw new Error(`compactRegion: start seq ${start} (position ${startIdx}) is after end seq ${end} (position ${endIdx}) on the surface`) + } + + // The region must never split a step's assistant-message tool-calls from + // their tool/results (which would orphan one side and produce a transcript + // every provider rejects). A region is safe iff BOTH its edges are balanced + // cuts: the cut before `start`, and the cut after `end`. A node that belongs + // to no step (pre-step user message, inter-step steering, injection context) + // is a balanced (free) boundary; an `end` inside an open (unclosed) tail step + // leaves the cut after it unbalanced (the open tool-call has no result yet), + // so it is rejected. See dsh-session's tool-pairing balance check. + const events = session.events + if (!isToolPairingBalanced(nodes, events, start)) { + throw new Error(`compactRegion: start seq ${start} is not a balanced boundary (would split a step's tool-call/result pair)`) + } + // The cut after `end` is named by `end`'s surface successor, or `null` when + // `end` is the tail. + // eslint-disable-next-line @typescript-eslint/no-non-null-assertion + const afterEnd: number | null = nodes[endIdx]!.next + if (!isToolPairingBalanced(nodes, events, afterEnd)) { + throw new Error(`compactRegion: end seq ${end} is not a balanced boundary (would split a step, or the step is still open)`) + } + + if (this._isCompactionInProgress(session)) { + throw new Error('compaction already in progress') + } + + // Compaction's events (compact/* and the replacement user/message) must be + // turn-enclosed: the session-log contract rejects any plugin event appended + // outside an open turn. Auto-compaction satisfies this — it runs on the + // `agent/pre-step` seam, after `turn/start` and before `step/start`, so + // strictly inside the open turn (but outside any step). A manual call on a + // fully-closed session has no turn to enclose the events, so reject rather + // than emit an un-enclosed run. + const openTurn = this._openTurn(session) + if (openTurn === null) { + throw new Error('compactRegion: no open turn — compaction events must be enclosed in a turn') + } + // Slice the ordered surface nodes [startIdx, endIdx] inclusive — the + // shadowed range is positional, so this is the set the replace op covers. + const shadowedSeqs = nodes.slice(startIdx, endIdx + 1).map(n => n.seq) + + // --- Acquire lock --- + const startEvent = session.append('compact/start', { turn: openTurn }) + + try { + // --- Extract text and summarize --- + const text = this._extractText(session, shadowedSeqs) + const summary = await this.summarize(text, agent, turn, step, signal) + + // Estimate token count of the shadowed content for provenance. + let shadowedTokenCount = 0 + for (const seq of shadowedSeqs) { + // seq comes from a surface node — always a valid log index by construction. + // eslint-disable-next-line @typescript-eslint/no-non-null-assertion + shadowedTokenCount += this.estimateEventTokens(session.events[seq]!) + } + const framedSummary = this._frameSummary(summary) + const framedSummaryTokenCount = this.estimateContentTokens(framedSummary) + if (framedSummaryTokenCount >= shadowedTokenCount) { + throw new Error( + `summary is not smaller than the shadowed content (${framedSummaryTokenCount} estimated framed tokens >= ${shadowedTokenCount})`, + ) + } + // --- Provenance record (log-only) --- + const summaryEvent = session.append('compact/summary', { + summary, + shadowedRange: { start, end }, + shadowedSeqs, + shadowedTokenCount, + }) + + // --- Surface replacement --- + // The user/message directly shadows all compacted surface nodes with a + // single replace op. It is the ONLY surface event in the compaction + // sequence — compact/start, compact/summary, and compact/end are log-only + // (surfaceOp is rejected by the compiler for non-SurfaceEventType). + // The landed content is FRAMED (checkpoint preamble + tag-wrapped summary); + // the compact/summary provenance event above holds the raw model output. + session.append('user/message', { + content: framedSummary, + source: { kind: 'plugin', plugin: 'compact' }, + }, { + surfaceOp: { op: 'replace', start, end }, + sourceEventSeqs: [startEvent.seq, summaryEvent.seq, ...shadowedSeqs], + }) + + // --- Release lock (log-only) --- + // Appended LAST so the lock brackets the WHOLE operation: a crash between + // compact/start and here leaves a detectable orphaned lock (a compact/start + // with no matching compact/end) rather than a compact/end that falsely + // claims compaction finished before the surface replacement landed. + const endEvent = session.append('compact/end', { turn: openTurn }) + + return { + startSeq: startEvent.seq, + summarySeq: summaryEvent.seq, + endSeq: endEvent.seq, + summary, + shadowedRange: { start, end }, + shadowedSeqs, + shadowedTokenCount, + } + } catch (error: unknown) { + // Always release the lock — append compact/end with the error so a + // wedged lock is impossible. + const msg = error instanceof Error ? error.message : String(error) + session.append('compact/end', { turn: openTurn, error: msg }) + throw error + } + } + + // ---- Internal helpers ---- + + /** + * Frame the raw summary blocks into the content that lands on the surface: + * a checkpoint preamble (so a resuming model reads it as a checkpoint, not a + * fresh user request) followed by the summary wrapped in + * {@link SUMMARY_OPEN_TAG}/{@link SUMMARY_CLOSE_TAG}. The tags make a prior + * checkpoint detectable in the transcript on the next compaction cycle, which + * triggers the merge rule in the summarization prompt. The raw, unframed + * `summary` is preserved separately on the `compact/summary` provenance event. + */ + private _frameSummary(summary: readonly ContentBlock[]): ContentBlock[] { + return [ + { type: 'text', text: `${CHECKPOINT_PREAMBLE}\n\n${SUMMARY_OPEN_TAG}` }, + ...summary, + { type: 'text', text: SUMMARY_CLOSE_TAG }, + ] + } + + /** + * Whether a compaction is currently in progress for `session` — an unmatched + * `compact/start` (no later `compact/end`) WITHIN the current turn. + * + * The scan is scoped to the current turn: walking back from the tail it stops + * at the first `turn/end` (the boundary closing the prior turn). A + * `compact/start` left orphaned by a crash mid-compaction lives in a turn that + * persistence repair then closes with a synthetic `turn/end`; scoping here so + * that a stale orphan from a PAST turn cannot wedge compaction forever (it sits + * before the nearest `turn/end`, so the scan never reaches it). An in-progress + * compaction's `compact/start` is always in the still-open current turn, + * before any `turn/end`, so it is still detected. + */ + private _isCompactionInProgress(session: Session): boolean { + const events = session.events + for (let i = events.length - 1; i >= 0; i--) { + // Index bounded by i >= 0 and i < events.length — never undefined. + // eslint-disable-next-line @typescript-eslint/no-non-null-assertion + const e = events[i]! + if (e.type === 'compact/start') return true + if (e.type === 'compact/end') break + // A turn/end bounds the scan: anything before it belongs to a prior + // (closed) turn and cannot be an in-progress compaction of THIS turn. + if (e.type === 'turn/end') break + } + return false + } + + /** Resolve the next head-anchored compactable surface range, or `null`. */ + private _compactableRange(session: Session): { start: number; end: number } | null { + const nodes = session.surface.nodes + if (nodes.length === 0) return null + + const events = session.events + const retainBudget = this.config.retainTokens + + // Walk tail→head summing per-node token estimates. `keepFromIdx` is the + // index of the OLDEST node we retain verbatim; everything strictly older + // (`[0, keepFromIdx - 1]`) is the compactable range. + let accumulated = 0 + let keepFromIdx = nodes.length // nothing retained yet + for (let i = nodes.length - 1; i >= 0; i--) { + // eslint-disable-next-line @typescript-eslint/no-non-null-assertion + const node = nodes[i]! + const event = events[node.seq] + /* v8 ignore next -- node.seq is a surface-node seq, always a valid log index by construction */ + if (event) accumulated += this.estimateEventTokens(event) + keepFromIdx = i + if (accumulated >= retainBudget) break + } + + // The whole surface fits the retain budget — nothing to compact. + if (keepFromIdx === 0) return null + + // Round the cutoff to a tool-pairing boundary: if the cut before + // `nodes[keepFromIdx]` is unbalanced (an unanswered tool-call sits before + // it — i.e. it is mid-step), extend the retained side head-ward until the + // cut is balanced, so the compacted range ends without splitting an + // assistant↔result pair. A node that belongs to no step is already a + // balanced (free) boundary. Decline if no balanced cut exists at or below + // `keepFromIdx` (the compactable range is only an un-splittable open tail + // step — retry once it closes). + while (keepFromIdx > 0) { + // eslint-disable-next-line @typescript-eslint/no-non-null-assertion + if (isToolPairingBalanced(nodes, events, nodes[keepFromIdx]!.seq)) break + keepFromIdx -= 1 + } + if (keepFromIdx === 0) return null + + // The compacted range is [head … keepFromIdx - 1], anchored at the head. + // eslint-disable-next-line @typescript-eslint/no-non-null-assertion + const firstSeq = nodes[0]!.seq + // eslint-disable-next-line @typescript-eslint/no-non-null-assertion + const cutoffSeq = nodes[keepFromIdx - 1]!.seq + return { start: firstSeq, end: cutoffSeq } + } + + /** + * Keep ONLY text blocks from the model-produced summary before storing it. + * + * The summary lands on the surface as a synthesized `user/message` (see + * {@link _frameSummary}), so the only block type that is both useful and safe + * there is `text`. A model assistant message can otherwise carry `reasoning` + * (private chain-of-thought, must not leak into the durable checkpoint) and + * `tool-call` blocks — and a surviving `tool-call` in a user message would be + * an orphaned call with no matching `tool-result`, exactly the tool-pairing + * breakage compaction works to avoid. Filtering to text drops both. + */ + private _textOnly(blocks: readonly ContentBlock[]): ContentBlock[] { + return blocks.filter((block): block is Extract => block.type === 'text') + } + + /** + * The turn number of the currently OPEN turn — a `turn/start` not yet + * followed by its `turn/end` — or `null` if the session has no open turn. + * + * Compaction's events must be enclosed in a turn, so scanning back from the + * tail: a `turn/start` means that turn is open (return it); a `turn/end` means + * the most recent turn already closed (return null). The whole compaction + * sequence (compact/start … compact/end) is stamped with this turn. + */ + private _openTurn(session: Session): number | null { + for (let i = session.events.length - 1; i >= 0; i--) { + // eslint-disable-next-line @typescript-eslint/no-non-null-assertion + const e = session.events[i]! + if (e.type === 'turn/start') return e.data.turn + if (e.type === 'turn/end') return null + } + return null + } + + /** + * Extract plain-text conversation from a set of surface node seqs, for + * feeding into the summarization model. Walks the seqs in the order given + * (surface order, as `compactRegion` slices the surface-node list) so the + * summary follows the conversation as the model sees it — which, after a + * `replace`, is NOT ascending log-seq order (a high-seq summary node heads the + * surface before older retained lower-seq nodes). + */ + private _extractText(session: Session, seqs: number[]): string { + const lines: string[] = [] + + // Walk seqs in the order given (surface order, as compactRegion slices the + // surface-node list) — NOT ascending log-seq order. After a replace the + // summary node carries a fresh high seq while sitting at the head of the + // surface before older retained lower-seq nodes, so a log-order scan would + // feed the transcript out of order and break the checkpoint-merge prompt. + for (const seq of seqs) { + const event = session.events[seq] + /* v8 ignore next -- seq is a surface-node seq, always a valid log index by construction */ + if (!event) continue + + switch (event.type) { + case 'user/message': { + const text = this._blocksToText(event.data.content) + if (text) lines.push(`User: ${text}`) + break + } + case 'assistant/message': { + const text = this._blocksToText(event.data.content) + if (text) lines.push(`Assistant: ${text}`) + break + } + case 'tool/result': { + const text = this._blocksToText(event.data.content) + const label = event.data.isError ? 'Tool error' : 'Tool result' + if (text) lines.push(`${label} (call ${event.data.callId}): ${text}`) + break + } + case 'context/message': { + const text = this._blocksToText(event.data.content) + if (text) lines.push(`[Context: ${text}]`) + break + } + case 'steering/message': { + const text = this._blocksToText(event.data.content) + if (text) lines.push(`[Steering: ${text}]`) + break + } + // SessionEventMap is merge-extensible — unknown types are + // non-message events that carry no extractable text. + /* v8 ignore next 2 -- seqs only name surface nodes, always one of the 5 handled SurfaceEventTypes; unreachable */ + default: + break + } + } + + return lines.join('\n\n') + } + + /** + * Render content blocks to a single plain-text string for the summarization + * prompt. Text and reasoning contribute their text; every other block type + * contributes a type-tagged placeholder (`[image]`, `[tool-call: name(args)]`, + * …) so the summarizer is told what non-text content existed in the region + * rather than silently losing it. Blocks join with newlines; empty-text + * blocks contribute nothing. + */ + private _blocksToText(blocks: readonly ContentBlock[]): string { + const parts: string[] = [] + for (const block of blocks) { + switch (block.type) { + case 'text': + if (block.text) parts.push(block.text) + break + case 'reasoning': + if (block.text) parts.push(`[reasoning: ${block.text}]`) + break + case 'tool-call': + parts.push(`[tool-call: ${block.name}(${block.arguments})]`) + break + case 'tool-result': { + const inner = this._blocksToText(block.content) + parts.push(inner ? `[tool-result: ${inner}]` : '[tool-result]') + break + } + case 'image': + parts.push('[image]') + break + // ContentBlockMap is merge-extensible — render an unknown block as a + // bare type-tagged placeholder so a plugin-added block type is still + // signalled to the summarizer rather than dropped. + default: + parts.push(`[${(block as ContentBlock).type}]`) + } + } + return parts.join('\n') + } +} + +export default BasicCompactService diff --git a/packages/compact/compact-basic/src/types.ts b/packages/compact/compact-basic/src/types.ts new file mode 100644 index 0000000000..98195d8883 --- /dev/null +++ b/packages/compact/compact-basic/src/types.ts @@ -0,0 +1,81 @@ +/** + * Configuration vocabulary for the basic compaction backend. + * + * Every tunable lives here, in the implementation — the abstract contract + * (`@deepseek-ai/dsh-compact`) carries no config, because thresholds and + * retention policy are HOW decisions a different backend would make + * differently. + * + * @module @deepseek-ai/dsh-compact-basic/types + */ + +/** + * Backend configuration. Every knob is REQUIRED except `auto`: there is no + * concrete data yet to justify default thresholds/budgets, so a consumer must + * state each value explicitly rather than inherit a guessed default. `auto` + * alone defaults to `true` (auto-compaction is the intended posture). + */ +export interface BasicCompactConfig { + /** Context window size in tokens. */ + contextWindow: number + /** Compact when estimated token usage exceeds this fraction of context window. */ + thresholdRatio: number + /** Number of tokens of recent context to retain during compaction. */ + retainTokens: number + /** Model to use for summarization (`''` — uses the agent's model). */ + summarizationModel: string + /** Provider generation cap for the summarization call. */ + maxTokens: number + /** Extra compaction attempts when the first compacted surface is still over threshold. */ + compactionRetries: number + /** Enable automatic compaction on the `agent/pre-step` seam (default true). */ + auto?: boolean +} + +/** Resolved config with `auto` defaulted. */ +export type ResolvedConfig = Required + +/** + * Default `auto` when unset and reject nonsensical numeric knobs. + * + * Convergence is not a static config invariant: provider generation caps can be + * spent on hidden or surfaced reasoning tokens, and the model may emit a summary + * of unpredictable size. The backend instead enforces convergence dynamically: + * each committed summary must be smaller than the content it shadows, and + * `compactIfNeeded` may re-compact up to `compactionRetries` extra times before + * throwing if the surface still exceeds the threshold. + */ +export function resolveConfig(config: BasicCompactConfig): ResolvedConfig { + const resolved: ResolvedConfig = { auto: true, ...config } + + assertPositiveInteger('contextWindow', resolved.contextWindow) + assertRatio('thresholdRatio', resolved.thresholdRatio) + assertNonNegativeInteger('retainTokens', resolved.retainTokens) + assertPositiveInteger('maxTokens', resolved.maxTokens) + assertNonNegativeInteger('compactionRetries', resolved.compactionRetries) + if (typeof resolved.summarizationModel !== 'string') { + throw new Error('BasicCompactConfig: summarizationModel must be a string.') + } + if (typeof resolved.auto !== 'boolean') { + throw new Error('BasicCompactConfig: auto must be a boolean.') + } + return resolved +} + +function assertPositiveInteger(name: string, value: number): void { + if (!Number.isInteger(value) || value <= 0) { + throw new Error(`BasicCompactConfig: ${name} (${value}) must be a positive integer.`) + } +} + +function assertNonNegativeInteger(name: string, value: number): void { + if (!Number.isInteger(value) || value < 0) { + throw new Error(`BasicCompactConfig: ${name} (${value}) must be a non-negative integer.`) + } +} + +function assertRatio(name: string, value: number): void { + if (typeof value !== 'number' || !Number.isFinite(value) || value <= 0 || value > 1) { + throw new Error(`BasicCompactConfig: ${name} (${value}) must be a number in (0, 1].`) + } +} diff --git a/packages/compact/compact-basic/tests/compact-basic.spec.ts b/packages/compact/compact-basic/tests/compact-basic.spec.ts new file mode 100644 index 0000000000..1929e656be --- /dev/null +++ b/packages/compact/compact-basic/tests/compact-basic.spec.ts @@ -0,0 +1,1707 @@ +import { describe, expect, it } from 'vitest' +import { Context } from 'cordis' +import { BasicCompactService } from '@deepseek-ai/dsh-compact-basic' +import type { BasicCompactConfig } from '@deepseek-ai/dsh-compact-basic' +import type { ContentBlock, GenerateOptions, Message, StreamChunk } from '@deepseek-ai/dsh-llm' +import { CallId, LlmAdapter, LlmService } from '@deepseek-ai/dsh-llm' +import SessionStore, { Session, SessionId } from '@deepseek-ai/dsh-session' +import type { SessionEvent, SurfaceEvent } from '@deepseek-ai/dsh-session' +import * as Invariants from '@deepseek-ai/dsh-invariants' +import type { Agent } from '@deepseek-ai/dsh-agent' + +/** A never-aborted signal for the required `compactIfNeeded`/listener arg. */ +const SIGNAL = new AbortController().signal + +/** + * Baseline config with every required knob set. `BasicCompactConfig` has no + * defaults for the numeric/model knobs (only `auto` defaults), so each test + * builds a complete config via `cfg()` and overrides only the knob under test. + */ +const TEST_CONFIG: BasicCompactConfig = { + contextWindow: 128000, + thresholdRatio: 0.8, + retainTokens: 20480, + summarizationModel: '', + maxTokens: 8192, + compactionRetries: 1, +} + +/** A complete config with `overrides` applied over the baseline. */ +function cfg(overrides: Partial = {}): BasicCompactConfig { + return { ...TEST_CONFIG, ...overrides } +} + +/** Long enough that the real checkpoint preamble is smaller than two fixture messages. */ +const LONG_FIXTURE_TEXT = ' Detailed fixture context that makes framed checkpoint compaction genuinely shrinking.'.repeat(20) + +/** + * A BasicCompactService with summarize() stubbed (no real model call) and a + * predictable token estimate, for deterministic unit tests of the algorithm. + */ +class TestCompactService extends BasicCompactService { + private readonly summaryOutputs = new WeakSet() + /** Boundary/unit tests use tiny fixtures; keep framing from dominating them unless a test opts out. */ + estimateFramedSummariesCheaply = true + /** Track calls to summarize for test assertions. */ + summarizeCalls: { text: string; model: string }[] = [] + /** The fixed summary to return. */ + mockSummary: ContentBlock[] = [{ type: 'text', text: 'Test summary of compacted content.' }] + /** Per-call summaries; when set, each summarize() call shifts one value. */ + mockSummaryQueue: ContentBlock[][] = [] + /** If set, summarize() throws this error. */ + summarizeError: Error | null = null + + override estimateContentTokens(blocks: readonly ContentBlock[]): number { + if (this.summaryOutputs.has(blocks)) return blocks.length * 2 + if (this.estimateFramedSummariesCheaply && isFramedCheckpoint(blocks)) return blocks.length * 2 + // 10 tokens per block — predictable for retention/threshold math. + return blocks.length * 10 + } + + override async summarize(text: string, agent: Agent): Promise { + const model = this.config.summarizationModel || agent.options.model || '' + this.summarizeCalls.push({ text, model }) + if (this.summarizeError) throw this.summarizeError + const summary = this.mockSummaryQueue.shift() ?? this.mockSummary + this.summaryOutputs.add(summary) + return summary + } +} + +function isFramedCheckpoint(blocks: readonly ContentBlock[]): boolean { + const first = blocks[0] + const last = blocks[blocks.length - 1] + return first?.type === 'text' + && first.text.includes('') + && last?.type === 'text' + && last.text === '' +} + +/** Create a test service with a throwaway context (auto disabled — no model). */ +function createTestService(overrides: Partial = {}): TestCompactService { + return new TestCompactService(new Context(), cfg({ auto: false, ...overrides })) +} + +/** + * Build a multi-turn session with surface markers (simulating real agent-loop + * output). Compaction always runs inside an OPEN turn (the loop fires the + * `agent/pre-step` seam after a turn's start and before a step's start), so by + * default the session is left with a trailing open turn: turns `1..turns` + * close, then one more `turn/start` opens with no matching `turn/end`. Pass + * `{ leaveOpen: false }` for a fully-closed session (e.g. to assert that manual + * compaction is rejected when no turn is open). + */ +function multiTurnSession(turns: number, messagesPerTurn: number = 2, opts: { leaveOpen?: boolean } = {}): Session { + const leaveOpen = opts.leaveOpen ?? true + const s = new Session(SessionId('test')) + for (let t = 1; t <= turns; t++) { + s.append('turn/start', { turn: t, trigger: { kind: 'message', source: { kind: 'user' } } }) + s.append('step/start', { turn: t, step: 1 }) + for (let m = 0; m < messagesPerTurn; m++) { + s.append('user/message', { + content: [{ type: 'text', text: `turn ${t} user message ${m + 1}.${LONG_FIXTURE_TEXT}` }], + source: { kind: 'user' }, + }, { surfaceOp: 'append' }) + s.append('assistant/message', { + turn: t, step: 1, + content: [{ type: 'text', text: `turn ${t} assistant response ${m + 1}.${LONG_FIXTURE_TEXT}` }], + }, { surfaceOp: 'append' }) + } + s.append('step/end', { turn: t, step: 1 }) + s.append('turn/end', { turn: t, reason: { kind: 'completed' } }) + } + // Open one more turn so compaction's events are turn-enclosed, as they are + // when the loop runs the auto-compaction listener mid-turn. + if (leaveOpen) { + s.append('turn/start', { turn: turns + 1, trigger: { kind: 'message', source: { kind: 'user' } } }) + } + return s +} + +/** Build a session with tool calls for richer extraction tests. */ +function sessionWithTools(): Session { + const s = new Session(SessionId('tools')) + s.append('turn/start', { turn: 1, trigger: { kind: 'message', source: { kind: 'user' } } }) + s.append('step/start', { turn: 1, step: 1 }) + s.append('user/message', { + content: [{ type: 'text', text: 'read file x' }], + source: { kind: 'user' }, + }, { surfaceOp: 'append' }) + s.append('assistant/message', { + turn: 1, step: 1, + content: [ + { type: 'text', text: 'Let me read that file.' }, + { type: 'tool-call', id: CallId('c1'), name: 'bash', arguments: '{"command":"cat x"}' }, + ], + }, { surfaceOp: 'append' }) + s.append('tool/call', { turn: 1, step: 1, callId: CallId('c1'), name: 'bash', arguments: '{"command":"cat x"}' }) + s.append('tool/result', { + turn: 1, step: 1, callId: CallId('c1'), + content: [{ type: 'text', text: 'hello world' }], + isError: false, + }, { surfaceOp: 'append' }) + s.append('assistant/message', { + turn: 1, step: 1, + content: [{ type: 'text', text: 'The file contains: hello world' }], + }, { surfaceOp: 'append' }) + s.append('step/end', { turn: 1, step: 1 }) + s.append('turn/end', { turn: 1, reason: { kind: 'completed' } }) + // Open a trailing turn so compaction's events are turn-enclosed (as they are + // when the loop runs the auto-compaction listener mid-turn). + s.append('turn/start', { turn: 2, trigger: { kind: 'message', source: { kind: 'user' } } }) + return s +} + +/** + * Build a session of `turns` turns, each a SINGLE step containing an + * assistant/message that issues a tool-call plus its tool/result — the real + * multi-node-step shape (a step is two surface nodes: the assistant and the + * result). Each turn is preceded by a user/message. Used to exercise + * step-alignment: a region boundary must not fall between the assistant and its + * result. + */ +function toolTurnSession(turns: number): Session { + const s = new Session(SessionId('tools-multi')) + for (let t = 1; t <= turns; t++) { + s.append('turn/start', { turn: t, trigger: { kind: 'message', source: { kind: 'user' } } }) + s.append('user/message', { + content: [{ type: 'text', text: `turn ${t} request` }], + source: { kind: 'user' }, + }, { surfaceOp: 'append' }) + s.append('step/start', { turn: t, step: 1 }) + s.append('assistant/message', { + turn: t, step: 1, + content: [ + { type: 'text', text: `turn ${t} calling tool` }, + { type: 'tool-call', id: CallId(`c${t}`), name: 'bash', arguments: '{"command":"ls"}' }, + ], + }, { surfaceOp: 'append' }) + s.append('tool/call', { turn: t, step: 1, callId: CallId(`c${t}`), name: 'bash', arguments: '{"command":"ls"}' }) + s.append('tool/result', { + turn: t, step: 1, callId: CallId(`c${t}`), + content: [{ type: 'text', text: `turn ${t} output` }], + isError: false, + }, { surfaceOp: 'append' }) + s.append('step/end', { turn: t, step: 1 }) + s.append('turn/end', { turn: t, reason: { kind: 'completed' } }) + } + // Open a trailing turn so compaction's events are turn-enclosed. + s.append('turn/start', { turn: turns + 1, trigger: { kind: 'message', source: { kind: 'user' } } }) + return s +} + +/** + * Assert the derived transcript has NO orphaned tool-result: every + * `tool-result` block's `toolCallId` must be matched by a preceding `tool-call` + * block in an earlier (assistant) message. A dangling tool-result is exactly + * what splitting a step at compaction produces, and every provider rejects it. + */ +function expectNoOrphanToolResults(messages: Message[]): void { + const seenCallIds = new Set() + for (const msg of messages) { + for (const block of msg.content) { + if (block.type === 'tool-call') seenCallIds.add(block.id) + if (block.type === 'tool-result') { + expect(seenCallIds.has(block.toolCallId), + `orphaned tool-result for callId ${block.toolCallId} (no preceding tool-call)`).toBe(true) + } + } + } +} + +describe('BasicCompactService step-alignment (never split a tool-call/result pair)', () => { + it('compactIfNeeded rounds the retained boundary head-ward to keep a whole step (no orphaned tool-result)', async () => { + // 3 turns, each one step = { assistant(tool-call), tool/result }. Surface + // (9 nodes): user1, asst1, res1, user2, asst2, res2, user3, asst3, res3 — + // 10/20/10 tokens. The tail→head walk retains by whole units; the compacted + // region always ends on a step boundary, so no step's tool-call is split + // from its result. retainTokens=55 keeps the recent tail; the older steps + // compact intact. + const svc = createTestService({ contextWindow: 280, thresholdRatio: 0.5, retainTokens: 55 }) + const session = toolTurnSession(3) + + const result = await compactIfNeeded(svc, session, '', 'm', SIGNAL) + expect(result).not.toBeNull() + expect(result!.shadowedSeqs.length).toBeGreaterThan(0) + // No dangling tool-result: every compacted/retained step stayed whole. + expectNoOrphanToolResults(session.deriveMessages()) + // The most-recent step's result is retained verbatim (still on the surface). + const lastResultSeq = session.events.findLast(e => e.type === 'tool/result')!.seq + expect(result!.shadowedSeqs).not.toContain(lastResultSeq) + }) + + it('compactIfNeeded returns null when the only compactable region is an un-splittable single step', async () => { + // The surface is exactly ONE step: [assistant(tool-call), tool/result]. Over + // threshold (by the derived role overhead), the tail→head walk stops with the + // retained boundary at the tool/result — which is NOT a step-aligned start (its + // issuing assistant precedes it in the same step). Rounding head-ward to find a + // clean boundary reaches index 0, so there is no step-aligned cutoff in the + // compactable range: compactIfNeeded declines rather than splitting the step. + const s = new Session(SessionId('one-step')) + s.append('turn/start', { turn: 1, trigger: { kind: 'message', source: { kind: 'user' } } }) + s.append('step/start', { turn: 1, step: 1 }) + s.append('assistant/message', { + turn: 1, step: 1, + content: [{ type: 'text', text: 'calling' }, { type: 'tool-call', id: CallId('c1'), name: 'bash', arguments: '{}' }], + }, { surfaceOp: 'append' }) + s.append('tool/call', { turn: 1, step: 1, callId: CallId('c1'), name: 'bash', arguments: '{}' }) + s.append('tool/result', { turn: 1, step: 1, callId: CallId('c1'), content: [{ type: 'text', text: 'out' }], isError: false }, { surfaceOp: 'append' }) + s.append('step/end', { turn: 1, step: 1 }) + // Turn stays open. + + const svc = createTestService({ contextWindow: 100, thresholdRatio: 0.1, retainTokens: 5 }) + const result = await compactIfNeeded(svc, s, '', 'm', SIGNAL) + expect(result).toBeNull() + expect(s.events.some(e => e.type === 'compact/start')).toBe(false) + }) + + it('compactRegion rejects a start that splits a step (unbalanced boundary)', async () => { + const svc = createTestService() + const session = toolTurnSession(1) + const nodes = session.surface.nodes // [user, asst(tool-call), result] + const userSeq = nodes[0]!.seq + const resultSeq = nodes[2]!.seq + // start = the tool/result: its issuing assistant precedes it IN THE SAME STEP, + // so starting here would orphan that assistant's tool-call. end is fine (user). + await expect(compactRegion(svc, session, resultSeq, resultSeq, 'm')) + .rejects.toThrow(/start seq .* is not a balanced boundary/) + expect(userSeq).toBeLessThan(resultSeq) // sanity: ordering as expected + }) + + it('compactRegion rejects an end that splits a step (unbalanced boundary)', async () => { + const svc = createTestService() + const session = toolTurnSession(1) + const nodes = session.surface.nodes + const userSeq = nodes[0]!.seq + const asstSeq = nodes[1]!.seq + // end = the assistant/message: its tool/result follows IN THE SAME STEP, so + // ending here would strand that result. start is fine (the pre-step user). + await expect(compactRegion(svc, session, userSeq, asstSeq, 'm')) + .rejects.toThrow(/end seq .* is not a balanced boundary/) + }) + + it('compactRegion rejects an end inside an open tail step', async () => { + const svc = createTestService() + const s = new Session(SessionId('open-tail')) + s.append('turn/start', { turn: 1, trigger: { kind: 'message', source: { kind: 'user' } } }) + s.append('user/message', { content: [{ type: 'text', text: 'go' }], source: { kind: 'user' } }, { surfaceOp: 'append' }) + s.append('step/start', { turn: 1, step: 1 }) + s.append('assistant/message', { + turn: 1, step: 1, + content: [{ type: 'tool-call', id: CallId('c1'), name: 'bash', arguments: '{}' }], + }, { surfaceOp: 'append' }) + const nodes = s.surface.nodes // [user, asst] + const userSeq = nodes[0]!.seq + const asstSeq = nodes[1]!.seq + await expect(compactRegion(svc, s, userSeq, asstSeq, 'm')) + .rejects.toThrow(/end seq .* is not a balanced boundary/) + }) + + it('compactRegion accepts step-aligned boundaries (pre-step user → last result of a closed step)', async () => { + const svc = createTestService() + const session = toolTurnSession(2) + const nodes = session.surface.nodes // [user1, asst1, res1, user2, asst2, res2] + const startSeq = nodes[0]!.seq // pre-step user1 (free boundary) + const endSeq = nodes[2]!.seq // res1 = last node of turn 1's closed step + const result = await compactRegion(svc, session, startSeq, endSeq, 'm') + expect(result.shadowedRange).toEqual({ start: startSeq, end: endSeq }) + expectNoOrphanToolResults(session.deriveMessages()) + }) + + it('compactRegion accepts a single inter-step node (start === end on a pre-step user/message)', async () => { + const svc = createTestService() + const session = toolTurnSession(1) + const nodes = session.surface.nodes + const userSeq = nodes[0]!.seq // pre-step user: free boundary both ways + const result = await compactRegion(svc, session, userSeq, userSeq, 'm') + expect(result.shadowedRange).toEqual({ start: userSeq, end: userSeq }) + }) + + it('compactRegion accepts an injection-turn context node (no step at all)', async () => { + const svc = createTestService() + const s = new Session(SessionId('inject')) + // An idle inject(): turn/start → context/message, NO step. A later turn is + // open so compaction's events are turn-enclosed. + s.append('turn/start', { turn: 1, trigger: { kind: 'injection', source: { kind: 'user' } } }) + s.append('context/message', { content: [{ type: 'text', text: 'ctx' }], source: { kind: 'user' } }, { surfaceOp: 'append' }) + s.append('turn/end', { turn: 1, reason: { kind: 'completed' } }) + s.append('turn/start', { turn: 2, trigger: { kind: 'message', source: { kind: 'user' } } }) + const nodes = s.surface.nodes + const ctxSeq = nodes[0]!.seq + const result = await compactRegion(svc, s, ctxSeq, ctxSeq, 'm') + expect(result.shadowedRange).toEqual({ start: ctxSeq, end: ctxSeq }) + }) +}) + +describe('BasicCompactService.estimateEventTokens', () => { + it('returns 0 for non-message events (boundary, chunk, step/end, tool/call)', () => { + const svc = createTestService() + expect(svc.estimateEventTokens({ type: 'turn/start', seq: 0, time: 1, data: { turn: 1, trigger: { kind: 'message', source: { kind: 'user' } } } })).toBe(0) + expect(svc.estimateEventTokens({ type: 'step/start', seq: 1, time: 2, data: { turn: 1, step: 1 } })).toBe(0) + expect(svc.estimateEventTokens({ type: 'assistant/chunk', seq: 2, time: 3, data: { turn: 1, step: 1, chunk: { type: 'text-delta', index: 0, text: 'h' } } })).toBe(0) + expect(svc.estimateEventTokens({ type: 'step/end', seq: 3, time: 4, data: { turn: 1, step: 1 } })).toBe(0) + expect(svc.estimateEventTokens({ type: 'tool/call', seq: 4, time: 5, data: { turn: 1, step: 1, callId: CallId('c1'), name: 'read', arguments: '{}' } })).toBe(0) + }) + + it('returns estimate for message-producing events', () => { + const svc = createTestService() + const userEvent: SessionEvent = { type: 'user/message', seq: 0, time: 1, data: { content: [{ type: 'text', text: 'hello' }], source: { kind: 'user' } } } + expect(svc.estimateEventTokens(userEvent)).toBe(10) + + const asstEvent: SessionEvent = { type: 'assistant/message', seq: 1, time: 2, data: { turn: 1, step: 1, content: [{ type: 'text', text: 'a' }, { type: 'text', text: 'b' }] } } + expect(svc.estimateEventTokens(asstEvent)).toBe(20) + + const toolEvent: SessionEvent = { type: 'tool/result', seq: 2, time: 3, data: { turn: 1, step: 1, callId: CallId('c1'), content: [{ type: 'text', text: 'output' }], isError: false } } + expect(svc.estimateEventTokens(toolEvent)).toBe(10) + }) +}) + +describe('BasicCompactService.estimateTokens', () => { + it('sums token estimates across messages', () => { + const svc = createTestService() + const messages: Message[] = [ + { role: 'user', content: [{ type: 'text', text: 'hello' }] }, + { role: 'assistant', content: [{ type: 'text', text: 'hi' }, { type: 'text', text: 'there' }] }, + ] + // 1 block * 10 + 4 (role) + 2 blocks * 10 + 4 (role) = 10 + 4 + 20 + 4 = 38 + expect(svc.estimateTokens(messages)).toBe(38) + }) + + it('includes system prompt in the estimate', () => { + const svc = createTestService() + const messages: Message[] = [ + { role: 'user', content: [{ type: 'text', text: 'hi' }] }, + ] + const systemPrompt = 'You are a helpful assistant.' + // 1 block * 10 + 4 (role) + ceil(28/4) = 10 + 4 + 7 = 21 + expect(svc.estimateTokens(messages, systemPrompt)).toBe(21) + }) +}) + +describe('BasicCompactService.compactRegion', () => { + it('shadows surface nodes and inserts a summary via user/message', async () => { + const svc = createTestService() + const session = multiTurnSession(3, 1) // 3 turns, 2 surface nodes each = 6 nodes + + const nodes = session.surface.nodes + expect(nodes.length).toBe(6) + + const firstSeq = nodes[0]!.seq + const secondSeq = nodes[1]!.seq + const result = await compactRegion(svc, session, firstSeq, secondSeq, 'test-model') + + expect(result.shadowedSeqs).toEqual([firstSeq, secondSeq]) + expect(result.shadowedRange.start).toBe(firstSeq) + expect(result.shadowedRange.end).toBe(secondSeq) + expect(result.summary).toEqual(svc.mockSummary) + + const events = session.events + const startEvent = events.findLast(e => e.type === 'compact/start') + const summaryEvent = events.findLast(e => e.type === 'compact/summary') + const endEvent = events.findLast(e => e.type === 'compact/end') + expect(startEvent).toBeDefined() + expect(summaryEvent).toBeDefined() + expect(endEvent).toBeDefined() + + // compact/* events are log-only — no surfaceOp (type system enforces this). + const startRaw = startEvent as unknown as { surfaceOp?: unknown } + expect(startRaw.surfaceOp).toBeUndefined() + + // The user/message carries the replace surfaceOp. + const userMsg = events.findLast(e => e.type === 'user/message')! + const surfaceUserMsg = userMsg as SurfaceEvent + expect(surfaceUserMsg.surfaceOp).toEqual({ op: 'replace', start: firstSeq, end: secondSeq }) + expect(surfaceUserMsg.sourceEventSeqs).toContain(startEvent!.seq) + expect(surfaceUserMsg.sourceEventSeqs).toContain(summaryEvent!.seq) + expect(surfaceUserMsg.sourceEventSeqs).toContain(firstSeq) + expect(surfaceUserMsg.sourceEventSeqs).toContain(secondSeq) + // compact/end is appended AFTER the replacement (the lock brackets the whole + // op), so the replacement cannot reference it — sourceEventSeqs may only + // reference earlier seqs. + expect(surfaceUserMsg.sourceEventSeqs).not.toContain(endEvent!.seq) + expect(endEvent!.seq).toBeGreaterThan(userMsg.seq) + + // Surface now has: summary user/message + retained 4 nodes = 5 nodes. + const newNodes = session.surface.nodes + expect(newNodes.length).toBe(5) + expect(newNodes[0]!.seq).toBe(userMsg.seq) + + // deriveMessages() produces the framed summary as a user-role message: + // a checkpoint preamble + tag-wrapped summary blocks. + const derived = session.deriveMessages() + expect(derived.length).toBe(5) + expect(derived[0]!.role).toBe('user') + const framed = derived[0]!.content + expect(framed[0]).toMatchObject({ type: 'text' }) + expect((framed[0] as { text: string }).text).toContain('') + expect(framed).toContainEqual(svc.mockSummary[0]) + expect((framed[framed.length - 1] as { text: string }).text).toBe('') + }) + + it('throws when start or end are not surface nodes', async () => { + const svc = createTestService() + const session = multiTurnSession(1, 1) + await expect(compactRegion(svc, session, 999, 1000, 'm')) + .rejects.toThrow(/start seq 999 not found in surface/) + }) + + it('throws when start is positioned after end on the surface', async () => { + const svc = createTestService() + const session = multiTurnSession(2, 1) + const nodes = session.surface.nodes + await expect(compactRegion(svc, session, nodes[1]!.seq, nodes[0]!.seq, 'm')) + .rejects.toThrow(/is after end seq .* on the surface/) + }) + + it('throws when compaction is already in progress', async () => { + const svc = createTestService() + const session = multiTurnSession(2, 1) + const nodes = session.surface.nodes + session.append('compact/start', { turn: 2 }) + await expect(compactRegion(svc, session, nodes[0]!.seq, nodes[1]!.seq, 'm')) + .rejects.toThrow(/compaction already in progress/) + }) + + it('appends compact/end with error on summarize failure', async () => { + const svc = createTestService() + svc.summarizeError = new Error('model unavailable') + const session = multiTurnSession(2, 1) + const nodes = session.surface.nodes + + await expect(compactRegion(svc, session, nodes[0]!.seq, nodes[1]!.seq, 'm')) + .rejects.toThrow('model unavailable') + + const endEvent = session.events.findLast(e => e.type === 'compact/end') + expect(endEvent).toBeDefined() + // multiTurnSession(2,…) closes turns 1-2 and leaves turn 3 open; compaction + // stamps the open turn. + expect(endEvent!.data).toMatchObject({ turn: 3, error: 'model unavailable' }) + + // No replace-op user/message was appended (summarize failed). + const userMsgsAfter = session.events.filter(e => e.type === 'user/message') + const replaceMsgs = userMsgsAfter.filter((e) => { + const se = e as unknown as { surfaceOp?: unknown } + return se.surfaceOp !== undefined && typeof se.surfaceOp !== 'string' + }) + expect(replaceMsgs.length).toBe(0) + }) + + it('extracts conversation text for summarization', async () => { + const svc = createTestService() + const session = multiTurnSession(1, 2) + const nodes = session.surface.nodes + + await compactRegion(svc, session, nodes[0]!.seq, nodes[nodes.length - 1]!.seq, 'm') + + expect(svc.summarizeCalls.length).toBe(1) + const { text, model } = svc.summarizeCalls[0]! + expect(model).toBe('m') + expect(text).toContain('User: turn 1 user message 1') + expect(text).toContain('Assistant: turn 1 assistant response 1') + }) + + it('frames the landed summary with a checkpoint preamble and tags, keeping raw provenance', async () => { + const svc = createTestService() + svc.mockSummary = [{ type: 'text', text: 'STRUCTURED SUMMARY' }] + const session = multiTurnSession(3, 1) + const nodes = session.surface.nodes + + const result = await compactRegion(svc, session, nodes[0]!.seq, nodes[1]!.seq, 'm') + + // Provenance (compact/summary) carries the RAW, unframed summary. + expect(result.summary).toEqual([{ type: 'text', text: 'STRUCTURED SUMMARY' }]) + const summaryEvent = session.events.findLast(e => e.type === 'compact/summary')! + expect(summaryEvent.data).toMatchObject({ summary: [{ type: 'text', text: 'STRUCTURED SUMMARY' }] }) + + // The landed surface node is framed: preamble + tag-wrapped summary. + const landed = session.deriveMessages()[0]!.content + expect((landed[0] as { text: string }).text).toContain('checkpoint') + expect((landed[0] as { text: string }).text).toContain('') + expect(landed).toContainEqual({ type: 'text', text: 'STRUCTURED SUMMARY' }) + expect((landed[landed.length - 1] as { text: string }).text).toBe('') + }) + + it('extracts tool-call and tool-result context', async () => { + const svc = createTestService() + const session = sessionWithTools() + const nodes = session.surface.nodes + + const firstSeq = nodes[0]!.seq + const lastSeq = nodes[nodes.length - 1]!.seq + await compactRegion(svc, session, firstSeq, lastSeq, 'm') + + expect(svc.summarizeCalls.length).toBe(1) + const { text } = svc.summarizeCalls[0]! + expect(text).toContain('read file x') + expect(text).toContain('bash') + expect(text).toContain('Tool result') + }) +}) + +describe('BasicCompactService.compactIfNeeded', () => { + it('returns null when tokens are under threshold', async () => { + const svc = createTestService({ contextWindow: 128000, thresholdRatio: 0.8 }) + const session = multiTurnSession(1, 1) + expect(await compactIfNeeded(svc, session, '', 'm', SIGNAL)).toBeNull() + }) + + it('compacts when tokens exceed threshold', async () => { + const svc = createTestService({ contextWindow: 100, thresholdRatio: 0.5, retainTokens: 10 }) + const session = multiTurnSession(3, 1) // 6 surface nodes, 10 tokens each = 60 + + const result = await compactIfNeeded(svc, session, '', 'm', SIGNAL) + expect(result).not.toBeNull() + expect(result!.shadowedSeqs.length).toBeGreaterThan(0) + }) + + it('returns the first compaction result when a zero-retry pass converges after the loop', async () => { + // With compactionRetries=0 there is no next-loop threshold check after the + // first mutation, so the success path is the post-loop `return result`. + const svc = createTestService({ + contextWindow: 100, + thresholdRatio: 0.7, + retainTokens: 10, + compactionRetries: 0, + }) + const session = multiTurnSession(3, 1) // 6 derived messages = 84 estimated tokens. + + const result = await compactIfNeeded(svc, session, '', 'm', SIGNAL) + + expect(result).not.toBeNull() + expect(session.events.filter(e => e.type === 'compact/summary')).toHaveLength(1) + expect(svc.estimateTokens(session.deriveMessages(), '')).toBeLessThan(70) + }) + + it('walks tail→head and retains nodes within token budget', async () => { + const svc = createTestService({ contextWindow: 350, thresholdRatio: 0.2, retainTokens: 15 }) + const session = multiTurnSession(5, 1) // 10 surface nodes = ~100 tokens + + const result = await compactIfNeeded(svc, session, '', 'm', SIGNAL) + expect(result).not.toBeNull() + const nodes = session.surface.nodes + expect(result!.shadowedSeqs.length).toBeGreaterThan(0) + expect(result!.shadowedSeqs).not.toContain(nodes[nodes.length - 1]!.seq) + }) + + it('returns null when the whole surface fits the retain budget (over threshold by role/system overhead)', async () => { + // threshold = floor(480*0.1) = 48. The 4 surface nodes weigh 10 each (raw 40 + // for the retention walk), but the derived estimate adds 4 role tokens per + // message → 56 ≥ 48, so the threshold check passes and the walk runs. The + // walk accumulates all 40 < retainTokens (45) without crossing the budget, + // so keepFromIdx reaches 0 and compaction declines. + const svc = createTestService({ contextWindow: 480, thresholdRatio: 0.1, retainTokens: 45 }) + const session = multiTurnSession(2, 1) + expect(await compactIfNeeded(svc, session, '', 'm', SIGNAL)).toBeNull() + }) + + it('compacts a runaway turn: its early CLOSED steps summarize while recent steps stay verbatim', async () => { + // The REGRESSION that motivated dropping turn-protection. A single in-flight + // (open) turn has grown past the threshold on its own: several CLOSED steps, + // each [assistant(tool-call), tool/result]. Retention is turn-agnostic, so + // the turn's OWN early closed steps are eligible — they compact while the + // recent tail stays verbatim, and the harness survives. + // + // On the OLD layer-2 code this test FAILS: the entire open turn was retained + // verbatim (protectedIdx = first open-turn node = 0), so compactIfNeeded + // returned null and shadowedSeqs would be empty — the runaway turn could + // never compact and the next model call would overflow the window. + const svc = createTestService({ contextWindow: 800, thresholdRatio: 0.1, retainTokens: 25 }) + const s = new Session(SessionId('runaway')) + // ONE open turn with 5 closed steps; each step is [asst(tool-call), result]. + s.append('turn/start', { turn: 1, trigger: { kind: 'message', source: { kind: 'user' } } }) + s.append('user/message', { content: [{ type: 'text', text: 'do a big multi-step task' }], source: { kind: 'user' } }, { surfaceOp: 'append' }) + for (let step = 1; step <= 5; step++) { + s.append('step/start', { turn: 1, step }) + s.append('assistant/message', { + turn: 1, step, + content: [{ type: 'text', text: `step ${step}` }, { type: 'tool-call', id: CallId(`c${step}`), name: 'bash', arguments: '{}' }], + }, { surfaceOp: 'append' }) + s.append('tool/call', { turn: 1, step, callId: CallId(`c${step}`), name: 'bash', arguments: '{}' }) + s.append('tool/result', { turn: 1, step, callId: CallId(`c${step}`), content: [{ type: 'text', text: `out ${step}` }], isError: false }, { surfaceOp: 'append' }) + s.append('step/end', { turn: 1, step }) + } + // The turn stays OPEN (no turn/end) — the model is mid-turn, about to run + // step 6. Surface: user + 5×[asst, result] = 11 nodes. + const nodesBefore = s.surface.nodes.length + expect(nodesBefore).toBe(11) + + const result = await compactIfNeeded(svc, s, '', 'm', SIGNAL) + expect(result).not.toBeNull() + // Early steps of the SAME open turn were shadowed (impossible under layer 2). + expect(result!.shadowedSeqs.length).toBeGreaterThan(0) + // The most-recent step's tool result is retained verbatim (still on surface). + const lastResultSeq = s.events.findLast(e => e.type === 'tool/result')!.seq + expect(result!.shadowedSeqs).not.toContain(lastResultSeq) + expect(s.surface.nodes.some(n => n.seq === lastResultSeq)).toBe(true) + // No orphaned tool-result survives (whole-step boundaries respected). + expectNoOrphanToolResults(s.deriveMessages()) + }) + + it('returns null for an empty surface', async () => { + const svc = createTestService({ contextWindow: 100, thresholdRatio: 0.5, retainTokens: 10 }) + const session = new Session(SessionId('empty')) + expect(await compactIfNeeded(svc, session, '', 'm', SIGNAL)).toBeNull() + }) + + it('compacts again after a prior summary node heads the surface (the summary stays eligible)', async () => { + // After the first compaction lands a replacement summary node at the head, + // a second compaction (still over threshold) re-consolidates it with newer + // context — head-anchoring means the prior checkpoint is always re-included, + // never stranded. retainTokens=25 leaves a couple of retained nodes after + // the first compaction (so the surface is [summary, …retained], not just + // [summary]). + const svc = createTestService({ contextWindow: 800, thresholdRatio: 0.1, retainTokens: 25 }) + const s = multiTurnSession(4, 1) // turns 1-4 closed, turn 5 open (no surface yet) + + const first = await compactIfNeeded(svc, s, '', 'm', SIGNAL) + expect(first).not.toBeNull() + // The summary node now heads the surface with a fresh high seq. + const summaryHeadSeq = s.surface.nodes[0]!.seq + const turn5StartSeq = s.events.filter(e => e.type === 'turn/start').at(-1)!.seq + expect(summaryHeadSeq).toBeGreaterThan(turn5StartSeq) + + // Append a verbatim node in the open turn (a step's output), still over + // threshold, then compact again — the older summary + closed turns compact, + // the fresh nodes are retained. + s.append('step/start', { turn: 5, step: 1 }) + s.append('user/message', { content: [{ type: 'text', text: 'turn 5 work' }], source: { kind: 'user' } }, { surfaceOp: 'append' }) + s.append('assistant/message', { turn: 5, step: 1, content: [{ type: 'text', text: 'reply 5' }] }, { surfaceOp: 'append' }) + s.append('step/end', { turn: 5, step: 1 }) + + const second = await compactIfNeeded(svc, s, '', 'm', SIGNAL) + expect(second).not.toBeNull() + expect(second!.shadowedSeqs.length).toBeGreaterThan(0) + // The fresh open-turn nodes were NOT compacted. + const turn5UserSeq = s.events.find(e => e.type === 'user/message' && e.data.content.some(b => b.type === 'text' && b.text === 'turn 5 work'))!.seq + expect(second!.shadowedSeqs).not.toContain(turn5UserSeq) + }) + + it('re-compacts smaller summaries until the post-compaction surface drops below threshold', async () => { + const svc = createTestService({ + contextWindow: 100, + thresholdRatio: 0.5, + retainTokens: 10, + compactionRetries: 2, + }) + svc.estimateFramedSummariesCheaply = false + svc.mockSummaryQueue = [ + Array.from({ length: 4 }, (_, index) => ({ type: 'text', text: `first ${index}` })), + [{ type: 'text', text: 'second' }], + ] + const session = multiTurnSession(4, 1) + + const result = await compactIfNeeded(svc, session, '', 'm', SIGNAL) + + expect(result).not.toBeNull() + expect(svc.summarizeCalls).toHaveLength(2) + expect(session.events.filter(e => e.type === 'compact/summary')).toHaveLength(2) + expect(svc.estimateTokens(session.deriveMessages(), '')).toBeLessThan(50) + }) + + it('throws after the configured re-compaction attempts still leave the surface above threshold', async () => { + const svc = createTestService({ + contextWindow: 100, + thresholdRatio: 0.5, + retainTokens: 10, + compactionRetries: 1, + }) + svc.estimateFramedSummariesCheaply = false + svc.mockSummaryQueue = [ + Array.from({ length: 4 }, (_, index) => ({ type: 'text', text: `first ${index}` })), + Array.from({ length: 3 }, (_, index) => ({ type: 'text', text: `second ${index}` })), + ] + const session = multiTurnSession(4, 1) + + await expect(compactIfNeeded(svc, session, '', 'm', SIGNAL)) + .rejects.toThrow(/still above threshold after 2 compaction attempts/) + expect(svc.summarizeCalls).toHaveLength(2) + }) +}) + +describe('BasicCompactService replay equivalence', () => { + it('produces identical deriveMessages() after seeding from compacted log', async () => { + const svc = createTestService() + const session = multiTurnSession(3, 1) + const nodes = session.surface.nodes + + await compactRegion(svc, session, nodes[0]!.seq, nodes[1]!.seq, 'm') + const derived = session.deriveMessages() + + const replayed = new Session(SessionId('replay'), [...session.events]) + expect(replayed.deriveMessages()).toEqual(derived) + }) +}) + +describe('BasicCompactService blocking (compaction in progress)', () => { + it('detects in-progress compaction from unmatched compact/start', async () => { + const svc = createTestService() + const session = multiTurnSession(1, 1) + session.append('compact/start', { turn: 1 }) + const nodes = session.surface.nodes + // Whole step (user → assistant) is a step-aligned region, so the call reaches + // the in-progress check rather than being rejected for splitting a step. + await expect(compactRegion(svc, session, nodes[0]!.seq, nodes[1]!.seq, 'm')) + .rejects.toThrow(/compaction already in progress/) + }) + + it('allows compaction after compact/end is appended', async () => { + const svc = createTestService() + const session = multiTurnSession(2, 1) + const nodes = session.surface.nodes + session.append('compact/start', { turn: 1 }) + session.append('compact/end', { turn: 1 }) + const result = await compactRegion(svc, session, nodes[0]!.seq, nodes[1]!.seq, 'm') + expect(result).toBeDefined() + }) + + it('is not wedged by an orphaned compact/start from a prior (now-closed) turn', async () => { + // A crash mid-compaction left a compact/start with no compact/end; the turn + // it lived in was later closed (persistence repair appends turn/end). A + // whole-log scan would treat that stale start as an active lock forever. The + // scan is scoped to the current turn, so a NEW turn compacts normally. + const svc = createTestService() + const s = new Session(SessionId('stale-lock')) + s.append('turn/start', { turn: 1, trigger: { kind: 'message', source: { kind: 'user' } } }) + s.append('step/start', { turn: 1, step: 1 }) + s.append('user/message', { content: [{ type: 'text', text: 'turn 1' }], source: { kind: 'user' } }, { surfaceOp: 'append' }) + s.append('assistant/message', { turn: 1, step: 1, content: [{ type: 'text', text: 'reply 1' }] }, { surfaceOp: 'append' }) + s.append('compact/start', { turn: 1 }) // ← orphaned: no matching compact/end + s.append('step/end', { turn: 1, step: 1 }) + s.append('turn/end', { turn: 1, reason: { kind: 'completed' } }) // repair closed the turn + // A new open turn. + s.append('turn/start', { turn: 2, trigger: { kind: 'message', source: { kind: 'user' } } }) + const nodes = s.surface.nodes + + // The stale start is before the turn/end, so it is NOT seen as in-progress. + const result = await compactRegion(svc, s, nodes[0]!.seq, nodes[1]!.seq, 'm') + expect(result).toBeDefined() + }) +}) + +describe('BasicCompactService token estimation (char/4 heuristic)', () => { + it('estimates text blocks with char/4 + overhead', () => { + const svc = new BasicCompactService(new Context(), cfg({ auto: false })) + // 'this is a somewhat longer text block' = 36 → ceil(36/4)+4 = 13; 'short' = 5 → 2+4 = 6 + const blocks: ContentBlock[] = [ + { type: 'text', text: 'this is a somewhat longer text block' }, + { type: 'text', text: 'short' }, + ] + expect(svc.estimateContentTokens(blocks)).toBe(19) + }) + + it('estimates reasoning blocks same as text', () => { + const svc = new BasicCompactService(new Context(), cfg({ auto: false })) + // 'thinking about this...' = 22 → ceil(22/4)+4 = 10 + expect(svc.estimateContentTokens([{ type: 'reasoning', text: 'thinking about this...' }])).toBe(10) + }) + + it('estimates tool-call blocks from name + arguments', () => { + const svc = new BasicCompactService(new Context(), cfg({ auto: false })) + // 'bash' = 4 → 1; '{"command":"ls"}' = 16 → 4; + 4 overhead = 9 + expect(svc.estimateContentTokens([ + { type: 'tool-call', id: CallId('c1'), name: 'bash', arguments: '{"command":"ls"}' }, + ])).toBe(9) + }) + + it('estimates tool-result blocks recursively', () => { + const svc = new BasicCompactService(new Context(), cfg({ auto: false })) + // inner text 5 → 2+4 = 6; outer 6 + 4 overhead = 10 + expect(svc.estimateContentTokens([ + { type: 'tool-result', toolCallId: CallId('c1'), content: [{ type: 'text', text: 'hello' }], isError: false }, + ])).toBe(10) + }) + + it('estimates image blocks at fixed 85 tokens', () => { + const svc = new BasicCompactService(new Context(), cfg({ auto: false })) + expect(svc.estimateContentTokens([{ type: 'image', url: 'https://example.com/img.png' }])).toBe(85) + }) + + it('returns 0 for empty content blocks', () => { + const svc = new BasicCompactService(new Context(), cfg({ auto: false })) + expect(svc.estimateContentTokens([])).toBe(0) + }) +}) + +describe('BasicCompactService HMR safety', () => { + it('registers as ctx.compact', () => { + const ctx = new Context() + void new BasicCompactService(ctx, cfg({ auto: false })) + expect(ctx.compact).toBeDefined() + expect(ctx.compact).toBeInstanceOf(BasicCompactService) + }) + + it('disposing the plugin fiber unregisters ctx.compact', async () => { + // Mount through the real plugin fiber (the Loader path), then dispose it and + // confirm the service registration is torn down. LlmService is mounted first + // so the service's `inject: ['llm']` resolves and the fiber activates. (The + // sibling-fiber ctx.llm resolution this same setup also exercises is covered + // under the "llm inject (real plugin-load path)" suite.) + const ctx = new Context() + await ctx.plugin(LlmService) + const fiber = await ctx.plugin(BasicCompactService, cfg({ auto: false })) + expect(ctx.get('compact')).toBeInstanceOf(BasicCompactService) + + await fiber.dispose() + expect(ctx.get('compact')).toBeUndefined() + }) +}) + +describe('BasicCompactService config validation', () => { + it('rejects invalid numeric config values', () => { + expect(() => new BasicCompactService(new Context(), cfg({ auto: false, contextWindow: 0 }))) + .toThrow(/contextWindow .* positive integer/) + expect(() => new BasicCompactService(new Context(), cfg({ auto: false, thresholdRatio: 0 }))).toThrow(/thresholdRatio .* \(0, 1\]/) + expect(() => new BasicCompactService(new Context(), cfg({ auto: false, thresholdRatio: 1.1 }))).toThrow(/thresholdRatio .* \(0, 1\]/) + expect(() => new BasicCompactService(new Context(), cfg({ auto: false, retainTokens: -1 }))) + .toThrow(/retainTokens .* non-negative integer/) + expect(() => new BasicCompactService(new Context(), cfg({ auto: false, maxTokens: 0 }))).toThrow(/maxTokens .* positive integer/) + expect(() => new BasicCompactService(new Context(), cfg({ auto: false, compactionRetries: -1 }))) + .toThrow(/compactionRetries .* non-negative integer/) + expect(() => new BasicCompactService( + new Context(), cfg({ auto: false, summarizationModel: 1 } as unknown as Partial), + )).toThrow(/summarizationModel must be a string/) + expect(() => new BasicCompactService(new Context(), cfg({ auto: 'no' } as unknown as Partial))) + .toThrow(/auto must be a boolean/) + }) + + it('accepts a large retain budget because convergence is enforced dynamically', () => { + expect(() => new BasicCompactService(new Context(), cfg({ + auto: false, + contextWindow: 1000, + thresholdRatio: 0.5, + retainTokens: 900, + }))).not.toThrow() + }) + + it('the default config is valid', () => { + expect(() => new BasicCompactService(new Context(), cfg({ auto: false }))).not.toThrow() + }) +}) + +/** An adapter that emits a fixed summary text, for exercising the real summarize() path. */ +class ScriptedAdapter extends LlmAdapter { + lastOptions: GenerateOptions | null = null + constructor(private summaryText: string) { + super() + } + + async * stream(options: GenerateOptions): AsyncIterable { + this.lastOptions = options + yield { type: 'block-start', index: 0, blockType: 'text' } + yield { type: 'text-delta', index: 0, text: this.summaryText } + yield { type: 'finish', reason: { kind: 'stop' } } + } +} + +/** An adapter that emits arbitrary content blocks, preserving reasoning/text shape. */ +class BlocksAdapter extends LlmAdapter { + lastOptions: GenerateOptions | null = null + constructor(private blocks: readonly ContentBlock[]) { + super() + } + + async * stream(options: GenerateOptions): AsyncIterable { + this.lastOptions = options + for (const [index, block] of this.blocks.entries()) { + yield { type: 'block-start', index, blockType: block.type } + switch (block.type) { + case 'text': + yield { type: 'text-delta', index, text: block.text } + break + case 'reasoning': + yield { type: 'reasoning-delta', index, text: block.text } + break + default: + yield { type: 'block-end', index, block } + } + } + yield { type: 'finish', reason: { kind: 'stop' } } + } +} + +/** Wire a real LlmService + arbitrary-block adapter into a context. */ +async function ctxWithBlocks(blocks: readonly ContentBlock[], model = 'test-model'): Promise<{ ctx: Context; adapter: BlocksAdapter }> { + const ctx = new Context() + await ctx.plugin(LlmService) + const adapter = new BlocksAdapter(blocks) + ctx.llm.registerAdapter([model], adapter) + return { ctx, adapter } +} + +/** Wire a real LlmService + scripted adapter into a context. */ +async function ctxWithModel(summaryText: string, model = 'test-model'): Promise<{ ctx: Context; adapter: ScriptedAdapter }> { + const ctx = new Context() + await ctx.plugin(LlmService) + const adapter = new ScriptedAdapter(summaryText) + ctx.llm.registerAdapter([model], adapter) + return { ctx, adapter } +} + +/** An adapter whose stream ends with a finish chunk of the given reason (no content). */ +class FinishOnlyAdapter extends LlmAdapter { + constructor(private reason: StreamChunk & { type: 'finish' }) { + super() + } + + async * stream(): AsyncIterable { + yield this.reason + } +} + +/** Wire a real LlmService + finish-only adapter into a context. */ +async function ctxWithFinish(reason: (StreamChunk & { type: 'finish' })['reason'], model = 'test-model'): Promise { + const ctx = new Context() + await ctx.plugin(LlmService) + ctx.llm.registerAdapter([model], new FinishOnlyAdapter({ type: 'finish', reason })) + return ctx +} + +/** A minimal Agent stub carrying just session + options (enough for the listeners). */ +function stubAgent(session: Session, model?: string): Agent { + return { session, options: { model } } as unknown as Agent +} + +function compactIfNeeded( + svc: BasicCompactService, + session: Session, + fullSystemPrompt: string, + model: string, + signal: AbortSignal, +) { + return svc.compactIfNeeded(stubAgent(session, model), 1, 1, fullSystemPrompt, signal) +} + +function compactRegion( + svc: BasicCompactService, + session: Session, + start: number, + end: number, + model: string, + signal?: AbortSignal, +) { + return svc.compactRegion(session, start, end, stubAgent(session, model), 1, 1, signal) +} + +function summarize(svc: BasicCompactService, text: string, model: string) { + return svc.summarize(text, stubAgent(new Session(SessionId('summary')), model), 1, 1) +} + +describe('BasicCompactService.summarize (real ctx.llm.stream)', () => { + it('summarizes via the registered adapter and returns its content', async () => { + const { ctx, adapter } = await ctxWithModel('SUMMARY TEXT') + const svc = new BasicCompactService(ctx, cfg({ auto: false, maxTokens: 512 })) + + const summary = await summarize(svc, 'User: hi\n\nAssistant: hello', 'test-model') + expect(summary).toEqual([{ type: 'text', text: 'SUMMARY TEXT' }]) + // The fixed system prompt and maxTokens flow through. + expect(adapter.lastOptions!.system).toContain('compaction engine') + expect(adapter.lastOptions!.system).toContain('## Next Step') + expect(adapter.lastOptions!.maxTokens).toBe(512) + expect(adapter.lastOptions!.sessionId).toBe(SessionId('summary')) + expect(adapter.lastOptions!.messages[0]!.content[0]).toMatchObject({ type: 'text' }) + }) + + it('uses maxTokens as the summarization provider cap', async () => { + const { ctx, adapter } = await ctxWithModel('SUMMARY TEXT') + const svc = new BasicCompactService(ctx, cfg({ + auto: false, + maxTokens: 50, + })) + + await summarize(svc, 'User: hi', 'test-model') + + expect(adapter.lastOptions!.maxTokens).toBe(50) + }) + + it('keeps only text blocks in the stored summary (drops reasoning and tool-call)', async () => { + const { ctx } = await ctxWithBlocks([ + { type: 'reasoning', text: 'private chain of thought' }, + { type: 'text', text: 'PUBLIC SUMMARY' }, + // A model reply can carry a tool-call; it must not survive into the + // synthesized user/message summary as an orphaned call. + { type: 'tool-call', id: CallId('c1'), name: 'bash', arguments: '{}' }, + ]) + const svc = new BasicCompactService(ctx, cfg({ auto: false })) + + const summary = await summarize(svc, 'User: hi', 'test-model') + + expect(summary).toEqual([{ type: 'text', text: 'PUBLIC SUMMARY' }]) + }) + + it('throws when no text block remains after filtering', async () => { + const { ctx } = await ctxWithBlocks([{ type: 'reasoning', text: 'private only' }]) + const svc = new BasicCompactService(ctx, cfg({ auto: false })) + + await expect(summarize(svc, 'User: hi', 'test-model')).rejects.toThrow(/no text summary content/) + }) + + it('throws when no model is provided', async () => { + const { ctx } = await ctxWithModel('x') + const svc = new BasicCompactService(ctx, cfg({ auto: false })) + await expect(summarize(svc, 'text', '')).rejects.toThrow(/no model available/) + }) + + it('rethrows when the stream ends with a finish-error chunk', async () => { + const ctx = await ctxWithFinish({ kind: 'error', message: 'provider 401', code: 'UNAUTHORIZED' }) + const svc = new BasicCompactService(ctx, cfg({ auto: false })) + await expect(summarize(svc, 'text', 'test-model')).rejects.toMatchObject({ message: 'provider 401', code: 'UNAUTHORIZED' }) + }) + + it('rethrows a finish-error chunk without a code (code stays undefined)', async () => { + const ctx = await ctxWithFinish({ kind: 'error', message: 'opaque failure' }) + const svc = new BasicCompactService(ctx, cfg({ auto: false })) + const error = await summarize(svc, 'text', 'test-model').then(() => null, (e: unknown) => e as Error & { code?: string }) + expect(error?.message).toBe('opaque failure') + expect(error?.code).toBeUndefined() + }) + + it('rethrows when the stream ends with a finish-aborted chunk', async () => { + const ctx = await ctxWithFinish({ kind: 'aborted' }) + const svc = new BasicCompactService(ctx, cfg({ auto: false })) + await expect(summarize(svc, 'text', 'test-model')).rejects.toMatchObject({ message: 'summarization stream aborted', code: 'ABORTED' }) + }) + + it('fails closed on a max-tokens finish (an incomplete checkpoint must not commit)', async () => { + const ctx = await ctxWithFinish({ kind: 'max-tokens' }) + const svc = new BasicCompactService(ctx, cfg({ auto: false })) + await expect(summarize(svc, 'text', 'test-model')).rejects.toMatchObject({ code: 'MAX_TOKENS' }) + }) + + it('compactRegion leaves the surface intact when summarization hits max-tokens', async () => { + const ctx = await ctxWithFinish({ kind: 'max-tokens' }) + const svc = new BasicCompactService(ctx, cfg({ auto: false })) + const session = multiTurnSession(2, 1) + const before = [...session.surface.nodes] + const nodes = session.surface.nodes + + await expect(compactRegion(svc, session, nodes[0]!.seq, nodes[1]!.seq, 'test-model')) + .rejects.toMatchObject({ code: 'MAX_TOKENS' }) + + // No replacement landed — the surface is byte-identical, and the lock was + // released with the error (compact/end carries it). + expect(session.surface.nodes).toEqual(before) + const endEvent = session.events.findLast(e => e.type === 'compact/end')! + const endData = endEvent.data as { error?: string } + expect(endData.error).toContain('truncated') + }) + + it('compactRegion uses the real summarizer end-to-end', async () => { + const { ctx } = await ctxWithModel('CONDENSED') + const svc = new BasicCompactService(ctx, cfg({ auto: false })) + const session = multiTurnSession(2, 1) + const nodes = session.surface.nodes + + const result = await compactRegion(svc, session, nodes[0]!.seq, nodes[1]!.seq, 'test-model') + expect(result.summary).toEqual([{ type: 'text', text: 'CONDENSED' }]) + // The raw summary is wrapped in the checkpoint framing on the surface. + expect(session.deriveMessages()[0]!.content).toContainEqual({ type: 'text', text: 'CONDENSED' }) + }) + + it('rejects a summary that is not smaller than the shadowed content', async () => { + const svc = createTestService({ auto: false }) + const session = multiTurnSession(2, 1) + const nodes = session.surface.nodes + svc.mockSummary = Array.from({ length: 20 }, (_, index) => ({ type: 'text', text: `large ${index}` })) + + await expect(compactRegion(svc, session, nodes[0]!.seq, nodes[1]!.seq, 'm')) + .rejects.toThrow(/summary is not smaller than the shadowed content/) + expect(session.events.some(e => e.type === 'compact/summary')).toBe(false) + }) + + it('rejects when the framed checkpoint is not smaller than the shadowed content', async () => { + const svc = createTestService({ auto: false }) + svc.estimateFramedSummariesCheaply = false + const session = new Session(SessionId('framed-nonshrinking')) + session.append('turn/start', { turn: 1, trigger: { kind: 'message', source: { kind: 'user' } } }) + session.append('step/start', { turn: 1, step: 1 }) + session.append('user/message', { content: [{ type: 'text', text: 'tiny user' }], source: { kind: 'user' } }, { surfaceOp: 'append' }) + session.append('assistant/message', { turn: 1, step: 1, content: [{ type: 'text', text: 'tiny assistant' }] }, { surfaceOp: 'append' }) + session.append('step/end', { turn: 1, step: 1 }) + session.append('turn/end', { turn: 1, reason: { kind: 'completed' } }) + session.append('turn/start', { turn: 2, trigger: { kind: 'message', source: { kind: 'user' } } }) + const before = [...session.surface.nodes] + const nodes = session.surface.nodes + + await expect(compactRegion(svc, session, nodes[0]!.seq, nodes[1]!.seq, 'm')) + .rejects.toThrow(/summary is not smaller than the shadowed content/) + expect(session.events.some(e => e.type === 'compact/summary')).toBe(false) + expect(session.surface.nodes).toEqual(before) + }) +}) + +describe('BasicCompactService auto-compaction (agent/pre-step listener)', () => { + /** Fire the agent/pre-step serial checkpoint as the loop does. */ + function firePreStep(ctx: Context, agent: Agent, step: number, fullSystemPrompt: string): Promise { + return ctx.serial('agent/pre-step', agent, 1, step, fullSystemPrompt, SIGNAL) + } + + it('compacts (mutating the surface) when over threshold', async () => { + const { ctx } = await ctxWithModel('SUMMARY') + void new BasicCompactService(ctx, cfg({ contextWindow: 200, thresholdRatio: 0.5, retainTokens: 20 })) + const session = multiTurnSession(5, 1) // 10 surface nodes + const agent = stubAgent(session, 'test-model') + const before = session.surface.nodes.length + + await firePreStep(ctx, agent, 1, '') + + // The surface shrank in place, and a summary checkpoint landed. + expect(session.surface.nodes.length).toBeLessThan(before) + expect(session.events.some(e => e.type === 'compact/summary')).toBe(true) + // The re-derived head message is the framed summary checkpoint. + expect(session.deriveMessages()[0]!.content).toContainEqual({ type: 'text', text: 'SUMMARY' }) + }) + + it('logs compaction details when auto-compaction returns a converged result', async () => { + const ctx = new Context() + const infos: string[] = [] + ctx.logger.info = ((msg: string) => void infos.push(msg)) as typeof ctx.logger.info + void new TestCompactService(ctx, cfg({ + contextWindow: 100, + thresholdRatio: 0.7, + retainTokens: 10, + compactionRetries: 0, + })) + const session = multiTurnSession(3, 1) + const agent = stubAgent(session, 'test-model') + + await firePreStep(ctx, agent, 1, '') + + expect(session.events.filter(e => e.type === 'compact/summary')).toHaveLength(1) + expect(infos.some(msg => msg.includes('compaction: shadowed'))).toBe(true) + expect(infos.some(msg => msg.includes('estimated tokens after compaction'))).toBe(true) + }) + + it('compacts mid-turn on steps after the first (the surface grows within a turn)', async () => { + const { ctx } = await ctxWithModel('SUMMARY') + void new BasicCompactService(ctx, cfg({ contextWindow: 100, thresholdRatio: 0.5, retainTokens: 10 })) + const session = multiTurnSession(3, 1) // over the 0.5 threshold + const agent = stubAgent(session, 'test-model') + + // A step-2 checkpoint (a tool-heavy turn's later step) must still compact — + // the surface accumulated assistant/message + tool/result nodes since step 1. + await firePreStep(ctx, agent, 2, '') + expect(session.events.some(e => e.type === 'compact/start')).toBe(true) + }) + + it('does nothing when under threshold', async () => { + const { ctx } = await ctxWithModel('SUMMARY') + void new BasicCompactService(ctx, cfg({ contextWindow: 128000, thresholdRatio: 0.8 })) + const session = multiTurnSession(1, 1) + const agent = stubAgent(session, 'test-model') + + await firePreStep(ctx, agent, 1, '') + expect(session.events.some(e => e.type === 'compact/start')).toBe(false) + }) + + it('leaves the surface intact when compaction fails (summarize rejects)', async () => { + // No adapter registered for this model → summarize() rejects → caught, the + // surface is untouched (the loop derives the full history). + const ctx = new Context() + await ctx.plugin(LlmService) + void new BasicCompactService(ctx, cfg({ contextWindow: 300, thresholdRatio: 0.1, retainTokens: 10 })) + const session = multiTurnSession(3, 1) + const agent = stubAgent(session, 'missing-model') + const before = session.surface.nodes.length + + await firePreStep(ctx, agent, 1, '') + // No summary landed; the surface is unchanged. + expect(session.events.some(e => e.type === 'compact/summary')).toBe(false) + expect(session.surface.nodes.length).toBe(before) + }) + + it('does not register the listener when auto is false', async () => { + const { ctx } = await ctxWithModel('SUMMARY') + void new BasicCompactService(ctx, cfg({ auto: false, contextWindow: 100, thresholdRatio: 0.1, retainTokens: 5 })) + const session = multiTurnSession(3, 1) + const agent = stubAgent(session, 'test-model') + + await firePreStep(ctx, agent, 1, '') + expect(session.events.some(e => e.type === 'compact/start')).toBe(false) + }) + + it('routes summarization through agent/request so router agents can choose the model', async () => { + const { ctx, adapter } = await ctxWithModel('ROUTED SUMMARY', 'routed-model') + ctx.on('agent/request', async (_agent, _turn, _step, options, next) => { + options.model = 'routed-model' + return next() + }) + void new BasicCompactService(ctx, cfg({ contextWindow: 200, thresholdRatio: 0.5, retainTokens: 20 })) + const session = multiTurnSession(5, 1) + const agent = stubAgent(session) + + await ctx.serial('agent/pre-step', agent, 1, 1, '', SIGNAL) + + expect(adapter.lastOptions?.model).toBe('routed-model') + expect(session.events.some(e => e.type === 'compact/summary')).toBe(true) + expect(session.deriveMessages()[0]!.content).toContainEqual({ type: 'text', text: 'ROUTED SUMMARY' }) + }) + + it('removes the auto pre-step listener when the plugin fiber is disposed', async () => { + const { ctx } = await ctxWithModel('SUMMARY') + const fiber = await ctx.plugin(BasicCompactService, cfg({ + contextWindow: 200, + thresholdRatio: 0.5, + retainTokens: 20, + })) + const session = multiTurnSession(5, 1) + const agent = stubAgent(session, 'test-model') + + await fiber.dispose() + await firePreStep(ctx, agent, 1, '') + + expect(session.events.some(e => e.type === 'compact/start')).toBe(false) + expect(ctx.get('compact')).toBeUndefined() + }) +}) + +describe('BasicCompactService._extractText branches', () => { + it('renders reasoning, context, and steering messages', async () => { + const svc = createTestService() + const s = new Session(SessionId('rich')) + s.append('turn/start', { turn: 1, trigger: { kind: 'message', source: { kind: 'user' } } }) + s.append('step/start', { turn: 1, step: 1 }) + s.append('context/message', { + content: [{ type: 'text', text: 'project context here' }], + source: { kind: 'user' }, + }, { surfaceOp: 'append' }) + s.append('assistant/message', { + turn: 1, step: 1, + content: [{ type: 'reasoning', text: 'thinking hard' }, { type: 'text', text: 'answer' }], + }, { surfaceOp: 'append' }) + s.append('steering/message', { + turn: 1, + content: [{ type: 'text', text: 'steer this way' }], + source: { kind: 'user' }, + }, { surfaceOp: 'append' }) + s.append('step/end', { turn: 1, step: 1 }) + s.append('turn/end', { turn: 1, reason: { kind: 'completed' } }) + s.append('turn/start', { turn: 2, trigger: { kind: 'message', source: { kind: 'user' } } }) + + const nodes = s.surface.nodes + await compactRegion(svc, s, nodes[0]!.seq, nodes[nodes.length - 1]!.seq, 'm') + + const { text } = svc.summarizeCalls[0]! + expect(text).toContain('[Context: project context here]') + expect(text).toContain('[reasoning: thinking hard]') + expect(text).toContain('[Steering: steer this way]') + }) + + it('labels tool errors distinctly from tool results', async () => { + const svc = createTestService() + const s = new Session(SessionId('toolerr')) + s.append('turn/start', { turn: 1, trigger: { kind: 'message', source: { kind: 'user' } } }) + s.append('step/start', { turn: 1, step: 1 }) + s.append('user/message', { content: [{ type: 'text', text: 'run it' }], source: { kind: 'user' } }, { surfaceOp: 'append' }) + s.append('assistant/message', { + turn: 1, step: 1, + content: [{ type: 'tool-call', id: CallId('c9'), name: 'bash', arguments: '{}' }], + }, { surfaceOp: 'append' }) + s.append('tool/call', { turn: 1, step: 1, callId: CallId('c9'), name: 'bash', arguments: '{}' }) + s.append('tool/result', { + turn: 1, step: 1, callId: CallId('c9'), + content: [{ type: 'text', text: 'boom failure' }], + isError: true, + }, { surfaceOp: 'append' }) + s.append('step/end', { turn: 1, step: 1 }) + s.append('turn/end', { turn: 1, reason: { kind: 'completed' } }) + s.append('turn/start', { turn: 2, trigger: { kind: 'message', source: { kind: 'user' } } }) + + const nodes = s.surface.nodes + await compactRegion(svc, s, nodes[0]!.seq, nodes[nodes.length - 1]!.seq, 'm') + expect(svc.summarizeCalls[0]!.text).toContain('Tool error (call c9): boom failure') + }) +}) + +describe('BasicCompactService edge cases', () => { + it('renders bare and nested tool-result placeholders and unknown blocks', async () => { + const svc = createTestService() + const s = new Session(SessionId('toolresult')) + s.append('turn/start', { turn: 1, trigger: { kind: 'message', source: { kind: 'user' } } }) + s.append('step/start', { turn: 1, step: 1 }) + // assistant/message carrying a nested tool-result block, an unknown block, + // and the tool-call that the following tool/result answers (so the surface + // is tool-pairing balanced). + s.append('assistant/message', { + turn: 1, step: 1, + content: [ + { type: 'tool-result', toolCallId: CallId('n1'), content: [{ type: 'image', url: 'https://x/n.png' }] }, + { type: 'custom-widget', payload: 'x' } as unknown as ContentBlock, + { type: 'tool-call', id: CallId('b1'), name: 'bash', arguments: '{}' }, + ], + }, { surfaceOp: 'append' }) + // tool/result whose content is itself only non-text → bare '[tool-result]'. + s.append('tool/call', { turn: 1, step: 1, callId: CallId('b1'), name: 'bash', arguments: '{}' }) + s.append('tool/result', { + turn: 1, step: 1, callId: CallId('b1'), + content: [{ type: 'tool-result', toolCallId: CallId('inner'), content: [] }], + isError: false, + }, { surfaceOp: 'append' }) + s.append('step/end', { turn: 1, step: 1 }) + s.append('turn/end', { turn: 1, reason: { kind: 'completed' } }) + s.append('turn/start', { turn: 2, trigger: { kind: 'message', source: { kind: 'user' } } }) + + const nodes = s.surface.nodes + await compactRegion(svc, s, nodes[0]!.seq, nodes[nodes.length - 1]!.seq, 'm') + const { text } = svc.summarizeCalls[0]! + expect(text).toContain('[tool-result: [image]]') // nested tool-result with content + expect(text).toContain('[custom-widget]') // unknown block placeholder + expect(text).toContain('Tool result (call b1): [tool-result]') // empty nested → bare placeholder + }) + + it('estimates unknown block types via JSON length (default branch)', () => { + const svc = new BasicCompactService(new Context(), cfg({ auto: false })) + // A block whose type is none of the known kinds — exercises the default arm. + const unknown = { type: 'custom-widget', payload: 'some data' } as unknown as ContentBlock + expect(svc.estimateContentTokens([unknown])).toBeGreaterThan(0) + }) + + it('auto-compaction reports bounded retry exhaustion after committing a smaller summary', async () => { + const { ctx } = await ctxWithModel('SUMMARY') + const warnings: string[] = [] + ctx.logger.warn = ((msg: string) => void warnings.push(msg)) as typeof ctx.logger.warn + void new BasicCompactService(ctx, cfg({ + contextWindow: 300, + thresholdRatio: 0.1, + retainTokens: 5, + compactionRetries: 0, + })) + const session = multiTurnSession(4, 1) + const agent = stubAgent(session, 'test-model') + + await ctx.serial('agent/pre-step', agent, 1, 1, '', SIGNAL) + expect(session.events.some(e => e.type === 'compact/summary')).toBe(true) + // The surface was mutated; the head message is the framed summary checkpoint. + expect(session.deriveMessages()[0]!.content).toContainEqual({ type: 'text', text: 'SUMMARY' }) + expect(warnings.some(w => w.includes('still above threshold after 1 compaction attempts'))).toBe(true) + }) + + it('rejects compaction when no turn is open (compaction events must be turn-enclosed)', async () => { + const svc = createTestService() + // A session whose only turn has CLOSED — scanning back from the tail hits + // turn/end before any turn/start, so there is no open turn to enclose + // compaction's compact/* + replacement events, which the log contract forbids. + const s = new Session(SessionId('noturn')) + s.append('turn/start', { turn: 1, trigger: { kind: 'message', source: { kind: 'user' } } }) + s.append('step/start', { turn: 1, step: 1 }) + s.append('user/message', { content: [{ type: 'text', text: 'orphan' }], source: { kind: 'user' } }, { surfaceOp: 'append' }) + s.append('assistant/message', { turn: 1, step: 1, content: [{ type: 'text', text: 'reply' }] }, { surfaceOp: 'append' }) + s.append('step/end', { turn: 1, step: 1 }) + s.append('turn/end', { turn: 1, reason: { kind: 'completed' } }) + const nodes = s.surface.nodes + + await expect(compactRegion(svc, s, nodes[0]!.seq, nodes[1]!.seq, 'm')) + .rejects.toThrow(/no open turn/) + // The lock was never acquired — no compact/start landed. + expect(s.events.some(e => e.type === 'compact/start')).toBe(false) + }) + + it('rejects compaction on a session with no turn boundaries at all', async () => { + const svc = createTestService() + // No turn events whatsoever — the open-turn scan falls through to the end + // of the log and finds none, so compaction is rejected (its events have no + // turn to enclose them). + const s = new Session(SessionId('turnless')) + s.append('user/message', { content: [{ type: 'text', text: 'orphan' }], source: { kind: 'user' } }, { surfaceOp: 'append' }) + const nodes = s.surface.nodes + + await expect(compactRegion(svc, s, nodes[0]!.seq, nodes[0]!.seq, 'm')) + .rejects.toThrow(/no open turn/) + expect(s.events.some(e => e.type === 'compact/start')).toBe(false) + }) + + it('compactIfNeeded returns null for empty surface even when over threshold', async () => { + const svc = createTestService({ contextWindow: 1000, thresholdRatio: 0.1, retainTokens: 5 }) + const session = new Session(SessionId('empty-but-pressured')) + // No surface nodes, but a large system prompt pushes the estimate over threshold. + const bigPrompt = 'x'.repeat(800) // ceil(800/4) = 200 tokens >> threshold 100 + expect(await compactIfNeeded(svc, session, bigPrompt, 'm', SIGNAL)).toBeNull() + }) + + it('compactRegion throws when end is not a surface node (start valid)', async () => { + const svc = createTestService() + const session = multiTurnSession(1, 1) + const nodes = session.surface.nodes + await expect(compactRegion(svc, session, nodes[0]!.seq, 9999, 'm')) + .rejects.toThrow(/end seq 9999 not found in surface/) + }) + + it('compactRegion stringifies a non-Error thrown by summarize', async () => { + const svc = createTestService() + // Throw a non-Error value to exercise the String(error) branch in the catch. + svc.summarizeError = 'plain string failure' as unknown as Error + const session = multiTurnSession(1, 1) + const nodes = session.surface.nodes + + // Whole step (user → assistant): a step-aligned region that reaches summarize. + await expect(compactRegion(svc, session, nodes[0]!.seq, nodes[1]!.seq, 'm')).rejects.toBe('plain string failure') + const endEvent = session.events.findLast(e => e.type === 'compact/end')! + expect(endEvent.data).toMatchObject({ error: 'plain string failure' }) + }) + + it('auto-compaction listener stringifies a non-Error and proceeds', async () => { + const { ctx } = await ctxWithModel('SUMMARY') + const warnings: string[] = [] + ctx.logger.warn = ((msg: string) => void warnings.push(msg)) as typeof ctx.logger.warn + const svc = new TestCompactService(ctx, cfg({ contextWindow: 300, thresholdRatio: 0.1, retainTokens: 10 })) + svc.summarizeError = 'boom' as unknown as Error + const session = multiTurnSession(3, 1) + const agent = stubAgent(session, 'test-model') + const before = session.surface.nodes.length + + await ctx.serial('agent/pre-step', agent, 1, 1, '', SIGNAL) + // The failure was swallowed; the surface is untouched and a warning logged. + expect(session.surface.nodes.length).toBe(before) + expect(session.events.some(e => e.type === 'compact/summary')).toBe(false) + expect(warnings.some(w => w.includes('compaction failed: boom'))).toBe(true) + }) + + it('auto-compaction listener takes the result-null branch (nothing to compact)', async () => { + const { ctx } = await ctxWithModel('SUMMARY') + // A large system prompt pushes the listener's estimate over threshold, but + // retainTokens is huge so compactIfNeeded walks everything and returns null. + // threshold = floor(2000*0.1) = 200; invariant: 5 + 150 = 155 ≤ 200. + const svc = new TestCompactService(ctx, cfg({ contextWindow: 2000, thresholdRatio: 0.1, retainTokens: 150 })) + const session = multiTurnSession(2, 1) + const agent = stubAgent(session, 'test-model') + const bigSystem = 'x'.repeat(900) // ceil(900/4)=225 > threshold 200 + + await ctx.serial('agent/pre-step', agent, 1, 1, bigSystem, SIGNAL) + expect(session.events.some(e => e.type === 'compact/start')).toBe(false) + expect(svc.summarizeCalls.length).toBe(0) + }) + + it('skips messages whose extracted text is empty across all kinds', async () => { + const svc = createTestService() + const s = new Session(SessionId('empties')) + s.append('turn/start', { turn: 1, trigger: { kind: 'message', source: { kind: 'user' } } }) + // Step 1: an empty-text user, an empty-reasoning assistant with NO tool-call + // (balanced: nothing to answer), and empty context/steering — all extract to + // nothing and are skipped. + s.append('step/start', { turn: 1, step: 1 }) + s.append('user/message', { content: [{ type: 'text', text: '' }], source: { kind: 'user' } }, { surfaceOp: 'append' }) + s.append('assistant/message', { turn: 1, step: 1, content: [{ type: 'reasoning', text: '' }] }, { surfaceOp: 'append' }) + s.append('context/message', { content: [], source: { kind: 'user' } }, { surfaceOp: 'append' }) + s.append('steering/message', { turn: 1, content: [{ type: 'text', text: '' }], source: { kind: 'user' } }, { surfaceOp: 'append' }) + s.append('step/end', { turn: 1, step: 1 }) + // Step 2: a tool exchange whose tool/result has empty content → empty + // extraction → skipped. The assistant carries the matching tool-call so the + // surface stays tool-pairing balanced; its text extracts to the tool-call + // placeholder (the one surviving line). + s.append('step/start', { turn: 1, step: 2 }) + s.append('assistant/message', { + turn: 1, step: 2, + content: [{ type: 'tool-call', id: CallId('z1'), name: 'bash', arguments: '{}' }], + }, { surfaceOp: 'append' }) + s.append('tool/call', { turn: 1, step: 2, callId: CallId('z1'), name: 'bash', arguments: '{}' }) + s.append('tool/result', { turn: 1, step: 2, callId: CallId('z1'), content: [], isError: false }, { surfaceOp: 'append' }) + s.append('step/end', { turn: 1, step: 2 }) + s.append('turn/end', { turn: 1, reason: { kind: 'completed' } }) + s.append('turn/start', { turn: 2, trigger: { kind: 'message', source: { kind: 'user' } } }) + + const nodes = s.surface.nodes + await compactRegion(svc, s, nodes[0]!.seq, nodes[nodes.length - 1]!.seq, 'm') + // Every empty-content message (user text, empty reasoning, empty-content + // tool/result, empty context, empty steering) extracted to nothing and was + // skipped — the only surviving line is the assistant's tool-call (which a + // balanced surface requires to answer the tool/result). + expect(svc.summarizeCalls[0]!.text).toBe('Assistant: [tool-call: bash({})]') + }) + + it('renders non-text blocks as type-tagged placeholders across all message kinds', async () => { + const svc = createTestService() + const s = new Session(SessionId('placeholders')) + s.append('turn/start', { turn: 1, trigger: { kind: 'message', source: { kind: 'user' } } }) + s.append('step/start', { turn: 1, step: 1 }) + // user/message with only an image block → '[image]' placeholder. + s.append('user/message', { content: [{ type: 'image', url: 'https://x/y.png' }], source: { kind: 'user' } }, { surfaceOp: 'append' }) + // assistant/message with an image block AND the tool-call its tool/result + // answers (so the surface is tool-pairing balanced) → '[image]' placeholder. + s.append('assistant/message', { + turn: 1, step: 1, + content: [ + { type: 'image', url: 'https://x/z.png' }, + { type: 'tool-call', id: CallId('e1'), name: 'bash', arguments: '{}' }, + ], + }, { surfaceOp: 'append' }) + // tool/result with an image block → '[image]' placeholder. + s.append('tool/call', { turn: 1, step: 1, callId: CallId('e1'), name: 'bash', arguments: '{}' }) + s.append('tool/result', { turn: 1, step: 1, callId: CallId('e1'), content: [{ type: 'image', url: 'https://x/r.png' }], isError: false }, { surfaceOp: 'append' }) + // context/message and steering/message with image content. + s.append('context/message', { content: [{ type: 'image', url: 'https://x/c.png' }], source: { kind: 'user' } }, { surfaceOp: 'append' }) + s.append('steering/message', { turn: 1, content: [{ type: 'image', url: 'https://x/s.png' }], source: { kind: 'user' } }, { surfaceOp: 'append' }) + s.append('step/end', { turn: 1, step: 1 }) + s.append('turn/end', { turn: 1, reason: { kind: 'completed' } }) + s.append('turn/start', { turn: 2, trigger: { kind: 'message', source: { kind: 'user' } } }) + + const nodes = s.surface.nodes + await compactRegion(svc, s, nodes[0]!.seq, nodes[nodes.length - 1]!.seq, 'm') + const { text } = svc.summarizeCalls[0]! + // Every non-text block surfaces as a placeholder rather than being dropped. + expect(text).toContain('User: [image]') + expect(text).toContain('Assistant: [image]') + expect(text).toContain('Tool result (call e1): [image]') + expect(text).toContain('[Context: [image]]') + expect(text).toContain('[Steering: [image]]') + }) + +}) + +describe('BasicCompactService positional range (surface seqs are not monotonic after a replace)', () => { + it('compacts a second region after the first replace lands a high-seq summary at the head position', async () => { + // A replace inserts the new summary node (a high seq) AT the shadowed + // range's surface position, so the surface becomes + // [highSeqSummary, …olderRetainedLowerSeqs]. A second compaction over a + // range whose start node has a HIGHER seq than its end node must still + // succeed — the range is positional, not a numeric seq interval. + const svc = createTestService({ auto: false }) + const session = multiTurnSession(4, 1) + + // First compaction: shadow the two oldest surface nodes. + const nodes0 = session.surface.nodes + const first = await compactRegion(svc, session, nodes0[0]!.seq, nodes0[1]!.seq, 'm') + + // The summary node now sits at the head with a seq HIGHER than the + // retained older nodes that follow it — the non-monotonic surface. (The + // head is the user/message replace node, appended after the compact/summary + // provenance event, so its seq is at least first.summarySeq.) + const nodes1 = session.surface.nodes + expect(nodes1[0]!.seq).toBeGreaterThanOrEqual(first.summarySeq) + expect(nodes1[0]!.seq).toBeGreaterThan(nodes1[1]!.seq) + + // Second compaction: shadow [summary(head) … turn-2's step end]. The start + // seq (the head summary node) is GREATER than the end seq (an older retained + // node), so the range is a SURFACE-POSITION span, not a numeric seq interval. + // The end must land on a step boundary (turn-2's assistant message closes + // its step). + const startSeq = nodes1[0]!.seq + const endSeq = nodes1[2]!.seq + expect(startSeq).toBeGreaterThan(endSeq) + const second = await compactRegion(svc, session, startSeq, endSeq, 'm') + + // Exactly the three nodes at surface positions [0..2] are shadowed, in + // surface order — the positional slice, regardless of their seq values. + expect(second.shadowedSeqs).toEqual([nodes1[0]!.seq, nodes1[1]!.seq, nodes1[2]!.seq]) + // The surface still derives cleanly: a new head replace node + the rest. + const finalNodes = session.surface.nodes + expect(finalNodes[0]!.seq).toBeGreaterThanOrEqual(second.summarySeq) + expect(session.deriveMessages().length).toBe(finalNodes.length) + }) + + it('extracts the second-compaction transcript in surface order, not log-seq order', async () => { + const svc = createTestService({ auto: false }) + const session = multiTurnSession(3, 1) + + // First compaction shadows the oldest two surface nodes, landing a high-seq + // summary node at the head. + const n0 = session.surface.nodes + await compactRegion(svc, session, n0[0]!.seq, n0[1]!.seq, 'm') + + // Second compaction spans [head summary … turn-2's step end]. The head's seq + // is higher than the older retained nodes' seqs, so a log-seq-order walk + // would emit the older messages BEFORE the checkpoint. + const n1 = session.surface.nodes + svc.summarizeCalls = [] + await compactRegion(svc, session, n1[0]!.seq, n1[2]!.seq, 'm') + + // The extracted transcript follows surface order: the checkpoint (head) + // first, then the older retained messages — matching deriveMessages(). + const { text } = svc.summarizeCalls[0]! + const checkpointIdx = text.indexOf('compacted-summary') + const olderIdx = text.indexOf('turn 2 user') + expect(checkpointIdx).toBeGreaterThanOrEqual(0) + expect(olderIdx).toBeGreaterThan(checkpointIdx) + }) +}) + +describe('BasicCompactService llm inject (real plugin-load path)', () => { + it('declares llm in static inject so a sibling fiber can resolve ctx.llm', () => { + // summarize() reads ctx.llm; the inject lets the cordis ctx proxy resolve a + // sibling LlmService when this service is mounted as its own plugin fiber. + // Asserting the declaration (and exercising the real mount below) guards the + // resolution that root-ctx unit tests cannot, since they share one fiber. + expect(BasicCompactService.inject).toContain('llm') + }) + + it('resolves ctx.llm and summarizes when mounted as a sibling plugin of LlmService', async () => { + const ctx = new Context() + await ctx.plugin(LlmService) + ctx.llm.registerAdapter(['test-model'], new ScriptedAdapter('CONDENSED')) + // Mount the service through its real plugin fiber (NOT new …(rootCtx)), so + // the sibling-fiber ctx.llm resolution actually exercises the inject. + const fiber = await ctx.plugin(BasicCompactService, cfg({ auto: false })) + + const svc = ctx.compact as BasicCompactService + const session = multiTurnSession(2, 1) + const nodes = session.surface.nodes + const result = await compactRegion(svc, session, nodes[0]!.seq, nodes[1]!.seq, 'test-model') + expect(result.summary).toEqual([{ type: 'text', text: 'CONDENSED' }]) + + // Tear the fiber down so this test owns no leaked registration; the + // dedicated cleanup assertion lives in the "HMR safety" suite. + await fiber.dispose() + expect(ctx.get('compact')).toBeUndefined() + }) +}) + +describe('BasicCompactService under the real invariants plugin', () => { + /** + * Drive compaction through a session whose `session/event` listeners include + * the real dev-mode invariants plugin (as a real app loads it via agent-core). + * The invariants throw on append, so a passing run proves the compaction + * sequence is contract-valid: every event is turn-enclosed, and the positional + * replace op is accepted even when the surface is no longer seq-ordered. + */ + async function setup(): Promise<{ ctx: Context; session: Session; svc: BasicCompactService }> { + const ctx = new Context() + await ctx.plugin(SessionStore) + await ctx.plugin(Invariants, {}) + await ctx.plugin(LlmService) + ctx.llm.registerAdapter(['test-model'], new ScriptedAdapter('CONDENSED')) + await ctx.plugin(BasicCompactService, cfg({ auto: false })) + const session = ctx.sessions.create() + return { ctx, session, svc: ctx.compact as BasicCompactService } + } + + /** Append one closed turn of [user, assistant] surface nodes via the store. */ + function closedTurn(session: Session, turn: number): void { + session.append('turn/start', { turn, trigger: { kind: 'message', source: { kind: 'user' } } }) + session.append('step/start', { turn, step: 1 }) + session.append('user/message', { content: [{ type: 'text', text: `turn ${turn} user.${LONG_FIXTURE_TEXT}` }], source: { kind: 'user' } }, { surfaceOp: 'append' }) + session.append('assistant/message', { turn, step: 1, content: [{ type: 'text', text: `turn ${turn} assistant.${LONG_FIXTURE_TEXT}` }] }, { surfaceOp: 'append' }) + session.append('step/end', { turn, step: 1 }) + session.append('turn/end', { turn, reason: { kind: 'completed' } }) + } + + it('runs a turn-enclosed compaction whose positional replace the invariants accept', async () => { + const { session, svc } = await setup() + closedTurn(session, 1) + closedTurn(session, 2) + // Open turn 3, as the loop has when the auto-compaction listener fires. + session.append('turn/start', { turn: 3, trigger: { kind: 'message', source: { kind: 'user' } } }) + + const nodes = session.surface.nodes + // No invariant throws here: compact/* + the replacement are all in turn 3. + const result = await compactRegion(svc, session, nodes[0]!.seq, nodes[1]!.seq, 'test-model') + expect(result.shadowedSeqs.length).toBe(2) + expect(session.surface.nodes[0]!.seq).toBeGreaterThan(session.surface.nodes[1]!.seq) + }) + + it('accepts a second compaction over the non-monotonic surface left by the first', async () => { + const { session, svc } = await setup() + closedTurn(session, 1) + closedTurn(session, 2) + closedTurn(session, 3) + session.append('turn/start', { turn: 4, trigger: { kind: 'message', source: { kind: 'user' } } }) + + const n0 = session.surface.nodes + await compactRegion(svc, session, n0[0]!.seq, n0[1]!.seq, 'test-model') + + // Surface head now carries a higher seq than the older retained nodes. A + // second compaction spanning [head … a later closed-step end] must pass the + // invariants' positional replace check even though startSeq > endSeq. + const n1 = session.surface.nodes + expect(n1[0]!.seq).toBeGreaterThan(n1[2]!.seq) + const second = await compactRegion(svc, session, n1[0]!.seq, n1[2]!.seq, 'test-model') + expect(second.shadowedSeqs).toEqual([n1[0]!.seq, n1[1]!.seq, n1[2]!.seq]) + }) +}) diff --git a/packages/compact/compact-basic/tests/compact-loop-repro.spec.ts b/packages/compact/compact-basic/tests/compact-loop-repro.spec.ts new file mode 100644 index 0000000000..319e30a73c --- /dev/null +++ b/packages/compact/compact-basic/tests/compact-loop-repro.spec.ts @@ -0,0 +1,156 @@ +import { describe, expect, it } from 'vitest' +import { Context } from 'cordis' +import LlmService from '@deepseek-ai/dsh-llm' +import type { ContentBlock, GenerateOptions, StreamChunk } from '@deepseek-ai/dsh-llm' +import { CallId, LlmAdapter } from '@deepseek-ai/dsh-llm' +import SessionStore from '@deepseek-ai/dsh-session' +import { isToolPairingBalanced } from '@deepseek-ai/dsh-session' +import SystemPrompt from '@deepseek-ai/dsh-system-prompt' +import ToolRegistry, { defineTool } from '@deepseek-ai/dsh-tools' +import AgentRegistry, { AgentId } from '@deepseek-ai/dsh-agent' +import AgentLoop, { ReactLoopAgent } from '@deepseek-ai/dsh-agent-loop' +import * as Invariants from '@deepseek-ai/dsh-invariants' +import { BasicCompactService } from '@deepseek-ai/dsh-compact-basic' +import type { SurfaceEvent } from '@deepseek-ai/dsh-session' + +/** + * CBR-001 regression: a compaction checkpoint that the REAL loop lands is a + * free surface boundary (it carries no tool-call/result pair), so it must be a + * valid region edge on BOTH sides. A surface-anchored balance check sees that; + * the abandoned log-position scan did not. + * + * The loop fires the compaction seam mid-flight, so the landed checkpoint + * `user/message{replace}` sits at a HIGH log seq positioned beside the current + * step even though its SURFACE position is the head. A log-position forward scan + * from the checkpoint reaches the step's own later `assistant/message` and + * wrongly reports the checkpoint as mid-step — refusing it as a region end. A + * SECOND compaction that re-summarizes just that head checkpoint (region end == + * checkpoint) therefore throws and is swallowed, so the surface never + * re-consolidates. + * + * This drives a real auto-compaction through the agent-loop and asserts the + * landed checkpoint balances on both sides AND that re-compacting it (end == + * checkpoint) succeeds. RED on the log-position predicates; GREEN once alignment + * is decided from surface tool-pairing balance. + */ + +const TOKENS_PER_BLOCK = 10 + +class ReproCompactService extends BasicCompactService { + override estimateContentTokens(blocks: readonly ContentBlock[]): number { + return blocks.length * TOKENS_PER_BLOCK + } + + override async summarize(): Promise { + return [{ type: 'text', text: 'CHECKPOINT SUMMARY' }] + } +} + +/** Each call emits one tool-call until exhausted, then a final text answer. */ +class StepwiseToolAdapter extends LlmAdapter { + calls = 0 + constructor(private toolSteps: number) { + super() + } + + async * stream(_options: GenerateOptions): AsyncIterable { + const n = this.calls + this.calls += 1 + if (n < this.toolSteps) { + const id = CallId(`c${n}`) + const args = `{"i":${n}}` + yield { type: 'block-start', index: 0, blockType: 'text' } + yield { type: 'block-end', index: 0, block: { type: 'text', text: `step ${n}` } } + yield { type: 'block-start', index: 1, blockType: 'tool-call' } + yield { type: 'block-end', index: 1, block: { type: 'tool-call', id, name: 'work', arguments: args } } + yield { type: 'finish', reason: { kind: 'tool-calls' } } + return + } + yield { type: 'block-start', index: 0, blockType: 'text' } + yield { type: 'block-end', index: 0, block: { type: 'text', text: 'all done' } } + yield { type: 'finish', reason: { kind: 'stop' } } + } +} + +async function harness(toolSteps: number): Promise<{ ctx: Context; compact: ReproCompactService }> { + const ctx = new Context() + await ctx.plugin(LlmService) + await ctx.plugin(SessionStore) + await ctx.plugin(Invariants, {}) + await ctx.plugin(SystemPrompt) + await ctx.plugin(ToolRegistry) + await ctx.plugin(AgentRegistry) + await ctx.plugin(AgentLoop, { agents: [] }) + ctx.llm.registerAdapter(['mock'], new StepwiseToolAdapter(toolSteps)) + ctx.tools.register(defineTool({ + name: 'work', + description: 'does work', + parameters: { i: { type: 'number' } }, + async execute() { + return [{ type: 'text', text: 'work result' }] + }, + })) + // Tiny window so a couple of tool steps cross the threshold and compaction + // fires within the runaway turn. + const compact = new ReproCompactService(ctx, { + auto: true, + contextWindow: 64, + thresholdRatio: 0.5, + retainTokens: 20, + summarizationModel: '', + maxTokens: 8192, + compactionRetries: 1, + }) + return { ctx, compact } +} + +function waitForIdle(ctx: Context, agent: ReactLoopAgent): Promise { + return new Promise((resolve) => { + const dispose = ctx.on('agent/status', (subject, status) => { + if (subject === agent && status === 'idle') { + dispose() + resolve() + } + }) + }) +} + +describe('CBR-001: a real-loop checkpoint is a valid boundary on both sides', () => { + it('the head checkpoint the loop lands is a balanced cut on both sides', async () => { + const { ctx } = await harness(8) + try { + const agent = ctx.agentLoop.create(AgentId('repro'), { model: 'mock' }) + agent.send([{ type: 'text', text: 'do a long multi-step task' }]) + await waitForIdle(ctx, agent) + + const events = [...agent.session.events] + // A compaction ran: at least one checkpoint landed on the surface. + const checkpoints = events.filter( + (e): e is SurfaceEvent => + e.type === 'user/message' + && typeof (e as SurfaceEvent).surfaceOp === 'object', + ) + expect(checkpoints.length).toBeGreaterThan(0) + + // The loop fired compaction mid-flight, so each landed checkpoint sits at a + // high log seq beside the step it landed in, even though its SURFACE + // position is the head of the range it shadowed. A checkpoint carries no + // tool-call/result pair (only summarized prose), so every checkpoint still + // on the surface must be a balanced cut on BOTH sides — the cut before it + // (region START) and the cut after it (region END). The abandoned + // log-position scan reported the END as mis-aligned because the forward log + // scan reached the neighbouring step's assistant/message. + const nodes = agent.session.surface.nodes + for (const cp of checkpoints) { + const node = nodes.find(n => n.seq === cp.seq) + if (!node) continue // shadowed by a later checkpoint — no longer an edge. + expect(isToolPairingBalanced(nodes, events, node.seq), + `checkpoint seq ${node.seq} must be a balanced region START`).toBe(true) + expect(isToolPairingBalanced(nodes, events, node.next), + `checkpoint seq ${node.seq} must be a balanced region END`).toBe(true) + } + } finally { + await ctx.fiber.dispose() + } + }) +}) diff --git a/packages/compact/compact-basic/tsconfig.json b/packages/compact/compact-basic/tsconfig.json new file mode 100644 index 0000000000..075c64cb61 --- /dev/null +++ b/packages/compact/compact-basic/tsconfig.json @@ -0,0 +1,16 @@ +{ + "extends": "../../../tsconfig.base.json", + "compilerOptions": { + "rootDir": "src", + "outDir": "lib/types" + }, + "include": ["src"], + "references": [ + { "path": "../../../vendor/cosmokit" }, + { "path": "../../../vendor/cordis" }, + { "path": "../../llm/llm" }, + { "path": "../../core/session" }, + { "path": "../../core/agent" }, + { "path": "../compact" } + ] +} diff --git a/packages/compact/compact/README.md b/packages/compact/compact/README.md index 9ef5b73005..b6f3cc0920 100644 --- a/packages/compact/compact/README.md +++ b/packages/compact/compact/README.md @@ -7,10 +7,10 @@ This package is the interface tier of the compaction capability, split so each c | Package | Role | |---|---| | `@deepseek-ai/dsh-compact` (this) | the interface: abstract service + `compact/*` events + `CompactionResult` | -| `@deepseek-ai/dsh-compact-basic` | a backend: char/4 estimation + token-budget retention + `llm.stream()` summarization | +| `@deepseek-ai/dsh-compact-basic` (deferred) | a backend: char/4 estimation + token-budget retention + `llm.stream()` summarization | | `@deepseek-ai/dsh-tool-compact` (deferred) | the model-facing `/compact` tool over `ctx.compact` | -Unlike the bash seam, this interface depends on `@deepseek-ai/dsh-session` and `@deepseek-ai/dsh-llm` — the contract's verbs are defined over a `Session` and its output is the `ContentBlock` vocabulary, so they cannot be expressed without naming those packages. That deviation from the "interface depends only on cordis" guidance is intentional and recorded in the [compaction capability-seam RFC](../../../docs/rfc/proposed/feature/2026-06-18-compaction-capability-seam.md). +Unlike the bash seam, this interface depends on `@deepseek-ai/dsh-session` and `@deepseek-ai/dsh-llm` — the contract's verbs are defined over a `Session` and its output is the `ContentBlock` vocabulary, so they cannot be expressed without naming those packages. That deviation from the "interface depends only on cordis" guidance is intentional and recorded in the [compaction capability-seam RFC](../../../docs/rfc/implemented/feature/2026-06-18-compaction-capability-seam.md). ## Service API (`ctx.compact`) @@ -18,10 +18,10 @@ Both methods are **abstract** — the backend owns the entire strategy (token es | Member | Semantics | |---|---| -| `compactIfNeeded(session, systemPrompt?, model?, signal?)` | Estimate the history size; if over the backend's threshold, compact an older range via `compactRegion`, keeping recent context intact. Returns the `CompactionResult`, or `null` if nothing needed compacting. | -| `compactRegion(session, start, end, model, signal?)` | Forcibly summarize surface nodes `[start, end]` (inclusive seqs) into a single replacement node. **Throws** if a compaction is already in progress, if `start`/`end` aren't surface nodes, or if `start > end`. | +| `compactIfNeeded(agent, turn, step, fullSystemPrompt, signal)` | Estimate the surface-derived history size; if over the backend's threshold, compact an older range via `compactRegion`, keeping recent context intact. Returns the `CompactionResult`, or `null` if nothing needed compacting. All parameters required — the loop's `agent/pre-step` checkpoint supplies the agent, lifecycle context, assembled `fullSystemPrompt`, and turn `signal`; router-aware summarizers can use the agent lifecycle context to route their own model call through `agent/request`. | +| `compactRegion(session, start, end, agent, turn, step, signal?)` | Forcibly summarize surface nodes `[start, end]` (inclusive seqs) into a single replacement node. **Throws** if a compaction is already in progress, if `start`/`end` aren't surface nodes, or if `start` is positioned after `end` on the surface. The range is a SURFACE-POSITION span, not a numeric seq interval — after a prior replace lands a fresh high-seq summary node at the shadowed range's position, surface order no longer tracks seq order. | -Both methods take an optional `signal: AbortSignal`. A backend that summarizes via `ctx.llm.stream()` **must** forward it into the call's `GenerateOptions.signal`, so an abort or fiber dispose tears down the in-flight summarization instead of leaving an orphaned model call running past the cancellation. The turn that the `compact/*` events belong to is not a parameter — it is recoverable from the log (the currently-open turn), so the backend stamps it without the caller supplying it. +`compactIfNeeded` takes a required `signal`; `compactRegion`'s is optional. A backend that summarizes via `ctx.llm.stream()` **must** forward it into the call's `GenerateOptions.signal`, so an abort or fiber dispose tears down the in-flight summarization instead of leaving an orphaned model call running past the cancellation. The session being compacted comes from the agent context; the turn that the `compact/*` events belong to is recoverable from the log (the currently-open turn), so the backend stamps it from the log rather than trusting a caller-supplied value. ## Surface contract @@ -53,4 +53,4 @@ The `compact/*` events extend `SessionEventMap` (merge-extensible) via declarati ## Implementing a backend -Subclass `CompactService`, implement `compactIfNeeded` and `compactRegion`, and load the subclass as a plugin — it registers as `ctx.compact`. See `@deepseek-ai/dsh-compact-basic` for the reference implementation. +Subclass `CompactService`, implement `compactIfNeeded` and `compactRegion`, and load the subclass as a plugin — it registers as `ctx.compact`. A tokenizer-, template-, or model-backed implementation can live as a sibling package without changing callers. diff --git a/packages/compact/compact/src/index.ts b/packages/compact/compact/src/index.ts index 9e58e5c905..f5d03fe2ac 100644 --- a/packages/compact/compact/src/index.ts +++ b/packages/compact/compact/src/index.ts @@ -6,17 +6,17 @@ * Implementations subclass {@link CompactService}, implement * {@link CompactService.compactIfNeeded} and {@link CompactService.compactRegion}, * and load as a plugin — registering as `ctx.compact` (one implementation per - * context). `@deepseek-ai/dsh-compact-basic` (char/4 estimation + token-budget - * retention + `ctx.llm.stream()` summarization) is the first. A tokenizer- or - * template-based backend swaps in without touching consumers. + * context). A tokenizer-, template-, or model-backed implementation can live + * as a sibling package; callers stay on the same `ctx.compact` seam without + * touching consumers. * * The split follows the capability-seams RFC — interface (this) / - * implementation (`dsh-compact-basic`) / consumer (a `/compact` tool, deferred) - * — modeled on the bash trio. Unlike `dsh-bash`, this interface necessarily + * implementation (deferred) / consumer (a `/compact` tool, deferred) — modeled + * on the bash trio. Unlike `dsh-bash`, this interface necessarily * depends on `dsh-session` and `dsh-llm`: the contract's verbs are defined over * a `Session` and its output is the `ContentBlock` vocabulary. That deviation * from the "interface depends only on cordis" guidance is intentional and - * recorded in the [compaction capability-seam RFC](../../../../docs/rfc/proposed/feature/2026-06-18-compaction-capability-seam.md). + * recorded in the [compaction capability-seam RFC](../../../../docs/rfc/implemented/feature/2026-06-18-compaction-capability-seam.md). * * @module @deepseek-ai/dsh-compact */ @@ -27,6 +27,12 @@ import type { CompactionResult } from './types.ts' export type { CompactionResult } from './types.ts' +/** Minimal agent context compaction needs without depending on the agent package. */ +export interface CompactAgentContext { + session: Session + options: { model?: string } +} + declare module 'cordis' { interface Context { compact: CompactService @@ -62,24 +68,44 @@ export abstract class CompactService extends Service { /** * Check token pressure and compact if the conversation is too large. * - * Estimates the current history size (optionally including a system prompt), - * and if it exceeds the backend's threshold, compacts an older range via - * {@link compactRegion}, keeping recent context intact. + * Estimates the current surface-derived history size (including the system + * prompt), and if it exceeds the backend's threshold, compacts an older range + * via {@link compactRegion}, keeping recent context intact. Returns `null` + * when no compaction is needed. * - * @param session - the session whose surface may be compacted. - * @param systemPrompt - optional system prompt, counted toward the estimate. - * @param model - optional summarization model (falls back to backend config). - * @param signal - optional cancellation signal. A backend that summarizes via + * Scope and guarantees a backend MUST honor: + * - **Surface-derived history only.** The decision is made against the history + * derived from the session surface — the only thing compaction can act on. + * Non-surface context injected downstream (into the request `messages` by a + * later listener) is out of this accounting by construction. + * - **Head-anchored, best-effort.** Auto-compaction consolidates from the + * surface HEAD up to a balanced tool-pairing cutoff, so a prior head + * checkpoint is + * re-summarized into one fresh checkpoint (the surface holds at most one + * auto-generated checkpoint, always at the head). It is best-effort over + * CLOSED steps: when the only compactable content left is an un-splittable + * open tail step, it declines (`null`) and retries once that step closes. + * - **Single-unit overflow is out of scope.** If a single retained unit (one + * closed step, or a large free node such as a pasted `user/message`) ALONE + * exceeds the budget, compaction cannot help and the call may go out + * over-budget. Bounding an individual unit's size is a separate concern. + * + * @param agent - agent context owning the session surface and model options. + * @param turn - turn number of the pre-step checkpoint. + * @param step - step number about to start. + * @param fullSystemPrompt - assembled system prompt, counted toward the estimate. + * @param signal - cancellation signal. A backend summarizing via * `ctx.llm.stream()` MUST forward this into the call's `GenerateOptions.signal` * so an abort/dispose tears down the in-flight summarization rather than * leaving an orphaned model call running past the cancellation. * @returns the compaction result, or `null` if no compaction was needed. */ abstract compactIfNeeded( - session: Session, - systemPrompt?: string, - model?: string, - signal?: AbortSignal, + agent: CompactAgentContext, + turn: number, + step: number, + fullSystemPrompt: string, + signal: AbortSignal, ): Promise /** @@ -89,22 +115,40 @@ export abstract class CompactService extends Service { * summarizes their content and appends a replacement surface node. Used by the * (future) `/compact` tool and internally by {@link compactIfNeeded}. * + * The region MUST NOT split a step's `assistant/message` tool-calls from their + * `tool/result`s, leaving the rehydrated transcript with a dangling tool-call + * or an orphaned tool-result that every provider rejects. A region is safe iff + * both its edges are balanced cuts on the surface: the cut before `start` and + * the cut after `end` each have no unanswered tool-call before them. A node + * that belongs to no step (a pre-step user message, inter-step steering, or an + * injection context message) is a balanced (free) boundary; an `end` inside an + * open (unclosed) tail step is invalid — its tool-calls have no results yet. + * `dsh-session` exports `isToolPairingBalanced` for this check. + * * @param session - the session whose surface is mutated. * @param start - inclusive seq of the first surface node to compact. * @param end - inclusive seq of the last surface node to compact. - * @param model - summarization model. + * @param agent - agent context used by router-aware summarizers. + * @param turn - lifecycle turn forwarded to request-routing seams. + * @param step - lifecycle step forwarded to request-routing seams. * @param signal - optional cancellation signal. A backend that summarizes via * `ctx.llm.stream()` MUST forward this into the call's `GenerateOptions.signal` * so an abort/dispose tears down the in-flight summarization rather than * leaving an orphaned model call running past the cancellation. - * @throws if compaction is already in progress, or if `start`/`end` are not - * valid surface nodes, or if `start > end`. + * @throws if compaction is already in progress, if `start`/`end` are not + * valid surface nodes, if `start` is positioned after `end` on the surface + * (the range is a surface-POSITION span, not a numeric seq interval — a + * prior replace can leave the surface non-monotonic in seq order), or if + * either boundary is not a balanced tool-pairing cut (would split a step's + * tool-call/result pair). */ abstract compactRegion( session: Session, start: number, end: number, - model: string, + agent: CompactAgentContext, + turn: number, + step: number, signal?: AbortSignal, ): Promise } diff --git a/packages/compact/compact/src/types.ts b/packages/compact/compact/src/types.ts index 36dd5fb629..ba886685d5 100644 --- a/packages/compact/compact/src/types.ts +++ b/packages/compact/compact/src/types.ts @@ -6,7 +6,7 @@ * events are log-only markers (lock + provenance); only the five * surface-eligible types can carry `surfaceOp`. The actual surface mutation is * performed by a separate `user/message` event carrying the summary (see the - * [compaction capability-seam RFC](../../../../docs/rfc/proposed/feature/2026-06-18-compaction-capability-seam.md)). + * [compaction capability-seam RFC](../../../../docs/rfc/implemented/feature/2026-06-18-compaction-capability-seam.md)). * * Configuration lives in the backend, not here: the contract states WHAT * compaction produces, while every tunable (context window, thresholds, @@ -48,9 +48,16 @@ export interface CompactionResult { endSeq: number /** The summary content blocks produced by the backend. */ summary: ContentBlock[] - /** The seq range that was shadowed [start, end] inclusive. */ + /** + * The surface-boundary pair that was shadowed: the seqs of the first + * (`start`) and last (`end`) surface nodes of the replaced range. A + * surface-POSITION span, not a numeric seq interval — after a prior replace + * lands a fresh high-seq summary node at an older range's position, `start` + * can be GREATER than `end`. {@link CompactionResult.shadowedSeqs} is the + * authoritative set of shadowed nodes, in surface order. + */ shadowedRange: { start: number; end: number } - /** The seq numbers of all shadowed surface nodes. */ + /** The seqs of all shadowed surface nodes, in surface order. */ shadowedSeqs: number[] /** Estimated token count of the shadowed content. */ shadowedTokenCount: number diff --git a/packages/compact/compact/tests/compact.spec.ts b/packages/compact/compact/tests/compact.spec.ts index b3ad9d1501..5b9e033fcc 100644 --- a/packages/compact/compact/tests/compact.spec.ts +++ b/packages/compact/compact/tests/compact.spec.ts @@ -3,6 +3,7 @@ import { Context } from 'cordis' import { CompactService } from '@deepseek-ai/dsh-compact' import type { CompactionResult } from '@deepseek-ai/dsh-compact' import { Session, SessionId } from '@deepseek-ai/dsh-session' +import type { CompactAgentContext } from '@deepseek-ai/dsh-compact' /** * A trivial concrete CompactService implementing the abstract contract. The @@ -15,10 +16,11 @@ class StubCompactService extends CompactService { lastSignal: AbortSignal | undefined override async compactIfNeeded( - _session: Session, - _systemPrompt?: string, - _model?: string, - signal?: AbortSignal, + _agent: CompactAgentContext, + _turn: number, + _step: number, + _fullSystemPrompt: string, + signal: AbortSignal, ): Promise { this.lastSignal = signal return null @@ -28,7 +30,9 @@ class StubCompactService extends CompactService { session: Session, start: number, end: number, - _model: string, + _agent: CompactAgentContext, + _turn: number, + _step: number, signal?: AbortSignal, ): Promise { this.lastSignal = signal @@ -54,6 +58,10 @@ class StubCompactService extends CompactService { } describe('CompactService seam', () => { + function stubAgent(session: Session, model?: string): CompactAgentContext { + return { session, options: model === undefined ? {} : { model } } + } + it('registers as ctx.compact', () => { const ctx = new Context() void new StubCompactService(ctx) @@ -72,7 +80,8 @@ describe('CompactService seam', () => { it('exposes the abstract contract methods', async () => { const ctx = new Context() const svc = new StubCompactService(ctx) - expect(await svc.compactIfNeeded(new Session(SessionId('s')))).toBeNull() + const session = new Session(SessionId('s')) + expect(await svc.compactIfNeeded(stubAgent(session), 1, 1, '', new AbortController().signal)).toBeNull() }) it('compact/* events merge into SessionEventMap and are log-only', async () => { @@ -80,7 +89,7 @@ describe('CompactService seam', () => { const svc = new StubCompactService(ctx) const session = new Session(SessionId('s')) - const result = await svc.compactRegion(session, 0, 0, 'm') + const result = await svc.compactRegion(session, 0, 0, stubAgent(session, 'm'), 1, 1) const startEvent = session.events.find(e => e.type === 'compact/start') expect(startEvent).toBeDefined() @@ -98,10 +107,10 @@ describe('CompactService seam', () => { const session = new Session(SessionId('s')) const controller = new AbortController() - await svc.compactRegion(session, 0, 0, 'm', controller.signal) + await svc.compactRegion(session, 0, 0, stubAgent(session, 'm'), 1, 1, controller.signal) expect(svc.lastSignal).toBe(controller.signal) - await svc.compactIfNeeded(session, undefined, undefined, controller.signal) + await svc.compactIfNeeded(stubAgent(session), 1, 1, '', controller.signal) expect(svc.lastSignal).toBe(controller.signal) }) }) diff --git a/packages/core/README.md b/packages/core/README.md index 8d8805471a..a3e93777ab 100644 --- a/packages/core/README.md +++ b/packages/core/README.md @@ -13,4 +13,4 @@ The packages every harness build is assembled from: the session log, the system- `agent-loop` is the one concrete implementation of the `agent` seam and lives here because it is the harness's default product loop; everything else in `core/` is interface/vocabulary. Plugins depend on the `agent` vocabulary, never on `agent-loop` directly, so the loop stays swappable. -`agent-core` is the composition counterpart: one bundle plugin that loads the whole providerless spine (`timer` + `llm` + sessions + system-prompt + tools + agents + invariants + `tool-bash` + `agent-loop`) and forwards `agent-loop`'s `agents` list as its own config. App packages (`ui/stdio-agent`, `ui/acp-agent`) consume it and add only a front door; a leaf adds only the swappable backends. It lives in `core/` because it composes exclusively `core/` + interface packages and ships no provider, executor, or UI of its own. +`agent-core` is the composition counterpart: one bundle plugin that loads the whole providerless spine (`timer` + `llm` + sessions + system-prompt + tools + agents + invariants + `tool-bash` + `agent-loop`) and forwards `agent-loop`'s `agents` list as its own config. App packages (`ui/stdio-agent`, `ui/acp-agent`) consume it and add only a front door; a leaf adds the swappable backends plus any optional product tools it wants to expose. It lives in `core/` because it composes exclusively `core/` + interface packages and ships no provider, executor, or UI of its own. diff --git a/packages/core/agent-loop/README.md b/packages/core/agent-loop/README.md index 4892b357bf..5932bc6741 100644 --- a/packages/core/agent-loop/README.md +++ b/packages/core/agent-loop/README.md @@ -12,10 +12,10 @@ This is the only package in the harness that contains concrete loop logic. Every `AgentLoop` also implements the `AgentFactory` seam and registers itself via `ctx.agents.setFactory(this)`, so plugins create/resume agents through `ctx.agents` (the interface): -- `ctx.agents.create({ agentId, sessionId, meta?, agentOptions? }): AgentHandle` — programmatic create on a caller-supplied `sessionId` (e.g. an ACP-generated id), NOT `${id}-session`. Returns an [`AgentHandle`](../agent/README.md) — the owner disposes it to tear down exactly this agent (stop loop + await quiescence + unregister + remove session). +- `ctx.agents.create({ agentId, sessionId, meta?, seed?, agentOptions? }): AgentHandle` — programmatic create on a caller-supplied `sessionId` (e.g. an ACP-generated id), NOT `${id}-session`; `meta` carries cwd/lineage/seed-boundary metadata and `seed` reconstructs a forked child prefix. Returns an [`AgentHandle`](../agent/README.md) — the owner disposes it to tear down exactly this agent (stop loop + await quiescence + unregister + remove session). - `ctx.agents.resume({ agentId, resumeSessionId, agentOptions? }): Promise` — load a persisted session via `ctx.sessionPersistence` ([session persistence](../../../docs/rfc/implemented/architecture/2026-06-14-session-persistence.md)) and resume an agent on it. The live session id is the resumed id; turn numbering and derived history continue from the loaded log. Requires a session-persistence backend (NOT hard-injected — non-persistent demos still work; `resume` rejects with a clear error when persistence is absent). Returns an `AgentHandle`. -The config-driven `ctx.agentLoop.create()` path keeps its agent owned by the loop fiber (it discards the handle) — only the programmatic factory callers (the ACP bridge) hold a handle and own per-agent teardown. +The config-driven `ctx.agentLoop.create()` path keeps its agent owned by the loop fiber (it discards the handle) — only the programmatic factory callers (the ACP bridge and in-process subagent backends) hold a handle and own per-agent teardown. ### Injected services @@ -52,6 +52,8 @@ forever: STEP loop: drain steering assembly = systemPrompt.assemble() + await serial agent/pre-step ⟵ surface mutation (compaction) outside the step + session('step/start') request = waterfall agent/request stream llm.stream(request) → session('assistant/chunk') message = waterfall agent/step-result @@ -73,9 +75,9 @@ Cancellation: `agent.cancel()` is the single public stop primitive — it clears ### What is NOT here Everything that goes beyond "call the model, run the tools, repeat" belongs to plugins listening on the event taxonomy: -- Hooks: `agent/request`, `agent/step-result`, `tools/execute`, `agent/turn-continuation` -- Compaction: `agent/request` +- Hooks: `agent/pre-step`, `agent/request`, `agent/step-result`, `tools/execute`, `agent/turn-continuation` +- Compaction: `agent/pre-step` - Sandbox, permission, plan mode: `tools/execute` -- Sub-agents: TODO seam on `AgentLoop.create()` +- Sub-agents: implemented outside the loop as `ctx.subagents` providers; in-process providers use `ctx.agents.create()` and owned `AgentHandle` teardown, while child streaming/progress and background/poll collection remain deferred. - Persistence: `session/event` + `session/flush` - UI: `agent/stream-chunk` + `agent/*` events diff --git a/packages/core/agent-loop/src/loop.ts b/packages/core/agent-loop/src/loop.ts index ceef1bca8e..200115af9a 100644 --- a/packages/core/agent-loop/src/loop.ts +++ b/packages/core/agent-loop/src/loop.ts @@ -12,6 +12,7 @@ import type { FinishReason, GenerateOptions, Message } from '@deepseek-ai/dsh-ll import { BlockAssembler, HarnessError } from '@deepseek-ai/dsh-llm' import type { Session, TurnEndReason, TurnTrigger } from '@deepseek-ai/dsh-session' import { renderPrompt } from '@deepseek-ai/dsh-system-prompt' +import type { PromptAssembly } from '@deepseek-ai/dsh-system-prompt' import type {} from '@deepseek-ai/dsh-tools' import type { ReactLoopAgent } from './agent.ts' @@ -147,10 +148,11 @@ export interface LoopHandle { * drain queued → 'turn/start' → session('user/message'…) → emit agent/turn-start * STEP loop: * drain steering → session('steering/message') ⟵ catches late steering - * session('step/start'); emit agent/step-start ⟵ append before emit (the event-sourcing RFC) * assembly = ctx.systemPrompt.assemble() ⟵ waterfall system-prompt/assemble + * await ctx.serial('agent/pre-step') ⟵ surface mutation (compaction) OUTSIDE the step + * session('step/start'); emit agent/step-start ⟵ append before emit (the event-sourcing RFC) * req = {model, system, tools, messages: session.deriveMessages(), signal} - * req = waterfall agent/request ⟵ hooks/compaction/model-switch + * req = waterfall agent/request ⟵ hooks/model-switch * stream ctx.llm.stream(req) ⟵ waterfall llm/stream (raw chunks) * session('assistant/chunk'); emit agent/stream-chunk * msg = waterfall agent/step-result ⟵ BEFORE the log append, so the @@ -386,30 +388,78 @@ async function runTurn(ctx: Context, agent: ReactLoopAgent, handle: LoopHandle, // (or turn-start listeners on the first step) joins before the request. drainSteering(ctx, agent, turn) + // The step's AbortController exists BEFORE any async pre-step work so a + // dispose() or cancel() — in a synchronous turn-start listener or an + // async listener whose effect fires before we block — always has an armed + // abort to cancel against. isDisposed below covers disposal, which does + // NOT set the cancel marker. Cleared on every exit path below. + const abort = new AbortController() + handle.setAbort(abort) + + // Assemble the system prompt for this step. Done HERE (before step/start) + // because the pre-step seam needs it: compaction measures token pressure + // against the system prompt (it counts toward the budget). runStep reuses + // this same assembly for the request, so the prompt is assembled once per + // step. + const assembly = await ctx.systemPrompt.assemble() + const fullSystemPrompt = [renderPrompt(assembly), agent.options.systemPrompt ?? ''] + .filter(text => text.length > 0) + .join('\n\n') + + // Interruption landing after assembly: dispose() or cancel() in a + // turn-start listener (or a listener whose promise resolved before the + // await above) arms either handle.isDisposed() or handle.isCancelled(). + // The Abort was created first, so any concurrent abort also lands on it. + // Drop the about-to-start step WITHOUT running the seam — no step is open + // yet, so end the turn accordingly (disposed wins for an unambiguous + // reason). + if (handle.isCancelled() || handle.isDisposed()) { + handle.setAbort(undefined) + reason = handle.isDisposed() ? { kind: 'disposed' } : { kind: 'aborted', reason: handle.cancelReason() } + break + } + + // Pre-step surface-mutation checkpoint (compaction), fired OUTSIDE the + // step: after `turn/start` (and the prior step's close) but before + // `step/start`, so a compaction's log-only `compact/*` records and its + // replacement node land cleanly outside any step (honest structure that + // crash-safety relies on — a dangling `compact/start` sits before the + // synthetic `turn/end` repair appends). Serial (awaited, in order, no + // veto): each listener completes its surface mutation before the next, so + // concurrent listeners cannot interleave their `session.append`s. A + // throwing listener escapes to the outer catch, which closes the (not-yet- + // open) step as a no-op and ends the turn via failTurn — a broken + // pre-step plugin ends the turn, not the loop. + await ctx.serial('agent/pre-step', agent, turn, step, fullSystemPrompt, abort.signal) + + // Interruption landing during the pre-step seam: do not open an empty + // step. `agent/step-start` listeners get their own check below because + // they necessarily run after step/start is appended/emitted. + if (handle.isCancelled() || handle.isDisposed()) { + handle.setAbort(undefined) + reason = handle.isDisposed() ? { kind: 'disposed' } : { kind: 'aborted', reason: handle.cancelReason() } + break + } + session.append('step/start', { turn, step }) stepOpen = true ctx.emit('agent/step-start', agent, turn, step) - const abort = new AbortController() - handle.setAbort(abort) - - // Cancel landing in the step-start window: a synchronous `agent/turn-start` - // or `agent/step-start` listener (both fire before this point) can have - // called `cancel()`, and `runStep` would otherwise run a full extra step - // with no AbortController having observed it. Check the marker AFTER - // setAbort (so the next-iteration drain sees a clean controller) and before - // `runStep`: drop the step, end the turn `aborted`. closeStep balances the - // already-appended step/start. - if (handle.isCancelled()) { + // Cancel landing in the step-start window: a synchronous + // `agent/step-start` listener can cancel after the step is already open. + // Check AFTER step/start append + emit and before `runStep`: drop the + // step, end the turn accordingly. closeStep balances the already-appended + // step/start. + if (handle.isCancelled() || handle.isDisposed()) { handle.setAbort(undefined) - reason = { kind: 'aborted', reason: handle.cancelReason() } + reason = handle.isDisposed() ? { kind: 'disposed' } : { kind: 'aborted', reason: handle.cancelReason() } closeStep() break } let stepOutcome: { hadToolCalls: boolean; finish: FinishReason } | { error: Error } try { - stepOutcome = await runStep(ctx, agent, turn, step, abort.signal) + stepOutcome = await runStep(ctx, agent, turn, step, assembly, fullSystemPrompt, abort.signal) } catch (error: unknown) { stepOutcome = { error: toError(error) } } finally { @@ -549,22 +599,22 @@ function drainSteering(ctx: Context, agent: ReactLoopAgent, turn: number): boole return messages.length > 0 } -/** One step: assemble request → stream model → record → execute tools. */ +/** One step: derive request from the (already pre-step-mutated) surface → + * stream model → record → execute tools. The caller assembles the system prompt + * and fires the `agent/pre-step` seam BEFORE opening the step, then passes the + * resulting `assembly`/`system` here, so the surface this step derives from + * already reflects any compaction. */ async function runStep( ctx: Context, agent: ReactLoopAgent, turn: number, step: number, + assembly: PromptAssembly, + system: string, signal: AbortSignal, ): Promise<{ hadToolCalls: boolean; finish: FinishReason }> { const { session, options } = agent - // --- Request assembly --- - const assembly = await ctx.systemPrompt.assemble() - const system = [renderPrompt(assembly), options.systemPrompt ?? ''] - .filter(text => text.length > 0) - .join('\n\n') - let request: GenerateOptions = { model: options.model ?? '', messages: session.deriveMessages(), diff --git a/packages/core/agent-loop/tests/cancel.spec.ts b/packages/core/agent-loop/tests/cancel.spec.ts index 9cdaa1973b..32c46f72f0 100644 --- a/packages/core/agent-loop/tests/cancel.spec.ts +++ b/packages/core/agent-loop/tests/cancel.spec.ts @@ -13,7 +13,7 @@ import { describe, expect, it } from 'vitest' import { Context } from 'cordis' import LlmService from '@deepseek-ai/dsh-llm' -import SessionStore, { TurnEndReason } from '@deepseek-ai/dsh-session' +import SessionStore, { SessionId, TurnEndReason } from '@deepseek-ai/dsh-session' import SystemPrompt from '@deepseek-ai/dsh-system-prompt' import ToolRegistry from '@deepseek-ai/dsh-tools' import AgentRegistry, { AgentId } from '@deepseek-ai/dsh-agent' @@ -194,6 +194,73 @@ describe('Agent.cancel()', () => { expect(reasons).toEqual([{ kind: 'aborted', reason: 'from turn-start' }]) }) + it('cancel from a synchronous agent/step-start listener drops the step (post-step-start window)', async () => { + const adapter = new MockAdapter([textResponse('should not stream')]) + const ctx = await harness(adapter) + const agent = ctx.agentLoop.create(AgentId('a1'), { model: 'mock' }) + + // A step-start listener fires AFTER step/start is appended (and after the + // pre-step seam), so cancelling there lands in the SECOND cancel check (the + // one that must closeStep() to balance the already-open step) — distinct + // from a turn-start cancel, which is caught before the step opens. + let streamed = false + ctx.on('agent/stream-chunk', () => { streamed = true }) + const dispose = ctx.on('agent/step-start', (subject) => { + if (subject === agent) agent.cancel('from step-start') + }) + + const reasons: TurnEndReason[] = [] + ctx.on('agent/turn-end', (_a, _t, reason) => void reasons.push(reason)) + + send(agent, 'go') + await waitForIdle(ctx, agent) + dispose() + + // No step streamed, the turn ended aborted with the caller's reason, and the + // log is balanced (the open step was closed by the cancel branch). + expect(streamed).toBe(false) + expect(reasons).toEqual([{ kind: 'aborted', reason: 'from step-start' }]) + const types = agent.session.events.map(e => e.type) + expect(types.filter(t => t === 'step/start').length).toBe(types.filter(t => t === 'step/end').length) + }) + + it('disposal from a synchronous agent/step-start listener closes the open step as disposed', async () => { + const adapter = new MockAdapter([textResponse('should not stream')]) + const ctx = new Context() + await ctx.plugin(LlmService) + await ctx.plugin(SessionStore) + await ctx.plugin(SystemPrompt) + await ctx.plugin(ToolRegistry) + await ctx.plugin(AgentRegistry) + await ctx.plugin(AgentLoop, { agents: [] }) + ctx.llm.registerAdapter(['mock'], adapter) + + const handle = ctx.agents.create({ + agentId: AgentId('a-dispose-step-start'), + sessionId: SessionId('dispose-step-start-session'), + agentOptions: { model: 'mock' }, + }) + const agent = handle.agent as ReactLoopAgent + + let disposalDone: Promise | undefined + let streamed = false + ctx.on('agent/stream-chunk', () => { streamed = true }) + ctx.on('agent/step-start', (subject) => { + if (subject === agent) disposalDone = handle.dispose() + }) + + send(agent, 'go') + await disposalDone + await agent.done + + expect(streamed).toBe(false) + expect(adapter.requests).toHaveLength(0) + const turnEnd = agent.session.events.findLast(e => e.type === 'turn/end') + expect(turnEnd?.type === 'turn/end' && turnEnd.data.reason).toEqual({ kind: 'disposed' }) + const types = agent.session.events.map(e => e.type) + expect(types.filter(t => t === 'step/start').length).toBe(types.filter(t => t === 'step/end').length) + }) + it('cancel during the continuation window ends the turn aborted and runs no further step', async () => { // A continuation-waterfall listener cancels DURING the continuation decision // (the finished step's AbortController is already cleared), and votes to diff --git a/packages/core/agent-loop/tests/loop.spec.ts b/packages/core/agent-loop/tests/loop.spec.ts index d018eff7a2..aa565f15b5 100644 --- a/packages/core/agent-loop/tests/loop.spec.ts +++ b/packages/core/agent-loop/tests/loop.spec.ts @@ -320,6 +320,110 @@ describe('agent loop', () => { expect(adapter.requests[0]!.model).toBe('other-model') }) + it('agent/pre-step fires once per step before the step is opened', async () => { + // Two steps (a tool call, then a final text turn) → two model calls → two + // pre-step fires, each carrying the assembled full system prompt, BEFORE + // the step is opened and its request is derived (the request the adapter + // sees reflects any surface state at fire time). + const adapter = new MockAdapter([ + toolCallResponse('c1', 'echo', {}, 'calling echo'), + textResponse('done'), + ]) + const ctx = await harness(adapter) + ctx.tools.register(defineTool({ + name: 'echo', description: 'echo', parameters: {}, + async execute() { return [{ type: 'text', text: 'echoed' }] }, + })) + const agent = ctx.agentLoop.create(AgentId('a1'), { model: 'mock' }) + + const fires: { turn: number; step: number; fullSystemPrompt: string }[] = [] + ctx.on('agent/pre-step', (subject, turn, step, fullSystemPrompt) => { + if (subject === agent) fires.push({ turn, step, fullSystemPrompt }) + }) + + send(agent, 'go') + await waitForIdle(ctx, agent) + + // One fire per step, in order, each with the assembled system prompt. + expect(fires).toEqual([ + { turn: 1, step: 1, fullSystemPrompt: '' }, + { turn: 1, step: 2, fullSystemPrompt: '' }, + ]) + }) + + it('agent/pre-step fires BEFORE the step it precedes opens (events land outside the step)', async () => { + // A listener appending a surface node in pre-step lands it BEFORE step/start + // in the log — proving the seam fires outside the step. The node is still in + // the derived request for that step (derive happens after step/start). + const adapter = new MockAdapter([textResponse('ok')]) + const ctx = await harness(adapter) + const agent = ctx.agentLoop.create(AgentId('a1'), { model: 'mock' }) + + let injected = false + ctx.on('agent/pre-step', (subject) => { + if (subject === agent && !injected) { + injected = true + subject.session.append('context/message', { + content: [{ type: 'text', text: 'INJECTED-IN-PRE-STEP' }], + source: { kind: 'plugin', plugin: 'test' }, + }, { surfaceOp: 'append' }) + } + }) + + send(agent, 'go') + await waitForIdle(ctx, agent) + + // The adapter's request includes the node injected during pre-step (derive + // reflects it). + const text = JSON.stringify(adapter.requests[0]!.messages) + expect(text).toContain('INJECTED-IN-PRE-STEP') + + // And the injected event sits BEFORE the first step/start in the log — + // the seam fired outside the step. + const events = agent.session.events + const injectedSeq = events.find(e => e.type === 'context/message')!.seq + const firstStepStartSeq = events.find(e => e.type === 'step/start')!.seq + expect(injectedSeq).toBeLessThan(firstStepStartSeq) + }) + + it('a throwing agent/pre-step listener ends the turn (error), not the loop', async () => { + // The seam fires before step/start, so a throw escapes to runTurn's outer + // catch: the not-yet-open step closes as a no-op, the failure surfaces via + // agent/error, and the turn ends `error` (recorded on the durable turn/end). + // The loop survives and a follow-up prompt still runs. + const adapter = new MockAdapter([textResponse('second turn ok')]) + const ctx = await harness(adapter) + const agent = ctx.agentLoop.create(AgentId('a1'), { model: 'mock' }) + + let throwOnce = true + ctx.on('agent/pre-step', () => { + if (throwOnce) { throwOnce = false; throw new Error('boom in pre-step') } + }) + + const errors: Error[] = [] + ctx.on('agent/error', (_a, _t, _s, error) => void errors.push(error)) + + send(agent, 'first') + await waitForIdle(ctx, agent) + // The first turn failed at step 1 (no model call happened), surfaced via + // agent/error, with the durable failure on turn/end.reason. + expect(errors).toHaveLength(1) + expect(errors[0]!.message).toContain('boom in pre-step') + expect(adapter.requests.length).toBe(0) + const firstTurnEnd = agent.session.events.find(e => e.type === 'turn/end') + expect(firstTurnEnd?.type === 'turn/end' && firstTurnEnd.data.reason).toMatchObject({ kind: 'error', step: 1 }) + // The step opened-and-closed count stays balanced even though it never ran. + const types = agent.session.events.map(e => e.type) + expect(types.filter(t => t === 'step/start').length).toBe(types.filter(t => t === 'step/end').length) + + // The loop survived: a second prompt runs a normal completed turn. + send(agent, 'second') + await waitForIdle(ctx, agent) + expect(adapter.requests.length).toBe(1) + const lastTurnEnd = agent.session.events.findLast(e => e.type === 'turn/end') + expect(lastTurnEnd?.type === 'turn/end' && lastTurnEnd.data.reason).toEqual({ kind: 'completed' }) + }) + it('cancel() mid-stream ends the turn with reason aborted', async () => { const adapter = new MockAdapter(['hang']) const ctx = await harness(adapter) diff --git a/packages/core/agent-loop/tests/review-fixes.spec.ts b/packages/core/agent-loop/tests/review-fixes.spec.ts index a092bd8419..346d70e2f7 100644 --- a/packages/core/agent-loop/tests/review-fixes.spec.ts +++ b/packages/core/agent-loop/tests/review-fixes.spec.ts @@ -1047,3 +1047,275 @@ describe('surface: assistant/message omits sourceEventSeqs when no chunks stream expect(JSON.stringify(agent.session.deriveMessages())).toContain('injected') }) }) + + + +describe('disposal/cancel honored during pre-step assembly (P1-1)', () => { + it('disposal during system-prompt assembly drops the about-to-start step as disposed', { timeout: 30000 }, async () => { + // Block `system-prompt/assemble` on a promise. Start disposal (which + // calls stop() synchronously, setting status=disposed), then release the + // block. The loop must check isDisposed() after assembly and end the turn + // `disposed` — no LLM call. Don't await fiber.dispose() before releasing + // the blocker: the dispose chain awaits agent.done, which hangs until the + // loop unblocks. + const adapter = new MockAdapter(['hang']) + let releaseAssemble!: () => void + const blocked = new Promise(r => void (releaseAssemble = r)) + + const ctx = new Context() + await ctx.plugin(LlmService) + await ctx.plugin(SessionStore) + await ctx.plugin(SystemPrompt) + await ctx.plugin(ToolRegistry) + await ctx.plugin(AgentRegistry) + await ctx.plugin(AgentLoop, { agents: [] }) + await ctx.plugin(Invariants, { freeze: false }) + ctx.llm.registerAdapter(['mock'], adapter) + + // Blocking listener on the parent context (survives fiber disposal). + const unlisten = ctx.on('system-prompt/assemble', async function (_assembly, next) { + await blocked + return next() + }) + + let agent!: ReactLoopAgent + const fiber = await ctx.plugin(Object.assign((inner: Context) => { + agent = inner.agentLoop.create(AgentId('a-dispose-assemble'), { model: 'mock' }) + }, { inject: ['agentLoop'] })) + + const reasons: TurnEndReason[] = [] + ctx.on('agent/turn-end', (_a, _t, reason) => void reasons.push(reason)) + + send(agent, 'go') + // Give the loop time to enter the step and reach assemble(). + await new Promise(r => setTimeout(r, 50)) + + // Start disposal — stop() sets status=disposed synchronously, then the + // disposer's await agent.done hangs because the loop is blocked in the + // waterfall. Do NOT await yet; release the blocker first. + const disposalDone = fiber.dispose() + + // Now release the blocked waterfall — the loop unblocks, checks + // isDisposed(), and exits, which resolves agent.done and disposalDone. + releaseAssemble() + await disposalDone + await agent.done + unlisten() + + const e = [...agent.session.events] + expect(e.filter(x => x.type === 'turn/start')).toHaveLength(1) + expect(e.filter(x => x.type === 'turn/end')).toHaveLength(1) + const turnEnd = e.findLast(x => x.type === 'turn/end') + expect(turnEnd?.type === 'turn/end' && turnEnd.data.reason).toEqual({ kind: 'disposed' }) + // No step was opened, no LLM call was made. + expect(e.some(x => x.type === 'step/start')).toBe(false) + expect(e.some(x => x.type === 'assistant/chunk')).toBe(false) + // agent/turn-end may not fire when disposal happens during assembly: the + // fiber's disposer (stop→status=disposed) runs before closeTurn(true)'s + // emit, and the LIFO chain disposes effects in reverse registration order. + // The turn/end durable record is the one that matters. + }) + + it('cancel during system-prompt assembly drops the about-to-start step as aborted', { timeout: 30000 }, async () => { + const adapter = new MockAdapter([textResponse('should not appear')]) + let releaseAssemble!: () => void + const blocker = new Promise(r => void (releaseAssemble = r)) + + const ctx = new Context() + await ctx.plugin(LlmService) + await ctx.plugin(SessionStore) + await ctx.plugin(SystemPrompt) + await ctx.plugin(ToolRegistry) + await ctx.plugin(AgentRegistry) + await ctx.plugin(AgentLoop, { agents: [] }) + await ctx.plugin(Invariants, { freeze: false }) + ctx.llm.registerAdapter(['mock'], adapter) + + const unlisten = ctx.on('system-prompt/assemble', async function (_assembly, next) { + await blocker + return next() + }) + + let agent!: ReactLoopAgent + const fiber = await ctx.plugin(Object.assign((inner: Context) => { + agent = inner.agentLoop.create(AgentId('a-cancel-assemble'), { model: 'mock' }) + }, { inject: ['agentLoop'] })) + + const reasons: TurnEndReason[] = [] + ctx.on('agent/turn-end', (_a, _t, reason) => void reasons.push(reason)) + + send(agent, 'go') + await new Promise(r => setTimeout(r, 50)) + agent.cancel('user cancelled during assembly') + + releaseAssemble() + await waitForIdle(ctx, agent) + await fiber.dispose() + await agent.done + unlisten() + + const e = [...agent.session.events] + expect(e.filter(x => x.type === 'turn/start')).toHaveLength(1) + expect(e.filter(x => x.type === 'turn/end')).toHaveLength(1) + const turnEnd = e.findLast(x => x.type === 'turn/end') + expect(turnEnd?.type === 'turn/end' && turnEnd.data.reason).toEqual({ + kind: 'aborted', + reason: 'user cancelled during assembly', + }) + expect(e.some(x => x.type === 'step/start')).toBe(false) + expect(e.some(x => x.type === 'assistant/chunk')).toBe(false) + expect(e.some(x => x.type === 'assistant/message')).toBe(false) + expect(adapter.requests).toHaveLength(0) + expect(reasons).toEqual([{ kind: 'aborted', reason: 'user cancelled during assembly' }]) + }) + + it('disposal during agent/pre-step seam ends the turn disposed', { timeout: 15000 }, async () => { + // Block the `agent/pre-step` serial seam on a promise we control, then + // dispose the agent's fiber. When the block releases, the loop must see + // isDisposed() at the post-seam check and end the turn disposed. + const adapter = new MockAdapter(['hang']) + let releasePreStep!: () => void + const blocker = new Promise(r => void (releasePreStep = r)) + + const ctx = new Context() + await ctx.plugin(LlmService) + await ctx.plugin(SessionStore) + await ctx.plugin(SystemPrompt) + await ctx.plugin(ToolRegistry) + await ctx.plugin(AgentRegistry) + await ctx.plugin(AgentLoop, { agents: [] }) + await ctx.plugin(Invariants, { freeze: false }) + ctx.llm.registerAdapter(['mock'], adapter) + + ctx.on('agent/pre-step', async () => { + await blocker + }) + + let agent!: ReactLoopAgent + const fiber = await ctx.plugin(Object.assign((inner: Context) => { + agent = inner.agentLoop.create(AgentId('a-dispose-prestep'), { model: 'mock' }) + }, { inject: ['agentLoop'] })) + + const reasons: TurnEndReason[] = [] + ctx.on('agent/turn-end', (_a, _t, reason) => void reasons.push(reason)) + + send(agent, 'go') + await new Promise(r => setTimeout(r, 50)) + + // Start disposal, then release the block, then await disposal. + const disposalDone = fiber.dispose() + releasePreStep() + await disposalDone + await agent.done + + // After the pre-step seam finishes, the post-seam cancel/dispose check + // catches disposal. The step was never opened, no LLM call was made. + const e = [...agent.session.events] + expect(e.filter(x => x.type === 'turn/start')).toHaveLength(1) + expect(e.filter(x => x.type === 'turn/end')).toHaveLength(1) + const turnEnd = e.findLast(x => x.type === 'turn/end') + // Disposal wins the post-seam check — reason is `disposed`. + expect(turnEnd?.type === 'turn/end' && turnEnd.data.reason).toEqual({ kind: 'disposed' }) + expect(e.some(x => x.type === 'step/start')).toBe(false) + expect(e.some(x => x.type === 'assistant/chunk')).toBe(false) + // agent/turn-end may not fire when disposal happens during pre-step: the + // fiber's disposer runs before closeTurn(true)'s emit. The durable turn/end + // is the authoritative record. + }) + + it('cancel during agent/pre-step seam ends the turn aborted', { timeout: 15000 }, async () => { + // Block `agent/pre-step`, then cancel() the agent. When the block releases, + // the post-seam check catches cancellation and ends the turn aborted. + const adapter = new MockAdapter(['hang']) + let releasePreStep!: () => void + const blocker = new Promise(r => void (releasePreStep = r)) + + const ctx = new Context() + await ctx.plugin(LlmService) + await ctx.plugin(SessionStore) + await ctx.plugin(SystemPrompt) + await ctx.plugin(ToolRegistry) + await ctx.plugin(AgentRegistry) + await ctx.plugin(AgentLoop, { agents: [] }) + await ctx.plugin(Invariants, { freeze: false }) + ctx.llm.registerAdapter(['mock'], adapter) + + ctx.on('agent/pre-step', async () => { + await blocker + }) + + let agent!: ReactLoopAgent + const fiber = await ctx.plugin(Object.assign((inner: Context) => { + agent = inner.agentLoop.create(AgentId('a-cancel-prestep'), { model: 'mock' }) + }, { inject: ['agentLoop'] })) + + const reasons: TurnEndReason[] = [] + ctx.on('agent/turn-end', (_a, _t, reason) => void reasons.push(reason)) + + send(agent, 'go') + await new Promise(r => setTimeout(r, 30)) + agent.cancel('user cancelled') + + releasePreStep() + await waitForIdle(ctx, agent) + await fiber.dispose() + await agent.done + + const e = [...agent.session.events] + expect(e.filter(x => x.type === 'turn/start')).toHaveLength(1) + expect(e.filter(x => x.type === 'turn/end')).toHaveLength(1) + const turnEnd = e.findLast(x => x.type === 'turn/end') + expect(turnEnd?.type === 'turn/end' && turnEnd.data.reason).toEqual({ kind: 'aborted', reason: 'user cancelled' }) + expect(e.some(x => x.type === 'step/start')).toBe(false) + expect(e.some(x => x.type === 'assistant/chunk')).toBe(false) + expect(reasons).toEqual([{ kind: 'aborted', reason: 'user cancelled' }]) + }) + + it('disposal during assembly does not leak an LLM call or append assistant/chunk', { timeout: 15000 }, async () => { + // The key assertion from the original bug report: after disposal, no + // assistant/chunk or assistant/message appears — the turn ends disposed + // before any model interaction. + const adapter = new MockAdapter([textResponse('should not appear')]) + let releaseAssemble!: () => void + const blocker = new Promise(r => void (releaseAssemble = r)) + + const ctx = new Context() + await ctx.plugin(LlmService) + await ctx.plugin(SessionStore) + await ctx.plugin(SystemPrompt) + await ctx.plugin(ToolRegistry) + await ctx.plugin(AgentRegistry) + await ctx.plugin(AgentLoop, { agents: [] }) + await ctx.plugin(Invariants, { freeze: false }) + ctx.llm.registerAdapter(['mock'], adapter) + + ctx.on('system-prompt/assemble', async function (_assembly, next) { + await blocker + return next() + }) + + let agent!: ReactLoopAgent + const fiber = await ctx.plugin(Object.assign((inner: Context) => { + agent = inner.agentLoop.create(AgentId('a-dispose-no-leak'), { model: 'mock' }) + }, { inject: ['agentLoop'] })) + + send(agent, 'go') + await new Promise(r => setTimeout(r, 50)) + + const disposalDone = fiber.dispose() + releaseAssemble() + await disposalDone + await agent.done + + const e = [...agent.session.events] + expect(e.filter(x => x.type === 'turn/start')).toHaveLength(1) + expect(e.filter(x => x.type === 'turn/end')).toHaveLength(1) + // The critical assertions: after disposal, the turn has no assistant + // artifacts — the turn ended disposed before the model was invoked. + expect(e.some(x => x.type === 'assistant/chunk')).toBe(false) + expect(e.some(x => x.type === 'assistant/message')).toBe(false) + expect(adapter.requests).toHaveLength(0) + // The durable turn/end reason is the authoritative record; agent/turn-end + // may not fire when disposal interleaves with closeTurn(true)'s emit. + }) +}) diff --git a/packages/core/agent/README.md b/packages/core/agent/README.md index d0ec0ee614..6a95d27572 100644 --- a/packages/core/agent/README.md +++ b/packages/core/agent/README.md @@ -17,10 +17,10 @@ Tracks live agents so UI, hook, and orchestrator plugins can find them without i Agent *creation* is provided by whichever plugin implements `AgentFactory` (phase 1: `dsh-agent-loop`), registered via `setFactory`. This keeps creation on the `dsh-agent` interface so consumers (UI, the ACP bridge) program against `ctx.agents` without depending on the concrete loop package. - `ctx.agents.setFactory(factory: AgentFactory): () => void` — register the creation factory (the loop calls this on construction). Throws on a second factory; the slot clears on dispose. -- `ctx.agents.create(options: CreateAgentOptions): AgentHandle` — construct, start, AND register a new agent on a caller-supplied `sessionId` (with optional `meta.cwd`). Distinct from `register` (which only records). Throws if no factory is registered. +- `ctx.agents.create(options: CreateAgentOptions): AgentHandle` — construct, start, AND register a new agent on a caller-supplied `sessionId` (with optional `meta.cwd`/`meta.parentSession`/`meta.seedLength` and optional `seed` events for forked children). Distinct from `register` (which only records). Throws if no factory is registered. - `ctx.agents.resume(options: ResumeAgentOptions): Promise` — load a persisted session ([session persistence](../../../docs/rfc/implemented/architecture/2026-06-14-session-persistence.md)) and resume an agent on it. Async; rejects if no factory is registered, or if the factory finds session persistence unconfigured. -`AgentHandle = { agent: Agent; dispose(): Promise }`. The disposer is a **capability** — only the holder can tear this agent down. `dispose()` stops the loop, `await`s its exit (quiescence — NOT just the `disposed` status flip), unregisters the agent, and removes its session from the store, in an order that captures the loop's final `session/flush` before the session is detached. `ctx.agents.get(id)` still returns a bare `Agent` — the handle is only for the OWNER that created it. The ACP bridge is the production consumer (one handle per session, disposed on disconnect/teardown); config-created agents are owned by the loop fiber and never need a handle. +`AgentHandle = { agent: Agent; dispose(): Promise }`. The disposer is a **capability** — only the holder can tear this agent down. `dispose()` stops the loop, `await`s its exit (quiescence — NOT just the `disposed` status flip), unregisters the agent, and removes its session from the store, in an order that captures the loop's final `session/flush` before the session is detached. `ctx.agents.get(id)` still returns a bare `Agent` — the handle is only for the OWNER that created it. The ACP bridge and in-process subagent backends are production consumers; config-created agents are owned by the loop fiber and never need a handle. ### Events @@ -37,11 +37,12 @@ The full `agent/*` event taxonomy is declared via declaration merging in `dsh-ag - `agent/turn-start`, `agent/turn-end` (carries `TurnEndReason`) - `agent/step-start`, `agent/step-end` -#### Interception seams (waterfall) +#### Interception seams -- `agent/request` — mutate `GenerateOptions` before the model call (hooks, compaction, model switching, tool filtering) -- `agent/step-result` — post-process the assembled assistant message before tool dispatch (validates what the log records) -- `agent/turn-continuation` — override the continue/stop decision (force-continue /loop, force-stop budget guard) +- `agent/pre-step` (serial) — mutate the session surface before the step opens and history is derived (compaction). Fires after `turn/start` and before `step/start`, so a listener's appended events land outside the step. +- `agent/request` (waterfall) — mutate `GenerateOptions` before the model call (hooks, model switching, tool filtering) +- `agent/step-result` (waterfall) — post-process the assembled assistant message before tool dispatch (validates what the log records) +- `agent/turn-continuation` (waterfall) — override the continue/stop decision (force-continue /loop, force-stop budget guard) #### Streaming + tool (emit) @@ -62,9 +63,10 @@ The handle every plugin programs against: ### Extension points -- Agent creation: `AgentLoop.create()` is the concrete implementation (in `dsh-agent-loop`). Replace the loop by implementing `Agent` and registering via `ctx.agents.register()`. +- Agent creation: `AgentLoop.create()` is the concrete config-path implementation (in `dsh-agent-loop`), while programmatic consumers create/resume owned agents through `ctx.agents.create()` / `ctx.agents.resume()`. Replace the loop by implementing `Agent` and registering via `ctx.agents.register()`. - Event listeners: all `agent/*` events are declared here — no dependency on the loop package needed. +- Subagent delegation: implemented by `@deepseek-ai/dsh-subagent`, not by a method on `Agent`; providers create or drive ordinary `Agent` handles through the factory seam, so spawn/fork/ACP transports stay outside the core agent interface. ### What is NOT here (TODO) -- **Sub-agent spawn/fork** — seam on `AgentLoop.create()`, semantics deferred. +- **Inter-agent channels beyond delegation** — shared state, streaming child output, and background/poll semantics remain outside the current synchronous `ctx.subagents` seam. diff --git a/packages/core/agent/src/types.ts b/packages/core/agent/src/types.ts index efe392155c..407cce5250 100644 --- a/packages/core/agent/src/types.ts +++ b/packages/core/agent/src/types.ts @@ -179,11 +179,45 @@ declare module 'cordis' { */ 'agent/step-end'(agent: Agent, turn: number, step: number): void - // ---- interception seams (waterfall) ---- + // ---- step/request extension seams (serial + waterfall) ---- + /** + * Awaited pre-step surface-mutation checkpoint, fired once per step AFTER + * `turn/start` (and after the prior step closed) but BEFORE this step's + * `step/start` — so anything a listener appends lands OUTSIDE the step, + * between `turn/start`/`step/end` and the upcoming `step/start`. `step` is + * the number of the step about to start. The loop awaits + * `ctx.serial('agent/pre-step', …)` after assembling the system prompt, then + * opens the step and derives the request history ONCE from whatever the + * surface now holds. This is where compaction belongs: it mutates the session + * surface in place (shadowing an older range with a summary node) with its + * log-only `compact/*` records cleanly outside any step, and the single + * subsequent derive reflects the mutation — so there is no double-derive and + * no listener can see (or be expected to act on) an assembled `messages` + * array that does not exist yet. + * + * Serial (awaited in registration order), not a waterfall: a listener + * mutates the surface as a side effect; there is nothing to transform, but + * the loop must wait for the mutation to complete before opening the step + * and deriving. Cordis `serial` bails early if a listener returns a bail + * value; this event is typed and documented as `void`, so listeners must not + * return a semantic veto value. `fullSystemPrompt` is the assembled prompt a + * listener needs to measure pressure (the system prompt counts toward the + * budget). `signal` cancels any in-flight work a listener starts (e.g. a + * summarization model call). + * @mode serial + */ + // TODO: `fullSystemPrompt` is a smell on a generic per-step seam — compaction + // is its only consumer, so a wide event carries a string just one listener + // reads. Revisit if no second consumer appears: e.g. hand listeners a lazy + // prompt provider, or move token-pressure measurement behind a + // compaction-specific seam instead of the shared pre-step checkpoint. + 'agent/pre-step'(agent: Agent, turn: number, step: number, fullSystemPrompt: string, signal: AbortSignal): Promise | void /** * Waterfall: mutate the fully-assembled {@link GenerateOptions} before the - * model call (hooks, compaction, model switching, tool filtering, …). Call - * `next()` to delegate, or return without it to short-circuit. + * model call (hooks, model switching, tool filtering, …). Call `next()` to + * delegate, or return without it to short-circuit. For surface mutation that + * must precede history derivation (compaction), use {@link agent/pre-step} + * instead — by the time this fires, `options.messages` is already derived. * @mode waterfall */ 'agent/request'(agent: Agent, turn: number, step: number, options: GenerateOptions, next: () => Promise): Promise diff --git a/packages/core/session/README.md b/packages/core/session/README.md index 2a90d0f792..45512c9d9e 100644 --- a/packages/core/session/README.md +++ b/packages/core/session/README.md @@ -8,7 +8,7 @@ Creates and holds event-sourced `Session` instances. Persistence is intentionall ### Public API -- `ctx.sessions.create(id?: SessionId, options?: { seed?: SessionEvent[]; meta?: { cwd?: string; parentSession?: SessionId; createdAt?: number } }): Session` — Create a session. `options.seed` replays/forks an existing event log; `options.meta` attaches creation metadata (validated absolute `cwd`, `parentSession` lineage) as the immutable `SessionHeader`. The store fills `version`/`id` and defaults `createdAt` to now; a caller reconstructing a persisted session passes the original `createdAt` to preserve it. Disposed with the calling fiber. +- `ctx.sessions.create(id?: SessionId, options?: { seed?: SessionEvent[]; meta?: { cwd?: string; parentSession?: SessionId; createdAt?: number; seedLength?: number } }): Session` — Create a session. `options.seed` replays/forks an existing event log; `options.meta` attaches creation metadata (validated absolute `cwd`, `parentSession` lineage, seed boundary) as the immutable `SessionHeader`. The store fills `version`/`id` and defaults `createdAt` to now; a caller reconstructing a persisted session passes the original `createdAt` and persisted `seedLength` to preserve them. Disposed with the calling fiber. - `ctx.sessions.get(id: SessionId): Session | undefined` - `ctx.sessions.list(): Session[]` @@ -38,7 +38,7 @@ Plain class (not a Cordis Service). Create via `ctx.sessions.create()`. - `session.deriveMessages(): Message[]` — derive the LLM message history by walking the surface linked list (skipping non-surface events like chunks and boundaries; a `replace` shadows the nodes it covers). The surface is the single source of derived history — there is no raw-log fallback. - `session.surface: SurfaceManager` — the derived surface, lazily rebuilt from `surfaceOp` markers in the log. Processes only new events (delta) on each access — the log is append-only, so prior events never change. - `session.events`, `session.seq`, `session.id` -- `session.header: SessionHeader` — immutable creation metadata (`version`, `id`, `createdAt`, optional `cwd`/`parentSession`). Kept out of the event log (a storage concern, not replayable state); a minimal header (stamped with the current `SESSION_FORMAT_VERSION`) is synthesized for bare `Session` construction. +- `session.header: SessionHeader` — immutable creation metadata (`version`, `id`, `createdAt`, optional `cwd`/`parentSession`/`seedLength`). Kept out of the event log (a storage concern, not replayable state); a minimal header (stamped with the current `SESSION_FORMAT_VERSION`) is synthesized for bare `Session` construction. ### Surface types @@ -49,26 +49,26 @@ Plain class (not a Cordis Service). Create via `ctx.sessions.create()`. ### Session event vocabulary (`types.ts`) -The append-only log: `turn/start`, `turn/end`, `step/start`, `step/end`, `user/message`, `assistant/message`, `assistant/chunk`, `tool/call`, `tool/result`, `steering/message`, `context/message`. Token usage rides on `assistant/message.usage`; an operational error's step is on `turn/end.reason` for `kind: 'error'`. +The append-only log: `turn/start`, `turn/end`, `step/start`, `step/end`, `user/message`, `assistant/message`, `assistant/chunk`, `tool/call`, `tool/result`, `steering/message`, `context/message`, `todo/write`. Token usage rides on `assistant/message.usage`; an operational error's step is on `turn/end.reason` for `kind: 'error'`. -Merge-extensible via `SessionEventMap` — a compaction plugin adds `compaction/marker`, etc. +Merge-extensible via `SessionEventMap` — the compaction seam adds `compact/start`, `compact/summary`, and `compact/end`. Also defines `TurnTriggerMap` and `TurnEndReasonMap` (merge-extensible sum types for typed turn boundaries — `kind`-tagged instead of strings). Every `SessionEvent` carries two optional top-level fields (structural metadata): -- `sourceEventSeqs?: number[]` — seq numbers of provenance sources (e.g., the `assistant/chunk` seqs behind an `assistant/message`, or the shadowed nodes behind a compaction marker). +- `sourceEventSeqs?: number[]` — seq numbers of provenance sources (e.g., the `assistant/chunk` seqs behind an `assistant/message`, or the shadowed nodes behind a compaction replace node). - `surfaceOp?: SurfaceOp` — how this event entered the surface. Absent for non-surface events (boundaries, chunks, usage, errors). ### Metadata types (`types.ts`) -- `SessionHeader` — immutable session metadata, written once: `{ version, id, createdAt, cwd?, parentSession? }`. Owned here (beside `SessionId`) because `Session.header` is typed by it; persistence backends re-export it rather than own it (which would force a package cycle). +- `SessionHeader` — immutable session metadata, written once: `{ version, id, createdAt, cwd?, parentSession?, seedLength? }`. Owned here (beside `SessionId`) because `Session.header` is typed by it; persistence backends re-export it rather than own it (which would force a package cycle). ### Extension points - Persistence plugins: subscribe to `session/event` (write-behind) and drain on `session/flush` (awaited) and fiber dispose. A durable backend reads the log and reloads it into a live session; the metadata seam (`SessionHeader`, `session.header`) is what such a backend stores beside the log. - Replay/fork: `ctx.sessions.create(id, { seed })` seeds a new session with an existing event log. The surface rebuilds deterministically from `surfaceOp` markers in the seeded events. The seed is validated to the SAME invariants `append` enforces — including that every surface-eligible event (`SurfaceEventType`) carries a `surfaceOp` marker — so a marker-less message event is rejected at construction rather than silently vanishing from `deriveMessages()` (the surface is the sole derivation path) on resume. -- Compaction: a future plugin appends a new event with `surfaceOp: { op: 'replace', start, end }` to shadow old surface nodes. +- Compaction: the `dsh-compact-basic` plugin appends a `user/message` with `surfaceOp: { op: 'replace', start, end }` to shadow old surface nodes behind a summary checkpoint. ### What is NOT here (TODO) diff --git a/packages/core/session/src/index.ts b/packages/core/session/src/index.ts index 2002c93051..cef5e93e91 100644 --- a/packages/core/session/src/index.ts +++ b/packages/core/session/src/index.ts @@ -19,6 +19,7 @@ export { isJsonValue } from './json.ts' export { interruptedTurnClosers } from './repair.ts' export type { SurfaceNode } from './surface.ts' export { isSurfaceEvent, isSurfaceEligibleType } from './surface.ts' +export { isToolPairingBalanced } from './tool-pairing.ts' declare module 'cordis' { interface Context { @@ -95,12 +96,12 @@ export class Session { } /** - * Immutable creation metadata (format version, cwd, lineage). Supplied by - * the store via `ctx.sessions.create()`. When a `Session` is constructed - * bare (tests, ad-hoc replay), a minimal header is synthesized (stamped with - * the current {@link SESSION_FORMAT_VERSION}) so `session.header` is always - * present. Kept out of the event log — it is a storage concern, not - * replayable conversation state. + * Immutable creation metadata (format version, cwd, lineage, seed boundary). + * Supplied by the store via `ctx.sessions.create()`. When a `Session` is + * constructed bare (tests, ad-hoc replay), a minimal header is synthesized + * (stamped with the current {@link SESSION_FORMAT_VERSION}) so + * `session.header` is always present. Kept out of the event log — it is a + * storage concern, not replayable conversation state. */ readonly header: SessionHeader diff --git a/packages/core/session/src/tool-pairing.ts b/packages/core/session/src/tool-pairing.ts new file mode 100644 index 0000000000..638daaf654 --- /dev/null +++ b/packages/core/session/src/tool-pairing.ts @@ -0,0 +1,100 @@ +/** + * Tool-pairing balance over a session's SURFACE: is a given cut point in the + * surface a safe edge for a collapsed region (e.g. compaction)? + * + * The invariant a consumer needs: a collapsed region must never separate an + * `assistant/message`'s `tool-call` blocks from their answering `tool/result`s + * — that would leave the rehydrated transcript with a dangling tool-call or an + * orphaned tool-result, which every provider rejects. (This is the + * compaction-time mirror of the crash-recovery imbalance that + * {@link interruptedTurnClosers} repairs on load.) Steps were once used as a + * proxy for this bracketing, but a compaction REWRITES the surface — it lands a + * replacement node at a high log seq whose SURFACE position is the head — so a + * scan over the LOG's `step/*` markers mis-reads such a node's neighbours. The + * pairing the invariant actually protects lives in the surface nodes' own + * content (a `tool-call` block's id, a `tool/result`'s `callId`), which travels + * with the node through any reshaping, so alignment is decided over the surface + * directly. + * + * A **cut** is a gap between two adjacent surface nodes (named by the node it + * sits immediately before), or the after-tail gap (`null`). Walking the surface + * head→tail and assigning each node a delta — `+1` per `tool-call` block on an + * `assistant/message`, `-1` per `tool/result`, `0` otherwise — the depth at a + * cut is the number of still-unanswered tool calls before it. A cut is + * **balanced** when that depth is `0`. A region `[start..end]` is safe to + * collapse iff BOTH its edges are balanced cuts: the cut before `start` and the + * cut after `end`. Nodes that belong to no step (a pre-step `user/message`, an + * inter-step `steering/message`, an injection `context/message`) carry no + * pairing, contribute `0`, and so are free boundaries — exactly as before, but + * now as a consequence of the balance rather than a special case. An open + * trailing step (an assistant whose `tool/result`s have not landed yet) keeps + * the depth positive through the tail, so no cut inside it is balanced — the + * old explicit open-step check falls out of the same counter. + * + * @module @deepseek-ai/dsh-session/tool-pairing + */ + +import type { SessionEvent } from './types.ts' +import type { SurfaceNode } from './surface.ts' + +/** + * The tool-pairing delta of a surface node: how it shifts the count of + * unanswered tool calls. An `assistant/message` opens one bracket per + * `tool-call` block; a `tool/result` closes one; every other surface node + * (`user/message`, `context/message`, `steering/message`, a usage-only + * `assistant/message` with no tool-call blocks) is pairing-neutral. + */ +function nodeDelta(event: SessionEvent): number { + switch (event.type) { + case 'assistant/message': + return event.data.content.filter(block => block.type === 'tool-call').length + case 'tool/result': + return -1 + // Non-pairing surface nodes and every non-surface event contribute nothing. + default: + return 0 + } +} + +/** + * Whether the surface prefix ending at the given cut has BALANCED tool-call / + * tool-result brackets — i.e. every `tool-call` block on the surface before the + * cut has its answering `tool/result` before the cut too, so the cut is a safe + * edge for a collapsed region (it cannot split an assistant↔result pair). + * + * `nodes` is the surface linked list in head→tail order (e.g. + * `session.surface.nodes`); `events` is the session log, used to look each + * node's event up by `seq`. `beforeSeq` names the cut by the surface node it + * sits immediately before; the after-tail cut (the whole surface) is `null`, + * as is any `beforeSeq` not present on the surface. + * + * A region `[start..end]` is collapsible iff both edges are balanced cuts: call + * `isToolPairingBalanced(nodes, events, start)` for the cut before `start`, and + * `isToolPairingBalanced(nodes, events, after)` — where `after` is `end`'s + * surface successor (`SurfaceNode.next`), or `null` when `end` is the tail — + * for the cut after `end`. + * + * @throws if the surface prefix drives the unanswered-call depth negative — a + * `tool/result` with no preceding open `tool-call` on the surface. That is a + * corrupt surface (a structural invariant violation), surfaced loudly here + * rather than silently mis-classifying a boundary. + */ +export function isToolPairingBalanced( + nodes: readonly SurfaceNode[], + events: readonly SessionEvent[], + beforeSeq: number | null, +): boolean { + let depth = 0 + for (const node of nodes) { + if (node.seq === beforeSeq) return depth === 0 + // node.seq is a surface-node seq, always a valid log index by construction. + // eslint-disable-next-line @typescript-eslint/no-non-null-assertion + depth += nodeDelta(events[node.seq]!) + if (depth < 0) { + throw new Error(`tool-pairing balance: tool/result at surface seq ${node.seq} has no matching tool-call (corrupt surface)`) + } + } + // Reached the after-tail cut (beforeSeq === null, or a seq not on the + // surface): the whole-surface prefix is balanced iff depth returned to 0. + return depth === 0 +} diff --git a/packages/core/session/src/types.ts b/packages/core/session/src/types.ts index 41f878892f..6ed8c4391e 100644 --- a/packages/core/session/src/types.ts +++ b/packages/core/session/src/types.ts @@ -149,6 +149,24 @@ export interface TurnEndReasonMap { export type TurnEndReason = TurnEndReasonMap[keyof TurnEndReasonMap] +/** + * One entry in an agent's todo list — the unit of the `todo/write` + * {@link SessionEventMap} event's whole-list snapshot. + * + * Deliberately minimal: a human-readable `content` line and a three-state + * `status`. No id, priority, or `activeForm` — the list is replaced wholesale + * on every write (last-write-wins), so entries need no stable identity, and the + * status triple is exactly the ACP `PlanEntryStatus`, so a UI bridge can map a + * todo list onto an ACP `plan` 1:1 (synthesizing the priority ACP additionally + * requires). + */ +export interface TodoItem { + /** What this task is — a short imperative line shown in the UI. */ + content: string + /** Lifecycle state. `in_progress` marks the single task being worked now. */ + status: 'pending' | 'in_progress' | 'completed' +} + /** * The session event vocabulary — the append-only source of truth for an * agent's whole interaction history. The LLM message history is *derived* @@ -156,7 +174,8 @@ export type TurnEndReason = TurnEndReasonMap[keyof TurnEndReasonMap] * same events; trace/telemetry = subscribe to the log. * * Merge-extensible: plugins declare extra event types via declaration merging - * (e.g. a compaction plugin adds `'compaction/marker'`). + * (e.g. the compaction plugin adds `'compact/start'`, `'compact/summary'`, + * `'compact/end'`). * * Durability contract (what a persistence backend relies on): the durable log * persists every event verbatim, INCLUDING `assistant/chunk` — `seq` must stay @@ -194,6 +213,20 @@ export interface SessionEventMap { 'tool/result': { turn: number; step: number; callId: CallId; content: ContentBlock[]; isError: boolean; error?: { name: string; code: string } } /** Steering content injected between steps of a running turn. */ 'steering/message': { turn: number; content: ContentBlock[]; source: MessageSource } + /** + * The agent's whole todo list, carried as a full snapshot and replaced + * wholesale on each write — the current list is the most recent `todo/write` + * (last-write-wins on replay, no fold). Appended by an owning agent via + * `session.append('todo/write', { todos })`. + * + * NOT a {@link SurfaceEventType}: it produces no LLM message and never reaches + * `deriveMessages()`, so it carries no `surfaceOp` and stays off the surface — + * it is durable, replayable UI state, distinct from the conversation history. + * It is a `SessionEventMap` member riding the existing `session/event` emit, + * not a first-class Cordis `interface Events` notification, so it has no + * cordis-catalog row. + */ + 'todo/write': { todos: TodoItem[] } } export type SessionEventType = keyof SessionEventMap @@ -279,7 +312,7 @@ export type SessionEvent = { /** * Seq numbers of events that are provenance sources of this event * (e.g. the `assistant/chunk` seqs that built an `assistant/message`, - * or the surface nodes shadowed by a compaction marker). + * or the surface nodes shadowed by a compaction replace node). */ sourceEventSeqs?: number[] /** How this event entered the surface; absent for non-surface events. */ diff --git a/packages/core/session/tests/session.spec.ts b/packages/core/session/tests/session.spec.ts index 6138a479f1..27f8b5d420 100644 --- a/packages/core/session/tests/session.spec.ts +++ b/packages/core/session/tests/session.spec.ts @@ -2,7 +2,7 @@ import { describe, expect, it } from 'vitest' import { Context } from 'cordis' import { CallId } from '@deepseek-ai/dsh-llm' import SessionStore, { SESSION_FORMAT_VERSION, Session, SessionEvent, SessionId } from '@deepseek-ai/dsh-session' -import type { SessionEventType } from '@deepseek-ai/dsh-session' +import type { SessionEventType, TodoItem } from '@deepseek-ai/dsh-session' describe('Session', () => { it('derives message history from the event log', () => { @@ -365,3 +365,63 @@ describe('SessionStore', () => { expect(events).toHaveLength(1) }) }) + +describe('todo/write event', () => { + it('appends the whole-list snapshot and isolates the log from later mutation', () => { + const session = new Session(SessionId('t1')) + const todos: TodoItem[] = [ + { content: 'plan the work', status: 'in_progress' }, + { content: 'write the code', status: 'pending' }, + ] + session.append('todo/write', { todos }) + + const event = session.events.findLast(e => e.type === 'todo/write')! + expect(event.type).toBe('todo/write') + expect(event.data.todos).toEqual(todos) + + // The append snapshots its input: mutating the caller's array afterward must + // not change what the log holds (the durable-source-of-truth contract). + todos.push({ content: 'sneak in', status: 'pending' }) + todos[0]!.status = 'completed' + expect(event.data.todos).toEqual([ + { content: 'plan the work', status: 'in_progress' }, + { content: 'write the code', status: 'pending' }, + ]) + }) + + it('is last-write-wins: the current list is the most recent todo/write', () => { + const session = new Session(SessionId('t2')) + session.append('todo/write', { todos: [{ content: 'first', status: 'pending' }] }) + session.append('todo/write', { todos: [ + { content: 'first', status: 'completed' }, + { content: 'second', status: 'in_progress' }, + ] }) + + const current = session.events.findLast(e => e.type === 'todo/write')!.data.todos + expect(current).toEqual([ + { content: 'first', status: 'completed' }, + { content: 'second', status: 'in_progress' }, + ]) + }) + + it('is NOT a surface event: it produces no derived message and joins no surface node', () => { + const session = new Session(SessionId('t3')) + session.append('user/message', { content: [{ type: 'text', text: 'q' }], source: { kind: 'user' } }, { surfaceOp: 'append' }) + const before = session.deriveMessages().length + session.append('todo/write', { todos: [{ content: 'a task', status: 'pending' }] }) + // The todo event must not add a message to the derived history… + expect(session.deriveMessages()).toHaveLength(before) + // …and must not appear on the surface linked list. + expect(session.surface.nodes.some(node => node.seq === session.seq - 1)).toBe(false) + }) + + it('round-trips through a seeded replay identically (durable, no surfaceOp needed)', () => { + const original = new Session(SessionId('t4')) + original.append('todo/write', { todos: [{ content: 'only', status: 'completed' }] }) + // Seeding a non-surface event with no surfaceOp must not throw. + const replayed = new Session(SessionId('t4-replay'), [...original.events]) + expect(replayed.events.findLast(e => e.type === 'todo/write')!.data.todos) + .toEqual([{ content: 'only', status: 'completed' }]) + expect(replayed.seq).toBe(original.seq) + }) +}) diff --git a/packages/core/session/tests/surface.spec.ts b/packages/core/session/tests/surface.spec.ts index 4e6bc1f433..6a37e6fcab 100644 --- a/packages/core/session/tests/surface.spec.ts +++ b/packages/core/session/tests/surface.spec.ts @@ -278,6 +278,23 @@ describe('Session.append surface opts', () => { // The string 'append' is a primitive — identity-preserving is fine. expect(event.surfaceOp).toBe('append') }) + + it('isSurfaceEvent rejects a surface-eligible type missing its surfaceOp marker', () => { + // A raw event (not built via append, which mandates the marker) of a + // surface-eligible type but with no surfaceOp must NOT narrow to a + // SurfaceEvent — it would otherwise be silently dropped from the surface. + const noMarker: SessionEvent = { + type: 'user/message', seq: 0, time: 1, + data: { content: [{ type: 'text', text: 'hi' }], source: { kind: 'user' } }, + } + expect(isSurfaceEvent(noMarker)).toBe(false) + // A non-surface type is rejected too (the type gate). + const boundary: SessionEvent = { type: 'turn/start', seq: 1, time: 1, data: { turn: 1, trigger: { kind: 'message', source: { kind: 'user' } } } } + expect(isSurfaceEvent(boundary)).toBe(false) + // A properly-marked surface event narrows. + const marked = { ...noMarker, surfaceOp: 'append' } as SurfaceEvent + expect(isSurfaceEvent(marked)).toBe(true) + }) }) describe('surface type guards', () => { diff --git a/packages/core/session/tests/tool-pairing.spec.ts b/packages/core/session/tests/tool-pairing.spec.ts new file mode 100644 index 0000000000..307b0d8658 --- /dev/null +++ b/packages/core/session/tests/tool-pairing.spec.ts @@ -0,0 +1,314 @@ +import { describe, expect, it } from 'vitest' +import { CallId } from '@deepseek-ai/dsh-llm' +import { Session, SessionId, isToolPairingBalanced } from '../src/index.ts' +import type { SessionEvent, SurfaceNode } from '../src/index.ts' + +/** + * Unit coverage for the tool-pairing balance check. It decides whether a CUT in + * the surface (a gap before a given surface node, or the after-tail gap) is a + * safe edge for a collapsed region (compaction): a region must never split an + * `assistant/message`'s tool-calls from their `tool/result`s. A cut is balanced + * when no unanswered tool-call sits before it on the surface. Nodes belonging to + * no step (pre-step user message, inter-step steering, injection context) are + * pairing-neutral, so their cuts are free boundaries. + * + * The fixtures are built through a real {@link Session} so the surface linked + * list is derived exactly as production does — including the non-monotonic + * surface a `replace` op leaves (a compaction checkpoint at a high log seq + * sitting at the surface head), which is the case the abandoned log-position + * scan mis-classified. + * + * Builders mirror the agent loop's real append order: queued user messages land + * BEFORE `step/start`; within a step the order is `assistant/message` then + * `tool/result`(s); injection turns are a bare `turn/start → context/message → + * turn/end` with no step. + */ + +const SURFACE = { surfaceOp: 'append' as const } + +/** Surface nodes + log for a session, the two args the balance check takes. */ +function surfaceOf(session: Session): { nodes: readonly SurfaceNode[]; events: readonly SessionEvent[] } { + return { nodes: session.surface.nodes, events: session.events } +} + +/** The cut BEFORE the surface node at `seq` is balanced (safe region start). */ +function startBalanced(session: Session, seq: number): boolean { + const { nodes, events } = surfaceOf(session) + return isToolPairingBalanced(nodes, events, seq) +} + +/** The cut AFTER the surface node at `seq` is balanced (safe region end). */ +function endBalanced(session: Session, seq: number): boolean { + const { nodes, events } = surfaceOf(session) + const node = nodes.find(n => n.seq === seq) + if (!node) throw new Error(`seq ${seq} is not a surface node`) + return isToolPairingBalanced(nodes, events, node.next) +} + +/** Surface seq of the nth (0-based) event of a given type. */ +function seqOf(s: Session, type: SessionEvent['type'], nth = 0): number { + return s.events.filter(e => e.type === type)[nth]!.seq +} + +/** A closed turn with one closed step holding an assistant + its tool result. */ +function toolStepSession(): Session { + const s = new Session(SessionId('tool-step')) + s.append('turn/start', { turn: 1, trigger: { kind: 'message', source: { kind: 'user' } } }) + s.append('user/message', { content: [{ type: 'text', text: 'go' }], source: { kind: 'user' } }, SURFACE) + s.append('step/start', { turn: 1, step: 1 }) + s.append('assistant/message', { + turn: 1, step: 1, + content: [ + { type: 'text', text: 'calling' }, + { type: 'tool-call', id: CallId('c1'), name: 'bash', arguments: '{}' }, + ], + }, SURFACE) + s.append('tool/call', { turn: 1, step: 1, callId: CallId('c1'), name: 'bash', arguments: '{}' }) + s.append('tool/result', { turn: 1, step: 1, callId: CallId('c1'), content: [{ type: 'text', text: 'out' }], isError: false }, SURFACE) + s.append('step/end', { turn: 1, step: 1 }) + s.append('turn/end', { turn: 1, reason: { kind: 'completed' } }) + return s +} + +describe('isToolPairingBalanced — region START (cut before a node)', () => { + it('is true for a pre-step user/message (belongs to no step)', () => { + const s = toolStepSession() + expect(startBalanced(s, seqOf(s, 'user/message'))).toBe(true) + }) + + it('is true for the first surface node of a step (the assistant/message)', () => { + // The cut before the assistant is balanced — nothing unanswered precedes it. + const s = toolStepSession() + expect(startBalanced(s, seqOf(s, 'assistant/message'))).toBe(true) + }) + + it('is false for a tool/result whose assistant/message precedes it in the same step', () => { + // The cut before the tool/result has one unanswered tool-call (the + // assistant's) → starting the region here would orphan that call. + const s = toolStepSession() + expect(startBalanced(s, seqOf(s, 'tool/result'))).toBe(false) + }) + + it('is true at the surface head (nothing precedes)', () => { + const s = new Session(SessionId('lone')) + s.append('turn/start', { turn: 1, trigger: { kind: 'message', source: { kind: 'user' } } }) + s.append('user/message', { content: [{ type: 'text', text: 'hi' }], source: { kind: 'user' } }, SURFACE) + expect(startBalanced(s, seqOf(s, 'user/message'))).toBe(true) + }) +}) + +describe('isToolPairingBalanced — region END (cut after a node)', () => { + it('is true for the last surface node of a closed step (the tool/result)', () => { + // After the tool/result the assistant's single call is answered → balanced. + const s = toolStepSession() + expect(endBalanced(s, seqOf(s, 'tool/result'))).toBe(true) + }) + + it('is false for an assistant/message with a later tool/result in the same step', () => { + // After the assistant its tool-call is still unanswered → ending here strands + // the result. + const s = toolStepSession() + expect(endBalanced(s, seqOf(s, 'assistant/message'))).toBe(false) + }) + + it('is true for a pre-step user/message', () => { + const s = toolStepSession() + expect(endBalanced(s, seqOf(s, 'user/message'))).toBe(true) + }) + + it('is false at the tail when the node is inside an open (unclosed) step', () => { + // step/start then an assistant tool-call, but no tool/result yet (mid-flight). + // The after-tail cut still has one unanswered call → not balanced. + const s = new Session(SessionId('open-step')) + s.append('turn/start', { turn: 1, trigger: { kind: 'message', source: { kind: 'user' } } }) + s.append('step/start', { turn: 1, step: 1 }) + s.append('assistant/message', { + turn: 1, step: 1, + content: [{ type: 'tool-call', id: CallId('c1'), name: 'bash', arguments: '{}' }], + }, SURFACE) + expect(endBalanced(s, seqOf(s, 'assistant/message'))).toBe(false) + }) + + it('is true at the tail when the node is a trailing inter-step node (step already closed)', () => { + // A steering message appended after step/end, at the tail. The prior step's + // pair is balanced and steering is neutral → the after-tail cut is balanced. + const s = new Session(SessionId('trailing-steer')) + s.append('turn/start', { turn: 1, trigger: { kind: 'message', source: { kind: 'user' } } }) + s.append('step/start', { turn: 1, step: 1 }) + s.append('assistant/message', { turn: 1, step: 1, content: [{ type: 'text', text: 'a' }] }, SURFACE) + s.append('step/end', { turn: 1, step: 1 }) + s.append('steering/message', { turn: 1, content: [{ type: 'text', text: 's' }], source: { kind: 'user' } }, SURFACE) + expect(endBalanced(s, seqOf(s, 'steering/message'))).toBe(true) + }) + + it('is true at the tail when no step ever opened', () => { + const s = new Session(SessionId('no-step')) + s.append('turn/start', { turn: 1, trigger: { kind: 'message', source: { kind: 'user' } } }) + s.append('user/message', { content: [{ type: 'text', text: 'hi' }], source: { kind: 'user' } }, SURFACE) + expect(endBalanced(s, seqOf(s, 'user/message'))).toBe(true) + }) +}) + +describe('isToolPairingBalanced — multiple tool calls in one assistant message', () => { + // An assistant message with two tool-calls needs BOTH results before the cut + // after it is balanced — depth +2, then -1, -1. + function twoCallStep(): Session { + const s = new Session(SessionId('two-call')) + s.append('turn/start', { turn: 1, trigger: { kind: 'message', source: { kind: 'user' } } }) + s.append('step/start', { turn: 1, step: 1 }) + s.append('assistant/message', { + turn: 1, step: 1, + content: [ + { type: 'tool-call', id: CallId('c1'), name: 'a', arguments: '{}' }, + { type: 'tool-call', id: CallId('c2'), name: 'b', arguments: '{}' }, + ], + }, SURFACE) + s.append('tool/result', { turn: 1, step: 1, callId: CallId('c1'), content: [{ type: 'text', text: '1' }], isError: false }, SURFACE) + s.append('tool/result', { turn: 1, step: 1, callId: CallId('c2'), content: [{ type: 'text', text: '2' }], isError: false }, SURFACE) + s.append('step/end', { turn: 1, step: 1 }) + s.append('turn/end', { turn: 1, reason: { kind: 'completed' } }) + return s + } + + it('is unbalanced after the first of two results (one call still open)', () => { + const s = twoCallStep() + expect(endBalanced(s, seqOf(s, 'tool/result', 0))).toBe(false) + }) + + it('is balanced after the second result (both calls answered)', () => { + const s = twoCallStep() + expect(endBalanced(s, seqOf(s, 'tool/result', 1))).toBe(true) + }) +}) + +describe('isToolPairingBalanced — a mid-step injection context/message', () => { + // A background task-done inject() lands a context/message INSIDE an open step, + // between the assistant (with a tool-call) and its tool/result. It is + // pairing-neutral, so the cut on EITHER side of it is unbalanced (the call is + // still open across it) — it is NOT a free boundary in this position. + function midStepInjection(): Session { + const s = new Session(SessionId('mid-inject')) + s.append('turn/start', { turn: 1, trigger: { kind: 'message', source: { kind: 'user' } } }) + s.append('step/start', { turn: 1, step: 1 }) + s.append('assistant/message', { + turn: 1, step: 1, + content: [{ type: 'tool-call', id: CallId('c1'), name: 'bash', arguments: '{}' }], + }, SURFACE) + s.append('context/message', { content: [{ type: 'text', text: 'bg task done' }], source: { kind: 'plugin', plugin: 'tool-bash' } }, SURFACE) + s.append('tool/result', { turn: 1, step: 1, callId: CallId('c1'), content: [{ type: 'text', text: 'out' }], isError: false }, SURFACE) + s.append('step/end', { turn: 1, step: 1 }) + s.append('turn/end', { turn: 1, reason: { kind: 'completed' } }) + return s + } + + it('start cut before the mid-step context/message is unbalanced (call still open)', () => { + const s = midStepInjection() + expect(startBalanced(s, seqOf(s, 'context/message'))).toBe(false) + }) + + it('end cut after the mid-step context/message is unbalanced (call still open)', () => { + const s = midStepInjection() + expect(endBalanced(s, seqOf(s, 'context/message'))).toBe(false) + }) +}) + +describe('isToolPairingBalanced on an injection turn (no step)', () => { + // An idle inject() wraps a context/message in a bare turn/start → + // context/message → turn/end with NO step. The context node is a free boundary + // both ways (pairing-neutral, nothing open around it). + function injectionSession(): Session { + const s = new Session(SessionId('injection')) + s.append('turn/start', { turn: 1, trigger: { kind: 'injection', source: { kind: 'user' } } }) + s.append('context/message', { content: [{ type: 'text', text: 'ctx' }], source: { kind: 'user' } }, SURFACE) + s.append('turn/end', { turn: 1, reason: { kind: 'completed' } }) + return s + } + + it('start: balanced', () => { + const s = injectionSession() + expect(startBalanced(s, seqOf(s, 'context/message'))).toBe(true) + }) + + it('end: balanced', () => { + const s = injectionSession() + expect(endBalanced(s, seqOf(s, 'context/message'))).toBe(true) + }) +}) + +describe('isToolPairingBalanced — CBR-001: a head checkpoint left by a replace op', () => { + // The case the log-position scan got wrong. After a compaction, a replacement + // user/message lands at a HIGH log seq but sits at the SURFACE head, beside + // the still-open step whose events follow it in the log. It carries no + // tool-call/result pair (just summarized prose), so it must be a balanced cut + // on BOTH sides regardless of its log neighbours. + function checkpointHeadedSession(): Session { + const s = new Session(SessionId('checkpoint')) + // A closed turn with a tool step → surface [u1, asst(call), result]. + s.append('turn/start', { turn: 1, trigger: { kind: 'message', source: { kind: 'user' } } }) + s.append('step/start', { turn: 1, step: 1 }) + s.append('user/message', { content: [{ type: 'text', text: 'u1' }], source: { kind: 'user' } }, SURFACE) + s.append('assistant/message', { + turn: 1, step: 1, + content: [{ type: 'tool-call', id: CallId('c1'), name: 'bash', arguments: '{}' }], + }, SURFACE) + s.append('tool/result', { turn: 1, step: 1, callId: CallId('c1'), content: [{ type: 'text', text: 'out' }], isError: false }, SURFACE) + s.append('step/end', { turn: 1, step: 1 }) + s.append('turn/end', { turn: 1, reason: { kind: 'completed' } }) + // An OPEN turn whose step is in progress (loop fires compaction here). + s.append('turn/start', { turn: 2, trigger: { kind: 'message', source: { kind: 'user' } } }) + s.append('step/start', { turn: 2, step: 1 }) + // Compaction replaces the whole turn-1 surface ([u1, asst, result]) with one + // summary user/message — appended now, so it carries a high log seq. + const u1 = seqOf(s, 'user/message') + const result = s.events.find(e => e.type === 'tool/result')!.seq + s.append('user/message', { + content: [{ type: 'text', text: 'CHECKPOINT' }], + source: { kind: 'plugin', plugin: 'compact' }, + }, { surfaceOp: { op: 'replace', start: u1, end: result } }) + // The step's own assistant/message lands AFTER the checkpoint in the log, + // still inside the open step. + s.append('assistant/message', { turn: 2, step: 1, content: [{ type: 'text', text: 'a2' }] }, SURFACE) + return s + } + + it('the head checkpoint sits at the surface head while a later surface node follows it in the log', () => { + const s = checkpointHeadedSession() + const nodes = s.surface.nodes + const checkpointSeq = nodes[0]!.seq + // The checkpoint heads the surface, yet a surface node (the open step's + // assistant) follows it in LOG order — the exact split between surface + // position and log position that the log-position scan tripped on. + const laterSurfaceInLog = s.events.find( + e => e.seq > checkpointSeq && nodes.some(n => n.seq === e.seq), + ) + expect(laterSurfaceInLog).toBeDefined() + expect(nodes[0]!.seq).toBe(checkpointSeq) + }) + + it('start cut before the head checkpoint is balanced (it is the head)', () => { + const s = checkpointHeadedSession() + expect(startBalanced(s, s.surface.nodes[0]!.seq)).toBe(true) + }) + + it('end cut after the head checkpoint is balanced (it carries no tool pair)', () => { + // This is the exact assertion the log-position scan failed: the forward log + // scan from the checkpoint reached the open step's assistant/message and + // wrongly reported mid-step. The surface balance sees a neutral node whose + // following cut closes no open call. + const s = checkpointHeadedSession() + expect(endBalanced(s, s.surface.nodes[0]!.seq)).toBe(true) + }) +}) + +describe('isToolPairingBalanced — corrupt surface guard', () => { + it('throws when a tool/result has no preceding tool-call (depth goes negative)', () => { + // A surface that opens with a tool/result (no assistant call before it) is + // structurally corrupt — surfaced loudly rather than mis-classified. + const s = new Session(SessionId('corrupt')) + s.append('turn/start', { turn: 1, trigger: { kind: 'message', source: { kind: 'user' } } }) + s.append('step/start', { turn: 1, step: 1 }) + s.append('tool/result', { turn: 1, step: 1, callId: CallId('c1'), content: [{ type: 'text', text: 'x' }], isError: false }, SURFACE) + const { nodes, events } = surfaceOf(s) + expect(() => isToolPairingBalanced(nodes, events, null)).toThrow(/no matching tool-call/) + }) +}) diff --git a/packages/core/system-prompt/README.md b/packages/core/system-prompt/README.md index 6f14e40f88..1c18bdf1d6 100644 --- a/packages/core/system-prompt/README.md +++ b/packages/core/system-prompt/README.md @@ -34,4 +34,4 @@ Merge-extensible: plugins can declare extra fields on `PromptAssembly` via decla ### What is NOT here - Any hardcoded prompt text — every section comes from plugins. -- Prompt compaction (belongs on the `agent/request` seam in `dsh-agent`). +- Prompt compaction (belongs on the `agent/pre-step` seam in `dsh-agent`). diff --git a/packages/session-persistence/session-persistence-jsonl/README.md b/packages/session-persistence/session-persistence-jsonl/README.md index b8df12e547..9a76381614 100644 --- a/packages/session-persistence/session-persistence-jsonl/README.md +++ b/packages/session-persistence/session-persistence-jsonl/README.md @@ -10,7 +10,7 @@ The JSONL durable session-persistence backend — a concrete `SessionPersistence .jsonl # header line + one SessionEvent per line (verbatim) ``` -- The first `.jsonl` line is the immutable `SessionHeader` tagged `{ type: 'session', version, id, cwd?, createdAt, parentSession? }`; every subsequent line is one `SessionEvent` JSON, **verbatim including `assistant/chunk`** so `seq` stays contiguous (`events[i].seq === i`). +- The first `.jsonl` line is the immutable `SessionHeader` tagged `{ type: 'session', version, id, cwd?, createdAt, parentSession?, seedLength? }`; every subsequent line is one `SessionEvent` JSON, **verbatim including `assistant/chunk`** so `seq` stays contiguous (`events[i].seq === i`). - Session ids are unvalidated branded strings, so they are percent-encoded to a single safe path segment before use (no traversal, no collision). ## Config diff --git a/packages/session-persistence/session-persistence/README.md b/packages/session-persistence/session-persistence/README.md index 928e7b033d..8bd3fed568 100644 --- a/packages/session-persistence/session-persistence/README.md +++ b/packages/session-persistence/session-persistence/README.md @@ -2,7 +2,7 @@ The abstract durable session-persistence seam (`ctx.sessionPersistence`). Defines WHAT a persistence backend does — durably store, reload, and list sessions — without saying HOW. Mirrors the `dsh-bash` capability-seam template ([capability seams](../../../docs/rfc/implemented/architecture/2026-06-13-capability-seams.md)): an abstract service here, a concrete implementation in a sibling package, consumers that inject the interface. -The persisted unit IS the existing `SessionEvent` (event-sourced model — the log is the single source of truth), so there is no parallel "persisted message" type. Metadata that is NOT replayable conversation state (format version, cwd, lineage) travels separately as `SessionHeader`, owned by `dsh-session` and re-exported here. +The persisted unit IS the existing `SessionEvent` (event-sourced model — the log is the single source of truth), so there is no parallel "persisted message" type. Metadata that is NOT replayable conversation state (format version, cwd, lineage, seed boundary) travels separately as `SessionHeader`, owned by `dsh-session` and re-exported here. ## Service API (`ctx.sessionPersistence`) @@ -44,8 +44,8 @@ The `tornMarker` is fully OPAQUE: the coordinator only tests `!== undefined` and Import `runPersistenceContract` from `tests/contract.ts` (the public-API contract) and `runCoordinatorContract` from `tests/coordinator-contract.ts` (the shared write-path orchestration: adoption, HMR, collision, dispose-drain, crash-tail repair) and call each with a fixture for your backend. Every backend is held to the same append-only / contiguous-seq / lazy-materialization / serializability semantics AND the same orchestration, so a backend's own spec is left with only storage-mechanics tests (path sanitization, fsync rollback; schema version, transaction rollback) on top. -Three backends run these suites: an in-memory reference (in `tests/`), `dsh-session-persistence-jsonl` (append-only file log) and `dsh-session-persistence-sqlite` (`node:sqlite`, each `SessionEvent` one row `(session_id, seq, type, time, data)`). All passing the same contract + coordinator suite is the proof that the seam is genuinely backend-agnostic — lazy materialization, crash-tail-on-load, and contiguous-seq hold identically over file bytes and over a transactional store. +Three backends run these suites: an in-memory reference (in `tests/`), `dsh-session-persistence-jsonl` (append-only file log) and `dsh-session-persistence-sqlite` (`node:sqlite`, each `SessionEvent` one row `(session_id, seq, type, time, data, source_event_seqs, surface_op)`). All passing the same contract + coordinator suite is the proof that the seam is genuinely backend-agnostic — lazy materialization, crash-tail-on-load, and contiguous-seq hold identically over file bytes and over a transactional store. ## Metadata types -Re-exported from `dsh-session`: `SessionHeader` (immutable session metadata: `version`, `id`, `createdAt`, `cwd?`, `parentSession?`). +Re-exported from `dsh-session`: `SessionHeader` (immutable session metadata: `version`, `id`, `createdAt`, `cwd?`, `parentSession?`, `seedLength?`). diff --git a/packages/session-persistence/session-persistence/src/index.ts b/packages/session-persistence/session-persistence/src/index.ts index a9ffd11792..f28bc06d5b 100644 --- a/packages/session-persistence/session-persistence/src/index.ts +++ b/packages/session-persistence/session-persistence/src/index.ts @@ -15,8 +15,8 @@ * parallel "persisted message" type the log must be converted to and from * (faithful to the event-sourced model: the log is the single source of * truth). Metadata that is NOT replayable conversation state (format version, - * cwd, lineage) travels separately as {@link SessionHeader}, which is owned by - * `dsh-session` and re-exported here. + * cwd, lineage, seed boundary) travels separately as {@link SessionHeader}, + * which is owned by `dsh-session` and re-exported here. * * @module @deepseek-ai/dsh-session-persistence */ diff --git a/packages/support/README.md b/packages/support/README.md index 52f7f6fe25..942734fc1e 100644 --- a/packages/support/README.md +++ b/packages/support/README.md @@ -7,5 +7,6 @@ Packages that exist to serve development, testing, and the examples rather than | `invariants/` | Dev-mode event-contract invariants + session-log freeze | (listens on `session/*`, `agent/*`) | | `ui-stdio/` | Minimal stdio (readline) UI plugin: renders `agent/*` events, feeds stdin lines to the agent | (drives `ctx.agents`) | | `llm-replay/` | Record/replay adapter: short-circuits `llm/stream` from a recorded session JSONL (keyless snapshot tests) | (listens on `llm/stream`) | +| `subagent-mock/` | Scripted `SubagentProvider` for deterministic seam/tool tests | (registers on `ctx.subagents`) | -`invariants` runs only in dev mode (contract checks, not runtime behavior). `ui-stdio` and `llm-replay` were extracted from the examples for reuse and to bring them under the per-file coverage gate; they back the demos and the snapshot test tier. A package graduates OUT of `support/` into a product group only when it gains documented product consumers. +`invariants` runs only in dev mode (contract checks, not runtime behavior). `ui-stdio` and `llm-replay` were extracted from the examples for reuse and to bring them under the per-file coverage gate; they back the demos and the snapshot test tier. `subagent-mock` exercises the real `ctx.subagents` load path without a model or child agent. A package graduates OUT of `support/` into a product group only when it gains documented product consumers. diff --git a/packages/support/invariants/src/index.ts b/packages/support/invariants/src/index.ts index b2372e28db..08fbf49b57 100644 --- a/packages/support/invariants/src/index.ts +++ b/packages/support/invariants/src/index.ts @@ -162,9 +162,6 @@ function checkEvent(trace: SessionTrace, event: SessionEvent): void { trace.surface.push(event.seq) } else { const { start, end } = se.surfaceOp - if (start > end) { - throw new InvariantError(`surface replace: start ${start} must be <= end ${end}`) - } const startIdx = trace.surface.indexOf(start) if (startIdx === -1) { throw new InvariantError(`surface replace: start seq ${start} is not on the surface`) diff --git a/packages/support/invariants/tests/invariants.spec.ts b/packages/support/invariants/tests/invariants.spec.ts index 6ae2b0edf1..d7e95c1103 100644 --- a/packages/support/invariants/tests/invariants.spec.ts +++ b/packages/support/invariants/tests/invariants.spec.ts @@ -530,16 +530,17 @@ describe('surface invariants', () => { }).toThrow(/unknown seq 2/) }) - it('rejects replace op with start > end', async () => { + it('rejects a replace whose start is positioned after its end on the surface', async () => { const { ctx } = await setup() const session = ctx.sessions.create() session.append('turn/start', { turn: 1, trigger: { kind: 'message', source: { kind: 'user' } } }) session.append('step/start', { turn: 1, step: 1 }) session.append('user/message', { content: [{ type: 'text', text: 'a' }], source: { kind: 'user' } }, { surfaceOp: 'append' }) // seq 2 - // start > end is invalid (reversed order). + session.append('user/message', { content: [{ type: 'text', text: 'b' }], source: { kind: 'user' } }, { surfaceOp: 'append' }) // seq 3 + // Reversed range: start seq 3 is at a later surface position than end seq 2. expect(() => { - session.append('assistant/message', { turn: 1, step: 1, content: [] }, { surfaceOp: { op: 'replace', start: 2, end: 1 }, sourceEventSeqs: [2] }) - }).toThrow(/must be <= end/) + session.append('assistant/message', { turn: 1, step: 1, content: [] }, { surfaceOp: { op: 'replace', start: 3, end: 2 }, sourceEventSeqs: [2, 3] }) + }).toThrow(/is after end seq 2 .* on the surface/) }) it('rejects a replace whose sourceEventSeqs omits a shadowed surface node', async () => { @@ -608,6 +609,23 @@ describe('surface invariants', () => { }).toThrow(/is after end seq 4 .* on the surface/) }) + it('accepts a replace whose start seq exceeds its end seq when the surface position order is valid', async () => { + const { ctx } = await setup() + const session = ctx.sessions.create() + session.append('turn/start', { turn: 1, trigger: { kind: 'message', source: { kind: 'user' } } }) + session.append('step/start', { turn: 1, step: 1 }) + session.append('user/message', { content: [{ type: 'text', text: 'a' }], source: { kind: 'user' } }, { surfaceOp: 'append' }) // seq 2 + session.append('user/message', { content: [{ type: 'text', text: 'b' }], source: { kind: 'user' } }, { surfaceOp: 'append' }) // seq 3 + // Replace node 2 (position 0) with seq 4 — surface becomes [4, 3], so the + // head seq (4) is numerically GREATER than the tail seq (3): the surface is + // not seq-ordered. A replace spanning start=4 (pos 0) … end=3 (pos 1) is + // valid positionally and must be accepted even though start seq > end seq. + session.append('assistant/message', { turn: 1, step: 1, content: [{ type: 'text', text: 's' }] }, { surfaceOp: { op: 'replace', start: 2, end: 2 }, sourceEventSeqs: [2] }) // seq 4 + expect(() => { + session.append('assistant/message', { turn: 1, step: 1, content: [] }, { surfaceOp: { op: 'replace', start: 4, end: 3 }, sourceEventSeqs: [4, 3] }) // seq 5 + }).not.toThrow() + }) + it('rejects a replace that omits sourceEventSeqs entirely', async () => { const { ctx } = await setup() const session = ctx.sessions.create() diff --git a/packages/support/ui-stdio/src/index.ts b/packages/support/ui-stdio/src/index.ts index 070e842094..edfa82285e 100644 --- a/packages/support/ui-stdio/src/index.ts +++ b/packages/support/ui-stdio/src/index.ts @@ -110,6 +110,13 @@ export function createStdioChat(ctx: Context, config: Config, runtime: StdioRunt const { content } = event.data const text = content.filter(block => block.type === 'text').map(block => block.text).join('') output.write(`\n [tool result] ${text}\n `) + } else if (event.type === 'todo/write') { + if (inReasoning) output.write('\x1B[0m') + inReasoning = false + const glyph = (status: string): string => + status === 'completed' ? '[x]' : status === 'in_progress' ? '[~]' : '[ ]' + const lines = event.data.todos.map(todo => ` ${glyph(todo.status)} ${todo.content}`).join('\n') + output.write(`\n [todos]\n${lines}\n `) } }) diff --git a/packages/support/ui-stdio/tests/ui-stdio.spec.ts b/packages/support/ui-stdio/tests/ui-stdio.spec.ts index bd5e0f7f91..7bd1fd4868 100644 --- a/packages/support/ui-stdio/tests/ui-stdio.spec.ts +++ b/packages/support/ui-stdio/tests/ui-stdio.spec.ts @@ -151,6 +151,35 @@ describe('createStdioChat rendering', () => { expect(out.text()).toContain('[tool result] file.txt') }) + it('renders a todo/write session event as a glyphed checklist', async () => { + const { ctx, out } = await setup() + const session = {} as Session + ctx.emit('session/event', session, { + type: 'todo/write', seq: 1, time: 0, + data: { todos: [ + { content: 'read the code', status: 'completed' }, + { content: 'write the fix', status: 'in_progress' }, + { content: 'run the tests', status: 'pending' }, + ] }, + } as SessionEvent) + const text = out.text() + expect(text).toContain('[todos]') + expect(text).toContain('[x] read the code') + expect(text).toContain('[~] write the fix') + expect(text).toContain('[ ] run the tests') + }) + + it('resets dim styling when a todo/write interrupts reasoning', async () => { + const { ctx, out } = await setup() + const agent = makeAgent('main') + ctx.emit('agent/stream-chunk', agent, 1, 0, { type: 'reasoning-delta', index: 0, text: 'r' }) + ctx.emit('session/event', {} as Session, { + type: 'todo/write', seq: 1, time: 0, + data: { todos: [{ content: 'a task', status: 'pending' }] }, + } as SessionEvent) + expect(out.text()).toContain('\x1B[2mr\x1B[0m') + }) + it('resets dim styling when a tool/call interrupts reasoning', async () => { const { ctx, out } = await setup() const agent = makeAgent('main') diff --git a/packages/todo/README.md b/packages/todo/README.md new file mode 100644 index 0000000000..df258fab0c --- /dev/null +++ b/packages/todo/README.md @@ -0,0 +1,9 @@ +# todo/ — todo / planning capability family + +The model-facing todo tool. A single **product** package — there is no interface/implementation seam here, because the list is single-owner session state (one agent session owns its own list), not a swappable capability. + +| Package | Role | ctx key | +|---|---|---| +| `tool-todo/` | Model-facing `todo_write` tool; writes the whole list to the session log (`todo/write`) | (registers on `ctx.tools`) | + +The list lives on the event-sourced session log (`SessionEventMap['todo/write']`, owned by [`dsh-session`](../core/session)); this package is the thin consumer that appends the snapshot. UIs render off `session/event`: the [stdio UI](../support/ui-stdio) prints the list, the [ACP bridge](../ui/acp) maps it to a `plan` sessionUpdate. diff --git a/packages/todo/tool-todo/README.md b/packages/todo/tool-todo/README.md new file mode 100644 index 0000000000..fc6dc94860 --- /dev/null +++ b/packages/todo/tool-todo/README.md @@ -0,0 +1,25 @@ +# @deepseek-ai/dsh-tool-todo + +The model-facing `todo_write` tool: the agent's whole task list, replaced wholesale on each call. + +## What it does + +Registers one tool, `todo_write(todos: [{ content, status }])`, on `ctx.tools`. The model sends the ENTIRE list every call — there are no partial updates or per-item edits. Each call appends a `todo/write` event (the full list snapshot) to the calling agent's session log via `agent.session.append('todo/write', { todos })`; the current list is the most recent such event (last-write-wins on replay). + +`status` is one of `pending`, `in_progress`, `completed` — exactly the ACP `PlanEntryStatus` triple. + +## Single owner + +The list belongs to the ONE agent session that called the tool. There is no subagent/shared/swarm scope: a non-agent caller (no `exec.agent`) has nowhere to write the list and is rejected. This is a deliberate scope limit — see the RFC. + +## Validation + +Beyond the schema's type/required/enum checks, `execute` rejects an empty or duplicate `content` and more than one `in_progress` task (a coherent plan has at most one task active). Ordering and the discipline of keeping the list current are left to the model via the tool description. + +## Rendering + +The tool writes only the session event; it does not render. UIs subscribe to `session/event` and render the `todo/write` data themselves: the [stdio UI](../../support/ui-stdio) prints a glyphed checklist, and the [ACP bridge](../../ui/acp) maps the list to a `plan` sessionUpdate (synthesizing the `priority` ACP requires). + +## Export shape + +A function/namespace plugin: it exports `name` / `inject` / `apply` and NO default. A stray `export default` would collapse the module via the Loader's `unwrapExports` and drop `inject` (see [docs/postmortem/0001](../../../docs/postmortem/0001-acp-default-export-drops-inject.md)). diff --git a/packages/todo/tool-todo/package.json b/packages/todo/tool-todo/package.json new file mode 100644 index 0000000000..f2d4344f99 --- /dev/null +++ b/packages/todo/tool-todo/package.json @@ -0,0 +1,39 @@ +{ + "name": "@deepseek-ai/dsh-tool-todo", + "description": "Model-facing todo_write tool over the DeepSeek Harness event-sourced session log", + "version": "0.0.1", + "private": true, + "type": "module", + "main": "lib/index.js", + "types": "lib/types/index.d.ts", + "exports": { + ".": { + "types": "./lib/types/index.d.ts", + "default": "./lib/index.js" + }, + "./src/*": "./src/*", + "./package.json": "./package.json" + }, + "files": [ + "lib/index.js", + "lib/types/**/*.d.ts", + "lib/types/**/*.d.ts.map", + "src" + ], + "license": "BSD-3-Clause", + "peerDependencies": { + "@deepseek-ai/dsh-agent": "^0.0.1", + "@deepseek-ai/dsh-session": "^0.0.1", + "@deepseek-ai/dsh-tools": "^0.0.1", + "cordis": "^4.0.0-rc.6" + }, + "devDependencies": { + "@deepseek-ai/dsh-agent": "workspace:^", + "@deepseek-ai/dsh-agent-loop": "workspace:^", + "@deepseek-ai/dsh-llm": "workspace:^", + "@deepseek-ai/dsh-session": "workspace:^", + "@deepseek-ai/dsh-system-prompt": "workspace:^", + "@deepseek-ai/dsh-tools": "workspace:^", + "cordis": "^4.0.0-rc.6" + } +} diff --git a/packages/todo/tool-todo/src/index.ts b/packages/todo/tool-todo/src/index.ts new file mode 100644 index 0000000000..912e08f269 --- /dev/null +++ b/packages/todo/tool-todo/src/index.ts @@ -0,0 +1,123 @@ +/** + * The model-facing `todo_write` tool: the agent's whole task list, replaced + * wholesale on each call. Every call appends a `todo/write` event (the full + * list snapshot) to the calling agent's session log via + * `exec.agent.session.append('todo/write', { todos })`; the current list is the + * most recent such event (last-write-wins on replay). UIs render off + * `session/event`: the stdio UI prints the checklist, the ACP bridge maps it to + * a `plan` sessionUpdate. + * + * Single owner: the list belongs to the ONE agent session that called the tool. + * There is no subagent/shared/swarm scope — a non-agent caller (no + * `exec.agent`) has nowhere to write the list and is rejected. + * + * Plugin export shape: named exports, NO default. The cordis Loader's + * `unwrapExports` does `exports.default ?? exports`, so a stray default would + * collapse the module to the bare `apply` and drop `inject`, crashing at load + * (see docs/postmortem/0001). + * + * @module @deepseek-ai/dsh-tool-todo + */ + +import type { Context } from 'cordis' +import type { ContentBlock } from '@deepseek-ai/dsh-llm' +import { defineTool } from '@deepseek-ai/dsh-tools' +import type { TodoItem } from '@deepseek-ai/dsh-session' + +export const name = 'tool-todo' +export const inject = ['tools'] + +/** The valid {@link TodoItem} statuses, as a runtime set for input narrowing. */ +const STATUSES = ['pending', 'in_progress', 'completed'] as const + +const DESCRIPTION = + 'Record and update a structured task list for the current work. Send the ENTIRE ' + + 'list every call — it REPLACES the previous list (there are no partial updates, ' + + 'no per-item edits). Use it to plan multi-step work and show progress: add one ' + + 'todo per concrete step before you start. Keep AT MOST ONE todo `in_progress` ' + + 'at a time; while work remains, exactly one active task should be ' + + '`in_progress`. Mark a todo `completed` the moment it is done (do not batch ' + + 'completions), and allow no `in_progress` item only once all work is complete. ' + + 'Skip the list for trivial single-step tasks. Statuses: `pending` ' + + '(not started), `in_progress` (being worked on now), `completed` (finished).' + +/** + * Validate the value constraints the SchemaSpec can't express and build the + * canonical {@link TodoItem}[]. + * + * `defineTool` already validates type/required/enum before `execute` runs (a + * bad `status` is rejected by the registry's `validateArgs`, never reaching + * here), so `status` is guaranteed to be one of the three enum literals. But + * `InferArgs` maps an `enum` string prop to plain `string`, so the compiler sees + * `args.todos` as `{ content: string; status: string }[]`; the + * `status as TodoItem['status']` narrowing records that registry guarantee + * rather than re-checking it (an unreachable re-check would be dead code — see + * AGENTS.md "don't validate scenarios that can't happen"). What remains is the + * value rules the DSL has no vocabulary for: non-empty unique content (stored + * trimmed, so the persisted value matches the dedupe/length key), and at most + * one `in_progress` task. + */ +function toTodoList(raw: { content: string; status: string }[]): TodoItem[] { + const todos: TodoItem[] = [] + const seen = new Set() + let inProgress = 0 + for (const item of raw) { + const content = item.content.trim() + if (content.length === 0) { + throw new Error('invalid todo: `content` must be a non-empty string') + } + if (seen.has(content)) { + throw new Error(`invalid todos: duplicate content ${JSON.stringify(content)}`) + } + seen.add(content) + const status = item.status as TodoItem['status'] + if (status === 'in_progress') inProgress++ + todos.push({ content, status }) + } + if (inProgress > 1) { + throw new Error(`invalid todos: at most one task may be in_progress, got ${inProgress}`) + } + return todos +} + +/** Register the `todo_write` tool on `ctx.tools`. */ +export function apply(ctx: Context): void { + ctx.tools.register(defineTool({ + name: 'todo_write', + description: DESCRIPTION, + parameters: { + todos: { + type: 'array', + required: true, + description: 'The COMPLETE task list, replacing any previous list.', + items: { + type: 'object', + properties: { + content: { type: 'string', required: true, description: 'What the task is — a short imperative line.' }, + status: { + type: 'string', + required: true, + enum: [...STATUSES], + description: 'pending (not started) | in_progress (now) | completed (done).', + }, + }, + }, + }, + }, + execute(args, exec): Promise { + const todos = toTodoList(args.todos) + if (!exec.agent) { + // The list is per-agent-session state; a non-agent caller (no owning + // session) has nowhere to write it. Reject rather than silently no-op. + throw new Error('todo_write requires an owning agent session') + } + exec.agent.session.append('todo/write', { todos }) + const count = (status: TodoItem['status']): number => todos.filter(t => t.status === status).length + return Promise.resolve([{ + type: 'text', + text: `Updated todo list: ${count('pending')} pending, ${count('in_progress')} in progress, ${count('completed')} completed.`, + }]) + }, + presentCall: args => ({ title: 'Update todo list', kind: 'other', rawInput: args.todos }), + })) +} diff --git a/packages/todo/tool-todo/tests/integration.spec.ts b/packages/todo/tool-todo/tests/integration.spec.ts new file mode 100644 index 0000000000..739367a699 --- /dev/null +++ b/packages/todo/tool-todo/tests/integration.spec.ts @@ -0,0 +1,107 @@ +import { describe, expect, it } from 'vitest' +import { Context } from 'cordis' +import LlmService from '@deepseek-ai/dsh-llm' +import SessionStore from '@deepseek-ai/dsh-session' +import type { SessionEvent } from '@deepseek-ai/dsh-session' +import SystemPrompt from '@deepseek-ai/dsh-system-prompt' +import ToolRegistry from '@deepseek-ai/dsh-tools' +import AgentRegistry, { AgentId } from '@deepseek-ai/dsh-agent' +import AgentLoop, { ReactLoopAgent } from '@deepseek-ai/dsh-agent-loop' +import * as ToolTodo from '@deepseek-ai/dsh-tool-todo' +import { MockAdapter, textResponse, toolCallResponse } from '../../../core/agent-loop/tests/mock-adapter.ts' + +/** + * Full-loop integration: a scripted mock model drives the REAL todo_write tool + * through the agent loop, exercising the same seams a live model would — the + * tool/call + tool/result session events AND the todo/write event the tool + * appends. Only the model is mocked; the tool and the session log are real. + */ +async function harness(adapter: MockAdapter): Promise { + const ctx = new Context() + await ctx.plugin(LlmService) + await ctx.plugin(SessionStore) + await ctx.plugin(SystemPrompt) + await ctx.plugin(ToolRegistry) + await ctx.plugin(AgentRegistry) + await ctx.plugin(AgentLoop, { agents: [] }) + await ctx.plugin(ToolTodo) + ctx.llm.registerAdapter(['mock'], adapter) + return ctx +} + +function waitForIdle(ctx: Context, agent: ReactLoopAgent): Promise { + return new Promise((resolve) => { + const dispose = ctx.on('agent/status', (subject, status) => { + if (subject === agent && status === 'idle') { + dispose() + resolve() + } + }) + }) +} + +function findEvent( + log: readonly SessionEvent[], + type: T, + position: 'first' | 'last' = 'first', +): Extract { + const found = position === 'first' + ? log.find(event => event.type === type) + : log.findLast(event => event.type === type) + if (!found) throw new Error(`no ${type} event in the session log`) + return found as Extract +} + +describe('todo_write tool through the agent loop', () => { + it('model calls todo_write: a tool/call, a non-error tool/result, and a todo/write snapshot land', async () => { + const adapter = new MockAdapter([ + toolCallResponse('call-1', 'todo_write', { + todos: [ + { content: 'read the code', status: 'in_progress' }, + { content: 'write the fix', status: 'pending' }, + ], + }, 'Planning the work.'), + textResponse('Plan recorded.'), + ]) + const ctx = await harness(adapter) + const agent = ctx.agentLoop.create(AgentId('it-todo'), { model: 'mock' }) + + agent.send([{ type: 'text', text: 'plan a two-step task' }]) + await waitForIdle(ctx, agent) + + const log = agent.session.events + expect(findEvent(log, 'tool/call').data.name).toBe('todo_write') + expect(findEvent(log, 'tool/result').data.isError).toBe(false) + + const todoEvent = findEvent(log, 'todo/write') + expect(todoEvent.data.todos).toEqual([ + { content: 'read the code', status: 'in_progress' }, + { content: 'write the fix', status: 'pending' }, + ]) + }) + + it('a second todo_write replaces the list (last-write-wins on the log)', async () => { + const adapter = new MockAdapter([ + toolCallResponse('call-1', 'todo_write', { todos: [{ content: 'step one', status: 'in_progress' }] }), + toolCallResponse('call-2', 'todo_write', { + todos: [ + { content: 'step one', status: 'completed' }, + { content: 'step two', status: 'in_progress' }, + ], + }), + textResponse('Done planning.'), + ]) + const ctx = await harness(adapter) + const agent = ctx.agentLoop.create(AgentId('it-todo-2'), { model: 'mock' }) + + agent.send([{ type: 'text', text: 'plan then update' }]) + await waitForIdle(ctx, agent) + + const todoEvents = agent.session.events.filter(e => e.type === 'todo/write') + expect(todoEvents).toHaveLength(2) + expect(findEvent(agent.session.events, 'todo/write', 'last').data.todos).toEqual([ + { content: 'step one', status: 'completed' }, + { content: 'step two', status: 'in_progress' }, + ]) + }) +}) diff --git a/packages/todo/tool-todo/tests/tool-todo.spec.ts b/packages/todo/tool-todo/tests/tool-todo.spec.ts new file mode 100644 index 0000000000..8b1c504840 --- /dev/null +++ b/packages/todo/tool-todo/tests/tool-todo.spec.ts @@ -0,0 +1,167 @@ +import { describe, expect, it } from 'vitest' +import { Context } from 'cordis' +import Loader from '@cordisjs/plugin-loader' +import { CallId } from '@deepseek-ai/dsh-llm' +import SystemPrompt from '@deepseek-ai/dsh-system-prompt' +import ToolRegistry from '@deepseek-ai/dsh-tools' +import { Session, SessionId } from '@deepseek-ai/dsh-session' +import type { TodoItem } from '@deepseek-ai/dsh-session' +import { AgentId, type Agent } from '@deepseek-ai/dsh-agent' +import * as tool from '../src/index.ts' + +/** + * Drives the REAL plugin body: mounts `dsh-tool-todo` on a real `ToolRegistry` + * and invokes the registered `todo_write` tool through `ctx.tools.execute`, + * with a fake parent Agent carrying a real `Session` — so the append the tool + * makes is observable on a genuine session log (only the agent wrapper is a + * stand-in; the session and the tool are the shipping code). + */ + +/** A parent Agent backed by a real Session — the tool reads `agent.session`. */ +function agentWithSession(id = 'parent-1'): Agent & { session: Session } { + const session = new Session(SessionId(id)) + return { id: AgentId(id), session } as unknown as Agent & { session: Session } +} + +async function setup(): Promise { + const ctx = new Context() + await ctx.plugin(SystemPrompt) + await ctx.plugin(ToolRegistry) + await ctx.plugin(tool) + return ctx +} + +let callCounter = 0 +function callTodo(ctx: Context, args: unknown, over: { agent?: Agent | undefined } = {}) { + const agent = 'agent' in over ? over.agent : agentWithSession() + return ctx.tools.execute({ + callId: CallId(`call-${++callCounter}`), + name: 'todo_write', + arguments: args, + ...agent ? { agent } : {}, + }) +} + +function text(result: { content: { type: string; text?: string }[] }): string { + return result.content.filter(b => b.type === 'text').map(b => b.text).join('') +} + +describe('dsh-tool-todo', () => { + it('registers a `todo_write` tool whose schema is an array of {content,status}', async () => { + const ctx = await setup() + const schema = ctx.tools.schemas().find(s => s.name === 'todo_write') + expect(schema).toBeDefined() + const props = (schema!.parameters as { properties?: Record }).properties ?? {} + expect(Object.keys(props)).toEqual(['todos']) + const todos = props.todos as { type: string; items?: { properties?: Record } } + expect(todos.type).toBe('array') + const itemProps = todos.items?.properties ?? {} + expect(Object.keys(itemProps).sort()).toEqual(['content', 'status']) + expect(itemProps.status?.enum).toEqual(['pending', 'in_progress', 'completed']) + }) + + it('appends a todo/write event carrying the whole list to the calling session', async () => { + const ctx = await setup() + const agent = agentWithSession('writer') + const todos: TodoItem[] = [ + { content: 'plan', status: 'in_progress' }, + { content: 'build', status: 'pending' }, + ] + const result = await callTodo(ctx, { todos }, { agent }) + expect(result.isError).toBe(false) + expect(text(result)).toContain('1 pending, 1 in progress, 0 completed') + + const event = agent.session.events.findLast(e => e.type === 'todo/write')! + expect(event.data.todos).toEqual(todos) + }) + + it('stores the trimmed content (the dedupe/length key), not the raw input', async () => { + const ctx = await setup() + const agent = agentWithSession('trim') + const result = await callTodo(ctx, { todos: [{ content: ' plan the work ', status: 'pending' }] }, { agent }) + expect(result.isError).toBe(false) + + const event = agent.session.events.findLast(e => e.type === 'todo/write')! + expect(event.data.todos).toEqual([{ content: 'plan the work', status: 'pending' }]) + }) + + it('replaces the list on a second call (last-write-wins on the log)', async () => { + const ctx = await setup() + const agent = agentWithSession('writer-2') + await callTodo(ctx, { todos: [{ content: 'a', status: 'pending' }] }, { agent }) + await callTodo(ctx, { todos: [ + { content: 'a', status: 'completed' }, + { content: 'b', status: 'in_progress' }, + ] }, { agent }) + + const current = agent.session.events.findLast(e => e.type === 'todo/write')!.data.todos + expect(current).toEqual([ + { content: 'a', status: 'completed' }, + { content: 'b', status: 'in_progress' }, + ]) + }) + + it('rejects a malformed status before execute runs (registry arg-validation)', async () => { + const ctx = await setup() + const result = await callTodo(ctx, { todos: [{ content: 'x', status: 'doing' }] }) + expect(result.isError).toBe(true) + }) + + it('rejects a non-array todos argument', async () => { + const ctx = await setup() + const result = await callTodo(ctx, { todos: 'nope' }) + expect(result.isError).toBe(true) + }) + + it.each([ + { label: 'empty content', todos: [{ content: ' ', status: 'pending' }], fragment: 'non-empty' }, + { label: 'duplicate content', todos: [{ content: 'dup', status: 'pending' }, { content: 'dup', status: 'completed' }], fragment: 'duplicate' }, + { label: 'two in_progress', todos: [{ content: 'a', status: 'in_progress' }, { content: 'b', status: 'in_progress' }], fragment: 'in_progress' }, + ])('rejects $label as an isError result', async ({ todos, fragment }) => { + const ctx = await setup() + const result = await callTodo(ctx, { todos }) + expect(result.isError).toBe(true) + expect(text(result)).toContain(fragment) + }) + + it('rejects a non-agent caller (the list has no owning session)', async () => { + const ctx = await setup() + const result = await callTodo(ctx, { todos: [{ content: 'a', status: 'pending' }] }, { agent: undefined }) + expect(result.isError).toBe(true) + expect(text(result)).toContain('owning agent session') + }) + + it('presents the call with a stable title and the list as raw input', async () => { + const ctx = await setup() + const def = ctx.tools.get('todo_write')! + const todos = [{ content: 'a', status: 'pending' }] + expect(def.presentCall?.({ todos })).toEqual({ title: 'Update todo list', kind: 'other', rawInput: todos }) + }) + + it('unregisters the tool when its contributing fiber is disposed (HMR-safety)', async () => { + const ctx = new Context() + await ctx.plugin(SystemPrompt) + await ctx.plugin(ToolRegistry) + const fiber = await ctx.plugin(tool) + expect(ctx.tools.schemas().some(s => s.name === 'todo_write')).toBe(true) + await fiber.dispose() + expect(ctx.tools.schemas().some(s => s.name === 'todo_write')).toBe(false) + }) + + it('has the namespace-plugin export shape (no stray default) so the Loader keeps name/inject/apply', () => { + // Postmortem 0001 guard: this plugin HAS `inject = ['tools']`, so a stray + // `export default apply` would collapse the module via `unwrapExports` + // (`exports.default ?? exports`), DROP `inject`, and crash at load with + // "cannot get property … without inject". Guard the shape directly. + expect('default' in tool).toBe(false) + expect(tool.name).toBe('tool-todo') + expect(tool.inject).toEqual(['tools']) + + const loader = Object.create(Loader.prototype) as Loader + const unwrapped = loader.unwrapExports(tool) as Record + expect(unwrapped).toBe(tool) + expect(unwrapped.name).toBe('tool-todo') + expect(unwrapped.inject).toEqual(['tools']) + expect(typeof unwrapped.apply).toBe('function') + }) +}) diff --git a/packages/todo/tool-todo/tsconfig.json b/packages/todo/tool-todo/tsconfig.json new file mode 100644 index 0000000000..adf2f25dec --- /dev/null +++ b/packages/todo/tool-todo/tsconfig.json @@ -0,0 +1,27 @@ +{ + "extends": "../../../tsconfig.base.json", + "compilerOptions": { + "rootDir": "src", + "outDir": "lib/types" + }, + "include": [ + "src" + ], + "references": [ + { + "path": "../../../vendor/cosmokit" + }, + { + "path": "../../../vendor/cordis" + }, + { + "path": "../../core/tools" + }, + { + "path": "../../core/agent" + }, + { + "path": "../../core/session" + } + ] +} diff --git a/packages/ui/README.md b/packages/ui/README.md index 075dfd524d..659519407c 100644 --- a/packages/ui/README.md +++ b/packages/ui/README.md @@ -10,4 +10,4 @@ Integrations that expose the agent to an external editor or client. These are ** A UI integration is a client-driver plugin, not a loop change and not a capability seam: it consumes the existing `agent/*` event taxonomy and the `dsh-agent` factory. The readline `ui-stdio` plugin is the unstructured analogue but lives in `support/` because it exists chiefly for the examples and the coverage gate — `ui/` is reserved for surfaces shipped as product. -`stdio-agent` and `acp-agent` are the two **app packages**: each composes the [`core/agent-core`](../core/agent-core/README.md) spine with its coupled front-door cluster (and owns the boot `bin`), so a leaf `cordis.yml` is just the swappable backends plus one app entry. They live in `ui/` because each IS a user-facing front door; the stdout-purity coupling (logger vs. no logger) becomes a property of the artifact rather than a leaf convention. +`stdio-agent` and `acp-agent` are the two **app packages**: each composes the [`core/agent-core`](../core/agent-core/README.md) spine with its coupled front-door cluster (and owns the boot `bin`), so a leaf `cordis.yml` is the swappable backends plus one app entry plus any optional product tools. They live in `ui/` because each IS a user-facing front door; the stdout-purity coupling (logger vs. no logger) becomes a property of the artifact rather than a leaf convention. diff --git a/packages/ui/acp-agent/src/index.ts b/packages/ui/acp-agent/src/index.ts index 625467cac2..c505e9bf41 100644 --- a/packages/ui/acp-agent/src/index.ts +++ b/packages/ui/acp-agent/src/index.ts @@ -14,11 +14,12 @@ * which this app does not prevent — so the rule "never add a stdout logger to an * ACP leaf" still stands; the app just gives the leaf nothing to misconfigure.) * - * The leaf supplies only the swappable backends: the LLM adapter (`llm-deepseek` - * for the real model, `llm-replay` for keyless snapshot replay) and the bash - * executor (`bash-local`). This app's {@link Config} (model, system prompt, - * persistence root) routes each value to where it is wired — model/prompt onto - * the bridge's per-session agent template, the root onto the JSONL backend. + * The leaf supplies the swappable backends: the LLM adapter (`llm-deepseek` for + * the real model, `llm-replay` for keyless snapshot replay), the bash executor + * (`bash-local`), and any optional product tools it wants to expose. This app's + * {@link Config} (model, system prompt, persistence root) routes each value to + * where it is wired — model/prompt onto the bridge's per-session agent + * template, the root onto the JSONL backend. * * Plugin export shape: named `name`/`Config`/`apply`, NO default export — the * cordis Loader's `unwrapExports` does `exports.default ?? exports`, so a stray diff --git a/packages/ui/acp/package.json b/packages/ui/acp/package.json index 6973dc5e20..52080be092 100644 --- a/packages/ui/acp/package.json +++ b/packages/ui/acp/package.json @@ -44,6 +44,7 @@ "@deepseek-ai/dsh-session-persistence-jsonl": "workspace:^", "@deepseek-ai/dsh-system-prompt": "workspace:^", "@deepseek-ai/dsh-tool-bash": "workspace:^", + "@deepseek-ai/dsh-tool-todo": "workspace:^", "@deepseek-ai/dsh-tools": "workspace:^", "cordis": "^4.0.0-rc.6" } diff --git a/packages/ui/acp/src/index.ts b/packages/ui/acp/src/index.ts index ec79e97443..754b1750be 100644 --- a/packages/ui/acp/src/index.ts +++ b/packages/ui/acp/src/index.ts @@ -53,6 +53,8 @@ import { type LoadSessionResponse, type NewSessionRequest, type NewSessionResponse, + type Plan, + type PlanEntry, type PromptRequest, type PromptResponse, type SessionNotification, @@ -64,7 +66,7 @@ import { CallId } from '@deepseek-ai/dsh-llm' import type { Agent, AgentStatus } from '@deepseek-ai/dsh-agent' import { AgentId } from '@deepseek-ai/dsh-agent' import { SessionId } from '@deepseek-ai/dsh-session' -import type { SessionEvent, TurnEndReason } from '@deepseek-ai/dsh-session' +import type { SessionEvent, TodoItem, TurnEndReason } from '@deepseek-ai/dsh-session' import type { ToolCallKind, ToolCallPresentation, ToolRegistry, ToolResultPresentation, ToolTerminal } from '@deepseek-ai/dsh-tools' // Side-effect type import: declaration-merges `ctx.sessionPersistence` onto // Context (the bridge injects it and reads `list()` for load cwd validation). @@ -873,6 +875,10 @@ export function streamSessionEventUpdate( }) return } + case 'todo/write': { + notify({ sessionId, update: { sessionUpdate: 'plan', ...todosToPlan(event.data.todos) } }) + return + } // turn/step boundaries, context/message, steering, // assistant/message — no direct ACP client update. default: @@ -880,6 +886,18 @@ export function streamSessionEventUpdate( } } +/** + * Map a harness todo list to an ACP `plan` body. ACP's `PlanEntry` requires + * `content` + `priority` + `status`, but a {@link TodoItem} carries no priority, + * so synthesize a constant `'medium'` on every entry; `status` maps 1:1 (the + * harness status triple IS `PlanEntryStatus`). The ACP client REPLACES its whole + * plan on each `plan` update, matching the harness's whole-list-replace + * semantics, so no per-entry diffing is needed. + */ +export function todosToPlan(todos: TodoItem[]): Plan { + return { entries: todos.map((todo): PlanEntry => ({ content: todo.content, priority: 'medium', status: todo.status })) } +} + /** * Per-connection terminal-rendering context threaded into * {@link streamSessionEventUpdate}: whether the client advertised the diff --git a/packages/ui/acp/tests/harness.ts b/packages/ui/acp/tests/harness.ts index 4f6b5ac17a..9077d106d0 100644 --- a/packages/ui/acp/tests/harness.ts +++ b/packages/ui/acp/tests/harness.ts @@ -20,6 +20,7 @@ import AgentLoop from '@deepseek-ai/dsh-agent-loop' import SessionPersistenceJsonl from '@deepseek-ai/dsh-session-persistence-jsonl' import { LocalBashExecutor } from '@deepseek-ai/dsh-bash-local' import * as ToolBash from '@deepseek-ai/dsh-tool-bash' +import * as ToolTodo from '@deepseek-ai/dsh-tool-todo' import { ClientSideConnection, ndJsonStream, @@ -158,6 +159,12 @@ export async function makeBridgeHarness(options: { * implementation over a mock in tests"). */ withBash?: boolean + /** + * Plug the REAL `dsh-tool-todo` tool so a test can drive `todo_write` through + * the bridge and assert the resulting `plan` sessionUpdate — the shipping + * tool + the bridge's own todo/write→plan mapping, not a stand-in. + */ + withTodo?: boolean } = { storageDir: '' }): Promise { const adapter = new MockAdapter(options.script ?? []) @@ -173,6 +180,9 @@ export async function makeBridgeHarness(options: { await ctx.plugin(LocalBashExecutor, { timeoutMs: 10_000 }) await ctx.plugin(ToolBash) } + if (options.withTodo) { + await ctx.plugin(ToolTodo) + } ctx.llm.registerAdapter(['mock'], adapter) // Two identity byte pipes cross-wired into the two ndJsonStreams: bytes the diff --git a/packages/ui/acp/tests/load.spec.ts b/packages/ui/acp/tests/load.spec.ts index a87fa49a9e..c07ef93274 100644 --- a/packages/ui/acp/tests/load.spec.ts +++ b/packages/ui/acp/tests/load.spec.ts @@ -94,6 +94,44 @@ describe('acp bridge — session/load replay', () => { expect(content[0]?.content.text).toBe('```console\nhello\n```') }) + it('replays a persisted todo/write as a plan sessionUpdate on load', async () => { + // A turn whose model called todo_write persists a todo/write event. A fresh + // bridge loading the session must re-emit the ACP `plan` update from the log + // (the load replay runs every event through streamSessionEventUpdate), so an + // editor reopening the session sees the current plan. + live = await makeBridgeHarness({ + storageDir, + withTodo: true, + script: [ + toolCallResponse('c1', 'todo_write', { + todos: [ + { content: 'first step', status: 'in_progress' }, + { content: 'second step', status: 'pending' }, + ], + }), + textResponse('planned'), + ], + }) + await live.client.initialize({ protocolVersion: PROTOCOL_VERSION, clientCapabilities: {} }) + const { sessionId } = await live.client.newSession({ cwd: process.cwd(), mcpServers: [] }) + await live.client.prompt({ sessionId, prompt: [{ type: 'text', text: 'plan it' }] }) + await live.dispose() + live = undefined + + loader = await makeBridgeHarness({ storageDir, withTodo: true, script: [] }) + await loader.client.initialize({ protocolVersion: PROTOCOL_VERSION, clientCapabilities: {} }) + await loader.client.loadSession({ sessionId, cwd: process.cwd(), mcpServers: [] }) + + const plan = loader.updates.find(u => u.sessionUpdate === 'plan') + expect(plan).toEqual({ + sessionUpdate: 'plan', + entries: [ + { content: 'first step', priority: 'medium', status: 'in_progress' }, + { content: 'second step', priority: 'medium', status: 'pending' }, + ], + }) + }) + it('replays a persisted bash call as a TERMINAL card when the loader advertises the capability', async () => { // The presentation is resolved at replay time, so a loader that advertised // _meta.terminal_output must reconstruct the terminal card (content + _meta) diff --git a/packages/ui/acp/tests/stream-update.spec.ts b/packages/ui/acp/tests/stream-update.spec.ts index 30cbd17c40..89d68fd1c0 100644 --- a/packages/ui/acp/tests/stream-update.spec.ts +++ b/packages/ui/acp/tests/stream-update.spec.ts @@ -3,7 +3,7 @@ import { CallId } from '@deepseek-ai/dsh-llm' import { SessionId, type SessionEvent } from '@deepseek-ai/dsh-session' import type { SessionNotification } from '@agentclientprotocol/sdk' import type { ToolDefinition, ToolRegistry } from '@deepseek-ai/dsh-tools' -import { streamSessionEventUpdate, agentOptions, ToolPresenter } from '../src/index.ts' +import { streamSessionEventUpdate, agentOptions, todosToPlan, ToolPresenter } from '../src/index.ts' /** Collect the updates a single event produces (no presenter → generic fallback). */ function updatesFor(event: SessionEvent): SessionNotification['update'][] { @@ -118,6 +118,43 @@ describe('streamSessionEventUpdate', () => { expect(updatesFor(evt('turn/end', { turn: 1, reason: { kind: 'completed' } }))).toEqual([]) expect(updatesFor(evt('step/start', { turn: 1, step: 1 }))).toEqual([]) }) + + it('maps todo/write to a plan sessionUpdate with priority synthesized as medium', () => { + expect(updatesFor(evt('todo/write', { + todos: [ + { content: 'plan the work', status: 'in_progress' }, + { content: 'write the code', status: 'pending' }, + { content: 'run the tests', status: 'completed' }, + ], + }))).toEqual([{ + sessionUpdate: 'plan', + entries: [ + { content: 'plan the work', priority: 'medium', status: 'in_progress' }, + { content: 'write the code', priority: 'medium', status: 'pending' }, + { content: 'run the tests', priority: 'medium', status: 'completed' }, + ], + }]) + }) + + it('maps an empty todo list to a plan with no entries', () => { + expect(updatesFor(evt('todo/write', { todos: [] }))).toEqual([{ sessionUpdate: 'plan', entries: [] }]) + }) +}) + +describe('todosToPlan', () => { + it('maps status 1:1 and stamps every entry priority medium', () => { + expect(todosToPlan([ + { content: 'a', status: 'pending' }, + { content: 'b', status: 'in_progress' }, + { content: 'c', status: 'completed' }, + ])).toEqual({ + entries: [ + { content: 'a', priority: 'medium', status: 'pending' }, + { content: 'b', priority: 'medium', status: 'in_progress' }, + { content: 'c', priority: 'medium', status: 'completed' }, + ], + }) + }) }) describe('ToolPresenter (tool-owned presentation via the tool registry)', () => { diff --git a/packages/ui/stdio-agent/src/index.ts b/packages/ui/stdio-agent/src/index.ts index c4b9ed202c..df42b3115b 100644 --- a/packages/ui/stdio-agent/src/index.ts +++ b/packages/ui/stdio-agent/src/index.ts @@ -6,9 +6,10 @@ * * The cluster is BAKED IN, not left to the leaf: a stdio app always logs to the * console (stdout is just the terminal) and always pre-creates the `main` agent - * `ui-stdio` sends to. The leaf supplies only the swappable backends (the LLM - * adapter, the bash executor), the optional `hmr` dev-reload plugin, and this - * app's {@link Config} (model, prompt, persistence root, welcome banner). + * `ui-stdio` sends to. The leaf supplies the swappable backends (the LLM + * adapter, the bash executor), optional product tools, the optional `hmr` + * dev-reload plugin, and this app's {@link Config} (model, prompt, persistence + * root, welcome banner). * * `hmr` is deliberately a LEAF entry, not baked in here: it is a Loader-only, * subprocess-only dev plugin (its constructor throws without `--expose-internals` diff --git a/pnpm-lock.yaml b/pnpm-lock.yaml index d7f84bd4fc..39fdcbd628 100644 --- a/pnpm-lock.yaml +++ b/pnpm-lock.yaml @@ -130,6 +130,36 @@ importers: specifier: ^4.0.0-rc.6 version: 4.0.0-rc.6(@cordisjs/plugin-include@1.0.4)(@cordisjs/plugin-loader@1.0.0-rc.4) + packages/compact/compact-basic: + devDependencies: + '@deepseek-ai/dsh-agent': + specifier: workspace:^ + version: link:../../core/agent + '@deepseek-ai/dsh-agent-loop': + specifier: workspace:^ + version: link:../../core/agent-loop + '@deepseek-ai/dsh-compact': + specifier: workspace:^ + version: link:../compact + '@deepseek-ai/dsh-invariants': + specifier: workspace:^ + version: link:../../support/invariants + '@deepseek-ai/dsh-llm': + specifier: workspace:^ + version: link:../../llm/llm + '@deepseek-ai/dsh-session': + specifier: workspace:^ + version: link:../../core/session + '@deepseek-ai/dsh-system-prompt': + specifier: workspace:^ + version: link:../../core/system-prompt + '@deepseek-ai/dsh-tools': + specifier: workspace:^ + version: link:../../core/tools + cordis: + specifier: ^4.0.0-rc.6 + version: 4.0.0-rc.6(@cordisjs/plugin-include@1.0.4)(@cordisjs/plugin-loader@1.0.0-rc.4) + packages/core/agent: devDependencies: '@deepseek-ai/dsh-brand': @@ -664,6 +694,30 @@ importers: specifier: ^4.0.0-rc.6 version: 4.0.0-rc.6(@cordisjs/plugin-include@1.0.4)(@cordisjs/plugin-loader@1.0.0-rc.4) + packages/todo/tool-todo: + devDependencies: + '@deepseek-ai/dsh-agent': + specifier: workspace:^ + version: link:../../core/agent + '@deepseek-ai/dsh-agent-loop': + specifier: workspace:^ + version: link:../../core/agent-loop + '@deepseek-ai/dsh-llm': + specifier: workspace:^ + version: link:../../llm/llm + '@deepseek-ai/dsh-session': + specifier: workspace:^ + version: link:../../core/session + '@deepseek-ai/dsh-system-prompt': + specifier: workspace:^ + version: link:../../core/system-prompt + '@deepseek-ai/dsh-tools': + specifier: workspace:^ + version: link:../../core/tools + cordis: + specifier: ^4.0.0-rc.6 + version: 4.0.0-rc.6(@cordisjs/plugin-include@1.0.4)(@cordisjs/plugin-loader@1.0.0-rc.4) + packages/ui/acp: dependencies: '@agentclientprotocol/sdk': @@ -703,6 +757,9 @@ importers: '@deepseek-ai/dsh-tool-bash': specifier: workspace:^ version: link:../../bash/tool-bash + '@deepseek-ai/dsh-tool-todo': + specifier: workspace:^ + version: link:../../todo/tool-todo '@deepseek-ai/dsh-tools': specifier: workspace:^ version: link:../../core/tools diff --git a/scripts/gen-cordis-catalog.ts b/scripts/gen-cordis-catalog.ts index 4d2598155b..68933a2a13 100644 --- a/scripts/gen-cordis-catalog.ts +++ b/scripts/gen-cordis-catalog.ts @@ -23,8 +23,8 @@ * * The HARNESS tier (the `@deepseek-ai/dsh-*` events + services) is rendered in * full from source: signature, the `@mode` badge, and the declaration's JSDoc. - * Every harness event MUST carry an `@mode emit|waterfall|parallel` tag — the - * generator hard-errors on a missing tag, and where the signature shape is + * Every harness event MUST carry an `@mode emit|waterfall|parallel|serial` tag + * — the generator hard-errors on a missing tag, and where the signature shape is * conclusive (a trailing `next: () => …` parameter is structurally a waterfall) * it asserts the tag agrees and hard-errors on a contradiction. The INHERITED * tier (cordis core + loader/hmr/timer) is pinned vendor source a plugin author @@ -48,7 +48,7 @@ const OUT = 'docs/cordis-catalog/events-and-services.md' const FENCE = 'ts cordis-catalog' /** A dispatch mode, rendered as the badge after an event name. */ -type Mode = 'emit' | 'waterfall' | 'parallel' +type Mode = 'emit' | 'waterfall' | 'parallel' | 'serial' /** * Cross-link map: a type name that appears in a signature → the @@ -175,7 +175,7 @@ function parseJsDoc(raw: string): { doc: string; mode: Mode | null } { para = [] } for (const line of inner) { - const m = /^@mode\s+(emit|waterfall|parallel)\s*$/.exec(line) + const m = /^@mode\s+(emit|waterfall|parallel|serial)\s*$/.exec(line) if (m) { mode = m[1] as Mode; continue } if (line.startsWith('@')) { flushPara(); continue } // other tags end the prose if (line.trim() === '') { flushPara(); continue } @@ -233,11 +233,11 @@ export function collectEvents(scanRoot: string = root): EventEntry[] { const { doc, mode } = parseJsDoc(rawJsDoc(text, member)) const src = pointer(rel, sf, member) if (!mode) { - throw new Error(`gen-cordis-catalog: event '${name}' (${src}) is missing an @mode tag. Add '@mode emit|waterfall|parallel' to its JSDoc (see AGENTS.md).`) + throw new Error(`gen-cordis-catalog: event '${name}' (${src}) is missing an @mode tag. Add '@mode emit|waterfall|parallel|serial' to its JSDoc (see AGENTS.md).`) } // Conclusive structural check: a trailing `next: () => …` parameter is a - // waterfall. (emit vs parallel is not structurally distinguishable, so - // it is trusted from the tag.) + // waterfall. (emit vs parallel vs serial is not structurally + // distinguishable, so it is trusted from the tag.) const last = member.parameters.at(-1) const hasNext = !!last && last.name.getText(sf) === 'next' if (hasNext && mode !== 'waterfall') { @@ -341,7 +341,7 @@ const INHERITED_EVENTS: InheritedEntry[] = [ const INHERITED_SERVICES: InheritedEntry[] = [ { name: 'ctx.on / ctx.once', summary: 'Register an event listener (disposable).', source: 'vendor/cordis/src/events.ts:29' }, - { name: 'ctx.emit / ctx.parallel / ctx.serial / ctx.bail / ctx.waterfall', summary: 'Dispatch an event (sync / awaited / first-non-nullish / veto-chain).', source: 'vendor/cordis/src/events.ts:29' }, + { name: 'ctx.emit / ctx.parallel / ctx.serial / ctx.bail / ctx.waterfall', summary: 'Dispatch an event (sync / awaited / first-bail / veto-chain).', source: 'vendor/cordis/src/events.ts:29' }, { name: 'ctx.plugin / ctx.inject', summary: 'Load a plugin / declare required services.', source: 'vendor/cordis/src/registry.ts:144' }, { name: 'ctx.effect', summary: 'Register a disposable side effect tied to the fiber.', source: 'vendor/cordis/src/fiber.ts:9' }, { name: 'ctx.get / ctx.set / ctx.provide / ctx.accessor / ctx.mixin', summary: 'Low-level service-store access and binding.', source: 'vendor/cordis/src/reflect.ts:7' }, @@ -404,7 +404,7 @@ function render(events: EventEntry[], services: ServiceEntry[]): string { '', '## Events', '', - `Dispatch modes: **emit** (fire-and-forget), **waterfall** (each listener gets \`next()\` and may transform or veto — see [waterfall semantics](../architecture.md#cordis-waterfall-semantics-important)), **parallel** (awaited fan-out, no veto). The harness declares ${events.length} events across ${new Set(events.map(e => e.scope)).size} scopes.`, + 'Dispatch modes: **emit** (fire-and-forget), **waterfall** (each listener gets `next()` and may transform or veto — see [waterfall semantics](../architecture.md#cordis-waterfall-semantics-important)), **parallel** (awaited fan-out; all listeners run), **serial** (awaited in registration order until one returns a bail value — anything other than `null`, `false`, or `undefined`).', '', ] const scopes = [...new Set(events.map(e => e.scope))].sort() @@ -417,7 +417,7 @@ function render(events: EventEntry[], services: ServiceEntry[]): string { lines.push( '## Services', '', - `The ${services.length} \`ctx.\` services the harness provides. An abstract seam (e.g. \`ctx.bash\`) is implemented by a separate package; the interface is what consumers code against.`, + 'The `ctx.` services the harness provides. An abstract seam (e.g. `ctx.bash`) is implemented by a separate package; the interface is what consumers code against.', '', ) for (const s of services) lines.push(...renderService(s)) diff --git a/scripts/type-equiv.manifest.json b/scripts/type-equiv.manifest.json index 68352a111b..6d584bc6c6 100644 --- a/scripts/type-equiv.manifest.json +++ b/scripts/type-equiv.manifest.json @@ -16,6 +16,7 @@ { "doc": "docs/core-data-structures/llm-streaming.md", "symbol": "ContentBlockMap", "source": "packages/llm/llm/src/types.ts" }, { "doc": "docs/core-data-structures/session.md", "symbol": "SessionEventMap", "source": "packages/core/session/src/types.ts" }, + { "doc": "docs/core-data-structures/session.md", "symbol": "TodoItem", "source": "packages/core/session/src/types.ts" }, { "doc": "docs/core-data-structures/session.md", "symbol": "SessionEvent", "source": "packages/core/session/src/types.ts" }, { "doc": "docs/core-data-structures/session.md", "symbol": "TurnTriggerMap", "source": "packages/core/session/src/types.ts" }, { "doc": "docs/core-data-structures/session.md", "symbol": "TurnEndReasonMap", "source": "packages/core/session/src/types.ts" }, diff --git a/scripts/verify-md-links.ts b/scripts/verify-md-links.ts index 57bb824b32..cbd913d5ca 100644 --- a/scripts/verify-md-links.ts +++ b/scripts/verify-md-links.ts @@ -20,12 +20,13 @@ * resolved against the linking file's directory, and the result must exist on * disk. This is checker, not fixer: it reports and never rewrites. * - * Scope is the other doc-sync gates' set plus the two AGENTS.md files AND the - * repo-authored agent-skill Markdown under `.agents/skills/` — those skill - * files cross-link into the docs tree (e.g. the dsh-code-review skill cites the - * RFC index), so a rename must not silently break them either: README.md, - * docs/** /*.md, packages/* /README.md, AGENTS.md, packages/AGENTS.md, - * .agents/skills/** /*.md. The root and packages/ CLAUDE.md are symlinks to the + * Scope is the other doc-sync gates' set plus example Markdown, AGENTS.md + * files in those checked trees, AND the repo-authored agent-skill Markdown under + * `.agents/skills/` — those skill files cross-link into the docs tree (e.g. the + * dsh-code-review skill cites the RFC index), so a rename must not silently + * break them either: README.md, docs/** /*.md, packages/* /README.md, + * examples/** /*.md, AGENTS.md, packages/AGENTS.md, .agents/skills/** /*.md. + * The root, packages/, and examples/ CLAUDE.md files are symlinks to the * AGENTS.md files, so they are deduped by real path. * * Run: `tsx scripts/verify-md-links.ts`. @@ -42,14 +43,15 @@ import type { Nodes } from 'mdast' const root = resolve(import.meta.dirname, '..') /** - * Files to check: doc-typecheck's scope, the AGENTS.md pair, and repo-authored - * agent-skill Markdown (which this repo's own docs reorg rewrites links in). + * Files to check: doc-typecheck's scope, example Markdown, the AGENTS.md pair, + * and repo-authored agent-skill Markdown. */ const PATTERNS = [ 'README.md', 'docs/**/*.md', 'packages/*/*.md', 'packages/*/*/*.md', + 'examples/**/*.md', 'AGENTS.md', 'packages/AGENTS.md', '.agents/skills/**/*.md', diff --git a/tsconfig.base.json b/tsconfig.base.json index 7ac09c1725..e9e1112823 100644 --- a/tsconfig.base.json +++ b/tsconfig.base.json @@ -46,6 +46,7 @@ "./packages/fs/*/src", "./packages/compact/*/src", "./packages/subagent/*/src", + "./packages/todo/*/src", "./packages/session-persistence/*/src", "./packages/ui/*/src", "./packages/util/*/src", diff --git a/tsconfig.build.json b/tsconfig.build.json index 106aa40883..a3334bb397 100644 --- a/tsconfig.build.json +++ b/tsconfig.build.json @@ -23,6 +23,7 @@ { "path": "./packages/core/agent-core" }, { "path": "./packages/bash/bash" }, { "path": "./packages/compact/compact" }, + { "path": "./packages/compact/compact-basic" }, { "path": "./packages/llm/llm-deepseek" }, { "path": "./packages/llm/llm-pi-ai" }, { "path": "./packages/bash/bash-local" }, @@ -43,6 +44,7 @@ { "path": "./packages/subagent/subagent-inprocess" }, { "path": "./packages/subagent/subagent-spawn" }, { "path": "./packages/subagent/subagent-fork" }, - { "path": "./packages/subagent/subagent-acp" } + { "path": "./packages/subagent/subagent-acp" }, + { "path": "./packages/todo/tool-todo" } ] } diff --git a/tsconfig.json b/tsconfig.json index 3ea12d7c6a..a192e9319e 100644 --- a/tsconfig.json +++ b/tsconfig.json @@ -42,6 +42,7 @@ { "path": "./packages/fs/file-context" }, { "path": "./packages/fs/tool-fs" }, { "path": "./packages/compact/compact" }, + { "path": "./packages/compact/compact-basic" }, { "path": "./packages/support/invariants" }, { "path": "./packages/ui/acp" }, { "path": "./packages/ui/acp-agent" }, @@ -54,6 +55,7 @@ { "path": "./packages/subagent/subagent-inprocess" }, { "path": "./packages/subagent/subagent-spawn" }, { "path": "./packages/subagent/subagent-fork" }, - { "path": "./packages/subagent/subagent-acp" } + { "path": "./packages/subagent/subagent-acp" }, + { "path": "./packages/todo/tool-todo" } ] }