The LLM service exposed three call surfaces (stream/streamBlocks/generate) but the only production consumer — the agent loop — uses stream() exclusively, feeding raw chunks through its own BlockAssembler for replay fidelity. Drop the speculative convenience surfaces and the registry-change event that no listener consumed, leaving stream() as the single model-call contract for both production and tests. - Remove LlmService.streamBlocks() and generate(), the llm/generate waterfall, and GenerateResult. - Remove the llm/adapter-change event (declaration + emits) and the listener-throw rollback ordering that existed only to protect it; keep the HMR rollback disposer. - Remove BlockAssembler.flushReady()/flushRemaining()/result() and the flushed cursor — the streaming-flush slice existed only for streamBlocks(). - Adapter tests drive a stream()+BlockAssembler helper (tests/assemble.ts) instead of generate(), exercising the same path production uses. - Land the AGENTS.md "RFCs are proposals, not golden truth" principle and move both RFCs proposed -> implemented. Implements: - docs/rfc/implemented/simplification/2026-06-20-drop-unconsumed-llm-adapter-change-event.md - docs/rfc/implemented/simplification/2026-06-20-drop-unconsumed-llm-assembled-surfaces.md
3.3 KiB
RFC: Property-based testing for protocol-shaped code
Status: implemented (proposed 2026-06-11, accepted 2026-06-14)
Merges the original proposal and the decision record for one topic. It found a real BlockAssembler duplicate-
block-endbug on first run.
Context
Example-based tests pin the cases we thought of. The harness's core is protocol-shaped — chunk streams, event logs, schema conversion, inbox scheduling — where the input space is combinatorial and the interesting bugs live in interleavings nobody wrote an example for. The motivating evidence: a block-assembly ordering bug once survived 100% line coverage of the happy paths. Per-file 100% coverage proves every line ran, not that every interleaving is correct.
Decision
Adopt fast-check (a root devDependency) with one tests/properties.spec.ts per protocol-shaped package, generators tuned for realistic-but-adversarial inputs (not uniform noise) and numRuns kept so the suite stays well under ~10s locally. Failures print a reproducible seed. (The original proposal also sketched a nightly CI job running 100× the iterations; that was not shipped — the property suite runs only in the normal push/pull_request CI, and a scheduled high-iteration job remains possible future work.)
- dsh-llm / BlockAssembler: arbitrary chunk streams (valid + malformed: duplicate indices, stragglers, missing block-start). Invariants: the blocks
push()returns incrementally are a prefix of the finalblocks(), in order; partial count ≤ distinct indices; re-assembly idempotent; streaming and one-shot consumers agree on usage and finish. - dsh-session: arbitrary event logs. Invariants:
deriveMessagesdeterministic; replay-from-seed identical; seq strictly monotonic; non-message events never affect derived history; derived content is decoupled from the log. - dsh-tools: arbitrary
SchemaSpec. Invariants: JSON Schemarequiredequals therequired:truekeys at every level; conversion total; and the composition with runtime arg validation — generated args satisfying a spec passvalidateArgs, and targeted corruptions (dropped required key, non-object top level) are rejected. This closes the validator/InferArgsdrift risk. - dsh-agent-loop: arbitrary send schedules against a never-exhausting adapter, driven through the
agent/statussettle signal (no wall-clock sleeps). Invariants: no message lost; turn numbers strictly increase; status transitions stay on the legal machine.
Consequences
- Generator quality is the value lever — the generators bias toward small index pools and short strings so collisions and interleavings are common.
- It already paid off: the BlockAssembler stream found a real bug — a duplicate
block-endat the same index overwrote an already-flushed block, so the streamed prefix disagreed with finalblocks(). Fixed (first close wins, matching the existing straggler rule) with a dedicated regression test. - A property flake from a timeout is a finding, not something to retry away. The loop properties are deterministic by construction (settle on
agent/status), so a hang is a real defect. - Property tests supplement, not replace, the example tests that pin specific branches for the 100%-coverage gate.