Files
deepseek-harness/docs/core-data-structures/llm-streaming.md
Tianyi Cui 8dd8ce21db docs: make the service map the front door of architecture.md
The Layering ASCII diagram was cluttered, lumped THE concrete loop plugin
undifferentiated into a grab-bag plugins box, omitted dsh-agent-core and the
app packages, and restated the dependency rule that packages/README.md
already owns — while being the first thing a reader linked from README hits.

- Drop the Layering section; the tier story becomes one intro sentence and
  the dependency rule a one-line Service map footer linking
  packages/README.md#dependencies.
- Split the Service map into the spine (packages/core/) vs the swappable
  capability seams; annotate ctx.agentLoop as THE concrete loop plugin.
- Promote Cordis waterfall semantics to a top-level section placed before
  first waterfall use; drop '(important)' from the heading (anchor swept:
  AGENTS.md, gen-cordis-catalog.ts + regenerated catalog).
- Promote Event taxonomy to a top-level section (anchor unchanged).
- Retitle 'The vocabulary (dsh-llm)' to 'Content blocks and streaming
  (dsh-llm)' so the heading names its content (citation swept:
  llm-streaming.md).

Net -29 words (1851 -> 1822); the 1890 ceiling stands, keeping the
manifest's working margin.
2026-07-05 01:22:17 +08:00

5.4 KiB

LLM Streaming

The wire-level streaming vocabulary of dsh-llm. core.md introduces StreamChunk, Message, and ContentBlock; this page owns the full chunk protocol, the adapter contract every adapter must obey, and the shared assembler.

Source: packages/llm/llm/src/types.ts

StreamChunk — the raw protocol

A streaming response interleaves several typed blocks (text, reasoning, multiple tool calls). index ties each delta to its block; block-end carries the fully-assembled ContentBlock so consumers don't have to re-assemble deltas themselves. It is a closed discriminated union — a switch over type ends with assertNever, so adding a variant breaks compilation at every consumer that must handle it.

type StreamChunk =
  | { type: 'block-start'; index: number; blockType: ContentBlockType }
  | { type: 'text-delta'; index: number; text: string }
  | { type: 'reasoning-delta'; index: number; text: string }
  | { type: 'tool-call-delta'; index: number; id: CallId; name?: string; argumentsDelta: string }
  | { type: 'block-end'; index: number; block: ContentBlock }
  | { type: 'usage'; usage: TokenUsage }
  | { type: 'finish'; reason: FinishReason }

The adapter contract

Every adapter MUST obey these, and every consumer may rely on them:

  • usage before finish, nothing after finish. Defer both to the provider's end-of-stream marker so a trailing usage-only chunk can't violate the ordering.
  • Tool-call arguments stay raw JSON strings end-to-end. Partial fragments stream via argumentsDelta; a provider that hands back parsed objects re-stringifies at block-end.
  • Two sanctioned error paths. A failure may either THROW from stream() (transport/protocol errors) or end the stream with finish {kind:'error'|'aborted'} (provider in-band errors, for adapters that can't throw mid-stream). Consumers must handle both. The agent loop translates a finish-error/aborted into a turn error — it never logs a normal completed assistant message for a failed step.
  • Every provider HTTP request carries the app-attribution header. Adapters send attributionHeaders() (below) - the User-Agent baseline - and prove it with a wire-level test (mock server asserting the received header, or the library's header hook for a library-backed adapter).

This contract is why two adapters exist as a deliberate pair: dsh-llm-deepseek (hand-rolled fetch/SSE) and dsh-llm-pi-ai (the same endpoint through @earendil-works/pi-ai). Two independent internals over one contract is what pinned the protocol down — the library-backed adapter can't throw mid-stream, so it exercises the finish-chunk error path the hand-rolled one might not.

AppIdentity — app attribution

The static public application identity every adapter sends to providers (packages/llm/llm/src/attribution.ts). attributionHeaders(identity?) maps it to the standard User-Agent header only; OpenRouter-specific app attribution headers are intentionally not supported by this contract. The default APP_IDENTITY sources its version from the package manifest; every field is a public product fact - no secrets, paths, session ids, or per-user identifiers, and nothing per-request may influence the values. Rationale: Mandatory User-Agent attribution.

interface AppIdentity {
  product: string
  version: string
  url: string
}

TokenUsage

Per-call token accounting. Counts are disjoint: inputTokens is uncached input only; cached input is reported separately, and billed input is the sum of the three. Adapters whose providers fold cache hits into a single prompt total (DeepSeek's prompt_tokens) subtract them back out.

interface TokenUsage {
  inputTokens: number
  outputTokens: number
  cacheReadTokens?: number
  cacheWriteTokens?: number
  reasoningTokens?: number
}

BlockAssembler

BlockAssembler (packages/llm/llm/src/assembler.ts) is the single shared implementation that folds a StreamChunk stream back into ContentBlocks and a final Message. The loop logs the raw chunks (for replay fidelity) while feeding the same chunks through an assembler — so the canonical log keeps token-level detail and the derived message is rebuilt deterministically. A consumer that needs the assembled result without re-implementing the fold uses this.

The seam

LlmAdapter is the provider seam: subclass, implement stream(), register with ctx.llm.registerAdapter(models, adapter). The block-start / block-end index correlation and the assembler together mean an adapter only has to emit well-formed chunks — block reassembly is not each adapter's problem. The consumer surface (ctx.llm.stream()) and the llm/stream waterfall are described in architecture.md § Content blocks and streaming.

ContentBlockType (the key set the index-correlated blocks carry) derives from ContentBlockMap:

interface ContentBlockMap {
  'text': TextBlock
  'reasoning': ReasoningBlock
  'tool-call': ToolCallBlock
  'tool-result': ToolResultBlock
}

See core.md § Content blocks and messages for the block interfaces.