Files
deepseek-harness/docs/rfc/implemented/architecture/2026-07-05-reconstructable-requests.md
Tianyi Cui c0808d5126 docs: the governing principle — every LLM request is reconstructable from the session log
The reconstructability RFC is the principle's home: model-visible ⟺
logged in both forms, the mechanism (boundary derivation + header
fold), the enforcement (write-time round-trip guard, the dev
invariant), the corollaries ranked (prefix-cache stability first), the
MiniCode lineage with the provenance arrow inverted, and the
alternatives it beat — including the stateful transmission client
whose three-design archaeology lives in PR #162.

Placements per the one-home-per-fact taxonomy: a standing-order line in
root AGENTS.md (with displacement trims to stay inside the 1,575-word
ceiling), the principle statement in architecture.md § Session Log and
its Turn Flow lines (condensed to the ratcheted 1,630 ceiling), the
request-envelope section in core-data-structures/core.md with the
LlmCallConfig paste, both review-requested FIXMEs
(FIXME(call-config-shape) beside the type, FIXME(catalog-verbs) at the
catalog's drift-gate note), cookbook rows redirected off agent/request
(tool filtering → system-prompt/assemble, plan-mode prompt → sections/
inject()), and the llm/stream JSDoc stating the frozen-request
contract. RFC index and all generated catalogs regenerated.
2026-07-06 03:49:35 +08:00

13 KiB

RFC: Every LLM request is reconstructable from the session log

Status: implemented

Problem

Two gaps shared one root. First, provider KV caching (DeepSeek context caching) is prefix-based — a request pays full price only for the tokens after the longest stored prefix it matches — yet nothing in the request pipeline stated, checked, or measured prefix stability: every registered PromptSection happened to be static, the tool set happened not to change mid-session, no listener happened to rewrite requests. A single time-interpolating section would have silently multiplied context cost, and no test or metric would have moved. Second, and deeper: the session log — the system's single source of truth — could not actually answer what the model saw. It recorded every message but never the system prompt, the tool schemas, or even which model; the mutable agent/request waterfall handed listeners the whole GenerateOptions to rewrite per call; replay equivalence was therefore a property of the plugin population, not of the design.

The reference shape for the happy path is MiniCode's LLMClient: a stateful conversation client, appended to — never rebuilt — as the conversation advances, resetting only when the system prompt, tool set, or compaction genuinely changes what the model must see. The design question this RFC answers is how to get that discipline without giving up event-sourcing.

Decision

The principle

Model-visible ⟺ logged. Anything that reaches a model request must be recorded in the session log. The checkable consequence: every conversation request the loop sends is a pure function of the session log — anyone holding the log reconstructs it byte-for-byte. Scope, stated precisely: the guarantee covers the loop-built GenerateOptions; provider wire bytes follow from it because both adapters' serialization is a pure per-message function at a pinned code version; direct one-shots (compaction's summarize call) log their envelope scalars (compact/summary.{model, maxTokens}) and their input is deterministic code over the logged region — reconstructable from log + code, outside the invariant by the unfrozen-request marker.

Prefix-cache stability is corollary #1, not the headline: an append-only log projected by a per-node pure function yields requests that are append-extensions of their predecessors whenever the header is unchanged — stability is emergent, not managed. Byte-exact audit/replay is corollary #2; resume and fork with attributable drift is corollary #3.

The mechanism

Messages. Session.deriveMessages() is cached: each surface node is projected exactly once, when first seen, through the public per-node function deriveEventMessage(event); a surface rewrite (a compaction replaceSurfaceManager.replaceGeneration) rebuilds. Callers get a fresh array per call over shared, deep-frozen messages: mutating logged history through a projection is unrepresentable (it throws), replacing the old clone-per-call isolation. External reconstructors fold the same public function over a log prefix, so no two paths can disagree.

The header. The request's non-content half — EpochHeader: call config (LlmCallConfig: model + sampling scalars), rendered system prompt, assembled tool schemas — is logged session state, in canonical form (empty system/tools ≡ absent). Two log-only, turn-enclosed events in dsh-session carry it: request/header, a full snapshot with reason 'initial' | 'resume' | 'fallback', and request/header-delta, an amendment (SystemDelta: a common-prefix/suffix line trim; ToolsDelta: name-keyed added/removed/changed; config: replaced whole). The pure trio foldRequestHeader / diffHeader / applyHeaderDelta reconstructs; the live session tracks the fold with the same lazy cursor as the message cache. Snapshots anchor the fold where a fold needs anchors — conversation birth and process boundaries — and each loop instance appends one on its first request ('initial' when the log has none, 'resume' otherwise, even when nothing changed: the boundary itself is a recorded fact, and cross-restart drift becomes attributable while an unchanged header resumes byte-identical). Deltas are an encoding optimization with a safety valve, never a correctness dependency: the writer verifies applyHeaderDelta(prev, delta) reproduces the new header exactly and records a 'fallback' snapshot when the encoding cannot express a change (a pure tool reordering), so a well-formed log always folds.

The loop, transmission-stateless. Per step: render assembly (every step — value comparison needs no change-signal discipline, and a section that varies per step surfaces as a logged header event per step instead of a silent bust) → agent/pre-step (compaction's surface mutations land before derivation) → messages snapshot, then step/start appended as the next operation in the same synchronous frame → seed the call config (first request of the instance: from AgentOptions, so explicit options always beat the logged baseline — fork model-overrides and resume reconfiguration stay correct; afterwards: from the folded header) → the agent/request waterfall, re-typed (agent, turn, step, config: LlmCallConfig, next) → LlmCallConfig — a frozen seed and a returned replacement are ALL a listener shapes; content flows through the log channels (inject(), steering, prompt-submit additionalContext, sections via system-prompt/assemble) — → the header event the request owes the log → build GenerateOptions from the snapshot + header, deep-freeze (deepFreeze exempts the AbortSignal, the one live control channel — freezing one breaks AbortController.abort()), dispatch. The loop's only in-process bookkeeping is one boolean: whether this instance has logged its anchoring snapshot.

The reconstruction boundary is step/start, unconditionally. A step's messages are the derivation over events[0..stepStartSeq). Because the snapshot precedes the step/start append in the same synchronous frame, nothing can enter this request past the boundary: an agent.inject() from an agent/request listener (or any concurrent task, or a session/event listener firing on step/start itself) lands in the log after the boundary and joins the NEXT request. For waterfall-window appends this matches the prior loop (it also derived before its waterfall); for a synchronous step/start listener it is a deliberate change — such a listener could previously reach the current request — and agent/pre-step is the sanctioned seam for content that must affect the CURRENT request. A step's header for reconstruction is the fold after its own request/header* event (which sits between its step/start and first response event) or the fold carried forward.

Enforcement. Dev-mode (dsh-invariants), on llm/stream: a frozen request with a live sessionId — the loop-built marker; hand-built one-shots are unfrozen and skipped — must carry messages deep-equal to the boundary derivation, rebuilt through a FRESH Session over events[0..stepStartSeq) so the live cache cannot vouch for itself, and header fields equal to foldRequestHeader over the log. There is no divergence allowance and nothing to allow: no seam can put unlogged content into a request. prepend: true only defends against the replay adapter's short-circuit (an append-registered listener); two prepended listeners have no defined mutual order in cordis, so correctness rests on the seq-bounded fold, never on listener timing. Measurement stays lean: the with-key e2e (request-cache.e2e.ts) proves usage.cacheReadTokens > 0 on every request after the first against the live API, and per-step usage in the log is the production observable — a header event or compaction shows up as a cache-read collapse on the next step.

The MiniCode shape: adopted, with the provenance arrow inverted

What survives from LLMClient: the conversation is maintained, not rebuilt — one projection per message, ever; requests advance append-only; resets happen only for a system-prompt/tool change, a config change, or compaction, each now a logged fact. What is deliberately inverted: MiniCode's client is the source of truth and its event stream derives from client appends (on_event(MessageAdded)), which suits an advisory event stream. Here the log is contractual — persistence, crash recovery, fork seeding, transcript rendering, and the snapshot harness all replay it — and it carries strictly more than a message list (turn/step boundaries, raw chunk streams, tool-call pairing, provenance, log-only records), so a message-list client cannot generate it. The arrow therefore points log → client: the conversation state IS the log plus two cached folds inside Session (messages, header), and the "client" the loop talks to is the session itself. What the inversion buys over the original: the reconstruction is checkable against an independent record on every request — MiniCode's client has nothing to check itself against.

Alternatives considered

  • Client as source of truth (literal MiniCode): a second operative truth beside the log — the two drift and nothing notices; see the section above.
  • A stateful transmission client mirroring the log (a PromptPrefix class holding committed/open message zones with an append/editTail/reset vocabulary, the log pushed into it per event): behaviorally equivalent on the happy path, but it duplicates conversation state outside the session, needs transactional rollback around listener seams, keeps an unlogged content-shaping surface (editTail) whose divergence the invariant must specially allow, and still cannot answer "what header did the model see" from the log. Dissolving it into the session's own caches plus logged header events made every one of those problems unrepresentable instead of guarded. (PR #162 is the archaeology of this alternative, three designs deep.)
  • Per-call request scalars (a freely mutable config handed to each agent/request dispatch): a listener flips the model per call with zero accounting, silently abandoning the provider cache this design exists to protect. Config is per-conversation logged state; the waterfall proposes, the log records.
  • Detect-and-report (compare consecutive requests, warn on divergence): catches violations after the fact; a violating request is still constructible and ships. Rejected for interface-level unrepresentability.
  • Event-driven assembly (re-render only on change signals): a missed-signal bug class — a tool registered mid-session emits tools/change, not system-prompt/change, and a third-party provider may emit nothing. Per-step render + value compare is robust with zero signal discipline.
  • Narrative fields on the header events (a reason/changed list on deltas): derivable by diffing consecutive events — one home per fact; snapshots carry a reason because an anchor's cause is NOT derivable from the data.

Consequences

  • A request that is not explained by the log cannot be constructed by accident — not by the loop, not by a listener; mutating a built request throws; every header change is a durable, diffable log event.
  • What still costs full price at the provider is inherent and logged: compaction (its compact/* events and replace node), a real prompt/tool change (request/header-delta), a config switch (ditto), a process boundary with drift ('resume' snapshot differing from its predecessor). The provider's own reasoning-content exclusion is managed server-side.
  • The step/start-listener behavior change (above) is the one observable semantics change for plugins; agent/pre-step is the current-request seam.
  • Tool-result trimming (planned) needs no new mechanism: a logged single-node surface replace (start === end) carrying a trimmed tool/result under the same callId — compaction-family, replay-correct, cache-bust batched by the same pressure logic.
  • Session logs grow one request/header snapshot per conversation (system + tool schemas: the dominant term), plus deltas on real changes — small next to assistant/chunk volume; SESSION_FORMAT_VERSION stays 0 (pre-release churn is absorbed, backends reject-not-migrate).
  • Snapshot goldens changed once (every transcript gains its header events); the fs-writing fixtures are stored in the normalized authored form with cwd-relative tool arguments, because replay only round-trips cwd-independent argument paths.
  • FIXME(call-config-shape): revisit LlmCallConfig's exact field set — which fields are genuinely epoch-level for cache purposes (model certainly; the sampling scalars sit there out of caution), and where provider-specific extras (reasoning options, extra body params) belong when an adapter needs them.