Files
deepseek-harness/docs/rfc/implemented/feature/2026-06-15-code-mode.md

34 KiB
Raw Blame History

RFC: Code Mode — the model writes TypeScript against the tool registry

Status: implemented

Problem

In the registry's native presentation, the agent loop advertises every visible capability as a JSON-schema function definition. ToolRegistry contributes its schemas to the system-prompt assembly, the assembly's tools land on the wire (and in the logged request header), the model invokes one tool-call block per step, and the loop dispatches each call through ctx.tools.execute() sequentially (parallel tool execution is an explicit open TODO in dsh-tools and docs/architecture.md), with every intermediate tool-result re-entering the model's context on the next request.

For multi-step tool work this is token-heavy and serial. The model cannot compose tools — loop over a result set, branch on an intermediate value, fan out, post-process — without a full model round-trip per call, and each round-trip drags the entire intermediate result back into context whether the model needs it or not.

Cloudflare's Code Mode proposes an alternative grounded in a simple observation: LLMs are better at writing code than at emitting tool calls, because they have seen millions of lines of real code and comparatively few contrived tool-calling traces. Instead of one tool call per step, the model writes a TypeScript program against a generated API over the tools, the program executes in a sandboxed runtime, and the model curates what comes back — only what it prints or returns — instead of every intermediate result.

Tool presentation belongs to the registry that owns tool visibility: implementing a second presentation as an after-the-fact waterfall transform would make correctness depend on listener order and fight reconstructable requests. The execution substrate is also part of the foundation rather than a placeholder: Node worker_threads provides a separate isolate, an empty environment, heap caps, and termination of a hot synchronous loop, while fitting the harness's existing trust model (§Trust posture).

Decision

Three decisions, each elaborated in its own section below:

  1. Code Mode is a first-class presentation mode of ToolRegistry (dsh-tools), selected by a validated mode config: 'native' (the default, contributing the visible capability schemas), 'code' (the registry contributes only its reserved run_code transport plus a generated SDK .d.ts in the system prompt), or 'both' (native schemas and the transport + SDK). The registry shapes its canonical contribution at the source; the cooperative prompt-assembly result remains authoritative, and the logged request header records exactly that returned presentation.
  2. Code execution is a capability seampackages/code-runtime/ contains the interface package @deepseek-ai/dsh-code-runtime, which owns ctx.codeRuntime (capability seams; consumer = dsh-tools, with core-consumes-a-seam precedent in agent-loopdsh-llm). The runtime knows nothing about tools: it is handed a program and named async bindings, runs the program, and reports { value, logs, error? }. Language and substrate are backend properties, so a future Python or container backend is another implementation package, not a redesign.
  3. The shipped implementation is @deepseek-ai/dsh-code-runtime-worker: one fresh Node worker thread per run, executing the model's TypeScript after type-strip, with bindings bridged over the message port, an empty environment, configurable heap/output/time caps, and hard termination. Its trust posture is bash-equivalent by design — no unsafe-acknowledgement flags — because the harness already ships dsh-bash-local, which executes arbitrary model-written shell commands with strictly more ambient authority.

The registry owns the mode

ToolRegistry gains a schemastery-validated config (static Config), its first: mode: 'native' | 'code' | 'both', default 'native'. A deployment flips it from cordis.yml (tools: { mode: code }) — no code edit, per the no-hardcoded-tunables convention.

Wire tool list = the registry's contribution before cooperative assembly. The registry feeds assembly through a mode-aware provider: 'native' contributes every capability visible to that assembly scope, 'code' contributes only run_code, and 'both' contributes both. Because PromptAssembly.tools is the single source the loop's request header snapshots, the final presentation is logged and reconstructable. The reserved transport is not a capability: it lives outside global/scoped registration and restriction layers, cannot be registered or shadowed there, and cannot be named by ctx.tools.restrict(). The mode governs this provider's input to assembly; other direct systemPrompt.tools() providers own their schemas, and the trusted assembly waterfall owns the returned wire list.

Interaction with toolOrder, stated up front: a configured systemPrompt.toolOrder naming native capabilities rejects every assembly under mode: 'code', because those names are outside that mode's wire-validation universe. This is correct behavior, not a bug: a deployment using Code Mode updates its order config or drops it.

The SDK prompt section. Under 'code' and 'both' the registry registers one lazy prompt section (tools:sdk, in the 100199 tool-guidance order band) whose thunk regenerates, for each assembly scope, a TypeScript declaration of every visible end-capability tool plus fixed usage instructions. It uses the same visibility resolver as lookup and execution, so scoped grants and shadows appear while restricted globals disappear; the reserved run_code transport itself is excluded. The thunk emits tools in lexicographic name order, so an unchanged visible set produces byte-identical text.

Assembly ownership. run_code and tools:sdk enter the trusted system-prompt/assemble waterfall as normal assembly inputs. A scoped tools:sdk section may shadow the global default before dispatch, and a listener may remove or replace either contribution. The waterfall's returned assembly is final, so whoever changes these inputs owns preserving a viable Code Mode protocol when the deployment expects Code Mode to remain usable; no restoration pass overrides deliberate composition.

Codegen. A pure jsonSchemaToTs(schema) module inside dsh-tools (sibling of json-schema.tsschemas() and the SDK are two projections of the same store) maps the JSON-Schema subset the defineTool DSL emits (object/string/number/boolean/array, properties, required, string enum → literal union, nested objects, array items, description → JSDoc) to a TS type literal. It is total: any construct outside that subset ($ref, oneOf/anyOf, integer, future MCP shapes, …) degrades to unknown without throwing. Because ToolSchema.name is an arbitrary string, the SDK is declared as one object constant — declare const tools: { "some-mcp-tool"(args: …): Promise<string>; bash(args: …): Promise<string>; … } — quoted keys make every name reachable with no sanitization or alias-collision logic. Typing is advisory (the runtime executes type-stripped JS); the instructions say so.

The run_code tool and the dispatch bridge

Under 'code' and 'both' the registry owns run_code as a reserved presentation transport with one required parameter, { code: string }. It is represented by a normal ToolDefinition for dispatch but stays outside the filterable capability layers, so restrictions cannot accidentally remove Code Mode's only entry point. Calls traverse the complete tool pipeline — tools/pre-execute → monotonic guards → tools/execute around dispatch → tools/post-execute → immutable tools/result notification — exactly like native calls; a permission plugin can inspect the program text before it runs, and final-result observers see the normalized outer outcome. Its execute(args, exec):

  1. Builds the bindings: the bridge owns a run-scoped AbortController whose signal follows exec.signal (an outer cancel propagates in) and which the bridge itself fires the moment the run settles for any reason — completion, program exception, computeMs/maxWallMs expiry, worker exit. For every visible capability tool, the binding is an async function that (a) checks the run signal before and after, (b) JSON-normalizes the argument — a JSON.parse(JSON.stringify(args)) round-trip, rejecting that one call with a descriptive Error when the value does not survive (BigInt, circular structures) — because the seam's structured-clone boundary is wider than JSON while the session log accepts only JSON, (c) awaits its turn on the per-run serialization queue (below), (d) calls this.execute({ callId, name, arguments, agent: exec.agent, parent: exec.token, signal: runSignal }) with a deterministic sub-id CallId(`${exec.callId}:code:${n}`), (e) appends a tool/code-dispatch session event, and (f) maps the result: success → the text-block contents joined as a string (non-text blocks become placeholders), isErrorthe binding rejects with an Error carrying the result text. The child's readonly parent is only the outer execution's frozen, property-free token, so commit-style observers can correlate outcomes without receiving a mutation path into the live run_code wrapper. Every sub-call still traverses the full pipeline under its own immutable identity and registry-assigned token. The run signal, rather than the bare outer one, lets budget expiry abort an in-flight sub-tool instead of orphaning it. Rejection gives programs ordinary try/catch and Promise.all failure semantics rather than a bespoke result envelope.
  2. Runs the program: ctx.codeRuntime.run({ program: args.code, bindings: [{ global: 'tools', functions }], signal: runController.signal }). The runtime receives the run-scoped signal, not only the caller's outer signal, so any way the outer run settles also aborts work inside the runtime.
  3. Surfaces the outcome — after reaching quiescence. When ctx.codeRuntime.run() settles, whether by fulfillment or rejection, the bridge fires the run-scoped abort (cancelling any in-flight sub-dispatch and abandoning queued-unstarted ones), then awaits the dispatch queue's drain before returning or propagating, per the dispose-to-quiescence rule in defensive patterns: an aborted in-flight sub-call still settles and logs its isError tool/code-dispatch event inside the open turn, and nothing can append after run_code settles. A successful result then returns one text block — the captured console/stdout output followed by the rendered return value (if any) — plus a meta payload (capped logs, dispatch count) for presentation. A fulfilled run with result.error throws a CodeRunFailedError extends HarnessError (code: 'CODE_RUN_FAILED', message = the error kind and text plus captured logs so the model can self-correct); a backend rejection propagates through the same registry error boundary. Both become structured isError tool results.

Sub-call additionalContext is suppressed, deliberately. A tools/post-execute hook may attach additionalContext to a call; for loop-dispatched calls the loop buffers those and appends each as a context/message only after the step's tool/results, preserving call/result adjacency. A sub-dispatch result's additionalContext has no such safe outlet from inside a running run_code: injecting immediately would land a context/message between the parent's tool/call and its tool/result (breaking the adjacency the buffering exists to protect), and PostToolDecision.additionalContext is singular where a program may produce many. The MVP therefore drops sub-call additionalContext, pinned by a test and stated in the hooks bridge's docs; the follow-up (a plural context channel or loop-level sub-dispatch buffering) is deferred until a real hook needs it through Code Mode.

Concurrency: serialized, enforced by the binding. The bindings are async, so a model writing Promise.all([tools.a(…), tools.b(…)]) starts both immediately — concurrent dispatch would be the default, while the tool contract carries no concurrency-safety metadata (the open parallel-execution TODO). Each run_code invocation therefore owns a dispatch queue and every binding call chains onto it, so even Promise.all executes the underlying ctx.tools.execute() calls one at a time in submission order; when the run settles, queued-but-unstarted dispatches are abandoned. Lifting this per tool remains tied to tools declaring themselves concurrency-safe.

Presentation. run_code's render intent is decided here per the render-intent RFC: presentCall → a generic card, kind: 'execute', title = the program text, rawInput = the same program text; presentResult → a generic card whose content is the captured output (from meta). The program is the title because ACP execute cards reliably render that field while some clients omit body and raw-input content. This is not a terminal card: that card's semantics are "a shell command in a working directory", which a program is not.

Observability: tool/code-dispatch

Each sub-dispatch appends one session event, declared by dsh-tools via SessionEventMap declaration merging (the map is merge-extensible for exactly this; todo/write is the log-only precedent): tool/code-dispatch with { parentCallId, subCallId, name, arguments, isError, resultSummary }arguments being the bridge's JSON-normalized value, the very one dispatched, so the append cannot fail on payload shape. It is log-only — deriveEventMessage() ignores unknown event types by design, so sub-calls never re-enter model context — but persistence and UIs get every call. As a log event it carries JSDoc prose but no @mode tag (that vocabulary belongs to cordis bus events; the persistence-catalog generator hard-errors on one) and lands in the regenerated docs/persistence-catalog.md; appends happen inside run_code's execution, so the turn-enclosure invariant is satisfied by construction. A run_code execution arriving without exec.agent (the loop always supplies it; direct programmatic calls may not) still runs and simply skips event logging, exactly as the ToolExecution contract allows.

The code-runtime seam

packages/code-runtime/code-runtime/@deepseek-ai/dsh-code-runtime, depending only on cordis. An abstract CodeRuntime extends Service (super(ctx, 'codeRuntime')) plus the vocabulary:

  • CodeRunRequest = { program: string; bindings: CodeBindingNamespace[]; signal?: AbortSignal }
  • CodeBindingNamespace = { global: string; functions: Record<string, (args: unknown) => Promise<unknown>> } — the runtime exposes each namespace as a global object of async functions inside the program; binding arguments and resolutions must be structured-cloneable (a runtime may cross a serialization boundary; ours does).
  • CodeRunResult = { value?: unknown; logs: CodeLogEntry[]; error?: CodeRunFailure } — program execution outcomes, including exception, timeout, abort, and worker exit, resolve as the error field. run() may reject only for caller/seam misuse (for example a duplicate binding namespace); consumers still contain a non-conforming backend rejection at their own error boundary.
  • CodeLogEntry = { source: 'console' | 'stdout' | 'stderr'; level?: 'log' | 'info' | 'warn' | 'error' | 'debug'; text: string }
  • CodeRunFailure = { kind: 'exception' | 'timeout' | 'abort' | 'worker-exit'; message: string } — orthogonal outcomes reported independently per defensive patterns; a timed-out run is not an exception, an abort is not a timeout.
  • Two readonly backend descriptors, informational not gating: language (what the program must be written in — 'typescript' for the shipped backend; a Python backend would say so, and pair with its own SDK generator on the presentation side) and isolation ('worker-thread' for the shipped backend; 'process', 'container', … for future ones). dsh-tools requires language === 'typescript' in the MVP — its codegen emits TS — and fails the assembly loudly otherwise, the same misconfiguration idiom as toolOrder violations (as when mode is non-native with no ctx.codeRuntime loaded at all).

Per explicit-over-implicit at seams, the request spells out everything the runtime acts on; defaulting (timeouts, caps) is the implementation's validated config, never a hidden ?? inside run(). Consumption uses the loop's optional-backend idiom: Cordis has no optional injection — every inject entry gates activation — so a static inject on the registry would hold ctx.tools (and every tool plugin behind it) hostage to a code runtime existing even under mode: 'native'; instead the registry reads ctx.get('codeRuntime') at use time, exactly as agent-loop consumes sessionPersistence, with absence failing loud in the provider thunk. The seam has concrete divergence on both axes: the worker-thread substrate can be replaced by a container or microVM implementation, and the TypeScript language contract can be paired with a language-specific SDK and runtime. dsh-tools consumes only the interface and tests against a trivial in-repo fake, exactly the interface/implementation/consumer shape of the bash template.

The worker-thread runtime

@deepseek-ai/dsh-code-runtime-worker, the second package of the packages/code-runtime/ group. Per run():

  1. Type-strip host-side with Node's built-in stripTypeScriptTypes (node:module; present across the repo's whole engines range, ^22.19.0 || >=24.0.0, and position-preserving, so runtime error line numbers match the model's source). Strip-only mode rejects non-erasable syntax (enum, namespaces) — that rejection returns as error.kind: 'exception' with Node's message, the SDK instructions say "erasable TypeScript only", and the model self-corrects like any other program error. A syntax-level failure never spawns a worker.
  2. Spawn one fresh Worker per run from the package's own bootstrap module: env: {} (truly empty — stronger than the scrubbed-env rule for spawned commands), resourceLimits from config, stdout/stderr captured into logs rather than inherited. No pooling and no cross-run state: a program's world dies with its worker, which keeps runs reconstructable from the log alone and makes state bleed unrepresentable.
  3. Execute in the bootstrap: the stripped program becomes the body of an AsyncFunction whose parameters are the binding globals and a capturing console shim, so top-level await and return work and the program's completion value is the run's value (structured-cloneable values cross as-is; anything else is replaced by its util.inspect rendering, documented).
  4. Bridge bindings over the message port: each binding function in the worker posts { id, global, name, args } and awaits the reply; the host validates the name against the request's bindings, invokes, and replies { id, ok, value } or { id, ok: false, message } (a host-side binding rejection becomes a program-side rejection). The worker-side namespace objects are built null-prototype via defineProperty, so a binding named __proto__, constructor, or toString is an ordinary own property, not a prototype collision. Unknown names, duplicate ids, and post-settlement messages are rejected or ignored — the port protocol assumes a hostile peer, because the peer runs model code.
  5. Enforce caps — two independent budgets, because the peer is hostile. The compute budget (computeMs) meters the worker's measured busy time via worker.performance.eventLoopUtilization() polling — not host-side "is an RPC pending" bookkeeping, which a program defeats by firing an un-awaited call at a slow tool and then spinning hot while the host thinks it is waiting. Measured busy time cannot be gamed: a hot loop accrues it whether or not a dispatch is in flight, and a program genuinely awaiting a slow tool accrues none, so a long-running bash sub-call still does not kill an innocent run. The wall ceiling (maxWallMs) never pauses for anything and backstops what busy-time cannot see (a program awaiting a promise nobody will resolve). Budget expiry, signal abort, and run completion all funnel into worker.terminate(), which ends hot synchronous loops too (measured; this was node:vm's unfixable gap); the failure reports which budget fired. Heap overflow surfaces as the worker's OOM exit → error.kind: 'worker-exit'. Log and value sizes are capped by config, truncation marked in-band. All caps are validated config fields with defaults (computeMs: 60_000, maxWallMs: 600_000, maxLogBytes: 65_536, maxValueBytes: 32_768, maxOldGenerationSizeMb: 512), changeable from cordis.yml.
  6. Dispose to quiescence: the service's own disposal terminates in-flight workers and awaits their exits before resolving, per defensive patterns.

Trust posture

The worker runtime is containment, not a security boundary. Model code in the worker can reach Node globals — fetch, process (with an empty env), dynamic import() of built-ins — so a deliberately adversarial program has ambient authority comparable to what the harness's own bash tool grants every model turn: dsh-bash-local runs arbitrary model-written commands with the host filesystem, network, and a scrubbed-but-populated environment. One asymmetry runs the other way: worker.terminate() ends the thread, not OS processes a program may have spawned via node:child_process — weaker than bash-local's process-group kill for direct children (equivalent for double-forked daemons, which survive both); the wall-clock ceiling bounds the worker itself, and orphan cleanup is the same deployment-level concern it is for bash. Code Mode is gated where bash is gated — tools/pre-execute, where permission/sandbox plugins veto or approve the program — and adds containment bash does not have: empty env, heap caps, hard termination of the program itself, and a separate isolate. A node:vm executor with no containment would need explicit unsafe acknowledgement; imposing that ceremony on the better-contained worker while bash needs none would be posture theater. A deployment that needs a hard boundary (untrusted multi-tenant input) needs it for bash too; that is a future isolation: 'container' backend, and the isolation descriptor lets deployments distinguish backends.

What the model sees

The tools:sdk section carries the .d.ts plus fixed instructions: the program is the body of an async TypeScript function (erasable syntax only — no enum/namespaces; type annotations are advisory); call tools as await tools.name(args) (quoted access for exotic names); a failed tool call rejects with an Error carrying the tool's error text — catch it to handle and continue; calls run sequentially even under Promise.all; emit results via return and/or console.log, and only that curated output returns to the context — intermediate tool results never do. That last line is the payoff the whole design serves: output-side context cost becomes the model's own editorial decision. On the input side the .d.ts is not free — for a large tool surface it can rival the native JSON schemas it replaces (and 'both' pays for the two side by side) — but it is prefix-stable, so provider prefix caching amortizes it; the win is workload-dependent and the RFC claims no more.

Consequences

The design consists of the dsh-code-runtime interface package, the dsh-code-runtime-worker backend, and the dsh-tools presentation and dispatch integration.

Shipped surface:

  • The seam: packages/code-runtime/@deepseek-ai/dsh-code-runtime (abstract CodeRuntime, the vocabulary above, ctx.codeRuntime) and @deepseek-ai/dsh-code-runtime-worker (the worker-thread backend, every cap a validated config field). Rows in the service map, capability-seams graph, config catalog, and cordis catalog.
  • The registry surface: ToolRegistry's mode config, mode-aware wire contribution, lazy tools:sdk section and reserved run_code transport, jsonSchemaToTs/renderToolsSdk (exported), the dispatch bridge and CodeRunFailedError, and the tool/code-dispatch log event (declaration-merged into SessionEventMap, regenerated into the persistence catalog; run_code in the tool catalog).
  • The composed surface: the tools config forwards through agent-core and both app packages (stdio-agent, acp-agent); demo:code-mode boots each UI example's code-mode.cordis.yml overlay (the worker runtime + mode: 'code' over the base tree); every program sub-dispatch resolves the same scoped capability view and re-enters the complete tool pipeline with an immutable link to its enclosing transport execution.
  • Interactions inherited by deployments: a toolOrder naming native tools rejects every assembly under 'code' (update or drop the order config when switching modes); restrictions can hide end capabilities but cannot remove the registry-owned presentation transport, while assembly listeners may rewrite the final model-visible surface and own its protocol integrity; sub-call additionalContext is dropped by the bridge (a plural context channel is deferred until a real hook needs it through Code Mode); sub-dispatch stays serialized until tools can declare concurrency safety — the same metadata the native parallel-dispatch TODO waits on.

Testing

What the suites pin, per tier:

  • Unit — worker runtime (real workers, no mocks): output/value capture and log-source attribution; error kinds (exception incl. non-erasable syntax, abort, worker-exit under OOM); the two budgets from both sides (a hot loop behind an un-awaited pending dispatch dies at computeMs busy time; a program idling on a slow binding outlives computeMs and dies only at maxWallMs); binding-bridge hostility (junk/forged port traffic incl. non-object messages and forged log/done cap bypass attempts, unknown names, duplicate ids, post-settlement replies, __proto__/constructor/toString binding names); structured-clone fallback and cap truncation; env emptiness verified from inside a program; dispose-awaits-exit. A real-load-path e2e runs the BUILT package under plain node so the worker entry resolves both unbuilt (tsx) and built — the published-artifact guard from docs/testing.md.
  • Unit — registry integration: the codegen table (DSL subset, quoted names, unknown degradation, byte-identical determinism); provider contribution per mode ('native' capabilities, 'code' exactly [run_code], 'both' capabilities + run_code); reserved-name, restriction, scoped shadowing, authoritative assembly transformation, and toolOrder × mode invariants; missing-runtime / wrong-language loud failures; full-pipeline and opaque parent-token behavior for sub-dispatches; serialization non-overlap (a probe tool records enter/exit under Promise.all); abort aborting the in-flight sub-dispatch and abandoning queued ones; binding rejection on isError and on JSON-unrepresentable arguments; CodeRunFailedError → structured isError carrying kind + logs; tool/code-dispatch payloads (JSON-normalized arguments identical to what dispatched); deriveMessages() ignoring the event; sub-call additionalContext suppression; HMR safety.
  • e2e (with-key, self-skips): a real model under mode: 'code' composes two bash calls in one program (examples/coding-agent/tests/code-mode.e2e.ts) — every logged request/header carries exactly [run_code], the dispatch events land under the parent call, the file the program wrote exists, and the final answer is the curated output.
  • Snapshot (keyless replay): goldens for a run_code turn under 'code' and 'both' (code-mode-turn, both-mode-turn), each its own header-pinning class — the SDK section text, the collapsed header tool list, the dispatch events, and the result card are committed and replayed.

Alternatives considered

An add-on consumer plugin with zero core changes. Rejected because agent/request is call-config-only under reconstructable requests, while transforming an assembled tool list would have to undo toolOrder canonicalization without owning its config and would depend on listener order. Which tools the model is offered, and in which representation, is the registry's single concern: native schemas and the SDK are two projections of one visible store.

node:vm as the reference runtime, with hardening deferred. Rejected: node:vm is not isolation (prototype-chain escapes reach the host realm) and cannot interrupt a hot loop. A worker thread provides a separate isolate, empty environment, resourceLimits, and reliable terminate() at bash-equivalent trust, so the reference and production implementation are one package without an unsafe-acknowledgement ceremony.

Result elision / summarization over native tool-calling. Addresses only the context-bloat half of the problem: trimming old tool-results is cheap to add as a logged surface replacement under reconstructable requests, but still pays one model round-trip per call and cannot express loops, branches, or joins. Complementary, not competing; it can layer under Code Mode for residual native calls.

Parallel native dispatch in the loop. The other answer to round-trip cost; still valid future work (the open TODO), still blocked on concurrency-safety metadata, and still no composition — it parallelizes calls the model already decided on in one step. Code Mode's serialized-queue decision keeps the two compatible: when the metadata lands, both native parallel dispatch and per-tool binding parallelism unlock together.

Always-exclusive (Cloudflare-faithful, no mode). Rejected for this SDK's primary consumer: a coding agent's bread-and-butter single calls (bash, read, edit) are already ideal as native calls, and forcing every edit through a program taxes the common case. The mode config keeps the faithful form ('code') one line away without imposing it.

Per-tool visibility tiers (this tool native, that tool code-only). Deferred: it needs per-tool metadata and a presentation split that 'native' | 'code' | 'both' does not, and its design depends on evidence about how models split usage under 'both'.

Sanitized identifier aliases in the SDK (my-toolmy_tool, Cloudflare's approach). Rejected: quoted keys on a declare const make every name reachable with zero alias-collision logic; models handle tools["my-tool"](…) fine.

A REPL-style persistent kernel (state survives across run_code calls). Rejected for the MVP: cross-call state would be invisible to the session log, breaking the reconstructability guarantee that every request is a pure function of the log; fresh-per-run keeps it. A kernel-style backend remains expressible behind the seam later, with its own logging story.

Risks

The worker is not a hard security boundary. Deliberate and documented (§Trust posture): posture equals the existing bash tool, containment exceeds it, gating uses the same seams. Deployments needing more need a future isolation: 'container' backend — tracked as the seam's designed extension, not a TODO on this design.

stripTypeScriptTypes is marked experimental. It is the same engine (amaro/swc) behind Node's own native .ts execution, exposed as an API across this repo's whole engines range. Mitigations: the runtime's unit suite pins the behaviors relied on (position preservation, erasable-only rejection message shape loosely), the call sits behind one private function, and amaro/sucrase are drop-in replacements if the API shifts. The erasable-only subset is a model-facing contract line, and the error path is a working feedback loop, not a dead end.

Prompt cost of the SDK, especially under 'both'. The .d.ts can rival the native schemas it complements; 'both' carries two representations. Prefix stability + provider caching amortize per-session cost; the mode is per-deployment; the RFC makes no unconditional-savings claim. Measured guidance (when to prefer which mode) is explicitly post-ship learning.

Registry scope growth. dsh-tools absorbs codegen, a tool, a bridge, and an event. Contained by module boundaries inside the package (ts-types.ts, code-mode.ts beside schema.ts/json-schema.ts/presentation.ts) and by the seam: everything substrate-shaped lives behind ctx.codeRuntime.

Structured-clone limits at the binding boundary. The seam's clone boundary admits values JSON does not (Date, Map, BigInt), and the session log accepts only JSON — left unhandled, a sub-call could execute and then fail at tool/code-dispatch append time. Closed by the bridge's JSON-normalization step (§ the dispatch bridge): what does not survive the round-trip rejects that binding call before dispatch, so every executed sub-call is loggable by construction. The seam itself keeps the wider structured-clone contract (it is about the port, and stated so a future binding producer cannot discover it in production); consumers with stricter payload needs enforce them at their own boundary, as the bridge does. Non-text sub-result content is reduced to placeholders — a known MVP limitation, recorded in the SDK instructions.

Serialized-only sub-dispatch. Promise.all gains no wall-clock parallelism yet, only fewer round-trips; models may over-expect. The instructions state it; lifting it is tied to the same concurrency-safety metadata the native parallel-dispatch TODO needs.

Budget metering reads the event loop, not a flag. Busy-time polling (eventLoopUtilization()) is coarser than an exact CPU meter — a budget expires up to one poll interval late — and its correctness claim ("a pending dispatch cannot pause it") is load-bearing against a hostile program. Both sides are unit-tested (hot loop with a pending decoy dispatch dies at computeMs; idle-on-slow-binding survives to maxWallMs), and the poll interval is an internal constant, not config — nothing a deployment could mis-tune into a bypass.