Files
deepseek-harness/docs/rfc/012-optional-code-mode.md
Tianyi Cui 1cc6e1caf7 docs: add RFC 012 (optional Code Mode for all tools)
Proposes an optional Code Mode where the model writes a TypeScript program
against a generated SDK wrapping every registered tool, instead of emitting
one native tool-call per step. Implemented Cordis-style as a capability-seam
trio (code-runtime interface / code-runtime-vm node:vm reference stub /
code-mode consumer plugin) with zero core-package changes; the hardened
execution substrate is deferred to a follow-up RFC.
2026-06-15 01:20:09 +08:00

22 KiB

RFC 012: Optional Code Mode — model writes TypeScript against an SDK of all tools

Status: proposed

Problem

Today the agent loop advertises every registered tool to the model as a native JSON-schema function definition. ToolRegistry feeds its schemas into ctx.systemPrompt, the loop puts them on GenerateOptions.tools, and the adapter serializes them to the provider's function-calling wire format. The model then invokes one tool-call block per step, the loop dispatches each call through ctx.tools.execute() sequentially (parallel tool execution is an explicit open TODO in dsh-tools and docs/architecture.md), and every intermediate tool-result re-enters the model's context on the next request.

For multi-step tool work this is token-heavy and serial. The model cannot compose tools — loop over a result set, branch on an intermediate value, fan out, post-process — without a full model round-trip per call, and each of those round-trips drags the entire intermediate result back into context whether the model needs it or not.

Cloudflare's Code Mode (shipped as the @cloudflare/codemode npm package) proposes an alternative grounded in a simple observation: LLMs are better at writing code than at emitting tool calls, because they have seen millions of lines of real code and comparatively few contrived tool-calling traces. Instead of one tool call per step, the model writes a TypeScript program against a generated SDK that wraps all the tools, and that program is executed. The model curates what comes back — only what it console.logs and/or returns — instead of every intermediate result. Because the SDK functions are async, the model can run independent calls concurrently.

This RFC proposes an optional Code Mode for the DeepSeek Harness, covering all tools uniformly — built-in and future MCP — with no per-tool work, implemented Cordis-style with zero core-package changes. It fully specifies the code-execution seam and the SDK-generation pipeline, but ships only a minimal node:vm reference stub for execution; the hardened, sandboxed execution substrate is deferred to a follow-up RFC (see Risks). This RFC does not change the agent loop, and it leaves native tool-calling exactly as it is — Code Mode is a plugin you load, not a replacement.

Proposal

The design follows the codebase's capability-seam pattern (ADR 0009, the bash template) as a three-package split, plus one consumer plugin. Nothing in dsh-session, dsh-agent, dsh-agent-loop, dsh-llm, dsh-tools, or dsh-system-prompt changes.

Prior art. @cloudflare/codemode validates this shape directly and several of its decisions are adopted below. Its Executor interface is deliberately tiny — execute(code, fns) → { result, error?, logs? } — with a production DynamicWorkerExecutor (isolated Workers) and a six-line NodeVMExecutor example as two implementations behind it: exactly the interface/implementation split ADR 0009 prescribes. It generates TypeScript type definitions from tools for the model's context and runs the generated JavaScript in a sandbox, capturing console output alongside the return value. It normalizes model output into an async arrow function via AST parsing (acorn) and sanitizes tool names into valid JS identifiers (my-toolmy_tool, deletedelete_). It blocks outbound network by default. The transferable lessons — minimal executor contract, host-side type generation costing zero prompt tokens, capture-output-and-return-value, name sanitization, AST-normalize the code, isolate by default — are folded into the design below. What does not transfer is the substrate: Cloudflare's isolation is Workers-specific; our equivalent hardened substrate is the deferred follow-up.

1. Interface package packages/code-runtime/ — a new package @deepseek-ai/dsh-code-runtime owning ctx.codeRuntime, depending only on cordis. It defines an abstract CodeRuntime extends Service plus the execution vocabulary. The runtime knows nothing about ctx.tools: it is handed a set of named async functions (the resolved SDK bindings), runs the program, and captures output. The result shape mirrors Cloudflare's proven-minimal contract so an error is a field on a resolved result, not a throw the runtime is expected to make:

  • CodeRunRequest = { code: string; sdk: SdkBinding[]; signal?: AbortSignal }
  • CodeRunResult = { result: unknown; logs: string[]; error?: string }
  • SdkBinding = { namespace: string; fns: Record<string, (args: unknown) => Promise<unknown>> }

Per the "explicit > implicit at seams" convention, the request spells out every field the runtime acts on; defaulting (e.g. an output cap, a timeout derived from signal) is the implementation's explicit job, not a hidden ?? default inside run(). The split into interface + implementation is justified under ADR 0009 because there is genuinely more than one planned implementation — the node:vm stub and the hardened substrate (a real isolate, or the generated program run as a sandboxed process through the existing ctx.bash seam) that is scheduled follow-up work, not speculative optionality. ADR 0009 warns against splitting preemptively when only one implementation is conceivable; here a second is not just conceivable but required before any untrusted use, so the seam earns its keep.

2. Implementation package packages/code-runtime-vm/ — a new package @deepseek-ai/dsh-code-runtime-vm, the node:vm reference stub. It type-erases the model's TypeScript via the compiler's transpileModule (or sucrase) — the types exist only to guide the model; the runtime is plain JS — then wraps the body in an async IIFE for top-level await (Cloudflare's NodeVMExecutor does literally new AsyncFunction("codemode", "return await (${code})()")), runs it in a vm.Context whose globals are a capturing console and the SDK namespace objects, awaits the IIFE, and captures the return value, the buffered logs, and any thrown error (as error: string). It applies an output cap (truncate captured logs) and a timeout tied to request.signal. These caps limit blast radius; they are not a security boundary. node:vm is not isolation: withholding require/process does not contain anything (code escapes via constructor/prototype reflection), and per AGENTS.md the harness must never hand model output the ambient environment. code-runtime-vm is therefore documented as reference / test-only / unsafe-for-untrusted-input, acceptable in the MVP only because the code runs at harness trust. Signal handling is best-effort: it aborts in-flight sub-dispatches but cannot reliably interrupt a hot synchronous loop (while(true){}) in node:vm — another reason the hardened substrate is deferred, not optional-forever.

3. Consumer plugin packages/code-mode/ — a new package @deepseek-ai/dsh-code-mode, the plugin that wires everything together. It declares inject = ['tools', 'systemPrompt', 'codeRuntime'] — Cordis throws on access to a service that is not injected, and keeps the plugin inactive until all three exist (the same pattern as tool-bash's inject = ['tools', 'bash']), which also gives correct load-ordering relative to code-runtime/code-runtime-vm. The plugin contributes four things, all through existing seams:

3a. Tool presentation — a lazy system-prompt section (the injection seam already exists). dsh-system-prompt already provides the Cordis-idiomatic way for any plugin to inject prompt snippets: ctx.systemPrompt.section({ name, order, text }), fiber-scoped and auto-disposed via ctx.effect(), where text may be a lazy () => string re-evaluated at each assembly. No new mechanism is needed or invented. Code Mode registers a lazy section (high order so it lands last) whose thunk reads ctx.tools.schemas() at assembly time and regenerates the SDK .d.ts plus usage instructions from the currently-registered tool set. Because the thunk reads the live registry, coverage of every tool — built-in, MCP, future — is automatic.

3b. Wire tool-list enforcement — an agent/request listener (the authoritative seam). The goal "exactly one tool reaches the wire" must be enforced where the wire request is finalized. The loop calls ctx.systemPrompt.assemble() first, then builds GenerateOptions (seeding tools from assembly.tools), then runs the agent/request waterfall, then calls ctx.llm.stream(). A system-prompt/assemble listener can only influence the seed; agent/request is the last seam before the model call, so it is authoritative. The plugin registers an agent/request listener that does const final = await next(); return { ...final, tools: [runCodeSchema] } — overriding the value returned by next(), not the inbound argument, so it dominates the cooperative request listeners it wraps. It registers with prepend: true to sit at the outer edge of the waterfall chain. One honest caveat, stated in the RFC body: ctx.llm.stream() itself runs a further llm/stream waterfall before the adapter, so the guarantee is "authoritative within the agent request pipeline," not an absolute wire invariant; if a hard invariant is ever required, a defensive llm/stream assertion with a spy adapter covers it in tests.

3c. The single tool — run_code. Registered normally in ctx.tools with one parameter { code: string (required) }. Because it is an ordinary tool, the unchanged loop dispatches it through the normal path — this is the crux of "zero loop changes." Its execute(args, exec):

  1. Builds the SDK bindings. For each real tool, an async invoke(callArgs) that checks exec.signal?.aborted (throwing if set) before and after calling ctx.tools.execute({ callId: <deterministic sub-id>, name, arguments: callArgs, agent: exec.agent, signal: exec.signal }), then maps the resulting ContentBlock[] to a simplified { output, isError } (text blocks for the MVP), and emits an observability event. The explicit abort check matters because ctx.tools.execute() catches thrown tool errors and converts them to isError results — without the check, an aborted sub-call would look like ordinary error data and the program would keep running instead of stopping. Sub-dispatch still flows through the tools/execute waterfall, so permission/sandbox/hook plugins apply to code-mode calls exactly as to native ones.
  2. Calls ctx.codeRuntime.run({ code: args.code, sdk: bindings, signal: exec.signal }).
  3. Surfaces the outcome. A successful run returns [{ type: 'text', text: <console logs + return value> }]. A runtime-error result cannot be reported by returning content, because a normal ToolDefinition.execute() returns only Promise<ContentBlock[]> and ToolRegistry.execute() hardcodes isError: false on any successful return — isError: true arises only from the registry's catch path. So on an error result the tool throws a CodeRunError extends HarnessError (HarnessError is exported from dsh-llm; the registry catch turns any throw into isError: true with the message as text, and a HarnessError additionally carries structured { name, code }). An alternative — registering run_code handling as a tools/execute listener that returns a full ToolExecutionResult and can set isError directly — is noted; the throw is simpler and preferred.

3d. Result discipline — what the model receives. The model gets back only the captured console output and/or the program's return value (the model chooses which to surface). Intermediate sub-call results are never returned to the model. This is the core context-saving benefit: the agent curates its own output, exactly as a script's stdout curates a pipeline's intermediate state.

Sub-call CallIds. Real tool calls dispatched from inside run_code need ids, but CallId is normally provider-issued (a branded string for correlating a call with its result — only brand-wrapped via CallId(), with no generator and no documented session-global-uniqueness guarantee). The plugin mints deterministic sub-ids scoped to the parent: `${exec.callId}:code:${n}` with a per-run counter n. These are unique within one run_code run (assuming the parent callId is unique, which the provider guarantees per turn); the code/dispatch event additionally carries the session log's seq so the UI and persistence can order and disambiguate globally without relying on the id alone. ToolExecution.agent is optional; the normal loop always supplies it (and with it exec.agent.session, the log code/dispatch appends to). A run_code execution arriving without exec.agent still runs (sub-calls propagate agent: undefined, exactly as the loop's own contract allows) but skips session-log observability — with no session to append to, those direct runs are simply not logged.

Observability without context cost. Each sub-dispatch emits a session event declared by the dsh-code-mode plugin itself via SessionEventMap declaration merging (the map is merge-extensible precisely so plugins can add events without touching dsh-session). Shape: code/dispatch with { parentCallId, subCallId, name, arguments (or redacted), isError, summary }, ordered by the session log's own seq. deriveMessages() does not translate it into a model message — an unknown event type falls through its default, per the merge-extensible-union convention — so the UI and persistence (RFC 009) can render every sub-call while the model's context only ever receives the single run_code tool-result. Because the event lives in the plugin, this adds no core change.

SDK codegen. A pure jsonSchemaToTs(schema) in code-mode maps the JSON-schema subset the defineTool DSL produces (object/string/number/boolean/array, properties, required[], enum → string-literal union, nested objects, array items) to a TS type literal. It is total: any unsupported construct ($ref, oneOf/anyOf, integer, null, additionalProperties, or any raw MCP shape it does not recognize) degrades to unknown without throwing — it never crashes codegen. Typing is best-effort, not a guarantee, because MCP tools accept arbitrary JSON Schema and ToolSchema.parameters is typed only as Record<string, unknown>. Because ToolSchema.name is an arbitrary string (not necessarily a valid TS identifier), the SDK is generated as a namespace with quoted access (e.g. tools["some-mcp-tool"](args)) plus safe camelCase aliases where the name is a clean identifier; alias collisions and TS reserved words fall back to quoted-only access (no duplicate alias emitted). This mirrors Cloudflare's sanitizeToolName. run_code itself is filtered out of the SDK. The MVP surfaces text content only; image and other block types in sub-results are deferred (noted as a limitation).

Concurrency — possible, not guaranteed safe. The SDK functions are async, so the model can write await Promise.all([...]) and Code Mode would dispatch those calls concurrently — the parallel-tools capability the loop itself still lacks. But the tool contract carries no concurrency-safety metadata today, and parallel tool execution is an open TODO in dsh-tools. So this RFC does not claim parallel SDK calls are free or naturally safe. MVP policy: document the capability with a footgun warning (concurrent calls share the same tools/execute seam and may race inside a not-yet-hardened tool), and make a future tool-metadata flag (read-only / concurrency-safe) the prerequisite for endorsing it. The MVP may serialize sub-dispatches; the SDK .d.ts can annotate which tools are safe to parallelize once that metadata exists.

Tool visibility tiers (design intentionally skipped). A natural extension is to mark each tool with a visibility tier: some tools "direct-call eligible" (still offered as native wire tools alongside run_code), some "code-mode only" (reachable solely from within a run_code program, never on the wire), and the default "both." This would let a deployment keep a few high-frequency or approval-gated tools as direct calls while routing the long tail through Code Mode, or hide composition-only primitives from the native surface entirely. This RFC notes the possibility but intentionally skips the detailed design — the per-tool metadata, how it interacts with the agent/request enforcement in 3b, and the presentation split in 3a are left to a follow-up. The MVP is the simple two-state model: Code Mode on (everything via run_code) or off (everything native).

Optionality / toggle. Loading the code-mode plugin enables Code Mode for that context; not loading it leaves today's native tool-calling untouched. The two are mutually exclusive within one ctx, because Code Mode rewrites the wire tool list down to [run_code]. Per-agent selection via ctx forks, and the visibility tiers above, are future work; the MVP toggle is plugin presence.

Plan

  1. Scaffold the interface package packages/code-runtime/ per the cookbook: abstract CodeRuntime extends Service (super(ctx, 'codeRuntime')), the declare module 'cordis' ctx key, the CodeRunRequest/CodeRunResult/SdkBinding vocabulary, method contracts documented in JSDoc (what run captures, abort semantics, that an error is a result field not a throw). HMR-safety test (dispose the contributing fiber, assert ctx.codeRuntime is gone).
  2. Scaffold the implementation package packages/code-runtime-vm/: the node:vm stub — transpile/type-erase, async-IIFE wrap, capturing console, SDK globals, return-value/logs/error capture, output cap, signal-tied timeout. Tests for output capture, return value, error-as-field, abort, and a README documenting the "not a sandbox, trusted-only" caveat prominently.
  3. Scaffold the consumer plugin packages/code-mode/: jsonSchemaToTs codegen with namespace/quoted-access + alias handling (unit tests, including non-identifier MCP names and unsupported-shape → unknown); the registered lazy ctx.systemPrompt.section() carrying the SDK .d.ts; the agent/request listener (prepend: true) collapsing request.tools to [run_code] after await next(); the run_code tool with the dispatch bridge (deterministic sub-call ids, before/after abort checks, CodeRunError on error results); and the code/dispatch event declared here via SessionEventMap merge. Declare inject = ['tools', 'systemPrompt', 'codeRuntime'].
  4. Tests: HMR-safety (dispose removes the tool, the section, and the listener); a waterfall test that the wire tool list is exactly [run_code] (spy adapter, asserting via agent/request and optionally llm/stream); an integration test that a program calling two tools returns only its printed/returned output (verify the world, not the self-report); deriveMessages() ignores code/dispatch; abort mid-program stops further dispatches; CodeRunError surfaces as isError: true.
  5. Wire an example: examples/coding-agent-code-mode (or a config flag on the existing example) loading the trio. Any example that runs a real model through code-runtime-vm is marked explicitly unsafe, or uses a mock model — never a real model against the unsandboxed stub by default. Add a yarn demo:* entry.
  6. Docs: update docs/architecture.md (a ctx.codeRuntime row in the service map, a Code Mode note under the tool pipeline / capability seams sections); add a cookbook note on writing a CodeRuntime backend; and file the follow-up RFC for the hardened execution substrate (the isolate/sandboxed-process design, plus the tool-visibility-tier design skipped here). Append the | 012 | … | proposed | row to the RFC index.

Risks

node:vm is not a sandbox. This is the single biggest caveat. Withholding require/process is not a boundary; the MVP runs at harness trust only; the hardened substrate is a hard prerequisite before any untrusted use and is the explicit subject of a follow-up RFC. The reference stub and every README around it say so loudly.

Wrong seam would leak tools. If the wire tool list were enforced only in system-prompt/assemble, a later agent/request listener could re-add tools. Mitigation: enforce request.tools = [run_code] in the agent/request waterfall (the authoritative seam, run last before llm.stream()) with prepend: true, and assert exactly one wire tool in tests. The residual llm/stream caveat is documented, not hidden.

Concurrency is not yet safe. The RFC does not endorse parallel SDK calls as free; it documents the footgun and makes a tool concurrency-safety metadata flag the prerequisite. The MVP may serialize sub-dispatches.

Two presentation modes to keep coherent. A tool added later must work in both native and Code Mode. Mitigation: both the codegen thunk and the agent/request listener read ctx.tools.schemas(), so coverage is automatic; a test asserts every registered schema produces valid .d.ts, including non-identifier MCP names via quoted access.

Type-erased runtime is not type-checked. The model can write code that type-checks against the advisory .d.ts but throws at runtime, and MCP-schema typing is best-effort. Mitigation: errors are captured as CodeRunResult.error and surfaced so the model can self-correct; the .d.ts is explicitly advisory.

Lost observability of sub-calls. Routing everything through one run_code result hides the individual calls from the model — and could hide them from operators too. Mitigation: the plugin-declared code/dispatch event keeps every sub-call in the session log and UI without polluting model context.

Abort granularity. node:vm cannot reliably interrupt hot synchronous code, and ctx.tools.execute() converts thrown aborts into isError data. Mitigation: the SDK bindings check signal.aborted and throw before/after each dispatch so an aborted sub-call stops the program; the vm stub wraps the run in a signal-tied timeout; the hardened substrate addresses the hot-loop case.

Unsafe example wiring. A demo running a real model through the node:vm stub would hand model output ambient authority. Mitigation: examples are mock-model or explicitly marked unsafe; code-runtime-vm is labeled reference/test-only.

Non-text sub-results dropped in the MVP. Image and other block types from sub-calls are not surfaced into the program yet. Mitigation: noted as a known limitation; block-type handling deferred.