Files
deepseek-harness/docs/cookbook/adding-a-tool.md
Tianyi Cui 7c400e9c02 docs: unify ADR/RFC trees into one lifecycle-organized RFC tree
Collapse docs/adr/ and docs/rfc/ into a single docs/rfc/ with proposed/,
implemented/, and rejected/ subfolders. Every file is renamed to
yyyy-mm-dd-topic-title.md, where the date is when the topic was first
proposed (from git history). ADRs and RFCs that covered exactly the same
topic are merged (property-based testing, session persistence); the
umbrella RFC 005 stays split across its three implemented decisions, and
RFC 006's deferred part-3 (API extractor reports) splits into its own
proposed RFC. All cross-references become machine-checkable relative
links instead of bare "ADR NNNN" / "RFC NNN" prose.

Add a verify-md-links doc-sync gate (scripts/verify-md-links.ts) that
checks every relative Markdown cross-link resolves, wired into doc-sync
alongside verify-md-wrap. This makes the reorganization self-verifying:
the same change that rewrote ~forty inter-doc links adds the check that
proves none dangle. Document the cross-link convention in a new
docs/AGENTS.md and record the gate as an implemented RFC.

doc-sync, typecheck, lint, and the full test suite (667) all pass.
2026-06-18 02:18:24 +08:00

3.6 KiB

Cookbook: adding a tool

How to give the model a new capability. Reference implementations: examples/echo-agent/src/echo-tool.ts (minimal) and packages/tool-bash (production-grade, three-package seam).

The minimal shape

import { readFile } from 'node:fs/promises'
import type { Context } from 'cordis'
import { defineTool } from '@deepseek-ai/dsh-tools'

export const name = 'my-tool'
export const inject = ['tools']

export function apply(ctx: Context) {
  ctx.tools.register(defineTool({
    name: 'read_file',
    description: 'Read a file from disk.',          // what the model sees
    parameters: {
      path: { type: 'string', required: true, description: 'Absolute path' },
      limit: { type: 'number' },                     // optional by default
    },
    async execute(args, exec) {
      // args is TYPED from the schema: { path: string; limit?: number }
      // exec carries { callId, name, arguments, agent?, signal? }
      return [{ type: 'text', text: await readFile(args.path, 'utf8') }]
    },
  }))
}

Registration is effect-based: disposing the plugin fiber unregisters the tool (write the HMR test). Schemas flow into the system-prompt assembly automatically.

Rules of the execute() contract

  • Args are validated for you. defineTool validates the model-generated arguments against the SchemaSpec before execute runs (type, required keys, enum membership, nested objects/arrays — runtime arg validation), so inside execute the args already match InferArgs. You still hand-check value constraints the DSL can't express (non-empty strings, positive numbers, cross-field rules); throw a descriptive Error for those. Raw JSON-Schema tools registered directly (MCP) are NOT validated by the harness — they validate their own input.
  • Throwing means isError. The registry catches anything execute() throws and returns {isError: true} to the model. Use that for infrastructure failures (bad input, spawn errors, aborts) — but REPORT domain failures in the result text instead (e.g. tool-bash returns [exit code: 9] with isError: false: the model decides what a failing command means).
  • Honor exec.signal. Cancel in-flight work when it fires.
  • Use exec.agent for async notifications. agent.inject(content, {source: {kind: 'plugin', plugin: '<name>'}}) appends durable context the NEXT model request sees — it is not a wake-up (an idle agent stays idle). Guard against disposed agents (try/catch).

Long-running work

Follow tool-bash's background pattern: a run_in_background flag returns a task id immediately; companion tools poll incrementally and kill; completion notices arrive via agent.inject(). Bound buffers and spill full output to disk so nothing is silently lost.

TODO: each tool reimplements this background pattern by hand today. At some point we need a generic long-running-tool layer that handles task ids, incremental polling, kill, and completion notices uniformly.

Permissions / sandboxing

Prefer not to build policy into the tool. The seam is the tools/execute waterfall (veto or wrap — see the permission-gate example in extension-cookbook.md), or a sandboxing implementation behind the tool's executor seam.

Tests every tool needs

Arg-validation rejections, result shaping for every outcome, the HMR disposal test, and — for tools with side effects — an integration spec that drives the tool through the agent loop with a scripted MockAdapter (packages/agent-loop/tests/mock-adapter.ts), asserting the tool/call / tool/result session events.