Files
deepseek-harness/packages/tools
Tianyi Cui 7803c38824 feat(acp): tool-owned tool-call UI presentation (title/command/output)
In Zed the tool-call card showed only "bash" — the bare tool name — instead
of what the command does. Fix it by letting each TOOL own how its calls render,
rather than the bridge special-casing names.

dsh-tools: add an optional two-state presentation seam to ToolDefinition /
defineTool — `presentCall(args)` (pending: title, kind, rawInput) and
`presentResult(args, result)` (completed: title?, content?). Provider-neutral
`ToolCallKind`/`ToolCallPresentation`/`ToolResultPresentation` vocabulary so
tools never depend on ACP. defineTool soft-validates args (display runs on log
replay, so a malformed/old shape returns undefined instead of throwing).

dsh-tool-bash: bash declares presentCall (model `description` → title, exact
`command` → rawInput, kind execute) and presentResult (wrap output in a fenced
```console block — a UI-only affordance kept out of the model-facing result);
bash_output/bash_kill present task-scoped titles.

dsh-acp: inject `tools`; a per-session `ToolPresenter` looks the tool up by name
and maps its neutral presentation to the ACP tool_call/tool_call_update wire
shape, with a generic fallback (title = name) for tools that declare nothing.
Because the `tool/result` event carries only {callId, content, isError}, the
presenter keeps a small bridge-local map of ONLY in-flight calls' (name, args),
keyed by callId and removed as each result is presented — no event-schema or
core change. Replay uses a throwaway presenter so loaded sessions render
identically to live ones.

Tests: dsh-tools defineTool presenters (typed args, soft-validate), tool-bash
bash/bash_output/bash_kill presenters, acp ToolPresenter (tool-owned mapping,
unknown-callId fallback, in-flight-only map), and an end-to-end turn through the
bridge. The key-gated e2e now asserts a real bash call's title is the model
description (not "bash") and rawInput is the command — verified against the real
DeepSeek model. The test harness derives its inject from the bridge's exported
`inject` so it can't drift again.
2026-06-18 09:01:36 +08:00
..
2026-06-16 14:55:37 +08:00

dsh-tools

Tool registry and execution waterfall. Tool plugins register their schemas and executors; the agent loop executes calls through the tools/execute waterfall.

Service: ToolRegistry (ctx key: tools)

Public API

  • ctx.tools.register(definition: ToolDefinition): () => void Register a tool. Disposed with the calling fiber.
  • ctx.tools.get(name: string): ToolDefinition | undefined
  • ctx.tools.schemas(): ToolSchema[] Schemas of all registered tools (without the execute functions).
  • ctx.tools.execute(exec: ToolExecution): Promise<ToolExecutionResult> Execute one tool call through the tools/execute waterfall.

Injected services

SystemPrompt — the registry automatically feeds its tool schemas into the system-prompt assembly via ctx.systemPrompt.tools().

Events

Event Mode Purpose
tools/execute waterfall Wrap/veto tool execution (sandbox, permission, hooks, plan mode)
tools/change emit A tool was registered or unregistered

Key types

  • ToolDefinitionToolSchema + execute(args, exec): Promise<ContentBlock[]>, plus optional presentCall(args) / presentResult(args, result) for tool-owned UI presentation (see below).
  • ToolExecution — one pending tool call: { callId, name, arguments, agent?, signal? }.
  • ToolExecutionResult — outcome: { callId, content, isError, error? }. On failure with a HarnessError, error: { name, code } carries the structured failure class alongside the model-facing text (the loop forwards it onto the tool/result session event for retry/sandbox plugins and replay).
  • ToolCallPresentation / ToolResultPresentation — provider-neutral shapes a tool returns from presentCall / presentResult to own how a UI renders ITS calls (see "Tool-owned UI presentation").

Extension points

  • Tool plugins call ctx.tools.register() — schemas flow into the assembly automatically.
  • The tools/execute waterfall is the single seam for sandbox, permission, hooks, and plan-mode plugins to wrap or veto a call. Listeners receive (exec, next): call next() to proceed, or return a result without calling next() to short-circuit (veto).
  • MCP servers: one plugin per server, discover tools, call ctx.tools.register() with the server's schemas.

Typed tool parameter schemas

First-party plugin authors can use the defineTool() helper (exported from this package) for typed tool parameter schemas:

import { readFile } from 'node:fs/promises'
import type { Context } from 'cordis'
import { defineTool } from '@deepseek-ai/dsh-tools'

declare const ctx: Context

ctx.tools.register(defineTool({
  name: 'read_file',
  description: 'Read a file from disk.',
  parameters: {
    path: { type: 'string', required: true, description: 'Absolute file path' },
    offset: { type: 'number' },
    limit: { type: 'number' },
  },
  async execute(args, exec) {
    // args is typed: { path: string; offset?: number; limit?: number }
    const text = await readFile(args.path, 'utf8')
    return [{ type: 'text', text }]
  },
}))

The helper converts the author-facing SchemaSpec (with required: true as a per-property boolean) to standard JSON Schema for the wire format. Raw JSON-Schema tool definitions (from MCP servers) are still accepted by the registry directly.

A defineTool tool also validates the model-generated arguments against its SchemaSpec before execute runs (validateArgs). The model's JSON is untrusted — InferArgs<S> is a compile-time claim, not a runtime guarantee — so on a mismatch (missing required key, wrong primitive, bad enum member, nested violation) the tool throws a ToolArgsError (code: 'INVALID_ARGS'); the registry turns it into an isError result whose text lists the violations, which the model sees and self-corrects from. Validation mirrors the JSON Schema conversion exactly: extra keys are allowed, default is not applied, and an object/array prop without properties/items only type-checks. Raw-registered tools (MCP) are not validated by the harness — they validate their own input.

See defineTool, validateArgs, ToolArgsError, SchemaSpec, InferArgs, and schemaSpecToJsonSchema in the public API for details.

Tool-owned UI presentation

A tool owns how ITS calls render in a UI (an editor's tool-call card, a CLI log line) — a UI plugin must NOT special-case tool names. A ToolDefinition may declare two optional, pure, display-only methods:

  • presentCall(args): ToolCallPresentation | undefined — the PENDING state: a human-readable title (always-visible label), an optional kind (read/edit/execute/… for icon/treatment, default other), and an optional rawInput (the salient input to show in a detail view — e.g. a shell command as a string, NOT the whole args object).
  • presentResult(args, result): ToolResultPresentation | undefined — the COMPLETED state, given the same args and the { content, isError } result: an optional replacement title and reformatted content (e.g. wrap command output in a fenced ```console block — a UI-only affordance that must NOT appear in the model-facing execute result).

Returning undefined (or omitting a method) tells a UI to fall back to a generic presentation (title = tool name, raw args as input, raw result content). Both methods must be pure and side-effect-free: a UI may call them during live streaming AND during a session-log replay, so they depend only on their arguments. With defineTool, args is the typed InferArgs<S> shape; the helper soft-validates before calling (a malformed/older logged arg shape yields undefined rather than throwing, since display must never crash a replay). The shapes are provider-neutral — the ACP bridge (dsh-acp) maps them to ACP tool_call/tool_call_update wire fields, and dsh-tool-bash is the reference implementation.

import { defineTool } from '@deepseek-ai/dsh-tools'

const bash = defineTool({
  name: 'bash',
  description: 'Run a shell command.',
  parameters: {
    command: { type: 'string', required: true, description: 'The command to run.' },
    description: { type: 'string', required: true, description: 'One-line summary shown in the UI.' },
  },
  async execute(args) {
    return [{ type: 'text', text: `ran: ${args.command}` }]
  },
  // The model-written description is the readable title; the command is the detail.
  presentCall: args => ({ title: args.description, kind: 'execute', rawInput: args.command }),
  // Wrap the output as a console block for the UI (not in the model-facing result).
  presentResult: (_args, result) => {
    const block = result.content.length === 1 ? result.content[0] : undefined
    if (block === undefined || block.type !== 'text') return undefined
    return { content: [{ type: 'text', text: '```console\n' + block.text + '\n```' }] }
  },
})

What is NOT here (TODO)

  • Tool shapes review — when real tools land (e.g. a concurrency-safety hint for parallel execution); phase 1 executes tool calls sequentially.
  • Parallel execution — the loop currently iterates tool calls sequentially.