mirror of
https://github.com/deepseek-ai/deepseek-harness
synced 2026-08-15 21:04:50 +00:00
docs: unwrap hard-wrapped Markdown to one line per paragraph
Hard line breaks mid-paragraph make docs harder to edit and diff — a one-word change reflows and re-diffs the whole paragraph. Reflow all tracked non-vendor Markdown (plus vendor/AGENTS.md) so each prose paragraph is a single line; soft-wrapping is the editor's job. Fenced code, tables, and list structure are preserved (wrapped list items fold to one line per bullet). Documents the convention in AGENTS.md.
This commit is contained in:
@@ -1,8 +1,6 @@
|
||||
# Cookbook: adding a tool
|
||||
|
||||
How to give the model a new capability. Reference implementations:
|
||||
`examples/echo-agent/src/echo-tool.ts` (minimal) and
|
||||
`packages/tool-bash` (production-grade, three-package seam).
|
||||
How to give the model a new capability. Reference implementations: `examples/echo-agent/src/echo-tool.ts` (minimal) and `packages/tool-bash` (production-grade, three-package seam).
|
||||
|
||||
## The minimal shape
|
||||
|
||||
@@ -30,33 +28,18 @@ export function apply(ctx: Context) {
|
||||
}
|
||||
```
|
||||
|
||||
Registration is effect-based: disposing the plugin fiber unregisters the
|
||||
tool (write the HMR test). Schemas flow into the system-prompt assembly
|
||||
automatically.
|
||||
Registration is effect-based: disposing the plugin fiber unregisters the tool (write the HMR test). Schemas flow into the system-prompt assembly automatically.
|
||||
|
||||
## Rules of the execute() contract
|
||||
|
||||
- **Validate args at runtime.** `defineTool`'s `InferArgs` typing is
|
||||
compile-time only; at runtime `arguments` is whatever JSON the model
|
||||
emitted. Check every field; throw a descriptive Error for bad input.
|
||||
- **Throwing means isError.** The registry catches anything `execute()`
|
||||
throws and returns `{isError: true}` to the model. Use that for
|
||||
infrastructure failures (bad input, spawn errors, aborts) — but REPORT
|
||||
domain failures in the result text instead (e.g. tool-bash returns
|
||||
`[exit code: 9]` with `isError: false`: the model decides what a failing
|
||||
command means).
|
||||
- **Validate args at runtime.** `defineTool`'s `InferArgs` typing is compile-time only; at runtime `arguments` is whatever JSON the model emitted. Check every field; throw a descriptive Error for bad input.
|
||||
- **Throwing means isError.** The registry catches anything `execute()` throws and returns `{isError: true}` to the model. Use that for infrastructure failures (bad input, spawn errors, aborts) — but REPORT domain failures in the result text instead (e.g. tool-bash returns `[exit code: 9]` with `isError: false`: the model decides what a failing command means).
|
||||
- **Honor `exec.signal`.** Cancel in-flight work when it fires.
|
||||
- **Use `exec.agent` for async notifications.** `agent.inject(content,
|
||||
{source: {kind: 'plugin', plugin: '<name>'}})` appends durable context the
|
||||
NEXT model request sees — it is not a wake-up (an idle agent stays idle).
|
||||
Guard against disposed agents (try/catch).
|
||||
- **Use `exec.agent` for async notifications.** `agent.inject(content, {source: {kind: 'plugin', plugin: '<name>'}})` appends durable context the NEXT model request sees — it is not a wake-up (an idle agent stays idle). Guard against disposed agents (try/catch).
|
||||
|
||||
## Long-running work
|
||||
|
||||
Follow tool-bash's background pattern: a `run_in_background` flag returns a
|
||||
task id immediately; companion tools poll incrementally and kill; completion
|
||||
notices arrive via `agent.inject()`. Bound buffers and spill full output to
|
||||
disk so nothing is silently lost.
|
||||
Follow tool-bash's background pattern: a `run_in_background` flag returns a task id immediately; companion tools poll incrementally and kill; completion notices arrive via `agent.inject()`. Bound buffers and spill full output to disk so nothing is silently lost.
|
||||
|
||||
> TODO: each tool reimplements this background pattern by hand today. At some
|
||||
> point we need a generic long-running-tool layer that handles task ids,
|
||||
@@ -64,14 +47,8 @@ disk so nothing is silently lost.
|
||||
|
||||
## Permissions / sandboxing
|
||||
|
||||
Prefer not to build policy into the tool. The seam is the `tools/execute` waterfall
|
||||
(veto or wrap — see the permission-gate example in docs/architecture.md), or
|
||||
a sandboxing implementation behind the tool's executor seam.
|
||||
Prefer not to build policy into the tool. The seam is the `tools/execute` waterfall (veto or wrap — see the permission-gate example in docs/architecture.md), or a sandboxing implementation behind the tool's executor seam.
|
||||
|
||||
## Tests every tool needs
|
||||
|
||||
Arg-validation rejections, result shaping for every outcome, the HMR
|
||||
disposal test, and — for tools with side effects — an integration spec that
|
||||
drives the tool through the agent loop with a scripted `MockAdapter`
|
||||
(`packages/agent-loop/tests/mock-adapter.ts`), asserting the `tool/call` /
|
||||
`tool/result` session events.
|
||||
Arg-validation rejections, result shaping for every outcome, the HMR disposal test, and — for tools with side effects — an integration spec that drives the tool through the agent loop with a scripted `MockAdapter` (`packages/agent-loop/tests/mock-adapter.ts`), asserting the `tool/call` / `tool/result` session events.
|
||||
|
||||
Reference in New Issue
Block a user