Files
deepseek-harness/docs/adr/0015-structured-error-taxonomy.md
Tianyi Cui 825b57aff9 feat(llm): structured error taxonomy with a shared HarnessError base (RFC 005 pt 2)
Introduce HarnessError in dsh-llm (the leaf package): a stable machine-routable
code distinct from the message, cause chaining, name from the subclass, plus
isHarnessError. LlmError, ToolArgsError, and InvariantError now extend it.

Tool failures carry the structure end-to-end: ToolExecutionResult gains
error: { name, code } (populated from a thrown HarnessError), and the loop
forwards it onto the tool/result session event (which gained the same optional
field) for retry/sandbox plugins and replay. The loop's toError wraps non-Error
throws in a HarnessError(code: UNKNOWN, cause) instead of a bare Error.

Landed last and in isolation so it's a pure upgrade over the plain Error+code
the earlier PRs used — independently revertible. Graduates RFC 005 pt 2 ->
ADR 0015; RFC 005 now fully implemented.
2026-06-14 01:07:28 +08:00

2.5 KiB

ADR 0015: Structured error taxonomy

Status: accepted (2026-06-14)

Context

Failures crossed seams as bare strings. A tool error flattened to a text block — name, code, and stack lost — so a future sandbox/retry plugin couldn't tell ENOENT from EACCES, and the model got less actionable feedback than it could. A non-Error throw degraded further: the loop wrapped it in new Error(String(x)), dropping any code. And LlmError was the only typed error in the system, with no shared base, so there was nothing for a consumer to instanceof against generically.

This is the last of the RFC 005 pieces and the one the user was most skeptical of, so it was deliberately built last and in isolation: the earlier PRs (arg validation, dev invariants) threw plain Errors with a code field, decoupled from any shared base, so this change is a pure upgrade and is independently revertible without unpicking them.

Decision

A single HarnessError extends Error base in dsh-llm (the leaf package every other imports — no new dependency edge): a stable code distinct from message, cause chaining via ErrorOptions, and name defaulting to the subclass. isHarnessError narrows at seams.

  • LlmError, ToolArgsError (dsh-tools), and InvariantError (dsh-invariants) now extend it, keeping their existing codes.
  • ToolExecutionResult gains optional error: { name, code }, populated in the registry's catch when the thrown value is a HarnessError. The agent loop forwards it onto the tool/result session event (which gained the same optional field), so the structured failure survives into the log for retry/sandbox plugins and replay. The model-facing text block is unchanged.
  • The loop's toError wraps a non-Error throw in a HarnessError (code: 'UNKNOWN', original chained as cause) instead of a bare Error, so even a bad throw carries a routable code into the session error event (which already surfaced code).

Consequences

  • Errors are machine-routable end-to-end: a plugin can branch on error.code rather than substring-matching a message.
  • One base class is imported widely, but it lives in the package everyone already depends on, so the cost is a single import, not a new edge.
  • deriveMessages does not surface error into model history — the model still sees the text block; the structured field is for code and replay.
  • Reverting this PR returns the earlier errors to plain Error+code form; nothing else in the stack depends on the shared base.