Files
deepseek-harness/docs/core-data-structures/persistence.zh.md
Ziya 5270dcd61d docs(i18n): core-data-structures and postmortem batch — 22 bilingual pairs
core-data-structures 18 篇(core.md 因超长仍在产出、随后补)、
postmortem 3 篇与 RFC 前门 README 配对;流水线 + 二遍校验产出。
生成文件 docs/rfc/INDEX.md(gen-rfc-index 产物)列入排除。中文侧
页内锚点统一指向英文侧锚名,满足配对门禁的链接目标一致规则。
2026-07-15 23:11:25 -07:00

6.3 KiB
Raw Blame History

会话持久化

English | 中文

事件日志的持久性 seamsession.md 描述了内存中的 Session:仅追加的 SessionEvent 日志即为真源。本页描述该日志如何被持久化:抽象的 SessionPersistence 服务、它的后端、flush 检查点、崩溃恢复,以及随日志一起存储的元数据头。日志所承载的事件词汇在生成的持久化日志事件目录中逐一列出。

该 seam 是教科书式的能力 seam:一个抽象服务(dsh-session-persistencectx.sessionPersistence)在既有的 SessionEvent 之上定义 create/append/load/list——没有平行的持久化类型——以及两个可互换的后端,它们通过同一套 runPersistenceContract 测试。见 session-persistence RFC

flush 检查点

session/event 是一个同步通知持久化插件对其进行缓冲write-behind并在 agent loop 于每个轮次结束时触发的 session/flush 检查点处排空缓冲区。flush 使用 ctx.parallel(被 await一个轮次的事件在下一个轮次开始前被持久提交轮次边界即提交边界。flush 失败时通过 agent/error 和 logger 报告,而非作为会话事件(那样会落在提交边界之后),因此后端保留其缓冲事件等待下一次 flush。

崩溃恢复保留被中断的轮次

后端重新加载一个在轮次中途崩溃的日志时,会发现一个已打开的 turn/start 而没有对应的 turn/end。它不会截断日志:在长周期任务中,单个轮次可能非常大(许多步骤、大量工具输出),而这些事件在崩溃前已被持久追加。后端改为用一个合成的 turn/end { reason: { kind: 'interrupted' } } 关闭这个遗留轮次,保持日志平衡与轮次封闭不变式完好。interrupted 是唯一一个 agent loop 不会自行发出的 TurnEndReason(见 session.md)。

SessionHeader:日志旁的元数据

每个会话的元数据与事件日志分开存储格式版本、cwd、血缘关系和 seed 边界属于存储关注点而非对话事件,因此它们不在 SessionEventMap 中,也不会进入 deriveMessages()。header 通过 session.header 附加到 Session 上。

源码:packages/core/session/src/types.ts

interface SessionHeader {
  /**
   * On-disk format version, stamped from {@link SESSION_FORMAT_VERSION} when the
   * session is created. A persistence backend rejects any other version on load
   * (no migration — see the constant).
   */
  readonly version: number
  /** The session's id (mirrors the {@link Session}'s id). */
  readonly id: SessionId
  /** Unix epoch milliseconds when the session was created. */
  readonly createdAt: number
  /** Absolute working directory the session was created in (if any). */
  readonly cwd?: string
  /** The session this one was forked from (seed lineage), if any. */
  readonly parentSession?: SessionId
  /**
   * How many leading events were INHERITED via a seed rather than produced by
   * this session — the seed boundary. Set when a fork seeds a child with a
   * prefix of the parent's log (= the seeded prefix length); absent/0 means the
   * session produced all its own events. Persisted so a reload reconstructs the
   * boundary instead of re-deriving it from the full stored log, and so a replay
   * harness can skip the inherited prefix when deriving the child's OWN script
   * (the seeded events are the parent's, not this child's model calls).
   */
  readonly seedLength?: number
}

CreateSessionOptionsseed 与元数据

通过 store 创建 Session 时接受 seed(回放/fork 一个已有事件日志)和 metastore 折叠进 SessionHeader 的存储级字段。store 填充 version/id 并为 createdAt 设默认值;调用方提供经过校验的绝对路径 cwdparentSession 血缘、seedLength seed 边界,以及仅在重建持久化会话时提供的原始 createdAt 以保留它。

interface CreateSessionOptions {
  /** Events to seed the new session with (replay/fork). */
  readonly seed?: readonly SessionEvent[]
  /**
   * Creation metadata. The store fills in `version`/`id` and defaults
   * `createdAt` to now; the caller supplies the storage-level fields (validated
   * absolute `cwd`, `parentSession` lineage, the seed boundary `seedLength`, and
   * — when reconstructing a persisted session — the original `createdAt` to
   * preserve it).
   *
   * `seedLength` is EXPLICIT, not inferred from `seed.length`: a reconstruction
   * (resume/load) seeds the WHOLE stored log, so its `seed.length` is the full
   * length, not the original boundary — the caller must pass the persisted
   * boundary back. A fresh fork passes its actual seeded-prefix length.
   */
  readonly meta?: {
    readonly cwd?: string
    readonly parentSession?: SessionId
    readonly createdAt?: number
    readonly seedLength?: number
  }
}

因此,回放/fork 是 ctx.sessions.create(id, { seed: seedEvents });将一个持久化会话恢复为活跃 agent 是 ctx.agents.resume({ resumeSessionId })

后端

两个后端实现同一个抽象 SessionPersistence(在 SessionEvent 之上的 create/append/load/list并通过 runPersistenceContract,证明该 seam 真正与后端无关:

  • dsh-session-persistence-jsonl:每个会话一个仅追加的 JSONL 日志,具备崩溃安全的原子写入、上述中断轮次崩溃恢复,以及读取/回放路径。
  • dsh-session-persistence-sqlite:基于 node:sqlite,每个 SessionEvent 一行。行结构 (session_id, seq, type, time, data, source_event_seqs, surface_op) 与事件 1:1 映射(包括可选的 surface 元数据),因此没有需要保持同步的平行持久化 schema。

多个后端共享同一个磁盘会话时,通过共享持久化写协调器协调写入。