- AgentLoop.resume uses `this.ctx.get('sessionPersistence')` (strict) instead
of the `, false` overload: still topology-independent, but an inactive/
absent backend reads as undefined (rejected by the existing guard) rather
than being handed back mid-teardown.
- Correct the bridge teardown comment: an ACP-created agent's registry entry
binds to the BRIDGE fiber (the factory is reached through the bridge's
traceable proxy, so AgentLoop.start's `this.ctx.effect` registration uses the
caller context), not the AgentLoop fiber — so an ACP-only HMR dispose
reclaims it. Add a regression test pinning that ownership.
- Sync the ctx.get guidance in the post-mortem, packages/AGENTS.md, and the
dsh-code-review skill to the strict form.
Post-mortems
Incident write-ups: a bug reached a place it shouldn't have (a real user, a merged PR, a release), and the interesting part is why our process let it through, not just the one-line fix.
A post-mortem is NOT an RFC (which records a deliberate design decision and its rejected alternatives, or proposes future work). It is a backward-looking record of a failure: what broke, the mechanism, why every safety net missed it, and the concrete guardrails added so the same class of bug fails loudly next time.
Write one when a bug is subtle (the mechanism is non-obvious and a careful engineer would re-derive it the hard way), systemic (the reason it escaped is a gap in tests/tooling/conventions, not a one-off typo), and costly to rediscover (it cost real debugging time, and would cost it again). Link the guardrails (tests, AGENTS.md rules, ADRs) the post-mortem motivated.
Every post-mortem opens with an Executive summary: one short paragraph a busy reader can absorb in thirty seconds — what broke, the root cause in plain terms, why it escaped, and the durable lesson — before the detailed Summary / Timeline / Root cause / Guardrails sections that follow.
| # | Title |
|---|---|
| 0001 | ACP server crashed on connect: export default dropped the plugin's inject |