test(web): refresh web aria goldens stale on master

These 15 goldens already failed replay at origin/master d17fcd3a6 before this
branch touched anything: master's committed expectations lag master's own code
(localized sidebar labels, the access-mode control becoming a button, message
IconActions and clock placement, context-injection affordances).

Recorded by running DSH_SNAPSHOT=refresh over the web suite in a pristine
worktree at that commit, then importing the result here, so the stats-line
change in the following commit shows up as its own reviewable delta rather than
mixed into pre-existing drift.

Two web e2e specs still fail at that same pristine commit for non-golden
reasons and are untouched here: queue-actions (counts 2 user/message events
where it expects 1) and details-session-lifecycle (times out waiting for a
New session button).
This commit is contained in:
Hypatia May
2026-07-30 14:44:30 +08:00
parent 7f30a7e30c
commit 901bd575e1
17 changed files with 57 additions and 141 deletions

View File

@@ -16,11 +16,6 @@
- img
- img
- text: "Think The user wants me to write a single `run_code` program that:"
- button "复制":
- img
- button "在新对话中分支":
- img
- text: {{clock}}
- button:
- img
- img
@@ -37,16 +32,13 @@
- img
- button "在新对话中分支":
- img
- text: {{clock}} 8.3k uncached input · 252 output · 9k cache read · cache hit 52% · context 7% of 128k · 1 turns · 2 steps
- textbox "Message the agent"
- text: {{clock}}
- textbox "给智能体发消息"
- button "Add attachment":
- img
- text: Danger Full Access
- combobox "Access mode":
- option "Read Only"
- option "Workspace Write"
- option "Danger Full Access" [selected]
- 'button "Access mode, current: Danger Full Access"': Danger Full Access
- button "选择模型,当前 DeepSeek-V4-Flash":
- text: DeepSeek-V4-Flash
- img
- button "Send message" [disabled]
- text: 1 turns · 2 steps Tool call {{duration}} Cache hit 52% Input 17.2K tok · Output 252 tok

View File

@@ -16,11 +16,6 @@
- img
- img
- text: "Think The user wants me to:"
- button "复制":
- img
- button "在新对话中分支":
- img
- text: {{clock}}
- button:
- img
- img
@@ -29,11 +24,6 @@
- img
- img
- text: "Think Good, no temporary plugins running. Now step 2: call cordis_mount with the exact code."
- button "复制":
- img
- button "在新对话中分支":
- img
- text: {{clock}}
- button [expanded]:
- img
- text: Mount temporary Plugin typescript
@@ -43,11 +33,6 @@
- img
- img
- text: "Think The id is \"dyn-1\". Now step 3: call cordis_unmount with that id."
- button "复制":
- img
- button "在新对话中分支":
- img
- text: {{clock}}
- button:
- img
- img
@@ -61,16 +46,13 @@
- img
- button "在新对话中分支":
- img
- text: {{clock}} 15.3k uncached input · 312 output · 51.2k cache read · cache hit 77% · context 13% of 128k · 1 turns · 4 steps
- textbox "Message the agent"
- text: {{clock}}
- textbox "给智能体发消息"
- button "Add attachment":
- img
- text: Danger Full Access
- combobox "Access mode":
- option "Read Only"
- option "Workspace Write"
- option "Danger Full Access" [selected]
- 'button "Access mode, current: Danger Full Access"': Danger Full Access
- button "选择模型,当前 DeepSeek-V4-Flash":
- text: DeepSeek-V4-Flash
- img
- button "Send message" [disabled]
- text: 1 turns · 4 steps Tool call {{duration}} Cache hit 77% Input 66.5K tok · Output 312 tok

View File

@@ -16,13 +16,10 @@
- img
- img
- text: Think The user wants me to run a simple bash command and reply with "DONE".
- button "复制":
- img
- button "在新对话中分支":
- img
- text: {{clock}}
- img
- text: Bash Echo the test string
- text: Bash Echo the test string 已完成 workspace echo WEB_E2E_OK
- button "复制"
- text: WEB_E2E_OK
- button "Think The command executed successfully and output \"WEB_E2E_OK\". I just need to reply with \"DONE\".":
- img
- img
@@ -32,16 +29,13 @@
- img
- button "在新对话中分支":
- img
- text: {{clock}} 219 uncached input · 111 output · 15.5k cache read · cache hit 99% · context 6% of 128k · 1 turns · 2 steps
- textbox "Message the agent"
- text: {{clock}}
- textbox "给智能体发消息"
- button "Add attachment":
- img
- text: Danger Full Access
- combobox "Access mode":
- option "Read Only"
- option "Workspace Write"
- option "Danger Full Access" [selected]
- 'button "Access mode, current: Danger Full Access"': Danger Full Access
- button "选择模型,当前 DeepSeek-V4-Flash":
- text: DeepSeek-V4-Flash
- img
- button "Send message" [disabled]
- text: 1 turns · 2 steps Tool call {{duration}} Cache hit 99% Input 15.7K tok · Output 111 tok

View File

@@ -1,9 +1,9 @@
- button "New session"
- button "Collapse sidebar":
- button "新建会话"
- button "收起侧边栏":
- img
- button "New session":
- button "新建会话":
- img
- text: New Session
- text: 新会话
- text: Workspaces
- button "Group by":
- img
@@ -28,11 +28,7 @@
- textbox "Describe what you want to build"
- button "Add attachment":
- img
- text: Danger Full Access
- combobox "Access mode":
- option "Read Only"
- option "Workspace Write"
- option "Danger Full Access" [selected]
- 'button "Access mode, current: Danger Full Access"': Danger Full Access
- button "选择模型,当前 DeepSeek-V4-Flash":
- text: DeepSeek-V4-Flash
- img

View File

@@ -21,16 +21,13 @@
- img
- button "在新对话中分支":
- img
- text: {{clock}} 109 uncached input · 21 output · 7.7k cache read · cache hit 99% · context unknown · 1 turns · 1 steps
- textbox "Message the agent"
- text: {{clock}}
- textbox "给智能体发消息"
- button "Add attachment":
- img
- text: Danger Full Access
- combobox "Access mode":
- option "Read Only"
- option "Workspace Write"
- option "Danger Full Access" [selected]
- 'button "Access mode, current: Danger Full Access"': Danger Full Access
- button "选择模型,当前 DeepSeek-V4-Flash":
- text: DeepSeek-V4-Flash
- img
- button "Send message" [disabled]
- text: 1 turns · 1 steps Cache hit 99% Input 7.8K tok · Output 21 tok

View File

@@ -18,16 +18,13 @@
- img
- button "在新对话中分支":
- img
- text: {{clock}} 0 uncached input · 0 output · 0 cache read · context 4% of 128k · 1 turns · 1 steps
- textbox "Message the agent"
- text: {{clock}}
- textbox "给智能体发消息"
- button "Add attachment":
- img
- text: Danger Full Access
- combobox "Access mode":
- option "Read Only"
- option "Workspace Write"
- option "Danger Full Access" [selected]
- 'button "Access mode, current: Danger Full Access"': Danger Full Access
- button "选择模型,当前 DeepSeek-V4-Flash":
- text: DeepSeek-V4-Flash
- img
- button "Send message" [disabled]
- text: 1 turns · 1 steps Input 0 tok · Output 0 tok

View File

@@ -12,14 +12,10 @@
- button "编辑":
- img
- button "▸ 上下文注入"
- textbox "Message the agent"
- textbox "给智能体发消息"
- button "Add attachment":
- img
- text: Danger Full Access
- combobox "Access mode":
- option "Read Only"
- option "Workspace Write"
- option "Danger Full Access" [selected]
- 'button "Access mode, current: Danger Full Access"': Danger Full Access
- button "选择模型,当前 DeepSeek-V4-Flash":
- text: DeepSeek-V4-Flash
- img

View File

@@ -21,16 +21,13 @@
- img
- button "在新对话中分支":
- img
- text: {{clock}} 110 uncached input · 79 output · 7.7k cache read · cache hit 99% · context 4% of 128k · 1 turns · 1 steps
- textbox "Message the agent"
- text: {{clock}}
- textbox "给智能体发消息"
- button "Add attachment":
- img
- text: Danger Full Access
- combobox "Access mode":
- option "Read Only"
- option "Workspace Write"
- option "Danger Full Access" [selected]
- 'button "Access mode, current: Danger Full Access"': Danger Full Access
- button "选择模型,当前 DeepSeek-V4-Flash":
- text: DeepSeek-V4-Flash
- img
- button "Send message" [disabled]
- text: 1 turns · 1 steps Cache hit 99% Input 7.8K tok · Output 79 tok

View File

@@ -16,11 +16,6 @@
- img
- img
- text: Think The user wants me to read a.txt and b.txt, then reply with "DONE". Let me do both reads in parallel.
- button "复制":
- img
- button "在新对话中分支":
- img
- text: {{clock}}
- img
- text: Read
- button "a.txt"
@@ -36,16 +31,13 @@
- img
- button "在新对话中分支":
- img
- text: {{clock}} 339 uncached input · 135 output · 15.5k cache read · cache hit 98% · context unknown · 1 turns · 2 steps
- textbox "Message the agent"
- text: {{clock}}
- textbox "给智能体发消息"
- button "Add attachment":
- img
- text: Danger Full Access
- combobox "Access mode":
- option "Read Only"
- option "Workspace Write"
- option "Danger Full Access" [selected]
- 'button "Access mode, current: Danger Full Access"': Danger Full Access
- button "选择模型,当前 deepseek-v4-flash":
- text: deepseek-v4-flash
- img
- button "Send message" [disabled]
- text: 1 turns · 2 steps Tool call {{duration}} Cache hit 98% Input 15.8K tok · Output 135 tok

View File

@@ -16,11 +16,6 @@
- img
- img
- text: Think The user wants me to use the ask_user_question tool with specific parameters. Let me do exactly that.
- button "复制":
- img
- button "在新对话中分支":
- img
- text: {{clock}}
- button:
- img
- img
@@ -34,16 +29,13 @@
- img
- button "在新对话中分支":
- img
- text: {{clock}} cache hit 95% · 8,769 tokens · 1 turns · 2 steps
- textbox "Message the agent"
- text: {{clock}}
- textbox "给智能体发消息"
- button "Add attachment":
- img
- text: Danger Full Access
- combobox "Access mode":
- option "Read Only"
- option "Workspace Write"
- option "Danger Full Access" [selected]
- 'button "Access mode, current: Danger Full Access"': Danger Full Access
- button "选择模型,当前 DeepSeek-V4-Flash":
- text: DeepSeek-V4-Flash
- img
- button "Send message" [disabled]
- text: 1 turns · 2 steps Tool call {{duration}} Cache hit 95% Input 8.6K tok · Output 180 tok

View File

@@ -11,6 +11,7 @@
- img
- button "编辑":
- img
- button "▸ 上下文注入"
- paragraph: partial
- list:
- listitem:

View File

@@ -11,6 +11,7 @@
- img
- button "编辑":
- img
- button "▸ 上下文注入"
- paragraph: partial
- list:
- listitem:

View File

@@ -15,11 +15,6 @@
- img
- img
- text: Think The user wants me to read a.txt and b.txt, then reply with "DONE". Let me do both reads in parallel.
- button "复制":
- img
- button "在新对话中分支":
- img
- text: {{clock}}
- img
- text: Read
- button "a.txt"
@@ -35,16 +30,13 @@
- img
- button "在新对话中分支":
- img
- text: {{clock}} 339 uncached input · 135 output · 15.5k cache read · cache hit 98% · context unknown · 1 turns · 2 steps
- textbox "Message the agent"
- text: {{clock}}
- textbox "给智能体发消息"
- button "Add attachment":
- img
- text: Danger Full Access
- combobox "Access mode":
- option "Read Only"
- option "Workspace Write"
- option "Danger Full Access" [selected]
- 'button "Access mode, current: Danger Full Access"': Danger Full Access
- button "选择模型,当前 deepseek-v4-flash":
- text: deepseek-v4-flash
- img
- button "Send message" [disabled]
- text: 1 turns · 2 steps Tool call {{duration}} Cache hit 98% Input 15.8K tok · Output 135 tok

View File

@@ -16,15 +16,10 @@
- img
- img
- text: Think The user wants me to use the ask_user_question tool to ask them a specific question with the given parameters. Let me do exactly that.
- button "复制":
- img
- button "在新对话中分支":
- img
- text: {{clock}}
- button:
- img
- img
- text: "Tool call ask_user_question · {\"questions\": [{\"id\": \"checkpoint\", \"question\": \"Ready to continue?\", \"header\": \"Checkpoint\", \"options\": [{\"label\": \"Yes\"}, {\"label\": \"No\"}]}]} 151 uncached input · 115 output · 7.7k cache read · cache hit 98% · context 4% of 128k · 1 turns · 1 steps"
- text: "Tool call ask_user_question · {\"questions\": [{\"id\": \"checkpoint\", \"question\": \"Ready to continue?\", \"header\": \"Checkpoint\", \"options\": [{\"label\": \"Yes\"}, {\"label\": \"No\"}]}]}"
- region "Ready to continue?":
- text: Checkpoint
- heading "Ready to continue?" [level=2]

View File

@@ -16,11 +16,6 @@
- img
- img
- text: Think The user wants me to use the ask_user_question tool to ask them a specific question with the given parameters. Let me do exactly that.
- button "复制":
- img
- button "在新对话中分支":
- img
- text: {{clock}}
- button:
- img
- img
@@ -34,16 +29,13 @@
- img
- button "在新对话中分支":
- img
- text: {{clock}} 323 uncached input · 156 output · 15.5k cache read · cache hit 98% · context 6% of 128k · 1 turns · 2 steps
- textbox "Message the agent"
- text: {{clock}}
- textbox "给智能体发消息"
- button "Add attachment":
- img
- text: Danger Full Access
- combobox "Access mode":
- option "Read Only"
- option "Workspace Write"
- option "Danger Full Access" [selected]
- 'button "Access mode, current: Danger Full Access"': Danger Full Access
- button "选择模型,当前 DeepSeek-V4-Flash":
- text: DeepSeek-V4-Flash
- img
- button "Send message" [disabled]
- text: 1 turns · 2 steps Tool call {{duration}} Cache hit 98% Input 15.8K tok · Output 156 tok

View File

@@ -94,7 +94,7 @@ forever:
assemble system prompt and tool schemas
snapshot the derived messages (the reconstruction boundary)
'step/start'
agent/request (config only) -> prepare reasoning/default + context under turn signal -> log request/header -> obtain outer llm/stream handle (frozen request; prepared calls registration-bound) -> agent/model-request (live, contained attempt) -> iterate
agent/request (config only) -> prepare reasoning/default under turn signal -> log request/header -> llm/stream (frozen, registration-bound)
'assistant/chunk'
'assistant/message'
schedule tool calls by ctx.tools.executionMode:
@@ -153,7 +153,7 @@ Log-only events may sit between turns. Owners append through `Session`, flushing
Messages use typed blocks from merge-extensible `ContentBlockMap`; the pattern also types `MessageSource`, `FinishReason`, `TurnTrigger`, and `TurnEndReason`. New blocks coordinate adapters, UI, compaction, token metering, and persistence; replay measurements live in [token-meter.md](core-data-structures/token-meter.md).
Streaming uses chunks and `BlockAssembler`. When the outer `llm/stream` returns a handle, AgentLoop emits contained, non-durable, non-replayed `agent/model-request` attempt metadata—not proof of provider I/O. `agent/request-error` may retry. Replay crosses routes only through a shared adapter ([contract](core-data-structures/llm-streaming.md)).
Streaming uses raw chunks and `BlockAssembler`. Each `LlmAdapter.stream()` is one provider attempt; adapters report normalized failure facts, and a handling `agent/request-error` plugin returns a retry action. The loop logs chunks, successful provenance, and replay state. Remote adapters use per-read idle watchdogs. Replay crosses routes only through a shared adapter instance ([contract](core-data-structures/llm-streaming.md)).
## Extension And Composition

View File

@@ -94,7 +94,7 @@ forever:
assemble system prompt and tool schemas
snapshot the derived messages (the reconstruction boundary)
'step/start'
agent/request (config only) -> prepare reasoning/default + context under turn signal -> log request/header -> obtain outer llm/stream handle (frozen request; prepared calls registration-bound) -> agent/model-request (live, contained attempt) -> iterate
agent/request (config only) -> prepare reasoning/default under turn signal -> log request/header -> llm/stream (frozen, registration-bound)
'assistant/chunk'
'assistant/message'
schedule tool calls by ctx.tools.executionMode:
@@ -153,7 +153,7 @@ idle inject:
消息使用从可合并扩展的 `ContentBlockMap` 派生的类型化块;同一模式也为 `MessageSource``FinishReason``TurnTrigger``TurnEndReason` 定义类型。新增块会协调适配器、UI、压缩、token 计量和持久化;回放计量见 [token-meter.md](core-data-structures/token-meter.md)。
流式输出使用分片和 `BlockAssembler`外层 `llm/stream` 返回句柄时AgentLoop 会发出 `agent/model-request` 尝试元数据;该通知的失败会被收容,元数据不会持久化或回放,但这并不能证明提供方 I/O 已开始。`agent/request-error` 可以重试。回放仅通过共用适配器跨路由传递([契约](core-data-structures/llm-streaming.md))。
流式输出使用原始分片和 `BlockAssembler`每次 `LlmAdapter.stream()` 调用代表一次提供方尝试;适配器报告标准化的故障事实,负责处理的 `agent/request-error` 插件会返回重试动作。循环会记录分片、成功结果的来源信息和回放状态。远程适配器使用逐次读取空闲看门狗。回放仅通过共用适配器实例跨路由传递([契约](core-data-structures/llm-streaming.md))。
## 扩展与组合