feat(core): recover truncated tool output via the Session scratchpad (#145)

Co-authored-by: Yaowei Zheng <hiyouga@buaa.edu.cn>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
Jingzhe Xu
2026-08-03 13:17:42 +08:00
committed by GitHub
parent 84949882ab
commit a798bca87f
30 changed files with 1393 additions and 47 deletions
@@ -161,6 +161,8 @@ An existing Agent always runs with its on-disk config verbatim — newer code de
`{{PROJECT_DIR}}` is surfaced to the model as the **App Data Dir**: PenguinHarness's application data root, holding every Agent's data files (`agents/<agent_id>/…`) and the project-level data — deliberately not described as a project or task directory, so the model does not mistake it for the task's working directory (`CWD`).
On Windows, `{{PROJECT_DIR}}` and `{{CWD}}` are injected with forward slashes — like every other path core composes for the model (attachment lines, the goal-file line, truncated-output recovery paths). The model re-emits these spellings into JSON tool arguments and shell commands; forward slashes are accepted by Node's fs APIs and the package's (Git) Bash tool shell, and avoid JSON backslash-escaping mistakes.
`agent_state/AGENTS.md` is the developer-editable instruction file, injected via `{{AGENTS_MD}}` and empty by default — it is also the file an optimizer edits most (see [Self-Improvement](/self-improvement)).
## Vault
@@ -161,6 +161,8 @@ compaction:
`{{PROJECT_DIR}}` 在提示词中以 **App Data Dir** 名义暴露给模型:PenguinHarness 的应用数据根目录,存放全部 Agent 的数据文件(`agents/<agent_id>/…`)与 Project 级数据——特意不以 Project/任务目录的口径描述,避免模型将其误认为本次任务的工作目录(`CWD`)。
Windows 上注入的 `{{PROJECT_DIR}}` 与 `{{CWD}}` 统一使用正斜杠——与 core 产出的其他模型可见路径(附件行、Goal file 行、截断输出 recovery 路径)同一拼写。模型会把这些拼写原样带入 JSON 工具参数和 Shell 命令;正斜杠被 Node 的 fs API 与包内 (Git) Bash 工具 Shell 接受,也避免 JSON 反斜杠转义出错。
`agent_state/AGENTS.md` 是开发者可编辑的指令文件,经 `{{AGENTS_MD}}` 注入系统提示词,缺省为空——它也是优化器最常改动的文件(见[自我进化](/self-improvement))。
## Vault
+14 -1
View File
@@ -110,7 +110,7 @@ interface EnvironmentInterface {
}
```
`executeTool` yields `partial_tool_call_output` fragments and ends with exactly one complete `tool_call_output`; `origin`-tagged nested messages (e.g. forwarded by `run_subagent`) pass through unchanged. Rendering is explicitly not this interface's concern — streaming rendering belongs to the CLI / Web front ends.
`executeTool` yields `partial_tool_call_output` fragments and ends with exactly one complete `tool_call_output`; `origin`-tagged nested messages (e.g. forwarded by `run_subagent`) pass through unchanged. The built-in Environment can keep truncated text in the Session scratchpad without exposing storage lifecycle hooks through this public interface. Its model-visible recovery path is a plain absolute path; on Windows it is written with forward slashes, which Node's fs APIs and the package's (Git) Bash tool shell both accept, so the same spelling works as a `read_file` argument and inside shell commands. Rendering is explicitly not this interface's concern — streaming rendering belongs to the CLI / Web front ends.
### ToolExecutionRequest and EnvironmentConfig
@@ -124,6 +124,7 @@ interface ToolExecutionRequest {
interface EnvironmentConfig {
workspaceDir: string;
toolConfig: ToolConfig; // { customTools: ToolDefinitionConfig[]; mcpServers: MCPServerConfig[] }
sessionScratchpadDir?: string; // this Session's scratchpad (scratchpad/<sessionId>); enables truncated-output recovery
services?: EnvironmentServices; // runtime services injected into individual tools
vault?: Record<string, string>; // Vault env vars, injected into exec_command / input_command subprocesses
}
@@ -141,6 +142,18 @@ interface MCPServerConfig {
}
```
`Agent.createSession()` and `resumeSession()` pass the Session scratchpad directory
automatically. A standalone embedder that owns a stable per-Session directory opts in by
supplying it — no archive-specific type is exposed:
```ts
const environment = new Environment({
workspaceDir,
toolConfig,
sessionScratchpadDir, // e.g. <dataRoot>/<project>/agents/<agent>/scratchpad/<sessionId>
});
```
### The inner tool contract: BuiltinTool
Inside the Environment, an individual tool follows a deliberately narrower contract ("loose tool, strict framework"):
+13 -1
View File
@@ -110,7 +110,7 @@ interface EnvironmentInterface {
}
```
`executeTool` 逐条产出 `partial_tool_call_output`,并以恰好一条完整 `tool_call_output` 收尾;带 `origin` 的嵌套消息(如 `run_subagent` 转发的子 Session 消息)原样透传。渲染不是本接口的职责——流式渲染由 CLI / Web 前端完成。
`executeTool` 逐条产出 `partial_tool_call_output`,并以恰好一条完整 `tool_call_output` 收尾;带 `origin` 的嵌套消息(如 `run_subagent` 转发的子 Session 消息)原样透传。内置 Environment 可以把被截断文本保存在 Session scratchpad 中,无需在此公共接口暴露存储生命周期钩子。其模型可见 recovery 路径是普通绝对路径;Windows 上统一写成正斜杠——Node 的 fs API 与包内 (Git) Bash 工具 Shell 都接受这种写法,同一拼写既可直接作 `read_file` 参数、也可用于 Shell 命令。渲染不是本接口的职责——流式渲染由 CLI / Web 前端完成。
### ToolExecutionRequest 与 EnvironmentConfig
@@ -124,6 +124,7 @@ interface ToolExecutionRequest {
interface EnvironmentConfig {
workspaceDir: string;
toolConfig: ToolConfig; // { customTools: ToolDefinitionConfig[]; mcpServers: MCPServerConfig[] }
sessionScratchpadDir?: string; // 本 Session 的 scratchpad(scratchpad/<sessionId>),提供后启用截断输出恢复
services?: EnvironmentServices; // 注入给个别工具的运行时服务
vault?: Record<string, string>; // Vault 环境变量,注入 exec_command / input_command 子进程
}
@@ -141,6 +142,17 @@ interface MCPServerConfig {
}
```
`Agent.createSession()` 与 `resumeSession()` 会自动传入 Session scratchpad 目录。自行管理稳定
per-Session 目录的独立 embedder 只需提供该目录即可启用,不暴露归档专用类型:
```ts
const environment = new Environment({
workspaceDir,
toolConfig,
sessionScratchpadDir, // 例如 <dataRoot>/<project>/agents/<agent>/scratchpad/<sessionId>
});
```
### 内层工具契约:BuiltinTool
Environment 之内,单个工具遵循更窄的契约(「松工具、紧框架」):
@@ -70,6 +70,8 @@ The head of a Trace (illustrative; one OmniMessage envelope per line):
{"timestamp":"…","type":"event_msg","payload":{"type":"token_usage","session":{…},"request":{…}}}
```
When tool output exceeds `maxOutputLength`, Trace records the same bounded head, truncation marker, and absolute Session recovery path seen by Web/CLI and the model; it does not separately duplicate the archived text. The path exposes the host data-root layout but remains valid across Tasks and Session resume because the unredacted recovery file lives in that Session's scratchpad. The existing explicit Session-deletion path removes the scratchpad and recovery file together. Trace replay therefore faithfully restores both what the model saw and a usable pointer for later follow-up.
## Session recovery
The Trace is the single source of truth for recovery — there is no separate session database to keep in sync. `resumeSession` works as follows:
@@ -68,6 +68,8 @@ Trace 是 append-only 的 JSON Lines 文件,每行一个 OmniMessage 信封(
{"timestamp":"…","type":"event_msg","payload":{"type":"token_usage","session":{…},"request":{…}}}
```
工具输出超过 `maxOutputLength` 时,Trace 与 Web/CLI、模型一样只记录有界头部、截断提示和绝对 Session recovery 路径,不重复保存归档正文。该路径会暴露宿主的数据根目录布局;未经脱敏的 recovery 文件位于该 Session 的 scratchpad,因此路径跨 Task 和 Session 恢复保持有效。用户明确删除 Session 时,现有删除路径会连同 scratchpad 和 recovery 文件一起清理。因而 Trace 重放既忠实恢复「模型当时看到了什么」,也为后续追问保留可用指针。
## Session 恢复
Trace 是恢复的唯一事实来源,没有独立的会话数据库需要与之对齐。`resumeSession` 的流程:
+10
View File
@@ -45,6 +45,16 @@ A tool only yields incremental `partial_tool_call_output` deltas; the Environmen
Tools and the Environment never throw into the engine: errors collapse into `tool_call_output` messages the model can read and react to. See the [OmniMessage Protocol](/omni-message) for message structure.
### Recovering oversized output
When tool text in an Agent Session exceeds `maxOutputLength`, the model and Web/CLI still receive the same head window, truncation marker, and terminal marker, and the streaming invariant that user-visible output equals model-visible output does not change. Environment also appends a short archive status/path note outside that visible-output cap and saves a Session-owned recovery file. The file is exact within the per-call archive budget and otherwise contains bounded head/tail windows. This is the complete text **received by Environment**: a producer such as a command or subagent session may already have replaced overflow with an `[..., N chars of earlier output dropped ...]` marker in its own bounded unread buffer, and the downstream archive cannot recover text lost before that point.
The Agent can inspect ordinary multiline archives with the existing `read_file` (`offset` / `limit`). For byte tails or very long lines, it must construct a targeted shell command such as `rg` / `tail`; no dedicated retrieval tool is added. The note carries a plain absolute path, always the last element inside the bracket. On Windows it is written with forward slashes: `exec_command` runs through (Git) Bash and Node's fs APIs accept them, so one spelling works in JSON tool arguments and shell commands alike; POSIX paths pass through unchanged, and Session paths are ordinary absolute paths (never `\\?\`-prefixed), so the separator swap is lossless. As with any path, quote it inside shell commands when it contains spaces. The same spelling rule covers every path core composes for the model — the system prompt's App Data Dir / CWD lines, `[attached image/file: …]` lines and the goal-file line (`modelVisiblePath` in the SDK).
Recovery files live under the Session's `scratchpad/<session-id>/truncated-tool-output/`, are created only after actual truncation, and use private permissions where the platform supports them. One call stores at most 8 MiB (the production byte limit is one byte lower so `read_file` remains below its 8 MiB scan cap); larger output keeps bounded head/tail windows in the file with an explicit middle-gap marker. The limit is per call only: a Session has no aggregate archive byte or file-count quota, and concurrent captures independently retain up to one call's budget. Files remain readable across Tasks, runtime disposal, and Session resume until explicit Session deletion removes the entire scratchpad; no separate archive cleanup lifecycle is added.
Recovery files contain the unredacted tool text received by Environment. Accidentally reading credentials or other sensitive data can therefore increase local at-rest retention from the visible head window to the archive budget. Trace does not duplicate those bytes, but it records the same absolute Session path shown to the model and Web/CLI, exposing the host's data-root layout. Archive-write failure never changes the original tool's `stop_reason`; the visible note and stderr warning carry only a short error code (and stderr's tool name), not the path or raw error message.
## Configuration fields
Each tool is described by one `ToolDefinitionConfig`:
+10
View File
@@ -45,6 +45,16 @@ interface ToolResult {
工具与 Environment 从不向引擎抛异常:错误一律折叠为 `tool_call_output` 消息,交给模型阅读并调整下一步。消息结构见 [OmniMessage 协议](/omni-message)。
### 过长输出恢复
Agent Session 中的工具文本超过 `maxOutputLength` 时,模型与 Web/CLI 仍只收到相同的头部窗口、截断提示与终止标记,「用户所见 = 模型所见」的流式契约也保持不变。Environment 还会在该可见输出上限之外追加一条简短的归档状态/路径 note,并保存归该 Session 所有的 recovery 文件:单次归档预算内保存完整文本,超出预算则保存有界头尾。这里的「完整」特指 **Environment 实际收到的文本**:命令或子 Agent Session 等生产者可能已在自身的有界未读缓冲区中用 `[..., N chars of earlier output dropped ...]` 标记替换溢出内容,下游归档无法恢复在此之前已经丢失的原文。
普通多行归档可用现有 `read_file`(`offset` / `limit`)查看;若要读取字节级尾部或超长单行,Agent 必须自行构造定向的 `rg` / `tail` 等 Shell 命令,不新增专用读取工具。note 中的路径是普通绝对路径,恒为括号内最后一个元素。Windows 上统一写成正斜杠:`exec_command` 经 (Git) Bash 执行、Node 的 fs API 也接受正斜杠,同一拼写在 JSON 工具参数与 Shell 命令中通用;POSIX 路径原样透传,且 Session 路径都是普通绝对路径(不会带 `\\?\` 前缀),分隔符替换无损。含空格的路径在 Shell 命令中照常引用即可。同一拼写规则覆盖 core 产出给模型的全部路径——系统提示词的 App Data Dir / CWD 行、`[attached image/file: …]` 行与 Goal file 行(SDK 中的 `modelVisiblePath`)。
Recovery 文件位于该 Session 的 `scratchpad/<session-id>/truncated-tool-output/`,仅在确实发生截断时创建;平台支持时使用仅当前用户可读写的私有权限。单次调用最多保存 8 MiB(生产字节上限少 1 byte,以保持低于 `read_file` 的 8 MiB 扫描上限);更大的输出在文件中保留有界头尾并写明中间被截。该限制仅针对单次调用:一个 Session 没有归档总字节数或文件数配额,并发捕获也各自最多保留一份单调用预算。文件跨 Task、运行时释放和 Session 恢复保持可读,直到用户明确删除 Session 时由现有路径连同整个 scratchpad 一起移除;不新增单独的归档清理生命周期。
Recovery 文件保存 Environment 收到的未经脱敏的工具文本。误读凭据或其他敏感数据会使本地静态留存量从可见头部扩大到归档预算。Trace 不重复保存这些正文,但会记录模型与 Web/CLI 看到的同一个绝对 Session 路径,因此会暴露宿主的数据根目录布局。归档写入失败不改变原工具的 `stop_reason`;双方可见的 note 与 stderr 警告只携带简短错误码(stderr 另含工具名),不携带路径或原始错误消息。
## 配置字段
每个工具由一条 `ToolDefinitionConfig` 描述: