feat: 完善两阶段 Agent 自进化 Pipeline、Skills 与 Benchmark (#129)

This commit is contained in:
Jingzhe Xu
2026-07-30 15:59:58 +08:00
committed by GitHub
parent 3bd0a7208a
commit 081fe2d172
34 changed files with 2093 additions and 763 deletions
+33 -18
View File
@@ -3,27 +3,40 @@ title: Self-Improvement
description: The Skill-orchestrated Benchmark and optimization loop: score, improve, snapshot, roll back.
---
Self-improvement in PenguinHarness is not carried by special-purpose engine code — it is carried by Skills orchestrating the ordinary Agent machinery: evaluations are ordinary Sessions, optimization is ordinary file editing, and orchestration uses the built-in `run_subagent` tool. The direct payoff is that the whole process shares the same observability and recovery machinery as everyday runs.
Self-improvement in PenguinHarness uses Skills to orchestrate the ordinary Agent machinery: evaluations are ordinary Sessions and optimization is ordinary file editing. Evaluation construction and optimization run in two independent top-level Sessions, while individual evaluations are delegated through the built-in `run_subagent` tool. Top-level prompts provide the Agent, Benchmark, capability, score, and round settings for the task; Skills own call relationships, calibration, Freeze, protocol, repair, rollback, and reporting.
## Three roles
## Roles and call relationships
| Role | Responsibility |
| --- | --- |
| Builder | Top-level Agent that directly follows `agent-creation` and then `benchmark-design` |
| Target Agent | The Agent being improved; runs evaluation tasks only inside its own Workspace |
| Evaluator | Runs and scores one Benchmark Case run |
| Optimizer | Drives the whole optimization loop |
| Evaluator | Leaf worker created through `run_subagent`; runs and scores one Benchmark Case run |
| Optimizer | New top-level Agent that directly follows `agent-optimization` |
The roles are defined by Skills, not hardcoded: the Evaluator follows the `agent-evaluation` Skill, the Optimizer follows the `agent-optimization` Skill. This applies the design principle stated in the [Configuration Reference](/configuration) — an Agent's behavior is editable files on disk, which is what makes Agents improvable by Agents.
The Builder and Optimizer directly follow their Skills in their own top-level Sessions. Evaluators are created through `run_subagent`; each follows `agent-evaluation` and uses the Penguin CLI to launch the specified Target Agent in an isolated Workspace identified by an absolute path. The Penguin CLI launches the Target Agent for the requested Case run.
## The loop
## Two independent steps
1. `benchmark-design` builds a multi-Case capability Benchmark: repeated independent runs, with a traceable baseline calibrated first;
2. The Optimizer orchestrates Evaluators in parallel via the `run_subagent` tool, covering the Case × runs matrix;
3. Scores plus their linked Traces show where points were lost;
4. The Optimizer edits the Target Agent's editable state — `AGENTS.md`, Skills, config — to produce version N+1;
5. A Snapshot is taken before each round; the candidate version is kept only if the total score strictly improves, otherwise rolled back.
The first top-level Session creates the Agent and its capability evaluation. The Builder first uses `agent-creation`, then uses `benchmark-design` to build a multi-Case Benchmark. It may build the complete initial Case set before Pilot 1 and may refine multiple Cases or difficulty dimensions in a later iteration. The evaluation contract and private standard must be clear and fixed, while the public Statement need not uniquely determine the Gold. A Benchmark may use incomplete public information, conflicting signals, and a fixed private decision standard when that standard expresses a reusable policy, priority, or inference boundary and is not rewritten after seeing the run's answer.
Benchmark optimization mode requires a complete baseline series in the scoreboard — without a calibrated baseline there is no improvement to compare against. Besides this loop, `agent-optimization` also supports a one-shot feedback mode: a concrete correction is applied directly as edits to the Target Agent's state, without going through the evaluation loop.
Before the first dispatch of every new or changed Case, the Builder checks that the Statement is internally coherent, the Rubric agrees with the current Statement and fixed private standard, and every scoring item relies only on defined, provided, or explicitly private premises; this does not require the public materials to reproduce the private standard. It repeats the full review across all Cases before Freeze. Most points should rest on decisions or concise artifacts for which the intended behavior and a plausible shortcut produce different results, rather than giving a high floor for format, evidence enumeration, or analysis completeness.
Before each calibration dispatch, the Builder predicts the result produced by the observed Trace strategy, the different result produced by the desired behavior, and the score range affected. Adding another public rule, exception, source, or check that the model can directly execute does not automatically increase difficulty. If both strategies still reach the same scored result, the Builder chooses another refinement.
The Pilot score is a desired target: meeting it permits an early Freeze; otherwise the Builder completes the configured number of valid Pilot iterations and freezes the lowest-scoring valid revision. The Builder temporarily retains only the current lowest valid revision, then removes that copy and other calibration scaffolding after recording the Formal Baseline. Freeze is followed by a fresh complete Formal matrix. Every valid Formal Baseline is recorded even when its score misses the desired target.
After the user confirms that step is complete, they start the second top-level Session in a new conversation. The Optimizer checks the Benchmark and its first complete Formal Baseline before following `agent-optimization`:
1. orchestrate Evaluators in parallel through `run_subagent`, covering the Case × runs matrix;
2. use scores and linked Traces to propose one bounded Candidate;
3. edit the Target Agent's editable state — `AGENTS.md`, Skills, config — to produce version N+1;
4. keep the Candidate only when its Evaluation score strictly improves; otherwise roll it back;
5. stop early when the desired score is reached, or complete the configured number of valid Candidate rounds and retain the highest-scoring Reference.
Invalid evaluations and correction reruns do not count toward the round limit. On an execution failure, the Optimizer keeps the same Candidate and repairs only the missing cell; it keeps trying while each attempt follows a new diagnosis and applies a distinct safe repair. Both Builder and Optimizer validate that the complete Evaluator response is plain protocol YAML before reading status or score; if formatting is invalid, that same Evaluator resends from its existing result without rerunning the Target Agent.
Every accepted Candidate is appended to and verified in the Scoreboard immediately. A strictly higher Evaluation score decides acceptance; whether the predicted Case behavior changed is reported separately so unrelated single-run variation is not presented as causal evidence. Agent optimization requires a complete Formal Baseline in the Scoreboard — without one there is no improvement to compare against.
## Benchmark storage
@@ -35,18 +48,20 @@ benchmarks/<id>/
├── <case-id>/
│ ├── statement/ # the task given to the Target Agent
│ └── rubric/ # private scoring rubric, isolated from the Target Agent
└── scoreboard.yaml # evaluation records (v2 format)
└── scoreboard.yaml # evaluation records (current format)
```
The separation of `rubric/` from `statement/` is deliberate: the Target Agent sees only the task statement and never touches the scoring rubric.
Each evaluation record in `scoreboard.yaml` (v2 format) is timestamped and carries:
Each evaluation record in `scoreboard.yaml` is timestamped and carries:
- the paired model reference `(provider, model_id)` used for the round;
- the evaluation runtime: a user-specified `(provider, model_id)` pair takes priority, otherwise the pair is inherited from the Builder Session; `thinking_level` is read from the Target Agent config and does not depend on Trace metadata;
- `summary_title` and `summary` (the round's conclusion and the hypothesis for the next one);
- total score, cost, and duration — Case-level metrics are the average over its runs, evaluation-level metrics are the sum over its Cases;
- Score, cost, and duration averages written by the model — Case-level values average Runs and Evaluation-level values average Cases; Run cost preserves its recorded precision, cost averages ignore `null` inputs and remain `null` only when every contributing cost is unknown; Score uses two decimals, cost averages use six decimals, and `duration_ms` is an integer;
- per-Case run details, each run recording `score`, `cost`, `duration_ms`, and `session_id`.
Every Run and every Case has a fixed maximum Score of 100, so Scoreboard entries do not carry `max_score`. The server and Web UI trust the stored aggregate values and do not recompute or cross-check them. Old Scoreboard formats are not migrated or backfilled.
The built-in `default_agent` ships with an example Benchmark (`packages/core/src/state/example-benchmark.ts`) so the evaluation pages have data out of the box; the whole directory can be deleted or replaced at any time.
## Snapshots and versions
@@ -57,7 +72,7 @@ Before each optimization round, the Agent State is packed into `snapshots/v<vers
- Every Evaluator run is an ordinary Session with a full Trace;
- Scoreboard records link back to those Sessions via `session_id`; see [Sessions & Traces](/sessions-and-traces);
- The Web evaluation pages are read-only views of these files; see the [Web App Guide](/web-app).
- The Web evaluation pages are read-only views of these files; the trend chart shows Score only, while the detail table shows model ID and thinking level as separate columns. See the [Web App Guide](/web-app).
Scores are not black-box output: every number can be traced back to the run that produced it.
@@ -68,6 +83,6 @@ Scores are not black-box output: every number can be traced back to the run that
| `agent-creation` | Turn a requirement into a working Agent: write its `AGENTS.md`, install the Skills it needs |
| `benchmark-design` | Design and calibrate a multi-Case capability Benchmark |
| `agent-evaluation` | Run and score one isolated Benchmark Case run |
| `agent-optimization` | Improve an Agent from feedback or Benchmark results |
| `agent-optimization` | Improve an Agent from Benchmark results |
How Skills are organized and installed is covered in the [Skill System](/skills).
+33 -18
View File
@@ -3,27 +3,40 @@ title: 自我进化
description: 由 Skill 编排的 Benchmark 评测与优化闭环:评分、改进、Snapshot 与回滚。
---
PenguinHarness 中的自我进化不依赖专用引擎代码,而是由 Skill 编排普通的 Agent 机制完成:评测是普通的 Session,优化是普通的文件编辑,编排靠内置的 `run_subagent` 工具。这样做的直接收益是——整个过程与日常运行共用同一套可观测性与恢复机制。
PenguinHarness 中的自我进化由 Skill 编排普通的 Agent 机制完成:评测是普通的 Session,优化是普通的文件编辑。创建评测与优化分别运行在两个独立的顶层 Session 中,单次评测通过内置的 `run_subagent` 工具委托。顶层 Prompt 提供 Agent、Benchmark、能力目标、分数和轮数等本次设定;调用关系、校准、Freeze、协议、重试、回滚和报告格式由 Skill 负责。
## 三个角色
## 角色与调用关系
| 角色 | 职责 |
| --- | --- |
| Builder | 顶层 Agent,依次直接执行 `agent-creation` 和 `benchmark-design` |
| Target Agent | 被改进的 Agent,只在自己的 Workspace 里执行评测任务 |
| Evaluator | 执行并评分一次 Benchmark Case 运行 |
| Optimizer | 驱动整个优化循环 |
| Evaluator | `run_subagent` 创建的叶子 Worker,执行并评分一次 Benchmark Case 运行 |
| Optimizer | 新顶层 Agent,直接执行 `agent-optimization` |
角色由 Skill 定义而非硬编码:Evaluator 遵循 `agent-evaluation` Skill,Optimizer 遵循 `agent-optimization` Skill。这正是[配置参考](/configuration)所述设计原则的应用——Agent 的行为是磁盘上的可编辑文件,所以 Agent 可以被 Agent 改进。
Builder 和 Optimizer 在各自的顶层 Session 中直接遵循对应 Skill。Evaluator 通过 `run_subagent` 创建,遵循 `agent-evaluation`,并通过 Penguin CLI 在绝对路径的隔离 Workspace 中启动指定的 Target Agent。Penguin CLI 在每次请求中启动 Target Agent 完成对应的 Case Run。
## 优化循环
## 两个独立步骤
1. `benchmark-design` 构建多 Case 的能力 Benchmark:重复独立运行,先校准出可追溯的基线;
2. Optimizer 通过 `run_subagent` 工具并行编排 Evaluator,覆盖 Case × 运行次数矩阵;
3. 得分与其关联的 Trace 共同指出失分位置;
4. Optimizer 编辑 Target Agent 的可编辑状态——`AGENTS.md`、Skills、配置——产出版本 N+1;
5. 每轮开始前先打 Snapshot;总分严格提升才保留候选版本,否则回滚。
第一个顶层 Session 创建 Agent 和能力评测。Builder 先使用 `agent-creation`,再使用 `benchmark-design` 构建多 Case Benchmark。初版 Cases 可以一次建好并形成完整 Pilot 1;后续每轮可以同时调整多个 Case 或难度维度。评测契约和私有标准必须明确、固定,公开 Statement 则不必唯一决定 Gold。Benchmark 可以通过公开信息不足、冲突信号和固定的私有决策标准形成信息差,只要该标准表达可复用的策略、优先级或推断边界,而且不会根据本次答案改写。
Benchmark 优化模式要求 scoreboard 中已有完整的基线序列——没有校准过的基线,就没有可比较的提升。除此之外 `agent-optimization` 还支持一次性反馈模式:把一条具体的纠正意见直接落实为对 Target Agent 状态的编辑,不经过评测循环。
每个新增或修改后的 Case 在首次派发前都要检查 Statement 自洽、Rubric 与当前 Statement 和固定私有标准一致,并确认评分项只依赖已定义、已提供或明确属于私有标准的前提;这不要求公开材料足以复现私有标准。Freeze 前再对所有 Case 完整检查一次。大部分分数应落在目标行为与合理捷径会产生不同结果的决定或简洁产物上,避免格式、证据罗列和分析完整度形成过高的保底分。
每轮校准都要在派发前预测:当前 Trace 中的策略会产生什么结果、期望行为会产生什么不同结果,以及会影响多少分。增加一条模型可以直接执行的公开规则、例外、来源或检查项并不会自动增加难度;如果两种策略仍会得到相同的计分结果,就应选择其他改法。
Pilot 分数是期望目标:达到后可以提前 Freeze;未达到时完成设定数量的有效 Pilot iteration,并选择其中分数最低的有效版本 Freeze。Builder 在临时目录只保留当前最低有效版本,Formal Baseline 记录后清理该副本和校准脚手架。Freeze 后必须运行全新完整的 Formal matrix;只要 Formal 有效就记录 Baseline,分数没有达到期望也不会使 Benchmark 作废。
用户确认第一步完成后,在新对话中启动第二个顶层 Session。Optimizer 先检查 Benchmark 和第一条完整 Formal Baseline,再使用 `agent-optimization`:
1. 通过 `run_subagent` 并行编排 Evaluator,覆盖 Case × 运行次数矩阵;
2. 根据得分和关联 Trace 提出一个有界 Candidate;
3. 编辑 Target Agent 的可编辑状态——`AGENTS.md`、Skills、配置——产出版本 N+1;
4. Evaluation 分数严格提升才保留 Candidate,否则回滚;
5. 达到期望分数时提前结束,否则完成本次设定数量的有效 Candidate round,并保留最高分 Reference。
无效评测和修复重跑不计入轮数。出现执行失败时,Optimizer 保持同一个 Candidate,只补齐失败单元;只要还能根据新诊断提出不同的安全修复,就继续尝试。Builder 和 Optimizer 都先验证 Evaluator 的完整响应是否为纯协议 YAML,再读取状态或分数;格式不合规时,由同一个 Evaluator 基于已有结果重发,不重新运行 Target Agent。
每个 Accepted Candidate 立即写入并校验 Scoreboard。Evaluation 分数严格提高决定是否接受;假设是否在预期 Case 上得到支持单独报告,避免把单次运行中的无关波动解释为改动因果。Agent 优化要求 Scoreboard 中已有完整 Formal Baseline——没有基线,就没有可比较的提升。
## Benchmark 存储
@@ -35,18 +48,20 @@ benchmarks/<id>/
├── <case-id>/
│ ├── statement/ # 交给 Target Agent 的任务描述
│ └── rubric/ # 私有评分标准,对 Target Agent 隔离
└── scoreboard.yaml # 评测记录(v2 格式)
└── scoreboard.yaml # 当前格式的评测记录
```
`rubric/` 与 `statement/` 的隔离是刻意设计:Target Agent 只能看到题面,永远接触不到评分标准。
`scoreboard.yaml`(v2 格式)中的每条评测记录带时间戳,并记录:
`scoreboard.yaml` 中的每条评测记录带时间戳,并记录:
- 本轮使用的模型成对引用 `(provider, model_id)`;
- 本轮 Runtime:用户显式指定的 `(provider, model_id)` 成对值优先,否则继承 Builder Session;`thinking_level` 从 Target Agent 配置读取,不依赖 Trace 元数据;
- `summary_title` 与 `summary`(本轮结论与下一轮假设);
- 总分、成本与耗时——Case 级指标是各次运行的平均值,评测级指标是各 Case 的加和;
- 由模型写入的 Score、成本与耗时平均值——Case 级对 Runs 求平均,Evaluation 级对 Cases 求平均;单次 Run 成本保留记录中的原始精度,成本平均值忽略 `null`,全部未知时才为 `null`;Score 保留两位小数,成本平均值保留六位小数,`duration_ms` 取整;
- 每个 Case 的逐次运行明细,每次运行含 `score`、`cost`、`duration_ms` 与 `session_id`。
每个 Run 和每个 Case 都固定满分 100,因此 Scoreboard 不再记录 `max_score`。服务端与 Web UI 直接信任已写入的聚合值,不重算、不交叉校验;旧 Scoreboard 不迁移、不回填。
内置的 `default_agent` 预置了一个示例 Benchmark(`packages/core/src/state/example-benchmark.ts`),评测页面开箱即有数据;整个目录可随时删除或替换。
## Snapshot 与版本
@@ -57,7 +72,7 @@ benchmarks/<id>/
- 每次 Evaluator 运行都是一个普通的 Session,留有完整 Trace;
- scoreboard 记录通过 `session_id` 链接回这些 Session,见 [Session 与 Trace](/sessions-and-traces);
- Web 的评测页面是这些文件的只读视图,见 [Web App 指南](/web-app)。
- Web 的评测页面是这些文件的只读视图;折线图只展示 Score,明细表将模型 ID 与推理强度分列显示。见 [Web App 指南](/web-app)。
分数不是黑盒输出:任何一个数字都可以回溯到产生它的那次运行。
@@ -68,6 +83,6 @@ benchmarks/<id>/
| `agent-creation` | 把需求变成可用的 Agent:撰写其 `AGENTS.md`、安装所需 Skill |
| `benchmark-design` | 设计并校准多 Case 的能力 Benchmark |
| `agent-evaluation` | 隔离执行并评分一次 Benchmark Case 运行 |
| `agent-optimization` | 根据反馈或 Benchmark 结果改进 Agent |
| `agent-optimization` | 根据 Benchmark 结果改进 Agent |
Skill 的组织与安装方式见[技能系统](/skills)。
+4 -4
View File
@@ -69,10 +69,10 @@ The built-in Skills, by group (the group manifest is `SKILL_GROUPS` in `packages
| | `vllm` | Deploy and serve LLMs with vLLM behind an OpenAI-compatible endpoint, with tool calling enabled for agent workloads |
| | `ollama` | Deploy and serve local models with Ollama: pull and run them, then expose the OpenAI-compatible endpoint to apps and agents |
| | `llamafactory` | Fine-tune LLMs with LlamaFactory: register datasets, train via YAML configs, merge LoRA adapters and serve the result |
| Agent Tuning | `agent-creation` | Turn a user requirement into a concrete agent: write the target agent's AGENTS.md and install the skills it needs |
| | `benchmark-design` | Design and calibrate a multi-Case capability Benchmark with repeated independent evaluations and a traceable baseline |
| | `agent-evaluation` | Run and score exactly one Benchmark Case run, with CLI execution, Trace provenance checks and private Rubric isolation |
| | `agent-optimization` | Improve an Agent State from direct feedback or versioned multi-Case Benchmark scores and score-linked Traces |
| Agent Tuning | `agent-creation` | Create or configure an Agent State from a user requirement by writing AGENTS.md, setting identity metadata and installing needed Skills |
| | `benchmark-design` | Design and calibrate a multi-Case capability Benchmark for a specified Agent and establish a traceable Formal Baseline |
| | `agent-evaluation` | Internal leaf worker that executes and privately scores exactly one Case run from a complete evaluation protocol |
| | `agent-optimization` | Improve a specified Agent from a complete current baseline on a frozen Benchmark |
## Writing and optimizing Skills
+4 -4
View File
@@ -69,10 +69,10 @@ Skill 库以 npm 包 `@prismshadow/penguin-skills` 发布,tarball 直接携带
| | `vllm` | 用 vLLM 部署与服务 LLM,提供 OpenAI 兼容端点,并为 Agent 负载启用工具调用 |
| | `ollama` | 用 Ollama 部署与运行本地模型,把 OpenAI 兼容端点接入应用与 Agent |
| | `llamafactory` | 用 LlamaFactory 微调 LLM:注册数据集、以 YAML 配置训练、合并 LoRA 适配器并部署产物 |
| Agent 调优 | `agent-creation` | 把用户需求变成具体的 Agent:撰写目标 Agent 的 AGENTS.md 并安装所需 Skill |
| | `benchmark-design` | 设计并校准多 Case 的能力评测 Benchmark,含重复独立评测与可追溯基线 |
| | `agent-evaluation` | 隔离执行并评分单个 Benchmark Case:CLI 执行、Trace 溯源检查、Rubric 私有隔离 |
| | `agent-optimization` | 依据直接反馈或带版本的多 Case Benchmark 分数与关联 Trace 改进 Agent State |
| Agent 调优 | `agent-creation` | 根据需求创建或配置 Agent State,编写 AGENTS.md、设置身份信息并安装所需 Skill |
| | `benchmark-design` | 为指定 Agent 设计并校准多 Case Benchmark,建立可追溯的 Formal Baseline |
| | `agent-evaluation` | 内部叶子执行器:根据完整评测协议隔离执行并私密评分一个 Case Run |
| | `agent-optimization` | 基于冻结 Benchmark 的完整当前基线改进指定 Agent |
## 编写与优化