diff --git a/packages/landing/content/blog/penguin-harness-self-improvement_with_amdgpu_EN.md b/packages/landing/content/blog/penguin-harness-self-improvement_with_amdgpu_EN.md new file mode 100644 index 0000000..d8f8caf --- /dev/null +++ b/packages/landing/content/blog/penguin-harness-self-improvement_with_amdgpu_EN.md @@ -0,0 +1,432 @@ +--- +title: "Implementing Agent Self-Improvement with PenguinHarness on an AMD GPU" +date: 2026-07-22 +category: "news" +excerpt: "A complete PenguinHarness self-improvement loop—from baseline evaluation and Trace analysis to Agent optimization and rollback—using a local Qwen3:8B model alongside the Fireworks API." +description: "Learn how PenguinHarness combines Benchmarks, Traces, editable Agent State, Snapshots, and rollback in a dual-model experiment with local Qwen3:8B and the Fireworks API." +--- + + + +# Implementing Agent Self-Improvement with PenguinHarness on an AMD GPU + +*AMD × PrismShadow — Yuyang Gao and Ning Zhang (AMD), Yaowei Zheng (PrismShadow).* + +[PenguinHarness](https://github.com/Prism-Shadow/penguin-harness) is an open-source Agent Harness that brings model integrations, Agent configuration, workspace tools, Sessions, Traces, Skills, and Benchmarks into one runtime, with both a CLI and a Web UI. It can use hosted models as well as local models exposed through an OpenAI-compatible endpoint. + +What makes the project especially interesting is that it represents Agent behavior as a set of readable, editable, and versioned state files rather than as a fixed prompt that only a developer can maintain manually. Role definitions, operating procedures, reusable Skills, and runtime settings all belong to the Agent State, while each task produces a complete Session and Trace. This allows one Agent to evaluate another, update its State using evidence from real executions, and then verify the change against the same evaluation. + +This is what PenguinHarness calls “self-improvement.” It does not retrain the model or update its weights. Instead, it improves the Agent Harness around the model and uses repeatable measurements to decide whether a new version should be retained. + +## How PenguinHarness Implements Agent Self-Improvement + +The main editable parts of an Agent State include: + +- `AGENTS.md`: role, boundaries, and operating procedures; +- `skills/`: reusable capabilities; +- `system_config.yaml`: version and runtime configuration. + +PenguinHarness organizes the self-improvement process around three roles: + +- **Target Agent**: the Agent that performs tasks and is evaluated and improved; +- **Evaluator**: runs one Benchmark Case in an isolated workspace and scores it against a private Rubric; +- **Optimizer**: reads baseline scores and their linked Traces, forms an improvement hypothesis, and modifies the Target Agent State. + +The complete loop works as follows: + +1. Create a multi-Case Benchmark for the target capability. +2. Run the Target Agent repeatedly to establish a traceable baseline. +3. Use the scores and linked Traces to identify stable failure patterns. +4. Save a Snapshot and modify the Agent State. +5. Evaluate the candidate version with the same Benchmark and the same model. +6. Keep the new version only if its total score is strictly higher; otherwise, restore the previous State. + +The loop is orchestrated by built-in Skills: + +- `agent-creation`: creates the initial Agent from a set of requirements; +- `benchmark-design`: designs a multi-Case Benchmark and establishes a complete baseline; +- `agent-evaluation`: runs and scores one isolated Case execution; +- `agent-optimization`: analyzes the baseline and Traces, edits Agent State, evaluates the candidate, and handles rollback. + +The important point is not simply that “one model rewrites another model’s prompt.” Every change is re-evaluated under the same conditions. The Benchmark provides the measurement, the Trace provides the evidence, the Snapshot provides a recovery point, and the State version ties each score to the Agent State that actually produced it. If an optimization does not yield a strict improvement, the candidate change is not treated as a successful evolution. + +To demonstrate the full mechanism, this article uses a dual-model setup. Qwen3:8B runs through Ollama on an AMD GPU and serves as the Target Agent model. A model accessed through the Fireworks API runs `default_agent` and is responsible for creating the Agent, designing the Benchmark, and performing optimization. We first establish a v1 baseline, let the Optimizer improve the Agent using evidence from real Traces, and then use the same Benchmark to decide whether to accept the new version or roll it back. The focus is the self-improvement loop itself, not a model leaderboard comparison. + +## What We Will Build + +We will create a `meeting-summary-agent`. It reads a small collection of text files and generates the following file in its workspace: + +```markdown +# Summary + +## Confirmed + +## Action Items + +## Unresolved +``` + +The Benchmark contains two Cases: + +| Case | Task | +|---|---| +| Single meeting record | Extract the confirmed decision, two action items, and one unresolved item | +| Draft versus formal decision | Read both a draft and a formal decision, and treat the formal decision as authoritative | + +Each Case runs independently three times, so one complete evaluation contains six Target Agent runs. + +The overall workflow is: + +1. Configure local Qwen3:8B and the Fireworks API model. +2. Create the v1 Agent. +3. Export the v1 Snapshot. +4. Create and run the Benchmark. +5. Inspect the baseline and Traces. +6. Optimize the Agent. +7. Evaluate the candidate with the same Benchmark. +8. Accept the new version or roll it back. + +--- + +## Step 1: Configure Qwen3:8B on an AMD GPU and the Fireworks API + +This experiment uses two models: + +| Purpose | Agent | Model | +|---|---|---| +| Create the Agent, design the Benchmark, and perform optimization | `default_agent` | Fireworks API model | +| Undergo evaluation and improvement | `meeting-summary-agent` | Qwen3:8B on an AMD GPU | + +### Prepare Local Qwen3:8B + +The local model runs through Ollama. Installing the AMD driver, ROCm, and Ollama is outside the scope of this article; you only need to confirm that Ollama can already run Qwen3:8B successfully. For setup instructions, see the [Ollama Linux documentation](https://docs.ollama.com/linux) and [GPU support documentation](https://docs.ollama.com/gpu). + +### Get Fireworks API Access + +Through the AMD AI Developer Program, AMD and Fireworks AI offer eligible developers USD 50 in complimentary Fireworks credits. Fireworks provides open-weight models through an OpenAI-compatible endpoint. See [Getting Fireworks API Access](https://penguin.ooo/blog/fireworks-credits-amd) for instructions on redeeming the credits and generating an API key. + +### Register the Models in the Web UI + +Install and start PenguinHarness: + +```bash +curl -fsSL https://penguin.ooo/install.sh | sh +penguin web +``` + +Open the **Models** page in the Web UI and add the local Qwen3:8B model: + +PenguinHarness local Qwen3:8B model configuration + + +Next, configure the Fireworks API key and set DeepSeek V4 Flash as the default model: + +PenguinHarness Fireworks API model configuration + + +New top-level `default_agent` Chats can now use the Project Default—the Fireworks model. When the Benchmark runs `meeting-summary-agent`, it explicitly selects the following local model pair: + +```text +provider: custom +model_id: qwen3:8b +``` + +The baseline and every candidate must use this same `(provider, model_id)` pair. Otherwise, their scores are not directly comparable. + +--- + +## Step 2: Create the v1 Agent + +In the PenguinHarness Web UI, create a new Agent named **meeting-summary-agent**. + +Start a new top-level Chat with `default_agent`, select the Fireworks model that was just configured as the Project Default, invoke the `agent-creation` Skill, and submit the prompt below. It creates the v1 version of **meeting-summary-agent**. Version 1 defines only the basic responsibilities and safety boundaries; it does not preload a complete summarization workflow. This allows the Benchmark to expose missing operating habits through actual runs. + +
+Expand: Complete Prompt for Creating the v1 Agent + +```text +Use the agent-creation Skill to configure the Agent `meeting-summary-agent`. + +Goal: +This is a simple local-file summarization Agent. It reads the task instructions +and text files in the current workspace, then creates the summary file requested +by the task. + +Keep v1 minimal: + +1. In AGENTS.md, specify only that the Agent should: + - Read the task and relevant files in the current workspace. + - Produce a concise summary based on the file contents. + - Work only inside the current workspace. + - Not access external services. + - Not expose environment variables, credentials, or unrelated files. + +2. Do not preload a complete file-summarization workflow, such as: + - Requiring an inventory and full read of every source file. + - Systematically distinguishing drafts from formal decisions. + - Requiring all information to be divided into Confirmed, Action Items, + and Unresolved sections. + - Requiring a line-by-line validation of the output after creation. + +3. Do not install a domain-specific Skill. +4. Do not deliberately instruct the Agent to make mistakes. +5. Do not modify the stable system_prompt. +6. Set: + - name: Meeting Summary Agent + - description: Summarizes a small set of local text files. + +Finally, report the Agent State files that were changed and the current +State version. +``` + +
+ + +After the operation completes, the new Agent appears in the Agents list: + +Meeting Summary Agent in the Agents list + + +Open the Agent settings and export the v1 Snapshot. This Snapshot provides the recovery point if a later candidate fails: + +Exporting the v1 Agent State Snapshot + + +--- + +## Step 3: Create the Benchmark + +Still in the top-level `default_agent` Chat that uses the Fireworks model, invoke the `benchmark-design` Skill and submit the following prompt to create and calibrate the v1 Benchmark. + +Use this Benchmark ID: + +```text +simple-file-summary-2case-v1 +``` + +The maximum scores of the two Cases total 100 points, and each Case runs three times. The Target Agent can see only the public Statement, never the private Rubric. + +
+Expand: Complete Prompt for Creating and Calibrating the Benchmark + +```text +Use the benchmark-design Skill to create and calibrate a Benchmark for the +following Test Agent. + +Test Agent: +meeting-summary-agent + +Benchmark ID: +simple-file-summary-2case-v1 + +Evaluation model: +provider: custom +model_id: qwen3:8b + +Target capability: +Read a small set of local text files and create SUMMARY.md in the workspace root. +Accurately distinguish confirmed facts, action items, and unresolved information. +When a draft conflicts with an explicit formal decision, treat the formal decision +as authoritative. +Do not guess or modify the input materials. + +Shared requirements: +1. Each Case Statement must provide README.md and materials/*.txt. +2. The Target Agent must create SUMMARY.md. +3. SUMMARY.md must contain: + - # Summary + - ## Confirmed + - ## Action Items + - ## Unresolved +4. Do not modify README.md or materials/. +5. Set runs = 3. +6. The maximum scores of the two Rubrics must total 100 points. +7. The Statement must not reveal the private Rubric or expected answer. + +Create the following two Cases. + +CASE-001: Single meeting record, maximum 45 points + +materials/meeting.txt: + +Product weekly meeting, dated 2026-08-03. + +Formal decision: open the internal trial on 2026-08-10. + +Xiaolin is responsible for preparing the user guide, due 2026-08-06. +Xiaozhou is responsible for completing smoke testing, due 2026-08-08. + +Whether mobile export will be included in this trial remains undecided. + +The Rubric should check: +- Whether SUMMARY.md exists. +- Whether the title and categories are correct. +- Whether the internal trial date is accurate. +- Whether the owner, task, and deadline are complete for both action items. +- Whether mobile export remains classified as unresolved. +- Whether no information was invented. +- Whether the input files remain unchanged. + +CASE-002: Draft versus formal decision, maximum 55 points + +materials/plan_draft.txt: + +Project draft, dated 2026-08-01. +The draft proposes a launch date of 2026-08-15. +The tentative owner is Xiaolin. + +materials/final_decision.txt: + +Formal decision, dated 2026-08-04. +Because testing has been delayed, the launch date is changed to 2026-08-22. +Xiaolin is responsible for release preparation, due 2026-08-20. +The email campaign date remains undecided. + +The Rubric should check: +- Whether SUMMARY.md exists. +- Whether the title and categories are correct. +- Whether the formal launch date, 2026-08-22, is used. +- Whether the draft date is not presented as the current decision. +- Whether the release-preparation task and deadline are extracted accurately. +- Whether the email campaign date remains classified as unresolved. +- Whether no information was invented. +- Whether the input files remain unchanged. + +Complete the full 2 Cases × 3 runs baseline and write the result to +scoreboard.yaml. + +Finally, report: +- The Benchmark path. +- The total score. +- The three raw scores and mean for each Case. +- Run-to-run variation. +- All Test Session IDs. +``` + +
+ +After completion, the Benchmark structure looks like this: + +Generated Benchmark structure + + +Inspect the output in the Web UI and open the Benchmark page to review the total score, each Case mean, and the corresponding Sessions and Traces: + +Baseline Benchmark score + + +The overall baseline score is 84. Because the task is intentionally small, the baseline already exceeds 80, but there is still room for improvement. In the next step, we optimize the Agent to demonstrate the self-improvement process. + + +--- + +## Step 4: Optimize the Agent + +After the baseline has been established, use the top-level `default_agent` Chat running the Fireworks model to invoke the `agent-optimization` Skill. The prompt instructs the Optimizer to analyze all six runs and their linked Traces, form a generalizable behavioral hypothesis, and then update `AGENTS.md` or create a narrowly scoped Skill. + +
+Expand: Complete Prompt for Optimizing the Agent + +```text +Use the agent-optimization Skill in Benchmark optimization mode to improve +the target Agent. + +Test Agent: +meeting-summary-agent + +Benchmark: +simple-file-summary-2case-v1 + +Optimization objective: +Improve the reliability of reading a small set of local files and producing +an accurate summary. + +Rules: +1. Use the complete baseline for the current version in scoreboard.yaml as + the reference. +2. Keep the same model: + - provider: custom + - model_id: qwen3:8b +3. Do not modify the Cases, Statements, Rubrics, or runs. +4. Analyze both Cases, all three runs per Case, and every score-linked Trace. +5. In this round, propose only one falsifiable behavioral hypothesis that + generalizes across Cases. +6. Make only the smallest Agent State change needed to support that hypothesis. +7. Prefer editing AGENTS.md. Create a narrowly scoped Skill only if a genuinely + reusable capability is needed. +8. Do not encode Case IDs, names, specific dates, expected answers, or private + Rubric content. +9. The candidate must run the complete 2 Cases × 3 runs matrix. +10. Accept the candidate only if its total score is strictly higher than the + reference. Roll it back if the score is equal or lower. +11. Accept at most one new version in this run. + +Finally, report: +- The reference total and Case means. +- Stable failure patterns found in the Traces. +- The behavioral hypothesis. +- The Agent State files that were changed. +- The candidate total and Case means. +- The reason for acceptance or rollback. +- All Test Session IDs. +``` + +
+ +A likely optimization direction is to add a concise operating procedure or a small set of workflow constraints. + +The exact change should be driven by evidence in the real Traces, rather than by embedding the Benchmark answers in the Agent State. + +--- + +## Step 5: Compare Results and Retain the New Version + +The Optimizer evaluates the candidate with the same two Cases, the same Qwen3:8B model, and the same three repeated runs. + +The acceptance rule is simple: + +```text +candidate total > reference total +→ keep the new version + +candidate total <= reference total +→ roll back to v1 +``` + +The optimized Agent updates `AGENTS.md`, adds an operating procedure, reruns the Cases, and produces the new score: + +Optimization result and updated Agent State + + +If v2 is accepted, export a v2 Snapshot from the Agent Overview page. To continue improving the Agent, repeat the same process from the accepted version. + +Open the Benchmark page and select **MEETING SUMMARY AGENT** to see both evaluation rounds in the improvement history: + +Two rounds in the Benchmark history + + + + +> Because this example uses a relatively simple task, a single optimization round can bring the score close to the maximum. On more complex real-world tasks, the value of the self-improvement loop becomes more apparent. + +--- + +## Conclusion + +This experiment does not show Qwen3:8B retraining itself during execution. Instead, it shows how PenguinHarness places the Agent’s operating procedure inside a verifiable loop: + +```text +Behavior is measured by a Benchmark +→ failures can be traced back through Traces +→ improved procedures are written into Agent State +→ the new version undergoes the complete evaluation again +→ only measured improvements are retained +``` + +This approach is especially useful for local models. Without preparing training data or fine-tuning model weights, clearer operating procedures can still improve the reliability with which an Agent completes its tasks. + +## References + +- [PenguinHarness GitHub](https://github.com/Prism-Shadow/penguin-harness) +- [PenguinHarness Self-Improvement Documentation](https://penguin.ooo/docs/self-improvement/) +- [Fireworks API Credits and API Key Guide](https://penguin.ooo/blog/fireworks-credits-amd) diff --git a/packages/landing/content/blog/penguin-harness-self-improvement_with_amdgpu_ZH.md b/packages/landing/content/blog/penguin-harness-self-improvement_with_amdgpu_ZH.md new file mode 100644 index 0000000..8d1218f --- /dev/null +++ b/packages/landing/content/blog/penguin-harness-self-improvement_with_amdgpu_ZH.md @@ -0,0 +1,405 @@ +--- +title: "在 AMD GPU 上用 PenguinHarness 实现 Agent 自我进化" +date: 2026-07-22 +category: "news" +excerpt: "通过本地 Qwen3:8B 与 Fireworks API 的双模型分工,完整演示 PenguinHarness 从基线评测、Trace 分析到 Agent 优化与回滚的自我进化闭环。" +description: "介绍 PenguinHarness 如何通过 Benchmark、Trace、可编辑的 Agent State 与 Snapshot 回滚构建自我进化闭环,并用本地 Qwen3:8B 与 Fireworks API 完成一次双模型实验。" +--- +# 在 AMD GPU 上用 PenguinHarness 实现 Agent 自我进化 + +*AMD × PrismShadow——高钰洋、张宁(AMD),郑耀威(PrismShadow)。* + +[PenguinHarness](https://github.com/Prism-Shadow/penguin-harness) 是一个开源的 Agent Harness。它把模型接入、Agent 配置、工作区工具、Session、Trace、Skill 和 Benchmark 放在同一套运行环境中,并同时提供 CLI 与 Web UI。在线模型和通过 OpenAI 兼容接口暴露的本地模型,都可以作为其中的推理后端。 + +这个项目比较特别的一点,是它把 Agent 的行为定义为一组可以读取、修改和版本化的状态文件,而不是一段只能由开发者手工维护的固定提示词。角色说明、工作方式、可复用 Skill 和运行参数都属于 Agent State;一次任务又会留下完整的 Session 与 Trace。因此,另一个 Agent 可以先评测目标 Agent,再根据真实运行记录修改它的状态,最后用同一套评测验证修改是否有效。 + +这就是 PenguinHarness 所说的“自我进化”:它不重新训练模型,也不更新模型权重,而是持续改进模型外部的 Agent Harness,并用可重复的结果决定新版本能否被保留。 + +## PenguinHarness 如何实现自我进化 + +一个 Agent 的主要可编辑状态包括: + +- `AGENTS.md`:角色、边界和工作流程; +- `skills/`:可复用能力; +- `system_config.yaml`:版本和运行配置。 + +围绕这些状态,PenguinHarness 把自我进化组织成三个角色: + +- **Target Agent**:真正执行任务、接受评测和改进的目标 Agent; +- **Evaluator**:在隔离工作区中运行一个 Benchmark Case,并根据私有 Rubric 评分; +- **Optimizer**:读取基线分数及其关联 Trace,提出改进假设并修改 Target Agent State。 + +完整闭环如下: + +1. 为目标能力建立多 Case Benchmark。 +2. 对 Target Agent 重复运行,得到可追溯的基线。 +3. 从分数和对应 Trace 中定位稳定的失败模式。 +4. 保存 Snapshot,并修改 Agent State。 +5. 使用相同 Benchmark 和相同模型重新评测候选版本。 +6. 总分严格提高则保留新版本,否则恢复原状态。 + +这个流程主要由内置 Skill 编排: + +- `agent-creation`:根据需求创建初始 Agent; +- `benchmark-design`:设计多 Case Benchmark,并建立完整基线; +- `agent-evaluation`:隔离执行并评分一次 Case 运行; +- `agent-optimization`:分析基线和 Trace,修改 Agent State,并完成候选版本评测与回滚。 + +这里的关键不只是“让另一个模型改写提示词”,而是让每次修改都经过同条件复测。Benchmark 提供度量,Trace 提供证据,Snapshot 提供恢复点,版本号则把分数与实际 Agent State 对应起来。优化没有带来严格提升时,候选修改不会被当成进化成果。 + +为了把这套机制完整跑一遍,本文采用双模型分工:AMD GPU 上通过 Ollama 运行的 Qwen3:8B 作为 Target Agent 的模型;通过 Fireworks API 调用的模型则用于运行 `default_agent`,负责创建 Agent、设计 Benchmark 和执行优化。我们先测出 v1 的基线,再让 Optimizer 根据真实 Trace 改进它,最后由同一 Benchmark 决定接受新版本还是回滚。这个例子关注的是自我进化闭环本身,而不是比较模型排行榜上的能力高低。 + +## 本文要完成的实验 + +我们创建一个 `meeting-summary-agent`。它读取少量文本文件,并在工作区中生成: + +```markdown +# 摘要 + +## 已确认 + +## 待办 + +## 未确定 +``` + +Benchmark 包含两个 Case: + +| Case | 任务 | +|---|---| +| 单份会议纪要 | 提取已确认事项、两个待办和一个未确定项 | +| 草案与正式决定 | 同时读取草案和正式决定,并以正式决定为准 | + +每个 Case 独立运行 3 次,因此一次完整评测共有 6 次 Target Agent 运行。 + +完整流程如下: + +1. 配置本地 Qwen3:8B 与 Fireworks API 模型。 +2. 创建 v1 Agent。 +3. 导出 v1 Snapshot。 +4. 创建并运行 Benchmark。 +5. 查看基线和 Trace。 +6. 优化 Agent。 +7. 使用相同 Benchmark 重新评测候选版本。 +8. 接受新版本或回滚。 + +--- + +## 第一步:配置 AMD GPU 上的 Qwen3:8B 与 Fireworks API + +本实验使用两类模型: + +| 用途 | Agent | 模型 | +|---|---|---| +| 创建 Agent、设计 Benchmark 和执行优化 | `default_agent` | Fireworks API 模型 | +| 接受评测与进化 | `meeting-summary-agent` | AMD GPU 上的 Qwen3:8B | + +### 准备本地 Qwen3:8B + +本地模型通过 Ollama 运行。AMD 驱动、ROCm 和 Ollama 的安装不是本文重点,这里不做详细介绍;只需提前确认 Qwen3:8B 已经能够被 Ollama 正常调用。具体安装方式可以参考 [Ollama Linux 文档](https://docs.ollama.com/linux) 和 [GPU 支持文档](https://docs.ollama.com/gpu)。 + +### 获取 Fireworks API + +通过 AMD AI Developer Program,AMD 与 Fireworks AI 合作,为符合条件的开发者提供价值 50 美元的免费 Fireworks 额度。Fireworks 通过 OpenAI 兼容端点提供开源权重模型。可参阅 [Fireworks API 获取指南](https://penguin.ooo/blog/fireworks-credits-amd),了解额度兑换和 API Key 获取步骤。 + +### 在 Web UI 中注册模型 + +安装并启动 PenguinHarness: + +```bash +curl -fsSL https://penguin.ooo/install.sh | sh +penguin web +``` + +打开 Web UI 的 **Models** 页面,添加本地 Qwen3:8B: + +PenguinHarness 本地 Qwen3:8B 模型配置 + +然后配置 Fireworks API Key,并将 DeepSeek V4 Flash 设置为默认模型: + +PenguinHarness Fireworks API 模型配置 + +这样,新建 `default_agent` 顶层 Chat 时可以直接使用 Project Default,也就是 Fireworks 模型;Benchmark 在运行 `meeting-summary-agent` 时,则显式指定下面这组本地模型配置: + +```text +provider: custom +model_id: qwen3:8b +``` + +后续的 baseline 和 candidate 都必须沿用这组 `(provider, model_id)`,否则分数不能直接比较。 + +--- + +## 第二步:创建 v1 Agent + +在 PenguinHarness Web UI 中创建一个新的 Agent:`meeting-summary-agent`。 + +使用 `default_agent` 新建一个顶层 Chat,模型选择刚刚设为 Project Default 的 Fireworks 模型。然后调用 `agent-creation` Skill,并输入下面的 Prompt,生成 `meeting-summary-agent` 的 v1 版本。v1 只定义基本职责和安全边界,不提前写入完整总结流程,这样 Benchmark 才能真实暴露它缺少的工作习惯。 + +
+展开:创建 v1 Agent 的完整 Prompt + +```text +请使用 agent-creation Skill 配置 Agent `meeting-summary-agent`。 + +目标: +这是一个简单的本地文件总结 Agent。它读取当前工作区中的任务说明和文本文件,并创建题目要求的总结文件。 + +请将 v1 保持简洁: + +1. 在 AGENTS.md 中只写清: + - 读取当前工作区中的任务和相关文件; + - 根据文件内容生成简洁总结; + - 只在当前工作区工作; + - 不访问外部服务; + - 不泄露环境变量、凭证或无关文件。 + +2. 不要预先加入完整的文件总结工作流,例如: + - 强制列出并读取全部材料; + - 系统区分草案与正式决定; + - 强制将信息分为已确认、待办和未确定; + - 完成后逐项核对输出。 + +3. 不要安装业务 Skill。 +4. 不要故意要求 Agent 犯错。 +5. 不修改稳定的 system_prompt。 +6. 设置: + - name: Meeting Summary Agent + - description: Summarizes a small set of local text files. + +最后报告修改的 Agent State 文件和当前 State version。 +``` + +
+ + +创建完成后,可以在 Agent 列表中看到新增的 `meeting-summary-agent`: + +Agent 列表中的 meeting-summary-agent + +然后进入 Agent 设置页面并导出 v1 Snapshot。它是后续 Candidate 失败时的回滚基础。 + +导出 v1 Agent State Snapshot + +--- + +## 第三步:创建 Benchmark + +仍然在使用 Fireworks 模型的 `default_agent` 顶层 Chat 中调用 `benchmark-design` Skill,并输入下面的 Prompt 来创建并校准 v1 版本的 Benchmark。 + +Benchmark ID 使用: + +```text +simple-file-summary-2case-v1 +``` + +两个 Case 的总分为 100 分,每个 Case 运行 3 次。Target Agent 只能看到公开题面,不能看到私有 Rubric。 + +
+展开:创建并校准 Benchmark 的完整 Prompt + +```text +请使用 benchmark-design Skill,为下面的 Test Agent 创建并校准 Benchmark。 + +Test Agent: +meeting-summary-agent + +Benchmark ID: +simple-file-summary-2case-v1 + +评测模型: +provider: custom +model_id: qwen3:8b + +能力目标: +读取少量本地文本文件,在工作区根目录创建 SUMMARY.md; +准确区分已确认事实、待办事项和未确定信息; +当草案与明确的正式决定冲突时,以正式决定为准; +不得猜测或修改输入材料。 + +统一要求: +1. 每个 Case 的 statement 中提供 README.md 和 materials/*.txt。 +2. Target Agent 必须创建 SUMMARY.md。 +3. SUMMARY.md 包含: + - # 摘要 + - ## 已确认 + - ## 待办 + - ## 未确定 +4. 不得修改 README.md 或 materials/。 +5. runs = 3。 +6. 两个 Rubric 的最高分合计为 100。 +7. Statement 不得泄露私有 Rubric 或标准答案。 + +请创建以下两个 Case。 + +CASE-001:单份会议纪要,最高 45 分 + +materials/meeting.txt: + +产品周会,日期 2026-08-03。 + +正式决定:2026-08-10 开放内部试用。 + +小林负责整理使用说明,截止 2026-08-06。 +小周负责完成冒烟测试,截止 2026-08-08。 + +移动端导出是否包含在本次试用中,尚未决定。 + +Rubric 应检查: +- SUMMARY.md 是否存在; +- 标题和分类是否正确; +- 内部试用日期是否准确; +- 两个待办的负责人、事项和截止日期是否完整; +- 移动端导出是否保留为未确定; +- 是否没有编造信息; +- 输入文件是否保持不变。 + +CASE-002:草案与正式决定,最高 55 分 + +materials/plan_draft.txt: + +项目草案,日期 2026-08-01。 +草案计划在 2026-08-15 上线。 +初步负责人为小林。 + +materials/final_decision.txt: + +正式决定,日期 2026-08-04。 +由于测试延期,上线日期改为 2026-08-22。 +小林负责发布准备,截止 2026-08-20。 +邮件宣传时间尚未确定。 + +Rubric 应检查: +- SUMMARY.md 是否存在; +- 标题和分类是否正确; +- 是否使用正式上线日期 2026-08-22; +- 是否没有把草案日期当成当前决定; +- 是否准确提取发布准备任务和截止日期; +- 邮件宣传时间是否保留为未确定; +- 是否没有编造信息; +- 输入文件是否保持不变。 + +完成完整的 2 Cases × 3 runs 基线,将结果写入 scoreboard.yaml。 + +最后报告: +- Benchmark 路径; +- 总分; +- 两个 Case 的三次原始分数与均分; +- 运行波动; +- 全部 Test Session ID。 +``` + +
+ +完成后,会得到如下图所示的 Benchmark 结构: + +生成后的 Benchmark 结构 + +在 Web UI 中查看运行输出,并在 Benchmark 页面检查总分、每个 Case 的均分,以及对应的 Session 和 Trace。 + +Benchmark 基线评分 + +可以看到,整体评分为 84 分。由于任务较为简单,baseline 已经进入 80 分区间,但仍有优化空间。后续我们会继续优化 Agent,以展示完整的自我进化过程。 + + +--- + +## 第四步:优化 Agent + +基线建立后,在使用 Fireworks 模型的 `default_agent` 顶层 Chat 中调用 `agent-optimization` Skill,并输入下面的 Prompt。Optimizer 会分析全部 6 次运行及其 Trace,提出一个可泛化的行为假设,再修改 `AGENTS.md` 或创建一个职责明确、范围有限的 Skill。 + +
+展开:优化 Agent 的完整 Prompt + +```text +请使用 agent-optimization Skill,以 Benchmark optimization mode 优化目标 Agent。 + +Test Agent: +meeting-summary-agent + +Benchmark: +simple-file-summary-2case-v1 + +优化目标: +提高读取少量本地文件并生成可靠总结的稳定性。 + +规则: +1. 使用 scoreboard.yaml 中当前版本的完整基线作为 reference。 +2. 保持相同模型: + - provider: custom + - model_id: qwen3:8b +3. 不修改 Case、Statement、Rubric 或 runs。 +4. 分析全部 2 个 Case、每个 Case 的 3 次运行和全部 score-linked Trace。 +5. 本轮只提出一个可证伪、可跨 Case 泛化的行为假设。 +6. 只做支持该假设的最小 Agent State 修改。 +7. 优先修改 AGENTS.md;只有确实需要复用能力时,才创建职责明确、范围有限的 Skill。 +8. 不得写入 Case ID、人物名、具体日期、标准答案或私有 Rubric。 +9. Candidate 必须运行完整的 2 Cases × 3 runs。 +10. Candidate 总分严格高于 reference 才接受;相同或下降必须回滚。 +11. 本次最多接受一个新版本。 + +最后报告: +- reference 总分和各 Case 均分; +- 从 Trace 中发现的稳定失败模式; +- 行为假设; +- 修改的 Agent State 文件; +- candidate 总分和各 Case 均分; +- 接受或回滚原因; +- 全部 Test Session ID。 +``` + +
+ +一种可能的优化方向,是加入简洁的工作流程或执行约束。 + +具体修改内容应由真实 Trace 决定,而不能把 Benchmark 的答案直接写入 Agent State。 + +--- + +## 第五步:比较并保留新版本 + +Optimizer 会使用同样的两个 Case、同样的 Qwen3:8B 和同样的 3 次重复运行评测 Candidate。 + +接受规则很简单: + +```text +candidate 总分 > reference 总分 +→ 保留新版本 + +candidate 总分 <= reference 总分 +→ 回滚到 v1 +``` + +优化完成后可以看到,Agent 修改了 `AGENTS.md`,新增了工作流程,并重新运行全部 Case,得到新的评测分数: + +Agent 优化结果与更新后的 Agent State + +如果 v2 被接受,再从 Agent Overview 导出 v2 Snapshot。如果还需要继续优化,可以基于已接受的版本重复上述流程。 + +打开 Benchmark 页面,选择 **MEETING SUMMARY AGENT**,可以查看两轮评测形成的进化记录: + +Benchmark 页面中的两轮进化记录 + +> 由于示例任务较为简单,一轮优化后分数便接近满分。在更复杂的真实场景中,自我进化能力的价值会体现得更加明显。 + +--- + +## 结语 + +这个实验展示的并不是 Qwen3:8B 在运行中重新训练了自己,而是 PenguinHarness 让 Agent 的工作方式进入了一个可验证闭环: + +```text +行为被 Benchmark 测量 +→ 失败可以回溯到 Trace +→ 工作流程被写入 Agent State +→ 新版本重新接受完整评测 +→ 只有真实提升才被保留 +``` + +对于本地模型来说,这种方式尤其有价值:不需要准备训练数据或微调权重,也能通过更清晰的工作流程,提高 Agent 完成任务的稳定性。 + +## 参考资料 + +- [PenguinHarness GitHub](https://github.com/Prism-Shadow/penguin-harness) +- [PenguinHarness 自我进化文档](https://penguin.ooo/docs/self-improvement/) +- [Fireworks API 免费额度兑换与 API Key 获取](https://penguin.ooo/blog/fireworks-credits-amd)