docs(blog): align intro with self-evolving example and soften exact scores (#42)

This commit is contained in:
Zhang Jason
2026-07-23 01:01:09 +08:00
committed by GitHub
parent c380113e26
commit 6d56ce9ccc
2 changed files with 8 additions and 8 deletions
@@ -3,12 +3,12 @@ title: "A closer look at PenguinHarness — and running a self-improving agent l
date: 2026-07-20
category: practice
author: Ning Zhang (AMD), Yuyang Gao (AMD), Yaowei Zheng (PrismShadow)
excerpt: What PenguinHarness actually is, the ideas behind its architecture, and a real end-to-end run — a local open-weight model on an AMD GPU that fails a scored task, then improves itself from 0 to 4 out of 5 by editing its own files, entirely on-device.
excerpt: What PenguinHarness actually is, the ideas behind its architecture, and a real end-to-end run — a local open-weight model on an AMD GPU that fails a scored task, then recursively improves itself from a middling score to near the ceiling by diagnosing and rewriting its own files, entirely on-device.
---
*AMD × PrismShadow — by Ning Zhang, Yuyang Gao (AMD) and Yaowei Zheng (PrismShadow).*
If you are new to PenguinHarness, this post is a guided tour: what the project is, the ideas that shape its architecture, and — to make it concrete — a real run where a fully local open-weight model on an AMD GPU *fails* a scored task and then *improves itself* from 0 to 4 out of 5 by editing its own files, without a single byte leaving the machine.
If you are new to PenguinHarness, this post is a guided tour: what the project is, the ideas that shape its architecture, and — to make it concrete — a real run where a fully local open-weight model on an AMD GPU fails a scored task and then recursively improves itself from a middling score to near the ceiling by diagnosing and rewriting its own files, without a single byte leaving the machine.
## What PenguinHarness is
@@ -169,7 +169,7 @@ penguin config model add \
**The task, and how it's scored.** We gave it a task that looks trivial: read a project notes file and produce a summary with a 2-sentence overview and exactly 3 key facts — *and follow the team's standard report format.* That last clause is the whole point. The "team format" is an arbitrary house convention — a specific marker line, a `# Report: <subject>` title, a `Classification: INTERNAL` line, a `Reviewed-by: Aurora Team` footer — that lives *only in the agent's `AGENTS.md`* and cannot be inferred from the task. The task carries a private rubric — a checklist the agent never sees — with 10 points: 5 for content any capable model earns from the task alone, and 5 for the convention, which is knowable only from `AGENTS.md`.
**The baseline — and why it isn't perfect.** Running locally on the AMD GPU with a blank `AGENTS.md`, the model wrote a perfectly reasonable summary — and stably lost all 5 convention points, landing around **4.6 / 10**. It could not have done otherwise: nothing in the task reveals the house convention, so this is an *information gap, not a capability gap* — which is exactly why a stronger model can't just "figure it out." And because every step is in the Trace, this is not guesswork — you can open the run and see precisely which points were missed.
**The baseline — and why it isn't perfect.** Running locally on the AMD GPU with a blank `AGENTS.md`, the model wrote a perfectly reasonable summary — and stably lost all 5 convention points, landing around **4.6 / 10** (exact numbers vary run to run). It could not have done otherwise: nothing in the task reveals the house convention, so this is an *information gap, not a capability gap* — which is exactly why a stronger model can't just "figure it out." And because every step is in the Trace, this is not guesswork — you can open the run and see precisely which points were missed.
That is the honest starting point: on local hardware, an out-of-the-box agent does *not* ace a task whose rules live in files it hasn't learned yet. Which is exactly what makes the next section interesting — a measured, auditable, *reproducible* failure is something the agent can systematically fix by teaching itself.
@@ -189,7 +189,7 @@ Let's improve the exact agent from the previous section — the same `qwen3.6:35
- **The edit (N → N+1), authored by the agent.** The agent writes the convention it just inferred into its own `AGENTS.md`. From a single example it correctly recovers the *structure*, but can't yet tell which tokens are fixed constants vs per-report fields — a single example is ambiguous — so it generalizes the marker to a placeholder. Re-evaluated, the score climbs to about **6.6 / 10**. No retraining, no code — the agent edited a text file it reads on every run.
- **The recursion (N+1 → N+2).** Now we show it *several* accepted reports from different projects that share the same marker and sign-off. The agent reasons that whatever is identical across all of them must be a fixed constant, reads its *own* N+1 `AGENTS.md`, and refines it — locking `<!-- ACME-DATA-PLATFORM -->` and `Reviewed-by: Aurora Team` to literals. Re-evaluated again: about **9.8 / 10**. That is recursion in the true sense: `state_{n+1} = agent.reflect(state_n, new_evidence)`.
That is a **4.6 → 6.6 → 9.8** climb on the same model and task, achieved purely by the agent editing a text file it reads — with the harness keeping each round only because the score strictly improved. That is the loop, self-driven: *what the agent can see, the agent can improve* — measured against a rubric, linked to a Trace, not vibes.
That is a **4.6 → 6.6 → 9.8** climb on the same model and task (numbers vary run to run), achieved purely by the agent editing a text file it reads — with the harness keeping each round only because the score strictly improved. That is the loop, self-driven: *what the agent can see, the agent can improve* — measured against a rubric, linked to a Trace, not vibes.
**Run it yourself.** This whole loop is a self-contained, SDK-driven script in the repo:
[`examples/self-improving-agent/`](https://github.com/Prism-Shadow/penguin-harness/tree/main/examples/self-improving-agent).
@@ -3,12 +3,12 @@ title: 深入了解 PenguinHarness——并在 AMD GPU 上本地跑通一个自
date: 2026-07-20
category: practice
author: 张宁(AMD)、高钰洋(AMD)、郑耀威(PrismShadow)
excerpt: PenguinHarness 到底是什么、它的架构背后有哪些理念,以及一次真实的端到端运行——一个完全本地、运行在 AMD GPU 上的开源权重模型,先在一个打分任务上失败,再通过编辑自己的文件把分数从 0 提升到 4(满分 5),全程数据不出本机。
excerpt: PenguinHarness 到底是什么、它的架构背后有哪些理念,以及一次真实的端到端运行——一个完全本地、运行在 AMD GPU 上的开源权重模型,先在一个打分任务上失败,再通过自己诊断并改写自身文件,把分数从大约半数递归自我提升到接近满分,全程数据不出本机。
---
*AMD × PrismShadow——张宁、高钰洋(AMD),郑耀威(PrismShadow)。*
如果你是第一次接触 PenguinHarness,这篇文章是一次完整的导览:这个项目是什么、它的架构背后有哪些设计理念,以及——为了让一切足够具体——一次真实的运行:一个完全本地、跑在 AMD GPU 上的开源权重模型,先在一个打分任务上*失败*,再通过编辑自己的文件把分数从 0 *自我提升*到 4(满分 5),全程没有一个字节离开本机。
如果你是第一次接触 PenguinHarness,这篇文章是一次完整的导览:这个项目是什么、它的架构背后有哪些设计理念,以及——为了让一切足够具体——一次真实的运行:一个完全本地、跑在 AMD GPU 上的开源权重模型,先在一个打分任务上失败,再由它自己诊断失分、改写自身文件,把分数从大约半数递归自我提升到接近满分,全程没有一个字节离开本机。
## PenguinHarness 是什么
@@ -167,7 +167,7 @@ penguin config model add \
**任务,以及如何评分。** 我们给它一个看似平平无奇的任务:读一个项目 notes 文件,写一份含 2 句概述和恰好 3 条要点的摘要——*并遵守团队的标准报告格式。* 玄机全在最后这句:这个“团队格式”是一套任意的内部约定——一行特定 marker、一个 `# Report: <subject>` 标题、一行 `Classification: INTERNAL`、一行 `Reviewed-by: Aurora Team` 页脚——它*只存在于 Agent 的 `AGENTS.md` 里*,无法从任务本身推断。任务附带一份私有 rubric——一份 Agent 永远看不到的清单——共 10 分:5 分是任何有能力的模型仅凭任务就能拿到的内容分,另外 5 分是只能从 `AGENTS.md` 得知的约定分。
**基线——以及它为何不是满分。** 跑在本地 AMD GPU 上、`AGENTS.md` 为空时,这个模型写出了一份相当合理的摘要——却稳定丢掉全部 5 个约定分,落在 **4.6 / 10** 左右。它别无选择:任务里没有任何东西透露那套内部约定,所以这是*信息缺口,不是能力缺口*——也正是为什么更强的模型也没法“自己想出来”。而因为每一步都在 Trace 里,这不是靠猜——你可以打开这次运行,精确看到哪些分被丢了。
**基线——以及它为何不是满分。** 跑在本地 AMD GPU 上、`AGENTS.md` 为空时,这个模型写出了一份相当合理的摘要——却稳定丢掉全部 5 个约定分,落在 **4.6 / 10** 左右(具体数字每次运行会有波动)。它别无选择:任务里没有任何东西透露那套内部约定,所以这是*信息缺口,不是能力缺口*——也正是为什么更强的模型也没法“自己想出来”。而因为每一步都在 Trace 里,这不是靠猜——你可以打开这次运行,精确看到哪些分被丢了。
这就是诚实的起点:在本地硬件上,一个开箱即用的 Agent,面对规则藏在它尚未学过的文件里的任务,并不能一次拿满分。而这恰恰让下一节变得有意思——一个可度量、可审计、*可复现*的失败,Agent 可以通过自学来系统性修复。
@@ -187,7 +187,7 @@ penguin config model add \
- **编辑(N → N+1),由 Agent 撰写。** Agent 把刚推断出的约定写进它自己的 `AGENTS.md`。从单一范例里它能正确还原出*结构*,但还分不清哪些 token 是固定常量、哪些是每份报告要替换的字段(单个范例本就有歧义),于是把 marker 泛化成了占位符。重新评估,分数升到约 **6.6 / 10**。没有重新训练、没有改代码——Agent 编辑的是一份它每次运行都会读的文本文件。
- **递归(N+1 → N+2)。** 现在我们再给它*多份*来自不同项目、却共享同一 marker 和签名的通过报告。Agent 推断出:凡是在所有样本里都完全相同的,必是固定常量;于是它读取自己*上一轮*写的 N+1 `AGENTS.md` 并精炼它——把 `<!-- ACME-DATA-PLATFORM -->` 和 `Reviewed-by: Aurora Team` 锁定为字面量。再评估一次:约 **9.8 / 10**。这才是真正意义上的递归:`state_{n+1} = agent.reflect(state_n, 新证据)`。
在同一个模型、同一个任务上,仅仅由 Agent 自己编辑一份它会读取的文本文件,分数就走出了 **4.6 → 6.6 → 9.8** 的轨迹——而且 harness 每一轮都只因分数严格提升才保留。这就是自驱动的循环:*Agent 能看到的,Agent 就能改进*——用一份 rubric 度量、关联着 Trace,不是拍脑袋。
在同一个模型、同一个任务上,仅仅由 Agent 自己编辑一份它会读取的文本文件,分数就走出了 **4.6 → 6.6 → 9.8** 的轨迹(数字每次运行会有波动)——而且 harness 每一轮都只因分数严格提升才保留。这就是自驱动的循环:*Agent 能看到的,Agent 就能改进*——用一份 rubric 度量、关联着 Trace,不是拍脑袋。
**你可以自己跑一遍。** 整个循环在仓库里有一个自包含、纯 SDK 驱动的脚本:
[`examples/self-improving-agent/`](https://github.com/Prism-Shadow/penguin-harness/tree/main/examples/self-improving-agent)。