docs(blog): local agents on AMD GPUs, with runnable examples (#14)

This commit is contained in:
Zhang Jason
2026-07-21 20:20:51 +08:00
committed by GitHub
parent 0e134e9c6c
commit d5c2bc001d
10 changed files with 1095 additions and 0 deletions
+77
View File
@@ -0,0 +1,77 @@
<!-- English | [简体中文](README.zh.md) -->
# Example: an Agent that improves itself (local, on an AMD GPU via Ollama)
This example is the **"Recursive Self-Improvement"** pillar in runnable code. Using only the
PenguinHarness SDK, it runs one turn of the self-improvement loop:
1. **Evaluate** a constrained writing task and score it against a rubric.
2. **Diagnose** which rubric points were lost (from the run itself).
3. **Edit** the agent's own `AGENTS.md` to address the failure (version N+1).
4. **Re-evaluate** the same task and keep the change only if the score improved.
Everything runs on a **local open-weight model** — `qwen3:8b` served by Ollama — so no cloud API
and no data leaving the machine. Ollama's ROCm backend runs this natively on AMD GPUs.
## Why a deterministic rubric, and why averaging
- The **rubric is plain code you can read** (`score()` in `self-improve.ts`): file actually
written · overview ≤ 2 sentences · exactly 3 bullets · under 60 words · key facts present. No
hidden judge — the before/after numbers are objective and reproducible. In the full product the
Evaluator is driven by the `agent-evaluation` skill against a *private* rubric; this example
distills that idea to its runnable core.
- A local model is **nondeterministic**, so the example runs each version several times and
averages — which is exactly why real benchmarks use a `runs` count per case. A single run can
swing; the mean is what tells you whether the edit actually helped.
## 1–2. Serve the model and point PenguinHarness at it
```bash
export HIP_VISIBLE_DEVICES=0 # optional: pin a specific AMD GPU
ollama serve &
ollama pull qwen3:8b
penguin config model add \
--model-id qwen3:8b \
--provider custom --client-type openai \
--base-url http://localhost:11434/v1 \
--api-key ollama --set-default
```
## 3. Run the example
```bash
pnpm install
pnpm build
pnpm --dir examples/self-improving-agent start
# or directly: npx tsx examples/self-improving-agent/self-improve.ts
```
## What you should see
```text
BASELINE (blank AGENTS.md): 3 runs
run 1: 0/5
run 2: 0/5
run 3: 0/5
BASELINE mean: 0.00/5
N+1 (with working discipline): 3 runs
run 1: 5/5
run 2: 5/5
run 3: 5/5
N+1 mean: 5.00/5
=== Self-improvement result ===
baseline: 0.00/5 → N+1: 5.00/5
Mean score improved — keep version N+1. ✔
```
With a blank `AGENTS.md`, `qwen3:8b` tends to *narrate* the summary in chat and never call the
tool to write the file — so the rubric scores it 0. Adding a short "working discipline" section
(read first, restate the constraints, actually write the file, self-check) flips that. Exact
numbers vary run to run; the averaged direction is the point.
## Notes
- Uses a dedicated agent id (`self-improve-demo`), created on the fly — your own agents are
untouched.
- Re-running updates that demo agent in place.
@@ -0,0 +1,72 @@
<!-- [English](README.md) | 简体中文 -->
# 示例:一个自我改进的 Agent(本地,经 Ollama 跑在 AMD GPU 上)
这个示例是 **“递归自我进化”** 支柱的可运行版本。仅用 PenguinHarness SDK,它会跑完自我进化循环的一轮:
1. **评估(Evaluate)** —— 让 Agent 完成一个约束型写作任务,并按 rubric 打分。
2. **诊断(Diagnose)** —— 从运行结果中看出哪些 rubric 项失了分。
3. **编辑(Edit)** —— 改写 Agent 自己的 `AGENTS.md` 来修复失分点(版本 N+1)。
4. **重新评估(Re-evaluate)** —— 重跑同一任务,只有分数提升时才保留这次改动。
全程跑在一个**本地开源权重模型**上——由 Ollama 提供的 `qwen3:8b`——所以不用任何云端 API、数据也不
离开本机。Ollama 的 ROCm 后端能原生在 AMD GPU 上运行它。
## 为什么用确定性 rubric,为什么要取平均
- **rubric 就是你能读懂的普通代码**(`self-improve.ts` 里的 `score()`):文件是否确实写出 · 概述
≤ 2 句 · 恰好 3 条要点 · 不超过 60 词 · 关键事实是否出现。没有隐藏的裁判——前后分数是客观、可复现的。
在完整产品里,Evaluator 由 `agent-evaluation` skill 按一份*私有* rubric 驱动;这个示例把那个理念
蒸馏成可运行的核心。
- 本地模型是**非确定性**的,所以示例对每个版本各跑多次取平均——这正是真实 benchmark 里为每个 Case
设置 `runs`(多次运行)的原因。单次可能上下波动;均值才能告诉你这次编辑到底有没有帮助。
## 1–2. 提供模型并把 PenguinHarness 指向它
```bash
export HIP_VISIBLE_DEVICES=0 # 可选:指定某张 AMD GPU
ollama serve &
ollama pull qwen3:8b
penguin config model add \
--model-id qwen3:8b \
--provider custom --client-type openai \
--base-url http://localhost:11434/v1 \
--api-key ollama --set-default
```
## 3. 运行示例
```bash
pnpm install
pnpm build
pnpm --dir examples/self-improving-agent start
# 或直接运行: npx tsx examples/self-improving-agent/self-improve.ts
```
## 你应当看到什么
```text
BASELINE (blank AGENTS.md): 3 runs
run 1: 0/5
run 2: 0/5
run 3: 0/5
BASELINE mean: 0.00/5
N+1 (with working discipline): 3 runs
run 1: 5/5
run 2: 5/5
run 3: 5/5
N+1 mean: 5.00/5
=== Self-improvement result ===
baseline: 0.00/5 → N+1: 5.00/5
Mean score improved — keep version N+1. ✔
```
当 `AGENTS.md` 为空时,`qwen3:8b` 往往会在对话里*叙述*摘要,却从不调用工具把文件写出来——于是
rubric 判它 0 分。加上一小段“任务纪律”(先读原文、把约束逐条列出、真正把文件写出来、结束前自检)就能
扭转这一点。具体数字每次运行会有波动;重点是取平均后的方向。
## 说明
- 使用一个专用 agent id(`self-improve-demo`),运行时即时创建——你自己的 Agent 不会被动到。
- 再次运行会就地更新这个 demo agent。
@@ -0,0 +1,14 @@
{
"name": "@prismshadow/example-self-improving-agent",
"private": true,
"type": "module",
"scripts": {
"start": "tsx self-improve.ts"
},
"dependencies": {
"@prismshadow/penguin-core": "workspace:*"
},
"devDependencies": {
"tsx": "^4.20.0"
}
}
@@ -0,0 +1,192 @@
/**
* Example: an Agent that improves itself — one turn of the self-improvement loop, in code.
*
* This is the "Recursive Self-Improvement" pillar made runnable. It runs entirely on a local
* open-weight model (Ollama serving qwen3:8b) — see README.md for the one-time setup.
*
* The loop, exactly as the docs describe it:
* 1. EVALUATE — run the agent on a constrained task, score it against a rubric.
* 2. DIAGNOSE — read the result to see which rubric points were lost.
* 3. EDIT — rewrite the agent's own AGENTS.md to address the failure (version N+1).
* 4. RE-EVALUATE — run the same task again; keep the change only if the score improved.
*
* The rubric here is a *deterministic, transparent* scorer (plain code you can read below), so the
* before/after numbers are objective and reproducible — no hidden judge. In the full product the
* Evaluator is driven by the `agent-evaluation` skill against a private rubric; this example
* distills that idea to its runnable core.
*
* It uses a dedicated agent id (`self-improve-demo`) created on the fly, so your existing agents
* are never touched.
*
* Run: pnpm --dir examples/self-improving-agent start
* or: npx tsx examples/self-improving-agent/self-improve.ts
*/
import { promises as fs } from "node:fs";
import os from "node:os";
import path from "node:path";
import {
createAgent,
userText,
isCompleteModelMessage,
type OmniMessage,
} from "@prismshadow/penguin-core";
const AGENT_ID = "self-improve-demo";
const PROJECT_ID = "default_project";
/** The constrained task the agent must complete. */
const TASK = `Read notes.txt in your workspace and write summary.md with:
(1) an overview of at most 2 sentences, and
(2) a bullet list of exactly 3 key facts.
The entire summary.md must be under 60 words total.`;
const NOTES = `Project Aurora — Internal Notes
Aurora is a real-time analytics platform launched in Q1 2026. At peak it ingests roughly
2 million events per second through a Kafka-based pipeline. The core query engine was rewritten
in Rust after the original Go version could not keep p99 latency under control; the rewrite cut
p99 from 800ms to 120ms.
The production deployment is single-region with no failover; a multi-region rollout is scheduled
for Q3 2026. Storage is the largest cost line at about $48,000/month, driven by a 90-day
hot-retention policy; cutting retention to 30 days would reduce storage cost by roughly 55%.
`;
/** The "working discipline" we install into AGENTS.md as the N→N+1 edit. */
const DISCIPLINE = `# Working discipline for constrained tasks
When a task lists explicit rules or an exact output format, follow this discipline:
1. Read the source first with a tool before producing any output; never guess file contents.
2. Restate every rule from the task as a short checklist before you start, so none is missed.
3. Actually call the tool to write the required file — do not emit the answer as chat text.
4. Before finishing, re-read your output file and verify it against every rule (file written,
sentence count, bullet count, word count), fixing any mismatch first.
`;
/** rootDir mirrors the SDK's resolveRoot(): PENGUIN_HOME or ~/.penguin/data. */
function rootDir(): string {
return process.env.PENGUIN_HOME ?? path.join(os.homedir(), ".penguin", "data");
}
function agentStateDir(): string {
return path.join(rootDir(), PROJECT_ID, "agents", AGENT_ID, "agent_state");
}
/** Deterministic rubric — 5 independent points, all checkable in code. Returns {score, detail}. */
function score(
summaryPath: string,
summaryText: string | null,
): { score: number; detail: string[] } {
const detail: string[] = [];
let s = 0;
const exists = summaryText !== null;
detail.push(`${exists ? "1" : "0"}/1 file summary.md was actually written`);
if (!exists) return { score: 0, detail };
s += 1;
const body = summaryText.replace(/^#.*$/gm, "").trim(); // drop a markdown title line if present
const overview = body.split(/\n\s*[-*]/)[0] ?? ""; // text before the first bullet
const sentences = (overview.match(/[.!?](\s|$)/g) ?? []).length;
const okSentences = sentences >= 1 && sentences <= 2;
detail.push(`${okSentences ? "1" : "0"}/1 overview is ≤ 2 sentences (found ${sentences})`);
if (okSentences) s += 1;
const bullets = (summaryText.match(/^\s*[-*]\s+/gm) ?? []).length;
const okBullets = bullets === 3;
detail.push(`${okBullets ? "1" : "0"}/1 exactly 3 bullet facts (found ${bullets})`);
if (okBullets) s += 1;
const words = body.split(/\s+/).filter(Boolean).length;
const okWords = words < 60;
detail.push(`${okWords ? "1" : "0"}/1 under 60 words (found ${words})`);
if (okWords) s += 1;
// "facts accurate": require at least two of the source's hard numbers to appear.
const anchors = ["120ms", "55%", "2 million", "$48,000", "Q3 2026", "800ms"];
const hits = anchors.filter((a) => summaryText.includes(a)).length;
const okFacts = hits >= 2;
detail.push(`${okFacts ? "1" : "0"}/1 key facts accurate (${hits} source figures present)`);
if (okFacts) s += 1;
void summaryPath;
return { score: s, detail };
}
async function drain(run: AsyncGenerator<OmniMessage>): Promise<void> {
for await (const msg of run) {
if (isCompleteModelMessage(msg) && msg.payload.type === "text") {
process.stdout.write(msg.payload.text + "\n");
}
}
}
/** Run the task once in a fresh temp workspace seeded with notes.txt; return the rubric score. */
async function runOnce(): Promise<{ score: number; detail: string[] }> {
const agent = await createAgent({ agentId: AGENT_ID });
const ws = await fs.mkdtemp(path.join(os.tmpdir(), "self-improve-"));
await fs.writeFile(path.join(ws, "notes.txt"), NOTES, "utf8");
const session = await agent.createSession({ workspaceDir: ws });
await drain(session.run([userText(TASK)], { approve: async () => "allow" }));
const summaryPath = path.join(ws, "summary.md");
let text: string | null = null;
try {
text = await fs.readFile(summaryPath, "utf8");
} catch {
text = null;
}
return score(summaryPath, text);
}
/**
* Evaluate a version by averaging over RUNS independent runs. Small local models are
* nondeterministic, so a single run is noisy — averaging is exactly why real benchmarks use a
* `runs` count per case. Returns the mean score.
*/
const RUNS = 3;
async function evaluate(label: string): Promise<number> {
console.log(`\n--- ${label}: ${RUNS} runs ---`);
const scores: number[] = [];
for (let i = 0; i < RUNS; i++) {
const r = await runOnce();
scores.push(r.score);
console.log(` run ${i + 1}: ${r.score}/5`);
}
const mean = scores.reduce((a, b) => a + b, 0) / scores.length;
console.log(` ${label} mean: ${mean.toFixed(2)}/5`);
return mean;
}
async function main(): Promise<void> {
// Ensure the demo agent exists, then start from a blank AGENTS.md (the baseline).
await createAgent({ agentId: AGENT_ID });
const agentsMd = path.join(agentStateDir(), "AGENTS.md");
await fs.writeFile(agentsMd, "", "utf8");
// --- Turn N: evaluate the baseline ---------------------------------------------------
const baseline = await evaluate("BASELINE (blank AGENTS.md)");
// --- Edit: install the working discipline (version N+1) ------------------------------
console.log("\n--- EDIT: writing a working-discipline section into the agent's AGENTS.md ---");
await fs.writeFile(agentsMd, DISCIPLINE, "utf8");
// --- Turn N+1: re-evaluate the same task --------------------------------------------
const improved = await evaluate("N+1 (with working discipline)");
// --- Keep-or-roll-back: the loop's decision rule ------------------------------------
console.log("\n=== Self-improvement result ===");
console.log(` baseline: ${baseline.toFixed(2)}/5 → N+1: ${improved.toFixed(2)}/5`);
if (improved > baseline) {
console.log(" Mean score improved — keep version N+1. ✔");
} else {
console.log(" No improvement — roll back to the baseline AGENTS.md.");
await fs.writeFile(agentsMd, "", "utf8");
}
}
main().catch((err) => {
console.error("Example failed:", err);
process.exitCode = 1;
});