docs(blog): local agents on AMD GPUs, with runnable examples (#14)
This commit is contained in:
@@ -0,0 +1,77 @@
|
||||
<!-- English | [简体中文](README.zh.md) -->
|
||||
|
||||
# Example: an Agent that improves itself (local, on an AMD GPU via Ollama)
|
||||
|
||||
This example is the **"Recursive Self-Improvement"** pillar in runnable code. Using only the
|
||||
PenguinHarness SDK, it runs one turn of the self-improvement loop:
|
||||
|
||||
1. **Evaluate** a constrained writing task and score it against a rubric.
|
||||
2. **Diagnose** which rubric points were lost (from the run itself).
|
||||
3. **Edit** the agent's own `AGENTS.md` to address the failure (version N+1).
|
||||
4. **Re-evaluate** the same task and keep the change only if the score improved.
|
||||
|
||||
Everything runs on a **local open-weight model** — `qwen3:8b` served by Ollama — so no cloud API
|
||||
and no data leaving the machine. Ollama's ROCm backend runs this natively on AMD GPUs.
|
||||
|
||||
## Why a deterministic rubric, and why averaging
|
||||
|
||||
- The **rubric is plain code you can read** (`score()` in `self-improve.ts`): file actually
|
||||
written · overview ≤ 2 sentences · exactly 3 bullets · under 60 words · key facts present. No
|
||||
hidden judge — the before/after numbers are objective and reproducible. In the full product the
|
||||
Evaluator is driven by the `agent-evaluation` skill against a *private* rubric; this example
|
||||
distills that idea to its runnable core.
|
||||
- A local model is **nondeterministic**, so the example runs each version several times and
|
||||
averages — which is exactly why real benchmarks use a `runs` count per case. A single run can
|
||||
swing; the mean is what tells you whether the edit actually helped.
|
||||
|
||||
## 1–2. Serve the model and point PenguinHarness at it
|
||||
|
||||
```bash
|
||||
export HIP_VISIBLE_DEVICES=0 # optional: pin a specific AMD GPU
|
||||
ollama serve &
|
||||
ollama pull qwen3:8b
|
||||
|
||||
penguin config model add \
|
||||
--model-id qwen3:8b \
|
||||
--provider custom --client-type openai \
|
||||
--base-url http://localhost:11434/v1 \
|
||||
--api-key ollama --set-default
|
||||
```
|
||||
|
||||
## 3. Run the example
|
||||
|
||||
```bash
|
||||
pnpm install
|
||||
pnpm build
|
||||
pnpm --dir examples/self-improving-agent start
|
||||
# or directly: npx tsx examples/self-improving-agent/self-improve.ts
|
||||
```
|
||||
|
||||
## What you should see
|
||||
|
||||
```text
|
||||
BASELINE (blank AGENTS.md): 3 runs
|
||||
run 1: 0/5
|
||||
run 2: 0/5
|
||||
run 3: 0/5
|
||||
BASELINE mean: 0.00/5
|
||||
N+1 (with working discipline): 3 runs
|
||||
run 1: 5/5
|
||||
run 2: 5/5
|
||||
run 3: 5/5
|
||||
N+1 mean: 5.00/5
|
||||
=== Self-improvement result ===
|
||||
baseline: 0.00/5 → N+1: 5.00/5
|
||||
Mean score improved — keep version N+1. ✔
|
||||
```
|
||||
|
||||
With a blank `AGENTS.md`, `qwen3:8b` tends to *narrate* the summary in chat and never call the
|
||||
tool to write the file — so the rubric scores it 0. Adding a short "working discipline" section
|
||||
(read first, restate the constraints, actually write the file, self-check) flips that. Exact
|
||||
numbers vary run to run; the averaged direction is the point.
|
||||
|
||||
## Notes
|
||||
|
||||
- Uses a dedicated agent id (`self-improve-demo`), created on the fly — your own agents are
|
||||
untouched.
|
||||
- Re-running updates that demo agent in place.
|
||||
@@ -0,0 +1,72 @@
|
||||
<!-- [English](README.md) | 简体中文 -->
|
||||
|
||||
# 示例:一个自我改进的 Agent(本地,经 Ollama 跑在 AMD GPU 上)
|
||||
|
||||
这个示例是 **“递归自我进化”** 支柱的可运行版本。仅用 PenguinHarness SDK,它会跑完自我进化循环的一轮:
|
||||
|
||||
1. **评估(Evaluate)** —— 让 Agent 完成一个约束型写作任务,并按 rubric 打分。
|
||||
2. **诊断(Diagnose)** —— 从运行结果中看出哪些 rubric 项失了分。
|
||||
3. **编辑(Edit)** —— 改写 Agent 自己的 `AGENTS.md` 来修复失分点(版本 N+1)。
|
||||
4. **重新评估(Re-evaluate)** —— 重跑同一任务,只有分数提升时才保留这次改动。
|
||||
|
||||
全程跑在一个**本地开源权重模型**上——由 Ollama 提供的 `qwen3:8b`——所以不用任何云端 API、数据也不
|
||||
离开本机。Ollama 的 ROCm 后端能原生在 AMD GPU 上运行它。
|
||||
|
||||
## 为什么用确定性 rubric,为什么要取平均
|
||||
|
||||
- **rubric 就是你能读懂的普通代码**(`self-improve.ts` 里的 `score()`):文件是否确实写出 · 概述
|
||||
≤ 2 句 · 恰好 3 条要点 · 不超过 60 词 · 关键事实是否出现。没有隐藏的裁判——前后分数是客观、可复现的。
|
||||
在完整产品里,Evaluator 由 `agent-evaluation` skill 按一份*私有* rubric 驱动;这个示例把那个理念
|
||||
蒸馏成可运行的核心。
|
||||
- 本地模型是**非确定性**的,所以示例对每个版本各跑多次取平均——这正是真实 benchmark 里为每个 Case
|
||||
设置 `runs`(多次运行)的原因。单次可能上下波动;均值才能告诉你这次编辑到底有没有帮助。
|
||||
|
||||
## 1–2. 提供模型并把 PenguinHarness 指向它
|
||||
|
||||
```bash
|
||||
export HIP_VISIBLE_DEVICES=0 # 可选:指定某张 AMD GPU
|
||||
ollama serve &
|
||||
ollama pull qwen3:8b
|
||||
|
||||
penguin config model add \
|
||||
--model-id qwen3:8b \
|
||||
--provider custom --client-type openai \
|
||||
--base-url http://localhost:11434/v1 \
|
||||
--api-key ollama --set-default
|
||||
```
|
||||
|
||||
## 3. 运行示例
|
||||
|
||||
```bash
|
||||
pnpm install
|
||||
pnpm build
|
||||
pnpm --dir examples/self-improving-agent start
|
||||
# 或直接运行: npx tsx examples/self-improving-agent/self-improve.ts
|
||||
```
|
||||
|
||||
## 你应当看到什么
|
||||
|
||||
```text
|
||||
BASELINE (blank AGENTS.md): 3 runs
|
||||
run 1: 0/5
|
||||
run 2: 0/5
|
||||
run 3: 0/5
|
||||
BASELINE mean: 0.00/5
|
||||
N+1 (with working discipline): 3 runs
|
||||
run 1: 5/5
|
||||
run 2: 5/5
|
||||
run 3: 5/5
|
||||
N+1 mean: 5.00/5
|
||||
=== Self-improvement result ===
|
||||
baseline: 0.00/5 → N+1: 5.00/5
|
||||
Mean score improved — keep version N+1. ✔
|
||||
```
|
||||
|
||||
当 `AGENTS.md` 为空时,`qwen3:8b` 往往会在对话里*叙述*摘要,却从不调用工具把文件写出来——于是
|
||||
rubric 判它 0 分。加上一小段“任务纪律”(先读原文、把约束逐条列出、真正把文件写出来、结束前自检)就能
|
||||
扭转这一点。具体数字每次运行会有波动;重点是取平均后的方向。
|
||||
|
||||
## 说明
|
||||
|
||||
- 使用一个专用 agent id(`self-improve-demo`),运行时即时创建——你自己的 Agent 不会被动到。
|
||||
- 再次运行会就地更新这个 demo agent。
|
||||
@@ -0,0 +1,14 @@
|
||||
{
|
||||
"name": "@prismshadow/example-self-improving-agent",
|
||||
"private": true,
|
||||
"type": "module",
|
||||
"scripts": {
|
||||
"start": "tsx self-improve.ts"
|
||||
},
|
||||
"dependencies": {
|
||||
"@prismshadow/penguin-core": "workspace:*"
|
||||
},
|
||||
"devDependencies": {
|
||||
"tsx": "^4.20.0"
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,192 @@
|
||||
/**
|
||||
* Example: an Agent that improves itself — one turn of the self-improvement loop, in code.
|
||||
*
|
||||
* This is the "Recursive Self-Improvement" pillar made runnable. It runs entirely on a local
|
||||
* open-weight model (Ollama serving qwen3:8b) — see README.md for the one-time setup.
|
||||
*
|
||||
* The loop, exactly as the docs describe it:
|
||||
* 1. EVALUATE — run the agent on a constrained task, score it against a rubric.
|
||||
* 2. DIAGNOSE — read the result to see which rubric points were lost.
|
||||
* 3. EDIT — rewrite the agent's own AGENTS.md to address the failure (version N+1).
|
||||
* 4. RE-EVALUATE — run the same task again; keep the change only if the score improved.
|
||||
*
|
||||
* The rubric here is a *deterministic, transparent* scorer (plain code you can read below), so the
|
||||
* before/after numbers are objective and reproducible — no hidden judge. In the full product the
|
||||
* Evaluator is driven by the `agent-evaluation` skill against a private rubric; this example
|
||||
* distills that idea to its runnable core.
|
||||
*
|
||||
* It uses a dedicated agent id (`self-improve-demo`) created on the fly, so your existing agents
|
||||
* are never touched.
|
||||
*
|
||||
* Run: pnpm --dir examples/self-improving-agent start
|
||||
* or: npx tsx examples/self-improving-agent/self-improve.ts
|
||||
*/
|
||||
import { promises as fs } from "node:fs";
|
||||
import os from "node:os";
|
||||
import path from "node:path";
|
||||
import {
|
||||
createAgent,
|
||||
userText,
|
||||
isCompleteModelMessage,
|
||||
type OmniMessage,
|
||||
} from "@prismshadow/penguin-core";
|
||||
|
||||
const AGENT_ID = "self-improve-demo";
|
||||
const PROJECT_ID = "default_project";
|
||||
|
||||
/** The constrained task the agent must complete. */
|
||||
const TASK = `Read notes.txt in your workspace and write summary.md with:
|
||||
(1) an overview of at most 2 sentences, and
|
||||
(2) a bullet list of exactly 3 key facts.
|
||||
The entire summary.md must be under 60 words total.`;
|
||||
|
||||
const NOTES = `Project Aurora — Internal Notes
|
||||
|
||||
Aurora is a real-time analytics platform launched in Q1 2026. At peak it ingests roughly
|
||||
2 million events per second through a Kafka-based pipeline. The core query engine was rewritten
|
||||
in Rust after the original Go version could not keep p99 latency under control; the rewrite cut
|
||||
p99 from 800ms to 120ms.
|
||||
|
||||
The production deployment is single-region with no failover; a multi-region rollout is scheduled
|
||||
for Q3 2026. Storage is the largest cost line at about $48,000/month, driven by a 90-day
|
||||
hot-retention policy; cutting retention to 30 days would reduce storage cost by roughly 55%.
|
||||
`;
|
||||
|
||||
/** The "working discipline" we install into AGENTS.md as the N→N+1 edit. */
|
||||
const DISCIPLINE = `# Working discipline for constrained tasks
|
||||
|
||||
When a task lists explicit rules or an exact output format, follow this discipline:
|
||||
|
||||
1. Read the source first with a tool before producing any output; never guess file contents.
|
||||
2. Restate every rule from the task as a short checklist before you start, so none is missed.
|
||||
3. Actually call the tool to write the required file — do not emit the answer as chat text.
|
||||
4. Before finishing, re-read your output file and verify it against every rule (file written,
|
||||
sentence count, bullet count, word count), fixing any mismatch first.
|
||||
`;
|
||||
|
||||
/** rootDir mirrors the SDK's resolveRoot(): PENGUIN_HOME or ~/.penguin/data. */
|
||||
function rootDir(): string {
|
||||
return process.env.PENGUIN_HOME ?? path.join(os.homedir(), ".penguin", "data");
|
||||
}
|
||||
|
||||
function agentStateDir(): string {
|
||||
return path.join(rootDir(), PROJECT_ID, "agents", AGENT_ID, "agent_state");
|
||||
}
|
||||
|
||||
/** Deterministic rubric — 5 independent points, all checkable in code. Returns {score, detail}. */
|
||||
function score(
|
||||
summaryPath: string,
|
||||
summaryText: string | null,
|
||||
): { score: number; detail: string[] } {
|
||||
const detail: string[] = [];
|
||||
let s = 0;
|
||||
const exists = summaryText !== null;
|
||||
detail.push(`${exists ? "1" : "0"}/1 file summary.md was actually written`);
|
||||
if (!exists) return { score: 0, detail };
|
||||
s += 1;
|
||||
|
||||
const body = summaryText.replace(/^#.*$/gm, "").trim(); // drop a markdown title line if present
|
||||
const overview = body.split(/\n\s*[-*]/)[0] ?? ""; // text before the first bullet
|
||||
const sentences = (overview.match(/[.!?](\s|$)/g) ?? []).length;
|
||||
const okSentences = sentences >= 1 && sentences <= 2;
|
||||
detail.push(`${okSentences ? "1" : "0"}/1 overview is ≤ 2 sentences (found ${sentences})`);
|
||||
if (okSentences) s += 1;
|
||||
|
||||
const bullets = (summaryText.match(/^\s*[-*]\s+/gm) ?? []).length;
|
||||
const okBullets = bullets === 3;
|
||||
detail.push(`${okBullets ? "1" : "0"}/1 exactly 3 bullet facts (found ${bullets})`);
|
||||
if (okBullets) s += 1;
|
||||
|
||||
const words = body.split(/\s+/).filter(Boolean).length;
|
||||
const okWords = words < 60;
|
||||
detail.push(`${okWords ? "1" : "0"}/1 under 60 words (found ${words})`);
|
||||
if (okWords) s += 1;
|
||||
|
||||
// "facts accurate": require at least two of the source's hard numbers to appear.
|
||||
const anchors = ["120ms", "55%", "2 million", "$48,000", "Q3 2026", "800ms"];
|
||||
const hits = anchors.filter((a) => summaryText.includes(a)).length;
|
||||
const okFacts = hits >= 2;
|
||||
detail.push(`${okFacts ? "1" : "0"}/1 key facts accurate (${hits} source figures present)`);
|
||||
if (okFacts) s += 1;
|
||||
|
||||
void summaryPath;
|
||||
return { score: s, detail };
|
||||
}
|
||||
|
||||
async function drain(run: AsyncGenerator<OmniMessage>): Promise<void> {
|
||||
for await (const msg of run) {
|
||||
if (isCompleteModelMessage(msg) && msg.payload.type === "text") {
|
||||
process.stdout.write(msg.payload.text + "\n");
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/** Run the task once in a fresh temp workspace seeded with notes.txt; return the rubric score. */
|
||||
async function runOnce(): Promise<{ score: number; detail: string[] }> {
|
||||
const agent = await createAgent({ agentId: AGENT_ID });
|
||||
const ws = await fs.mkdtemp(path.join(os.tmpdir(), "self-improve-"));
|
||||
await fs.writeFile(path.join(ws, "notes.txt"), NOTES, "utf8");
|
||||
|
||||
const session = await agent.createSession({ workspaceDir: ws });
|
||||
await drain(session.run([userText(TASK)], { approve: async () => "allow" }));
|
||||
|
||||
const summaryPath = path.join(ws, "summary.md");
|
||||
let text: string | null = null;
|
||||
try {
|
||||
text = await fs.readFile(summaryPath, "utf8");
|
||||
} catch {
|
||||
text = null;
|
||||
}
|
||||
return score(summaryPath, text);
|
||||
}
|
||||
|
||||
/**
|
||||
* Evaluate a version by averaging over RUNS independent runs. Small local models are
|
||||
* nondeterministic, so a single run is noisy — averaging is exactly why real benchmarks use a
|
||||
* `runs` count per case. Returns the mean score.
|
||||
*/
|
||||
const RUNS = 3;
|
||||
async function evaluate(label: string): Promise<number> {
|
||||
console.log(`\n--- ${label}: ${RUNS} runs ---`);
|
||||
const scores: number[] = [];
|
||||
for (let i = 0; i < RUNS; i++) {
|
||||
const r = await runOnce();
|
||||
scores.push(r.score);
|
||||
console.log(` run ${i + 1}: ${r.score}/5`);
|
||||
}
|
||||
const mean = scores.reduce((a, b) => a + b, 0) / scores.length;
|
||||
console.log(` ${label} mean: ${mean.toFixed(2)}/5`);
|
||||
return mean;
|
||||
}
|
||||
|
||||
async function main(): Promise<void> {
|
||||
// Ensure the demo agent exists, then start from a blank AGENTS.md (the baseline).
|
||||
await createAgent({ agentId: AGENT_ID });
|
||||
const agentsMd = path.join(agentStateDir(), "AGENTS.md");
|
||||
await fs.writeFile(agentsMd, "", "utf8");
|
||||
|
||||
// --- Turn N: evaluate the baseline ---------------------------------------------------
|
||||
const baseline = await evaluate("BASELINE (blank AGENTS.md)");
|
||||
|
||||
// --- Edit: install the working discipline (version N+1) ------------------------------
|
||||
console.log("\n--- EDIT: writing a working-discipline section into the agent's AGENTS.md ---");
|
||||
await fs.writeFile(agentsMd, DISCIPLINE, "utf8");
|
||||
|
||||
// --- Turn N+1: re-evaluate the same task --------------------------------------------
|
||||
const improved = await evaluate("N+1 (with working discipline)");
|
||||
|
||||
// --- Keep-or-roll-back: the loop's decision rule ------------------------------------
|
||||
console.log("\n=== Self-improvement result ===");
|
||||
console.log(` baseline: ${baseline.toFixed(2)}/5 → N+1: ${improved.toFixed(2)}/5`);
|
||||
if (improved > baseline) {
|
||||
console.log(" Mean score improved — keep version N+1. ✔");
|
||||
} else {
|
||||
console.log(" No improvement — roll back to the baseline AGENTS.md.");
|
||||
await fs.writeFile(agentsMd, "", "utf8");
|
||||
}
|
||||
}
|
||||
|
||||
main().catch((err) => {
|
||||
console.error("Example failed:", err);
|
||||
process.exitCode = 1;
|
||||
});
|
||||
Reference in New Issue
Block a user