docs(blog): local agents on AMD GPUs, with runnable examples (#14)

This commit is contained in:
Zhang Jason
2026-07-21 20:20:51 +08:00
committed by GitHub
parent 0e134e9c6c
commit d5c2bc001d
10 changed files with 1095 additions and 0 deletions
+75
View File
@@ -0,0 +1,75 @@
<!-- English | [简体中文](README.zh.md) -->
# Example: an Agent that builds another Agent (local, on an AMD GPU via Ollama)
This example is the **"Harness for Building Agents"** pillar in runnable code. Using only the
PenguinHarness SDK, it:
1. **builds** a brand-new agent (`commit-helper`) by driving `default_agent` with the
`agent-creation` skill from a plain-language requirement, then
2. **runs** that freshly-created agent to prove the generated `AGENTS.md` actually shapes its
behavior (it writes a Conventional Commits message).
Everything runs on a **local open-weight model** — `qwen3.6:35b` served by Ollama — so no cloud
API and no data leaving the machine. Ollama's ROCm backend runs this natively on AMD GPUs
(from Radeon PRO workstation cards up to Instinct accelerators).
## 1. Serve the model locally with Ollama
```bash
# Ollama detects an AMD GPU (ROCm) automatically; pin a specific card if you like:
export HIP_VISIBLE_DEVICES=0
ollama serve & # if not already running as a service
ollama pull qwen3.6:35b
```
## 2. Point PenguinHarness at it (once)
```bash
penguin config model add \
--model-id qwen3.6:35b \
--provider custom --client-type openai \
--base-url http://localhost:11434/v1 \
--api-key ollama --set-default
```
This writes the model into `~/.penguin/data/default_project/.project_config.toml`. The example
uses the project's default model, so no model id is hard-coded in the script.
## 3. Run the example
From the repo root (build the workspace first so `@prismshadow/penguin-core` resolves to its
`dist/`):
```bash
pnpm install
pnpm build
pnpm --dir examples/build-agent-with-agent start
# or directly: npx tsx examples/build-agent-with-agent/build-agent.ts
```
## What you should see
- **Phase 1** — `default_agent` scaffolds `agents/commit-helper/` under the project: its
directory layout, a copied `system_config.yaml` (with name + description), and an `AGENTS.md`
encoding the Conventional Commits rules.
- **Phase 2** — the new `commit-helper` agent, following only that generated `AGENTS.md`,
produces something like:
```text
fix(payment): add retry-with-backoff for transient gateway 503 errors
Transient 503 responses from the payment gateway were causing checkout
failures during peak traffic. Retry with exponential backoff gives the
gateway time to recover, preventing spurious user-facing errors.
```
## Notes
- Output quality tracks the model. `qwen3.6:35b` handles this task well; much smaller models
may not follow the tool-calling protocol reliably.
- A capable model occasionally introduces a mechanical slip when re-serializing the base config
(e.g. an invalid YAML escape) — every file is plain text and every run is traced, so such
slips are quick to spot and fix. This is the realistic shape of "agents building agents":
the requirement→`AGENTS.md` heavy lifting is automated; a human reviews the mechanical edges.
- Re-running Phase 1 will update the existing `commit-helper` agent in place.
@@ -0,0 +1,70 @@
<!-- [English](README.md) | 简体中文 -->
# 示例:用一个 Agent 构建另一个 Agent(本地,经 Ollama 跑在 AMD GPU 上)
这个示例是 **“构建 Agent 的 Harness”** 支柱的可运行版本。仅用 PenguinHarness SDK,它会:
1. **构建**一个全新的 Agent(`commit-helper`)——用 `agent-creation` skill 驱动 `default_agent`,
根据一句大白话需求把它搭建出来;然后
2. **运行**这个刚创建的 Agent,证明生成的 `AGENTS.md` 确实塑造了它的行为(它会写出一条
Conventional Commits 提交信息)。
全程跑在一个**本地开源权重模型**上——由 Ollama 提供的 `qwen3.6:35b`——所以不用任何云端 API、
数据也不离开本机。Ollama 的 ROCm 后端能原生在 AMD GPU 上运行它(从 Radeon PRO 工作站显卡一直到
Instinct 加速卡)。
## 1. 用 Ollama 在本地提供模型
```bash
# Ollama 会自动识别 AMD GPU(ROCm);也可以指定某张卡:
export HIP_VISIBLE_DEVICES=0
ollama serve & # 若尚未作为服务运行
ollama pull qwen3.6:35b
```
## 2. 把 PenguinHarness 指向它(一次即可)
```bash
penguin config model add \
--model-id qwen3.6:35b \
--provider custom --client-type openai \
--base-url http://localhost:11434/v1 \
--api-key ollama --set-default
```
这会把模型写进 `~/.penguin/data/default_project/.project_config.toml`。示例使用项目的默认模型,
所以脚本里没有硬编码任何 model id。
## 3. 运行示例
从仓库根目录(先构建 workspace,好让 `@prismshadow/penguin-core` 解析到它的 `dist/`):
```bash
pnpm install
pnpm build
pnpm --dir examples/build-agent-with-agent start
# 或直接运行: npx tsx examples/build-agent-with-agent/build-agent.ts
```
## 你应当看到什么
- **第一阶段** —— `default_agent` 在项目下搭建出 `agents/commit-helper/`:它的目录布局、一份复制
来的 `system_config.yaml`(含 name + description),以及一份编码了 Conventional Commits 规则的
`AGENTS.md`。
- **第二阶段** —— 这个新的 `commit-helper` Agent,只依据那份生成的 `AGENTS.md`,产出类似这样的结果:
```text
fix(payment): add retry-with-backoff for transient gateway 503 errors
Transient 503 responses from the payment gateway were causing checkout
failures during peak traffic. Retry with exponential backoff gives the
gateway time to recover, preventing spurious user-facing errors.
```
## 说明
- 输出质量取决于模型。`qwen3.6:35b` 能很好地完成这个任务;更小的模型可能无法可靠地遵循工具调用协议。
- 有能力的模型偶尔会在重新序列化 base config 时引入一个机械性的小失误(例如非法的 YAML 转义)——
因为每个文件都是纯文本、每次运行都被追踪,这类失误很容易被发现并修复。这正是“用 Agent 造 Agent”
真实的样子:把需求变成 `AGENTS.md` 的重活被自动化了,而机械层面的毛刺由人类来审阅。
- 再次运行第一阶段会就地更新已存在的 `commit-helper` Agent。
@@ -0,0 +1,82 @@
/**
* Example: an Agent that builds another Agent — programmatically, via the SDK.
*
* This is the "Harness for Building Agents" pillar in code. It runs entirely on a local
* open-weight model (Ollama serving qwen3.6:35b) — see README.md for the one-time setup.
*
* Two phases, both driven purely through the SDK:
* 1. BUILD — drive `default_agent` with the `agent-creation` skill to scaffold a brand-new
* agent (`commit-helper`) from a plain-language requirement.
* 2. RUN — load the freshly-created agent and have it do its job (write a commit message),
* proving the generated AGENTS.md actually shapes its behavior.
*
* Run: pnpm --dir examples/build-agent-with-agent start
* or: npx tsx examples/build-agent-with-agent/build-agent.ts
*/
import {
createAgent,
userText,
isCompleteModelMessage,
type OmniMessage,
} from "@prismshadow/penguin-core";
/** Stream a Session run to stdout and collect the assistant's final text. */
async function runToStdout(run: AsyncGenerator<OmniMessage>): Promise<string> {
let finalText = "";
for await (const msg of run) {
if (isCompleteModelMessage(msg) && msg.payload.type === "text") {
finalText += msg.payload.text;
process.stdout.write(msg.payload.text + "\n");
}
}
return finalText;
}
const BUILD_REQUEST = `Use the agent-creation skill to create a brand-new agent in this project.
Requirement: an agent called "commit-helper" that writes high-quality git commit messages.
Given a diff or a description of changes, it must produce a Conventional Commits message: a
\`type(scope): subject\` header (type one of feat/fix/docs/refactor/test/chore), subject in
imperative mood and under 50 characters, then a blank line and a short body explaining the "why".
Do everything the agent-creation skill specifies: create the agent directory layout, copy the
base system_config.yaml, write a concise AGENTS.md capturing this requirement, and set the new
agent's name and description in system_config.yaml. Report the files you created when done.`;
const COMMIT_TASK = `Write a commit message for this change: I added retry-with-backoff logic to
the payment API client because transient 503s from the gateway were causing checkout failures.
Touched packages/core/src/payment/client.ts.`;
async function main(): Promise<void> {
// --- Phase 1: an Agent builds an Agent ------------------------------------------------
console.log(
"=== Phase 1: default_agent is building a new agent via the agent-creation skill ===\n",
);
const builder = await createAgent({ agentId: "default_agent" });
const buildSession = await builder.createSession({ workspaceDir: process.cwd() });
// approve: () => "allow" lets the builder run its shell tool calls unattended. In a real
// integration you would inspect each tool_call and decide — that is the whole point of the
// per-call approval callback.
await runToStdout(buildSession.run([userText(BUILD_REQUEST)], { approve: async () => "allow" }));
// --- Phase 2: run the agent that was just built ---------------------------------------
console.log("\n=== Phase 2: running the freshly-created commit-helper agent ===\n");
const helper = await createAgent({ agentId: "commit-helper" });
const helperSession = await helper.createSession({ workspaceDir: process.cwd() });
const commitMessage = await runToStdout(
helperSession.run([userText(COMMIT_TASK)], { approve: async () => "allow" }),
);
console.log("\n=== Done. commit-helper produced the message above. ===");
if (!commitMessage.trim()) {
console.error("(No text output — check that Ollama is running and the model is configured.)");
process.exitCode = 1;
}
}
main().catch((err) => {
console.error("Example failed:", err);
process.exitCode = 1;
});
@@ -0,0 +1,14 @@
{
"name": "@prismshadow/example-build-agent-with-agent",
"private": true,
"type": "module",
"scripts": {
"start": "tsx build-agent.ts"
},
"dependencies": {
"@prismshadow/penguin-core": "workspace:*"
},
"devDependencies": {
"tsx": "^4.20.0"
}
}
+77
View File
@@ -0,0 +1,77 @@
<!-- English | [简体中文](README.zh.md) -->
# Example: an Agent that improves itself (local, on an AMD GPU via Ollama)
This example is the **"Recursive Self-Improvement"** pillar in runnable code. Using only the
PenguinHarness SDK, it runs one turn of the self-improvement loop:
1. **Evaluate** a constrained writing task and score it against a rubric.
2. **Diagnose** which rubric points were lost (from the run itself).
3. **Edit** the agent's own `AGENTS.md` to address the failure (version N+1).
4. **Re-evaluate** the same task and keep the change only if the score improved.
Everything runs on a **local open-weight model** — `qwen3:8b` served by Ollama — so no cloud API
and no data leaving the machine. Ollama's ROCm backend runs this natively on AMD GPUs.
## Why a deterministic rubric, and why averaging
- The **rubric is plain code you can read** (`score()` in `self-improve.ts`): file actually
written · overview ≤ 2 sentences · exactly 3 bullets · under 60 words · key facts present. No
hidden judge — the before/after numbers are objective and reproducible. In the full product the
Evaluator is driven by the `agent-evaluation` skill against a *private* rubric; this example
distills that idea to its runnable core.
- A local model is **nondeterministic**, so the example runs each version several times and
averages — which is exactly why real benchmarks use a `runs` count per case. A single run can
swing; the mean is what tells you whether the edit actually helped.
## 1–2. Serve the model and point PenguinHarness at it
```bash
export HIP_VISIBLE_DEVICES=0 # optional: pin a specific AMD GPU
ollama serve &
ollama pull qwen3:8b
penguin config model add \
--model-id qwen3:8b \
--provider custom --client-type openai \
--base-url http://localhost:11434/v1 \
--api-key ollama --set-default
```
## 3. Run the example
```bash
pnpm install
pnpm build
pnpm --dir examples/self-improving-agent start
# or directly: npx tsx examples/self-improving-agent/self-improve.ts
```
## What you should see
```text
BASELINE (blank AGENTS.md): 3 runs
run 1: 0/5
run 2: 0/5
run 3: 0/5
BASELINE mean: 0.00/5
N+1 (with working discipline): 3 runs
run 1: 5/5
run 2: 5/5
run 3: 5/5
N+1 mean: 5.00/5
=== Self-improvement result ===
baseline: 0.00/5 → N+1: 5.00/5
Mean score improved — keep version N+1. ✔
```
With a blank `AGENTS.md`, `qwen3:8b` tends to *narrate* the summary in chat and never call the
tool to write the file — so the rubric scores it 0. Adding a short "working discipline" section
(read first, restate the constraints, actually write the file, self-check) flips that. Exact
numbers vary run to run; the averaged direction is the point.
## Notes
- Uses a dedicated agent id (`self-improve-demo`), created on the fly — your own agents are
untouched.
- Re-running updates that demo agent in place.
@@ -0,0 +1,72 @@
<!-- [English](README.md) | 简体中文 -->
# 示例:一个自我改进的 Agent(本地,经 Ollama 跑在 AMD GPU 上)
这个示例是 **“递归自我进化”** 支柱的可运行版本。仅用 PenguinHarness SDK,它会跑完自我进化循环的一轮:
1. **评估(Evaluate)** —— 让 Agent 完成一个约束型写作任务,并按 rubric 打分。
2. **诊断(Diagnose)** —— 从运行结果中看出哪些 rubric 项失了分。
3. **编辑(Edit)** —— 改写 Agent 自己的 `AGENTS.md` 来修复失分点(版本 N+1)。
4. **重新评估(Re-evaluate)** —— 重跑同一任务,只有分数提升时才保留这次改动。
全程跑在一个**本地开源权重模型**上——由 Ollama 提供的 `qwen3:8b`——所以不用任何云端 API、数据也不
离开本机。Ollama 的 ROCm 后端能原生在 AMD GPU 上运行它。
## 为什么用确定性 rubric,为什么要取平均
- **rubric 就是你能读懂的普通代码**(`self-improve.ts` 里的 `score()`):文件是否确实写出 · 概述
≤ 2 句 · 恰好 3 条要点 · 不超过 60 词 · 关键事实是否出现。没有隐藏的裁判——前后分数是客观、可复现的。
在完整产品里,Evaluator 由 `agent-evaluation` skill 按一份*私有* rubric 驱动;这个示例把那个理念
蒸馏成可运行的核心。
- 本地模型是**非确定性**的,所以示例对每个版本各跑多次取平均——这正是真实 benchmark 里为每个 Case
设置 `runs`(多次运行)的原因。单次可能上下波动;均值才能告诉你这次编辑到底有没有帮助。
## 1–2. 提供模型并把 PenguinHarness 指向它
```bash
export HIP_VISIBLE_DEVICES=0 # 可选:指定某张 AMD GPU
ollama serve &
ollama pull qwen3:8b
penguin config model add \
--model-id qwen3:8b \
--provider custom --client-type openai \
--base-url http://localhost:11434/v1 \
--api-key ollama --set-default
```
## 3. 运行示例
```bash
pnpm install
pnpm build
pnpm --dir examples/self-improving-agent start
# 或直接运行: npx tsx examples/self-improving-agent/self-improve.ts
```
## 你应当看到什么
```text
BASELINE (blank AGENTS.md): 3 runs
run 1: 0/5
run 2: 0/5
run 3: 0/5
BASELINE mean: 0.00/5
N+1 (with working discipline): 3 runs
run 1: 5/5
run 2: 5/5
run 3: 5/5
N+1 mean: 5.00/5
=== Self-improvement result ===
baseline: 0.00/5 → N+1: 5.00/5
Mean score improved — keep version N+1. ✔
```
当 `AGENTS.md` 为空时,`qwen3:8b` 往往会在对话里*叙述*摘要,却从不调用工具把文件写出来——于是
rubric 判它 0 分。加上一小段“任务纪律”(先读原文、把约束逐条列出、真正把文件写出来、结束前自检)就能
扭转这一点。具体数字每次运行会有波动;重点是取平均后的方向。
## 说明
- 使用一个专用 agent id(`self-improve-demo`),运行时即时创建——你自己的 Agent 不会被动到。
- 再次运行会就地更新这个 demo agent。
@@ -0,0 +1,14 @@
{
"name": "@prismshadow/example-self-improving-agent",
"private": true,
"type": "module",
"scripts": {
"start": "tsx self-improve.ts"
},
"dependencies": {
"@prismshadow/penguin-core": "workspace:*"
},
"devDependencies": {
"tsx": "^4.20.0"
}
}
@@ -0,0 +1,192 @@
/**
* Example: an Agent that improves itself — one turn of the self-improvement loop, in code.
*
* This is the "Recursive Self-Improvement" pillar made runnable. It runs entirely on a local
* open-weight model (Ollama serving qwen3:8b) — see README.md for the one-time setup.
*
* The loop, exactly as the docs describe it:
* 1. EVALUATE — run the agent on a constrained task, score it against a rubric.
* 2. DIAGNOSE — read the result to see which rubric points were lost.
* 3. EDIT — rewrite the agent's own AGENTS.md to address the failure (version N+1).
* 4. RE-EVALUATE — run the same task again; keep the change only if the score improved.
*
* The rubric here is a *deterministic, transparent* scorer (plain code you can read below), so the
* before/after numbers are objective and reproducible — no hidden judge. In the full product the
* Evaluator is driven by the `agent-evaluation` skill against a private rubric; this example
* distills that idea to its runnable core.
*
* It uses a dedicated agent id (`self-improve-demo`) created on the fly, so your existing agents
* are never touched.
*
* Run: pnpm --dir examples/self-improving-agent start
* or: npx tsx examples/self-improving-agent/self-improve.ts
*/
import { promises as fs } from "node:fs";
import os from "node:os";
import path from "node:path";
import {
createAgent,
userText,
isCompleteModelMessage,
type OmniMessage,
} from "@prismshadow/penguin-core";
const AGENT_ID = "self-improve-demo";
const PROJECT_ID = "default_project";
/** The constrained task the agent must complete. */
const TASK = `Read notes.txt in your workspace and write summary.md with:
(1) an overview of at most 2 sentences, and
(2) a bullet list of exactly 3 key facts.
The entire summary.md must be under 60 words total.`;
const NOTES = `Project Aurora — Internal Notes
Aurora is a real-time analytics platform launched in Q1 2026. At peak it ingests roughly
2 million events per second through a Kafka-based pipeline. The core query engine was rewritten
in Rust after the original Go version could not keep p99 latency under control; the rewrite cut
p99 from 800ms to 120ms.
The production deployment is single-region with no failover; a multi-region rollout is scheduled
for Q3 2026. Storage is the largest cost line at about $48,000/month, driven by a 90-day
hot-retention policy; cutting retention to 30 days would reduce storage cost by roughly 55%.
`;
/** The "working discipline" we install into AGENTS.md as the N→N+1 edit. */
const DISCIPLINE = `# Working discipline for constrained tasks
When a task lists explicit rules or an exact output format, follow this discipline:
1. Read the source first with a tool before producing any output; never guess file contents.
2. Restate every rule from the task as a short checklist before you start, so none is missed.
3. Actually call the tool to write the required file — do not emit the answer as chat text.
4. Before finishing, re-read your output file and verify it against every rule (file written,
sentence count, bullet count, word count), fixing any mismatch first.
`;
/** rootDir mirrors the SDK's resolveRoot(): PENGUIN_HOME or ~/.penguin/data. */
function rootDir(): string {
return process.env.PENGUIN_HOME ?? path.join(os.homedir(), ".penguin", "data");
}
function agentStateDir(): string {
return path.join(rootDir(), PROJECT_ID, "agents", AGENT_ID, "agent_state");
}
/** Deterministic rubric — 5 independent points, all checkable in code. Returns {score, detail}. */
function score(
summaryPath: string,
summaryText: string | null,
): { score: number; detail: string[] } {
const detail: string[] = [];
let s = 0;
const exists = summaryText !== null;
detail.push(`${exists ? "1" : "0"}/1 file summary.md was actually written`);
if (!exists) return { score: 0, detail };
s += 1;
const body = summaryText.replace(/^#.*$/gm, "").trim(); // drop a markdown title line if present
const overview = body.split(/\n\s*[-*]/)[0] ?? ""; // text before the first bullet
const sentences = (overview.match(/[.!?](\s|$)/g) ?? []).length;
const okSentences = sentences >= 1 && sentences <= 2;
detail.push(`${okSentences ? "1" : "0"}/1 overview is ≤ 2 sentences (found ${sentences})`);
if (okSentences) s += 1;
const bullets = (summaryText.match(/^\s*[-*]\s+/gm) ?? []).length;
const okBullets = bullets === 3;
detail.push(`${okBullets ? "1" : "0"}/1 exactly 3 bullet facts (found ${bullets})`);
if (okBullets) s += 1;
const words = body.split(/\s+/).filter(Boolean).length;
const okWords = words < 60;
detail.push(`${okWords ? "1" : "0"}/1 under 60 words (found ${words})`);
if (okWords) s += 1;
// "facts accurate": require at least two of the source's hard numbers to appear.
const anchors = ["120ms", "55%", "2 million", "$48,000", "Q3 2026", "800ms"];
const hits = anchors.filter((a) => summaryText.includes(a)).length;
const okFacts = hits >= 2;
detail.push(`${okFacts ? "1" : "0"}/1 key facts accurate (${hits} source figures present)`);
if (okFacts) s += 1;
void summaryPath;
return { score: s, detail };
}
async function drain(run: AsyncGenerator<OmniMessage>): Promise<void> {
for await (const msg of run) {
if (isCompleteModelMessage(msg) && msg.payload.type === "text") {
process.stdout.write(msg.payload.text + "\n");
}
}
}
/** Run the task once in a fresh temp workspace seeded with notes.txt; return the rubric score. */
async function runOnce(): Promise<{ score: number; detail: string[] }> {
const agent = await createAgent({ agentId: AGENT_ID });
const ws = await fs.mkdtemp(path.join(os.tmpdir(), "self-improve-"));
await fs.writeFile(path.join(ws, "notes.txt"), NOTES, "utf8");
const session = await agent.createSession({ workspaceDir: ws });
await drain(session.run([userText(TASK)], { approve: async () => "allow" }));
const summaryPath = path.join(ws, "summary.md");
let text: string | null = null;
try {
text = await fs.readFile(summaryPath, "utf8");
} catch {
text = null;
}
return score(summaryPath, text);
}
/**
* Evaluate a version by averaging over RUNS independent runs. Small local models are
* nondeterministic, so a single run is noisy — averaging is exactly why real benchmarks use a
* `runs` count per case. Returns the mean score.
*/
const RUNS = 3;
async function evaluate(label: string): Promise<number> {
console.log(`\n--- ${label}: ${RUNS} runs ---`);
const scores: number[] = [];
for (let i = 0; i < RUNS; i++) {
const r = await runOnce();
scores.push(r.score);
console.log(` run ${i + 1}: ${r.score}/5`);
}
const mean = scores.reduce((a, b) => a + b, 0) / scores.length;
console.log(` ${label} mean: ${mean.toFixed(2)}/5`);
return mean;
}
async function main(): Promise<void> {
// Ensure the demo agent exists, then start from a blank AGENTS.md (the baseline).
await createAgent({ agentId: AGENT_ID });
const agentsMd = path.join(agentStateDir(), "AGENTS.md");
await fs.writeFile(agentsMd, "", "utf8");
// --- Turn N: evaluate the baseline ---------------------------------------------------
const baseline = await evaluate("BASELINE (blank AGENTS.md)");
// --- Edit: install the working discipline (version N+1) ------------------------------
console.log("\n--- EDIT: writing a working-discipline section into the agent's AGENTS.md ---");
await fs.writeFile(agentsMd, DISCIPLINE, "utf8");
// --- Turn N+1: re-evaluate the same task --------------------------------------------
const improved = await evaluate("N+1 (with working discipline)");
// --- Keep-or-roll-back: the loop's decision rule ------------------------------------
console.log("\n=== Self-improvement result ===");
console.log(` baseline: ${baseline.toFixed(2)}/5 → N+1: ${improved.toFixed(2)}/5`);
if (improved > baseline) {
console.log(" Mean score improved — keep version N+1. ✔");
} else {
console.log(" No improvement — roll back to the baseline AGENTS.md.");
await fs.writeFile(agentsMd, "", "utf8");
}
}
main().catch((err) => {
console.error("Example failed:", err);
process.exitCode = 1;
});