diff --git a/examples/build-agent-with-agent/README.md b/examples/build-agent-with-agent/README.md new file mode 100644 index 0000000..975d09e --- /dev/null +++ b/examples/build-agent-with-agent/README.md @@ -0,0 +1,75 @@ + + +# Example: an Agent that builds another Agent (local, on an AMD GPU via Ollama) + +This example is the **"Harness for Building Agents"** pillar in runnable code. Using only the +PenguinHarness SDK, it: + +1. **builds** a brand-new agent (`commit-helper`) by driving `default_agent` with the + `agent-creation` skill from a plain-language requirement, then +2. **runs** that freshly-created agent to prove the generated `AGENTS.md` actually shapes its + behavior (it writes a Conventional Commits message). + +Everything runs on a **local open-weight model** — `qwen3.6:35b` served by Ollama — so no cloud +API and no data leaving the machine. Ollama's ROCm backend runs this natively on AMD GPUs +(from Radeon PRO workstation cards up to Instinct accelerators). + +## 1. Serve the model locally with Ollama + +```bash +# Ollama detects an AMD GPU (ROCm) automatically; pin a specific card if you like: +export HIP_VISIBLE_DEVICES=0 +ollama serve & # if not already running as a service +ollama pull qwen3.6:35b +``` + +## 2. Point PenguinHarness at it (once) + +```bash +penguin config model add \ + --model-id qwen3.6:35b \ + --provider custom --client-type openai \ + --base-url http://localhost:11434/v1 \ + --api-key ollama --set-default +``` + +This writes the model into `~/.penguin/data/default_project/.project_config.toml`. The example +uses the project's default model, so no model id is hard-coded in the script. + +## 3. Run the example + +From the repo root (build the workspace first so `@prismshadow/penguin-core` resolves to its +`dist/`): + +```bash +pnpm install +pnpm build +pnpm --dir examples/build-agent-with-agent start +# or directly: npx tsx examples/build-agent-with-agent/build-agent.ts +``` + +## What you should see + +- **Phase 1** — `default_agent` scaffolds `agents/commit-helper/` under the project: its + directory layout, a copied `system_config.yaml` (with name + description), and an `AGENTS.md` + encoding the Conventional Commits rules. +- **Phase 2** — the new `commit-helper` agent, following only that generated `AGENTS.md`, + produces something like: + + ```text + fix(payment): add retry-with-backoff for transient gateway 503 errors + + Transient 503 responses from the payment gateway were causing checkout + failures during peak traffic. Retry with exponential backoff gives the + gateway time to recover, preventing spurious user-facing errors. + ``` + +## Notes + +- Output quality tracks the model. `qwen3.6:35b` handles this task well; much smaller models + may not follow the tool-calling protocol reliably. +- A capable model occasionally introduces a mechanical slip when re-serializing the base config + (e.g. an invalid YAML escape) — every file is plain text and every run is traced, so such + slips are quick to spot and fix. This is the realistic shape of "agents building agents": + the requirement→`AGENTS.md` heavy lifting is automated; a human reviews the mechanical edges. +- Re-running Phase 1 will update the existing `commit-helper` agent in place. diff --git a/examples/build-agent-with-agent/README.zh.md b/examples/build-agent-with-agent/README.zh.md new file mode 100644 index 0000000..5512473 --- /dev/null +++ b/examples/build-agent-with-agent/README.zh.md @@ -0,0 +1,70 @@ + + +# 示例:用一个 Agent 构建另一个 Agent(本地,经 Ollama 跑在 AMD GPU 上) + +这个示例是 **“构建 Agent 的 Harness”** 支柱的可运行版本。仅用 PenguinHarness SDK,它会: + +1. **构建**一个全新的 Agent(`commit-helper`)——用 `agent-creation` skill 驱动 `default_agent`, + 根据一句大白话需求把它搭建出来;然后 +2. **运行**这个刚创建的 Agent,证明生成的 `AGENTS.md` 确实塑造了它的行为(它会写出一条 + Conventional Commits 提交信息)。 + +全程跑在一个**本地开源权重模型**上——由 Ollama 提供的 `qwen3.6:35b`——所以不用任何云端 API、 +数据也不离开本机。Ollama 的 ROCm 后端能原生在 AMD GPU 上运行它(从 Radeon PRO 工作站显卡一直到 +Instinct 加速卡)。 + +## 1. 用 Ollama 在本地提供模型 + +```bash +# Ollama 会自动识别 AMD GPU(ROCm);也可以指定某张卡: +export HIP_VISIBLE_DEVICES=0 +ollama serve & # 若尚未作为服务运行 +ollama pull qwen3.6:35b +``` + +## 2. 把 PenguinHarness 指向它(一次即可) + +```bash +penguin config model add \ + --model-id qwen3.6:35b \ + --provider custom --client-type openai \ + --base-url http://localhost:11434/v1 \ + --api-key ollama --set-default +``` + +这会把模型写进 `~/.penguin/data/default_project/.project_config.toml`。示例使用项目的默认模型, +所以脚本里没有硬编码任何 model id。 + +## 3. 运行示例 + +从仓库根目录(先构建 workspace,好让 `@prismshadow/penguin-core` 解析到它的 `dist/`): + +```bash +pnpm install +pnpm build +pnpm --dir examples/build-agent-with-agent start +# 或直接运行: npx tsx examples/build-agent-with-agent/build-agent.ts +``` + +## 你应当看到什么 + +- **第一阶段** —— `default_agent` 在项目下搭建出 `agents/commit-helper/`:它的目录布局、一份复制 + 来的 `system_config.yaml`(含 name + description),以及一份编码了 Conventional Commits 规则的 + `AGENTS.md`。 +- **第二阶段** —— 这个新的 `commit-helper` Agent,只依据那份生成的 `AGENTS.md`,产出类似这样的结果: + + ```text + fix(payment): add retry-with-backoff for transient gateway 503 errors + + Transient 503 responses from the payment gateway were causing checkout + failures during peak traffic. Retry with exponential backoff gives the + gateway time to recover, preventing spurious user-facing errors. + ``` + +## 说明 + +- 输出质量取决于模型。`qwen3.6:35b` 能很好地完成这个任务;更小的模型可能无法可靠地遵循工具调用协议。 +- 有能力的模型偶尔会在重新序列化 base config 时引入一个机械性的小失误(例如非法的 YAML 转义)—— + 因为每个文件都是纯文本、每次运行都被追踪,这类失误很容易被发现并修复。这正是“用 Agent 造 Agent” + 真实的样子:把需求变成 `AGENTS.md` 的重活被自动化了,而机械层面的毛刺由人类来审阅。 +- 再次运行第一阶段会就地更新已存在的 `commit-helper` Agent。 diff --git a/examples/build-agent-with-agent/build-agent.ts b/examples/build-agent-with-agent/build-agent.ts new file mode 100644 index 0000000..32569e3 --- /dev/null +++ b/examples/build-agent-with-agent/build-agent.ts @@ -0,0 +1,82 @@ +/** + * Example: an Agent that builds another Agent — programmatically, via the SDK. + * + * This is the "Harness for Building Agents" pillar in code. It runs entirely on a local + * open-weight model (Ollama serving qwen3.6:35b) — see README.md for the one-time setup. + * + * Two phases, both driven purely through the SDK: + * 1. BUILD — drive `default_agent` with the `agent-creation` skill to scaffold a brand-new + * agent (`commit-helper`) from a plain-language requirement. + * 2. RUN — load the freshly-created agent and have it do its job (write a commit message), + * proving the generated AGENTS.md actually shapes its behavior. + * + * Run: pnpm --dir examples/build-agent-with-agent start + * or: npx tsx examples/build-agent-with-agent/build-agent.ts + */ +import { + createAgent, + userText, + isCompleteModelMessage, + type OmniMessage, +} from "@prismshadow/penguin-core"; + +/** Stream a Session run to stdout and collect the assistant's final text. */ +async function runToStdout(run: AsyncGenerator): Promise { + let finalText = ""; + for await (const msg of run) { + if (isCompleteModelMessage(msg) && msg.payload.type === "text") { + finalText += msg.payload.text; + process.stdout.write(msg.payload.text + "\n"); + } + } + return finalText; +} + +const BUILD_REQUEST = `Use the agent-creation skill to create a brand-new agent in this project. + +Requirement: an agent called "commit-helper" that writes high-quality git commit messages. +Given a diff or a description of changes, it must produce a Conventional Commits message: a +\`type(scope): subject\` header (type one of feat/fix/docs/refactor/test/chore), subject in +imperative mood and under 50 characters, then a blank line and a short body explaining the "why". + +Do everything the agent-creation skill specifies: create the agent directory layout, copy the +base system_config.yaml, write a concise AGENTS.md capturing this requirement, and set the new +agent's name and description in system_config.yaml. Report the files you created when done.`; + +const COMMIT_TASK = `Write a commit message for this change: I added retry-with-backoff logic to +the payment API client because transient 503s from the gateway were causing checkout failures. +Touched packages/core/src/payment/client.ts.`; + +async function main(): Promise { + // --- Phase 1: an Agent builds an Agent ------------------------------------------------ + console.log( + "=== Phase 1: default_agent is building a new agent via the agent-creation skill ===\n", + ); + const builder = await createAgent({ agentId: "default_agent" }); + const buildSession = await builder.createSession({ workspaceDir: process.cwd() }); + + // approve: () => "allow" lets the builder run its shell tool calls unattended. In a real + // integration you would inspect each tool_call and decide — that is the whole point of the + // per-call approval callback. + await runToStdout(buildSession.run([userText(BUILD_REQUEST)], { approve: async () => "allow" })); + + // --- Phase 2: run the agent that was just built --------------------------------------- + console.log("\n=== Phase 2: running the freshly-created commit-helper agent ===\n"); + const helper = await createAgent({ agentId: "commit-helper" }); + const helperSession = await helper.createSession({ workspaceDir: process.cwd() }); + + const commitMessage = await runToStdout( + helperSession.run([userText(COMMIT_TASK)], { approve: async () => "allow" }), + ); + + console.log("\n=== Done. commit-helper produced the message above. ==="); + if (!commitMessage.trim()) { + console.error("(No text output — check that Ollama is running and the model is configured.)"); + process.exitCode = 1; + } +} + +main().catch((err) => { + console.error("Example failed:", err); + process.exitCode = 1; +}); diff --git a/examples/build-agent-with-agent/package.json b/examples/build-agent-with-agent/package.json new file mode 100644 index 0000000..ad1b6ea --- /dev/null +++ b/examples/build-agent-with-agent/package.json @@ -0,0 +1,14 @@ +{ + "name": "@prismshadow/example-build-agent-with-agent", + "private": true, + "type": "module", + "scripts": { + "start": "tsx build-agent.ts" + }, + "dependencies": { + "@prismshadow/penguin-core": "workspace:*" + }, + "devDependencies": { + "tsx": "^4.20.0" + } +} diff --git a/examples/self-improving-agent/README.md b/examples/self-improving-agent/README.md new file mode 100644 index 0000000..26a39e2 --- /dev/null +++ b/examples/self-improving-agent/README.md @@ -0,0 +1,77 @@ + + +# Example: an Agent that improves itself (local, on an AMD GPU via Ollama) + +This example is the **"Recursive Self-Improvement"** pillar in runnable code. Using only the +PenguinHarness SDK, it runs one turn of the self-improvement loop: + +1. **Evaluate** a constrained writing task and score it against a rubric. +2. **Diagnose** which rubric points were lost (from the run itself). +3. **Edit** the agent's own `AGENTS.md` to address the failure (version N+1). +4. **Re-evaluate** the same task and keep the change only if the score improved. + +Everything runs on a **local open-weight model** — `qwen3:8b` served by Ollama — so no cloud API +and no data leaving the machine. Ollama's ROCm backend runs this natively on AMD GPUs. + +## Why a deterministic rubric, and why averaging + +- The **rubric is plain code you can read** (`score()` in `self-improve.ts`): file actually + written · overview ≤ 2 sentences · exactly 3 bullets · under 60 words · key facts present. No + hidden judge — the before/after numbers are objective and reproducible. In the full product the + Evaluator is driven by the `agent-evaluation` skill against a *private* rubric; this example + distills that idea to its runnable core. +- A local model is **nondeterministic**, so the example runs each version several times and + averages — which is exactly why real benchmarks use a `runs` count per case. A single run can + swing; the mean is what tells you whether the edit actually helped. + +## 1–2. Serve the model and point PenguinHarness at it + +```bash +export HIP_VISIBLE_DEVICES=0 # optional: pin a specific AMD GPU +ollama serve & +ollama pull qwen3:8b + +penguin config model add \ + --model-id qwen3:8b \ + --provider custom --client-type openai \ + --base-url http://localhost:11434/v1 \ + --api-key ollama --set-default +``` + +## 3. Run the example + +```bash +pnpm install +pnpm build +pnpm --dir examples/self-improving-agent start +# or directly: npx tsx examples/self-improving-agent/self-improve.ts +``` + +## What you should see + +```text +BASELINE (blank AGENTS.md): 3 runs + run 1: 0/5 + run 2: 0/5 + run 3: 0/5 + BASELINE mean: 0.00/5 +N+1 (with working discipline): 3 runs + run 1: 5/5 + run 2: 5/5 + run 3: 5/5 + N+1 mean: 5.00/5 +=== Self-improvement result === + baseline: 0.00/5 → N+1: 5.00/5 + Mean score improved — keep version N+1. ✔ +``` + +With a blank `AGENTS.md`, `qwen3:8b` tends to *narrate* the summary in chat and never call the +tool to write the file — so the rubric scores it 0. Adding a short "working discipline" section +(read first, restate the constraints, actually write the file, self-check) flips that. Exact +numbers vary run to run; the averaged direction is the point. + +## Notes + +- Uses a dedicated agent id (`self-improve-demo`), created on the fly — your own agents are + untouched. +- Re-running updates that demo agent in place. diff --git a/examples/self-improving-agent/README.zh.md b/examples/self-improving-agent/README.zh.md new file mode 100644 index 0000000..1913d1a --- /dev/null +++ b/examples/self-improving-agent/README.zh.md @@ -0,0 +1,72 @@ + + +# 示例:一个自我改进的 Agent(本地,经 Ollama 跑在 AMD GPU 上) + +这个示例是 **“递归自我进化”** 支柱的可运行版本。仅用 PenguinHarness SDK,它会跑完自我进化循环的一轮: + +1. **评估(Evaluate)** —— 让 Agent 完成一个约束型写作任务,并按 rubric 打分。 +2. **诊断(Diagnose)** —— 从运行结果中看出哪些 rubric 项失了分。 +3. **编辑(Edit)** —— 改写 Agent 自己的 `AGENTS.md` 来修复失分点(版本 N+1)。 +4. **重新评估(Re-evaluate)** —— 重跑同一任务,只有分数提升时才保留这次改动。 + +全程跑在一个**本地开源权重模型**上——由 Ollama 提供的 `qwen3:8b`——所以不用任何云端 API、数据也不 +离开本机。Ollama 的 ROCm 后端能原生在 AMD GPU 上运行它。 + +## 为什么用确定性 rubric,为什么要取平均 + +- **rubric 就是你能读懂的普通代码**(`self-improve.ts` 里的 `score()`):文件是否确实写出 · 概述 + ≤ 2 句 · 恰好 3 条要点 · 不超过 60 词 · 关键事实是否出现。没有隐藏的裁判——前后分数是客观、可复现的。 + 在完整产品里,Evaluator 由 `agent-evaluation` skill 按一份*私有* rubric 驱动;这个示例把那个理念 + 蒸馏成可运行的核心。 +- 本地模型是**非确定性**的,所以示例对每个版本各跑多次取平均——这正是真实 benchmark 里为每个 Case + 设置 `runs`(多次运行)的原因。单次可能上下波动;均值才能告诉你这次编辑到底有没有帮助。 + +## 1–2. 提供模型并把 PenguinHarness 指向它 + +```bash +export HIP_VISIBLE_DEVICES=0 # 可选:指定某张 AMD GPU +ollama serve & +ollama pull qwen3:8b + +penguin config model add \ + --model-id qwen3:8b \ + --provider custom --client-type openai \ + --base-url http://localhost:11434/v1 \ + --api-key ollama --set-default +``` + +## 3. 运行示例 + +```bash +pnpm install +pnpm build +pnpm --dir examples/self-improving-agent start +# 或直接运行: npx tsx examples/self-improving-agent/self-improve.ts +``` + +## 你应当看到什么 + +```text +BASELINE (blank AGENTS.md): 3 runs + run 1: 0/5 + run 2: 0/5 + run 3: 0/5 + BASELINE mean: 0.00/5 +N+1 (with working discipline): 3 runs + run 1: 5/5 + run 2: 5/5 + run 3: 5/5 + N+1 mean: 5.00/5 +=== Self-improvement result === + baseline: 0.00/5 → N+1: 5.00/5 + Mean score improved — keep version N+1. ✔ +``` + +当 `AGENTS.md` 为空时,`qwen3:8b` 往往会在对话里*叙述*摘要,却从不调用工具把文件写出来——于是 +rubric 判它 0 分。加上一小段“任务纪律”(先读原文、把约束逐条列出、真正把文件写出来、结束前自检)就能 +扭转这一点。具体数字每次运行会有波动;重点是取平均后的方向。 + +## 说明 + +- 使用一个专用 agent id(`self-improve-demo`),运行时即时创建——你自己的 Agent 不会被动到。 +- 再次运行会就地更新这个 demo agent。 diff --git a/examples/self-improving-agent/package.json b/examples/self-improving-agent/package.json new file mode 100644 index 0000000..90dc91f --- /dev/null +++ b/examples/self-improving-agent/package.json @@ -0,0 +1,14 @@ +{ + "name": "@prismshadow/example-self-improving-agent", + "private": true, + "type": "module", + "scripts": { + "start": "tsx self-improve.ts" + }, + "dependencies": { + "@prismshadow/penguin-core": "workspace:*" + }, + "devDependencies": { + "tsx": "^4.20.0" + } +} diff --git a/examples/self-improving-agent/self-improve.ts b/examples/self-improving-agent/self-improve.ts new file mode 100644 index 0000000..a052cce --- /dev/null +++ b/examples/self-improving-agent/self-improve.ts @@ -0,0 +1,192 @@ +/** + * Example: an Agent that improves itself — one turn of the self-improvement loop, in code. + * + * This is the "Recursive Self-Improvement" pillar made runnable. It runs entirely on a local + * open-weight model (Ollama serving qwen3:8b) — see README.md for the one-time setup. + * + * The loop, exactly as the docs describe it: + * 1. EVALUATE — run the agent on a constrained task, score it against a rubric. + * 2. DIAGNOSE — read the result to see which rubric points were lost. + * 3. EDIT — rewrite the agent's own AGENTS.md to address the failure (version N+1). + * 4. RE-EVALUATE — run the same task again; keep the change only if the score improved. + * + * The rubric here is a *deterministic, transparent* scorer (plain code you can read below), so the + * before/after numbers are objective and reproducible — no hidden judge. In the full product the + * Evaluator is driven by the `agent-evaluation` skill against a private rubric; this example + * distills that idea to its runnable core. + * + * It uses a dedicated agent id (`self-improve-demo`) created on the fly, so your existing agents + * are never touched. + * + * Run: pnpm --dir examples/self-improving-agent start + * or: npx tsx examples/self-improving-agent/self-improve.ts + */ +import { promises as fs } from "node:fs"; +import os from "node:os"; +import path from "node:path"; +import { + createAgent, + userText, + isCompleteModelMessage, + type OmniMessage, +} from "@prismshadow/penguin-core"; + +const AGENT_ID = "self-improve-demo"; +const PROJECT_ID = "default_project"; + +/** The constrained task the agent must complete. */ +const TASK = `Read notes.txt in your workspace and write summary.md with: +(1) an overview of at most 2 sentences, and +(2) a bullet list of exactly 3 key facts. +The entire summary.md must be under 60 words total.`; + +const NOTES = `Project Aurora — Internal Notes + +Aurora is a real-time analytics platform launched in Q1 2026. At peak it ingests roughly +2 million events per second through a Kafka-based pipeline. The core query engine was rewritten +in Rust after the original Go version could not keep p99 latency under control; the rewrite cut +p99 from 800ms to 120ms. + +The production deployment is single-region with no failover; a multi-region rollout is scheduled +for Q3 2026. Storage is the largest cost line at about $48,000/month, driven by a 90-day +hot-retention policy; cutting retention to 30 days would reduce storage cost by roughly 55%. +`; + +/** The "working discipline" we install into AGENTS.md as the N→N+1 edit. */ +const DISCIPLINE = `# Working discipline for constrained tasks + +When a task lists explicit rules or an exact output format, follow this discipline: + +1. Read the source first with a tool before producing any output; never guess file contents. +2. Restate every rule from the task as a short checklist before you start, so none is missed. +3. Actually call the tool to write the required file — do not emit the answer as chat text. +4. Before finishing, re-read your output file and verify it against every rule (file written, + sentence count, bullet count, word count), fixing any mismatch first. +`; + +/** rootDir mirrors the SDK's resolveRoot(): PENGUIN_HOME or ~/.penguin/data. */ +function rootDir(): string { + return process.env.PENGUIN_HOME ?? path.join(os.homedir(), ".penguin", "data"); +} + +function agentStateDir(): string { + return path.join(rootDir(), PROJECT_ID, "agents", AGENT_ID, "agent_state"); +} + +/** Deterministic rubric — 5 independent points, all checkable in code. Returns {score, detail}. */ +function score( + summaryPath: string, + summaryText: string | null, +): { score: number; detail: string[] } { + const detail: string[] = []; + let s = 0; + const exists = summaryText !== null; + detail.push(`${exists ? "1" : "0"}/1 file summary.md was actually written`); + if (!exists) return { score: 0, detail }; + s += 1; + + const body = summaryText.replace(/^#.*$/gm, "").trim(); // drop a markdown title line if present + const overview = body.split(/\n\s*[-*]/)[0] ?? ""; // text before the first bullet + const sentences = (overview.match(/[.!?](\s|$)/g) ?? []).length; + const okSentences = sentences >= 1 && sentences <= 2; + detail.push(`${okSentences ? "1" : "0"}/1 overview is ≤ 2 sentences (found ${sentences})`); + if (okSentences) s += 1; + + const bullets = (summaryText.match(/^\s*[-*]\s+/gm) ?? []).length; + const okBullets = bullets === 3; + detail.push(`${okBullets ? "1" : "0"}/1 exactly 3 bullet facts (found ${bullets})`); + if (okBullets) s += 1; + + const words = body.split(/\s+/).filter(Boolean).length; + const okWords = words < 60; + detail.push(`${okWords ? "1" : "0"}/1 under 60 words (found ${words})`); + if (okWords) s += 1; + + // "facts accurate": require at least two of the source's hard numbers to appear. + const anchors = ["120ms", "55%", "2 million", "$48,000", "Q3 2026", "800ms"]; + const hits = anchors.filter((a) => summaryText.includes(a)).length; + const okFacts = hits >= 2; + detail.push(`${okFacts ? "1" : "0"}/1 key facts accurate (${hits} source figures present)`); + if (okFacts) s += 1; + + void summaryPath; + return { score: s, detail }; +} + +async function drain(run: AsyncGenerator): Promise { + for await (const msg of run) { + if (isCompleteModelMessage(msg) && msg.payload.type === "text") { + process.stdout.write(msg.payload.text + "\n"); + } + } +} + +/** Run the task once in a fresh temp workspace seeded with notes.txt; return the rubric score. */ +async function runOnce(): Promise<{ score: number; detail: string[] }> { + const agent = await createAgent({ agentId: AGENT_ID }); + const ws = await fs.mkdtemp(path.join(os.tmpdir(), "self-improve-")); + await fs.writeFile(path.join(ws, "notes.txt"), NOTES, "utf8"); + + const session = await agent.createSession({ workspaceDir: ws }); + await drain(session.run([userText(TASK)], { approve: async () => "allow" })); + + const summaryPath = path.join(ws, "summary.md"); + let text: string | null = null; + try { + text = await fs.readFile(summaryPath, "utf8"); + } catch { + text = null; + } + return score(summaryPath, text); +} + +/** + * Evaluate a version by averaging over RUNS independent runs. Small local models are + * nondeterministic, so a single run is noisy — averaging is exactly why real benchmarks use a + * `runs` count per case. Returns the mean score. + */ +const RUNS = 3; +async function evaluate(label: string): Promise { + console.log(`\n--- ${label}: ${RUNS} runs ---`); + const scores: number[] = []; + for (let i = 0; i < RUNS; i++) { + const r = await runOnce(); + scores.push(r.score); + console.log(` run ${i + 1}: ${r.score}/5`); + } + const mean = scores.reduce((a, b) => a + b, 0) / scores.length; + console.log(` ${label} mean: ${mean.toFixed(2)}/5`); + return mean; +} + +async function main(): Promise { + // Ensure the demo agent exists, then start from a blank AGENTS.md (the baseline). + await createAgent({ agentId: AGENT_ID }); + const agentsMd = path.join(agentStateDir(), "AGENTS.md"); + await fs.writeFile(agentsMd, "", "utf8"); + + // --- Turn N: evaluate the baseline --------------------------------------------------- + const baseline = await evaluate("BASELINE (blank AGENTS.md)"); + + // --- Edit: install the working discipline (version N+1) ------------------------------ + console.log("\n--- EDIT: writing a working-discipline section into the agent's AGENTS.md ---"); + await fs.writeFile(agentsMd, DISCIPLINE, "utf8"); + + // --- Turn N+1: re-evaluate the same task -------------------------------------------- + const improved = await evaluate("N+1 (with working discipline)"); + + // --- Keep-or-roll-back: the loop's decision rule ------------------------------------ + console.log("\n=== Self-improvement result ==="); + console.log(` baseline: ${baseline.toFixed(2)}/5 → N+1: ${improved.toFixed(2)}/5`); + if (improved > baseline) { + console.log(" Mean score improved — keep version N+1. ✔"); + } else { + console.log(" No improvement — roll back to the baseline AGENTS.md."); + await fs.writeFile(agentsMd, "", "utf8"); + } +} + +main().catch((err) => { + console.error("Example failed:", err); + process.exitCode = 1; +}); diff --git a/packages/landing/content/blog/local-agents-on-amd-gpus.en.md b/packages/landing/content/blog/local-agents-on-amd-gpus.en.md new file mode 100644 index 0000000..4976add --- /dev/null +++ b/packages/landing/content/blog/local-agents-on-amd-gpus.en.md @@ -0,0 +1,253 @@ +--- +title: "A closer look at PenguinHarness — and running a self-improving agent locally on an AMD GPU" +date: 2026-07-20 +category: news +excerpt: What PenguinHarness actually is, the ideas behind its architecture, and a real end-to-end run — a local open-weight model on an AMD GPU that fails a scored task, then improves itself from 0 to 4 out of 5 by editing its own files, entirely on-device. +--- + +*AMD × PrismShadow — by Ning Zhang, Yuyang Gao (AMD) and Yaowei Zheng (PrismShadow).* + +If you are new to PenguinHarness, this post is a guided tour: what the project is, the ideas that shape its architecture, and — to make it concrete — a real run where a fully local open-weight model on an AMD GPU *fails* a scored task and then *improves itself* from 0 to 4 out of 5 by editing its own files, without a single byte leaving the machine. + +## What PenguinHarness is + +PenguinHarness is an open-source AI Agent harness — a complete TypeScript stack for *building* and *evolving* agents, not a single app. It deploys fully locally, can run on as little as a single CPU, and reaches 1000+ online and local models through one unified gateway. Its purpose fits in one line: + +> Efficient Self-Improving Harness for Everyone. + +The word "harness" is deliberate. This is not a heavyweight framework you build *on top of*; it is a thin, reliable, observable substrate that an agent can *stand inside* — and, crucially, one that an agent can reach back into and improve. Three pillars carry that idea: + +| Pillar | Meaning | +| --- | --- | +| **Simplest Is the Best** | A deliberately minimal toolset over clean low-level interfaces: fewer tool calls, fewer tokens, complex tasks done efficiently. | +| **Harness for Building Agents** | Either build one programmatically with the SDK (`createAgent` → `createSession` → `run`), or have an Agent build a whole new Agent for you from a plain-language requirement. | +| **Harness for Recursive Self-Improvement** | With Skills, an Agent evaluates and optimizes *itself*, improving recursively over time. | + +For the latter two, PenguinHarness is the first open-source implementation of its kind. + +## The architecture, and why it is shaped this way + +One install gives you four layers that share one data directory and one message protocol: + +```text +┌─────────────┐ ┌─────────────────────────────┐ +│ CLI │ │ Web App (React SPA) │ +│ (penguin) │ │ ↑ OmniMessage over SSE │ +│ │ │ Server (Hono + SQLite) │ +└──────┬──────┘ └──────────────┬──────────────┘ + │ session.run(...) │ ← Human boundary +┌──────┴────────────────────────┴──────────────┐ +│ core: context_engine (ReAct loop) │ +│ ├── LLMInterface ──→ AgentHub ──→ models │ +│ ├── EnvironmentInterface ──→ builtin tools│ +│ ├── Agent State (editable files) │ +│ └── Trace (append-only JSONL) │ +└──────────────────────────────────────────────┘ +``` + +The center is the execution engine in `@prismshadow/penguin-core`. The CLI, the Server, and the Web App are simply different "Human implementations" of that same engine. This single decision — one kernel, many front-ends — is what keeps the whole system coherent, and it comes from a small set of design tenets worth understanding on their own. + +### One protocol, three jobs — OmniMessage + +Everything the system does is expressed as one message type, OmniMessage. It is simultaneously: + +- the SDK's external interface (what you send in and stream back out), +- the on-disk Trace format, and +- the engine's internal currency. + +In other words, *what streams live, what is stored on disk, and what the model sees are literally the same object*. There is no translation layer silently reshaping your data between "what happened" and "what was recorded." That identity is the foundation for the observability and recoverability everything else relies on. + +### The three-interface boundary + +The engine speaks only OmniMessage and orchestrates the flow between exactly three boundaries: + +- **Human** — the user side. Notably *not* a class: the SDK's single entry point `session.run(newMessages, { approve, signal })` *is* the Human boundary. Input is a list of messages plus an approval callback; output is a stream of messages. The CLI and the Server are its two shipped implementations. +- **LLM** — the model side (`LLMInterface`). All provider-specific protocol adaptation lives in the AgentHub gateway; the core never imports a vendor SDK. This is exactly why any OpenAI-compatible endpoint — including a local one — just works. +- **Environment** — the tool side (`EnvironmentInterface`). Runs approved tool calls and streams results back. + +Because the kernel contains no provider, tool, or UI specifics, each side swaps by configuration alone. Today's local shell can become tomorrow's sandbox; a CLI caller can become a web caller — the core never changes. + +### Agents are editable data, not code + +An Agent's entire behavior — its prompt, its Skills, its runtime config — lives as editable files on disk (`agent_state/`), not as hardcoded constants. This is the quiet key to the whole project: *what you can see, an Agent can improve*. Self-improvement is not a special engine feature; it is an agent editing the same files you would edit by hand, then re-evaluating itself. + +### A few more tenets that run through everything + +- **Errors converge into messages.** Model and tool failures never throw into the engine; they become messages the model can react to. Robustness is a property of the protocol, not of scattered try/catch. +- **Everything is observable.** Every request, tool call, and approval decision is appended to the Trace; a Session restores fully from it. +- **Streaming first.** Text streams token by token; tool calls and results appear live. +- **Model ↔ Agent decoupling.** An Agent never binds to a model — you choose one per Session. The same Agent can run different sessions on different models. + +The one-line summary of the layering: *what is editable or recorded lives in files; what makes messages flow lives in the SDK; what needs a resident process and multiple users lives in the Server; the rest is rendering.* + +## Simplest is the best — six tools, and why that matters + +The first pillar is the easiest to overlook and the one you feel most in practice: the toolset is deliberately tiny. PenguinHarness ships exactly six built-in tools: + +| Tool | Purpose | +| --- | --- | +| `exec_command` | Run a shell command in the workspace (streams stdout/stderr) | +| `input_command` | Drive a running command — write stdin, send Ctrl-C, poll output | +| `run_subagent` | Delegate a self-contained subtask to a child agent | +| `input_subagent` | Poll or follow up with a background subagent | +| `read_image` | Read an image as image content (vision models) | +| `describe_image` | Have a vision model describe an image for text-only models | + +Notice what is *not* there: no `read_file`, no `write_file`, no `edit_file`, no `list_dir`, no `grep` tool. That is intentional — the shell is the universal interface, so reading, writing, and editing files all go through `exec_command` (`cat`, `>`, `sed`, and so on). The principle is "the simplest is the best": every extra tool is more schema in the prompt, more tokens on every call, and one more thing the model can pick wrong. Fewer tools means fewer wrong calls and less token overhead — complex tasks done with less ceremony. + +This is not just theory — you can read it straight out of the Trace. Here are the only three tool calls the agent made to complete the entire CSV-cleanup case, all of them `exec_command`: + +```bash +# 1. read the input — no read_file tool, just cat +cat users.csv + +# 2. do the work — the shell lets the model reach for Python +python3 -c " +import csv +rows = list(csv.DictReader(open('users.csv', newline=''))) +cleaned = [r for r in rows if (r.__setitem__('email', r['email'].strip().lower()) or r['email'])] +seen, out = set(), [] +for r in cleaned: + key = tuple(r.values()) + if key not in seen: seen.add(key); out.append(r) +# ... write users_clean.csv, columns unchanged ... +" + +# 3. verify by reading the result back — again just cat +cat users_clean.csv +``` + +Read the file, transform it, check the result — three calls, one tool, no special file machinery. And notice the second step: because the interface is a shell, the model naturally reached for Python to express the dedup logic, something no fixed `edit_file` tool could have done. A minimal toolset is not a limitation the model works around; it is *why* a capable model can finish a real task in so few steps. + +## Agents that build agents — a worked example + +This is the second pillar — *the harness for building agents* — made concrete. It has two faces. The first is the SDK: you embed an agent in your own program with a few lines — `createAgent()` → `createSession()` → `session.run(...)` (see the snippet at the end). The second is the more striking one, and the flip side of "agents are editable data": if an Agent is just files, then *an Agent can write those files for you*. That is what the built-in `agent-creation` skill does — given a plain-language requirement, an agent scaffolds a new agent: its directory layout, its `system_config.yaml` (name and description), and above all its `AGENTS.md`, the file that turns the requirement into behavior. To show that second face end-to-end, we did exactly this on the local AMD GPU stack. + +**The request.** We asked the local `qwen3.6:35b` agent, using the `agent-creation` skill: + +> Create a new agent called `commit-helper` that writes Conventional Commits messages — a `type(scope): subject` header (type from feat/fix/docs/…), imperative subject under ~50 chars, a blank line, then a body explaining the *why*. + +**What it produced.** Working on its own, the agent created the new agent's directory, copied a base config, set its name and description, and wrote a genuinely good `AGENTS.md` — encoding the header format, a type enum, a subject-length rule, "explain the *why*, not the *what*" for the body, an optional `BREAKING CHANGE`/`Closes #` footer, and even a heuristic for inferring the type from a diff (e.g. "renames → refactor, not chore"). No hand-holding on the content. + +**Then we ran the agent it built.** Pointed at a change description ("added retry-with-backoff to the payment client because transient 503s broke checkout"), the freshly-created `commit-helper` — following only the `AGENTS.md` written for it — produced: + +```text +fix(payment): add retry-with-backoff for transient gateway 503 errors + +Transient 503 responses from the payment gateway were causing +checkout failures for users during peak traffic. Retry with +exponential backoff gives the gateway time to recover, preventing +spurious user-facing errors without requiring manual retries. +``` + +It even reasoned out loud about whether the change was a `fix` or a `feat` before committing to `fix` — behavior that came entirely from the AGENTS.md its parent agent had written. + +**Run it yourself.** The whole flow above is a self-contained, SDK-driven script in the repo: +[`examples/build-agent-with-agent/`](https://github.com/Prism-Shadow/penguin-harness/tree/main/examples/build-agent-with-agent). +Phase 1 uses `createAgent`/`createSession`/`run` to drive `default_agent` into building +`commit-helper`; phase 2 loads the new agent and runs it — all on the local Ollama + qwen3.6:35b +setup, no cloud API. See its `README.md` for the one-time Ollama configuration. + +## Self-improvement, in one line + +The third pillar builds on the same idea. Because agents are editable data and everything is traced, an agent can *measure itself and get better* — a loop of benchmark → evaluate → find where points were lost → edit the agent's own files → keep the change only if the score improves. There is no special engine code behind it; it is ordinary agent machinery orchestrated by Skills, and every number on the scoreboard links back to the exact Session that produced it. We walk through the full loop — with a real before/after — after the run below. + +## A real local run on an AMD GPU + +Design tenets are easy to claim. Here is an end-to-end run that exercises them, entirely on fully local, open-weight infrastructure with an AMD GPU — every token generated on-device, nothing sent to a cloud API. + +**The setup.** An AMD GPU running ROCm, serving `qwen3:8b` — a popular, everyday local open-weight model (about 5 GB) — through Ollama's OpenAI-compatible endpoint. Ollama detected the AMD GPU natively (no architecture override needed) and loaded the model into GPU memory. This same path spans AMD's ROCm-supported lineup — from Radeon PRO workstation cards such as the W7900 (48 GB, RDNA3) up to datacenter Instinct accelerators. Wiring it into PenguinHarness took a single command, precisely because the core treats any OpenAI-compatible endpoint the same way: + +```bash +penguin config model add \ + --model-id qwen3:8b \ + --provider custom --client-type openai \ + --base-url http://localhost:11434/v1 \ + --api-key ollama --set-default +``` + +**The task, and how it's scored.** We gave it a constrained writing task in the style of the built-in `example-benchmark`: read a project notes file and produce a summary with an overview of at most 2 sentences, exactly 3 key facts, under 60 words total. The task carries a private rubric — a scoring checklist the agent never sees, so it cannot "teach to the test." The rubric turns each requirement into one point (file actually written · ≤2-sentence overview · exactly 3 bullets · under 60 words · facts accurate), for a score out of 5. + +**The baseline — and why it isn't perfect.** Running locally on the AMD GPU, the model read the file, composed a perfectly reasonable summary… and then never wrote it to disk. It emitted the answer as chat text instead of calling the write tool. No deliverable means the rubric scores it **0 / 5**. And because every step is in the Trace, this is not guesswork — you can open the run and see exactly where it went wrong. + +That is the honest starting point: on local hardware, an out-of-the-box model does *not* ace a constrained task. Which is exactly what makes the next section interesting — a measured, auditable failure is something you can systematically fix. + +## How self-improvement actually works — and a real before/after + +The 0 / 5 above is an *evaluation* — a snapshot of the agent as it stands. The more interesting question is how PenguinHarness turns that snapshot into progress. This is the recursive self-improvement loop, and it is worth understanding as a mechanism, because there is no magic engine code behind it — it is ordinary agent machinery, orchestrated by Skills: + +1. **Benchmark** — define capability cases, each with a private rubric (as above). +2. **Evaluate** — run the agent over the cases and score against the rubrics. Each run is an ordinary, fully-traced Session. +3. **Read the Trace to find where points were lost** — because every score links back to the exact run, you can see *why* a point was missed, not just that it was. +4. **Edit the Agent's state** — the agent's behavior lives in editable files (`AGENTS.md`, Skills, config). You (or an Optimizer agent) change those to address the failure, producing version N+1. +5. **Snapshot & keep-or-roll-back** — snapshot before each round; keep N+1 only if the score strictly improves, otherwise roll back. + +Let's improve the exact agent from the previous section — the same `qwen3:8b` on the same AMD GPU, the same task that just scored **0 / 5**. + +- **The diagnosis (from the Trace).** The run failed for one concrete reason: the model produced the summary as chat text and never called the tool to write `summary.md`. The Trace shows it — a `cat` to read the input, then prose, and no write. So the fix is not "make the model smarter"; it is "make the behavior reliable." +- **The edit (N → N+1).** We added a short *working-discipline* section to the agent's `AGENTS.md`: *read the source first, restate the task's constraints as a checklist, actually call the tool to write the file, and re-read the output to self-check before finishing.* That is the entire change — a few sentences in an editable text file the agent reads on every run. No retraining, no code. +- **The re-evaluation.** Same model, same task, run again: this time it wrote a correct, on-topic `summary.md` to disk, with a 2-sentence overview and exactly 3 key facts drawn from the source — **4 / 5** (the one remaining miss: it ran to 70 words, over the 60-word budget). + +That is a jump on the same model and the same task, achieved purely by editing a text file the agent reads. That is the loop in miniature: *what you can see, you (or an Optimizer agent) can improve* — and because the score comes from a rubric and links to a Trace, the improvement is measured, not vibes. This is exactly the "keep the change only if the score strictly improves" loop, one turn of the crank. + +**Run it yourself.** This whole loop is a self-contained, SDK-driven script in the repo: +[`examples/self-improving-agent/`](https://github.com/Prism-Shadow/penguin-harness/tree/main/examples/self-improving-agent). +It uses a deterministic, readable rubric (file written · ≤2 sentences · exactly 3 bullets · under +60 words · facts accurate) and — because a local model is nondeterministic — averages several runs +per version, which is exactly why real benchmarks use a `runs` count. In our runs the averaged +score moved from **0.00/5** (blank `AGENTS.md`) to **5.00/5** (with the working-discipline edit), +and the script keeps N+1 only because the score strictly improved. It runs on the same local +Ollama + qwen3:8b setup and uses a dedicated agent id, so your own agents are untouched. + +## No AMD GPU yet? Free cloud compute from AMD + Fireworks + +Not everyone has an AMD GPU under their desk — and you do not need one to try this. Through the AMD AI Developer Program, AMD partners with Fireworks AI to hand eligible developers **$50 in free Fireworks credits**. Fireworks serves open-weight models over an OpenAI-compatible endpoint, so — just like the local Ollama setup above — it is a single-line change to point PenguinHarness at it. + +Getting the credits (approval typically takes 2–3 business days): + +1. Sign up at the [AMD AI Developer Program](https://developer.amd.com/ai-developer-program/). +2. Open **Member Perks → Cloud Credit Options → Request Cloud Credits**. +3. In the form, choose **Fireworks AI** as the product needed, add at least one public profile link (GitHub, LinkedIn, etc.), and submit. +4. AMD emails you a coupon code. Redeem it at [fireworks.ai](https://fireworks.ai/) via **Redeem Promo**, then generate an API key. + +Then wire it into PenguinHarness exactly like any other endpoint: + +```bash +penguin config model add --model-id \ + --provider custom --client-type openai \ + --base-url https://api.fireworks.ai/inference/v1 \ + --api-key --set-default +``` + +Same harness, same one-line swap — whether the tokens are generated on your own AMD GPU or on AMD-backed cloud credits. (Keep your coupon code and key private; program terms may change, so check the official page and approval email for current details.) + +## Why this matters + +- **Local-first is not a slogan.** A complete build → run → self-evaluate loop ran on-device, on an AMD GPU, with no data leaving the machine — a real answer for privacy-sensitive and enterprise settings. And because it rides on ROCm + Ollama, the same setup runs across AMD's GPU range: a single Radeon PRO workstation card (e.g. the 48 GB W7900) comfortably runs models from 8B up to 30B-plus, while Instinct accelerators scale it further. +- **The thin model layer pays off.** Because provider adaptation lives entirely in the gateway, "a local Ollama model" and "a frontier cloud API" are the same one-line change. You are never locked to a vendor — or to a GPU vendor. We ran this on an AMD GPU (ROCm), but nothing here is AMD-specific: the very same steps work on an NVIDIA GPU (Ollama's CUDA backend) or on Apple Silicon — only the Ollama runtime underneath changes, while the harness, the commands, and the examples stay identical. +- **Observability is built in, everywhere.** The local run produced the same append-only Trace and scoreboard linkage as any cloud run. Evaluation is auditable by construction. + +## Get started + +```bash +curl -fsSL https://github.com/Prism-Shadow/penguin-harness/releases/latest/download/install.sh | sh + +# Point at any OpenAI-compatible endpoint — a local Ollama model included +penguin config model add --model-id \ + --provider custom --client-type openai \ + --base-url http://localhost:11434/v1 --api-key ollama --set-default + +penguin web # or: penguin run --approve allow-all --message "..." +``` + +Prefer to embed it in your own program? That is the SDK face of the "building agents" pillar — the core loop is three calls: + +```ts +const agent = await createAgent({ agentId: "default_agent" }); +const session = await agent.createSession({ workspaceDir: process.cwd() }); +for await (const out of session.run([userText("...")], { approve: async () => "allow" })) { /* stream */ } +``` + +Complete, runnable versions live in [`examples/`](https://github.com/Prism-Shadow/penguin-harness/tree/main/examples) — including an agent that builds another agent and an agent that improves itself, both on local Ollama. + +Whether you run frontier cloud models or open weights on your own AMD GPU, PenguinHarness gives you the same minimal, observable, self-improving substrate. Follow us on [GitHub](https://github.com/Prism-Shadow/penguin-harness) and open your first issue. diff --git a/packages/landing/content/blog/local-agents-on-amd-gpus.zh.md b/packages/landing/content/blog/local-agents-on-amd-gpus.zh.md new file mode 100644 index 0000000..22eef79 --- /dev/null +++ b/packages/landing/content/blog/local-agents-on-amd-gpus.zh.md @@ -0,0 +1,246 @@ +--- +title: 深入了解 PenguinHarness——并在 AMD GPU 上本地跑通一个自我进化的 Agent +date: 2026-07-20 +category: news +excerpt: PenguinHarness 到底是什么、它的架构背后有哪些理念,以及一次真实的端到端运行——一个完全本地、运行在 AMD GPU 上的开源权重模型,先在一个打分任务上失败,再通过编辑自己的文件把分数从 0 提升到 4(满分 5),全程数据不出本机。 +--- + +*AMD × PrismShadow——张宁、高钰洋(AMD),郑耀威(PrismShadow)。* + +如果你是第一次接触 PenguinHarness,这篇文章是一次完整的导览:这个项目是什么、它的架构背后有哪些设计理念,以及——为了让一切足够具体——一次真实的运行:一个完全本地、跑在 AMD GPU 上的开源权重模型,先在一个打分任务上*失败*,再通过编辑自己的文件把分数从 0 *自我提升*到 4(满分 5),全程没有一个字节离开本机。 + +## PenguinHarness 是什么 + +PenguinHarness 是一个开源的 AI Agent Harness——一套用于*构建*和*进化* Agent 的完整 TypeScript 技术栈,而不是某一个具体应用。它可以完全本地部署,最低单 CPU 即可运行,并通过一个统一网关触达 1000+ 在线与本地模型。它的主旨只有一句话: + +> Efficient Self-Improving Harness for Everyone.(为每个人打造的、高效的自我进化 Harness。) + +"Harness"(挽具 / 骨架)这个词是刻意选择的。它不是一个让你*在其之上*层层搭建的重型框架,而是一层轻薄、可靠、可观测的底座——Agent 可以*站在其中*运行,更关键的是,Agent 可以反过来触及并改进它自己。三大支柱承载了这一理念: + +| 支柱 | 含义 | +| --- | --- | +| **Simplest Is the Best(大道至简)** | 在干净的底层接口之上,提供刻意精简的工具集:更少的工具调用、更少的 Token、复杂任务高效完成。 | +| **构建 Agent 的 Harness** | 既可用 SDK 编程构建(`createAgent` → `createSession` → `run`),也可由一个 Agent 根据一句需求描述,自动造出一个全新 Agent。 | +| **递归自我进化的 Harness** | 借助 Skills,一个 Agent 能评估并优化*它自己*,随时间递归改进。 | + +后两点,PenguinHarness 是业界首个开源实现。 + +## 架构,以及它为何长成这样 + +一次安装即可获得共享同一个数据目录和同一套消息协议的四层: + +```text +┌─────────────┐ ┌─────────────────────────────┐ +│ CLI │ │ Web App (React SPA) │ +│ (penguin) │ │ ↑ OmniMessage over SSE │ +│ │ │ Server (Hono + SQLite) │ +└──────┬──────┘ └──────────────┬──────────────┘ + │ session.run(...) │ ← Human 边界 +┌──────┴────────────────────────┴──────────────┐ +│ core: context_engine(ReAct 循环) │ +│ ├── LLMInterface ──→ AgentHub ──→ 模型 │ +│ ├── EnvironmentInterface ──→ 内置工具 │ +│ ├── Agent State(可编辑文件) │ +│ └── Trace(只追加 JSONL) │ +└──────────────────────────────────────────────┘ +``` + +系统的中心是 `@prismshadow/penguin-core` 里的执行引擎。CLI、Server、Web App 只是同一个引擎的不同"Human 实现"。这一个决定——单一内核、多前端——正是让整个系统保持自洽的关键,而它源自一组值得单独理解的设计信条。 + +### 一套协议,三重身份——OmniMessage + +系统所做的一切都用同一种消息类型 OmniMessage 表达。它同时是: + +- SDK 的对外接口(你输入的、以及流式返回的), +- 磁盘上的 Trace 格式, +- 引擎内部流转的"通货"。 + +换句话说,*实时流式的内容、落盘存储的内容、模型看到的内容,是字面意义上的同一个对象*。中间没有一个隐形的转换层在"实际发生了什么"和"被记录了什么"之间悄悄改写数据。正是这种"三位一体",构成了其余一切可观测性与可恢复性的根基。 + +### 三接口边界 + +引擎只讲 OmniMessage,并在恰好三个边界之间编排信息流: + +- **Human**——用户侧。值得注意的是它*不是*一个类:SDK 唯一的入口 `session.run(newMessages, { approve, signal })` *就是* Human 边界。输入是一组新消息加一个审批回调,输出是一串流式消息。CLI 和 Server 是它的两个官方实现。 +- **LLM**——模型侧(`LLMInterface`)。所有与厂商相关的协议适配都活在 AgentHub 网关里;core 从不导入任何厂商 SDK。这正是为什么任何 OpenAI 兼容端点——包括本地端点——都能开箱即用。 +- **Environment**——工具侧(`EnvironmentInterface`)。执行已审批的工具调用并把结果流式返回。 + +因为内核不含任何 provider、工具或 UI 的具体细节,每一侧都仅靠配置即可替换。今天的本地 shell 可以变成明天的沙箱;一个 CLI 调用方可以换成 Web 调用方——core 始终不变。 + +### Agent 是可编辑的数据,不是代码 + +一个 Agent 的全部行为——它的提示词、Skills、运行参数——都以磁盘上的可编辑文件形式存在(`agent_state/`),而非硬编码常量。这是整个项目安静而关键的一点:*你能看到的,Agent 就能改进*。自我进化不是引擎的某个特殊功能,而是一个 Agent 去编辑那些你本可以手动编辑的同一批文件,然后重新评估自己。 + +### 贯穿始终的其余信条 + +- **错误收敛为消息。** 模型与工具的失败从不向引擎抛异常;它们会变成模型可以据此反应的消息。健壮性是协议的属性,而非散落各处的 try/catch。 +- **一切皆可观测。** 每一次请求、工具调用、审批决策都会追加进 Trace;一个 Session 可完整从中恢复。 +- **流式优先。** 文本逐 Token 流出;工具调用与结果实时呈现。 +- **模型与 Agent 解耦。** Agent 从不绑定模型——在创建每个 Session 时选择。同一个 Agent 可以在不同 Session 上跑不同模型。 + +分层的一句话总结:*可编辑或需记录的活在文件里;让消息流动的活在 SDK 里;需要常驻进程与多用户的活在 Server 里;其余的都是渲染。* + +## 大道至简——只有 6 个工具,以及这为何重要 + +第一根支柱最容易被忽略,却是你在实际使用中感受最深的:工具集刻意做得极小。PenguinHarness 只内置恰好 6 个工具: + +| 工具 | 用途 | +| --- | --- | +| `exec_command` | 在工作区里执行 shell 命令(流式返回 stdout/stderr) | +| `input_command` | 驱动一个运行中的命令——写 stdin、发 Ctrl-C、轮询输出 | +| `run_subagent` | 把一个自包含的子任务委派给子 Agent | +| `input_subagent` | 轮询后台子 Agent,或在其空闲后追加提示 | +| `read_image` | 把图片作为图像内容读入(视觉模型) | +| `describe_image` | 让视觉模型把图片描述成文本(供纯文本模型使用) | + +注意这里*没有*什么:没有 `read_file`、没有 `write_file`、没有 `edit_file`、没有 `list_dir`、没有 `grep` 工具。这是刻意的——shell 就是通用接口,所以文件的读、写、改全部通过 `exec_command` 走(`cat`、`>`、`sed` 等等)。它遵循的原则就是"大道至简":每多一个工具,就是提示词里多一段 schema、每次调用多一些 token、模型多一个可能选错的选项。工具越少,选错越少、token 开销越小——复杂任务反而做得更利落。 + +这不只是理论——你可以直接从 Trace 里读出来。下面就是那个 CSV 清洗 Case,Agent 完成整个任务所做的全部 3 次工具调用,全都是 `exec_command`: + +```bash +# 1. 读输入——没有 read_file 工具,直接 cat +cat users.csv + +# 2. 干活——shell 让模型自然地用上 Python +python3 -c " +import csv +rows = list(csv.DictReader(open('users.csv', newline=''))) +cleaned = [r for r in rows if (r.__setitem__('email', r['email'].strip().lower()) or r['email'])] +seen, out = set(), [] +for r in cleaned: + key = tuple(r.values()) + if key not in seen: seen.add(key); out.append(r) +# ... 写出 users_clean.csv,保持列序不变 ... +" + +# 3. 读回结果做自检——同样只是 cat +cat users_clean.csv +``` + +读文件、做转换、核对结果——3 次调用、1 个工具、没有任何专用文件机制。再留意第二步:正因为接口是一个 shell,模型很自然地用起了 Python 来表达去重逻辑——这是任何固定的 `edit_file` 工具都做不到的。精简的工具集不是模型需要绕过的限制;它恰恰*正是*一个有能力的模型能用极少步骤完成真实任务的原因。 + +## 用 Agent 造 Agent——一个完整的例子 + +这就是第二根支柱——构建 Agent 的 Harness——的具体呈现。它有两面。第一面是 SDK:你用几行代码就能把一个 Agent 嵌进自己的程序——`createAgent()` → `createSession()` → `session.run(...)`(文末有代码片段)。第二面更惊艳,也正是"Agent 是可编辑的数据"的另一面:如果一个 Agent 不过是一堆文件,那么*一个 Agent 就能替你把这些文件写出来*。这正是内置 `agent-creation` skill 所做的事——给它一句大白话需求,一个 Agent 就能搭建出一个全新的 Agent:它的目录布局、它的 `system_config.yaml`(名称与描述),以及最关键的 `AGENTS.md`——那个把需求变成行为的文件。为了把这第二面端到端地展示出来,我们就在本地 AMD GPU 这套技术栈上真跑了一次。 + +**需求。** 我们让本地的 `qwen3.6:35b`,通过 `agent-creation` skill: + +> 创建一个叫 `commit-helper` 的新 Agent,专门写 Conventional Commits 提交信息——`type(scope): subject` 的标题(type 取自 feat/fix/docs/…)、祈使句、subject 控制在约 50 字符内,空一行,然后是一段解释"为什么"的正文。 + +**它产出了什么。** 这个 Agent 自主完成了:创建新 Agent 的目录、复制一份 base config、设置名称与描述,并写出了一份质量相当高的 `AGENTS.md`——涵盖了标题格式、type 枚举、subject 长度规则、正文"解释 why 而非 what"、可选的 `BREAKING CHANGE`/`Closes #` 脚注,甚至还有一条"从 diff 推断 type"的启发式(例如"重命名 → refactor 而非 chore")。内容层面完全没有需要人工提点。 + +**然后我们运行了它造出来的这个 Agent。** 给刚创建的 `commit-helper` 一段改动描述("给支付客户端加了带退避的重试,因为网关偶发的 503 导致下单失败"),它——只依据为它写的那份 `AGENTS.md`——产出了: + +```text +fix(payment): add retry-with-backoff for transient gateway 503 errors + +Transient 503 responses from the payment gateway were causing +checkout failures for users during peak traffic. Retry with +exponential backoff gives the gateway time to recover, preventing +spurious user-facing errors without requiring manual retries. +``` + +它甚至先"出声"权衡了这个改动到底算 `fix` 还是 `feat`,才最终定为 `fix`——这个行为完全来自它的父 Agent 为它写的 AGENTS.md。 + +**你可以自己跑一遍。** 上面整个流程在仓库里有一个自包含、纯 SDK 驱动的脚本: +[`examples/build-agent-with-agent/`](https://github.com/Prism-Shadow/penguin-harness/tree/main/examples/build-agent-with-agent)。 +第一阶段用 `createAgent`/`createSession`/`run` 驱动 `default_agent` 造出 `commit-helper`;第二阶段加载这个新 Agent 并运行它——全部跑在本地 Ollama + qwen3.6:35b 上,不用任何云端 API。一次性的 Ollama 配置见其 `README.md`。 + +## 自我进化,一句话概括 + +第三根支柱建立在同一个理念之上。因为 Agent 是可编辑的数据、且一切皆被追踪,一个 Agent 可以*度量自己并变得更好*——这是一个"定基准 → 评估 → 找到失分点 → 编辑 Agent 自己的文件 → 只有分数提升才保留"的循环。它背后没有任何专用引擎代码,就是由 Skills 编排的普通 Agent 机制,而 scoreboard 上的每一个数字都能追回到产生它的那次具体 Session。我们会在下面这次运行之后,用一次真实的前后对比,把整个循环走一遍。 + +## 在 AMD GPU 上的一次真实本地运行 + +设计信条说起来容易。下面是一次真正践行了这些信条的端到端运行,全程在完全本地、开源权重、AMD GPU 的环境上进行——每一个 Token 都在本机生成,不向任何云端 API 发送。 + +**环境。** 一块 AMD GPU 运行 ROCm,通过 Ollama 的 OpenAI 兼容端点提供 `qwen3:8b`——一个开发者常用的本地开源权重模型(约 5 GB)。Ollama 原生识别到该 AMD GPU(无需任何架构 override),并把模型加载进显存。这条路径覆盖 AMD 受 ROCm 支持的整个产品线——从 Radeon PRO 工作站显卡(如 W7900,48 GB,RDNA3)一直到数据中心的 Instinct 加速卡。把它接入 PenguinHarness 只需一条命令,正因为 core 对任何 OpenAI 兼容端点一视同仁: + +```bash +penguin config model add \ + --model-id qwen3:8b \ + --provider custom --client-type openai \ + --base-url http://localhost:11434/v1 \ + --api-key ollama --set-default +``` + +**任务,以及如何评分。** 我们给它一个仿照内置 `example-benchmark` 风格的约束型写作任务:读一个项目 notes 文件,写一份摘要——概述至多 2 句、恰好 3 条要点、全文不超过 60 词。任务附带一份私有 rubric——一份 Agent 永远看不到的评分清单,这样它就无法"照着答案作答"。rubric 把每条要求变成一分(文件确实写出 · 概述 ≤2 句 · 恰好 3 条要点 · 不超过 60 词 · 事实准确),满分 5 分。 + +**基线——以及它为何不是满分。** 跑在本地 AMD GPU 上,这个模型读了文件、组织出了一份相当合理的摘要……然后却从没把它写到磁盘。它把答案当成对话文本吐了出来,而没有调用写文件工具。没有交付物 → rubric 判它 **0 / 5**。而因为每一步都在 Trace 里,这不是靠猜——你可以打开这次运行,精确看到它错在哪。 + +这就是诚实的起点:在本地硬件上,一个开箱即用的模型面对约束型任务并不能一次拿满分。而这恰恰让下一节变得有意思——一个可度量、可审计的失败,是可以被系统性修复的。 + +## 自我进化到底是怎么运作的——附一次真实的前后对比 + +上面那个 0 / 5 是一次*评估*——是这个 Agent 当前状态的一张快照。更有意思的问题是:PenguinHarness 如何把这张快照变成进步。这就是递归自我进化循环,值得作为一种"机制"来理解,因为它背后没有任何魔法引擎代码——它就是由 Skills 编排的普通 Agent 机制: + +1. **Benchmark(定基准)**——定义能力 Case,每个都配一份私有 rubric(如上)。 +2. **Evaluate(评估)**——让 Agent 跑这些 Case 并按 rubric 打分。每一次运行都是一个普通的、被完整追踪的 Session。 +3. **读 Trace 定位失分点**——因为每个分数都能追回到那次具体运行,你能看到某一分*为什么*被扣,而不只是"被扣了"。 +4. **编辑 Agent 的状态**——Agent 的行为活在可编辑文件里(`AGENTS.md`、Skills、config)。你(或一个 Optimizer Agent)针对失分点去改这些文件,产出版本 N+1。 +5. **快照 & 保留或回滚**——每轮前先打快照;只有当分数*严格提升*时才保留 N+1,否则回滚。 + +我们就来改进上一节里那个 Agent——同一个 `qwen3:8b`、同一块 AMD GPU、同一个刚拿了 0 / 5 的任务。 + +- **诊断(来自 Trace)。** 这次失败只有一个具体原因:模型把摘要当对话文本产出,却从没调用工具去写 `summary.md`。Trace 把它摆在眼前——一次 `cat` 读入、然后是正文、没有写入。所以要修的不是"让模型更聪明",而是"让行为更可靠"。 +- **编辑(N → N+1)。** 我们往这个 Agent 的 `AGENTS.md` 里加了一小段*任务纪律*:*先读原文、把任务约束逐条列成清单、真正调用工具把文件写出来、结束前重新读一遍输出做自检。* 改动仅此而已——几句话,写进一份 Agent 每次运行都会读的可编辑文本文件。没有重新训练、没有改代码。 +- **重新评估。** 同一个模型、同一个任务,再跑一次:这次它把一份准确、切题的 `summary.md` 写到了磁盘,有一段 2 句的概述、恰好 3 条取自原文的要点——**4 / 5**(唯一还差的一分:它写到了 70 词,超出 60 词的预算)。 + +在同一个模型、同一个任务上,仅仅通过编辑一份 Agent 会读取的文本文件,分数就提升了。这就是循环的微缩版:*你能看到的,你(或一个 Optimizer Agent)就能改进*——而且因为分数来自一份 rubric、又关联着 Trace,这份改进是被度量出来的,不是拍脑袋。这正是"只有分数严格提升才保留"的循环,转动了一格。 + +**你可以自己跑一遍。** 整个循环在仓库里有一个自包含、纯 SDK 驱动的脚本: +[`examples/self-improving-agent/`](https://github.com/Prism-Shadow/penguin-harness/tree/main/examples/self-improving-agent)。 +它用一份确定性的、可读的 rubric(文件是否写出 · 概述 ≤2 句 · 恰好 3 条要点 · 不超过 60 词 · 事实准确),并且——因为本地模型有随机性——对每个版本跑多次取平均分,这正是真实 benchmark 里 `runs`(多次运行)的意义所在。在我们的运行中,平均分从 **0.00/5**(空 `AGENTS.md`)提升到 **5.00/5**(加入"任务纪律"后),脚本仅因分数严格提升才保留 N+1。它跑在同一套本地 Ollama + qwen3:8b 环境上,并使用一个专用 agent id,所以你自己的 Agent 不会被动到。 + +## 还没有 AMD GPU?来自 AMD + Fireworks 的免费云端算力 + +不是每个人桌下都有一块 AMD GPU——而要试用这一切,你并不需要它。通过 AMD AI Developer Program,AMD 与 Fireworks AI 合作,向符合条件的开发者提供**价值 50 美元的免费 Fireworks 额度**。Fireworks 通过 OpenAI 兼容端点提供开源权重模型,因此——和上面的本地 Ollama 一样——把 PenguinHarness 指向它也只是一行配置的事。 + +领取额度(审核通常需要 2–3 个工作日): + +1. 在 [AMD AI Developer Program](https://developer.amd.com/ai-developer-program/) 注册。 +2. 进入 **Member Perks → Cloud Credit Options → Request Cloud Credits**。 +3. 在表单中把"所需产品"选为 **Fireworks AI**,附上至少一个公开主页链接(GitHub、LinkedIn 等),提交。 +4. AMD 会把优惠码发到你的邮箱。到 [fireworks.ai](https://fireworks.ai/) 通过 **Redeem Promo** 兑换,然后生成 API Key。 + +之后像接入任何其他端点一样把它接进 PenguinHarness: + +```bash +penguin config model add --model-id \ + --provider custom --client-type openai \ + --base-url https://api.fireworks.ai/inference/v1 \ + --api-key <你的-fireworks-key> --set-default +``` + +同一套 Harness、同样的一行切换——无论 Token 是在你自己的 AMD GPU 上生成,还是用 AMD 支持的云端额度生成。(请妥善保管优惠码与 Key;项目条款可能变动,具体以官方页面和审核邮件为准。) + +## 这为什么重要 + +- **本地优先不是口号。** 一个完整的"构建 → 运行 → 自我评估"闭环在本机、在一块 AMD GPU 上跑通,没有数据离开机器——这是对隐私敏感与企业场景的真实回答。而且因为它跑在 ROCm + Ollama 之上,同一套配置可覆盖 AMD 的整个 GPU 阵容:单块 Radeon PRO 工作站显卡(如 48 GB 的 W7900)从 8B 到 30B 以上的模型都能从容运行,而 Instinct 加速卡则可进一步向上扩展。 +- **薄模型层带来实际收益。** 因为 provider 适配完全活在网关里,"一个本地 Ollama 模型"和"一个前沿云端 API"只是同一处一行配置的差别。你不会被某个厂商锁定——也不会被某个 GPU 厂商锁定。我们这次跑在 AMD GPU(ROCm)上,但这里没有任何 AMD 专属的东西:同样的步骤在 NVIDIA GPU(Ollama 的 CUDA 后端)或 Apple Silicon 上一样成立——底层变的只是 Ollama 运行时,而 Harness、命令、example 全都保持不变。 +- **可观测性内建于每一处。** 这次本地运行产生了与任何云端运行相同的只追加 Trace 与 scoreboard 关联。评估在设计上就是可审计的。 + +## 立即开始 + +```bash +curl -fsSL https://github.com/Prism-Shadow/penguin-harness/releases/latest/download/install.sh | sh + +# 指向任意 OpenAI 兼容端点——包括一个本地 Ollama 模型 +penguin config model add --model-id <你的模型> \ + --provider custom --client-type openai \ + --base-url http://localhost:11434/v1 --api-key ollama --set-default + +penguin web # 或:penguin run --approve allow-all --message "..." +``` + +想把它嵌进自己的程序?这就是"构建 Agent"支柱的 SDK 那一面——核心循环就是三次调用: + +```ts +const agent = await createAgent({ agentId: "default_agent" }); +const session = await agent.createSession({ workspaceDir: process.cwd() }); +for await (const out of session.run([userText("...")], { approve: async () => "allow" })) { /* 流式处理 */ } +``` + +完整、可运行的版本见 [`examples/`](https://github.com/Prism-Shadow/penguin-harness/tree/main/examples)——包括一个"用 Agent 造 Agent"和一个"自我改进的 Agent",都跑在本地 Ollama 上。 + +无论你跑的是前沿云端模型,还是自己 AMD GPU 上的开源权重模型,PenguinHarness 都为你提供同一套精简、可观测、可自我进化的底座。在 [GitHub](https://github.com/Prism-Shadow/penguin-harness) 关注我们,并提交你的第一个 issue。