Branch-length batch covering tooling, the model layer, the Web App and the public surfaces. Highlights: - Changelog: a per-release `changelog/<version>/` tree, grouped by the surface each change touches, with a root CHANGELOG.md holding one line per release. - Dev startup: `scripts/dev-prebuild.mjs` serializes the skills+core prebuild behind a lock and keeps `pnpm install` current; `pnpm dev` runs server+web together. - AgentHub 0.3.3 -> 0.4.0: OmniMessage complete payloads carry one opaque `fidelity` object in place of item-level `signature`/`phase`, threaded verbatim through Trace, replay and resume; malformed classification adapted to the new error types. - Model layer: a model is always referenced by an explicit `(provider, model_id)` pair. The provider is never inferred, guessed or defaulted -- both the catalog inference and the unique-match config resolution are gone, and CLI, SDK, server routes and run_subagent all require the complete pair. Catalog gains the Qwen Token Plan, Qwen Pay-As-You-Go and Fireworks AI gateways, plus an expanded OpenRouter group. - Web App: catalog preset sync and per-group speed test on the Models page, positional slash commands, a markdown renderer, skill-library update reminders, and a vertically centred draft page whose upward menus size themselves to the room available. - Public surfaces: restructured READMEs, the penguin.ooo landing site and blog, refreshed benchmark results for both suites, and the demo videos playing on the landing page. Includes the fixes from a full review of the branch: 23 confirmed findings, among them a provider-inference bug that could send one vendor's API key to another vendor's endpoint, and an Escape handler that destroyed the composer's contents unrecoverably. Verified on the branch head: pnpm test (1127 passing, 7 packages), pnpm typecheck and pnpm format:check clean, Playwright e2e 14/14.
7.5 KiB
title, date, category, excerpt
| title | date | category | excerpt |
|---|---|---|---|
| Introducing PenguinHarness: agents that build agents | 2026-07-17 | news | We proved agents can self-evolve in our GDPevo Benchmark — now we are bringing that capability to everyone. The first open-source harness with recursive self-improvement covers everything from one-sentence agent construction to continuous self-evolution. |
Today we are releasing PenguinHarness — an open-source harness built for constructing and evolving agents: a zero-code Harness CLI and Web UI, connected to 1000+ models. The story it tells fits in one line:
With LangChain, you build agents by hand — at 1× speed. With PenguinHarness, agents build agents — at 100×.
From GDPevo to PenguinHarness: why we built this
Before PenguinHarness, our team published the GDPevo Benchmark. In GDPevo we systematically verified one thing: agents can self-evolve — an Agent can score its own performance, find where the points were lost, rewrite its own prompts and Skills, and climb version after version.
With the capability proven, the question became: how does everyone get to use it? Self-evolution should not stay a curve in a paper — it should be infrastructure that works out of the box on every developer's desk. Bringing an efficient self-improving harness to everyone is why we built PenguinHarness — and it is right there in the name: Efficient Self-Improving Harness for Everyone.
Why PenguinHarness
Three reasons, in deliberate order — from task quality, to how agents get built, to how they keep improving.
1. Better on complex tasks, at lower cost
A deliberately minimal toolset over clean low-level interfaces: fewer tool calls, fewer Tokens, deeply tuned for open models like DeepSeek. Each harness runs the model it is normally paired with — the comparison is between the products as people actually use them — head-to-head on two suites:
Complex data analysis (15 tasks, single run; PenguinHarness and Codex at thinking xhigh, Claude Code at max):
| Framework | Model | Accuracy (%) | Tokens (M) | Cost ($) |
|---|---|---|---|---|
| PenguinHarness | DeepSeek V4 Pro | 66.67 | 18.04 | 0.55 |
| Claude Code | Claude Opus 4.8 | 53.33 | 22.20 | 38.48 |
| OpenAI Codex | GPT-5.5 | 53.33 | 13.72 | 19.41 |
Coding tasks (40 tasks × 2 runs; accuracy is over all 80 outcomes):
| Framework | Model | Accuracy (%) | Tokens (M) | Cost ($) |
|---|---|---|---|---|
| PenguinHarness | DeepSeek V4 Pro | 71.25 | 200.00 | 3.81 |
| Claude Code | Claude Opus 4.8 | 86.25 | 151.61 | 146.97 |
| OpenAI Codex | GPT-5.5 | 71.25 | 251.20 | 220.08 |
Tokens and cost are suite totals, not per-run means. On data analysis we take the highest accuracy of the three — 66.67% against 53.33% for both — while spending 1/35 of what Codex spent and 1/70 of Claude Code's. On coding we tie Codex at 71.25% and trail Claude Code's 86.25%, but the whole suite cost us $3.81 against their $220.08 and $146.97: comparable work, one to two orders of magnitude apart on the bill.
2. One sentence, and an Agent builds your Agent app
Type one sentence, and an Agent builds the complete Agent application for you — scaffold, code, and run instructions, end to end:
Collect the docs from https://github.com/ericbuess/claude-code-docs and build a RAG app that answers Claude Code questions as a configuration expert, citing its sources.
And this is the finished product — a docs expert with retrieval, cited sources that link to the original files, and example questions built in:
And generating this entire RAG app burned just $0.02 (¥0.2) of tokens — on DeepSeek V4 Pro.
3. Self-evolution: it gets stronger with use
With PenguinHarness Skills, an Agent evaluates and optimizes itself: the Optimizer orchestrates multiple Evaluators to score in parallel, uses the scores and run traces to find where points were lost, and upgrades the Agent from version N to N+1 — with a snapshot before every round, and every request replayable in the Trace view. A self-evolution demo video is coming soon.
Evolution within bounds, security first
The biggest worry about self-evolution is losing control. PenguinHarness answers with a contract (CONTRACT.md):
- Evolution is strictly confined to Workspace and Skills — the harness core security boundary is never modified;
- Tool calls require approval first, and every approval leaves an audit record;
- Risky changes are preceded by version snapshots, so any round of evolution can be rolled back;
- Fully open source and locally deployed — data never leaves your machine, meeting enterprise data-security requirements.
Supported models
| Model | Providers |
|---|---|
| DeepSeek V4 | DeepSeek, OpenRouter, Fireworks AI, SiliconFlow, Qwen Token Plan |
| Kimi K3 | Moonshot AI, OpenRouter, Qwen Pay-As-You-Go |
| GLM 5.2 | Z.AI, OpenRouter, Fireworks AI, SiliconFlow, Qwen Token Plan, Qwen Pay-As-You-Go |
| Hunyuan 3 | OpenRouter |
| Qwen 3.8 Max | Qwen Token Plan (preview) |
| GPT 5.5 | OpenAI, OpenRouter |
| Gemini 3.5 Flash | Google Gemini, OpenRouter |
| Claude Opus 4.8 | Anthropic, OpenRouter |
Any OpenAI-protocol endpoint is supported: pick a preset above, or point a custom endpoint at any of the 1000+ online and local models.
How to use it
Install with one command (Linux / macOS, x64 / arm64, bundled Node runtime), then launch the Web UI:
curl -fsSL https://penguin.ooo/install.sh | sh
penguin web # opens http://127.0.0.1:7364 (first login: admin / penguin-2026)
Open the Models page, paste an API key under the DeepSeek or OpenRouter group and set it as default; then head back to Chat and hand the Agent its first task — e.g. "Analyze data.csv and summarize quarterly sales".
What's next
- Public release of the benchmark suite;
- A desktop app;
- Windows support;
- More to come.
Join the community and build with us
A self-improving harness needs a community that improves with it. Come discuss, request, and contribute — your first Issue is the best way to start:
- Discord: chat with us and other developers in real time;
- X (Twitter): follow the latest updates;
- WeChat group: Chinese community discussions;
- GitHub: stars, Issues, and PRs all welcome.
Self-evolving agent infrastructure, for everyone — starting today.
