Files
penguin-harness/README.md
T
Yaowei Zheng d4faee3a1e Changelog, dev startup, README, AgentHub 0.4.0, model catalog, and landing site (#7)
Branch-length batch covering tooling, the model layer, the Web App and the public
surfaces. Highlights:

- Changelog: a per-release `changelog/<version>/` tree, grouped by the surface each
  change touches, with a root CHANGELOG.md holding one line per release.
- Dev startup: `scripts/dev-prebuild.mjs` serializes the skills+core prebuild behind a
  lock and keeps `pnpm install` current; `pnpm dev` runs server+web together.
- AgentHub 0.3.3 -> 0.4.0: OmniMessage complete payloads carry one opaque `fidelity`
  object in place of item-level `signature`/`phase`, threaded verbatim through Trace,
  replay and resume; malformed classification adapted to the new error types.
- Model layer: a model is always referenced by an explicit `(provider, model_id)` pair.
  The provider is never inferred, guessed or defaulted -- both the catalog inference and
  the unique-match config resolution are gone, and CLI, SDK, server routes and
  run_subagent all require the complete pair. Catalog gains the Qwen Token Plan, Qwen
  Pay-As-You-Go and Fireworks AI gateways, plus an expanded OpenRouter group.
- Web App: catalog preset sync and per-group speed test on the Models page, positional
  slash commands, a markdown renderer, skill-library update reminders, and a vertically
  centred draft page whose upward menus size themselves to the room available.
- Public surfaces: restructured READMEs, the penguin.ooo landing site and blog, refreshed
  benchmark results for both suites, and the demo videos playing on the landing page.

Includes the fixes from a full review of the branch: 23 confirmed findings, among them a
provider-inference bug that could send one vendor's API key to another vendor's endpoint,
and an Escape handler that destroyed the composer's contents unrecoverably.

Verified on the branch head: pnpm test (1127 passing, 7 packages), pnpm typecheck and
pnpm format:check clean, Playwright e2e 14/14.
2026-07-21 17:43:31 +08:00

8.3 KiB
Raw Blame History

PenguinHarness logo

PenguinHarness

With LangChain, you build agents by hand — at 1× speed.
With PenguinHarness, agents build agents — at 100×.

A zero-code Harness CLI and Web UI, connected to 1000+ models.

CI Deploy Site License: Apache-2.0 Node >= 24

Website Docs Blog

Discord X (Twitter) WeChat

English | 简体中文

Why PenguinHarness

Three reasons, in deliberate order — from task quality, to how agents get built, to how they keep improving.

1. 🏆 Comparable quality, one to two orders of magnitude cheaper

A deliberately minimal toolset over clean low-level interfaces: fewer tool calls, fewer tokens — deeply tuned for open models like DeepSeek. Each harness on the model it is normally paired with, same tasks, head-to-head:

Benchmark: PenguinHarness leads the data-analysis suite and ties OpenAI Codex on coding, at a small fraction of both rivals' cost

Best accuracy on data analysis — at 1/70 of Claude Code's cost.

2. ⚡ One sentence, and an Agent builds your Agent app

Type one sentence, and an Agent builds the complete Agent application for you — scaffold, code, and run instructions, end to end:

Collect the docs from https://github.com/ericbuess/claude-code-docs and build a RAG app that answers Claude Code questions as a configuration expert, citing its sources.

And this is the finished product — a docs expert with retrieval, cited sources that link to the original files, and example questions built in:

https://github.com/user-attachments/assets/9b7033e8-f08a-4c3f-bd33-547896664e6e

And generating this entire RAG app burned just $0.02 (¥0.2) of tokens — on DeepSeek V4 Pro.

3. 🧬 Self-evolution: it gets stronger with use

With PenguinHarness Skills, an Agent evaluates and optimizes itself: run the benchmark, find the lost points, ship version N+1 — with a snapshot before every round, and every request observable in the Trace view.

https://github.com/user-attachments/assets/922d13a6-5ffc-4685-9a39-352f02f9afc0

Supported Models

Model Providers
DeepSeek V4 DeepSeek, OpenRouter, Fireworks AI, SiliconFlow, Qwen Token Plan
Kimi K3 OpenRouter, Qwen Pay-As-You-Go
Kimi K2.6 Moonshot AI
GLM 5.2 Z.AI, OpenRouter, Fireworks AI, SiliconFlow, Qwen Token Plan, Qwen Pay-As-You-Go
Hunyuan 3 OpenRouter
Qwen 3.8 Max Qwen Token Plan (preview)
GPT 5.5 OpenAI, OpenRouter
Gemini 3.5 Flash Google Gemini, OpenRouter
Claude Opus 4.8 Anthropic, OpenRouter

Any OpenAI-protocol endpoint is supported: pick a preset above, or point a custom endpoint at any of the 1000+ online and local models.

Requirements

Requirement Supported
OS Linux, macOS
Architecture x64, arm64
Runtime bundled by the one-line installer (npm installs need Node >= 24)
Model an API key for at least one model

Installation

🌐 Web App — for humans

🚀 Install and launch the full experience (multi-session chat, Agent/skill/model management, usage stats, Trace observability, evaluation center):

curl -fsSL https://penguin.ooo/install.sh | sh
penguin web        # start the service and open http://127.0.0.1:7364 (first login: admin / penguin-2026)

📦 Or via npm: npm install -g @prismshadow/penguin-cli. Configure models on the in-app Models page, then chat.

🤖 CLI & SDK — for agents

The same engine, scriptable — made to be driven by agents (and agents building agents):

penguin config model add --provider deepseek --model-id deepseek-v4-pro --api-key sk-... --set-default
penguin run -m "Create hello.txt containing Hello, Penguin"   # one-shot task
penguin chat       # interactive REPL (/compact, /exit, Ctrl-C to interrupt)
penguin server     # headless service (same API the Web App uses)
import { createAgent, isCompleteModelMessage, userText } from "@prismshadow/penguin-core";

const agent = await createAgent({ agentId: "default_agent" });
const session = await agent.createSession({ workspaceDir: process.cwd() });

for await (const output of session.run([userText("Create hello.txt containing hi")], {
  approve: async () => "allow", // per-tool-call approval
})) {
  if (isCompleteModelMessage(output) && output.payload.type === "text") {
    console.log(output.payload.text);
  }
}

Roadmap

  • Public release of the benchmark suite
  • Desktop app
  • Windows support
  • More to come…

Development

pnpm install && pnpm build   # build first: core's exports point at dist/
pnpm dev                     # backend + web app together (prefixed logs, deps built once)

See CONTRIBUTING.md for the full workspace guide: dev commands, quality gates, repo layout, and the changelog rule.

Citation

If you use PenguinHarness in your research, please cite:

@software{penguinharness2026,
  author  = {{PrismShadow Team}},
  title   = {PenguinHarness: Efficient Self-Improving Harness for Everyone},
  year    = {2026},
  url     = {https://github.com/Prism-Shadow/penguin-harness},
  license = {Apache-2.0}
}

License

Apache-2.0 © 2026 Prism Shadow

Built with ❤️ by Yaowei Zheng (author of LlamaFactory), the PrismShadow AI Team, and Fable 5.