Adds three bilingual posts, each grounded in primary sources rather than secondary summaries, plus a new blog category to hold them. Simple Harness Is All You Need — opens on the counter-intuitive result in the Databricks coding-agent benchmark: on their cost-versus-pass-rate Pareto chart, the highest score on the board belongs to Opus 4.8 on the minimal Pi harness, ahead of the same model on Claude Code at maximum effort for roughly half the cost per task. Keeps Databricks' own caution and notes that Pi at max effort lands well below Claude Code at comparable spend. Maps the result onto PenguinHarness's measured design: six built-in tools with no file tools, a 72-line system prompt, a 16,000-character output cap, and compaction into a fresh context. The Easiest Way to Build AI Agents in 2026 — argues the cost of building an agent has moved out of the agent and into the stack around it: LangChain to build, LangGraph to orchestrate, LangSmith or Langfuse to observe and evaluate, LangGraph Platform to deploy. Five products, two or three vendors, and a person who becomes the optimization loop. AI Infrastructure: Past, Present, and Future — the stack used to build AI (PyTorch, vLLM, Ollama, LlamaFactory) assumes a human operator who carries state in their head and treats errors as a starting point. It needs no reinventing for agents; what was missing is the operating knowledge, which the ollama, vllm and llamafactory skills encode. Both benchmark comparisons disclose the results we lose as well as the ones we win, and the framework post names two cases where you should pick something else. Also adds the Perspectives / 观点 category with a teal badge, filter chip and both dictionaries, keeping Tech practice for the hands-on AMD walkthroughs.
PenguinHarness
With LangChain, you build agents by hand — at 1× speed.
With PenguinHarness, agents build agents — at 100×.
A zero-code Harness CLI and Web UI, connected to 1000+ models.
English | 简体中文
Why PenguinHarness
Three reasons, in deliberate order — from task quality, to how agents get built, to how they keep improving.
1. 🏆 Comparable quality, one to two orders of magnitude cheaper
A deliberately minimal toolset over clean low-level interfaces: fewer tool calls, fewer tokens — deeply tuned for open models like DeepSeek. Each harness on the model it is normally paired with, same tasks, head-to-head:
Best accuracy on data analysis — at 1/70 of Claude Code's cost.
2. ⚡ One sentence, and an Agent builds your Agent app
Type one sentence, and an Agent builds the complete Agent application for you — scaffold, code, and run instructions, end to end:
Collect the docs from https://github.com/ericbuess/claude-code-docs and build a RAG app that answers Claude Code questions as a configuration expert, citing its sources.
And this is the finished product — a docs expert with retrieval, cited sources that link to the original files, and example questions built in:
https://github.com/user-attachments/assets/9b7033e8-f08a-4c3f-bd33-547896664e6e
And generating this entire RAG app burned just $0.02 (¥0.2) of tokens — on DeepSeek V4 Pro.
3. 🧬 Self-evolution: it gets stronger with use
With PenguinHarness Skills, an Agent evaluates and optimizes itself: run the benchmark, find the lost points, ship version N+1 — with a snapshot before every round, and every request observable in the Trace view.
https://github.com/user-attachments/assets/922d13a6-5ffc-4685-9a39-352f02f9afc0
Built-in Skills
Four Skill groups ship in the box (docs); Agents can also write and optimize their own:
| Group | Skills |
|---|---|
| Office Productivity | data-analysis, firecrawl |
| Software Development | web-design, software-engineering |
| AI App Development | penguin-sdk, penguin-cli, agenthub-models, vllm, ollama, llamafactory |
| Agent Tuning | agent-creation, benchmark-design, agent-evaluation, agent-optimization |
Supported Models
| Model | Providers |
|---|---|
| DeepSeek V4 | DeepSeek, OpenRouter, Fireworks AI, SiliconFlow, Qwen Token Plan |
| Kimi K3 | Moonshot AI, OpenRouter, Qwen Pay-As-You-Go |
| GLM 5.2 | Z.AI, OpenRouter, Fireworks AI, SiliconFlow, Qwen Token Plan, Qwen Pay-As-You-Go |
| Hunyuan 3 | OpenRouter |
| Qwen 3.8 Max | Qwen Token Plan (preview) |
| GPT 5.6 | OpenRouter |
| Gemini 3.6 Flash | Google Gemini, OpenRouter |
| Claude 5 | Anthropic, OpenRouter |
Each family's latest generation only — the app's Models page lists every built-in preset, and any OpenAI-protocol endpoint works too: pick a preset, or point a custom endpoint at any of the 1000+ online and local models.
Requirements
| Requirement | Supported |
|---|---|
| OS | Linux, macOS |
| Architecture | x64, arm64 |
| Runtime | bundled by the one-line installer (npm installs need Node >= 24) |
| Model | an API key for at least one model |
Installation
🌐 Web App — for humans
🚀 Install and launch the full experience (multi-session chat, Agent/skill/model management, usage stats, Trace observability, evaluation center):
curl -fsSL https://penguin.ooo/install.sh | sh
penguin web # start the service and open http://127.0.0.1:7364 (first login: admin / penguin-2026)
📦 Or via npm: npm install -g @prismshadow/penguin-cli. Configure models on the in-app Models page, then chat.
🤖 CLI & SDK — for agents
The same engine, scriptable — made to be driven by agents (and agents building agents):
penguin config model add --provider deepseek --model-id deepseek-v4-pro --api-key sk-... --set-default
penguin run -m "Create hello.txt containing Hello, Penguin" # one-shot task
penguin chat # interactive REPL (/compact, /exit, Ctrl-C to interrupt)
penguin server # headless service (same API the Web App uses)
import { createAgent, isCompleteModelMessage, userText } from "@prismshadow/penguin-core";
const agent = await createAgent({ agentId: "default_agent" });
const session = await agent.createSession({ workspaceDir: process.cwd() });
for await (const output of session.run([userText("Create hello.txt containing hi")], {
approve: async () => "allow", // per-tool-call approval
})) {
if (isCompleteModelMessage(output) && output.payload.type === "text") {
console.log(output.payload.text);
}
}
Roadmap
- Public release of the benchmark suite
- Desktop app
- Windows support
- Agent company and templates
- Company-level self evolving
- More to come…
Development
pnpm install && pnpm build # build first: core's exports point at dist/
pnpm dev # backend + web app together (prefixed logs, deps built once)
See CONTRIBUTING.md for the full workspace guide: dev commands, quality gates, repo layout, and the changelog rule.
Citation
If you use PenguinHarness in your research, please cite:
@software{penguinharness2026,
author = {{PrismShadow Team}},
title = {PenguinHarness: Efficient Self-Improving Harness for Everyone},
year = {2026},
url = {https://github.com/Prism-Shadow/penguin-harness},
license = {Apache-2.0}
}
License
Apache-2.0 © 2026 Prism Shadow
Built with ❤️ by Yaowei Zheng (author of LlamaFactory), the PrismShadow AI Team, and Fable 5.