Changelog, dev startup, README, AgentHub 0.4.0, model catalog, and landing site (#7)

Branch-length batch covering tooling, the model layer, the Web App and the public
surfaces. Highlights:

- Changelog: a per-release `changelog/<version>/` tree, grouped by the surface each
  change touches, with a root CHANGELOG.md holding one line per release.
- Dev startup: `scripts/dev-prebuild.mjs` serializes the skills+core prebuild behind a
  lock and keeps `pnpm install` current; `pnpm dev` runs server+web together.
- AgentHub 0.3.3 -> 0.4.0: OmniMessage complete payloads carry one opaque `fidelity`
  object in place of item-level `signature`/`phase`, threaded verbatim through Trace,
  replay and resume; malformed classification adapted to the new error types.
- Model layer: a model is always referenced by an explicit `(provider, model_id)` pair.
  The provider is never inferred, guessed or defaulted -- both the catalog inference and
  the unique-match config resolution are gone, and CLI, SDK, server routes and
  run_subagent all require the complete pair. Catalog gains the Qwen Token Plan, Qwen
  Pay-As-You-Go and Fireworks AI gateways, plus an expanded OpenRouter group.
- Web App: catalog preset sync and per-group speed test on the Models page, positional
  slash commands, a markdown renderer, skill-library update reminders, and a vertically
  centred draft page whose upward menus size themselves to the room available.
- Public surfaces: restructured READMEs, the penguin.ooo landing site and blog, refreshed
  benchmark results for both suites, and the demo videos playing on the landing page.

Includes the fixes from a full review of the branch: 23 confirmed findings, among them a
provider-inference bug that could send one vendor's API key to another vendor's endpoint,
and an Escape handler that destroyed the composer's contents unrecoverably.

Verified on the branch head: pnpm test (1127 passing, 7 packages), pnpm typecheck and
pnpm format:check clean, Playwright e2e 14/14.
This commit is contained in:
Yaowei Zheng
2026-07-21 17:43:31 +08:00
committed by GitHub
parent abf0a2f248
commit d4faee3a1e
214 changed files with 9406 additions and 1688 deletions
+6
View File
@@ -0,0 +1,6 @@
# Changelog
One brief line per release. Per-release detail lives in [`changelog/<version>/`](changelog/).
- **0.0.2** — unreleased. ([details](changelog/0.0.2/README.md))
- **0.0.1** — 2026-07-19. First tagged release; changelog history starts after this tag.
+100
View File
@@ -0,0 +1,100 @@
# Contributing to PenguinHarness
Thanks for helping build PenguinHarness! This guide covers the workspace setup, daily
commands, quality gates, and the repo's working rules.
## Prerequisites
- Node >= 24
- pnpm 10 (`corepack enable` or `npm install -g pnpm`)
## Setup and daily commands
```bash
pnpm install
pnpm build # build first: core's exports point at dist/
pnpm dev # backend + web app together (prefixed logs, deps built once)
pnpm dev:server # backend at 127.0.0.1:7364
pnpm dev:web # web app (Vite) at 127.0.0.1:7365, /api proxied
pnpm dev:docs # docs site (Vite) at 127.0.0.1:7367
pnpm dev:landing # landing page (Vite) at 127.0.0.1:7366
BASE_PATH=/ pnpm build:site # assemble landing + docs exactly like the Pages deploy
```
Every dev command runs `scripts/dev-prebuild.mjs` first, which (behind a lock that
serializes concurrent invocations) **keeps `pnpm install` current automatically** — a
fresh clone or a pulled lockfile change installs before starting, and an up-to-date tree
pays nothing (the lockfile hash is stamped) — then prebuilds the workspace deps (skills,
core) with back-to-back builds deduped: starting `dev:server` and `dev:web` at the same
time (or just `pnpm dev`) installs and builds exactly once. `dev:docs` / `dev:landing`
run the install check only (`--install-only`).
Copy `.env.example` to `.env` for model credentials in development.
## Repo layout
A pnpm monorepo (TypeScript, Node >= 24). One install ships four layers that share a
single data directory (`~/.penguin/data`) and a single message protocol (OmniMessage):
| Package | Name | Role |
| ------------------------------------ | ----------------------------- | ------------------------------------------------------------------------------------------------------- |
| [`packages/core`](packages/core) | `@prismshadow/penguin-core` | SDK & engine: ReAct loop, OmniMessage protocol, LLM/Environment interface contracts, Agent State, Trace |
| [`packages/cli`](packages/cli) | `@prismshadow/penguin-cli` | The `penguin` command: REPL, one-shot runs, model & vault config, service launcher |
| [`packages/server`](packages/server) | `@prismshadow/penguin-server` | Web backend: HTTP API + SSE streaming, multi-user auth, Project authorization, usage stats |
| [`packages/web`](packages/web) | `@prismshadow/penguin-web` | Web App: multi-session chat, Agent/skill/model management, Trace observability, evaluation center |
| [`packages/skills`](packages/skills) | `@prismshadow/penguin-skills` | Built-in skill library (agent creation, benchmarking, evaluation, optimization, …) |
| [`packages/landing`](packages/landing) | — | Product landing page (this repo's website) |
| [`packages/docs`](packages/docs) | — | Documentation site (bilingual, deployed under `/docs/`) |
Responsibilities split by source of truth: the **SDK** owns protocol and execution
(message parsing, the agent loop, tools), the **Server** owns the multi-user runtime
(auth, SSE streaming, scheduled tasks), and the **file layer** under `~/.penguin/data`
owns everything editable and recorded (prompts, Skills, secrets, Traces). The full map
is in [Architecture → Division of responsibilities](https://penguin.ooo/docs/architecture).
## Quality gates
CI runs all of these on every PR — run them locally before pushing:
```bash
pnpm format:check # prettier
pnpm typecheck
pnpm test # unit suites for every package
```
End-to-end suites (optional locally, slower):
```bash
npx playwright install chromium # once
pnpm --filter @prismshadow/penguin-web test:e2e # browser e2e against a mock LLM
pnpm test:e2e # core live-model e2e, needs DEEPSEEK_API_KEY
```
## Working rules
- **English is the repository's working language** — code, comments, error/log messages,
test names and fixtures, package metadata, and developer docs. Chinese appears only
where it is the content itself: zh i18n catalogs and fields (`strings.ts` dictionaries,
CLI `i18n.ts`, `titleZh`, `short_description_zh`), `*.zh.md` documents, and test
literals that assert zh i18n output or exercise CJK-specific behavior.
- **Every change ships with a changelog entry**: add
`changelog/<version>/YYYY-MM-DD-<semantic-id>.md` under the next unreleased version
(released versions' folders are frozen) — an H1 title, a one-sentence summary
paragraph, then details — and add a one-line link for it to that version's index,
`changelog/<version>/README.md`. The layout is documented in
[`changelog/README.md`](changelog/README.md). Related changes may share one entry
file (extending its details) instead of opening a new file per small change.
- README assets under `assets/readme/` are generated — the benchmark charts from the
landing benchmark data, and the demo screenshots via
`node packages/landing/scripts/capture-readme-demo.mjs` (build first; needs Playwright
chromium). Regenerate rather than hand-editing.
## Pull requests
- Branch from `main`; keep PRs focused on one topic.
- Make sure CI is green (build, format, typecheck, tests) and describe user-visible
changes in the PR body.
- New user-facing behavior should come with tests, and with docs updates when it changes
documented behavior (README, docs site).
+102 -65
View File
@@ -4,12 +4,9 @@
<h1 align="center">PenguinHarness</h1>
<p align="center"><b>Efficient Self-Improving Harness for Everyone</b></p>
<p align="center"><b>With LangChain, you build agents by hand — at 1× speed.<br />With PenguinHarness, agents build agents — at 100×.</b></p>
<p align="center">
Open-source, local-first infrastructure that builds AI agents for you —
from automatic agent construction to recursive self-improvement.
</p>
<p align="center">A zero-code Harness CLI and Web UI, connected to 1000+ models.</p>
<p align="center">
<a href="https://github.com/Prism-Shadow/penguin-harness/actions/workflows/ci.yml"><img src="https://github.com/Prism-Shadow/penguin-harness/actions/workflows/ci.yml/badge.svg" alt="CI" /></a>
@@ -19,58 +16,105 @@
</p>
<p align="center">
English | <a href="README.zh.md">简体中文</a> ·
<a href="https://penguin.ooo/">Website</a> ·
<a href="https://penguin.ooo/docs/">Docs</a> ·
<a href="https://penguin.ooo/blog">Blog</a>
<a href="https://penguin.ooo/"><img src="https://img.shields.io/badge/Website-penguin.ooo-1f6feb?logo=googlechrome&logoColor=white" alt="Website" /></a>
<a href="https://penguin.ooo/docs/"><img src="https://img.shields.io/badge/Docs-penguin.ooo%2Fdocs-1f6feb?logo=readthedocs&logoColor=white" alt="Docs" /></a>
<a href="https://penguin.ooo/blog"><img src="https://img.shields.io/badge/Blog-penguin.ooo%2Fblog-1f6feb?logo=rss&logoColor=white" alt="Blog" /></a>
</p>
<p align="center">
<picture>
<source media="(prefers-color-scheme: dark)" srcset="packages/landing/src/assets/shots/chat-en-dark.webp" />
<img src="packages/landing/src/assets/shots/chat-en-light.webp" alt="PenguinHarness Web App — multi-session chat with live streaming tool calls" width="920" />
</picture>
<a href="https://discord.gg/eFHKqqcU3D"><img src="https://img.shields.io/badge/Discord-join%20chat-5865F2?logo=discord&logoColor=white" alt="Discord" /></a>
<a href="https://x.com/code_hiyouga"><img src="https://img.shields.io/badge/X-code%5Fhiyouga-000000?logo=x&logoColor=white" alt="X (Twitter)" /></a>
<a href="https://github.com/Prism-Shadow/penguin-harness-community/blob/main/wechat/group.jpg"><img src="https://img.shields.io/badge/WeChat-user%20group-07C160?logo=wechat&logoColor=white" alt="WeChat" /></a>
</p>
---
<p align="center">English | <a href="README.zh.md">简体中文</a></p>
## Why PenguinHarness
- **Simplest Is the Best** — a deliberately minimal toolset over clean low-level interfaces: fewer tool calls, fewer tokens, complex tasks done efficiently.
- **Harness for Building Agents** — with the PenguinHarness SDK, an Agent builds complete Agent applications for you, autonomously, from scratch.
- **Harness for Recursive Self-Improvement** — with PenguinHarness Skills, an Agent evaluates and optimizes itself: benchmark, find the lost points, ship version N+1, snapshot before every round.
- **Local-first and lightweight** — 100% open source, runs on a single CPU, your data never leaves the machine. 1000+ online and local models reachable through one gateway.
- **Everything observable** — every request, tool call and approval decision lands in an append-only Trace; any Session can be resumed from it.
Three reasons, in deliberate order — from task quality, to how agents get built, to how they keep improving.
## Quickstart
### 1. 🏆 Comparable quality, one to two orders of magnitude cheaper
Install with one command (Linux / macOS, x64 / arm64, bundled Node runtime):
A deliberately minimal toolset over clean low-level interfaces: fewer tool calls, fewer tokens — deeply tuned for open models like DeepSeek. Each harness on the model it is normally paired with, same tasks, head-to-head:
```bash
curl -fsSL https://github.com/Prism-Shadow/penguin-harness/releases/latest/download/install.sh | sh
<p align="center">
<picture>
<source media="(prefers-color-scheme: dark)" srcset="assets/readme/benchmark-dark.svg" />
<img src="assets/readme/benchmark-light.svg" alt="Benchmark: PenguinHarness leads the data-analysis suite and ties OpenAI Codex on coding, at a small fraction of both rivals' cost" width="920" />
</picture>
</p>
**Best accuracy on data analysis — at 1/70 of Claude Code's cost.**
### 2. ⚡ One sentence, and an Agent builds your Agent app
Type one sentence, and an Agent builds the complete Agent application for you — scaffold, code, and run instructions, end to end:
```text
Collect the docs from https://github.com/ericbuess/claude-code-docs and build a RAG app that answers Claude Code questions as a configuration expert, citing its sources.
```
Or via npm (requires Node >= 24; the command it installs is `penguin`):
And this is the finished product — a docs expert with retrieval, cited sources that link to the original files, and example questions built in:
https://github.com/user-attachments/assets/9b7033e8-f08a-4c3f-bd33-547896664e6e
**And generating this entire RAG app burned just $0.02 (¥0.2) of tokens — on DeepSeek V4 Pro.**
### 3. 🧬 Self-evolution: it gets stronger with use
With PenguinHarness Skills, an Agent evaluates and optimizes itself: run the benchmark, find the lost points, ship version N+1 — with a snapshot before every round, and every request observable in the Trace view.
https://github.com/user-attachments/assets/922d13a6-5ffc-4685-9a39-352f02f9afc0
## Supported Models
| Model | Providers |
| ---------------- | -------------------------------------------------------------------------------- |
| DeepSeek V4 | DeepSeek, OpenRouter, Fireworks AI, SiliconFlow, Qwen Token Plan |
| Kimi K3 | OpenRouter, Qwen Pay-As-You-Go |
| Kimi K2.6 | Moonshot AI |
| GLM 5.2 | Z.AI, OpenRouter, Fireworks AI, SiliconFlow, Qwen Token Plan, Qwen Pay-As-You-Go |
| Hunyuan 3 | OpenRouter |
| Qwen 3.8 Max | Qwen Token Plan (preview) |
| GPT 5.5 | OpenAI, OpenRouter |
| Gemini 3.5 Flash | Google Gemini, OpenRouter |
| Claude Opus 4.8 | Anthropic, OpenRouter |
Any OpenAI-protocol endpoint is supported: pick a preset above, or point a custom endpoint at any of the 1000+ online and local models.
## Requirements
| Requirement | Supported |
| ------------ | -------------------------------------------------------------------------- |
| OS | Linux, macOS |
| Architecture | x64, arm64 |
| Runtime | bundled by the one-line installer (npm installs need Node >= 24) |
| Model | an API key for at least one model |
## Installation
### 🌐 Web App — for humans
🚀 Install and launch the full experience (multi-session chat, Agent/skill/model management, usage stats, Trace observability, evaluation center):
```bash
npm install -g @prismshadow/penguin-cli
curl -fsSL https://penguin.ooo/install.sh | sh
penguin web # start the service and open http://127.0.0.1:7364 (first login: admin / penguin-2026)
```
Then launch the Web App — or stay in the terminal:
📦 Or via npm: `npm install -g @prismshadow/penguin-cli`. Configure models on the in-app Models page, then chat.
### 🤖 CLI & SDK — for agents
The same engine, scriptable — made to be driven by agents (and agents building agents):
```bash
penguin web # start the service and open http://127.0.0.1:7364 (first login: admin / admin123)
penguin server # same service, headless
# configure a model once (or use the in-app Models page)
penguin config model add --model-id deepseek-v4-pro --api-key sk-... --set-default
penguin config model add --provider deepseek --model-id deepseek-v4-pro --api-key sk-... --set-default
penguin run -m "Create hello.txt containing Hello, Penguin" # one-shot task
penguin chat # interactive REPL (/compact, /exit, Ctrl-C to interrupt)
penguin server # headless service (same API the Web App uses)
```
Using the SDK directly:
```ts
import { createAgent, isCompleteModelMessage, userText } from "@prismshadow/penguin-core";
@@ -86,45 +130,38 @@ for await (const output of session.run([userText("Create hello.txt containing hi
}
```
## What's inside
## Roadmap
A pnpm monorepo (TypeScript, Node >= 24). One install ships four layers that share a single data directory (`~/.penguin/data`) and a single message protocol (OmniMessage):
| Package | Name | Role |
| --- | --- | --- |
| [`packages/core`](packages/core) | `@prismshadow/penguin-core` | SDK & engine: ReAct loop, OmniMessage protocol, LLM/Environment interface contracts, Agent State, Trace |
| [`packages/cli`](packages/cli) | `@prismshadow/penguin-cli` | The `penguin` command: REPL, one-shot runs, model & vault config, service launcher |
| [`packages/server`](packages/server) | `@prismshadow/penguin-server` | Web backend: HTTP API + SSE streaming, multi-user auth, Project authorization, usage stats |
| [`packages/web`](packages/web) | `@prismshadow/penguin-web` | Web App: multi-session chat, Agent/skill/model management, Trace observability, evaluation center |
| [`packages/skills`](packages/skills) | `@prismshadow/penguin-skills` | Built-in skill library (agent creation, benchmarking, evaluation, optimization, …) |
| [`packages/landing`](packages/landing) | — | Product landing page (this repo's website) |
| [`packages/docs`](packages/docs) | — | Documentation site (bilingual, deployed under `/docs/`) |
Responsibilities split by source of truth: the **SDK** owns protocol and execution (message parsing, the agent loop, tools), the **Server** owns the multi-user runtime (auth, SSE streaming, scheduled tasks), and the **file layer** under `~/.penguin/data` owns everything editable and recorded (prompts, Skills, secrets, Traces). The full design-by-design map is in [Architecture → Division of responsibilities](https://penguin.ooo/docs/architecture).
## Documentation
The docs site covers both usage and design: [Introduction](https://penguin.ooo/docs/) · [Quickstart](https://penguin.ooo/docs/quickstart) · [Architecture](https://penguin.ooo/docs/architecture) · [The OmniMessage Protocol](https://penguin.ooo/docs/omni-message) · [Core Interfaces](https://penguin.ooo/docs/interfaces) · [The Agent Loop](https://penguin.ooo/docs/agent-loop) · [CLI Reference](https://penguin.ooo/docs/cli) · [Server API](https://penguin.ooo/docs/server-api) · [Configuration](https://penguin.ooo/docs/configuration)
Every doc page has a "Copy Markdown" button, so you can paste it straight into a model context.
- [ ] Public release of the benchmark suite
- [ ] Desktop app
- [ ] Windows support
- More to come…
## Development
```bash
pnpm install
pnpm build # build first: core's exports point at dist/
pnpm typecheck
pnpm test
pnpm dev:server # backend at 127.0.0.1:7364
pnpm dev:web # web app (Vite) at 127.0.0.1:7365, /api proxied
pnpm dev:docs # docs site (Vite) at 127.0.0.1:7367
BASE_PATH=/ pnpm build:site # assemble landing + docs exactly like the Pages deploy
pnpm install && pnpm build # build first: core's exports point at dist/
pnpm dev # backend + web app together (prefixed logs, deps built once)
```
Copy `.env.example` to `.env` for model credentials in development. E2E tests run against a live model (`pnpm test:e2e`, needs `DEEPSEEK_API_KEY`).
See [CONTRIBUTING.md](CONTRIBUTING.md) for the full workspace guide: dev commands, quality gates, repo layout, and the changelog rule.
## Citation
If you use PenguinHarness in your research, please cite:
```bibtex
@software{penguinharness2026,
author = {{PrismShadow Team}},
title = {PenguinHarness: Efficient Self-Improving Harness for Everyone},
year = {2026},
url = {https://github.com/Prism-Shadow/penguin-harness},
license = {Apache-2.0}
}
```
## License
[Apache-2.0](LICENSE) © 2026 Prism Shadow
Built with ❤️ by [Yaowei Zheng](https://github.com/hiyouga) (author of [LlamaFactory](https://github.com/hiyouga/LlamaFactory)), the [PrismShadow AI Team](https://github.com/Prism-Shadow), and [Fable 5](https://www.anthropic.com/news/claude-fable-5-mythos-5).
+108 -70
View File
@@ -4,11 +4,9 @@
<h1 align="center">PenguinHarness</h1>
<p align="center"><b>Efficient Self-Improving Harness for Everyone</b></p>
<p align="center"><b>使用 LangChain,以 1 倍速度人工构建 Agent;<br />使用 PenguinHarness,以 100 倍速度用 Agent 构建 Agent。</b></p>
<p align="center">
开源、本地优先的 AI Agent 基础设施——从自动构建 Agent 到递归自我进化。
</p>
<p align="center">零代码 Harness CLI 与 Web UI,连接 1000+ 模型。</p>
<p align="center">
<a href="https://github.com/Prism-Shadow/penguin-harness/actions/workflows/ci.yml"><img src="https://github.com/Prism-Shadow/penguin-harness/actions/workflows/ci.yml/badge.svg" alt="CI" /></a>
@@ -18,66 +16,113 @@
</p>
<p align="center">
<a href="README.md">English</a> | 简体中文 ·
<a href="https://penguin.ooo/">官网</a> ·
<a href="https://penguin.ooo/docs/">文档</a> ·
<a href="https://penguin.ooo/blog">博客</a>
<a href="https://penguin.ooo/"><img src="https://img.shields.io/badge/%E5%AE%98%E7%BD%91-penguin.ooo-1f6feb?logo=googlechrome&logoColor=white" alt="官网" /></a>
<a href="https://penguin.ooo/docs/"><img src="https://img.shields.io/badge/%E6%96%87%E6%A1%A3-penguin.ooo%2Fdocs-1f6feb?logo=readthedocs&logoColor=white" alt="文档" /></a>
<a href="https://penguin.ooo/blog"><img src="https://img.shields.io/badge/%E5%8D%9A%E5%AE%A2-penguin.ooo%2Fblog-1f6feb?logo=rss&logoColor=white" alt="博客" /></a>
</p>
<p align="center">
<picture>
<source media="(prefers-color-scheme: dark)" srcset="packages/landing/src/assets/shots/chat-zh-dark.webp" />
<img src="packages/landing/src/assets/shots/chat-zh-light.webp" alt="PenguinHarness Web App——多 Session 对话与实时流式工具调用" width="920" />
</picture>
<a href="https://discord.gg/eFHKqqcU3D"><img src="https://img.shields.io/badge/Discord-%E5%8A%A0%E5%85%A5%E8%AE%A8%E8%AE%BA-5865F2?logo=discord&logoColor=white" alt="Discord" /></a>
<a href="https://x.com/code_hiyouga"><img src="https://img.shields.io/badge/X-code%5Fhiyouga-000000?logo=x&logoColor=white" alt="X(Twitter)" /></a>
<a href="https://github.com/Prism-Shadow/penguin-harness-community/blob/main/wechat/group.jpg"><img src="https://img.shields.io/badge/%E5%BE%AE%E4%BF%A1-%E4%BA%A4%E6%B5%81%E7%BE%A4-07C160?logo=wechat&logoColor=white" alt="微信群" /></a>
</p>
---
<p align="center"><a href="README.md">English</a> | 简体中文</p>
## 为什么选择 PenguinHarness
- **Simplest Is the Best**——在干净的底层接口之上刻意保持极简的工具集:更少的工具调用、更少的 Token,高效完成复杂任务。
- **Harness for Building Agents**——基于 PenguinHarness SDK,由一个 Agent 从零开始为你自主构建完整的 Agent 应用。
- **Harness for Recursive Self-Improvement**——基于 PenguinHarness Skills,Agent 评估并优化自己:跑 Benchmark、找失分点、产出 N+1 版本,每轮之前先做快照。
- **本地优先且轻量**——100% 开源,一颗 CPU 即可运行,数据不出机器;经统一网关可接入 1000+ 在线与本地模型。
- **全量可观测**——每次请求、工具调用与审批决策都以追加方式写入 Trace,任何 Session 均可从 Trace 恢复。
三个递进的理由——从任务效果,到构建方式,再到进化能力。
## 快速开始
### 1. 🏆 效果同级,成本低一到两个数量级
一行命令安装(Linux / macOS,x64 / arm64,内嵌 Node 运行时,解压即用):
刻意精简的工具集配合干净的底层接口:更少的工具调用、更少的 Token,对 DeepSeek 等开放模型深度适配。各自搭配常用模型、同一批任务,正面对比:
```bash
curl -fsSL https://github.com/Prism-Shadow/penguin-harness/releases/latest/download/install.sh | sh
<p align="center">
<picture>
<source media="(prefers-color-scheme: dark)" srcset="assets/readme/benchmark-dark.svg" />
<img src="assets/readme/benchmark-light.svg" alt="Benchmark:PenguinHarness 在数据分析题库准确率最高、编程题库与 OpenAI Codex 持平,成本仅为两者的零头" width="920" />
</picture>
</p>
**数据分析准确率最高——成本只有 Claude Code 的 1/70。**
### 2. ⚡ 一句话,让 Agent 构建 Agent 应用
输入一句话,Agent 为你构建完整的 Agent 应用——脚手架、代码、运行说明,一步到位:
```text
收集 https://github.com/ericbuess/claude-code-docs 的文档,做一个化身 Claude Code 配置专家、回答带来源引用的 RAG 问答应用。
```
或经 npm 安装(需系统 Node >= 24,安装后的命令为 `penguin`):
这是做出来的成品——一个文档专家:检索增强、引用可点击直达原文、内置示例问题:
https://github.com/user-attachments/assets/604eb626-0a5d-4a62-87e3-14ebade1cd5f
**而生成整个 RAG 应用,仅消耗了 0.2 元($0.02)的 token——使用 DeepSeek V4 Pro 模型。**
### 3. 🧬 自进化,越用越强
借助 PenguinHarness 技能库,Agent 自己评估、自己优化:跑 Benchmark、找失分点、发布 N+1 版——每轮之前自动快照,每个请求都可在轨迹观测中回放。
https://github.com/user-attachments/assets/aec49ae9-b743-467b-b247-37bedfeaa36e
## 支持的模型
| 模型 | 可用供应商 |
| ---------------- | -------------------------------------------------------------------------------- |
| DeepSeek V4 | DeepSeek, OpenRouter, Fireworks AI, SiliconFlow, Qwen Token Plan |
| Kimi K3 | OpenRouter, Qwen Pay-As-You-Go |
| Kimi K2.6 | Moonshot AI |
| GLM 5.2 | Z.AI, OpenRouter, Fireworks AI, SiliconFlow, Qwen Token Plan, Qwen Pay-As-You-Go |
| Hunyuan 3 | OpenRouter |
| Qwen 3.8 Max | Qwen Token Plan(预览) |
| GPT 5.5 | OpenAI, OpenRouter |
| Gemini 3.5 Flash | Google Gemini, OpenRouter |
| Claude Opus 4.8 | Anthropic, OpenRouter |
只要是 OpenAI 协议的端点都可以接入:从上表选择预置,或用自定义端点连接 1000+ 在线与本地模型。
## 系统需求
| 需求项 | 支持情况 |
| -------- | ------------------------------------------------- |
| 操作系统 | Linux、macOS |
| 架构 | x64、arm64 |
| 运行时 | 一行安装器自带(经 npm 安装需 Node >= 24) |
| 模型 | 至少一个模型的 API key |
## 安装
### 🌐 Web 应用——面向人
🚀 一行安装,启动完整体验(多会话对话、Agent / 技能 / 模型管理、用量统计、轨迹观测、评估中心):
```bash
npm install -g @prismshadow/penguin-cli
curl -fsSL https://penguin.ooo/install.sh | sh
penguin web # 启动服务并打开 http://127.0.0.1:7364(首次登录:admin / penguin-2026)
```
然后启动 Web App,或直接留在终端:
📦 或经 npm 安装:`npm install -g @prismshadow/penguin-cli`。在应用内模型页配置模型后即可对话。
### 🤖 CLI 与 SDK——面向 Agent
同一引擎、可脚本化——为被 Agent 驱动而生(以及让 Agent 构建 Agent):
```bash
penguin web # 启动服务并打开 http://127.0.0.1:7364(初始账号 admin / admin123)
penguin server # 同一服务,无头运行
# 先配置一次模型(也可在 Web 的模型页完成)
penguin config model add --model-id deepseek-v4-pro --api-key sk-... --set-default
penguin run -m "创建 hello.txt,内容为 Hello, Penguin" # 单次任务
penguin chat # 交互式 REPL(/compact、/exit,Ctrl-C 中断)
penguin config model add --provider deepseek --model-id deepseek-v4-pro --api-key sk-... --set-default
penguin run -m "Create hello.txt containing Hello, Penguin" # 单次任务
penguin chat # 交互式 REPL(/compact、/exit、Ctrl-C 中断)
penguin server # 无界面服务(与 Web 应用同一套 API)
```
直接使用 SDK:
```ts
import { createAgent, isCompleteModelMessage, userText } from "@prismshadow/penguin-core";
const agent = await createAgent({ agentId: "default_agent" });
const session = await agent.createSession({ workspaceDir: process.cwd() });
for await (const output of session.run([userText("创建 hello.txt 并写入 hi")], {
approve: async () => "allow", // 逐个工具审批
for await (const output of session.run([userText("Create hello.txt containing hi")], {
approve: async () => "allow", // 按工具调用逐个审批
})) {
if (isCompleteModelMessage(output) && output.payload.type === "text") {
console.log(output.payload.text);
@@ -85,45 +130,38 @@ for await (const output of session.run([userText("创建 hello.txt 并写入 hi"
}
```
## 仓库结构
## 路线图
pnpm monorepo(TypeScript,Node >= 24)。一次安装交付四层组件,共享同一数据目录(`~/.penguin/data`)与同一消息协议(OmniMessage):
- [ ] Benchmark 套件正式发布
- [ ] 桌面端应用
- [ ] Windows 系统支持
- 更多规划,敬请期待……
| 目录 | 包名 | 职责 |
| --- | --- | --- |
| [`packages/core`](packages/core) | `@prismshadow/penguin-core` | SDK 与引擎:ReAct 循环、OmniMessage 协议、LLM/Environment 接口契约、Agent State、Trace |
| [`packages/cli`](packages/cli) | `@prismshadow/penguin-cli` | `penguin` 命令:REPL、单次运行、模型与 Vault 配置、服务启动 |
| [`packages/server`](packages/server) | `@prismshadow/penguin-server` | Web 服务端:HTTP API + SSE 流式、多用户认证、Project 授权、用量统计 |
| [`packages/web`](packages/web) | `@prismshadow/penguin-web` | Web App:多 Session 对话、Agent/技能/模型管理、Trace 观测、评估中心 |
| [`packages/skills`](packages/skills) | `@prismshadow/penguin-skills` | 内置技能库(Agent 创建、Benchmark 设计、评估、优化等) |
| [`packages/landing`](packages/landing) | — | 产品落地页(本仓库官网) |
| [`packages/docs`](packages/docs) | — | 文档站(双语,部署于 `/docs/` 路径) |
职责按事实来源划分:**SDK** 负责协议与执行(消息解析、运行循环、工具),**Server** 负责多用户运行时(认证、SSE 流式、定时任务),`~/.penguin/data` 下的**文件层**承载一切可编辑与被记录的状态(Prompt、Skill、密钥、Trace)。逐项对应表见[架构总览 → 职责划分](https://penguin.ooo/docs/architecture)。
## 文档
文档站覆盖使用与设计两个层面:[产品介绍](https://penguin.ooo/docs/) · [快速开始](https://penguin.ooo/docs/quickstart) · [架构总览](https://penguin.ooo/docs/architecture) · [OmniMessage 协议](https://penguin.ooo/docs/omni-message) · [接口契约](https://penguin.ooo/docs/interfaces) · [Agent 运行循环](https://penguin.ooo/docs/agent-loop) · [CLI 参考](https://penguin.ooo/docs/cli) · [Server API](https://penguin.ooo/docs/server-api) · [配置参考](https://penguin.ooo/docs/configuration)
每页文档都带「复制 Markdown」按钮,可直接粘贴进模型上下文。
## 本地开发
## 参与开发
```bash
pnpm install
pnpm build # 先构建:core 的导出指向 dist/
pnpm typecheck
pnpm test
pnpm dev:server # 服务端 127.0.0.1:7364
pnpm dev:web # Web App(Vite)127.0.0.1:7365,/api 代理到服务端
pnpm dev:docs # 文档站(Vite)127.0.0.1:7367
BASE_PATH=/ pnpm build:site # 按 Pages 部署的方式组装 落地页 + 文档
pnpm install && pnpm build # 先构建:core 的导出指向 dist/
pnpm dev # 服务端 + Web 一起启动(带前缀日志,依赖只构建一次)
```
开发态模型凭据可复制 `.env.example` 为 `.env` 填写。E2E 测试走真实模型(`pnpm test:e2e`,需要 `DEEPSEEK_API_KEY`)。
完整工作区指南见 [CONTRIBUTING.md](CONTRIBUTING.md):开发命令、质量门禁、仓库结构与 changelog 规则。
## 许可证
## 引用
如果 PenguinHarness 对你的研究有帮助,请引用:
```bibtex
@software{penguinharness2026,
author = {{PrismShadow Team}},
title = {PenguinHarness: Efficient Self-Improving Harness for Everyone},
year = {2026},
url = {https://github.com/Prism-Shadow/penguin-harness},
license = {Apache-2.0}
}
```
## 协议
[Apache-2.0](LICENSE) © 2026 Prism Shadow
由 [LlamaFactory](https://github.com/hiyouga/LlamaFactory) 作者 [Yaowei Zheng](https://github.com/hiyouga)、[PrismShadow AI Team](https://github.com/Prism-Shadow) 与 [Fable 5](https://www.anthropic.com/news/claude-fable-5-mythos-5) 共同用 ❤️ 构建。
+48
View File
@@ -0,0 +1,48 @@
<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 920 368" width="920" height="368" role="img" aria-label="Benchmark: PenguinHarness vs Claude Code vs OpenAI Codex on two suites — comparable accuracy at a small fraction of the cost">
<text x="148" y="36" font-size="12" font-weight="400" fill="#898781" text-anchor="start" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">Accuracy · suite total · higher is better</text>
<text x="596" y="36" font-size="12" font-weight="400" fill="#898781" text-anchor="start" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">Total cost (USD) · lower is better</text>
<text x="24" y="68" font-size="13" font-weight="600" fill="#f0f3f6" text-anchor="start" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">Data analysis — 15 tasks, single run</text>
<rect x="147" y="78" width="1" height="92" fill="#383835"/>
<text x="136" y="94.5" font-size="12" font-weight="600" fill="#f0f3f6" text-anchor="end" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">PenguinHarness</text>
<path d="M149 82 h172.0 a4 4 0 0 1 4 4 v8 a4 4 0 0 1 -4 4 h-172.0 z" fill="#3987e5"/>
<text x="337" y="94.5" font-size="12" font-weight="600" fill="#f0f3f6" text-anchor="start" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">66.67%</text>
<text x="136" y="124.5" font-size="12" font-weight="400" fill="#c3c2b7" text-anchor="end" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">Claude Code</text>
<path d="M149 112 h136.8 a4 4 0 0 1 4 4 v8 a4 4 0 0 1 -4 4 h-136.8 z" fill="#6b6a64"/>
<text x="301.7841607919604" y="124.5" font-size="12" font-weight="400" fill="#f0f3f6" text-anchor="start" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">53.33%</text>
<text x="136" y="154.5" font-size="12" font-weight="400" fill="#c3c2b7" text-anchor="end" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">OpenAI Codex</text>
<path d="M149 142 h136.8 a4 4 0 0 1 4 4 v8 a4 4 0 0 1 -4 4 h-136.8 z" fill="#6b6a64"/>
<text x="301.7841607919604" y="154.5" font-size="12" font-weight="400" fill="#f0f3f6" text-anchor="start" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">53.33%</text>
<rect x="595" y="78" width="1" height="92" fill="#383835"/>
<text x="584" y="94.5" font-size="12" font-weight="600" fill="#f0f3f6" text-anchor="end" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">PenguinHarness</text>
<path d="M597 82 h10.0 a4 4 0 0 1 4 4 v8 a4 4 0 0 1 -4 4 h-10.0 z" fill="#3987e5"/>
<text x="623" y="94.5" font-size="12" font-weight="600" fill="#f0f3f6" text-anchor="start" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">$0.55</text>
<text x="584" y="124.5" font-size="12" font-weight="400" fill="#c3c2b7" text-anchor="end" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">Claude Code</text>
<path d="M597 112 h172.0 a4 4 0 0 1 4 4 v8 a4 4 0 0 1 -4 4 h-172.0 z" fill="#6b6a64"/>
<text x="785" y="124.5" font-size="12" font-weight="400" fill="#f0f3f6" text-anchor="start" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">$38.48</text>
<text x="584" y="154.5" font-size="12" font-weight="400" fill="#c3c2b7" text-anchor="end" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">OpenAI Codex</text>
<path d="M597 142 h84.8 a4 4 0 0 1 4 4 v8 a4 4 0 0 1 -4 4 h-84.8 z" fill="#6b6a64"/>
<text x="697.7945961503353" y="154.5" font-size="12" font-weight="400" fill="#f0f3f6" text-anchor="start" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">$19.41</text>
<text x="24" y="210" font-size="13" font-weight="600" fill="#f0f3f6" text-anchor="start" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">Coding — 40 tasks × 2 runs</text>
<rect x="147" y="220" width="1" height="92" fill="#383835"/>
<text x="136" y="236.5" font-size="12" font-weight="600" fill="#f0f3f6" text-anchor="end" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">PenguinHarness</text>
<path d="M149 224 h141.4 a4 4 0 0 1 4 4 v8 a4 4 0 0 1 -4 4 h-141.4 z" fill="#3987e5"/>
<text x="306.3913043478261" y="236.5" font-size="12" font-weight="600" fill="#f0f3f6" text-anchor="start" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">71.25%</text>
<text x="136" y="266.5" font-size="12" font-weight="400" fill="#c3c2b7" text-anchor="end" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">Claude Code</text>
<path d="M149 254 h172.0 a4 4 0 0 1 4 4 v8 a4 4 0 0 1 -4 4 h-172.0 z" fill="#6b6a64"/>
<text x="337" y="266.5" font-size="12" font-weight="400" fill="#f0f3f6" text-anchor="start" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">86.25%</text>
<text x="136" y="296.5" font-size="12" font-weight="400" fill="#c3c2b7" text-anchor="end" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">OpenAI Codex</text>
<path d="M149 284 h141.4 a4 4 0 0 1 4 4 v8 a4 4 0 0 1 -4 4 h-141.4 z" fill="#6b6a64"/>
<text x="306.3913043478261" y="296.5" font-size="12" font-weight="400" fill="#f0f3f6" text-anchor="start" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">71.25%</text>
<rect x="595" y="220" width="1" height="92" fill="#383835"/>
<text x="584" y="236.5" font-size="12" font-weight="600" fill="#f0f3f6" text-anchor="end" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">PenguinHarness</text>
<path d="M597 224 h10.0 a4 4 0 0 1 4 4 v8 a4 4 0 0 1 -4 4 h-10.0 z" fill="#3987e5"/>
<text x="623" y="236.5" font-size="12" font-weight="600" fill="#f0f3f6" text-anchor="start" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">$3.81</text>
<text x="584" y="266.5" font-size="12" font-weight="400" fill="#c3c2b7" text-anchor="end" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">Claude Code</text>
<path d="M597 254 h113.5 a4 4 0 0 1 4 4 v8 a4 4 0 0 1 -4 4 h-113.5 z" fill="#6b6a64"/>
<text x="726.5332606324973" y="266.5" font-size="12" font-weight="400" fill="#f0f3f6" text-anchor="start" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">$146.97</text>
<text x="584" y="296.5" font-size="12" font-weight="400" fill="#c3c2b7" text-anchor="end" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">OpenAI Codex</text>
<path d="M597 284 h172.0 a4 4 0 0 1 4 4 v8 a4 4 0 0 1 -4 4 h-172.0 z" fill="#6b6a64"/>
<text x="785" y="296.5" font-size="12" font-weight="400" fill="#f0f3f6" text-anchor="start" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">$220.08</text>
<text x="24" y="338" font-size="12" font-weight="400" fill="#898781" text-anchor="start" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">Each harness runs the model it is normally paired with: PenguinHarness on DeepSeek V4 Pro, Claude Code on Claude Opus 4.8,</text>
<text x="24" y="356" font-size="12" font-weight="400" fill="#898781" text-anchor="start" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">OpenAI Codex on GPT-5.5. Accuracy, Tokens and cost are suite totals at official pricing.</text>
</svg>

After

Width:  |  Height:  |  Size: 6.9 KiB

+48
View File
@@ -0,0 +1,48 @@
<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 920 368" width="920" height="368" role="img" aria-label="Benchmark: PenguinHarness vs Claude Code vs OpenAI Codex on two suites — comparable accuracy at a small fraction of the cost">
<text x="148" y="36" font-size="12" font-weight="400" fill="#898781" text-anchor="start" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">Accuracy · suite total · higher is better</text>
<text x="596" y="36" font-size="12" font-weight="400" fill="#898781" text-anchor="start" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">Total cost (USD) · lower is better</text>
<text x="24" y="68" font-size="13" font-weight="600" fill="#1f2328" text-anchor="start" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">Data analysis — 15 tasks, single run</text>
<rect x="147" y="78" width="1" height="92" fill="#c3c2b7"/>
<text x="136" y="94.5" font-size="12" font-weight="600" fill="#1f2328" text-anchor="end" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">PenguinHarness</text>
<path d="M149 82 h172.0 a4 4 0 0 1 4 4 v8 a4 4 0 0 1 -4 4 h-172.0 z" fill="#2a78d6"/>
<text x="337" y="94.5" font-size="12" font-weight="600" fill="#1f2328" text-anchor="start" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">66.67%</text>
<text x="136" y="124.5" font-size="12" font-weight="400" fill="#52514e" text-anchor="end" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">Claude Code</text>
<path d="M149 112 h136.8 a4 4 0 0 1 4 4 v8 a4 4 0 0 1 -4 4 h-136.8 z" fill="#898781"/>
<text x="301.7841607919604" y="124.5" font-size="12" font-weight="400" fill="#1f2328" text-anchor="start" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">53.33%</text>
<text x="136" y="154.5" font-size="12" font-weight="400" fill="#52514e" text-anchor="end" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">OpenAI Codex</text>
<path d="M149 142 h136.8 a4 4 0 0 1 4 4 v8 a4 4 0 0 1 -4 4 h-136.8 z" fill="#898781"/>
<text x="301.7841607919604" y="154.5" font-size="12" font-weight="400" fill="#1f2328" text-anchor="start" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">53.33%</text>
<rect x="595" y="78" width="1" height="92" fill="#c3c2b7"/>
<text x="584" y="94.5" font-size="12" font-weight="600" fill="#1f2328" text-anchor="end" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">PenguinHarness</text>
<path d="M597 82 h10.0 a4 4 0 0 1 4 4 v8 a4 4 0 0 1 -4 4 h-10.0 z" fill="#2a78d6"/>
<text x="623" y="94.5" font-size="12" font-weight="600" fill="#1f2328" text-anchor="start" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">$0.55</text>
<text x="584" y="124.5" font-size="12" font-weight="400" fill="#52514e" text-anchor="end" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">Claude Code</text>
<path d="M597 112 h172.0 a4 4 0 0 1 4 4 v8 a4 4 0 0 1 -4 4 h-172.0 z" fill="#898781"/>
<text x="785" y="124.5" font-size="12" font-weight="400" fill="#1f2328" text-anchor="start" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">$38.48</text>
<text x="584" y="154.5" font-size="12" font-weight="400" fill="#52514e" text-anchor="end" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">OpenAI Codex</text>
<path d="M597 142 h84.8 a4 4 0 0 1 4 4 v8 a4 4 0 0 1 -4 4 h-84.8 z" fill="#898781"/>
<text x="697.7945961503353" y="154.5" font-size="12" font-weight="400" fill="#1f2328" text-anchor="start" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">$19.41</text>
<text x="24" y="210" font-size="13" font-weight="600" fill="#1f2328" text-anchor="start" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">Coding — 40 tasks × 2 runs</text>
<rect x="147" y="220" width="1" height="92" fill="#c3c2b7"/>
<text x="136" y="236.5" font-size="12" font-weight="600" fill="#1f2328" text-anchor="end" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">PenguinHarness</text>
<path d="M149 224 h141.4 a4 4 0 0 1 4 4 v8 a4 4 0 0 1 -4 4 h-141.4 z" fill="#2a78d6"/>
<text x="306.3913043478261" y="236.5" font-size="12" font-weight="600" fill="#1f2328" text-anchor="start" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">71.25%</text>
<text x="136" y="266.5" font-size="12" font-weight="400" fill="#52514e" text-anchor="end" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">Claude Code</text>
<path d="M149 254 h172.0 a4 4 0 0 1 4 4 v8 a4 4 0 0 1 -4 4 h-172.0 z" fill="#898781"/>
<text x="337" y="266.5" font-size="12" font-weight="400" fill="#1f2328" text-anchor="start" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">86.25%</text>
<text x="136" y="296.5" font-size="12" font-weight="400" fill="#52514e" text-anchor="end" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">OpenAI Codex</text>
<path d="M149 284 h141.4 a4 4 0 0 1 4 4 v8 a4 4 0 0 1 -4 4 h-141.4 z" fill="#898781"/>
<text x="306.3913043478261" y="296.5" font-size="12" font-weight="400" fill="#1f2328" text-anchor="start" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">71.25%</text>
<rect x="595" y="220" width="1" height="92" fill="#c3c2b7"/>
<text x="584" y="236.5" font-size="12" font-weight="600" fill="#1f2328" text-anchor="end" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">PenguinHarness</text>
<path d="M597 224 h10.0 a4 4 0 0 1 4 4 v8 a4 4 0 0 1 -4 4 h-10.0 z" fill="#2a78d6"/>
<text x="623" y="236.5" font-size="12" font-weight="600" fill="#1f2328" text-anchor="start" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">$3.81</text>
<text x="584" y="266.5" font-size="12" font-weight="400" fill="#52514e" text-anchor="end" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">Claude Code</text>
<path d="M597 254 h113.5 a4 4 0 0 1 4 4 v8 a4 4 0 0 1 -4 4 h-113.5 z" fill="#898781"/>
<text x="726.5332606324973" y="266.5" font-size="12" font-weight="400" fill="#1f2328" text-anchor="start" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">$146.97</text>
<text x="584" y="296.5" font-size="12" font-weight="400" fill="#52514e" text-anchor="end" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">OpenAI Codex</text>
<path d="M597 284 h172.0 a4 4 0 0 1 4 4 v8 a4 4 0 0 1 -4 4 h-172.0 z" fill="#898781"/>
<text x="785" y="296.5" font-size="12" font-weight="400" fill="#1f2328" text-anchor="start" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">$220.08</text>
<text x="24" y="338" font-size="12" font-weight="400" fill="#898781" text-anchor="start" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">Each harness runs the model it is normally paired with: PenguinHarness on DeepSeek V4 Pro, Claude Code on Claude Opus 4.8,</text>
<text x="24" y="356" font-size="12" font-weight="400" fill="#898781" text-anchor="start" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">OpenAI Codex on GPT-5.5. Accuracy, Tokens and cost are suite totals at official pricing.</text>
</svg>

After

Width:  |  Height:  |  Size: 6.9 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 70 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 76 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 64 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 71 KiB

@@ -0,0 +1,84 @@
# Landing site
The penguin.ooo landing page: story, structure, animations, navigation, and the install domain.
## Landing, README, and blog retell one story, with penguin.ooo/install.sh
The marketing surfaces now tell the same 1x/100x story, persuade through one numbered "Why" section, and install from the site's own domain.
- **penguin.ooo/install.sh** — the landing site ships a thin `public/install.sh` that forwards to the latest GitHub release installer (GitHub Pages cannot serve real redirects), so `curl -fsSL https://penguin.ooo/install.sh | sh` works everywhere; the hero command, README (en/zh), docs installation/quickstart (en/zh), and blog posts all switch to it.
- **Landing** — the hero replaces the rotating-word headline with the two-line story ("With LangChain, you build agents by hand — at 1x; with PenguinHarness, agents build agents — at 100x", the 100x fragment in brand color) plus the "zero-code Harness CLI and Web UI, 1000+ models" subtitle. One numbered Why section (1 benchmark suites + DeepSeek tuning, 2 one-sentence build demo with the real RAG shot, 3 self-evolution loop with a "demo video coming soon" pill) absorbs the former Pillars / Showcase / Benchmark / SelfImprove sections; UI screenshots leave the home page; nav anchors become why / quickstart / contract / features (reason 1 keeps the #benchmark anchor). The blog already shares the landing Nav, theme, and language settings.
- **README (en/zh)** — subtitle now says "Harness CLI and Web UI"; website/docs/blog/community links render as shields.io badges; the top screenshot and the Changelog-Blog-Docs section are gone; the three reasons sit numbered with emoji under "Why PenguinHarness"; the models table keeps only a one-line "any OpenAI-protocol endpoint works" note; requirements become a table; install sections gain emoji; the roadmap adds a desktop app and Windows support; the credits link LlamaFactory, the PrismShadow AI Team, and Fable 5.
- **Blog** — the introduction post (en/zh) leads with the same story line and phrasing.
## Landing back to the classic structure, with tabbed features/cases and a docs-matching nav
The previous landing overhaul went too far; the classic section set returns, with the new material folded in as tabs and side sections instead of replacements.
- **Restored as they were**: the rotating hero headline ("Efficient Self-Improving Harness for Developers/Enterprises"), the three pillars, the self-improvement loop (now with a "demo video coming soon" pill), Quickstart, the standalone Benchmark section, CONTRACT.md, and Security.
- **LangChain comparison moved, not deleted**: the 1x/100x story lives in its own compact section after the pillars — two cards (LangChain, hand-built, 1x, de-emphasized vs PenguinHarness, agents building agents, 100x, brand-emphasized) with the story sentence as the subtitle.
- **Features become switchable tabs**: all nine feature descriptions remain, each as a tab with its icon; the three views with real captures — multi-session chat, Trace view, Agent evaluation — show their locale/theme-matched screenshot in the active panel, bringing the Trace and benchmark shots back onto the page.
- **Use cases become tabs too**: a Cases section with a tab bar holds the RAG case only for now (prompt + captured result); future cases append as tabs.
- **Community section added at the end**: Discord / X / WeChat group / GitHub as outbound cards after the CTA.
- **Docs nav now mirrors the landing nav exactly** — same link row (Highlights / Quick start / Benchmark / CONTRACT.md / Features / Blog / Docs) with the same sliding hover pill, anchoring into the landing one level up, so the two sites link seamlessly; the old standalone "Website" link is absorbed by the row.
- **Launch blog post back to its original skeleton**: the "Why PenguinHarness" three-pillar bullets, the benchmark section wording, and the closing "start right now" paragraph return, with the GDPevo origin story, images, models table, roadmap, and community call-to-action kept as additions rather than replacements.
## Landing trace screenshots: same opened timeline in English and Chinese
The English trace screenshots showed the empty "select a Session" state while the Chinese ones showed a full trace; all four now capture the same opened trace with stats and an execution timeline containing tool calls.
- The capture script navigated to `/traces?sessionId=...` — but the traces page only honors the session deep link when `agentId=` is present (the product's own links always carry both), so selection relied on a fragile title click that silently failed for English via `.catch()`. The script now uses the canonical `?agentId=default_agent&sessionId=...` deep link and waits for the timeline's `exec_command` lanes before shooting, so an empty capture fails loudly instead of shipping.
- The scripted English session title was 31 chars and core clips titles at `TITLE_MAX_CHARS = 30`, producing "…Agent ap" in the shots; the mock title is now "Build a data-analysis Agent" (27 chars).
- Regenerated `traces-{en,zh}-{light,dark}.webp`; chat and benchmark shots are unchanged.
## Landing polish: sliding announcements, construction animation, trend value labels, feature split
A polish pass over the landing page's motion and grouping, plus web-first usage guidance in the blog.
- **Announcement bar** — light brand-tinted background; both entries now link to blog posts (the new-models entry to the launch post, the credits entry to the AMD post); switching SLIDES horizontally (auto-advance, hover pause) with dot indicators instead of arrows; the trailing arrow icon is gone.
- **LangChain comparison** — a looping "construction" animation on a shared cycle: the LangChain card lays one gray block at a time and never tops out before the cycle resets, while the PenguinHarness card raises a whole brand-blue skyline in under two seconds and holds it. Reduced-motion shows both skylines complete.
- **Self-improvement trends** — the three outcome charts are now driven by one rAF clock, and a value label rides each curve's moving head (score climbing, cost and time falling), changing as the line draws; reduced-motion shows the finished line with its final value.
- **Features split by capture** — only the three features with real screenshots (multi-session chat, trace view, agent evaluation) form the tab bar; the remaining six sit below as the classic card grid, closed with an "and more…" card.
- **Blog usage goes web-only** — the launch post's getting-started and the AMD post's setup now guide through `penguin web` and the Models page only; the CLI config/run alternatives are removed to funnel users into the Web UI.
## Announcement carousel: one-direction slide, arrow kept, dots removed
The announcement bar keeps the small arrow marking each entry as a click-through link, drops the dot switch buttons on the right, and auto-rotates by sliding in ONE direction only: a clone of the first slide follows the last, and once the clone is fully in view the track snaps back without animation — the bar never visibly slides backwards. Hovering still pauses the rotation.
## Feature tabs drop the description card
The three screenshot-backed feature tabs no longer render an icon card above the capture; the feature's one-line description now sits as centered text directly above the image (the tab chip already carries the title and icon). The no-capture features keep their card grid below.
## Trend value labels carry units with animated decimals
The self-improvement loop's three outcome labels now read as real measurements while they ride the curve head: score as a percentage with one decimal (79.0%), cost in dollars with two decimals ($0.25), and time in seconds with one decimal (83.0s) — the decimal part animates continuously along with the line.
## Fix the nav hover highlight sweeping in from the edge
The landing and docs navs' sliding hover pill animated from its hidden state at the nav's left edge, so the first hover sent a gray pill flying across the whole link row (and it slid back on leave).
The pill now appears IN PLACE under the first link it lands on — the position jumps with only the fade animating — slides while moving between links, and fades out where it is when the pointer leaves. Applied identically to the landing nav and the landing-parity docs nav.
## Merge main to restore the persistent nav active highlight
The branch predated main's PR #6 ("correct navigation state and anchor scrolling"), so it was missing the nav's persistent active highlight — the black chip that marks the current section or route — and would have reverted that fix on merge. Merging origin/main brings it back and reconciles it with the branch's nav work.
- Landing nav: section links route through `/#id` again with `getActiveNavItem` tracking (new `lib/nav-state.ts` + tests), the active link keeps its black chip with `aria-current`, and the branch's in-place hover pill behavior is preserved on top.
- Also restored from #6: anchor targets scroll below the sticky header via `.section-anchor` (ids moved onto the section's inner div), footer/hero/CTA anchors as router links, and the docs brand logo returning to the landing site.
- Docs nav (landing-parity) additionally marks its own "Docs" link as the current page with the same black chip.
## Nav highlight follows the live scroll position
The nav's black current-item chip previously updated only from the URL hash (i.e. on click); while scrolling, it stayed stale. A scroll-spy hook now measures the five section anchors on every scroll frame (rAF-throttled, document-level capture so any scrolling container works) and lights up the LAST section whose top has crossed the activation line under the sticky header — null above the first section, and in-between sections keep the previous anchor lit. On the home page the highlight is fully live; other routes keep route-based state (Blog).
## The Cases tab shows the finished RAG app with a localized prompt
The RAG case now displays the condensed claude-code-docs configuration-expert prompt from the strings catalog (localized zh/en) and the finished-product app screenshot matched to the visitor's locale and theme, replacing the shared English prompt and the PenguinHarness chat capture.
## The penguin sled game joins the cases, with a $0.02 cost hook on the RAG demo
The draft-screen game example becomes a cute Antarctic penguin sledding game, lands on the landing page as a second case after RAG, and every RAG demo now leads with how little it cost to generate.
- The draft-screen example card (zh/en) swaps the 2D motocross runner for a penguin sledding game: Space to jump the rocks on the ice, speed and difficulty ramping up, live scoring, one-click restart, cute cartoon look — the full detailed prompt still gets submitted as-is.
- The landing Cases section gains a "Penguin sled game" tab after the RAG one, showing the finished play screen per language and theme (light = polar day, dark = polar night with an aurora) from the new dependency-free `penguin-game-mockup.html` + `capture-game-mockup.mjs` pipeline.
- README (en/zh), the launch blog post (en/zh), and the landing RAG case now carry an emphasized hook under the finished-product shot: generating the entire RAG app burned just $0.02 (¥0.2) of tokens on DeepSeek V4 Pro.
@@ -0,0 +1,259 @@
# Model catalog, Models page, and credential handling
Preset provider groups, catalog entries and ordering, and the Models page features built around them.
## Add the Qwen Token Plan provider group to the model catalog
The built-in catalog gains a Qwen Token Plan subscription gateway group (OpenAI-compatible,
preset base URL) with five models — qwen3.8-max-preview, qwen3.7-max, qwen3.7-plus, glm-5.2,
and deepseek-v4-pro — plus a custom provider logo.
## Details
- New provider `qwen-token-plan` ("Qwen Token Plan"), placed with the gateway cluster after
SiliconFlow: OpenAI-compatible endpoint preset to
`https://token-plan.cn-beijing.maas.aliyuncs.com/compatible-mode/v1`, API-key page
`https://platform.qianwenai.com/pricing/token-plan`, model-id docs page
`https://platform.qianwenai.com/docs/token-plan/personal/token-plan-personal-overview`;
env fallback is `OPENAI_API_KEY`/`OPENAI_BASE_URL` like the other gateways.
- Five catalog entries with `client_type: openai` and the inlined base URL. Vision flags per
the plan's supported-model table (qwen3.8-max-preview and qwen3.7-plus see images; the
rest do not). Pricing and context windows come from each model's page at
`www.qianwenai.com/models/<id>` (official CNY list prices; limited-time promotions are not
stored): qwen3.7-max ¥2.4/¥12/¥36, qwen3.7-plus ¥0.4/¥2/¥8, glm-5.2 ¥2/¥8/¥28,
deepseek-v4-pro ¥1/¥12/¥24 (cache-hit/input/output per M tokens), windows 1M (glm-5.2:
1.04M). qwen3.8-max-preview is preview-only with a quota-multiplier promotion and no
per-token list price, so it alone carries no pricing (costs read as 0, same as unpriced
user models); the pricing invariant in the catalog tests is scoped to that one entry.
- Catalog invariant updated: bare model ids may now repeat across providers (the gateway
resells vendor models under their upstream ids, e.g. `glm-5.2` / `deepseek-v4-pro`);
uniqueness is the `(provider, model_id)` pair, matching the catalog's sole lookup key.
- New provider logo: the official Qwen wordmark (icon + lettering) from the brand SVG,
gradient fills flattened to currentColor monochrome, coordinates rounded to 2dp.
- Docs (models.en/zh) provider table and gateway notes updated.
## Add the Qwen Pay-As-You-Go provider group
A pay-per-token gateway group below Qwen Token Plan (DashScope's OpenAI-compatible
endpoint) with four preset models — qwen3.7-max, qwen3.7-plus, and the resold
vendor-prefixed kimi/kimi-k3 and ZHIPU/GLM-5.2 — priced from each model's official page.
## Details
- New provider `qwen-pay-as-you-go` ("Qwen Pay-As-You-Go") right after Qwen Token Plan in
the gateway cluster: preset base URL `https://dashscope.aliyuncs.com/compatible-mode/v1`,
API-key link `https://platform.qianwenai.com/docs/api-reference/preparation/api-key`,
models page `https://www.qianwenai.com/models`; `OPENAI_*` env fallback like the other
gateways.
- Four entries (`client_type: openai` + inlined endpoint), official CNY list prices and
specs from each model's page: kimi/kimi-k3 ¥2/¥20/¥100 (1.04M, vision), qwen3.7-max
¥2.4/¥12/¥36 (1M), ZHIPU/GLM-5.2 ¥2/¥8/¥28 (1.04M), qwen3.7-plus ¥0.4/¥2/¥8 (1M,
vision). Resold third-party models keep their vendor-prefixed upstream ids.
- The group shares the Qwen emblem (glyph extracted into a shared constant), and
`modelHomepageUrl` URL-encodes the slash-prefixed ids
(`.../models/ZHIPU%2FGLM-5.2`). Docs (models.en/zh) provider tables and gateway notes
updated.
## Add the Fireworks AI provider group
An OpenAI-protocol gateway group below OpenRouter with five preset models (GLM-5.2,
Kimi K2.7 Code, DeepSeek V4 Pro, MiniMax M3, DeepSeek V4 Flash), priced from each model's
Fireworks page.
## Details
- New provider `fireworks` ("Fireworks AI") right after OpenRouter in the gateway cluster:
preset base URL `https://api.fireworks.ai/inference/v1`, API-key page
`https://app.fireworks.ai/settings/users/api-keys`, models page
`https://app.fireworks.ai/models`; `OPENAI_*` env fallback like the other gateways.
- Five entries (`client_type: openai` + inlined endpoint) with standard-serverless USD
pricing (cached input / uncached input / output per M tokens) and specs from each page:
glm-5p2 $0.14/$1.40/$4.40 (1M), kimi-k2p7-code $0.19/$0.95/$4.00 (262K, vision),
deepseek-v4-pro $0.15/$1.74/$3.48 (1M), minimax-m3 $0.06/$0.30/$1.20 (512K, vision),
deepseek-v4-flash $0.03/$0.14/$0.28 (1M). API model ids use Fireworks' full
`accounts/fireworks/models/<slug>` form (sent verbatim).
- `modelHomepageUrl` maps the `accounts/<owner>/models/<slug>` id to the model page
(`app.fireworks.ai/models/<owner>/<slug>`), falling back to the models listing for
nonconforming user-added ids. A simplified starburst glyph approximates the brand mark
(same approach as Z.AI). Docs (models.en/zh) provider tables and gateway notes updated.
## Expand the OpenRouter catalog with twelve models
The OpenRouter gateway group grows from 4 to 16 entries, adding the current flagship and
free tiers with pricing, context windows, and vision flags taken from each model's
OpenRouter page.
## Details
- Added (ordered by output price): anthropic/claude-fable-5 ($10/$50), openai/gpt-5.6-sol
($5/$30), openai/gpt-5.5 ($5/$30), anthropic/claude-opus-4.8 ($5/$25),
anthropic/claude-opus-4.7 ($5/$25), moonshotai/kimi-k3 ($3/$15), openai/gpt-5.6-terra
($2.50/$15), anthropic/claude-sonnet-5 ($2/$10), z-ai/glm-5.2 ($0.93/$3),
deepseek/deepseek-v4-pro ($0.435/$0.87), deepseek/deepseek-v4-flash ($0.09/$0.18), and
nvidia/nemotron-3-ultra-550b-a55b:free (input/output per M tokens; all 1M context).
- Vision per the pages: the Claude models, GPT-5.5, and Kimi K3 accept image input; the
rest do not.
- None of these pages list cache pricing, so cache_read carries the standard input price
(no discount). The :free tier stores a genuine $0 price — not "unknown" — so costs
correctly compute to 0; the catalog pricing invariant gains a free-tier case.
## Add Grok 4.5 to the OpenRouter catalog
`x-ai/grok-4.5` joins the OpenRouter group: $2 input / $6 output per M tokens (no cache
price listed, so cache_read carries the input price), 500K context, vision-capable —
inserted at its dictionary position.
## Add Gemini 3.5 Flash to the OpenRouter catalog
`google/gemini-3.5-flash` joins the OpenRouter group: $1.50 input / $9 output per M tokens
(no cache price listed, so cache_read carries the input price), 1M context, vision-capable
— at its dictionary position; the README model table's Gemini row now lists OpenRouter too.
## Order catalog models by dictionary, newer versions first
Within each provider group, catalog entries are now in dictionary order by model id, except
that newer versions of the same series come first (gpt-5.6-* before gpt-5.5,
claude-opus-4.8 before 4.7, glm-5.2 before glm-5) — precomputed by hand in the catalog
literal, with no runtime sorting anywhere.
## Details
- Every provider section of MODEL_CATALOG is hand-reordered: dictionary order
(case-insensitive) across families and tiers; within a version series, the newest version
block leads, tiers inside a version staying alphabetical. Section comments and the
exact-order test assertions are updated to match.
- The order flows everywhere in-group order is preserved: new Projects' preset config, the
models page cards, and the chat model dropdown (orderModelsLikeLibrary). Existing Project
configs keep their stored order — the sync-presets merge deliberately preserves local
positions.
## Catalog data: official Fireworks logo, two SiliconFlow models, per-model vendor pages
The Fireworks group now wears the official burst mark, SiliconFlow gains
moonshotai/Kimi-K2.7-Code and deepseek-ai/DeepSeek-V4-Flash, and Z.AI / Moonshot models
link to their per-model docs pages.
## Details
- Fireworks logo: the official three-stroke burst mark (viewBox 0 0 638 315, currentColor)
replaces the interim starburst approximation.
- SiliconFlow entries (official CNY pricing, dictionary position):
deepseek-ai/DeepSeek-V4-Flash ¥0.02/¥1/¥2 (1M) and moonshotai/Kimi-K2.7-Code
¥1.3/¥6.5/¥27 (262K, vision).
- modelHomepageUrl: zhipu -> `docs.z.ai/guides/llm/<model_id>`; moonshot ->
`platform.kimi.com/docs/pricing/chat-k<version-without-dot>` (kimi-k2.6 -> chat-k26),
nonconforming ids falling back to the group's models page.
## Add a sync-presets button to the Models page
A small owner-only button next to the Models page search box merges the built-in catalog
into the Project's model table: catalog entries missing locally are added, entries present
on both sides are reset to the catalog's fields, locally added models and API keys stay
untouched.
## Details
- Union semantics (`catalog-sync.ts`, pure and unit-tested): keyed by the
`(provider, model_id)` pair. Catalog-only entries are appended (gateway base URLs preset);
intersecting entries take the catalog's context window, pricing (including removal when
the catalog carries none, e.g. the unpriced preview model), protocol, base URL, and vision
flag — the catalog wins wherever the two differ; local-only models (including
user-defined groups) are kept verbatim and in place.
- Credentials are structurally untouched: merged rows submit no `apiKey` (the PUT
full-table replace keeps stored keys when the field is absent), and existing rows keep
their credential state; a user base-URL override on a preset model is reset to the
catalog's (the API-key carve-out is the only one).
- Feedback via toasts: "Presets synced: N added, M updated", or "already up to date"
without a PUT when nothing differs. Strings added to both locales.
- The Qwen Token Plan provider logo is trimmed to the official emblem only (the wordmark
lettering dropped), on a square viewBox.
## Model test no longer fails on thinking-only responses
Testing a reasoning-heavy model (e.g. qwen3.8-max-preview) failed with "OpenaiClient
returned no content other than thinking (finish_reason=\"length\")": the connectivity
probe's tiny output cap was burned entirely on thinking. The probe now counts a streamed
thinking-only ending as reachable — the endpoint, credential, and model id all
demonstrably work.
## Details
- The probe deliberately sends one "ping" with `maxTokens: 16` and thinking disabled
(single-digit token cost by design). Reasoning models behind OpenAI-compatible endpoints
can ignore the disabled thinking level, hit `finish_reason=length` with no text, and
AgentHub 0.4 raises `EmptyResponseError` — collapsed to a `malformed` outcome, which the
probe previously reported as a test failure.
- `testModel` now tracks whether genuine model content (thinking or text, partial or
complete) was streamed, and a `malformed` ending after streamed content passes the test;
timeouts, auth/parameter failures, and malformed endings with nothing received still
fail. The logic lives in two pure functions (`isProbeContent` / `probeVerdict`) with unit
tests, including the exact qwen3.8-max-preview case.
## Group speed test on the Models page
Each model group header gains an owner-only speed-test action: after a quota warning it
probes the group's models one at a time, measuring time-to-first-token and output rate, and
writes tone-colored badges (green / yellow / red) onto each card; the model-homepage link
moves from the card corner into the config dialog.
## Details
- Server: the model-test endpoint gains a `speed` flag — the probe's output cap rises from
16 to 64 tokens so a real streaming window exists, and the response now carries `ttftMs`
(request start -> first streamed content) and `tps` (output tokens over the streaming
window, from the completed stream's usage report; thinking-only endings carry TTFT but no
rate). The plain connectivity test is unchanged.
- Web: a gauge button on each group header (owner-only) opens a confirmation dialog warning
that one real request per model will consume API quota; on confirm the group is tested
**strictly sequentially** (concurrent probes trip provider rate limits), each result
landing on its card as it finishes. Badges: clock icon + ms for TTFT (green < 1s, yellow
<= 3s, red beyond), zap icon + tok/s for TPS (green >= 40, yellow >= 15, red below);
failures show a red "test failed" with the reason on hover. Thresholds live in a pure,
unit-tested helper; results are session-scoped.
- The model-homepage link moves off the card corner into the config dialog next to the
"get model ids" link (the card stays a single clickable surface; the freed corner hosts
the speed badges).
- Refinements: the group-header actions (add model / bulk API key / speed test) are all
icon + text buttons; the speed badges live on the card's meta line in their own
non-shrinking slot (the numbers never crowd or wrap the title row); the probe prompt
discourages reasoning and ends with an empty `<think></think>` block so reasoning models
skip their thinking phase instead of burning the probe budget on it.
## Model page refinements: draft follows the new default, ordered dropdown, GPT vision, homepage links
Four refinements: changing the Project's default model now resets the stored draft's model
selection to follow the new default; the chat model dropdown lists models in the same order
as the model library page; all GPT models are marked vision-capable; and model cards link to
each model's homepage.
## Details
- Draft follows the default: when saving a default-model change on the Models page, the
current user's stored draft for that Project drops its `modelRef` — the draft chat then
resolves the (new) default live instead of pinning the old model forever.
- Dropdown order: a new `orderModelsLikeLibrary` helper (unit-tested) flattens the library
grouping — built-in provider groups in MODEL_PROVIDERS order, user-defined groups after,
custom last, in-group order preserved — and the chat model dropdown now uses it.
- GPT vision: `openai/gpt-5.6-sol` and `openai/gpt-5.6-terra` are flipped to
vision-capable — GPT models are uniformly multimodal (OpenAI product-line policy) even
where the gateway page omits the modality.
- Homepage links: a new `modelHomepageUrl` helper (unit-tested) — OpenRouter and Qwen Token
Plan have stable per-model URL patterns (working for user-added ids in those groups too;
the unpaged Token Plan preview model falls back to the plan overview), direct vendors link
to their model docs page, custom/user-defined groups have none. Model cards show the link
as a corner external-link icon (a sibling of the clickable card, since interactive
elements must not nest).
## penguin-sdk and agenthub-models keep model keys project-local
Both skills now spell out where model API keys belong: in the project under the working directory, never in the user's global `~/.penguin`.
- penguin-sdk (v6) and agenthub-models (v3) instruct configuring keys with the penguin CLI into the app's own data root under CWD (`penguin config model add --root <data_dir> …`), or relying on vault-injected environment variables; reading, copying or falling back to model keys stored in the global `~/.penguin` directory is explicitly forbidden — that config belongs to the person running Penguin, not to the app being built.
- The no-key path stays as before: stop and ask the user to open the agent's settings via the gear icon on its card and update the key vault.
## Vault edits take effect on the next task; global keys don't count as usable
A vault save now invalidates the Agent's cached Session runtimes so the next task runs with the new values, and the AI-app skills stop counting keys from the global `~/.penguin` as usable.
- Server: `PUT /agents/:agentId/vault` bumps the Agent's config generation in the session manager; every runtime built before the update is discarded on its next idle access and re-resumed via the loader (resume re-reads `agent_state/.vault.toml`; history is preserved through the Trace). A task already in flight keeps the values it started with and rebuilds on the first access after it finishes. Unit and integration tests cover idle/busy entries and the HTTP wiring; the configuration docs (en/zh) and the web Vault tab hint document the new semantics.
- penguin-sdk (v9) and agenthub-models (v6): only two sources count as a usable credential — a vault-injected environment variable, or a key configured in the app's own data root (`penguin config model list --root <data_dir>`). Keys in the global `~/.penguin` (what a bare `penguin config model list` without `--root` reads — the CLI defaults to the global root) or any other `.penguin` directory never count and must never be used or copied; when no counted key is usable, stop immediately and ask the user to configure one instead of building or retrying.
@@ -0,0 +1,74 @@
# READMEs, blog, and docs site
## READMEs
The repository READMEs (en/zh).
### Restructure the README around the product story
The README now leads with the agents-build-agents pitch and community links, then three
feature showcases (benchmark chart, one-sentence RAG demo, self-evolution), followed by
changelog/blog/docs, supported models, human-first installation, a roadmap, CONTRIBUTING,
a citation, and credits.
### Details
- New narrative header: "With LangChain, you build agents by hand — at 1x speed. With
PenguinHarness, agents build agents — at 100x." with the subtitle "A zero-code CLI and
Web UI, connected to 1000+ models," plus community links (Discord / X / WeChat).
- Feature 1 "Simple and Efficient": light/dark benchmark bar charts generated from the
landing benchmark data (accuracy and cost per run vs Claude Code and OpenAI Codex, all
driven by DeepSeek V4 Pro), committed as `assets/readme/benchmark-{light,dark}.svg`.
- Feature 2 "Build an Agent in One Sentence": the RAG one-sentence prompt plus a real
product screenshot captured by the new `packages/landing/scripts/capture-readme-demo.mjs`
(same real-server + mock-LLM pipeline as the landing shots), committed as
`assets/readme/rag-demo-{light,dark}.webp`.
- Feature 3 "Self-Evolution": copy describing the evaluate-optimize-snapshot loop with an
HTML-comment placeholder for the upcoming demo video.
- New sections: Changelog / Blog / Docs links, a supported-models table (DeepSeek V4,
Kimi K3, GLM 5.2, Hunyuan 3, Qwen 3.8 Max, GPT 5.5, Gemini 3.5 Flash, Claude Opus 4.8
with their providers, plus the 1000+-via-gateways note), Requirements and Installation
split into "Web App — for humans" and "CLI & SDK — for agents", a Roadmap (benchmark
suite release), a BibTeX citation ({PrismShadow Team}), and the license/credits footer.
- New `CONTRIBUTING.md` absorbs the developer content: dev commands, repo layout table,
quality gates, the English-only and changelog working rules, and the README-asset
regeneration notes; the README's Development section now points there.
- `README.zh.md` mirrors the new structure in Chinese.
### Refresh the README model table against the current catalog
The supported-models table (the same eight models) becomes two columns — model on the
left, the comma-separated providers it's available from on the right (per today's catalog)
— and the note below now names all five OpenAI-compatible gateways.
### Details
- Availability per the catalog: DeepSeek V4 in five groups, GLM 5.2 in six, Kimi K3 via
OpenRouter and Qwen Pay-As-You-Go, Qwen 3.8 Max as the Token Plan preview, GPT 5.5 and
Claude Opus 4.8 native + OpenRouter, Hunyuan 3 via OpenRouter, Gemini 3.5 Flash native.
- The gateway note lists OpenRouter, Fireworks AI, SiliconFlow, Qwen Token Plan, and
Qwen Pay-As-You-Go. README.zh.md mirrors.
### Split badge rows and showcase the finished RAG app
The site badges (Website / Docs / Blog) and the community badges (Discord / X / WeChat) now sit on separate lines. The one-sentence example becomes the condensed claude-code-docs configuration-expert prompt (the chat page's example task carries the full version), and the demo image shows the FINISHED PRODUCT — the generated docs-expert app with cited clickable sources and example questions (`assets/readme/rag-app-<lang>-<theme>.webp`, per-language shots; the zh README shows the Chinese prompt and shots) — instead of the PenguinHarness chat UI. The mockup renderer (`rag-app-mockup.html` + rewritten `capture-readme-demo.mjs`, no server needed) comes along so the assets stay regenerable.
## Blog and docs site
Blog posts and the docs site.
### Announcement bar, AMD Fireworks-credits blog post, and the GDPevo launch story
The site gains a rotating announcement bar, a new campaign post, and a launch post that finally tells the whole story.
- **Announcement bar** — a switchable bar above the nav (auto-rotates every 6s, paused on hover, prev/next chevrons): entry 1 announces Kimi K3 and Qwen 3.8 Max availability (links to the models docs), entry 2 the $50 Fireworks credits campaign (links to the new post). Bilingual, on every page, scrolls away with the page.
- **New blog post `fireworks-credits-amd` (en/zh)** — announces the AMD AI Developer Program partnership bringing free Fireworks redemption codes: step-by-step application (join ADP, Member Perks, form with Fireworks AI selected, review, coupon email, redeem + API key, screenshots adapted from WhatGhost's guides with credit), then a three-step PenguinHarness setup (install, Fireworks group bulk-key + presets + speed test, run).
- **Launch post rewrite (en/zh)** — now opens with the GDPevo origin story: self-evolution was validated in the team's GDPevo Benchmark (linked), and bringing it to everyone is why PenguinHarness exists. The rest mirrors the README: three numbered reasons with the benchmark chart and RAG demo images (served from the site's own /blog-assets/), the security contract, the models table with the any-OpenAI-protocol note, install/usage steps, the roadmap (benchmark suite, desktop app, Windows), and a closing community call-to-action (Discord / X / WeChat / GitHub).
### The launch post settles on the numbered three-reasons structure
The launch blog post (en/zh) keeps the GDPevo origin story and the numbered "Why PenguinHarness" structure — ### 1 better on complex tasks at lower cost (benchmark chart + tables), ### 2 one-sentence Agent-builds-your-app (prompt + demo shot), ### 3 self-evolution — followed by the security contract, the models table, a Web-only "How to use it" (install + penguin web + Models page; no CLI commands), the roadmap, and the community call-to-action.
### The launch post shows the finished RAG app with the condensed prompt
The one-sentence build section now uses the condensed claude-code-docs configuration-expert prompt (Chinese in the zh post) and the finished-product screenshot of the generated docs-expert app (per-language, served from /blog-assets/), replacing the PenguinHarness chat capture.
+30
View File
@@ -0,0 +1,30 @@
# Skill library and the AI-app skills
## One-shot RAG app flow: skills rewrite and a draft-screen example task
PR #11's squash — briefly lost to a force-push race during the branch rebase — is restored: the skills library rewrite that makes one-sentence RAG apps one-shot, a draft-screen example task card, and the finished-product showcase pipeline.
- penguin-sdk rewritten around a complete RAG recipe (corpus collection, heading chunking, local BM25 retrieval, per-request Session SSE answers with citations, run-and-verify checklist); web-design gains the Penguin visual language and chat/RAG layout recipes; agent-creation covers skill bundles; agenthub-dev retires and a firecrawl skill joins.
- The Web draft screen gains example task cards (full prompts submitted as-is), with chat stream rendering refinements (markdown module, work summaries, stream follow) and colorful skill icons.
- The landing capture pipeline drives the docs-expert conversation and renders the finished-app mockup; README/Cases assets and prompts match the earlier finished-product switch.
## Skill library regrouped by audience, with square card actions
The library's six ad-hoc groups become four audience-oriented ones, and the card actions move into a tidy column.
- New groups, in order: Office Productivity / 办公效率 (`data-analysis`, `firecrawl`), Software Development / 软件开发 (`web-design`, `software-engineering`), AI App Development / AI 应用开发 (`penguin-sdk`, `penguin-cli`, `agenthub-models`), and Agent Tuning / Agent 调优 (`agent-creation`, `benchmark-design`, `agent-evaluation`, `agent-optimization`). Docs tables, skills/server tests, and the skills e2e spec follow.
- Skill card actions (update nudge / quick invoke / manage installs) are now equal squares stacked in one column, vertically centered at the card's right edge; the metadata line moves under the header.
## Skills: pin the CLI --root to the app dir and fast-stop when no key
Three skills gain two hard rules for AI-app development: always target the app's own data root with the penguin CLI, and stop asking for help the moment no API key is usable.
- penguin-sdk (v7), agenthub-models (v4), penguin-cli (v2): when building an AI app, `penguin config ...` must always pass `--root <data_dir>` pointing at the app's data directory inside the current working directory (the same path given to `createAgent({ root })`, e.g. `./penguin_data`); running without `--root` writes to the global `~/.penguin/data`, which belongs to the person running Penguin, not the app.
- penguin-sdk (v7) and agenthub-models (v4): when no API key is usable, stop immediately and ask the user for help instead of looping on tool calls — re-running `env`, re-checking the vault, or retrying the build wastes turns and money; one clear check, then hand back to the user (open the agent's settings via the gear icon and add a key to the key vault).
## System prompt exposes Provider and Model ID; skills state the key-first premise
The Agent system prompt's Environment section now carries the session model reference, and the AI-app skills spell out the key-and-root prerequisites up front.
- The `# Environment` section gains `Provider: {{PROVIDER}}` and `Model ID: {{MODEL_ID}}` placeholders (the session model's provider group and upstream id, filled from the resolved model entry) right before `Session ID`, and the fields are reordered to `Project Dir → Agent ID → CWD → Provider → Model ID → Session ID`. `assembleSystemPrompt` / `SessionEnvironmentValues` / `sessionEnvironment` thread the two new values through; the configuration docs' placeholder table (en/zh) follows.
- penguin-sdk (v8) and agenthub-models (v5) open with the prerequisite: for AI-app development, have the user add the model API key in **this agent's key vault** *before* building, and keep the app's Penguin data root **inside the CWD workspace** (`--root ./penguin_data`), never `~/.penguin`; model ids can come from the penguin CLI catalog.
+156
View File
@@ -0,0 +1,156 @@
# Tooling, language, and infrastructure
Repository language, dev commands, the changelog layout itself, and dependency upgrades.
## Make English the repository working language
Translated all residual non-i18n Chinese (comments, error/log messages, test titles and
fixtures, package metadata, e2e mock content) to English; Chinese remains only in i18n
catalogs, zh documents, and CJK-purpose test fixtures.
## Details
- Package metadata: `package.json` descriptions (root/server/web/landing) and the `link:cli`
fallback echo.
- Source comments: all remaining Chinese comments, including the SQL comment set in
`packages/server/src/db/schema.ts`.
- Hardcoded non-i18n user-facing strings (user-visible language change; error codes,
placeholders, and limits untouched): server HTTP/validation/409 messages, schedule TOML
validation errors, server logs, core SDK errors, CLI fallback errors, web provider
invariants.
- Tests: all Chinese describe/it titles and language-irrelevant fixtures translated;
assertions realigned to the new English source strings.
- e2e: mock LLM conversation content and cross-file branch markers translated consistently
(the substring relationships the mock's `includes()` dispatch relies on are preserved);
zh-locale UI assertions kept; `e2e/README.md` translated.
- Kept as required Chinese: zh i18n catalogs and fields (`strings.ts` dictionaries, CLI
`i18n.ts`, `titleZh`, `short_description_zh`, the zh language-name label, `*.zh.md` docs) and
test literals that assert zh i18n output or exercise CJK-specific behavior (display width,
CJK id validation, title-language consistency, zh-CN e2e UI assertions).
- Known gap: web `formatTaskStats` still hardcodes zh stat-line fragments; the proper fix is
routing them through the locale catalogs.
PR: [#5](https://github.com/Prism-Shadow/penguin-harness/pull/5) (merged 2026-07-20)
## Reorganize the changelog by release version
The top level now holds only version folders; each version's README carries the
one-line-per-change index, and detail files live inside the folder (mirrors agenthub PR
#162).
## Details
- `changelog/<version>/README.md` is the release summary: one line per change with its
title and one-sentence summary, linking the detail file relatively; `0.0.2/README.md`
absorbs the section previously kept in the top-level index.
- `changelog/README.md` now documents only the layout and the entry conventions (no
per-file listing).
- The working rules (AGENTS.md, CONTRIBUTING.md) point index updates at the version
folder's README instead of the top-level file.
## Serialize the dev prebuild to fix concurrent dev:server / dev:web clobbering
dev:server and dev:web now share a lock-serialized, deduplicated prebuild of skills and core,
so launching both at the same time no longer corrupts dist/ (previously two tsup builds with
clean:true raced in the same output directories).
## Details
- New `scripts/dev-prebuild.mjs`: takes an exclusive on-disk lock (atomic mkdir under
`node_modules/`) around `pnpm --filter skills --filter core build`; concurrent invocations
wait instead of clobbering. Locks left by crashed runs are stolen when the holder PID is
dead or the lock is older than 10 minutes.
- A 5-second success stamp collapses duplicate builds: starting dev:server and dev:web
simultaneously builds once (the waiter skips), and dev:server's inner re-invocation is a
no-op. The window is deliberately tiny so an edit-then-restart cycle always rebuilds —
the "never start on stale deps" behavior is unchanged.
- Root `dev:server` now just delegates to the server package's `dev` script (which prebuilds
via the shared script), removing the historical double build of core; `dev:web` prebuilds
via the script and then starts Vite.
- `pnpm dev` run standalone from `packages/server` now also builds skills (it previously
built only core, even though the server imports skills at runtime).
- Verified: concurrent invocations wait/skip correctly and release the lock; a simultaneous
dev:server + dev:web start performs a single build with both servers coming up cleanly.
## Add a combined pnpm dev and a dev:landing shortcut
`pnpm dev` now starts the backend and the web app together with prefixed logs (workspace deps
built once via the shared prebuild lock), and `pnpm dev:landing` serves the landing page dev
server from the repo root.
## Details
- `pnpm dev` runs `concurrently -n server,web "pnpm dev:server" "pnpm dev:web"`. Merging the
two was previously unsafe: each command prebuilt skills/core with tsup `clean: true` into
the same dist/ directories and the parallel builds clobbered each other. The lock-serialized
prebuild (see the serialize-dev-prebuild entry) removed that race, and its success stamp
collapses the two prebuilds into a single build on a combined start.
- `pnpm dev:landing` delegates to the landing package's Vite dev server (port 7366, completing
the 7364/7365/7367 dev-port family). The landing package has no workspace deps, so no
prebuild is involved.
- `concurrently` added as a root devDependency; the dev-command lists in README.md and
README.zh.md now cover `dev` and `dev:landing`.
- Verified: a combined start performs exactly one skills+core build (the other prebuild
waits and skips), with the server on 127.0.0.1:7364 and Vite on localhost:7365; the
landing dev server responds on localhost:7366. Note Vite binds localhost (IPv6 ::1) —
use `localhost`, not `127.0.0.1`, when probing the Vite ports with curl.
## Dev commands keep pnpm install current automatically
Forgetting `pnpm install` before a dev command no longer breaks the start: every dev
command's prestep checks install freshness (lockfile-hash stamp) and runs `pnpm install`
itself when node_modules is missing or the lockfile changed.
## Details
- `scripts/dev-prebuild.mjs` gains an install-freshness step inside its existing lock: the
pnpm-lock.yaml content hash is stamped after a successful install, so the usual dev start
pays nothing; a fresh clone or a pulled lockfile change triggers `pnpm install`
automatically (concurrent dev commands still install/build exactly once).
- The lock and stamps move from `node_modules/` to the OS temp directory keyed by the repo
path — they must exist before the first install does (the old location crashed on a fresh
clone), and per-checkout isolation comes from the key.
- `dev:docs` / `dev:landing` have no workspace deps to build but still need current
installs: they now run the prestep with `--install-only`.
## Upgrade AgentHub to 0.4.0 and adopt the opaque fidelity payload
@prismshadow/agenthub 0.3.3 -> 0.4.0 (agenthub PR #159): content items replace the item-level
`signature`/`phase` fields with one opaque `fidelity` object, and OmniMessage now carries it
verbatim end to end — Trace, replay, and resume included; the `agenthub-dev` skill joins the
built-in library.
## Details
- OmniMessage complete payloads (text, thinking, inline_data, inline_thinking, tool_call)
replace `signature?: string` / `phase?: string | null` with `fidelity?: Record<string,
unknown>` — an opaque wire-fidelity payload written to the Trace as-is and passed back
verbatim on replay (Claude thinking signatures, GPT-5 encrypted reasoning `{id,
encrypted_content}` and `{phase}` markers, the OpenAI-compatible `{reasoning_field}` name).
Builders take the object directly; an empty object is treated as absent.
- GenerativeModel's streaming translator mirrors AgentHub baseClient aggregation: a thinking
block is closed by its fidelity payload and a run of equal fidelity is one block (the
OpenAI-compatible clients stamp every thinking delta with the same `{reasoning_field}`,
which must not split blocks — this carries agenthub's reasoning-field replay fix through
PenguinHarness so multi-turn conversations against strict OpenAI-compatible upstreams
survive); a text segment splits on a differing `fidelity.phase` and closes on
`fidelity.signature`, merging fidelity keys. Text-phase stickiness across segments is gone
(mirrors baseClient).
- The `agenthub-dev` skill (AgentHub's own model-support development workflow) is installed
into the built-in library under the Penguin Development group, completed to the library
contract (version/updated frontmatter, short descriptions, a "Before you start" section,
and a custom icon).
- Malformed-classification fix for 0.4.x: agenthub now surfaces truncated streamed tool-call
arguments as its own `ToolCallArgumentParseError` (previously a raw `SyntaxError`) and
thinking-only completions as `EmptyResponseError`; `isMalformedJsonParseError` recognizes
both (instanceof + name fallback + cause chain), so these still end as `malformed` and the
engine reconnects and retries instead of failing the turn (caught by the malformed e2e).
- Docs (omni-message, interfaces, sessions-and-traces; en + zh) and the design specs updated
to the fidelity semantics.
- Traces written before this change carried `signature`/`phase` on payloads; per the
pre-release no-migration policy they are not converted (old fields are ignored on replay —
resume of such Sessions loses provider fidelity; delete and recreate if needed).
## The changelog folder merges related entries
One file per change had grown to 35 fragment files for this release alone; related changes are now merged into six date-slug entry files (models-and-catalog, web-app, landing-site, blog-and-docs, readme, tooling), each keeping the standard entry shape — H1 title, one-sentence summary, then per-change details — with the version README listing one line per entry file. The layout convention itself is unchanged; the branch was also rebased onto main to linearize away the merge commit with the final tree kept byte-identical.
+99
View File
@@ -0,0 +1,99 @@
# Web App
Chat input, session titles, and the skill library in the Web App.
## Chat input: positional slash, skill chips above the input, wider skill menu
The `/` command menu now opens from any caret position (like `@` mentions) and running a
command removes just the token; selected skills display as chips above the input next to
the agent chip; chip remove buttons recolor on hover instead of washing a background; the
skill dropdown widens for readable descriptions while staying inside phone screens.
## Details
- Positional slash (`slash-token.ts`, pure + unit-tested): a `/` at the start of the text
or after whitespace opens the command menu from wherever the caret is (paths/URLs never
trigger it), mirroring the existing positional `@` mention matching; running a command
removes only the `start..end` token and keeps the rest of the text. Send-time `@`
semantics are unchanged — only a leading `@` hands off.
- Selected skills render as chips above the input in the same row as the `@` handoff chip
(skill icon + monospace name + remove x), appearing as you pick them from the dropdown;
the toolbar count badge stays in sync.
- Chip remove buttons (agent and skill) lose the hover background — the x recolors instead.
- The skill dropdown widens (26rem, clamped to the viewport) so descriptions stay readable
on desktop without overflowing phones.
- The models-page group speed button collapses to icon-only below the sm breakpoint (three
labeled actions don't fit a 390px header — caught by the layout e2e).
## Start title generation after 1000 chars of body text
Session titles start generating as soon as ~1000 characters of main-session body text have
streamed, instead of waiting for the whole Task to finish — long answers no longer overrun
the title material — and the title prompt now suppresses chain-of-thought.
## Details
- Server: the output relay fires the title generator mid-run once EARLY_TITLE_BODY_CHARS
(1000) of main-session complete body text have streamed (sub-session text doesn't count);
the generator self-guards (NULL title, single flight), and the Task-completion trigger
stays as the short-answer fallback. Covered by a gated mid-run test proving the early
fire happens while the run is still in flight.
- Core: the captured assistant-side title material is capped at 1000 chars (the user side
keeps 2000) — a title only needs the opening of the answer.
- The title prompt adds an explicit "answer immediately — do not think aloud or produce
chain-of-thought" rule and ends with an empty `<think></think>` block, which many
reasoning models treat as an already-closed thinking phase (same trick as the model
probe), keeping the one-off request's budget on the title itself.
## Skill library: update reminder, borderless groups, accent icons
The skill library reminds you when an Agent's installed copy of a skill is older than the
library's version — an accent rotate button on the card updates every outdated Agent in one
click, and the manage-installs dialog marks outdated rows with an Update button. Group
sections lose their border, and card icons sit on a theme-accent background.
## Details
- The installed-skills snapshot now tracks each installed copy's `version` (the read API
already returns it — the installed SKILL.md frontmatter is the source). A pure
`outdatedAgentIds` helper (unit-tested) flags Agents whose copy is strictly older than
the library's; not-installed Agents and locally *newer* copies never trigger it.
- Card footer: when any Agent is outdated, an accent rotate button appears
("有新版本:更新 N 个 Agent 的安装" / "Update available"); clicking reinstalls the current
library copy on every outdated Agent (install-again-is-update semantics) with one batch
success toast; partial failures keep the succeeded Agents and toast the first error.
- Manage-installs dialog: outdated rows show an accent "更新"/"Update" button next to
"已安装"/"Installed", updating just that Agent.
- Styling: group sections are borderless (header background alone carries the grouping);
card icons use the theme accent (`--accent-bg`/`--accent-fg`, following the theme-color
setting) instead of the gray bordered tile.
## Titles strip skill-invocation markers; Agent cards show the installed skill count
Two Web App additions: generated titles no longer leak the `<use_skills>` block, and Agent cards display how many Skills are installed.
- Session titles strip machine markers before generation and in the fallback path: `stripConversationMarkers` (core, exported) removes `<use_skills>…</use_skills>` and the handoff / scheduled-task marker blocks from the material sent to the model, from the model output (via `sanitizeTitle`), and from the server's first-line fallback — so a skill invocation's marker can never become the title. Ordinary angle-bracket text (e.g. `<div>`) is left untouched.
- The Agent card's stats line gains an installed-skill count (open-book icon): the agents API/service compute `skillCount` from the `agent_state/skills/<name>/SKILL.md` directories, alongside the existing session / tool / vault-key / schedule counts.
## Slash menu stays on screen; skill card buttons go horizontal and light
Two Web App polish fixes: the slash command menu could grow past the top of the viewport, and the skill card actions change arrangement again.
- The slash menu (which lists /compact plus every installed skill) now caps its height at min(20rem, 40vh) with internal scrolling, so its top edge never leaves the screen; the active row keeps itself scrolled into view for keyboard navigation.
- Skill card actions return to a single horizontal row (still equal squares, vertically centered at the card's right edge), and all three buttons now wear the light secondary background.
## Password field polish, model-key visibility, and borderless skill cards
Several small UX fixes across the password fields, the model key input, and the skill library cards, plus a redrawn penguin game shot.
- The change-password dialog's current-password field now shows a hint naming the built-in admin's default initial password (penguin-2026), so a user who forgot it can still get in.
- The password show/hide toggle is removed from the tab order (tabIndex -1): Tab now moves between fields instead of landing on the reveal button.
- The model API key input gains the same show/hide toggle (it reuses PasswordInput), so a pasted key can be verified before saving.
- Skill library cards drop their border; hover now tints the whole card with a light gray background instead.
- The penguin sled game mockup is redrawn so the penguin stays a single connected cartoon shape (it had looked fragmented), and the game example is emphasized as 2D with a smoother, gentler difficulty ramp in the draft-screen card, its prompt, and the landing tab label.
## The admin initial password becomes penguin-2026
`admin123` sits in every breach corpus, so Chrome flags the first login as a compromised password; the seeded admin now starts as `penguin-2026` — unflagged, brand-related, all lowercase plus a hyphen so it stays easy to type.
The value swaps everywhere it appears: the server seed (`ADMIN_INITIAL_PASSWORD`), the release installer's first-login hint, READMEs, docs (quickstart / web-app / server-api), blog posts, landing quickstart copy, the screenshot capture scripts, and the e2e auth helper. The change-it-soon banner semantics are unchanged.
@@ -0,0 +1,17 @@
# Explicit model reference pairs, and a batch of review fixes
The provider group is no longer resolved automatically anywhere — a model is referenced by the complete `(provider, model_id)` pair or not at all — and a code review of the release branch turned up a further set of correctness, accessibility and documentation defects, all fixed here.
- **The provider is never inferred, guessed, or defaulted (core, CLI, server, web).** Two mechanisms used to pick a group on the user's behalf, and both are gone. Catalog inference (`inferProviderForUpstream`) took the first bare-id match, and gateways resell vendor models under their upstream ids — `glm-5.2`, `qwen3.7-max`, `qwen3.7-plus` and `deepseek-v4-pro` each belong to two groups now — so `penguin config model add --model-id glm-5.2 --api-key <key>` wrote a Z.AI credential onto the Qwen Token Plan entry and sent it to Alibaba's endpoint (on the previous release the same command resolved to `zhipu`). Config resolution (`resolveModelRef`) separately matched a bare `model_id` against the Project's configured models whenever it happened to be unique. Both are deleted, along with `inferProviderForUpstream` and `providersForUpstream`. The rule everywhere is now **both halves or neither**: `--provider` is required on `config model add`; on `run` / `chat` the pair as a whole stays optional (omit both for the Project default) but supplying one half is an error; `addModel` and `Agent.createSession` require the pair at the type level; `POST /sessions` and the schedule routes return 400 for a half reference, and a schedule file persisted with only a `modelId` is reported invalid instead of silently binding to a guessed group. `run_subagent` gains a `provider` argument beside `model_id` and refuses half a reference, so a delegating model cannot land a subagent on the wrong vendor either. The trace viewer stops attributing cost to a first-match guess for pre-tag traces, showing no cost instead. CLI help, the CLI/models/server-api/quickstart/configuration/interfaces docs (en+zh), the READMEs and every skill that scripts the CLI were updated to pass the pair.
- **Escape no longer wipes the composer (Web App).** The slash menu's Escape handler cleared the whole input — correct back when the menu only opened on a draft starting with `/`, but the menu became positional this release, so dismissing it from a `/token` mid-draft destroyed everything typed so far, unrecoverably (the textarea is controlled, so undo could not bring it back). Escape now only dismisses the menu, mirroring `@` mentions via a `slashDismissed` offset. Separately, `removeSlashToken` collapsed *every* run of two or more spaces in the surviving draft rather than the join seam, silently reflowing pasted code, YAML, aligned tables and indentation; it now touches the seam alone.
- **The penguin.ooo installer forwards its arguments.** `https://penguin.ooo/install.sh` piped the real installer straight into `sh`, so `--universal` and `--version` were silently discarded — leaving unsupported-architecture users with advice ("re-run with `--universal`") that could not work through the documented URL, and turning version pinning into "latest". The forwarder now downloads to a temp file, runs it with `"$@"`, and propagates the exit code; this also removes the partial-execution hazard of piping a truncated download into a script that deletes the old install before moving the new one in.
- **Speed test reports throughput or nothing.** Speed mode raised the output cap to 64 tokens but reused the connectivity prompt, which demands a single word — so the rate was computed over a 1-3 token sample dominated by the closing usage chunk's round trip, and the same model graded green or yellow on 30ms of jitter. Speed mode now has its own prompt that keeps generating into the cap, and the new pure `probeTps` drops the rate below meaningful sample floors (16 tokens, 100ms), leaving TTFT alone rather than a fabricated number.
- **Announcement bar (landing).** Rotation pauses on keyboard focus as well as hover, so the focused link is no longer left inside an `aria-hidden` slide translated out of view (an ARIA violation and a WCAG 2.2.2 failure); under `prefers-reduced-motion` the carousel no longer strands itself on the clone slide.
- **Streaming fidelity rules are pinned by tests.** The two rules added with the AgentHub 0.4.0 payload — a `fidelity.signature` closes the open text segment, and fidelity keys merge rather than replace across deltas — had no coverage; both mutations passed the whole suite. Tests now fail on either, guarding against a replay handing the provider a signature covering text it never signed, and against GPT-5 phase segmentation being lost on resume.
- **New benchmark results, and a comparison that changed shape.** Both suites are refreshed, and the design of the comparison changed with them: instead of running all three harnesses on one shared DeepSeek V4 Pro, each now runs the model it is normally paired with — PenguinHarness on DeepSeek V4 Pro, Claude Code on Claude Opus 4.8, OpenAI Codex on GPT-5.5 — so the model column is part of the result rather than a constant, and Tokens/cost are suite totals rather than per-run means. Data analysis (15 tasks, single run): 66.67% for PenguinHarness against 53.33% for both rivals, at $0.55 versus $38.48 and $19.41. Coding (40 tasks × 2 runs, accuracy over all 80 outcomes): 71.25%, level with Codex and behind Claude Code's 86.25%, at $3.81 versus $220.08 and $146.97. The copy moved with the data — the section no longer claims "same model" or "highest accuracy of the three", both of which the new results contradict, and leads on cost-per-quality instead; the 35×/70×/58×/39× multiples in that copy are now asserted against the data by unit test so they cannot drift on the next refresh. The README and blog charts are generated from the same source by `render-benchmark-svg.mjs`. Bars stay linear and zero-based, but carry a minimum length: at true scale the cost panels put PenguinHarness at ~2px, which reads as an empty slot rather than as the winning series, so the floor keeps it visible while still plainly the shortest bar — with the exact figure printed beside every one. The README's chart caption is now a single-line conclusion rather than a paragraph of run settings, matching the RAG cost hook.
- **The demo videos reach the landing page.** The self-improvement section's "demo video coming soon" pill is replaced by the recording itself, and the RAG case tab plays the finished app instead of only picturing it — both following the READMEs, both locale-matched (`evo_zh`/`evo_en`, `rag_zh`/`rag_en`). The files stay out of this repo: at ~9 MB each against a ~17 MB history, committing four would triple what every contributor clones for assets only the marketing site shows, so they are served from the sibling `penguin-harness-community` repo via `raw.githubusercontent`, which answers with `accept-ranges: bytes` (seeking works) and `access-control-allow-origin: *`. Its `application/octet-stream` type does not block playback — `nosniff` is enforced for scripts and styles, not media — and this was verified end to end in Chromium rather than assumed. Every embed pairs a poster with `preload="none"`, so a page view downloads one ~37 KB still and no video byte until someone presses play (asserted by watching the network on load: zero mp4 requests). The standalone embed deliberately drops the BrowserFrame chrome the screenshots use, because the self-improvement recording is a narrated deck rather than the product's own UI; the RAG case keeps the frame, because there it really is the app.
- **The draft page is genuinely centred, and the upward menus stop clipping.** The brand block sat inside the upper flex spacer, so the gap above it was shorter than the gap below the example cards by exactly the brand's own height — the whole page rode high, and the slash menu, which opens upward from the composer, ran into the top of the scroll area. The brand now lives in the same centred block as the composer, pills and example tasks, making the two gaps identical (measured: 134/134 at 1000px tall, 84/84 at 900, 34/34 at 800). The slash and `@` menus additionally cap their height to the room actually measured between the composer and the nearest clipping ancestor instead of a fixed `40vh`, which could not know that distance: previously the menu was cut by ~50px at 900px tall and still 5px at 800px, and now it fits at every height down to 620px, shrinking and scrolling internally instead of being clipped.
- **Web polish.** Speed-test badges reset when the active Project changes (results are keyed by `(provider, model_id)`, so they used to leak across Projects); the work-group header ticks and settles on the same span, so its duration never jumps backwards; the markdown renderer's component map has a stable identity, so streaming no longer remounts closed code blocks and drops text selection or Copy state.
- **Sources are diffable again.** Four files embedded a literal NUL byte as a composite-key separator instead of the `\0` escape, which made git treat them as binary — `catalog-sync.ts`, the model-catalog and models tests, and the scheduler produced no reviewable diff, no `git grep` output, and no textual merge. The runtime value is unchanged.
- **Docs and metadata.** The workspace `files/content` preview mode and its response headers are documented in the server API reference (en/zh); the skills README lists the groups that actually exist; the skill-library page's docstrings describe the per-skill palette and two-column grid it really renders; `penguin-cli`'s `updated` date matches its version bump; the penguin-sdk RAG recipe stops installing the entire skill library into the app it generates; the README model table no longer lists Kimi K3 under Moonshot, which does not serve it; and a Chinese comment in the catalog's pricing block is translated.
- **A test that did not test.** The early-title case asserting that sub-session output does not count toward the 1000-character trigger stayed green with the origin guard removed; it now fails without it.
+17
View File
@@ -0,0 +1,17 @@
# Version 0.0.2
Unreleased.
- [2026-07-21] A model is always referenced by an explicit `(provider, model_id)` pair — the provider is never inferred, guessed, or defaulted — together with the fixes from a full review of the release branch, refreshed benchmark results, and the demo videos on the landing page. ([details](2026-07-21-model-reference-and-review-fixes.md))
- [2026-07-20] Model catalog: preset provider groups, catalog entries and ordering, the Models page features built around them, and where model credentials are allowed to live. ([details](2026-07-20-models-and-credentials.md))
- [2026-07-20] Web App: chat input and slash menu, session titles, the skill library cards, password fields, and the seeded admin password. ([details](2026-07-20-web-app.md))
- [2026-07-20] Landing site: the story it tells, section structure, animations, navigation, the install domain, and the case gallery. ([details](2026-07-20-landing-site.md))
- [2026-07-20] Skill library regrouped by audience, and the AI-app skills that turn one sentence into a running RAG app. ([details](2026-07-20-skills.md))
- [2026-07-20] The repository READMEs, the blog posts, and the docs site. ([details](2026-07-20-readme-blog-and-docs.md))
- [2026-07-20] Repository language, dev commands, the changelog layout itself, and dependency upgrades. ([details](2026-07-20-tooling.md))
+10
View File
@@ -0,0 +1,10 @@
# Changelog Details
The root [../CHANGELOG.md](../CHANGELOG.md) keeps one brief line per release. Each release owns a folder here:
- `<version>/README.md` — the release summary: one brief line per change, each linking its detail file, e.g. `- [YYYY-MM-DD] Brief description. ([details](YYYY-MM-DD-short-slug.md))`
- `<version>/YYYY-MM-DD-short-slug.md` — one detail file per entry, named by the entry date: an H1 title, then what changed and why, using `##` sections once the entry covers more than one thing.
- Entries are grouped by the surface they change — Models, Web App, landing site, skills, docs, tooling — rather than one file per commit. A related change extends the existing file and its summary line instead of opening a new one.
- Changes not yet released go into the upcoming version's folder; during release preparation, rename the folder if the number changed and add the release line to the root file. Released folders are frozen.
Written in English. History starts after the v0.0.1 release (2026-07-19); earlier changes are not backfilled.
+1 -1
View File
@@ -149,5 +149,5 @@ echo "PenguinHarness $installed_version installed to $INSTALL_DIR"
echo ""
echo "Get started:"
echo " penguin --help # all commands"
echo " penguin web # start the Web UI at http://127.0.0.1:7364 (initial login: admin / admin123)"
echo " penguin web # start the Web UI at http://127.0.0.1:7364 (initial login: admin / penguin-2026)"
echo " penguin server # headless server (PORT / HOST to override)"
+6 -3
View File
@@ -16,15 +16,18 @@
"build": "pnpm -r build && pnpm link:cli",
"link:cli": "pnpm --dir packages/cli link --global || echo '[link:cli] pnpm global directory not configured; skipping the global penguin link (run pnpm setup once first)'",
"penguin": "tsx packages/cli/src/index.ts",
"dev:server": "pnpm --filter @prismshadow/penguin-skills --filter @prismshadow/penguin-core build && pnpm --filter @prismshadow/penguin-server dev",
"dev:web": "pnpm --filter @prismshadow/penguin-skills --filter @prismshadow/penguin-core build && pnpm --filter @prismshadow/penguin-web dev",
"dev:docs": "pnpm --filter @prismshadow/penguin-docs dev",
"dev": "concurrently -n server,web -c cyan,magenta \"pnpm dev:server\" \"pnpm dev:web\"",
"dev:server": "pnpm --filter @prismshadow/penguin-server dev",
"dev:web": "node scripts/dev-prebuild.mjs && pnpm --filter @prismshadow/penguin-web dev",
"dev:docs": "node scripts/dev-prebuild.mjs --install-only && pnpm --filter @prismshadow/penguin-docs dev",
"dev:landing": "node scripts/dev-prebuild.mjs --install-only && pnpm --filter @prismshadow/penguin-landing dev",
"build:site": "node scripts/build-site.mjs"
},
"devDependencies": {
"@prismshadow/penguin-cli": "workspace:*",
"@prismshadow/penguin-core": "workspace:*",
"@types/node": "^24.0.0",
"concurrently": "^10.0.3",
"prettier": "^3.9.4",
"tsx": "^4.20.0",
"typescript": "^5.6.0",
+1 -1
View File
@@ -22,7 +22,7 @@
"penguin": "tsx src/index.ts"
},
"dependencies": {
"@prismshadow/agenthub": "^0.3.3",
"@prismshadow/agenthub": "^0.4.0",
"@prismshadow/penguin-core": "workspace:*",
"@prismshadow/penguin-server": "workspace:*",
"@prismshadow/penguin-skills": "workspace:*",
+15 -2
View File
@@ -1,12 +1,14 @@
/**
* `penguin chat` — interactive REPL.
*
* penguin chat [--model-id <id>] [--provider <group>] [--project-id <id>] [--agent-id <id>]
* penguin chat [--model-id <id> --provider <group>] [--project-id <id>] [--agent-id <id>]
* [--workspace <path>] [--approve <allow-all|deny-all|read-only|always-ask>]
*
* Each line of input starts one conversation turn; `/compact` proactively compacts the
* context (reason=manual); `/exit` or `/quit` exits.
* Uses the current directory when no Workspace is specified.
* Uses the current directory when no Workspace is specified. A model reference is always an
* explicit `(provider, model_id)` pair, so `--model-id` and `--provider` must be given
* together; giving neither uses the Project's default model.
*
* Multi-line input: trailing `\` continues the line; when the terminal supports bracketed
* paste, a multi-line paste is treated as a single message (sent on Enter).
@@ -65,6 +67,17 @@ export function registerChatCommand(program: Command, t: Messages): void {
.option("--approve <mode>", t.common.approve)
.option("--resume [sessionId]", t.chat.resume)
.action(async (opts) => {
// The model reference is a pair: commander can only require each option on its own,
// so the "both or neither" rule is enforced here. Giving neither is the normal case
// and falls back to the Project's default model. Skipped under --resume, which
// rejects both options outright further down with a more specific message.
// Usage errors go to stderr (as in `run` and `config model add`), unlike this file's
// informational messages, which the REPL writes to stdout.
if (opts.resume === undefined && Boolean(opts.modelId) !== Boolean(opts.provider)) {
process.stderr.write(`${t.error(t.modelRefIncomplete())}\n`);
process.exitCode = 1;
return;
}
const mode = resolveApprovalMode(opts.approve, t);
const out = process.stdout;
+12 -13
View File
@@ -2,7 +2,7 @@
* `penguin config` — manages a Project's model credentials, default model, model list,
* Agent-level vault environment variables, and UI language.
*
* penguin config model add --model-id <upstream id> [--provider <group>] [--api-key <key>] [--context-window <n>] [--set-default] [--root <dir>]
* penguin config model add --model-id <upstream id> --provider <group> [--api-key <key>] [--context-window <n>] [--set-default] [--root <dir>]
* penguin config model default --model-id <upstream id> --provider <group> [--root <dir>]
* penguin config model vision --model-id <upstream id> --provider <group> [--root <dir>]
* penguin config model list [--root <dir>]
@@ -13,12 +13,12 @@
*
* `--model-id` always takes the **upstream id** (the request id sent to AgentHub verbatim),
* which together with `--provider` forms a `(provider, model_id)` paired reference —
* **no string concatenation is ever performed**. For `model add`, --provider defaults to
* an inference from the built-in catalog (falling back to custom when inference fails);
* a new entry's client_type defaults according to the group's semantics (not set for
* first-party vendors; openai for custom / self-hosted groups / gateways, with the
* gateway's endpoint base URL pre-filled). For `model default` / `model vision`,
* --provider is **required**; core validation raises an error when the reference is not
* **no string concatenation is ever performed**. `--provider` is **required** on all three
* model subcommands: the group is never guessed, so `--api-key` can never land on a vendor
* the user did not name. For `model add`, a new entry's client_type defaults according to
* the group's semantics (not set for first-party vendors; openai for custom / self-hosted
* groups / gateways, with the gateway's endpoint base URL pre-filled). For `model default`
* / `model vision`, core validation raises an error when the reference is not
* found in models. `--root` specifies the data root directory (priority: option >
* PENGUIN_HOME > ~/.penguin/data). The UI language is controlled by the PENGUIN_LANG
* environment variable; `config lang` writes it into the shell startup file and restarts
@@ -39,7 +39,6 @@ import {
catalogEntryFor,
formatModelRef,
getModel,
inferProviderForUpstream,
loadAgentVault,
loadProjectConfig,
providerInfo,
@@ -117,7 +116,7 @@ export function registerConfigCommand(program: Command, t: Messages): void {
.command("add")
.description(t.config.addDesc)
.requiredOption("--model-id <id>", t.config.addModelId)
.option("--provider <group>", t.config.addProvider)
.requiredOption("--provider <group>", t.config.addProvider)
.option("--api-key <key>", t.config.addApiKey)
.option("--base-url <url>", t.config.addBaseUrl)
.option("--context-window <n>", t.config.addContextWindow, parseIntArg)
@@ -133,11 +132,11 @@ export function registerConfigCommand(program: Command, t: Messages): void {
.option("--root <dir>", t.common.root)
.action(async (opts) => {
const root = resolveRootOption(opts.root);
// --model-id takes the upstream id, paired with --provider as a reference
// (--provider defaults to catalog-based inference, falling back to custom); no
// concatenation is performed.
// --model-id takes the upstream id, paired with the required --provider as a
// reference; the group is never guessed, so --api-key can only ever land on the
// vendor the user named. No concatenation is performed.
const modelId: string = opts.modelId;
const provider: string = opts.provider ?? inferProviderForUpstream(modelId);
const provider: string = opts.provider;
const ref: ModelRef = { provider, model_id: modelId };
const before = await loadProjectConfig(root, opts.projectId);
const existed = getModel(before, ref) !== undefined;
+12 -4
View File
@@ -1,14 +1,14 @@
/**
* `penguin run` — send a single Task in one shot.
*
* penguin run -m <msg> [--model-id <id>] [--provider <group>] [--workspace <path>]
* penguin run -m <msg> [--model-id <id> --provider <group>] [--workspace <path>]
* [--project-id <id>] [--agent-id <id>]
* [--approve <allow-all|deny-all|read-only|always-ask>]
*
* Uses the current directory when Workspace is unspecified; uses the Project's default model
* when model is unspecified. `--provider` is optional: when omitted, `--model-id` is resolved
* via resolveModelRef semantics (only matches when the exact value is globally unique in the
* config; ambiguity is an error). Defaults to interactive per-call approval; `--approve`
* when model is unspecified. A model reference is always an explicit `(provider, model_id)`
* pair, so `--model-id` and `--provider` must be given together — giving only one of them is
* an error, never a lookup. Defaults to interactive per-call approval; `--approve`
* selects the permission mode.
* Docs: /docs/cli § "penguin run".
*/
@@ -31,6 +31,14 @@ export function registerRunCommand(program: Command, t: Messages): void {
.option("--workspace <path>", t.common.workspace)
.option("--approve <mode>", t.common.approve)
.action(async (opts) => {
// The model reference is a pair: commander can only require each option on its own,
// so the "both or neither" rule is enforced here. Giving neither is the normal case
// and falls back to the Project's default model.
if (Boolean(opts.modelId) !== Boolean(opts.provider)) {
process.stderr.write(`${t.error(t.modelRefIncomplete())}\n`);
process.exitCode = 1;
return;
}
const mode = resolveApprovalMode(opts.approve, t);
const agent = await createAgent({
+13 -7
View File
@@ -23,7 +23,7 @@ export interface Messages {
projectId: string;
agentId: string;
modelId: string;
/** run/chat's --provider: pairs with --model-id; when omitted, resolved by unique match (ambiguity is an error). */
/** run/chat's --provider: must be given together with --model-id (the group is never inferred). */
provider: string;
/** Data root directory option (priority: --root > PENGUIN_HOME > ~/.penguin/data). */
root: string;
@@ -113,6 +113,8 @@ export interface Messages {
approveModeInvalid(value: string): string;
/** Render label for an approval decision (frontend renders the approval_decision event; one label each for allow/deny). */
approvalDecision(decision: "allow" | "deny"): string;
/** run/chat given only one of --model-id / --provider: a model reference is always an explicit pair, never a lookup. */
modelRefIncomplete(): string;
/** --resume is mutually exclusive with --workspace/--model-id (neither can change once the Session is created). */
resumeNoOverride(): string;
/** --resume given without a session id, and the current Agent has no Session at all. */
@@ -156,7 +158,7 @@ const en: Messages = {
agentId: "Agent id",
modelId: "Model to use (upstream model id; defaults to the Project default model)",
provider:
"Provider of --model-id; when omitted, the model id must match exactly one configured entry (ambiguity is an error)",
"Provider group of --model-id; required whenever --model-id is given (the group is never inferred)",
root: "Data root directory (overrides PENGUIN_HOME and ~/.penguin/data)",
workspace: "Workspace directory; must already exist (defaults to the current directory)",
approve:
@@ -168,11 +170,11 @@ const en: Messages = {
addDesc: "Add or update a model, optionally writing a credential",
addModelId: "Upstream model id sent to AgentHub as-is (e.g. claude-sonnet-4-6)",
addProvider:
"Provider group stored alongside model_id; inferred from the builtin catalog when omitted, else custom",
"Provider group stored alongside model_id; required, never inferred (use custom for anything without a vendor group)",
addApiKey: "API key, stored inline in the Project's hidden .project_config.toml",
addBaseUrl: "Custom base URL",
addContextWindow: "Context window size (tokens)",
addClientType: "AgentHub client type (e.g. openai); inferred from model id when omitted",
addClientType: "AgentHub client type (e.g. openai); defaults by provider group when omitted",
addVision: "Mark the model as supporting image input (vision)",
addNoVision: "Mark the model as NOT supporting image input; omit both to keep current",
addPriceCacheRead: "Price per 1M tokens: cache read (USD)",
@@ -235,6 +237,8 @@ const en: Messages = {
approveModeInvalid: (value) =>
`Invalid approval mode "${value}". Use allow-all, deny-all, read-only, or always-ask.`,
approvalDecision: (decision) => (decision === "allow" ? "✓ [approved]" : "× [denied]"),
modelRefIncomplete: () =>
"--model-id and --provider must be given together: a model reference is always an explicit (provider, model_id) pair. Omit both to use the Project default model.",
resumeNoOverride: () =>
"--resume does not accept --workspace, --model-id or --provider: they follow the original Session and cannot change.",
resumeNoSession: () => "No session to resume: this agent has no recorded sessions yet.",
@@ -268,7 +272,7 @@ const zh: Messages = {
projectId: "Project id",
agentId: "Agent id",
modelId: "本次使用的模型(上游模型 id;默认 Project 默认模型)",
provider: "--model-id 的 provider 分组;省略时 model id 须在配置中精确唯一命中(歧义报错)",
provider: "--model-id 的 provider 分组;给出 --model-id 时必须一并给出(分组不作任何推断)",
root: "数据根目录(优先于 PENGUIN_HOME 与 ~/.penguin/data)",
workspace: "Workspace 目录,须为已存在目录(默认当前目录)",
approve:
@@ -279,11 +283,11 @@ const zh: Messages = {
modelDesc: "管理模型 credential 与默认模型",
addDesc: "新增或更新一个模型,并可写入 credential",
addModelId: "上游模型 id(如 claude-sonnet-4-6,原样发给 AgentHub)",
addProvider: "与 model_id 分列存储的 provider 分组;缺省按内置目录推断,推断不出为 custom",
addProvider: "与 model_id 分列存储的 provider 分组;必填,不作推断(无厂商分组时填 custom)",
addApiKey: "API key,内联存入 Project 的隐藏文件 .project_config.toml",
addBaseUrl: "自定义 base url",
addContextWindow: "上下文窗口大小(token 数)",
addClientType: "AgentHub 客户端协议(如 openai);缺省由 model id 推断",
addClientType: "AgentHub 客户端协议(如 openai);缺省按 provider 分组的语义取值",
addVision: "标注该模型支持图片输入(视觉)",
addNoVision: "标注该模型不支持图片输入;两者都不给则保留原值",
addPriceCacheRead: "每百万 token 价格:缓存读取(USD)",
@@ -345,6 +349,8 @@ const zh: Messages = {
approveModeInvalid: (value) =>
`无效的审批模式 "${value}"。请使用 allow-all、deny-all、read-only 或 always-ask。`,
approvalDecision: (decision) => (decision === "allow" ? "✓ [已批准]" : "× [已拒绝]"),
modelRefIncomplete: () =>
"--model-id 与 --provider 必须成对给出:模型引用始终是显式的 (provider, model_id) 组合。两者都不给则使用 Project 默认模型。",
resumeNoOverride: () =>
"--resume 不接受 --workspace、--model-id 与 --provider:均沿用原 Session,创建后不可更换。",
resumeNoSession: () => "没有可恢复的 Session:当前 Agent 还没有任何会话记录。",
+37 -9
View File
@@ -1,10 +1,10 @@
/**
* Integration tests for `penguin config model add|default|vision|list` (run through
* commander's parseAsync for the full command path): --model-id always takes the
* upstream id, paired with --provider to form a (provider, model_id) reference (add's
* --provider defaults to catalog-based inference, falling back to custom when
* inference fails; default / vision require --provider and raise an error when the
* reference isn't found in models — no string concatenation is ever performed); --root
* upstream id, paired with --provider to form a (provider, model_id) reference
* (--provider is required on all three subcommands — the group is never inferred — and
* default / vision raise an error when the reference isn't found in models; no string
* concatenation is ever performed); --root
* specifies the data root directory (takes priority over PENGUIN_HOME); persisted to a
* single hidden .project_config.toml (mode 0600, credentials inline, provider and
* model_id as separate columns); list displays provider and model_id as separate
@@ -84,13 +84,15 @@ describe("penguin config model add/list (--root plus provider / model_id stored
"add",
"--model-id",
"my-own-model",
"--provider",
"custom",
"--api-key",
"sk-root-secret-1",
"--root",
tmpRoot,
]);
expect(add.code).toBe(0);
// Catalog inference fails -> falls back to the custom group (provider is a separate field, never concatenated into the id).
// The named group is stored as a separate field, never concatenated into the id.
expect(add.out).toContain("Added model (provider=custom, model_id=my-own-model).");
const file = projectConfigPath(tmpRoot, DEFAULT_PROJECT_ID);
@@ -119,11 +121,13 @@ describe("penguin config model add/list (--root plus provider / model_id stored
expect(list.out).not.toContain("sk-root-secret-1");
});
it("built-in catalog infers the grouping: an upstream id matching the catalog lands under that provider; --set-default writes a pair reference", async () => {
it("naming an existing pair updates that preset entry in place; --set-default writes a pair reference", async () => {
const add = await runModel([
"add",
"--model-id",
"claude-sonnet-4-6",
"--provider",
"anthropic",
"--set-default",
"--root",
tmpRoot,
@@ -168,8 +172,16 @@ describe("penguin config model add/list (--root plus provider / model_id stored
});
it("client_type defaults by grouping semantics (PRN-021): custom / self-hosted / gateway get openai, first-party providers get none", async () => {
// custom (catalog inference fails) and self-hosted groups (--provider not a catalog value): default to client_type=openai.
await runModel(["add", "--model-id", "my-openai-proxy", "--root", tmpRoot]);
// The custom group and self-hosted groups (--provider not a catalog value): default to client_type=openai.
await runModel([
"add",
"--model-id",
"my-openai-proxy",
"--provider",
"custom",
"--root",
tmpRoot,
]);
await runModel(["add", "--model-id", "in-house-1", "--provider", "mylab", "--root", tmpRoot]);
// A non-catalog id under a first-party vendor group: client_type is not set (AgentHub auto-routes by upstream id).
await runModel([
@@ -241,13 +253,29 @@ describe("penguin config model add/list (--root plus provider / model_id stored
});
});
describe("model default/vision: --provider is required, (provider, model_id) pair reference", () => {
describe("model add/default/vision: --provider is required, (provider, model_id) pair reference", () => {
it("missing --provider: commander usage error, nonzero exit code", async () => {
const bad = await runModel(["default", "--model-id", "deepseek-v4-flash", "--root", tmpRoot]);
expect(bad.code).not.toBe(0);
expect(bad.err).toContain("--provider");
});
it("add without --provider is a usage error too: the group is never inferred, so no config is written", async () => {
const bad = await runModel([
"add",
"--model-id",
"claude-sonnet-4-6",
"--api-key",
"sk-never-stored",
"--root",
tmpRoot,
]);
expect(bad.code).not.toBe(0);
expect(bad.err).toContain("--provider");
// The credential must not have landed on a guessed vendor: nothing was persisted at all.
await expect(fs.access(projectConfigPath(tmpRoot, DEFAULT_PROJECT_ID))).rejects.toThrow();
});
it("dangling reference: the pair is not in models; the error carries the pair reference and a model list hint", async () => {
const bad = await runModel([
"default",
+106
View File
@@ -0,0 +1,106 @@
/**
* `penguin run` / `penguin chat`: a model reference is always an explicit
* (provider, model_id) pair. Commander can only mark each option required on its own, so
* the "both or neither" rule is enforced inside the action — supplying exactly one of
* --model-id / --provider is a usage error (never a lookup against the configured models),
* while supplying neither falls back to the Project's default model.
*/
import fs from "node:fs/promises";
import os from "node:os";
import path from "node:path";
import { afterEach, beforeEach, describe, expect, it, vi } from "vitest";
import { Command } from "commander";
import { registerChatCommand } from "../src/commands/chat.js";
import { registerRunCommand } from "../src/commands/run.js";
import { getMessages } from "../src/i18n.js";
// The --resume case below reaches createAgent, which initializes Agent state on disk; point
// the data root at a throwaway directory so no test ever writes to the real ~/.penguin/data.
let tmpHome: string;
let prevHome: string | undefined;
beforeEach(async () => {
prevHome = process.env.PENGUIN_HOME;
tmpHome = await fs.mkdtemp(path.join(os.tmpdir(), "penguin-cli-pairing-"));
process.env.PENGUIN_HOME = tmpHome;
});
afterEach(async () => {
if (prevHome === undefined) delete process.env.PENGUIN_HOME;
else process.env.PENGUIN_HOME = prevHome;
await fs.rm(tmpHome, { recursive: true, force: true });
});
/** Runs one command, capturing stdout / stderr and the exit code (without exiting the process). */
async function runCommand(
register: (program: Command, t: ReturnType<typeof getMessages>) => void,
args: string[],
): Promise<{ out: string; err: string; code: number }> {
const program = new Command();
program.exitOverride();
register(program, getMessages("en"));
const out: string[] = [];
const err: string[] = [];
const outSpy = vi.spyOn(process.stdout, "write").mockImplementation((chunk) => {
out.push(String(chunk));
return true;
});
const errSpy = vi.spyOn(process.stderr, "write").mockImplementation((chunk) => {
err.push(String(chunk));
return true;
});
const prevExitCode = process.exitCode;
process.exitCode = undefined;
try {
await program.parseAsync(["node", "penguin", ...args]);
return { out: out.join(""), err: err.join(""), code: Number(process.exitCode ?? 0) };
} catch (e) {
const exitCode = (e as { exitCode?: number }).exitCode;
return { out: out.join(""), err: err.join(""), code: exitCode || 1 };
} finally {
outSpy.mockRestore();
errSpy.mockRestore();
process.exitCode = prevExitCode;
}
}
describe("run: --model-id and --provider must be given together", () => {
it("--model-id without --provider: error on stderr, exit code 1", async () => {
const bad = await runCommand(registerRunCommand, [
"run",
"-m",
"hi",
"--model-id",
"deepseek-v4-flash",
]);
expect(bad.code).toBe(1);
expect(bad.err).toContain("--model-id and --provider must be given together");
});
it("--provider without --model-id: same error (the pair is never half-specified)", async () => {
const bad = await runCommand(registerRunCommand, ["run", "-m", "hi", "--provider", "deepseek"]);
expect(bad.code).toBe(1);
expect(bad.err).toContain("--model-id and --provider must be given together");
});
});
describe("chat: --model-id and --provider must be given together", () => {
it("--model-id without --provider: error on stderr, exit code 1", async () => {
const bad = await runCommand(registerChatCommand, ["chat", "--model-id", "deepseek-v4-flash"]);
expect(bad.code).toBe(1);
expect(bad.err).toContain("--model-id and --provider must be given together");
});
it("--resume plus a lone --model-id keeps the more specific resume error", async () => {
const bad = await runCommand(registerChatCommand, [
"chat",
"--resume",
"sess-1",
"--model-id",
"deepseek-v4-flash",
]);
expect(bad.code).toBe(1);
expect(bad.out).toContain("--resume does not accept");
expect(`${bad.out}${bad.err}`).not.toContain("must be given together");
});
});
+1 -1
View File
@@ -36,7 +36,7 @@
"build": "tsup"
},
"dependencies": {
"@prismshadow/agenthub": "^0.3.3",
"@prismshadow/agenthub": "^0.4.0",
"@prismshadow/penguin-skills": "workspace:*",
"smol-toml": "^1.3.0",
"yaml": "^2.5.0"
+28 -16
View File
@@ -79,12 +79,14 @@ export interface CreateAgentOptions {
export interface CreateSessionOptions {
/** Workspace for this run; if unspecified, a temporary Workspace is created under the Agent directory. */
workspaceDir?: string;
/** Model used for this Session (upstream model_id); if unspecified, uses the Project's default Model. */
/**
* Model used for this Session (upstream model_id); must be given together with `provider`.
* Omit both to use the Project's default Model.
*/
modelId?: string;
/**
* Provider grouping for `modelId` (a paired reference); if omitted, resolved via
* `resolveModelRef` semantics — `model_id` only resolves if it is a globally unique
* exact match in the config; zero or multiple matches produce a clear error.
* Provider grouping for `modelId`: the two together form the paired reference. It is never
* inferred, so it's "both or neither" — either half alone is an error, not a lookup.
*/
provider?: string;
/** Explicit credentials; if unspecified, falls back to credentials in the Project config, then to AgentHub reading environment variables. */
@@ -134,18 +136,20 @@ export class Agent {
*/
async createSession(opts: CreateSessionOptions = {}): Promise<Session> {
// Model is validated first (before creating the Workspace, so failure leaves no
// temp directory behind): the reference must resolve to an entry in the Project
// config (the (provider, model_id) pair is the unique key); a reference
// outside the config throws immediately rather than passing silently — otherwise
// credentials, pricing, and the context window would all be unavailable.
if (opts.modelId === undefined && opts.provider !== undefined) {
// temp directory behind): the reference must be the complete (provider, model_id)
// pair — the config's unique key — and must name an entry in the Project config; a
// reference outside the config throws immediately rather than passing silently,
// otherwise credentials, pricing, and the context window would all be unavailable.
// Half a reference is always an error: the missing half is never inferred, since a
// guessed provider would send the entry's credential to a vendor nobody named.
if ((opts.modelId === undefined) !== (opts.provider === undefined)) {
throw new Error(
"provider was specified without modelId: a model reference must be given as a pair (provider cannot be used alone).",
"A model reference must be given as a (provider, model_id) pair: both must be specified, or neither (to use the Project's default model).",
);
}
let ref: ModelRef;
if (opts.modelId !== undefined) {
// The only entry point for resolving an "omitted provider" reference (resolveModelRef): three branches — unique match / zero matches / ambiguous.
if (opts.modelId !== undefined && opts.provider !== undefined) {
// Pair validation only (resolveModelRef): the pair either names a configured entry or errors.
ref = resolveModelRef(this.projectConfig, opts.modelId, opts.provider);
} else if (this.projectConfig.default_model) {
ref = this.projectConfig.default_model;
@@ -211,6 +215,8 @@ export class Agent {
sessionEnvironment(workspaceDir, sessionId, {
agentId: this.state.agentId,
projectDir: projectDir(this.state.root, this.state.projectId),
provider: modelEntry.provider,
modelId: modelEntry.model_id,
}),
Object.keys(vault),
installedSkills,
@@ -445,9 +451,12 @@ export class Agent {
const { workspaceDir, modelEntry, apiKey, baseUrl, systemPrompt, subagentDepth, vault } = args;
// Child-Agent runner: injected into the run_subagent tool so it doesn't need to
// depend on Agent/Session (breaking a circular dependency). The model can
// optionally choose agentId (omitted = call the current Agent) and modelId
// (omitted = Project default). Precheck errors (depth limit exceeded / agent
// doesn't exist) are expressed as throws, which the Environment collapses to failed.
// optionally choose agentId (omitted = call the current Agent) and the child
// Session's model (omitted = Project default); the model reference is forwarded to
// createSession as-is and must be a complete (provider, model_id) pair — half a
// reference is rejected there rather than being guessed here. Precheck errors (depth
// limit exceeded / agent doesn't exist) are expressed as throws, which the
// Environment collapses to failed.
// Docs: /docs/interfaces § "Subagent interfaces"
const parentAgent = this;
const { root, projectId, agentId: parentAgentId } = this.state;
@@ -455,7 +464,7 @@ export class Agent {
// Spawn and run are separate: the same child Session can run for multiple turns
// (continuing via input_subagent appending a prompt); resource cleanup is
// consolidated in handle.dispose (called by the managing ManagedSubagentSession).
async spawn({ agentId, modelId }) {
async spawn({ agentId, modelId, provider }) {
if (subagentDepth >= MAX_SUBAGENT_DEPTH) {
throw new Error(
`subagent depth limit ${MAX_SUBAGENT_DEPTH} reached; not spawning another subagent`,
@@ -475,9 +484,12 @@ export class Agent {
agentId !== undefined && agentId !== parentAgentId
? await createAgent({ root, projectId, agentId })
: parentAgent;
// The model reference is forwarded as a whole: createSession rejects half a pair, so a
// caller that named only one side gets the same error the CLI and HTTP layers give.
const childSession = await childAgent.createSession({
workspaceDir,
...(modelId !== undefined ? { modelId } : {}),
...(provider !== undefined ? { provider } : {}),
subagentDepth: subagentDepth + 1,
});
// All child-session messages are tagged with an origin (the child Session id,
@@ -82,6 +82,15 @@ export function createSubagentTool(
}
const agentId = typeof args.agent_id === "string" ? args.agent_id : undefined;
const modelId = typeof args.model_id === "string" ? args.model_id : undefined;
const provider = typeof args.provider === "string" ? args.provider : undefined;
// A model is referenced by the complete (provider, model_id) pair — never half of one.
// Caught here rather than in createSession so the model is told which half it left out.
if ((modelId === undefined) !== (provider === undefined)) {
yield* fail(
"[run_subagent error: `model_id` and `provider` must be given together (a model reference is the pair), or both omitted to use the Project default model]",
);
return { stopReason: "failed" };
}
const yieldMs = clampYield(
args.yield_time_ms,
DEFAULT_SUBAGENT_YIELD_MS,
@@ -106,6 +115,7 @@ export function createSubagentTool(
const handle = await runner.spawn({
...(agentId !== undefined ? { agentId } : {}),
...(modelId !== undefined ? { modelId } : {}),
...(provider !== undefined ? { provider } : {}),
});
session = new ManagedSubagentSession(handle);
} catch (err) {
+8 -2
View File
@@ -9,7 +9,8 @@
*
* ```ts
* const agent = await createAgent({ agentId: "default_agent" });
* const session = await agent.createSession({ workspaceDir, modelId });
* // A model reference is always the (provider, model_id) pair; omit both for the Project default.
* const session = await agent.createSession({ workspaceDir, provider, modelId });
* for await (const output of session.run([userText("...")])) { ... }
* ```
*/
@@ -36,7 +37,12 @@ export type {
} from "./engine/context-engine.js";
export { Session } from "./session.js";
export type { SessionConfig } from "./session.js";
export { buildTitlePrompt, generateTitleWithLLM, sanitizeTitle } from "./session-title.js";
export {
buildTitlePrompt,
generateTitleWithLLM,
sanitizeTitle,
stripConversationMarkers,
} from "./session-title.js";
export type { SessionTitleResult } from "./session-title.js";
export { Agent, createAgent } from "./agent.js";
export type { CreateAgentOptions, CreateSessionOptions, ResumeSessionOptions } from "./agent.js";
+7 -1
View File
@@ -200,8 +200,14 @@ export interface SubagentRunner {
spawn(input: {
/** The child Agent's agentId; if omitted, reuses the current Agent (self-invocation). */
agentId?: string;
/** The Model used by the child Session; if omitted, uses the Project's default Model. */
/**
* Upstream model id for the child Session, paired with `provider` — a model reference is
* always the complete pair. Omit both to use the Project's default Model; supplying one
* half without the other is rejected.
*/
modelId?: string;
/** Provider group for `modelId`; required whenever `modelId` is given. */
provider?: string;
}): Promise<SubagentHandle>;
}
@@ -35,7 +35,7 @@ export function formatSessionId(date: Date = new Date()): string {
export function sessionEnvironment(
workspaceDir: string,
sessionId: string,
ids: { agentId: string; projectDir: string },
ids: { agentId: string; projectDir: string; provider: string; modelId: string },
date = new Date(),
): SessionEnvironment {
return {
@@ -43,6 +43,8 @@ export function sessionEnvironment(
cwd: workspaceDir,
agentId: ids.agentId,
projectDir: ids.projectDir,
provider: ids.provider,
modelId: ids.modelId,
platform: process.platform,
osVersion: getOsVersion(),
date: formatLocalDate(date),
+120 -65
View File
@@ -21,7 +21,12 @@
* `context_engine` only consumes OmniMessage; all Uni* protocol details are encapsulated here.
* Docs: /docs/interfaces § "The built-in implementation: GenerativeModel".
*/
import { AutoLLMClient, ThinkingLevel } from "@prismshadow/agenthub";
import {
AutoLLMClient,
EmptyResponseError,
ThinkingLevel,
ToolCallArgumentParseError,
} from "@prismshadow/agenthub";
import type {
ContentItem,
FinishReason,
@@ -45,6 +50,7 @@ import {
} from "../omnimessage/index.js";
import type {
CompleteModelPayload,
Fidelity,
OmniMessage,
StopReason,
TokenCounts,
@@ -75,21 +81,43 @@ function parseToolArguments(raw: string): Record<string, unknown> {
return parsed !== null && typeof parsed === "object" ? (parsed as Record<string, unknown>) : {};
}
/**
* Whether a fidelity payload carries at least one key (mirrors AgentHub's baseClient helper —
* an absent and an empty fidelity are equivalent).
*/
function hasFidelity(fidelity?: Fidelity): boolean {
return fidelity != null && Object.keys(fidelity).length > 0;
}
/**
* Compare two fidelity payloads by value (mirrors AgentHub's baseClient helper). Fidelity
* objects are built with a stable key order by each AgentHub client, so JSON serialization is
* a faithful equality check.
*/
function fidelityEquals(a?: Fidelity, b?: Fidelity): boolean {
return JSON.stringify(a ?? {}) === JSON.stringify(b ?? {});
}
/** Spread helper: attach `fidelity` only when it carries at least one key. */
function fidelityProp(fidelity?: Fidelity): { fidelity?: Fidelity } {
return hasFidelity(fidelity) ? { fidelity } : {};
}
/**
* Maps a complete OmniMessage payload to an AgentHub `ContentItem`.
* Only complete model_msg payloads are supported; `partial_*` is an output-only protocol.
*/
function payloadToContentItem(payload: CompleteModelPayload): ContentItem {
// Provider-fidelity fields (signature / phase) are restored verbatim — some models require
// them when history is replayed back (e.g. Claude thinking signatures, GPT-5 encrypted
// reasoning and phase segmentation); losing them would break Session recovery.
// The provider-fidelity payload is opaque and restored verbatim — some models require it
// when history is replayed back (e.g. Claude thinking signatures, GPT-5 encrypted reasoning
// and phase segmentation, the OpenAI-compatible reasoning field name); losing it would break
// Session recovery.
switch (payload.type) {
case "text":
return {
type: "text",
text: payload.text,
...(payload.phase != null ? { phase: payload.phase } : {}),
...(payload.signature !== undefined ? { signature: payload.signature } : {}),
...fidelityProp(payload.fidelity),
};
case "image_url":
return { type: "image_url", image_url: payload.image_url };
@@ -98,20 +126,20 @@ function payloadToContentItem(payload: CompleteModelPayload): ContentItem {
type: "inline_data",
data: Buffer.from(payload.data, "base64"),
mime_type: payload.mime_type,
...(payload.signature !== undefined ? { signature: payload.signature } : {}),
...fidelityProp(payload.fidelity),
};
case "inline_thinking":
return {
type: "inline_thinking",
data: Buffer.from(payload.data, "base64"),
mime_type: payload.mime_type,
...(payload.signature !== undefined ? { signature: payload.signature } : {}),
...fidelityProp(payload.fidelity),
};
case "thinking":
return {
type: "thinking",
thinking: payload.thinking,
...(payload.signature !== undefined ? { signature: payload.signature } : {}),
...fidelityProp(payload.fidelity),
};
case "tool_call":
return {
@@ -121,7 +149,7 @@ function payloadToContentItem(payload: CompleteModelPayload): ContentItem {
arguments: parseToolArguments(payload.arguments),
// On the way back, strip the uniqueness suffix to restore the provider's original id (see tool-call-ids.ts).
tool_call_id: stripToolCallIdSuffix(payload.tool_call_id),
...(payload.signature !== undefined ? { signature: payload.signature } : {}),
...fidelityProp(payload.fidelity),
};
case "tool_call_output":
return {
@@ -240,8 +268,8 @@ interface ToolCallAccumulator {
toolCallId: string;
/** The original tool_call_id reported by the provider (the attribution key for inbound events). */
providerKey: string;
/** Provider-fidelity field: signature (kept verbatim, produced alongside the complete tool_call). */
signature: string | undefined;
/** Provider-fidelity payload (kept verbatim, produced alongside the complete tool_call). */
fidelity: Fidelity | undefined;
/** Whether this tool_call's complete message has already been emitted eagerly in `pushEvent` (avoids duplicate emission in finish). */
emitted: boolean;
}
@@ -274,12 +302,12 @@ export class EventTranslator {
// Buffers needed for the complete message.
private textBuffer = "";
private thinkingBuffer = "";
// Provider-fidelity fields: the thinking block's signature (a signature
// marks the end of a block) and the current text segment's phase (sticky across segments;
// a differing phase marker starts a new segment) and signature.
private thinkingSignature: string | undefined;
private textPhase: string | null = null;
private textSignature: string | undefined;
// Provider-fidelity payloads of the currently open segments. Segmentation mirrors AgentHub's
// baseClient aggregation: a thinking block is closed by its fidelity payload (a run of equal
// fidelity is one block); a text segment is closed by a `fidelity.signature` and split by a
// differing `fidelity.phase`.
private thinkingFidelity: Fidelity | undefined;
private textFidelity: Fidelity | undefined;
/** Provider id keys saved in order of appearance, so complete tool_calls are emitted in a stable order. */
private toolOrder: string[] = [];
/** provider's original tool_call_id → the accumulator for the **latest** call under that id. */
@@ -306,17 +334,26 @@ export class EventTranslator {
for (const item of event.content_items) {
switch (item.type) {
case "text": {
// Provider fidelity: a phase marker can arrive as an increment with **empty text**
// (e.g. GPT-5's segment markers). A phase differing from the current segment's phase
// starts a new segment — providers split by phase when replaying history, so mixing
// segments would break fidelity. Phase is sticky across segments (a subsequent segment
// with the same phase isn't re-marked).
if (item.phase != null && item.phase !== this.textPhase) {
// Mirrors AgentHub baseClient text aggregation. A `fidelity.phase` marker can arrive
// as an increment with **empty text** (e.g. GPT-5's segment markers): a phase
// differing from the current segment's phase starts a new segment — providers split
// by phase when replaying history, so mixing segments would break fidelity. A
// `fidelity.signature` closes the segment: any further content starts a new one.
// On merge, fidelity keys accumulate ({...current, ...incoming}).
const curPhase = (this.textFidelity as { phase?: unknown } | undefined)?.phase ?? null;
const inPhase = (item.fidelity as { phase?: unknown } | undefined)?.phase ?? null;
const signatureClosed =
(this.textFidelity as { signature?: unknown } | undefined)?.signature != null;
if (
(signatureClosed && (item.text || hasFidelity(item.fidelity))) ||
(inPhase != null && inPhase !== curPhase)
) {
yield* this.flushThinking("completed");
yield* this.flushText("completed");
this.textPhase = item.phase;
}
if (item.signature) this.textSignature = item.signature;
if (hasFidelity(item.fidelity)) {
this.textFidelity = { ...this.textFidelity, ...item.fidelity };
}
if (!item.text) break;
// Type boundary: before a text segment starts, flush any unclosed thinking
// segment, so the complete-message order matches generation order (thinking → text).
@@ -332,16 +369,22 @@ export class EventTranslator {
break;
}
case "thinking": {
// Provider fidelity: a thinking block ends with a signature (Claude's signature_delta
// is empty text + signature; redacted blocks carry sentinel text + signature; GPT-5
// encrypted reasoning is empty text + signature). If thinking content/a new signature
// arrives after a signature is already set, that's a new block — close the current
// segment first, so each block's signature stays independently faithful and blocks
// don't bleed into each other when history is replayed.
if (this.thinkingSignature !== undefined && (item.thinking || item.signature)) {
// Mirrors AgentHub baseClient thinking aggregation: a thinking block is closed by its
// fidelity payload (Claude's signature_delta is empty text + fidelity{signature};
// redacted blocks carry sentinel text + fidelity; GPT-5 encrypted reasoning is empty
// text + fidelity{id, encrypted_content}), and **a run of equal fidelity is one
// block** — OpenAI-compatible clients stamp every delta with the same
// fidelity{reasoning_field}, which must not split blocks. Content arriving after a
// different fidelity is set starts a new block, so each block's fidelity stays
// independently faithful when history is replayed.
if (
hasFidelity(this.thinkingFidelity) &&
!fidelityEquals(this.thinkingFidelity, item.fidelity) &&
(item.thinking || hasFidelity(item.fidelity))
) {
yield* this.flushThinking("completed");
}
if (item.signature) this.thinkingSignature = item.signature;
if (hasFidelity(item.fidelity)) this.thinkingFidelity = item.fidelity;
if (!item.thinking) break;
// Type boundary: before a thinking segment starts, flush any unclosed text
// segment, so the complete-message order matches generation order (text → thinking).
@@ -365,7 +408,7 @@ export class EventTranslator {
if (item.tool_call_id) this.activeToolCallId = item.tool_call_id;
const acc = this.ensureTool(providerKey, item.name);
if (item.name) acc.name = item.name;
if (item.signature) acc.signature = item.signature;
if (hasFidelity(item.fidelity)) acc.fidelity = item.fidelity;
// Externally always use the uniqueness-resolved id (with a `#n` suffix on provider id collisions), matching the complete tool_call.
if (!this.toolStarted.has(acc.toolCallId)) {
// Type boundary: before a new tool_call starts, flush any unclosed thinking/text
@@ -418,7 +461,7 @@ export class EventTranslator {
if (!acc || acc.emitted) acc = this.createTool(item.tool_call_id, item.name);
acc.name = item.name;
acc.completeArgs = JSON.stringify(item.arguments ?? {});
if (item.signature) acc.signature = item.signature;
if (hasFidelity(item.fidelity)) acc.fidelity = item.fidelity;
yield* this.emitCompleteTool(acc);
break;
}
@@ -507,7 +550,7 @@ export class EventTranslator {
arguments: acc.completeArgs ?? acc.argsBuffer,
toolCallId: acc.toolCallId,
stopReason,
...(acc.signature !== undefined ? { signature: acc.signature } : {}),
...(acc.fidelity !== undefined ? { fidelity: acc.fidelity } : {}),
});
// activeToolCallId holds the provider id key (used to attribute id-less deltas); reset it by providerKey.
if (this.activeToolCallId === acc.providerKey) {
@@ -536,14 +579,12 @@ export class EventTranslator {
yield partialThinking("stop", "", stopReason);
this.thinkingStarted = false;
}
// A thinking block with empty text but a signature (GPT-5 encrypted reasoning) still
// produces a complete message — the signature is required when replaying history.
if (this.thinkingBuffer || this.thinkingSignature !== undefined) {
yield thinkingMessage(this.thinkingBuffer, stopReason, {
...(this.thinkingSignature !== undefined ? { signature: this.thinkingSignature } : {}),
});
// A thinking block with empty text but a fidelity payload (GPT-5 encrypted reasoning)
// still produces a complete message — the fidelity is required when replaying history.
if (this.thinkingBuffer || hasFidelity(this.thinkingFidelity)) {
yield thinkingMessage(this.thinkingBuffer, stopReason, this.thinkingFidelity);
this.thinkingBuffer = "";
this.thinkingSignature = undefined;
this.thinkingFidelity = undefined;
}
}
@@ -559,18 +600,14 @@ export class EventTranslator {
yield partialText("stop", "", stopReason);
this.textStarted = false;
}
// A text segment with empty text but a signature (e.g. Gemini carrying a thoughtSignature on
// a text part) still produces a complete message — aligned with flushThinking, so the
// signature isn't lost or leaked into a later segment just because the buffer is empty.
if (this.textBuffer || this.textSignature !== undefined) {
yield assistantText(this.textBuffer, stopReason, {
...(this.textPhase != null ? { phase: this.textPhase } : {}),
...(this.textSignature !== undefined ? { signature: this.textSignature } : {}),
});
// A text segment with empty text but a fidelity payload (e.g. Gemini carrying a
// thoughtSignature on a text part, or a GPT-5 phase marker with no text) still produces a
// complete message — aligned with flushThinking, so the fidelity isn't lost or leaked into
// a later segment just because the buffer is empty.
if (this.textBuffer || hasFidelity(this.textFidelity)) {
yield assistantText(this.textBuffer, stopReason, this.textFidelity);
this.textBuffer = "";
this.textSignature = undefined;
// textPhase is sticky across segments: a later segment with the same phase isn't
// re-marked; it's updated when a different phase marker appears.
this.textFidelity = undefined;
}
}
@@ -601,7 +638,7 @@ export class EventTranslator {
completeArgs: null,
toolCallId: this.toolCallIds.allocate(providerKey),
providerKey,
signature: undefined,
fidelity: undefined,
emitted: false,
};
this.tools.set(providerKey, acc);
@@ -665,19 +702,37 @@ export function translateEvents(
// ---------------------------------------------------------------------------
/**
* Determines whether an error is an AgentHub / Provider response JSON parse error.
* Determines whether an error is an AgentHub / Provider "response delivered but unusable"
* parse or validation error. Two shapes (@prismshadow/agenthub 0.4.x):
*
* AgentHub uses `JSON.parse` internally to parse response bodies, and a parse failure throws a
* `SyntaxError`; hence we judge directly by exception type (the `name` check also covers
* cross-realm or deserialization-reconstructed errors, and we probe down the `cause` chain for
* wrapped errors). This is not an auth/parameter failure but an incomplete LLM Request, and
* should end with `malformed` and be handed to the engine to retry.
* - A raw `SyntaxError` from `JSON.parse` on a response body;
* - AgentHub's own error classes: `ToolCallArgumentParseError` (streamed tool-call arguments
* are not valid JSON — e.g. a stream truncated mid-arguments) and `EmptyResponseError`
* (a completed response carrying thinking only, which cannot be replayed).
*
* In every case the turn was **not committed** to AgentHub history: this is not an
* auth/parameter failure but an incomplete LLM Request, and should end with `malformed` and
* be handed to the engine to reconnect and retry. Judged by exception type with a `name`
* fallback (covers cross-realm or deserialization-reconstructed errors), probing down the
* `cause` chain for wrapped errors.
*/
export function isMalformedJsonParseError(error: unknown): boolean {
if (error == null) return false;
if (error instanceof SyntaxError) return true;
if (
error instanceof SyntaxError ||
error instanceof ToolCallArgumentParseError ||
error instanceof EmptyResponseError
) {
return true;
}
const err = error as { name?: string; cause?: unknown };
if (err.name === "SyntaxError") return true;
if (
err.name === "SyntaxError" ||
err.name === "ToolCallArgumentParseError" ||
err.name === "EmptyResponseError"
) {
return true;
}
if (err.cause && err.cause !== error) {
return isMalformedJsonParseError(err.cause);
}
@@ -692,7 +747,7 @@ export function isMalformedJsonParseError(error: unknown): boolean {
* usage_metadata|finish_reason"). This is not an auth/parameter failure but an incomplete LLM
* Request, and should end with `malformed` and be handed to the engine to reconnect and retry.
* AgentHub doesn't provide an error type for this, so we match by message prefix
* (@prismshadow/agenthub 0.3.x), probing down the `cause` chain.
* (verified against @prismshadow/agenthub 0.4.x), probing down the `cause` chain.
*/
export function isIncompleteStreamError(error: unknown): boolean {
if (error == null) return false;
+19 -20
View File
@@ -13,6 +13,7 @@ import type {
CompactionMode,
CompactionReason,
EventMessage,
Fidelity,
ImageUrlPayload,
InlineDataPayload,
InlineThinkingPayload,
@@ -61,30 +62,28 @@ export function sessionMeta(payload: SessionMetaPayload): SessionMetaMessage {
// Complete model_msg -----------------------------------------------------------
/**
* Provider fidelity fields: kept as-is and restored verbatim on replay.
* Builder convention: positional-argument-style builders carry these in a trailing `fidelity`
* object (narrowed via Pick per payload type — e.g. thinking only has signature); object-argument-
* style builders (toolCall) flatten `fidelity` fields into the parameter object alongside
* `stopReason`, mirroring the payload structure directly.
* Provider-fidelity payload (opaque, see `Fidelity` in types.ts): kept as-is and restored
* verbatim on replay. Builder convention: positional-argument-style builders take it as a
* trailing `fidelity` object; object-argument-style builders (toolCall) carry it in the
* parameter object alongside `stopReason`. An empty object is treated as absent — the
* payload field is only set when the fidelity carries at least one key.
*/
export interface FidelityFields {
phase?: string | null;
signature?: string;
function fidelityProp(fidelity?: Fidelity): { fidelity?: Fidelity } {
return fidelity !== undefined && Object.keys(fidelity).length > 0 ? { fidelity } : {};
}
export function textMessage(
role: Role,
text: string,
stopReason: StopReason = "completed",
fidelity?: FidelityFields,
fidelity?: Fidelity,
): OmniMessage<TextPayload> {
return model({
type: "text",
role,
text,
stop_reason: stopReason,
...(fidelity?.phase != null ? { phase: fidelity.phase } : {}),
...(fidelity?.signature !== undefined ? { signature: fidelity.signature } : {}),
...fidelityProp(fidelity),
});
}
@@ -93,7 +92,7 @@ export const userText = (text: string): OmniMessage<TextPayload> => textMessage(
export const assistantText = (
text: string,
stopReason: StopReason = "completed",
fidelity?: FidelityFields,
fidelity?: Fidelity,
): OmniMessage<TextPayload> => textMessage("assistant", text, stopReason, fidelity);
export function imageUrlMessage(imageUrl: string): OmniMessage<ImageUrlPayload> {
@@ -109,7 +108,7 @@ export function inlineData(
role: Role,
data: string,
mimeType: string,
fidelity?: Pick<FidelityFields, "signature">,
fidelity?: Fidelity,
): OmniMessage<InlineDataPayload> {
return model({
type: "inline_data",
@@ -117,28 +116,28 @@ export function inlineData(
data,
mime_type: mimeType,
stop_reason: "completed",
...(fidelity?.signature !== undefined ? { signature: fidelity.signature } : {}),
...fidelityProp(fidelity),
});
}
export function thinkingMessage(
thinking: string,
stopReason: StopReason = "completed",
fidelity?: Pick<FidelityFields, "signature">,
fidelity?: Fidelity,
): OmniMessage<ThinkingPayload> {
return model({
type: "thinking",
role: "assistant",
thinking,
stop_reason: stopReason,
...(fidelity?.signature !== undefined ? { signature: fidelity.signature } : {}),
...fidelityProp(fidelity),
});
}
export function inlineThinking(
data: string,
mimeType: string,
fidelity?: Pick<FidelityFields, "signature">,
fidelity?: Fidelity,
): OmniMessage<InlineThinkingPayload> {
return model({
type: "inline_thinking",
@@ -146,7 +145,7 @@ export function inlineThinking(
data,
mime_type: mimeType,
stop_reason: "completed",
...(fidelity?.signature !== undefined ? { signature: fidelity.signature } : {}),
...fidelityProp(fidelity),
});
}
@@ -155,7 +154,7 @@ export function toolCall(args: {
arguments: string;
toolCallId: string;
stopReason?: StopReason;
signature?: string;
fidelity?: Fidelity;
}): OmniMessage<ToolCallPayload> {
return model({
type: "tool_call",
@@ -164,7 +163,7 @@ export function toolCall(args: {
arguments: args.arguments,
tool_call_id: args.toolCallId,
stop_reason: args.stopReason ?? "completed",
...(args.signature !== undefined ? { signature: args.signature } : {}),
...fidelityProp(args.fidelity),
});
}
+23 -14
View File
@@ -99,15 +99,23 @@ export interface SessionMetaPayload {
// ---------------------------------------------------------------------------
// Docs: /docs/omni-message § "model_msg: complete payloads"
/**
* Provider-fidelity payload (mirrors AgentHub's `Fidelity`): an arbitrary JSON-style object of
* wire-level data the LLM client records to reproduce the original message on replay — thinking
* signatures, phase labels, encrypted reasoning, the upstream reasoning field name, etc. Opaque
* to PenguinHarness: written to the Trace as-is and passed back verbatim; some models **require**
* it when history is replayed (e.g. Claude thinking signatures, GPT-5 encrypted reasoning) —
* losing it breaks Session resumption.
*/
export type Fidelity = Record<string, unknown>;
export interface TextPayload {
type: "text";
role: Role;
text: string;
stop_reason?: StopReason;
/** Provider fidelity field: text phase marker (e.g. GPT-5 segments by phase), kept as-is and restored verbatim. */
phase?: string | null;
/** Provider fidelity field: signature, kept as-is and restored verbatim. */
signature?: string;
/** Provider-fidelity payload (e.g. `phase` for GPT-5 segment markers, `signature`), kept as-is and restored verbatim. */
fidelity?: Fidelity;
}
export interface ImageUrlPayload {
@@ -125,8 +133,8 @@ export interface InlineDataPayload {
data: string;
mime_type: string;
stop_reason?: StopReason;
/** Provider fidelity field: signature, kept as-is and restored verbatim. */
signature?: string;
/** Provider-fidelity payload, kept as-is and restored verbatim. */
fidelity?: Fidelity;
}
export interface ThinkingPayload {
@@ -135,11 +143,12 @@ export interface ThinkingPayload {
thinking: string;
stop_reason?: StopReason;
/**
* Provider fidelity field: thinking-block signature (Claude thinking blocks / redacted
* thinking, GPT-5 encrypted reasoning, etc. — **required** when some models replay history),
* kept as-is and restored verbatim — losing it breaks Session resumption.
* Provider-fidelity payload closing the thinking block (Claude thinking signatures / redacted
* thinking, GPT-5 encrypted reasoning, the OpenAI-compatible reasoning field name, etc. —
* **required** when some models replay history), kept as-is and restored verbatim — losing it
* breaks Session resumption.
*/
signature?: string;
fidelity?: Fidelity;
}
export interface InlineThinkingPayload {
@@ -149,8 +158,8 @@ export interface InlineThinkingPayload {
data: string;
mime_type: string;
stop_reason?: StopReason;
/** Provider fidelity field: signature, kept as-is and restored verbatim. */
signature?: string;
/** Provider-fidelity payload, kept as-is and restored verbatim. */
fidelity?: Fidelity;
}
export interface ToolCallPayload {
@@ -161,8 +170,8 @@ export interface ToolCallPayload {
arguments: string;
tool_call_id: string;
stop_reason?: StopReason;
/** Provider fidelity field: signature, kept as-is and restored verbatim. */
signature?: string;
/** Provider-fidelity payload, kept as-is and restored verbatim. */
fidelity?: Fidelity;
}
export interface ToolCallOutputPayload {
+37 -4
View File
@@ -22,6 +22,31 @@ const EXCERPT_MAX_CHARS = 2000;
/** Cap on title length (fallback truncation for when the model occasionally ignores the constraint). */
const TITLE_MAX_CHARS = 30;
/**
* Special message markers that must never leak into a title. These are machine-inserted
* XML-ish blocks (a skill invocation wraps the body in `<use_skills>…</use_skills>`, a
* subagent handoff / scheduled task prepend their own blocks) — meaningful to the runtime,
* noise in a title. The list is a fixed allowlist so ordinary angle-bracket text (e.g. a
* user pasting `<div>`) is left untouched.
*/
const MARKER_TAGS = ["use_skills", "handoff_from", "scheduled_task"];
/**
* Strips machine-inserted marker blocks (see MARKER_TAGS) from conversation text so titles are
* built from the human-meaningful body only — both the material sent to the model and the
* fallback derived from the raw first message. Removes the paired `<tag>…</tag>` block (any
* inner content, across lines) and any stray unpaired `<tag>` / `</tag>` left behind.
*/
export function stripConversationMarkers(text: string): string {
let out = text;
for (const tag of MARKER_TAGS) {
out = out
.replace(new RegExp(`<${tag}>[\\s\\S]*?</${tag}>`, "g"), "")
.replace(new RegExp(`</?${tag}>`, "g"), "");
}
return out.trim();
}
export interface SessionTitleResult {
/** The sanitized title; null when material is insufficient, the request fails, or the output is empty. */
title: string | null;
@@ -44,6 +69,7 @@ export function buildTitlePrompt(userExcerpt: string, assistantExcerpt: string):
"- Write the title in the SAME language the user is using.",
"- Keep it short: at most 6 words, or ~16 characters for CJK.",
"- Output ONLY the title text — no quotes, no trailing punctuation, no explanation.",
"- Answer immediately — do not think aloud or produce chain-of-thought.",
"",
"[User]",
clip(userExcerpt),
@@ -51,12 +77,15 @@ export function buildTitlePrompt(userExcerpt: string, assistantExcerpt: string):
if (assistantExcerpt.trim()) {
lines.push("", "[Assistant]", clip(assistantExcerpt));
}
// The trailing empty think block makes many reasoning models treat their thinking phase
// as already closed, so the one-off request spends its budget on the title itself.
lines.push("", "<think></think>");
return lines.join("\n");
}
/** Sanitizes model output into a title: strips leading/trailing quotes/brackets and trailing punctuation (until stable), collapses whitespace, and truncates if too long; returns null for an empty result. */
/** Sanitizes model output into a title: strips any leaked marker blocks and leading/trailing quotes/brackets and trailing punctuation (until stable), collapses whitespace, and truncates if too long; returns null for an empty result. */
export function sanitizeTitle(raw: string): string | null {
let t = raw.replace(/\s+/g, " ").trim();
let t = stripConversationMarkers(raw).replace(/\s+/g, " ").trim();
// Stripping quotes can expose more punctuation underneath (or vice versa), so strip repeatedly until stable.
for (let prev = ""; prev !== t;) {
prev = t;
@@ -81,10 +110,14 @@ export async function generateTitleWithLLM(
llm: LLMInterface,
args: { userText: string; assistantText: string; signal?: AbortSignal },
): Promise<SessionTitleResult> {
if (!args.userText.trim()) {
// Strip machine markers before the model ever sees the material, so a skill-invocation
// block (or handoff / scheduled-task marker) can't bleed into the generated title.
const userMaterial = stripConversationMarkers(args.userText);
const assistantMaterial = stripConversationMarkers(args.assistantText);
if (!userMaterial.trim()) {
return { title: null, usage: null };
}
const prompt = buildTitlePrompt(args.userText, args.assistantText);
const prompt = buildTitlePrompt(userMaterial, assistantMaterial);
const gen = llm.streamGenerate({
newMessages: [userText(prompt)],
...(args.signal ? { signal: args.signal } : {}),
+27 -6
View File
@@ -65,15 +65,26 @@ export interface SessionConfig {
inputImagesDir?: string;
}
/** Cap on captured title material (chars per side, matching buildTitlePrompt's truncation); stops accumulating once exceeded. */
const TITLE_MATERIAL_LIMIT = 2000;
/**
* Caps on captured title material (chars per side); accumulation stops once exceeded. The
* assistant body is capped tighter: a title only needs the opening of the answer, and hosts
* may start generating as soon as this much body text has streamed (see the Web server's
* early trigger) — a long answer would otherwise overrun the material.
*/
const TITLE_USER_MATERIAL_LIMIT = 2000;
const TITLE_ASSISTANT_MATERIAL_LIMIT = 1000;
/**
* Accumulates title material: the body text of complete text messages from the main session
* (no origin) — thinking and tool calls naturally don't count — and stops once the cap is hit.
*/
function appendTitleText(base: string, msg: OmniMessage, role: "user" | "assistant"): string {
if (base.length >= TITLE_MATERIAL_LIMIT) return base;
function appendTitleText(
base: string,
msg: OmniMessage,
role: "user" | "assistant",
limit: number,
): string {
if (base.length >= limit) return base;
if (msg.origin && msg.origin.length > 0) return base;
const p = msg.payload as { type?: string; role?: string; text?: string };
if (msg.type !== "model_msg" || p.type !== "text" || p.role !== role || !p.text) return base;
@@ -153,12 +164,22 @@ export class Session {
const capture = !this.titleMaterialFrozen;
if (capture) {
for (const m of newMessages) {
this.titleUserText = appendTitleText(this.titleUserText, m, "user");
this.titleUserText = appendTitleText(
this.titleUserText,
m,
"user",
TITLE_USER_MATERIAL_LIMIT,
);
}
}
for await (const msg of this.engine.run(newMessages, opts)) {
if (capture) {
this.titleAssistantText = appendTitleText(this.titleAssistantText, msg, "assistant");
this.titleAssistantText = appendTitleText(
this.titleAssistantText,
msg,
"assistant",
TITLE_ASSISTANT_MATERIAL_LIMIT,
);
}
yield msg;
}
+10
View File
@@ -26,6 +26,8 @@ import {
VAULT_KEYS_PLACEHOLDER,
SKILL_METADATA_PLACEHOLDER,
CWD_PLACEHOLDER,
PROVIDER_PLACEHOLDER,
MODEL_ID_PLACEHOLDER,
DATE_PLACEHOLDER,
defaultAgentsMd,
defaultSystemConfig,
@@ -83,6 +85,10 @@ export interface SessionEnvironmentValues {
agentId: string;
/** Absolute path to this Project's directory (system Prompt placeholder {{PROJECT_DIR}}; Agent State/scratchpad paths are derived from it). */
projectDir: string;
/** The session model's provider group (system Prompt placeholder {{PROVIDER}}; paired with modelId to form the model reference). */
provider: string;
/** The session model's upstream model id (system Prompt placeholder {{MODEL_ID}}). */
modelId: string;
platform: string;
osVersion: string;
date: string;
@@ -371,6 +377,10 @@ export function assembleSystemPrompt(
.join(sessionEnvironment?.sessionId ?? "")
.split(CWD_PLACEHOLDER)
.join(sessionEnvironment?.cwd ?? "")
.split(PROVIDER_PLACEHOLDER)
.join(sessionEnvironment?.provider ?? "")
.split(MODEL_ID_PLACEHOLDER)
.join(sessionEnvironment?.modelId ?? "")
.split(PLATFORM_PLACEHOLDER)
.join(sessionEnvironment?.platform ?? "")
.split(OS_VERSION_PLACEHOLDER)
+12 -3
View File
@@ -28,6 +28,8 @@ export const SESSION_ID_PLACEHOLDER = "{{SESSION_ID}}";
export const CWD_PLACEHOLDER = "{{CWD}}";
export const AGENT_ID_PLACEHOLDER = "{{AGENT_ID}}";
export const PROJECT_DIR_PLACEHOLDER = "{{PROJECT_DIR}}";
export const PROVIDER_PLACEHOLDER = "{{PROVIDER}}";
export const MODEL_ID_PLACEHOLDER = "{{MODEL_ID}}";
export const PLATFORM_PLACEHOLDER = "{{PLATFORM}}";
export const OS_VERSION_PLACEHOLDER = "{{OS_VERSION}}";
export const DATE_PLACEHOLDER = "{{DATE}}";
@@ -141,9 +143,11 @@ Skills are reusable instruction packages stored under <project_dir>/agents/<agen
- Platform: {{PLATFORM}}
- OS Version: {{OS_VERSION}}
- Date: {{DATE}}
- CWD: {{CWD}}
- Agent ID: {{AGENT_ID}}
- Project Dir: {{PROJECT_DIR}}
- Agent ID: {{AGENT_ID}}
- CWD: {{CWD}}
- Provider: {{PROVIDER}}
- Model ID: {{MODEL_ID}}
- Session ID: {{SESSION_ID}}`;
/**
@@ -247,7 +251,12 @@ function defaultBuiltinTools(): ToolDefinitionConfig[] {
model_id: {
type: "string",
description:
"Which model the subagent should use; defaults to the Project default model when omitted.",
"Which model the subagent should use, as the upstream model id. Must be given together with provider — a model is always referenced by the pair. Omit both to use the Project default model.",
},
provider: {
type: "string",
description:
"The provider group that model_id belongs to (see the Environment section's Provider). Required whenever model_id is given.",
},
yield_time_ms: {
type: "number",
+522 -138
View File
@@ -1,7 +1,7 @@
/**
* Built-in model catalog (single source of truth): official chat models that AgentHub can
* auto-route, shared by core's default config, server's initial config, and web/cli display.
* Data verified as of 2026-07-10.
* Data verified as of 2026-07-10 (Qwen Token Plan entries: 2026-07-20, per the plan's docs).
* Docs: packages/docs/content/models.{zh,en}.md (site path /docs/models) documents the
* provider groups and credential resolution described here.
*
@@ -17,10 +17,10 @@
*
* Scope: excludes deepseek-chat / deepseek-reasoner legacy aliases that AgentHub cannot
* auto-route (deprecated 2026-07-24), glm-5v-turbo (image input unsupported by AgentHub's GLM
* client), non-chat models (embedding / image generation / TTS), and Bedrock plus
* OpenRouter / SiliconFlow gateway mirror ids. Every model id in this catalog can be
* auto-routed by AgentHub via substring matching, so none set client_type; only custom
* OpenAI-protocol models need `client_type: "openai"`.
* client), non-chat models (embedding / image generation / TTS), and Bedrock. Direct-vendor
* ids are auto-routed by AgentHub and leave client_type unset; gateway entries (OpenRouter /
* SiliconFlow / Qwen Token Plan) can't be auto-routed, so they set `client_type: "openai"`
* and inline their preset base URL.
*
* This file imports no Node built-ins (type-only imports only), so it can be bundled directly
* for the browser.
@@ -41,8 +41,9 @@ export interface ModelProviderInfo {
/** Vendor's model list / docs page URL (frontend's "add model" dialog links this as "get model id"); none for custom. */
modelsUrl?: string;
/**
* Gateway's OpenAI-compatible endpoint (openrouter / siliconflow): used by the frontend's
* "add model" dialog to prefill base URL by group; left blank for direct vendors and custom.
* Gateway's OpenAI-compatible endpoint (openrouter / siliconflow / qwen-token-plan): used by
* the frontend's "add model" dialog to prefill base URL by group; left blank for direct
* vendors and custom.
*/
gatewayBaseUrl?: string;
}
@@ -66,11 +67,15 @@ export interface ModelCatalogEntry {
/** Each gateway's OpenAI-compatible endpoint (preset base URL for gateway models; also used as the provider's gatewayBaseUrl). */
const OPENROUTER_BASE_URL = "https://openrouter.ai/api/v1";
const SILICONFLOW_BASE_URL = "https://api.siliconflow.cn/v1";
const QWEN_TOKEN_PLAN_BASE_URL =
"https://token-plan.cn-beijing.maas.aliyuncs.com/compatible-mode/v1";
const QWEN_PAYG_BASE_URL = "https://dashscope.aliyuncs.com/compatible-mode/v1";
const FIREWORKS_BASE_URL = "https://api.fireworks.ai/inference/v1";
/**
* Provider list (web model page groups in this order): DeepSeek first (the default model's
* provider), followed by the OpenRouter and SiliconFlow gateways, then Google Gemini before
* Anthropic; custom groups custom OpenAI-protocol models and comes last.
* provider), followed by the OpenRouter, SiliconFlow, and Qwen Token Plan gateways, then
* Google Gemini before Anthropic; custom groups custom OpenAI-protocol models and comes last.
*/
export const MODEL_PROVIDERS: ModelProviderInfo[] = [
{
@@ -94,6 +99,15 @@ export const MODEL_PROVIDERS: ModelProviderInfo[] = [
modelsUrl: "https://openrouter.ai/models",
gatewayBaseUrl: OPENROUTER_BASE_URL,
},
{
id: "fireworks",
label: "Fireworks AI",
envKey: "OPENAI_API_KEY",
envBaseUrlKey: "OPENAI_BASE_URL",
apiKeyUrl: "https://app.fireworks.ai/settings/users/api-keys",
modelsUrl: "https://app.fireworks.ai/models",
gatewayBaseUrl: FIREWORKS_BASE_URL,
},
{
id: "siliconflow",
label: "SiliconFlow",
@@ -103,6 +117,25 @@ export const MODEL_PROVIDERS: ModelProviderInfo[] = [
modelsUrl: "https://cloud.siliconflow.cn/models",
gatewayBaseUrl: SILICONFLOW_BASE_URL,
},
{
id: "qwen-token-plan",
label: "Qwen Token Plan",
envKey: "OPENAI_API_KEY",
envBaseUrlKey: "OPENAI_BASE_URL",
apiKeyUrl: "https://platform.qianwenai.com/pricing/token-plan",
modelsUrl:
"https://platform.qianwenai.com/docs/token-plan/personal/token-plan-personal-overview",
gatewayBaseUrl: QWEN_TOKEN_PLAN_BASE_URL,
},
{
id: "qwen-pay-as-you-go",
label: "Qwen Pay-As-You-Go",
envKey: "OPENAI_API_KEY",
envBaseUrlKey: "OPENAI_BASE_URL",
apiKeyUrl: "https://platform.qianwenai.com/docs/api-reference/preparation/api-key",
modelsUrl: "https://www.qianwenai.com/models",
gatewayBaseUrl: QWEN_PAYG_BASE_URL,
},
{
id: "google",
label: "Google Gemini",
@@ -161,17 +194,14 @@ function usd(cacheRead: number, cacheWrite: number, output: number): ModelPricin
return { unit: "usd_per_mtok", cache_read: cacheRead, cache_write: cacheWrite, output };
}
/** Built-in model catalog (clustered by provider; within each provider, ordered by capability/price, highest first). */
/**
* Built-in model catalog, clustered by provider. Within each provider, entries are in
* dictionary order by model id, except that newer versions of the same series come first
* (e.g. gpt-5.6-* before gpt-5.5, claude-opus-4.8 before 4.7, glm-5.2 before glm-5). The
* order is precomputed by hand right here — no runtime sorting anywhere.
*/
export const MODEL_CATALOG: ModelCatalogEntry[] = [
// -- DeepSeek (official CNY pricing: cache hit / cache miss / output) --
{
modelId: "deepseek-v4-pro",
displayName: "DeepSeek V4 Pro",
provider: "deepseek",
contextWindow: 1000000,
pricing: cny(0.025, 3, 6),
supportsVision: false,
},
{
modelId: "deepseek-v4-flash",
displayName: "DeepSeek V4 Flash",
@@ -180,7 +210,438 @@ export const MODEL_CATALOG: ModelCatalogEntry[] = [
pricing: cny(0.02, 1, 2),
supportsVision: false,
},
// —— Anthropic ——
{
modelId: "deepseek-v4-pro",
displayName: "DeepSeek V4 Pro",
provider: "deepseek",
contextWindow: 1000000,
pricing: cny(0.025, 3, 6),
supportsVision: false,
},
// -- OpenRouter (gateway: OpenAI-compatible protocol, preset base URL). Entries added
// 2026-07-20 list no cache pricing on their OpenRouter pages, so cache_read carries the
// standard input price; their :free tier stores a genuine $0 price (not "unknown"), so
// costs correctly compute to 0. GPT models are uniformly vision-capable (OpenAI
// product-line policy) even where the gateway page omits the modality. --
{
modelId: "anthropic/claude-fable-5",
displayName: "Claude Fable 5",
provider: "openrouter",
contextWindow: 1000000,
pricing: usd(10, 10, 50),
supportsVision: true,
clientType: "openai",
baseUrl: OPENROUTER_BASE_URL,
},
{
modelId: "anthropic/claude-opus-4.8",
displayName: "Claude Opus 4.8",
provider: "openrouter",
contextWindow: 1000000,
pricing: usd(5, 5, 25),
supportsVision: true,
clientType: "openai",
baseUrl: OPENROUTER_BASE_URL,
},
{
modelId: "anthropic/claude-opus-4.7",
displayName: "Claude Opus 4.7",
provider: "openrouter",
contextWindow: 1000000,
pricing: usd(5, 5, 25),
supportsVision: true,
clientType: "openai",
baseUrl: OPENROUTER_BASE_URL,
},
{
modelId: "anthropic/claude-sonnet-5",
displayName: "Claude Sonnet 5",
provider: "openrouter",
contextWindow: 1000000,
pricing: usd(2, 2, 10),
supportsVision: true,
clientType: "openai",
baseUrl: OPENROUTER_BASE_URL,
},
{
modelId: "deepseek/deepseek-v4-flash",
displayName: "DeepSeek V4 Flash",
provider: "openrouter",
contextWindow: 1000000,
pricing: usd(0.09, 0.09, 0.18),
supportsVision: false,
clientType: "openai",
baseUrl: OPENROUTER_BASE_URL,
},
{
modelId: "deepseek/deepseek-v4-pro",
displayName: "DeepSeek V4 Pro",
provider: "openrouter",
contextWindow: 1000000,
pricing: usd(0.435, 0.435, 0.87),
supportsVision: false,
clientType: "openai",
baseUrl: OPENROUTER_BASE_URL,
},
{
modelId: "google/gemini-3.5-flash",
displayName: "Gemini 3.5 Flash",
provider: "openrouter",
contextWindow: 1000000,
pricing: usd(1.5, 1.5, 9),
supportsVision: true,
clientType: "openai",
baseUrl: OPENROUTER_BASE_URL,
},
{
// No official separate cache price published: cache_read uses the standard input price (no discount assumed).
modelId: "minimax/minimax-m3",
displayName: "MiniMax M3",
provider: "openrouter",
contextWindow: 1048576,
pricing: usd(0.06, 0.3, 1.2),
supportsVision: true,
clientType: "openai",
baseUrl: OPENROUTER_BASE_URL,
},
{
modelId: "moonshotai/kimi-k3",
displayName: "Kimi K3",
provider: "openrouter",
contextWindow: 1000000,
pricing: usd(3, 3, 15),
supportsVision: true,
clientType: "openai",
baseUrl: OPENROUTER_BASE_URL,
},
{
modelId: "nvidia/nemotron-3-ultra-550b-a55b:free",
displayName: "Nemotron 3 Ultra (free)",
provider: "openrouter",
contextWindow: 1000000,
pricing: usd(0, 0, 0),
supportsVision: false,
clientType: "openai",
baseUrl: OPENROUTER_BASE_URL,
},
{
modelId: "openai/gpt-5.6-sol",
displayName: "GPT-5.6 Sol",
provider: "openrouter",
contextWindow: 1000000,
pricing: usd(5, 5, 30),
supportsVision: true,
clientType: "openai",
baseUrl: OPENROUTER_BASE_URL,
},
{
modelId: "openai/gpt-5.6-terra",
displayName: "GPT-5.6 Terra",
provider: "openrouter",
contextWindow: 1000000,
pricing: usd(2.5, 2.5, 15),
supportsVision: true,
clientType: "openai",
baseUrl: OPENROUTER_BASE_URL,
},
{
modelId: "openai/gpt-5.5",
displayName: "GPT-5.5",
provider: "openrouter",
contextWindow: 1000000,
pricing: usd(5, 5, 30),
supportsVision: true,
clientType: "openai",
baseUrl: OPENROUTER_BASE_URL,
},
{
// No official separate cache price published: cache_read uses the standard input price.
modelId: "stepfun/step-3.7-flash",
displayName: "Step 3.7 Flash",
provider: "openrouter",
contextWindow: 256000,
pricing: usd(0.04, 0.2, 1.15),
supportsVision: true,
clientType: "openai",
baseUrl: OPENROUTER_BASE_URL,
},
{
modelId: "tencent/hy3",
displayName: "Hy3",
provider: "openrouter",
contextWindow: 262144,
pricing: usd(0.035, 0.14, 0.58),
supportsVision: false,
clientType: "openai",
baseUrl: OPENROUTER_BASE_URL,
},
{
modelId: "x-ai/grok-4.5",
displayName: "Grok 4.5",
provider: "openrouter",
contextWindow: 500000,
pricing: usd(2, 2, 6),
supportsVision: true,
clientType: "openai",
baseUrl: OPENROUTER_BASE_URL,
},
{
modelId: "xiaomi/mimo-v2.5",
displayName: "MiMo-V2.5",
provider: "openrouter",
contextWindow: 1048576,
pricing: usd(0.0028, 0.14, 0.28),
supportsVision: true,
clientType: "openai",
baseUrl: OPENROUTER_BASE_URL,
},
{
modelId: "z-ai/glm-5.2",
displayName: "GLM-5.2",
provider: "openrouter",
contextWindow: 1000000,
pricing: usd(0.93, 0.93, 3),
supportsVision: false,
clientType: "openai",
baseUrl: OPENROUTER_BASE_URL,
},
// -- Fireworks AI (gateway, standard serverless USD pricing: cached input / uncached
// input / output from each model's page; API ids use the accounts/fireworks/models/<slug>
// form) --
{
modelId: "accounts/fireworks/models/deepseek-v4-flash",
displayName: "DeepSeek V4 Flash",
provider: "fireworks",
contextWindow: 1000000,
pricing: usd(0.03, 0.14, 0.28),
supportsVision: false,
clientType: "openai",
baseUrl: FIREWORKS_BASE_URL,
},
{
modelId: "accounts/fireworks/models/deepseek-v4-pro",
displayName: "DeepSeek V4 Pro",
provider: "fireworks",
contextWindow: 1000000,
pricing: usd(0.15, 1.74, 3.48),
supportsVision: false,
clientType: "openai",
baseUrl: FIREWORKS_BASE_URL,
},
{
modelId: "accounts/fireworks/models/glm-5p2",
displayName: "GLM-5.2",
provider: "fireworks",
contextWindow: 1000000,
pricing: usd(0.14, 1.4, 4.4),
supportsVision: false,
clientType: "openai",
baseUrl: FIREWORKS_BASE_URL,
},
{
modelId: "accounts/fireworks/models/kimi-k2p7-code",
displayName: "Kimi K2.7 Code",
provider: "fireworks",
contextWindow: 262144,
pricing: usd(0.19, 0.95, 4),
supportsVision: true,
clientType: "openai",
baseUrl: FIREWORKS_BASE_URL,
},
{
modelId: "accounts/fireworks/models/minimax-m3",
displayName: "MiniMax M3",
provider: "fireworks",
contextWindow: 524288,
pricing: usd(0.06, 0.3, 1.2),
supportsVision: true,
clientType: "openai",
baseUrl: FIREWORKS_BASE_URL,
},
// -- SiliconFlow (gateway, official CNY pricing: cache hit / input / output) --
{
modelId: "deepseek-ai/DeepSeek-V4-Flash",
displayName: "DeepSeek V4 Flash",
provider: "siliconflow",
contextWindow: 1000000,
pricing: cny(0.02, 1, 2),
supportsVision: false,
clientType: "openai",
baseUrl: SILICONFLOW_BASE_URL,
},
{
modelId: "deepseek-ai/DeepSeek-V4-Pro",
displayName: "DeepSeek V4 Pro",
provider: "siliconflow",
contextWindow: 1000000,
pricing: cny(0.1, 12, 24),
supportsVision: false,
clientType: "openai",
baseUrl: SILICONFLOW_BASE_URL,
},
{
modelId: "meituan-longcat/LongCat-2.0",
displayName: "LongCat 2.0",
provider: "siliconflow",
contextWindow: 1000000,
pricing: cny(0.1, 5, 20),
supportsVision: false,
clientType: "openai",
baseUrl: SILICONFLOW_BASE_URL,
},
{
modelId: "moonshotai/Kimi-K2.7-Code",
displayName: "Kimi K2.7 Code",
provider: "siliconflow",
contextWindow: 262144,
pricing: cny(1.3, 6.5, 27),
supportsVision: true,
clientType: "openai",
baseUrl: SILICONFLOW_BASE_URL,
},
{
modelId: "zai-org/GLM-5.2",
displayName: "GLM-5.2",
provider: "siliconflow",
contextWindow: 1000000,
pricing: cny(2, 8, 28),
supportsVision: false,
clientType: "openai",
baseUrl: SILICONFLOW_BASE_URL,
},
// -- Qwen Token Plan (subscription gateway; vision flags per the plan's supported-model
// table). Pricing and context windows from each model's page at
// www.qianwenai.com/models/<id> (official CNY list prices; limited-time promotions such as
// the 20%/50% off discounts are not stored). qwen3.8-max-preview is preview-only with a
// quota-multiplier promotion and publishes no per-token list price nor a context window, so
// it carries no pricing and uses its family's 1M window. --
{
modelId: "deepseek-v4-pro",
displayName: "DeepSeek V4 Pro",
provider: "qwen-token-plan",
contextWindow: 1000000,
pricing: cny(1, 12, 24),
supportsVision: false,
clientType: "openai",
baseUrl: QWEN_TOKEN_PLAN_BASE_URL,
},
{
modelId: "glm-5.2",
displayName: "GLM-5.2",
provider: "qwen-token-plan",
contextWindow: 1048576,
pricing: cny(2, 8, 28),
supportsVision: false,
clientType: "openai",
baseUrl: QWEN_TOKEN_PLAN_BASE_URL,
},
{
modelId: "qwen3.8-max-preview",
displayName: "Qwen 3.8 Max Preview",
provider: "qwen-token-plan",
contextWindow: 1000000,
supportsVision: true,
clientType: "openai",
baseUrl: QWEN_TOKEN_PLAN_BASE_URL,
},
{
modelId: "qwen3.7-max",
displayName: "Qwen 3.7 Max",
provider: "qwen-token-plan",
contextWindow: 1000000,
pricing: cny(2.4, 12, 36),
supportsVision: false,
clientType: "openai",
baseUrl: QWEN_TOKEN_PLAN_BASE_URL,
},
{
modelId: "qwen3.7-plus",
displayName: "Qwen 3.7 Plus",
provider: "qwen-token-plan",
contextWindow: 1000000,
pricing: cny(0.4, 2, 8),
supportsVision: true,
clientType: "openai",
baseUrl: QWEN_TOKEN_PLAN_BASE_URL,
},
// -- Qwen Pay-As-You-Go (DashScope's OpenAI-compatible pay-per-token marketplace; official
// CNY list prices and specs from each model's page at www.qianwenai.com/models/<id> —
// resold third-party models keep their vendor-prefixed upstream ids) --
{
modelId: "kimi/kimi-k3",
displayName: "Kimi K3",
provider: "qwen-pay-as-you-go",
contextWindow: 1048576,
pricing: cny(2, 20, 100),
supportsVision: true,
clientType: "openai",
baseUrl: QWEN_PAYG_BASE_URL,
},
{
modelId: "qwen3.7-max",
displayName: "Qwen 3.7 Max",
provider: "qwen-pay-as-you-go",
contextWindow: 1000000,
pricing: cny(2.4, 12, 36),
supportsVision: false,
clientType: "openai",
baseUrl: QWEN_PAYG_BASE_URL,
},
{
modelId: "qwen3.7-plus",
displayName: "Qwen 3.7 Plus",
provider: "qwen-pay-as-you-go",
contextWindow: 1000000,
pricing: cny(0.4, 2, 8),
supportsVision: true,
clientType: "openai",
baseUrl: QWEN_PAYG_BASE_URL,
},
{
modelId: "ZHIPU/GLM-5.2",
displayName: "GLM-5.2",
provider: "qwen-pay-as-you-go",
contextWindow: 1048576,
pricing: cny(2, 8, 28),
supportsVision: false,
clientType: "openai",
baseUrl: QWEN_PAYG_BASE_URL,
},
// -- Google Gemini (official USD pricing) --
{
modelId: "gemini-3.5-flash",
displayName: "Gemini 3.5 Flash",
provider: "google",
contextWindow: 1048576,
pricing: usd(0.15, 1.5, 9),
supportsVision: true,
},
{
modelId: "gemini-3.1-flash-lite",
displayName: "Gemini 3.1 Flash-Lite",
provider: "google",
contextWindow: 1048576,
pricing: usd(0.025, 0.25, 1.5),
supportsVision: true,
},
{
// ≤200K input tier; >200K has official surcharge pricing (see file header comment).
modelId: "gemini-3.1-pro-preview",
displayName: "Gemini 3.1 Pro (Preview)",
provider: "google",
contextWindow: 1048576,
pricing: usd(0.2, 2, 12),
supportsVision: true,
},
{
modelId: "gemini-3-flash-preview",
displayName: "Gemini 3 Flash (Preview)",
provider: "google",
contextWindow: 1048576,
pricing: usd(0.05, 0.5, 3),
supportsVision: true,
},
// -- Anthropic (official USD pricing; cache write = 1.25 x input) --
{
modelId: "claude-opus-4-8",
displayName: "Claude Opus 4.8",
@@ -205,7 +666,7 @@ export const MODEL_CATALOG: ModelCatalogEntry[] = [
pricing: usd(0.3, 3.75, 15),
supportsVision: true,
},
// —— OpenAI ——
// -- OpenAI (official USD pricing) --
{
modelId: "gpt-5.5",
displayName: "GPT-5.5",
@@ -256,41 +717,7 @@ export const MODEL_CATALOG: ModelCatalogEntry[] = [
pricing: usd(30, 30, 180),
supportsVision: true,
},
// —— Google Gemini ——
{
// ≤200K input tier; >200K has official surcharge pricing (see file header comment).
modelId: "gemini-3.1-pro-preview",
displayName: "Gemini 3.1 Pro (Preview)",
provider: "google",
contextWindow: 1048576,
pricing: usd(0.2, 2, 12),
supportsVision: true,
},
{
modelId: "gemini-3.5-flash",
displayName: "Gemini 3.5 Flash",
provider: "google",
contextWindow: 1048576,
pricing: usd(0.15, 1.5, 9),
supportsVision: true,
},
{
modelId: "gemini-3-flash-preview",
displayName: "Gemini 3 Flash (Preview)",
provider: "google",
contextWindow: 1048576,
pricing: usd(0.05, 0.5, 3),
supportsVision: true,
},
{
modelId: "gemini-3.1-flash-lite",
displayName: "Gemini 3.1 Flash-Lite",
provider: "google",
contextWindow: 1048576,
pricing: usd(0.025, 0.25, 1.5),
supportsVision: true,
},
// —— Z.AI (GLM) ——
// -- Z.AI (GLM) --
{
modelId: "glm-5.2",
displayName: "GLM-5.2",
@@ -332,80 +759,6 @@ export const MODEL_CATALOG: ModelCatalogEntry[] = [
pricing: cny(0.7, 4, 21),
supportsVision: true,
},
// -- OpenRouter (gateway: uses OpenAI-compatible protocol, preset base URL) --
{
modelId: "xiaomi/mimo-v2.5",
displayName: "MiMo-V2.5",
provider: "openrouter",
contextWindow: 1048576,
pricing: usd(0.0028, 0.14, 0.28),
supportsVision: true,
clientType: "openai",
baseUrl: OPENROUTER_BASE_URL,
},
{
modelId: "tencent/hy3",
displayName: "Hy3",
provider: "openrouter",
contextWindow: 262144,
pricing: usd(0.035, 0.14, 0.58),
supportsVision: false,
clientType: "openai",
baseUrl: OPENROUTER_BASE_URL,
},
{
// No official separate cache price published: cache_read uses the standard input price (no discount assumed).
modelId: "minimax/minimax-m3",
displayName: "MiniMax M3",
provider: "openrouter",
contextWindow: 1048576,
pricing: usd(0.06, 0.3, 1.2),
supportsVision: true,
clientType: "openai",
baseUrl: OPENROUTER_BASE_URL,
},
{
// No official separate cache price published: cache_read uses the standard input price.
modelId: "stepfun/step-3.7-flash",
displayName: "Step 3.7 Flash",
provider: "openrouter",
contextWindow: 256000,
pricing: usd(0.04, 0.2, 1.15),
supportsVision: true,
clientType: "openai",
baseUrl: OPENROUTER_BASE_URL,
},
// -- SiliconFlow (gateway, official CNY pricing: cache hit / input / output) --
{
modelId: "zai-org/GLM-5.2",
displayName: "GLM-5.2",
provider: "siliconflow",
contextWindow: 1000000,
pricing: cny(2, 8, 28),
supportsVision: false,
clientType: "openai",
baseUrl: SILICONFLOW_BASE_URL,
},
{
modelId: "deepseek-ai/DeepSeek-V4-Pro",
displayName: "DeepSeek V4 Pro",
provider: "siliconflow",
contextWindow: 1000000,
pricing: cny(0.1, 12, 24),
supportsVision: false,
clientType: "openai",
baseUrl: SILICONFLOW_BASE_URL,
},
{
modelId: "meituan-longcat/LongCat-2.0",
displayName: "LongCat 2.0",
provider: "siliconflow",
contextWindow: 1000000,
pricing: cny(0.1, 5, 20),
supportsVision: false,
clientType: "openai",
baseUrl: SILICONFLOW_BASE_URL,
},
];
/** Looks up a catalog entry by (provider, upstream id) pair (**the sole catalog-matching entry point**); returns undefined if not in the catalog. */
@@ -416,15 +769,6 @@ export function catalogEntryFor(
return MODEL_CATALOG.find((m) => m.provider === provider && m.modelId === upstreamId);
}
/**
* Infers the provider for an upstream id from the built-in catalog (used to default
* `provider` on `model add`): if it matches a catalog entry, use that entry's provider
* (upstream ids are globally unique within the catalog); otherwise custom.
*/
export function inferProviderForUpstream(upstreamId: string): string {
return MODEL_CATALOG.find((m) => m.modelId === upstreamId)?.provider ?? "custom";
}
/** Looks up provider info by provider id; returns undefined for an unknown id. */
export function providerInfo(providerId: string): ModelProviderInfo | undefined {
return MODEL_PROVIDERS.find((p) => p.id === providerId);
@@ -487,3 +831,43 @@ export function presetModelEntries(): ModelEntry[] {
...(m.baseUrl !== undefined ? { base_url: m.baseUrl } : {}),
}));
}
/**
* The model's own homepage/detail page for the frontend's model-card link. Gateway groups
* have a stable per-model URL pattern (works for user-added ids in those groups too);
* direct-vendor models link to the vendor's model list/docs page; the Token Plan preview
* model has no dedicated page and links to the plan's model overview; custom and
* user-defined groups have no page to vouch for.
*/
export function modelHomepageUrl(provider: string, modelId: string): string | undefined {
if (provider === "openrouter") return `https://openrouter.ai/${modelId}`;
if (provider === "qwen-token-plan") {
return modelId === "qwen3.8-max-preview"
? providerInfo(provider)?.modelsUrl
: `https://www.qianwenai.com/models/${modelId}`;
}
if (provider === "fireworks") {
// API id "accounts/<owner>/models/<slug>" -> page "app.fireworks.ai/models/<owner>/<slug>";
// nonconforming (user-added) ids fall back to the models listing.
const m = /^accounts\/([^/]+)\/models\/(.+)$/.exec(modelId);
return m
? `https://app.fireworks.ai/models/${m[1]}/${m[2]}`
: providerInfo(provider)?.modelsUrl;
}
if (provider === "qwen-pay-as-you-go") {
return `https://www.qianwenai.com/models/${encodeURIComponent(modelId)}`;
}
if (provider === "zhipu") {
// Z.AI's per-model guide pages use the bare model id as the slug.
return `https://docs.z.ai/guides/llm/${modelId}`;
}
if (provider === "moonshot") {
// Moonshot's pricing pages: kimi-k2.6 -> chat-k26 (dot dropped); other ids fall back.
const m = /^kimi-k(\d+)\.(\d+)$/.exec(modelId);
return m
? `https://platform.kimi.com/docs/pricing/chat-k${m[1]}${m[2]}`
: providerInfo(provider)?.modelsUrl;
}
if (provider === "custom") return undefined;
return providerInfo(provider)?.modelsUrl;
}
+26 -34
View File
@@ -17,11 +17,17 @@
* the unique key — string concatenation like `<provider>/<id>` is forbidden anywhere in the
* pipeline. `model_id` is the upstream request id, sent to AgentHub unchanged; `default_model` /
* `vision_model` are paired `{ provider, model_id }` references (a TOML inline table).
*
* A caller always supplies the **complete pair**: `provider` is never guessed from the builtin
* catalog and never derived from whichever configured entry happens to carry the same
* `model_id`. Both halves or neither — a `model_id` without a `provider` is an error, not a
* lookup, because resolving it would silently point credentials and pricing at a vendor the
* caller never named.
*/
import fs from "node:fs/promises";
import path from "node:path";
import { parse as parseToml, stringify as stringifyToml } from "smol-toml";
import { inferProviderForUpstream, presetModelEntries } from "./model-catalog.js";
import { presetModelEntries } from "./model-catalog.js";
import { projectConfigPath } from "./paths.js";
/** Model reference: a `(provider, model_id)` pair (never string-concatenated anywhere). */
@@ -247,9 +253,11 @@ export async function saveProjectConfig(
/**
* Adds or updates a Model:
* - Upserts into `models`, deduplicated by the `(provider, model_id)` pair (provider may be
* omitted — the builtin catalog is used to infer the upstream id's group, falling back to
* custom if it can't be inferred);
* - Upserts into `models`, deduplicated by the `(provider, model_id)` pair — both halves are
* supplied by the caller, since the group is never guessed from the builtin catalog (a
* gateway reselling a vendor model keeps the vendor's upstream id, so a bare id names no
* single group, and guessing wrong files the caller's api_key under a vendor they never
* picked); a model outside every known group is added under `"custom"` explicitly;
* - If `api_key`/`base_url` are provided, they're written inline into the entry;
* - Set as the default Model (a paired reference) when `opts.setDefault` is true.
* Reads the existing config (or the default), saves after the change, and returns the updated
@@ -259,8 +267,8 @@ export async function addModel(
root: string,
projectId: string,
entry: {
/** provider group; inferred from the builtin catalog when omitted (`inferProviderForUpstream`, falling back to custom if it can't be inferred). */
provider?: string;
/** provider group (required; never inferred — pass `"custom"` for a model outside the known groups). */
provider: string;
/** Upstream model id (sent to AgentHub unchanged). */
model_id: string;
context_window?: number;
@@ -275,7 +283,7 @@ export async function addModel(
opts?: { setDefault?: boolean },
): Promise<ProjectConfig> {
const cfg = await loadProjectConfig(root, projectId);
const provider = entry.provider ?? inferProviderForUpstream(entry.model_id);
const { provider } = entry;
// upsert: layers new fields on top of the existing entry; fields not explicitly provided
// (e.g. context_window) keep their existing value, so a call like "just add an api_key"
@@ -400,35 +408,19 @@ export function getModel(cfg: ProjectConfig, ref: ModelRef): ModelEntry | undefi
}
/**
* Resolves a model reference (the **single entry point for "provider omitted"**, shared by core
* and CLI/server — never set up a second one):
* - `provider` given: validated for existence by exact paired reference;
* - `provider` omitted: an exact-match lookup on `model_id` (no fuzzy matching of any kind) —
* resolvable only when **exactly one** entry matches; 0 or multiple matches always report a
* clear error (an ambiguity error lists the candidate paired references).
* Validates a `(provider, model_id)` pair against the Project config and returns it as a
* `ModelRef` (the **single validation entry point**, shared by core and CLI/server — never set
* up a second one). Both halves are required: this only ever checks that the exact pair is
* configured, it never searches for a group to attach to a bare `model_id`. A pair the config
* doesn't have throws — a reference outside the config would leave credentials, pricing, and
* the context window unavailable at request time.
*/
export function resolveModelRef(cfg: ProjectConfig, modelId: string, provider?: string): ModelRef {
if (provider !== undefined) {
const ref: ModelRef = { provider, model_id: modelId };
if (!getModel(cfg, ref)) {
throw new Error(
`Model is not in the Project config: ${formatModelRef(ref)}. Use \`penguin config model list\` to see the configured models, or \`penguin config model add\` to add one.`,
);
}
return ref;
}
const candidates = cfg.models.filter((m) => m.model_id === modelId);
if (candidates.length === 1) {
return { provider: candidates[0]!.provider, model_id: modelId };
}
if (candidates.length === 0) {
export function resolveModelRef(cfg: ProjectConfig, modelId: string, provider: string): ModelRef {
const ref: ModelRef = { provider, model_id: modelId };
if (!getModel(cfg, ref)) {
throw new Error(
`Model is not in the Project config: no entry has model_id ${modelId}. Use \`penguin config model list\` to see the configured models, or \`penguin config model add\` to add one.`,
`Model is not in the Project config: ${formatModelRef(ref)}. Use \`penguin config model list\` to see the configured models, or \`penguin config model add\` to add one.`,
);
}
throw new Error(
`Ambiguous model reference: model_id ${modelId} matches multiple entries: ${candidates
.map((m) => formatModelRef({ provider: m.provider, model_id: m.model_id }))
.join(", ")}. Specify provider to give a paired reference.`,
);
return ref;
}
+48 -41
View File
@@ -80,14 +80,18 @@ describe("Agent.createSession workspace handling", () => {
expect(path.isAbsolute(session.workspaceDir)).toBe(true);
});
it("rejects a modelId that is not in the Project config with a clear error", async () => {
it("rejects a model reference that is not in the Project config with a clear error", async () => {
const agent = await createAgent();
const ws = path.join(tmpRoot, "ws-bad-model");
await fs.mkdir(ws, { recursive: true });
// A reference outside the config is not silently allowed (the unique key is provider +
// model_id); the error is thrown before creating the temp Workspace.
await expect(
agent.createSession({ workspaceDir: ws, modelId: "not-configured-model" }),
agent.createSession({
workspaceDir: ws,
modelId: "not-configured-model",
provider: "custom",
}),
).rejects.toThrow(/is not in the Project config/);
await expect(
agent.createSession({ workspaceDir: ws, modelId: "deepseek-v4-pro", provider: "openai" }),
@@ -129,59 +133,62 @@ describe("Agent.createSession model reference ((provider, model_id) pair)", () =
}
});
it("resolves a unique bare model_id and accepts an explicit pair", async () => {
const agent = await createAgent();
const ws = path.join(tmpRoot, "ws-ref-pair");
await fs.mkdir(ws, { recursive: true });
// Provider omitted: model_id is a globally unique exact match in the config -> resolves to that entry.
const bare = await agent.createSession({ workspaceDir: ws, modelId: "deepseek-v4-flash" });
try {
expect(bare.provider).toBe("deepseek");
expect(bare.modelId).toBe("deepseek-v4-flash");
} finally {
bare.dispose();
}
const paired = await agent.createSession({
workspaceDir: ws,
modelId: "claude-sonnet-4-6",
provider: "anthropic",
});
try {
expect(paired.provider).toBe("anthropic");
expect(paired.modelId).toBe("claude-sonnet-4-6");
} finally {
paired.dispose();
}
});
it("rejects an ambiguous bare model_id and a provider without modelId", async () => {
// Two providers coexist with the same model_id: omitting provider throws an ambiguity
// error (listing the candidate pair references).
it("selects the entry named by the pair, even when a second group sells the same model_id", async () => {
// A user-run proxy resells claude-sonnet-4-6 under the same upstream id: the two entries
// coexist and the pair — not the bare id — decides which one (and therefore which
// credential and base_url) the Session runs on.
await addModel(tmpRoot, DEFAULT_PROJECT_ID, {
provider: "myproxy",
model_id: "claude-sonnet-4-6",
});
const agent = await createAgent();
const ws = path.join(tmpRoot, "ws-ref-ambiguous");
const ws = path.join(tmpRoot, "ws-ref-pair");
await fs.mkdir(ws, { recursive: true });
await expect(
agent.createSession({ workspaceDir: ws, modelId: "claude-sonnet-4-6" }),
).rejects.toThrow(/Ambiguous.*\(provider=anthropic, model_id=claude-sonnet-4-6\)/);
// Adding provider resolves it.
const session = await agent.createSession({
const vendor = await agent.createSession({
workspaceDir: ws,
modelId: "claude-sonnet-4-6",
provider: "anthropic",
});
try {
expect(vendor.provider).toBe("anthropic");
expect(vendor.modelId).toBe("claude-sonnet-4-6");
} finally {
vendor.dispose();
}
const proxied = await agent.createSession({
workspaceDir: ws,
modelId: "claude-sonnet-4-6",
provider: "myproxy",
});
try {
expect(session.provider).toBe("myproxy");
expect(proxied.provider).toBe("myproxy");
expect(proxied.modelId).toBe("claude-sonnet-4-6");
} finally {
proxied.dispose();
}
});
it("rejects half a reference: modelId without provider, and provider without modelId", async () => {
const agent = await createAgent();
const ws = path.join(tmpRoot, "ws-ref-half");
await fs.mkdir(ws, { recursive: true });
// A bare model_id is never resolved against the config, not even when exactly one entry
// carries it (deepseek-v4-flash is unique here): the group is the caller's to name.
await expect(
agent.createSession({ workspaceDir: ws, modelId: "deepseek-v4-flash" }),
).rejects.toThrow(/must be given as a \(provider, model_id\) pair/);
// The mirror case: provider alone is not a reference either.
await expect(agent.createSession({ workspaceDir: ws, provider: "deepseek" })).rejects.toThrow(
/must be given as a \(provider, model_id\) pair/,
);
// Neither half given is the documented "use the Project default" path, not an error.
const session = await agent.createSession({ workspaceDir: ws });
try {
expect(session.provider).toBe("deepseek");
expect(session.modelId).toBe("deepseek-v4-pro");
} finally {
session.dispose();
}
// provider cannot be used alone (the reference must be a pair).
await expect(agent.createSession({ workspaceDir: ws, provider: "anthropic" })).rejects.toThrow(
/provider cannot be used alone/,
);
});
});
@@ -122,6 +122,8 @@ describe("Project Dir / Agent ID placeholders", () => {
cwd: "/tmp/ws",
agentId: "env_agent",
projectDir: "/tmp/proj",
provider: "deepseek",
modelId: "deepseek-v4-pro",
platform: "linux",
osVersion: "test",
date: "2026-07-08",
+167 -41
View File
@@ -10,7 +10,11 @@
* determination.
*/
import { describe, expect, it } from "vitest";
import { ThinkingLevel } from "@prismshadow/agenthub";
import {
EmptyResponseError,
ThinkingLevel,
ToolCallArgumentParseError,
} from "@prismshadow/agenthub";
import type { UniEvent, UniMessage, UsageMetadata } from "@prismshadow/agenthub";
import type { LLMOutcome } from "../src/interfaces.js";
@@ -1105,6 +1109,37 @@ describe("isMalformedJsonParseError", () => {
);
expect(isMalformedJsonParseError(new Error("socket hang up"))).toBe(false);
});
it("detects AgentHub 0.4 parse/validation error classes (truncated tool args, thinking-only)", () => {
// A stream truncated mid-arguments surfaces as ToolCallArgumentParseError since agenthub
// 0.4 (previously a raw SyntaxError) — must stay malformed so the engine reconnects.
expect(
isMalformedJsonParseError(
new ToolCallArgumentParseError({
client: "Claude5Client",
toolName: "exec_command",
toolCallId: "toolu_broken_1",
rawArguments: '{"cmd": "ec',
reason: "Unterminated string in JSON at position 11",
}),
),
).toBe(true);
// A completed thinking-only response cannot be replayed (400 on the next turn): retrying
// via malformed gives the model another chance instead of failing the turn.
expect(
isMalformedJsonParseError(
new EmptyResponseError({ client: "Claude5Client", finishReason: "stop" }),
),
).toBe(true);
// Also detectable via the name fallback and the cause chain.
expect(
isMalformedJsonParseError(
new Error("request failed", {
cause: new EmptyResponseError({ client: "GPT5_5Client", finishReason: null }),
}),
),
).toBe(true);
});
});
describe("isIncompleteStreamError", () => {
@@ -1337,35 +1372,40 @@ describe("GenerativeModel.streamGenerate outcome classification (PRN-013)", () =
});
});
describe("provider fidelity fields (signature / phase)", () => {
describe("provider fidelity payloads (opaque, AgentHub 0.4 semantics)", () => {
const complete = (messages: ReturnType<typeof translateEvents>["messages"]) =>
messages.filter((m) => !(m.payload as { type: string }).type.startsWith("partial_"));
it("captures the thinking signature arriving as an empty-text delta (Claude signature_delta)", () => {
it("captures the thinking fidelity arriving as an empty-text delta (Claude signature_delta)", () => {
const { messages } = translateEvents([
ev({ content_items: [{ type: "thinking", thinking: "let me think" }] }),
ev({ content_items: [{ type: "thinking", thinking: "", signature: "sig-abc" }] }),
ev({
content_items: [{ type: "thinking", thinking: "", fidelity: { signature: "sig-abc" } }],
}),
ev({ content_items: [{ type: "text", text: "answer" }] }),
ev({ event_type: "stop", content_items: [], finish_reason: "stop" }),
]);
const thinking = complete(messages).find(
(m) => (m.payload as { type: string }).type === "thinking",
)!;
expect((thinking.payload as { thinking: string; signature?: string }).thinking).toBe(
"let me think",
);
expect((thinking.payload as { signature?: string }).signature).toBe("sig-abc");
const p = thinking.payload as { thinking: string; fidelity?: Record<string, unknown> };
expect(p.thinking).toBe("let me think");
expect(p.fidelity).toEqual({ signature: "sig-abc" });
});
it("splits adjacent thinking blocks on signature (redacted + normal keep their own signatures)", () => {
it("splits adjacent thinking blocks on differing fidelity (redacted + normal keep their own)", () => {
const { messages } = translateEvents([
// A redacted block: sentinel text + signature arrive together (Claude content_block_start).
// A redacted block: sentinel text + fidelity arrive together (Claude content_block_start).
ev({
content_items: [{ type: "thinking", thinking: "_REDACTED_THINKING", signature: "sig-red" }],
content_items: [
{ type: "thinking", thinking: "_REDACTED_THINKING", fidelity: { signature: "sig-red" } },
],
}),
// The next, ordinary thinking block.
ev({ content_items: [{ type: "thinking", thinking: "visible" }] }),
ev({ content_items: [{ type: "thinking", thinking: "", signature: "sig-vis" }] }),
ev({
content_items: [{ type: "thinking", thinking: "", fidelity: { signature: "sig-vis" } }],
}),
ev({ event_type: "stop", content_items: [], finish_reason: "stop" }),
]);
const thinkings = complete(messages).filter(
@@ -1373,49 +1413,125 @@ describe("provider fidelity fields (signature / phase)", () => {
);
expect(
thinkings.map((m) => {
const p = m.payload as { thinking: string; signature?: string };
return [p.thinking, p.signature];
const p = m.payload as { thinking: string; fidelity?: Record<string, unknown> };
return [p.thinking, p.fidelity];
}),
).toEqual([
["_REDACTED_THINKING", "sig-red"],
["visible", "sig-vis"],
["_REDACTED_THINKING", { signature: "sig-red" }],
["visible", { signature: "sig-vis" }],
]);
});
it("emits an empty-text thinking with signature (GPT-5 encrypted reasoning)", () => {
it("keeps a run of equal fidelity as one thinking block (OpenAI-compatible reasoning_field per delta)", () => {
const rf = { reasoning_field: "reasoning_content" };
const { messages } = translateEvents([
ev({ content_items: [{ type: "thinking", thinking: "", signature: '{"id":"rs_1"}' }] }),
ev({ content_items: [{ type: "thinking", thinking: "step 1, ", fidelity: { ...rf } }] }),
ev({ content_items: [{ type: "thinking", thinking: "step 2, ", fidelity: { ...rf } }] }),
ev({ content_items: [{ type: "thinking", thinking: "done", fidelity: { ...rf } }] }),
ev({ content_items: [{ type: "text", text: "answer" }] }),
ev({ event_type: "stop", content_items: [], finish_reason: "stop" }),
]);
const thinking = complete(messages).find(
const thinkings = complete(messages).filter(
(m) => (m.payload as { type: string }).type === "thinking",
)!;
expect((thinking.payload as { thinking: string }).thinking).toBe("");
expect((thinking.payload as { signature?: string }).signature).toBe('{"id":"rs_1"}');
);
expect(
thinkings.map((m) => {
const p = m.payload as { thinking: string; fidelity?: Record<string, unknown> };
return [p.thinking, p.fidelity];
}),
).toEqual([["step 1, step 2, done", rf]]);
});
it("splits text segments on phase markers arriving as empty-text deltas (GPT-5)", () => {
it("emits an empty-text thinking with fidelity (GPT-5 encrypted reasoning) and splits on the next one", () => {
const { messages } = translateEvents([
ev({ content_items: [{ type: "text", text: "", phase: "planning" }] }),
ev({
content_items: [
{ type: "thinking", thinking: "", fidelity: { id: "rs_1", encrypted_content: "aaa" } },
],
}),
ev({
content_items: [
{ type: "thinking", thinking: "", fidelity: { id: "rs_2", encrypted_content: "bbb" } },
],
}),
ev({ content_items: [{ type: "text", text: "answer" }] }),
ev({ event_type: "stop", content_items: [], finish_reason: "stop" }),
]);
const thinkings = complete(messages).filter(
(m) => (m.payload as { type: string }).type === "thinking",
);
expect(
thinkings.map((m) => {
const p = m.payload as { thinking: string; fidelity?: Record<string, unknown> };
return [p.thinking, p.fidelity];
}),
).toEqual([
["", { id: "rs_1", encrypted_content: "aaa" }],
["", { id: "rs_2", encrypted_content: "bbb" }],
]);
});
it("splits text segments on fidelity.phase markers arriving as empty-text deltas (GPT-5)", () => {
const { messages } = translateEvents([
ev({ content_items: [{ type: "text", text: "", fidelity: { phase: "planning" } }] }),
ev({ content_items: [{ type: "text", text: "plan..." }] }),
ev({ content_items: [{ type: "text", text: "", phase: "answer" }] }),
ev({ content_items: [{ type: "text", text: "", fidelity: { phase: "answer" } }] }),
ev({ content_items: [{ type: "text", text: "final" }] }),
ev({ event_type: "stop", content_items: [], finish_reason: "stop" }),
]);
const texts = complete(messages).filter((m) => (m.payload as { type: string }).type === "text");
expect(
texts.map((m) => {
const p = m.payload as { text: string; phase?: string | null };
return [p.text, p.phase];
const p = m.payload as { text: string; fidelity?: Record<string, unknown> };
return [p.text, p.fidelity];
}),
).toEqual([
["plan...", "planning"],
["final", "answer"],
["plan...", { phase: "planning" }],
["final", { phase: "answer" }],
]);
});
it("carries the tool_call signature through to the complete message", () => {
it("closes a text segment on fidelity.signature: later text becomes its own segment (the signature must not cover unsigned text)", () => {
const { messages } = translateEvents([
// Gemini stamps a thoughtSignature on the text part it signed; whatever follows is
// unsigned. Merging them would replay a signature covering text the provider never
// signed, and the provider rejects the resumed turn.
ev({ content_items: [{ type: "text", text: "part one", fidelity: { signature: "sigA" } }] }),
ev({ content_items: [{ type: "text", text: "part two" }] }),
ev({ event_type: "stop", content_items: [], finish_reason: "stop" }),
]);
const texts = complete(messages).filter((m) => (m.payload as { type: string }).type === "text");
expect(
texts.map((m) => {
const p = m.payload as { text: string; fidelity?: Record<string, unknown> };
return [p.text, p.fidelity];
}),
).toEqual([
["part one", { signature: "sigA" }],
["part two", undefined],
]);
});
it("accumulates fidelity keys across deltas of one text segment (phase marker + trailing signature)", () => {
const { messages } = translateEvents([
// GPT-5 opens the segment with a bare phase marker and signs it only at the end; a
// replacing (rather than merging) assignment would drop the phase and lose the
// segmentation on replay.
ev({ content_items: [{ type: "text", text: "", fidelity: { phase: "answer" } }] }),
ev({ content_items: [{ type: "text", text: "final" }] }),
ev({ content_items: [{ type: "text", text: "", fidelity: { signature: "sigB" } }] }),
ev({ event_type: "stop", content_items: [], finish_reason: "stop" }),
]);
const texts = complete(messages).filter((m) => (m.payload as { type: string }).type === "text");
expect(
texts.map((m) => {
const p = m.payload as { text: string; fidelity?: Record<string, unknown> };
return [p.text, p.fidelity];
}),
).toEqual([["final", { phase: "answer", signature: "sigB" }]]);
});
it("carries the tool_call fidelity through to the complete message", () => {
const { messages } = translateEvents([
ev({
content_items: [
@@ -1424,7 +1540,7 @@ describe("provider fidelity fields (signature / phase)", () => {
name: "exec_command",
arguments: { cmd: "ls" },
tool_call_id: "tc1",
signature: "sig-tool",
fidelity: { signature: "sig-tool" },
},
],
}),
@@ -1433,32 +1549,42 @@ describe("provider fidelity fields (signature / phase)", () => {
const tc = complete(messages).find(
(m) => (m.payload as { type: string }).type === "tool_call",
)!;
expect((tc.payload as { signature?: string }).signature).toBe("sig-tool");
expect((tc.payload as { fidelity?: Record<string, unknown> }).fidelity).toEqual({
signature: "sig-tool",
});
});
it("round-trips fidelity fields back to UniMessage content items (setHistory path)", () => {
it("round-trips fidelity payloads back to UniMessage content items (setHistory path)", () => {
const uni = mergeOmniToUniMessage([
thinkingMessage("deep", "completed", { signature: "sig-1" }),
assistantText("hi", "completed", { phase: "answer", signature: "sig-2" }),
toolCall({ name: "t", arguments: "{}", toolCallId: "tc1", signature: "sig-3" }),
toolCall({ name: "t", arguments: "{}", toolCallId: "tc1", fidelity: { signature: "sig-3" } }),
]);
expect(uni.content_items).toEqual([
{ type: "thinking", thinking: "deep", signature: "sig-1" },
{ type: "text", text: "hi", phase: "answer", signature: "sig-2" },
{ type: "tool_call", name: "t", arguments: {}, tool_call_id: "tc1", signature: "sig-3" },
{ type: "thinking", thinking: "deep", fidelity: { signature: "sig-1" } },
{ type: "text", text: "hi", fidelity: { phase: "answer", signature: "sig-2" } },
{
type: "tool_call",
name: "t",
arguments: {},
tool_call_id: "tc1",
fidelity: { signature: "sig-3" },
},
]);
});
});
describe("flushText signature parity (PR #39 review)", () => {
it("emits an empty-text message carrying a text signature instead of dropping it", () => {
describe("flushText fidelity parity (PR #39 review)", () => {
it("emits an empty-text message carrying a text fidelity instead of dropping it", () => {
const { messages } = translateEvents([
ev({ content_items: [{ type: "text", text: "", signature: "sig-t" }] }),
ev({ content_items: [{ type: "text", text: "", fidelity: { signature: "sig-t" } }] }),
ev({ event_type: "stop", content_items: [], finish_reason: "stop" }),
]);
const text = messages.find((m) => (m.payload as { type: string }).type === "text")!;
expect((text.payload as { text: string }).text).toBe("");
expect((text.payload as { signature?: string }).signature).toBe("sig-t");
expect((text.payload as { fidelity?: Record<string, unknown> }).fidelity).toEqual({
signature: "sig-t",
});
});
});
+175 -27
View File
@@ -6,24 +6,31 @@ import { describe, expect, it } from "vitest";
import {
MODEL_CATALOG,
MODEL_PROVIDERS,
modelHomepageUrl,
catalogEntryFor,
inferProviderForUpstream,
presetModelEntries,
providerInfo,
resolveModelEnv,
} from "../src/state/index.js";
describe("model-catalog", () => {
it("model ids are globally unique; DeepSeek comes first (the default model's provider)", () => {
it("(provider, model_id) pairs are unique; DeepSeek comes first (the default model's provider)", () => {
// Bare model ids may repeat across providers (a gateway reselling a vendor model keeps the
// vendor's upstream id, e.g. Qwen Token Plan's glm-5.2 / deepseek-v4-pro) — uniqueness is
// the (provider, model_id) pair, matching the catalog's sole lookup key (catalogEntryFor).
const pairs = MODEL_CATALOG.map((m) => `${m.provider}\0${m.modelId}`);
expect(new Set(pairs).size).toBe(pairs.length);
const ids = MODEL_CATALOG.map((m) => m.modelId);
expect(new Set(ids).size).toBe(ids.length);
expect(MODEL_CATALOG[0]!.provider).toBe("deepseek");
// Group order: DeepSeek first, followed by the OpenRouter and SiliconFlow gateways,
// then Google Gemini before Anthropic, with custom last.
// Group order: DeepSeek first, followed by the OpenRouter, SiliconFlow, and Qwen Token
// Plan gateways, then Google Gemini before Anthropic, with custom last.
expect(MODEL_PROVIDERS.map((p) => p.id)).toEqual([
"deepseek",
"openrouter",
"fireworks",
"siliconflow",
"qwen-token-plan",
"qwen-pay-as-you-go",
"google",
"anthropic",
"openai",
@@ -61,13 +68,26 @@ describe("model-catalog", () => {
}
});
it("all three price buckets are positive; context_window is a positive integer", () => {
it("price buckets are positive (preview models without a list price omit pricing); context_window is a positive integer", () => {
for (const m of MODEL_CATALOG) {
expect(m.pricing, m.modelId).toBeDefined();
expect(m.pricing!.unit).toBe("usd_per_mtok");
expect(m.pricing!.cache_read).toBeGreaterThan(0);
expect(m.pricing!.cache_write).toBeGreaterThan(0);
expect(m.pricing!.output).toBeGreaterThan(0);
if (m.provider === "qwen-token-plan" && m.modelId === "qwen3.8-max-preview") {
// Preview-only model: the plan runs a quota-multiplier promotion and publishes no
// per-token list price, so the entry carries none and costs read as 0 (same as
// unpriced user models).
expect(m.pricing, m.modelId).toBeUndefined();
} else if (m.modelId.endsWith(":free")) {
// Free-tier gateway model: a genuine $0 price (not "unknown"), so costs compute to 0.
expect(m.pricing, m.modelId).toBeDefined();
expect([m.pricing!.cache_read, m.pricing!.cache_write, m.pricing!.output]).toEqual([
0, 0, 0,
]);
} else {
expect(m.pricing, m.modelId).toBeDefined();
expect(m.pricing!.unit).toBe("usd_per_mtok");
expect(m.pricing!.cache_read).toBeGreaterThan(0);
expect(m.pricing!.cache_write).toBeGreaterThan(0);
expect(m.pricing!.output).toBeGreaterThan(0);
}
expect(Number.isInteger(m.contextWindow)).toBe(true);
expect(m.contextWindow!).toBeGreaterThan(0);
}
@@ -78,9 +98,12 @@ describe("model-catalog", () => {
expect(providerInfo("nonexistent")).toBeUndefined();
});
it("pair matching and provider-grouping inference: catalogEntryFor / inferProviderForUpstream", () => {
// catalogEntryFor is the sole catalog lookup entry point: it matches on (group, upstream id)
// pairs, so an identically named upstream id never matches across the wrong group.
it("catalogEntryFor is the sole lookup and always takes the (provider, model_id) pair", () => {
// It matches on (group, upstream id) pairs, so an identically named upstream id never
// matches across the wrong group. There is no bare-id lookup at all: a gateway reselling a
// vendor model keeps the vendor's upstream id, so a bare id names no single catalog entry
// and the catalog never offers to pick one (`glm-5.2`, `qwen3.7-max`, `qwen3.7-plus` and
// `deepseek-v4-pro` each appear under two groups).
expect(catalogEntryFor("anthropic", "claude-sonnet-4-6")?.displayName).toBe(
"Claude Sonnet 4.6",
);
@@ -88,11 +111,11 @@ describe("model-catalog", () => {
// The upstream id itself may contain / (gateway models); it is never split apart.
expect(catalogEntryFor("openrouter", "xiaomi/mimo-v2.5")?.displayName).toBe("MiMo-V2.5");
expect(catalogEntryFor("custom", "my-own")).toBeUndefined();
// Inference for `model add` when --provider is omitted: a catalog hit yields its provider, otherwise custom.
expect(inferProviderForUpstream("deepseek-v4-pro")).toBe("deepseek");
expect(inferProviderForUpstream("xiaomi/mimo-v2.5")).toBe("openrouter");
expect(inferProviderForUpstream("my-own-model")).toBe("custom");
// Each group's entry for a resold id is reached only through that group.
expect(catalogEntryFor("zhipu", "glm-5.2")?.contextWindow).toBe(1000000);
expect(catalogEntryFor("qwen-token-plan", "glm-5.2")?.contextWindow).toBe(1048576);
expect(catalogEntryFor("deepseek", "deepseek-v4-pro")?.provider).toBe("deepseek");
expect(catalogEntryFor("qwen-token-plan", "deepseek-v4-pro")?.provider).toBe("qwen-token-plan");
});
it("presetModelEntries: provider and bare upstream model_id are separate fields; gateway models inline base_url", () => {
@@ -115,42 +138,124 @@ describe("model-catalog", () => {
}
});
it("gateway models (OpenRouter / SiliconFlow): openai protocol + preset base URL; env fallback is OPENAI_API_KEY", () => {
it("gateway models (OpenRouter / SiliconFlow / Qwen Token Plan): openai protocol + preset base URL; env fallback is OPENAI_API_KEY", () => {
const or = MODEL_CATALOG.filter((m) => m.provider === "openrouter");
// Dictionary order, newer versions of a series first (gpt-5.6-* before gpt-5.5,
// opus-4.8 before 4.7) — precomputed in the catalog, no runtime sorting.
expect(or.map((m) => m.modelId)).toEqual([
"xiaomi/mimo-v2.5",
"tencent/hy3",
"anthropic/claude-fable-5",
"anthropic/claude-opus-4.8",
"anthropic/claude-opus-4.7",
"anthropic/claude-sonnet-5",
"deepseek/deepseek-v4-flash",
"deepseek/deepseek-v4-pro",
"google/gemini-3.5-flash",
"minimax/minimax-m3",
"moonshotai/kimi-k3",
"nvidia/nemotron-3-ultra-550b-a55b:free",
"openai/gpt-5.6-sol",
"openai/gpt-5.6-terra",
"openai/gpt-5.5",
"stepfun/step-3.7-flash",
"tencent/hy3",
"x-ai/grok-4.5",
"xiaomi/mimo-v2.5",
"z-ai/glm-5.2",
]);
for (const m of or) {
expect(m.clientType).toBe("openai");
expect(m.baseUrl).toBe("https://openrouter.ai/api/v1");
}
const fw = MODEL_CATALOG.filter((m) => m.provider === "fireworks");
expect(fw.map((m) => [m.modelId, m.supportsVision])).toEqual([
["accounts/fireworks/models/deepseek-v4-flash", false],
["accounts/fireworks/models/deepseek-v4-pro", false],
["accounts/fireworks/models/glm-5p2", false],
["accounts/fireworks/models/kimi-k2p7-code", true],
["accounts/fireworks/models/minimax-m3", true],
]);
for (const m of fw) {
expect(m.clientType).toBe("openai");
expect(m.baseUrl).toBe("https://api.fireworks.ai/inference/v1");
}
const sf = MODEL_CATALOG.filter((m) => m.provider === "siliconflow");
expect(sf.map((m) => m.modelId)).toEqual([
"zai-org/GLM-5.2",
"deepseek-ai/DeepSeek-V4-Flash",
"deepseek-ai/DeepSeek-V4-Pro",
"meituan-longcat/LongCat-2.0",
"moonshotai/Kimi-K2.7-Code",
"zai-org/GLM-5.2",
]);
for (const m of sf) {
expect(m.clientType).toBe("openai");
expect(m.baseUrl).toBe("https://api.siliconflow.cn/v1");
}
const qtp = MODEL_CATALOG.filter((m) => m.provider === "qwen-token-plan");
expect(qtp.map((m) => m.modelId)).toEqual([
"deepseek-v4-pro",
"glm-5.2",
"qwen3.8-max-preview",
"qwen3.7-max",
"qwen3.7-plus",
]);
for (const m of qtp) {
expect(m.clientType).toBe("openai");
expect(m.baseUrl).toBe("https://token-plan.cn-beijing.maas.aliyuncs.com/compatible-mode/v1");
}
// Vision flags per the plan's supported-model table: 3.8-max-preview and 3.7-plus see images.
expect(qtp.map((m) => [m.modelId, m.supportsVision])).toEqual([
["deepseek-v4-pro", false],
["glm-5.2", false],
["qwen3.8-max-preview", true],
["qwen3.7-max", false],
["qwen3.7-plus", true],
]);
const qpayg = MODEL_CATALOG.filter((m) => m.provider === "qwen-pay-as-you-go");
expect(qpayg.map((m) => [m.modelId, m.supportsVision])).toEqual([
["kimi/kimi-k3", true],
["qwen3.7-max", false],
["qwen3.7-plus", true],
["ZHIPU/GLM-5.2", false],
]);
for (const m of qpayg) {
expect(m.clientType).toBe("openai");
expect(m.baseUrl).toBe("https://dashscope.aliyuncs.com/compatible-mode/v1");
}
// Routed through AgentHub's OpenAI client -> when the credential is left blank it reads OPENAI_API_KEY (not the provider's own env var name).
for (const id of ["openrouter", "siliconflow", "custom"]) {
for (const id of [
"openrouter",
"fireworks",
"siliconflow",
"qwen-token-plan",
"qwen-pay-as-you-go",
"custom",
]) {
expect(providerInfo(id)!.envKey).toBe("OPENAI_API_KEY");
expect(providerInfo(id)!.envBaseUrlKey).toBe("OPENAI_BASE_URL");
}
// gatewayBaseUrl (prefilled by group in the frontend's "add model" dialog) is only carried by the two gateway providers.
// gatewayBaseUrl (prefilled by group in the frontend's "add model" dialog) is only carried by the gateway providers.
expect(providerInfo("openrouter")!.gatewayBaseUrl).toBe("https://openrouter.ai/api/v1");
expect(providerInfo("siliconflow")!.gatewayBaseUrl).toBe("https://api.siliconflow.cn/v1");
expect(providerInfo("qwen-token-plan")!.gatewayBaseUrl).toBe(
"https://token-plan.cn-beijing.maas.aliyuncs.com/compatible-mode/v1",
);
expect(providerInfo("qwen-pay-as-you-go")!.gatewayBaseUrl).toBe(
"https://dashscope.aliyuncs.com/compatible-mode/v1",
);
expect(providerInfo("fireworks")!.gatewayBaseUrl).toBe("https://api.fireworks.ai/inference/v1");
const GATEWAYS = [
"openrouter",
"fireworks",
"siliconflow",
"qwen-token-plan",
"qwen-pay-as-you-go",
];
for (const p of MODEL_PROVIDERS) {
if (p.id !== "openrouter" && p.id !== "siliconflow") {
if (!GATEWAYS.includes(p.id)) {
expect(p.gatewayBaseUrl, p.id).toBeUndefined();
}
}
const gateway = [...or, ...sf];
const gateway = [...or, ...fw, ...sf, ...qtp, ...qpayg];
// Pricing (USD): MiMo v2.5 and Hy3.
const mimo = MODEL_CATALOG.find((m) => m.modelId === "xiaomi/mimo-v2.5")!.pricing!;
expect([mimo.cache_read, mimo.cache_write, mimo.output]).toEqual([0.0028, 0.14, 0.28]);
@@ -209,4 +314,47 @@ describe("resolveModelEnv (PRN-021: env fallback resolved by AgentHub routing ru
}
}
});
it("modelHomepageUrl: gateway per-model pages, vendor docs fallback, none for custom groups", () => {
// Gateway URL patterns work for user-added ids in those groups too (not catalog-gated).
expect(modelHomepageUrl("openrouter", "anthropic/claude-fable-5")).toBe(
"https://openrouter.ai/anthropic/claude-fable-5",
);
expect(modelHomepageUrl("openrouter", "someone/new-model")).toBe(
"https://openrouter.ai/someone/new-model",
);
expect(modelHomepageUrl("qwen-token-plan", "qwen3.7-plus")).toBe(
"https://www.qianwenai.com/models/qwen3.7-plus",
);
// Fireworks maps the accounts/<owner>/models/<slug> API id to its page path; other ids
// fall back to the models listing.
expect(modelHomepageUrl("fireworks", "accounts/fireworks/models/glm-5p2")).toBe(
"https://app.fireworks.ai/models/fireworks/glm-5p2",
);
expect(modelHomepageUrl("fireworks", "my-own-id")).toBe("https://app.fireworks.ai/models");
// Pay-as-you-go resells third-party models under slash-prefixed ids: the id is URL-encoded.
expect(modelHomepageUrl("qwen-pay-as-you-go", "ZHIPU/GLM-5.2")).toBe(
"https://www.qianwenai.com/models/ZHIPU%2FGLM-5.2",
);
// The preview model has no dedicated page: falls back to the plan's model overview.
expect(modelHomepageUrl("qwen-token-plan", "qwen3.8-max-preview")).toBe(
providerInfo("qwen-token-plan")!.modelsUrl,
);
// Direct vendors link to the vendor's model docs page.
expect(modelHomepageUrl("deepseek", "deepseek-v4-pro")).toBe(
"https://api-docs.deepseek.com/quick_start/pricing",
);
// Z.AI and Moonshot have per-model pages (Moonshot drops the dot: kimi-k2.6 -> chat-k26).
expect(modelHomepageUrl("zhipu", "glm-5.2")).toBe("https://docs.z.ai/guides/llm/glm-5.2");
expect(modelHomepageUrl("moonshot", "kimi-k2.6")).toBe(
"https://platform.kimi.com/docs/pricing/chat-k26",
);
expect(modelHomepageUrl("moonshot", "kimi-k2.5")).toBe(
"https://platform.kimi.com/docs/pricing/chat-k25",
);
expect(modelHomepageUrl("moonshot", "my-own")).toBe("https://platform.kimi.com/docs/pricing");
// Custom and user-defined groups have no page to vouch for.
expect(modelHomepageUrl("custom", "my-model")).toBeUndefined();
expect(modelHomepageUrl("my-own-gateway", "x")).toBeUndefined();
});
});
+23
View File
@@ -10,6 +10,7 @@ import {
generateTitleWithLLM,
sanitizeTitle,
Session,
stripConversationMarkers,
thinkingMessage,
tokenUsage,
userText,
@@ -119,6 +120,24 @@ describe("session-title", () => {
expect(sanitizeTitle("『标题』!")).toBe("标题");
expect(sanitizeTitle(" \n ")).toBeNull();
expect(sanitizeTitle("x".repeat(50))).toHaveLength(30);
// A leaked <use_skills> block is stripped from the model output.
expect(sanitizeTitle("<use_skills>\nskills: web-design\n</use_skills>\n构建落地页")).toBe(
"构建落地页",
);
});
it("stripConversationMarkers: removes machine marker blocks, keeps the human body", () => {
// The skill-invocation block that wraps a first user message must not reach the title.
expect(
stripConversationMarkers(
"<use_skills>\nskills: penguin-sdk, web-design\n</use_skills>\n做一个 RAG 应用",
),
).toBe("做一个 RAG 应用");
// Handoff and scheduled-task markers are stripped too; ordinary angle-bracket text stays.
expect(stripConversationMarkers("<handoff_from>data_analyst</handoff_from>继续分析")).toBe(
"继续分析",
);
expect(stripConversationMarkers("render a <div> element")).toBe("render a <div> element");
});
it("Session.generateTitle: sends via createBareLLM; returns null when no factory is provided", async () => {
@@ -162,6 +181,10 @@ describe("session-title", () => {
// Material = the first Task's user text + model text (thinking does not count), matching
// buildTitlePrompt's shape.
expect(seen[0]).toBe(buildTitlePrompt("user question", "answer body"));
// Anti-CoT shape: an explicit no-thinking rule, and the prompt ends with an empty think
// block so reasoning models treat their thinking phase as already closed.
expect(seen[0]).toContain("do not think aloud");
expect(seen[0]!.endsWith("<think></think>")).toBe(true);
// No request is sent when no material has been collected (run was never called).
const idle = new Session({
+78 -31
View File
@@ -41,6 +41,7 @@ import {
skillsDir,
systemConfigPath,
toolsDir,
type ModelRef,
type ProjectConfig,
} from "../src/state/index.js";
import { sessionEnvironment } from "../src/internal/session-support.js";
@@ -288,6 +289,8 @@ describe("assembleSystemPrompt", () => {
sessionEnvironment("/tmp/penguin-ws", "session-test-1", {
agentId: DEFAULT_AGENT_ID,
projectDir: "/tmp/proj",
provider: "deepseek",
modelId: "deepseek-v4-pro",
}),
);
expect(prompt).toContain("AGENTS.md");
@@ -341,6 +344,8 @@ describe("assembleSystemPrompt", () => {
cwd: "/tmp/ws",
agentId: "agent-x",
projectDir: "/tmp/proj",
provider: "deepseek",
modelId: "deepseek-v4-pro",
platform: "darwin",
osVersion: "Darwin 25.0.0",
date: "2026-06-30",
@@ -393,6 +398,8 @@ describe("assembleSystemPrompt", () => {
cwd: "/tmp/ws",
agentId: "agent-x",
projectDir: "/tmp/proj",
provider: "deepseek",
modelId: "deepseek-v4-pro",
platform: "darwin",
osVersion: "Darwin 25.0.0",
date: "2026-06-30",
@@ -466,7 +473,7 @@ describe("assembleSystemPrompt", () => {
const env = sessionEnvironment(
"/tmp/penguin-ws",
"session-test-1",
{ agentId: "agent-x", projectDir: "/tmp/proj" },
{ agentId: "agent-x", projectDir: "/tmp/proj", provider: "openai", modelId: "gpt-5.5" },
new Date("2026-06-30T00:00:00"),
);
const prompt = assembleSystemPrompt(state, env);
@@ -476,15 +483,19 @@ describe("assembleSystemPrompt", () => {
expect(prompt).toContain("CWD: /tmp/penguin-ws");
expect(prompt).toContain("Agent ID: agent-x");
expect(prompt).toContain("Project Dir: /tmp/proj");
expect(prompt).toContain("Provider: openai");
expect(prompt).toContain("Model ID: gpt-5.5");
expect(prompt).toContain("Platform:");
expect(prompt).toContain("OS Version:");
expect(prompt).toContain("Date: 2026-06-30");
expect(prompt.indexOf("Platform:")).toBeLessThan(prompt.indexOf("OS Version:"));
expect(prompt.indexOf("OS Version:")).toBeLessThan(prompt.indexOf("Date:"));
expect(prompt.indexOf("Date:")).toBeLessThan(prompt.indexOf("CWD:"));
expect(prompt.indexOf("CWD:")).toBeLessThan(prompt.indexOf("Agent ID:"));
expect(prompt.indexOf("Agent ID:")).toBeLessThan(prompt.indexOf("Project Dir:"));
expect(prompt.indexOf("Project Dir:")).toBeLessThan(prompt.indexOf("Session ID:"));
expect(prompt.indexOf("Date:")).toBeLessThan(prompt.indexOf("Project Dir:"));
expect(prompt.indexOf("Project Dir:")).toBeLessThan(prompt.indexOf("Agent ID:"));
expect(prompt.indexOf("Agent ID:")).toBeLessThan(prompt.indexOf("CWD:"));
expect(prompt.indexOf("CWD:")).toBeLessThan(prompt.indexOf("Provider:"));
expect(prompt.indexOf("Provider:")).toBeLessThan(prompt.indexOf("Model ID:"));
expect(prompt.indexOf("Model ID:")).toBeLessThan(prompt.indexOf("Session ID:"));
});
});
@@ -531,15 +542,41 @@ describe("project-config round trip", () => {
expect(getModel(loaded, { provider: "custom", model_id: "unknown-model" })).toBeUndefined();
});
it("infers the provider from the builtin catalog when addModel omits it", async () => {
await addModel(tmpRoot, DEFAULT_PROJECT_ID, { model_id: "claude-sonnet-4-6" });
await addModel(tmpRoot, DEFAULT_PROJECT_ID, { model_id: "my-own-model" });
it("addModel files the entry under the provider it was given, never one of its own choosing", async () => {
// provider is a required field: nothing is inferred from the builtin catalog, so a model
// outside the known groups is filed under custom only because the caller said so. glm-5.2
// is sold by both the Qwen Token Plan gateway and Zhipu — the caller names which one, and
// the entry (with its api_key) lands in exactly that group.
await addModel(tmpRoot, DEFAULT_PROJECT_ID, {
provider: "anthropic",
model_id: "claude-sonnet-4-6",
});
await addModel(tmpRoot, DEFAULT_PROJECT_ID, { provider: "custom", model_id: "my-own-model" });
await addModel(tmpRoot, DEFAULT_PROJECT_ID, {
provider: "zhipu",
model_id: "glm-5.2",
api_key: "sk-zhipu",
});
const loaded = await loadProjectConfig(tmpRoot, DEFAULT_PROJECT_ID);
// A catalog hit gets its provider; otherwise it falls back to custom.
expect(
getModel(loaded, { provider: "anthropic", model_id: "claude-sonnet-4-6" }),
).toBeDefined();
expect(getModel(loaded, { provider: "custom", model_id: "my-own-model" })).toBeDefined();
expect(getModel(loaded, { provider: "zhipu", model_id: "glm-5.2" })?.api_key).toBe("sk-zhipu");
// The key never leaks into the other group that resells the same bare id.
expect(
getModel(loaded, { provider: "qwen-token-plan", model_id: "glm-5.2" })?.api_key,
).toBeUndefined();
});
it("addModel requires an explicit provider: a bare model_id does not type-check", () => {
// Compile-time contract (asserted by `pnpm typecheck`, which includes this file): with the
// catalog inference gone there is nothing to fall back to, so the entry's provider field is
// required rather than optional. vitest only checks that the call expression exists.
const bare = { model_id: "glm-5.2" };
// @ts-expect-error provider is required: a model reference is always a (provider, model_id) pair.
const call = (): Promise<ProjectConfig> => addModel(tmpRoot, DEFAULT_PROJECT_ID, bare);
expect(call).toBeTypeOf("function");
});
it("upserts by the (provider, model_id) pair; same model_id under two providers co-exists", async () => {
@@ -706,7 +743,10 @@ describe("project-config round trip", () => {
(c) => c.provider === entry.provider && c.modelId === entry.model_id,
)!;
expect(entry.vision).toBe(cat.supportsVision ? undefined : false);
expect(entry.pricing?.unit).toBe("usd_per_mtok");
// A catalog entry without a list price (the Token Plan preview model) presets no
// pricing; every other catalog entry stores USD pricing.
if (cat.pricing === undefined) expect(entry.pricing).toBeUndefined();
else expect(entry.pricing?.unit).toBe("usd_per_mtok");
// A model that auto-routes leaves client_type unset; a gateway model (OpenRouter)
// explicitly sets it to openai.
expect(entry.client_type).toBe(cat.clientType);
@@ -800,7 +840,7 @@ describe("project-config round trip", () => {
});
});
describe('resolveModelRef (the single "provider omitted" resolution entry point)', () => {
describe("resolveModelRef (validates a (provider, model_id) pair against the config)", () => {
const cfg: ProjectConfig = {
models: [
{ provider: "deepseek", model_id: "deepseek-v4-pro" },
@@ -809,37 +849,44 @@ describe('resolveModelRef (the single "provider omitted" resolution entry point)
],
};
it("resolves a globally unique model_id without provider (exact match only)", () => {
expect(resolveModelRef(cfg, "deepseek-v4-pro")).toEqual({
it("returns the pair when it names a configured entry", () => {
expect(resolveModelRef(cfg, "deepseek-v4-pro", "deepseek")).toEqual({
provider: "deepseek",
model_id: "deepseek-v4-pro",
});
// Exact-match lookup, no fuzzy/prefix matching.
expect(() => resolveModelRef(cfg, "deepseek-v4")).toThrow(/is not in the Project config/);
});
it("validates the exact pair when provider is given", () => {
// The same bare model_id under two groups is never ambiguous: each pair names its own entry.
expect(resolveModelRef(cfg, "shared-id", "siliconflow")).toEqual({
provider: "siliconflow",
model_id: "shared-id",
});
// A wrong provider grouping likewise fails: the error includes the pair reference.
expect(resolveModelRef(cfg, "shared-id", "openrouter")).toEqual({
provider: "openrouter",
model_id: "shared-id",
});
});
it("throws when the pair is not configured; the error carries the pair reference", () => {
// Wrong group for a configured model_id: no fallback to "the entry that happens to have
// this id" — a pair the config doesn't have simply isn't a model.
expect(() => resolveModelRef(cfg, "shared-id", "openai")).toThrow(
/\(provider=openai, model_id=shared-id\)/,
/is not in the Project config.*\(provider=openai, model_id=shared-id\)/,
);
// Unknown model_id, and exact matching only (no fuzzy/prefix matching).
expect(() => resolveModelRef(cfg, "no-such-model", "deepseek")).toThrow(
/\(provider=deepseek, model_id=no-such-model\)/,
);
expect(() => resolveModelRef(cfg, "deepseek-v4", "deepseek")).toThrow(
/is not in the Project config/,
);
});
it("errors clearly on zero matches", () => {
expect(() => resolveModelRef(cfg, "no-such-model")).toThrow(
/is not in the Project config.*no-such-model/,
);
});
it("errors on ambiguity, listing the candidate pair references", () => {
expect(() => resolveModelRef(cfg, "shared-id")).toThrow(/Ambiguous/);
expect(() => resolveModelRef(cfg, "shared-id")).toThrow(
/\(provider=siliconflow, model_id=shared-id\).*\(provider=openrouter, model_id=shared-id\)/,
);
it("requires provider: a bare model_id does not type-check (no resolution path left)", () => {
// The pair is enforced by the type checker — asserted by `pnpm typecheck`, which includes
// this file; the unused-directive error is the failure mode if the parameter ever goes
// optional again. vitest only checks that the call expression exists.
// @ts-expect-error provider is required: a model reference is always a (provider, model_id) pair.
const call = (): ModelRef => resolveModelRef(cfg, "deepseek-v4-pro");
expect(call).toBeTypeOf("function");
});
});
+50 -4
View File
@@ -56,7 +56,7 @@ type RunInput = { prompt: string; signal?: AbortSignal; approve?: ApproveFn };
/** Builds a SubagentRunner from a run implementation (spawn arguments observed via a spy). */
function runnerOf(
run: (input: RunInput) => AsyncGenerator<OmniMessage>,
spawnSpy?: (input: { agentId?: string; modelId?: string }) => void,
spawnSpy?: (input: { agentId?: string; modelId?: string; provider?: string }) => void,
): SubagentRunner {
return {
async spawn(input) {
@@ -136,7 +136,8 @@ afterEach(() => {
describe("run_subagent tool (foreground)", () => {
it("forwards stamped child messages and mirrors child text as its own output deltas", async () => {
const seen: Array<{ prompt?: string; agentId?: string; modelId?: string }> = [];
const seen: Array<{ prompt?: string; agentId?: string; modelId?: string; provider?: string }> =
[];
const runner = runnerOf(
async function* (input) {
seen[0] = { ...seen[0], prompt: input.prompt };
@@ -150,7 +151,10 @@ describe("run_subagent tool (foreground)", () => {
const { services } = makeServices(runner);
const tool = createSubagentTool(DEF, services);
const { out, result } = await collectWithReturn(
tool.execute({ prompt: "world", agent_id: "researcher", model_id: "m1" }, CTX),
tool.execute(
{ prompt: "world", agent_id: "researcher", model_id: "m1", provider: "p1" },
CTX,
),
);
// Child session messages pass through verbatim (with origin).
@@ -163,7 +167,49 @@ describe("run_subagent tool (foreground)", () => {
expect(result?.stopReason).toBe("completed");
// The model is free to choose the agent and model (spawn arguments); the prompt is
// handed to run.
expect(seen[0]).toEqual({ prompt: "world", agentId: "researcher", modelId: "m1" });
expect(seen[0]).toEqual({
prompt: "world",
agentId: "researcher",
modelId: "m1",
provider: "p1",
});
});
it("rejects half a model reference in either direction, and never spawns", async () => {
// A model is referenced by the (provider, model_id) pair; the tool refuses half of one
// rather than letting the session layer guess a group for a bare upstream id.
for (const args of [
{ prompt: "x", model_id: "m1" },
{ prompt: "x", provider: "p1" },
]) {
const spawned: unknown[] = [];
const runner = runnerOf(
async function* () {},
(input) => spawned.push(input),
);
const { services } = makeServices(runner);
const tool = createSubagentTool(DEF, services);
const { out, result } = await collectWithReturn(tool.execute(args, CTX));
expect(result?.stopReason).toBe("failed");
expect(ownDeltas(out)).toContain("must be given together");
expect(spawned).toHaveLength(0);
}
});
it("uses the Project default model when neither half of the reference is given", async () => {
const seen: Array<{ modelId?: string; provider?: string }> = [];
const runner = runnerOf(
async function* () {
yield withOrigin(partialText("delta", "ok"), HOP);
},
(input) => seen.push(input),
);
const { services } = makeServices(runner);
const tool = createSubagentTool(DEF, services);
const { result } = await collectWithReturn(tool.execute({ prompt: "x" }, CTX));
expect(result?.stopReason).toBe("completed");
expect(seen[0]?.modelId).toBeUndefined();
expect(seen[0]?.provider).toBeUndefined();
});
it("does not mirror deeper-nested (origin.length > 1) text into its own output", async () => {
+5 -5
View File
@@ -7,7 +7,7 @@ The CLI ships as the npm package `@prismshadow/penguin-cli`; the command is `pen
## Global conventions
- Model references: a model's identity is always the `(provider, model_id)` pair. `--model-id` takes the upstream model id and pairs with `--provider`. When `run` / `chat` omit `--provider`, the `--model-id` matches only if it is globally unique in the configuration; ambiguity is an error.
- Model references: a model's identity is always the `(provider, model_id)` pair. `--model-id` takes the upstream model id and `--provider` the group it belongs to; the provider is never inferred, guessed, or defaulted. On `run` / `chat` the pair as a whole is optional — pass both to pick a model, or neither to use the Project's default model — but passing one without the other is an error.
- Data root: `--root <dir>` overrides the data root directory. Priority: `--root` > the `PENGUIN_HOME` env var > `~/.penguin/data`.
## penguin run
@@ -21,8 +21,8 @@ penguin run -m "Summarize the code structure of this directory"
| Option | Description |
| --- | --- |
| `-m, --message <message>` | Required; the message to send |
| `--model-id <id>` | Model to use; defaults to the Project's default model |
| `--provider <group>` | Provider group of the model |
| `--model-id <id>` | Upstream id of the model to use; requires `--provider`. Omit both to use the Project's default model |
| `--provider <group>` | Provider group of the model; required whenever `--model-id` is given |
| `--project-id <id>` | Project to use |
| `--agent-id <id>` | Agent to use |
| `--workspace <path>` | Workspace directory; defaults to the current directory and must exist |
@@ -74,13 +74,13 @@ Manages a Project's model configuration, per-Agent vault environment variables,
Add or update a model entry:
```bash
penguin config model add --model-id deepseek-v4-pro --api-key sk-... --set-default
penguin config model add --provider deepseek --model-id deepseek-v4-pro --api-key sk-... --set-default
```
| Option | Description |
| --- | --- |
| `--model-id <id>` | Required; the upstream model id |
| `--provider <group>` | Provider group; inferred from the built-in catalog when omitted |
| `--provider <group>` | Required; the provider group the entry belongs to. It is never derived from the model id: gateways resell vendor models under their upstream ids, so a guessed group would write the credential onto another vendor's endpoint. Use `custom` for any endpoint outside the built-in groups. |
| `--api-key <key>` | API key, stored inline in the Project's hidden `.project_config.toml` |
| `--base-url <url>` | Custom endpoint base URL |
| `--context-window <n>` | Context window size |
+5 -5
View File
@@ -7,7 +7,7 @@ CLI 由 npm 包 `@prismshadow/penguin-cli` 提供,命令为 `penguin`。不带
## 全局约定
- 模型引用:模型身份始终是 `(provider, model_id)` 二元组。`--model-id` 填上游模型 id,与 `--provider` 组成配对引用。`run` / `chat` 省略 `--provider` 时,仅当该 `--model-id` 在配置中全局唯一才会匹配,存在歧义则报错。
- 模型引用:模型身份始终是 `(provider, model_id)` 二元组。`--model-id` 填上游模型 id,`--provider` 填其所属分组;provider 绝不推断、绝不猜测、也没有缺省值。`run` / `chat` 上这对参数整体可选——两个都给即指定模型,两个都不给则使用 Project 默认模型——但只给其中一个是错误。
- 数据根目录:`--root <dir>` 覆盖数据根目录,优先级为 `--root` > 环境变量 `PENGUIN_HOME` > `~/.penguin/data`。
## penguin run
@@ -21,8 +21,8 @@ penguin run -m "总结当前目录的代码结构"
| 选项 | 说明 |
| --- | --- |
| `-m, --message <message>` | 必填,要发送的消息 |
| `--model-id <id>` | 指定模型,缺省使用 Project 默认模型 |
| `--provider <group>` | 模型所属 Provider 分组 |
| `--model-id <id>` | 指定模型的上游 id,须与 `--provider` 同时给出;两者都不给时使用 Project 默认模型 |
| `--provider <group>` | 模型所属 Provider 分组,给出 `--model-id` 时必填 |
| `--project-id <id>` | 指定 Project |
| `--agent-id <id>` | 指定 Agent |
| `--workspace <path>` | Workspace 目录,默认当前目录,必须已存在 |
@@ -74,13 +74,13 @@ Ctrl-C 的行为依状态而定:
新增或更新模型条目:
```bash
penguin config model add --model-id deepseek-v4-pro --api-key sk-... --set-default
penguin config model add --provider deepseek --model-id deepseek-v4-pro --api-key sk-... --set-default
```
| 选项 | 说明 |
| --- | --- |
| `--model-id <id>` | 必填,上游模型 id |
| `--provider <group>` | Provider 分组,缺省时根据内置目录推断 |
| `--provider <group>` | 必填,条目所属的 Provider 分组。它绝不由模型 id 推导:网关会以上游 id 转售厂商模型,猜错分组会把凭据写到另一家厂商的接口上。内置分组之外的接口一律用 `custom`。 |
| `--api-key <key>` | API Key,内联存入 Project 隐藏文件 `.project_config.toml` |
| `--base-url <url>` | 自定义接口地址 |
| `--context-window <n>` | 上下文窗口大小 |
+7 -4
View File
@@ -35,7 +35,7 @@ The openrouter, siliconflow, and custom groups speak the OpenAI-compatible proto
## Project config
`<root>/<project>/.project_config.toml` is the Project's single config file: a hidden file written with mode 0600, with credentials inlined on the model entries. Model identity is always the `(provider, model_id)` pair — string concatenation is forbidden everywhere.
`<root>/<project>/.project_config.toml` is the Project's single config file: a hidden file written with mode 0600, with credentials inlined on the model entries. Model identity is always the `(provider, model_id)` pair — string concatenation is forbidden everywhere, and every reference into this file carries both halves: the provider is never inferred from a bare `model_id`.
| Field | Description |
| --- | --- |
@@ -143,9 +143,11 @@ compaction:
| `{{PLATFORM}}` | Runtime platform |
| `{{OS_VERSION}}` | Operating system version |
| `{{DATE}}` | Current date |
| `{{CWD}}` | Workspace path |
| `{{AGENT_ID}}` | Agent id |
| `{{PROJECT_DIR}}` | Project directory |
| `{{AGENT_ID}}` | Agent id |
| `{{CWD}}` | Workspace path |
| `{{PROVIDER}}` | Model provider group |
| `{{MODEL_ID}}` | Upstream model id |
| `{{SESSION_ID}}` | Session id |
`agent_state/AGENTS.md` is the developer-editable instruction file, injected via `{{AGENTS_MD}}` and empty by default — it is also the file an optimizer edits most (see [Self-Improvement](/self-improvement)).
@@ -157,6 +159,7 @@ compaction:
- Key names must match `^[A-Za-z_][A-Za-z0-9_]*$` (shell environment variable naming rules);
- Values are injected only into tool subprocess environments and never enter the model context or the Trace;
- Only key names are disclosed in the system prompt via `{{VAULT_KEYS}}`;
- Saving through the Web/API invalidates the Agent's cached Session runtimes: the next Task on any of its Sessions re-resumes and runs with the new values; a Task already in flight keeps the values it started with (a direct CLI file edit reaches a running server only when a Session is next created or resumed);
- Managed via `penguin config vault set/list/remove` or the Web Vault tab.
## Schedules
@@ -172,7 +175,7 @@ Each file `agent_state/schedule/<name>.toml` describes one scheduled task (the f
| `end_at` | no | End time; must be later than `start_at` |
| `session_id` | no | Bind to an existing Session; mutually exclusive with the three fields below |
| `workspace` | no | Workspace for new-Session mode |
| `provider` / `model_id` | no | Paired model reference for new-Session mode |
| `provider` / `model_id` | no | Paired model reference for new-Session mode; write both or neither — a lone `model_id` is rejected, and with neither the Project's default model is used |
```toml
prompt = "Check yesterday's builds and summarize the failures"
+7 -4
View File
@@ -35,7 +35,7 @@ openrouter、siliconflow 与 custom 分组走 OpenAI 兼容协议,因此复用
## Project 配置
`<root>/<project>/.project_config.toml` 是 Project 唯一的配置文件:隐藏文件,落盘权限 0600,凭证内联在模型条目上。模型身份始终是 `(provider, model_id)` 成对引用,禁止任何形式的字符串拼接。
`<root>/<project>/.project_config.toml` 是 Project 唯一的配置文件:隐藏文件,落盘权限 0600,凭证内联在模型条目上。模型身份始终是 `(provider, model_id)` 成对引用,禁止任何形式的字符串拼接;指向本文件的每一处引用都要带上两半,provider 绝不由裸 `model_id` 推断。
| 字段 | 说明 |
| --- | --- |
@@ -143,9 +143,11 @@ compaction:
| `{{PLATFORM}}` | 运行平台 |
| `{{OS_VERSION}}` | 操作系统版本 |
| `{{DATE}}` | 当前日期 |
| `{{CWD}}` | Workspace 路径 |
| `{{AGENT_ID}}` | Agent id |
| `{{PROJECT_DIR}}` | Project 目录 |
| `{{AGENT_ID}}` | Agent id |
| `{{CWD}}` | Workspace 路径 |
| `{{PROVIDER}}` | 模型 provider 分组 |
| `{{MODEL_ID}}` | 上游模型 id |
| `{{SESSION_ID}}` | Session id |
`agent_state/AGENTS.md` 是开发者可编辑的指令文件,经 `{{AGENTS_MD}}` 注入系统提示词,缺省为空——它也是优化器最常改动的文件(见[自我进化](/self-improvement))。
@@ -157,6 +159,7 @@ compaction:
- 键名须匹配 `^[A-Za-z_][A-Za-z0-9_]*$`(shell 环境变量命名规则);
- 值只注入工具子进程的环境变量,永远不进入模型上下文与 Trace;
- 系统提示词中经 `{{VAULT_KEYS}}` 只披露键名;
- 经 Web/API 保存会使该 Agent 已缓存的 Session 运行时失效:其任意 Session 的下一个任务会重新恢复(resume)并使用新值;进行中的任务保持其启动时的值(CLI 直接改文件对运行中的 server 则要等 Session 下次创建或恢复时生效);
- 通过 CLI `penguin config vault set/list/remove` 或 Web 的 Vault 标签页管理。
## 定时任务
@@ -172,7 +175,7 @@ compaction:
| `end_at` | 否 | 结束时刻,须晚于 `start_at` |
| `session_id` | 否 | 绑定既有 Session;与下列三项互斥 |
| `workspace` | 否 | 新建 Session 模式的 Workspace |
| `provider` / `model_id` | 否 | 新建 Session 模式的模型成对引用 |
| `provider` / `model_id` | 否 | 新建 Session 模式的模型成对引用;要写就两个都写,只写 `model_id` 会被拒绝,两个都不写则使用 Project 默认模型 |
```toml
prompt = "检查昨日构建结果并汇总失败原因"
+2 -2
View File
@@ -13,7 +13,7 @@ description: Install PenguinHarness via the install script, npm, or from source.
On Linux / macOS:
```bash
curl -fsSL https://github.com/Prism-Shadow/penguin-harness/releases/latest/download/install.sh | sh
curl -fsSL https://penguin.ooo/install.sh | sh
```
The script downloads the matching `penguin-{linux,darwin}-{x64,arm64}.tar.gz`, which bundles an official Node.js runtime. Other platforms do **not** fall back automatically: the script exits and asks you to install Node.js >= 24 and re-run with `--universal`, which selects the runtime-less `penguin-universal.tar.gz`.
@@ -34,7 +34,7 @@ penguin -v
| Integrity check | Downloads are sha256-verified when the Release ships checksum assets |
| Upgrade | Re-run the install script; files are swapped atomically |
Script flags are passed as `curl ... | sh -s -- --universal`.
Script flags go after `sh -s --`, e.g. `curl -fsSL https://penguin.ooo/install.sh | sh -s -- --universal`.
### Data directory
+2 -2
View File
@@ -13,7 +13,7 @@ description: 通过安装脚本、npm 或源码安装 PenguinHarness。
在 Linux / macOS 上执行:
```bash
curl -fsSL https://github.com/Prism-Shadow/penguin-harness/releases/latest/download/install.sh | sh
curl -fsSL https://penguin.ooo/install.sh | sh
```
脚本按平台下载 `penguin-{linux,darwin}-{x64,arm64}.tar.gz`,其中捆绑了官方 Node.js 运行时。其他平台**不会自动回退**:脚本会退出并提示先安装 Node.js >= 24、再携带 `--universal` 重新执行,改用不含运行时的 `penguin-universal.tar.gz`。
@@ -34,7 +34,7 @@ penguin -v
| 完整性校验 | Release 提供 checksum 资产时自动进行 sha256 校验 |
| 升级 | 重新执行安装脚本即可,文件原子替换 |
脚本参数通过 `curl ... | sh -s -- --universal` 的形式传入。
脚本参数写在 `sh -s --` 之后,例如 `curl -fsSL https://penguin.ooo/install.sh | sh -s -- --universal`。
### 数据目录
+2 -2
View File
@@ -89,7 +89,7 @@ interface GenerativeModelConfig {
`GenerativeModel` (`packages/core/src/llm/generative-model.ts`) grounds the contract on the `AutoLLMClient` of the `@prismshadow/agenthub` model gateway:
- the gateway maintains conversation history **statefully**, receiving only new messages each turn; resuming a Session replays committed history through a one-time `setHistory`;
- an internal `EventTranslator` translates gateway stream events into `partial_*` fragments plus complete messages, preserving the `signature` / `phase` fidelity fields; complete messages settle in thinking → text → tool_call order;
- an internal `EventTranslator` translates gateway stream events into `partial_*` fragments plus complete messages, preserving each item's opaque `fidelity` payload verbatim; segmentation mirrors the gateway's own aggregation — a thinking block is closed by its fidelity payload and a run of equal fidelity stays one block (OpenAI-compatible clients stamp every delta with the same `{ reasoning_field }`, which must not split blocks), while a text segment splits on a differing `fidelity.phase` and closes on a `fidelity.signature`, fidelity keys accumulating on merge; complete messages settle in thinking → text → tool_call order;
- `ToolCallIdAllocator` disambiguates providers that use the function name as the call id (append `#n` inbound, strip outbound), scoped to the whole Session;
- provider differences (tool-call formats, reasoning content, streaming events) are absorbed entirely inside the gateway — see [Models & Providers](/models).
@@ -169,7 +169,7 @@ A tool emits only content deltas; framing, timeouts, truncation, `stop_reason` p
Human is deliberately not an interface class. The SDK caller *is* the Human:
```ts
const session = await agent.createSession({ workspaceDir, modelId });
const session = await agent.createSession({ workspaceDir, provider, modelId });
session.run(
newMessages: OmniMessage[], // input: the Prompt
+2 -2
View File
@@ -89,7 +89,7 @@ interface GenerativeModelConfig {
`GenerativeModel`(`packages/core/src/llm/generative-model.ts`)把契约落到模型网关 `@prismshadow/agenthub` 的 `AutoLLMClient` 上:
- 网关**有状态**地维护会话历史,每轮只接收新消息;恢复 Session 时经一次性的 `setHistory` 重放已提交历史;
- 内部的 `EventTranslator` 把网关流式事件翻译为 `partial_*` 分片 + 完整消息,保留 `signature` / `phase` 保真字段,完整消息按 thinking → text → tool_call 顺序落盘;
- 内部的 `EventTranslator` 把网关流式事件翻译为 `partial_*` 分片 + 完整消息,逐条原样保留不透明的 `fidelity` 保真负载;分段与网关自身的聚合一致——thinking 块由其 fidelity 负载闭合,连续相同的 fidelity 归为同一块(OpenAI 兼容客户端给每条增量盖同一个 `{ reasoning_field }`,不能因此切块),text 段遇到不同的 `fidelity.phase` 即切分、遇到 `fidelity.signature` 即闭合,合并时 fidelity 键累积;完整消息按 thinking → text → tool_call 顺序落盘;
- `ToolCallIdAllocator` 处理个别 Provider 用函数名充当调用 id 的情况(入站追加 `#n`、出站剥离),作用域覆盖整个 Session;
- Provider 协议差异(工具调用格式、思考内容、流式事件)全部在网关内抹平,见[模型与 Provider](/models)。
@@ -169,7 +169,7 @@ interface ToolDefinitionConfig {
Human 刻意不设计为接口类。SDK 的调用方就是 Human:
```ts
const session = await agent.createSession({ workspaceDir, modelId });
const session = await agent.createSession({ workspaceDir, provider, modelId });
session.run(
newMessages: OmniMessage[], // 输入:Prompt
+7 -2
View File
@@ -11,6 +11,8 @@ All model access goes through one gateway library: `@prismshadow/agenthub` (Auto
A model's identity is always the `(provider, model_id)` pair: `provider` is a config group name, `model_id` the upstream request id sent to AgentHub unchanged. The two are independent fields — concatenating them into one string is forbidden anywhere in the pipeline.
Every interface that names a model takes the complete pair: the CLI, the HTTP API, and the SDK all reject half a reference instead of completing it. The provider is never inferred from the model id and has no default, because gateways resell vendor models under their upstream ids — a guessed group would send the entry's credential to a vendor nobody named. Where a model reference is optional at all (`penguin run` / `chat`, Session creation, Schedules), the choice is between the whole pair and nothing: omit both halves to take the Project's default model.
## The per-Project model table
Each Project's available models are recorded in the hidden `.project_config.toml`, maintained via the CLI (`penguin config model add / default / list`, see [CLI Reference](/cli)) or the Web UI — never hand-edited. `ModelEntry` fields:
@@ -57,7 +59,10 @@ Built-in groups and their env-var fallbacks (catalog source: `packages/core/src/
| --- | --- | --- |
| deepseek | `DEEPSEEK_API_KEY` | Group of the default model |
| openrouter | `OPENAI_API_KEY` | OpenAI-compatible gateway, preset base URL `https://openrouter.ai/api/v1` |
| fireworks | `OPENAI_API_KEY` | Fireworks AI (OpenAI-compatible), preset base URL `https://api.fireworks.ai/inference/v1`; API model ids look like `accounts/fireworks/models/<slug>` |
| siliconflow | `OPENAI_API_KEY` | OpenAI-compatible gateway, preset base URL `https://api.siliconflow.cn/v1` |
| qwen-token-plan | `OPENAI_API_KEY` | Qwen Token Plan subscription gateway, preset base URL `https://token-plan.cn-beijing.maas.aliyuncs.com/compatible-mode/v1`; pricing from each model page's official list price (the preview model has only a quota-multiplier promo, no list price) |
| qwen-pay-as-you-go | `OPENAI_API_KEY` | Qwen pay-as-you-go (DashScope's OpenAI-compatible endpoint), preset base URL `https://dashscope.aliyuncs.com/compatible-mode/v1`; resold third-party models keep vendor-prefixed ids (e.g. `kimi/kimi-k3`) |
| google | `GEMINI_API_KEY` | |
| anthropic | `ANTHROPIC_API_KEY` | |
| openai | `OPENAI_API_KEY` | |
@@ -65,9 +70,9 @@ Built-in groups and their env-var fallbacks (catalog source: `packages/core/src/
| moonshot | `MOONSHOT_API_KEY` | |
| custom | `OPENAI_API_KEY` | Any OpenAI-protocol endpoint |
The gateway groups (openrouter / siliconflow) go through AgentHub's OpenAI client, so with blank credentials they read `OPENAI_API_KEY` — not a gateway-specific variable.
The gateway groups (openrouter / fireworks / siliconflow / qwen-token-plan / qwen-pay-as-you-go) go through AgentHub's OpenAI client, so with blank credentials they read `OPENAI_API_KEY` — not a gateway-specific variable.
Some models in the preset catalog: deepseek-v4-pro / deepseek-v4-flash, gemini-3.1-pro-preview, claude-opus-4-8 / claude-sonnet-4-6, gpt-5.5, glm-5.2, kimi-k2.6 (not exhaustive).
Some models in the preset catalog: deepseek-v4-pro / deepseek-v4-flash, gemini-3.1-pro-preview, claude-opus-4-8 / claude-sonnet-4-6, gpt-5.5, glm-5.2, kimi-k2.6, qwen3.8-max-preview (not exhaustive).
## Thinking levels
+7 -2
View File
@@ -11,6 +11,8 @@ description: 经 AgentHub 单一网关接入模型,以 (provider, model_id)
模型身份永远是 `(provider, model_id)` 成对表示:`provider` 是配置分组名,`model_id` 是原样发给上游的请求 id。二者是两个独立字段,任何环节都不允许拼接成一个字符串。
所有涉及模型的接口都要求完整的二元组:CLI、HTTP API 与 SDK 都会拒绝半个引用,而不会替你补全。provider 绝不由模型 id 推断,也没有缺省值——网关会以上游 id 转售厂商模型,猜出来的分组会把该条目的凭据发往无人指定的厂商。凡是模型引用本身可省略之处(`penguin run` / `chat`、创建 Session、定时任务),可选的是整对:两半都省略即使用 Project 默认模型。
## Project 模型表
每个 Project 的可用模型记录在隐藏文件 `.project_config.toml` 中,由 CLI(`penguin config model add / default / list`,见 [CLI 参考](/cli))或 Web 界面维护,不手工编辑。`ModelEntry` 字段:
@@ -57,7 +59,10 @@ api_key = "sk-..."
| --- | --- | --- |
| deepseek | `DEEPSEEK_API_KEY` | 默认模型所在分组 |
| openrouter | `OPENAI_API_KEY` | OpenAI 兼容网关,预置 base URL `https://openrouter.ai/api/v1` |
| fireworks | `OPENAI_API_KEY` | Fireworks AI(OpenAI 兼容),预置 base URL `https://api.fireworks.ai/inference/v1`;API 模型 id 形如 `accounts/fireworks/models/<slug>` |
| siliconflow | `OPENAI_API_KEY` | OpenAI 兼容网关,预置 base URL `https://api.siliconflow.cn/v1` |
| qwen-token-plan | `OPENAI_API_KEY` | Qwen Token Plan 订阅网关,预置 base URL `https://token-plan.cn-beijing.maas.aliyuncs.com/compatible-mode/v1`;定价取各模型页官方牌价(预览模型仅配额倍率促销、无牌价) |
| qwen-pay-as-you-go | `OPENAI_API_KEY` | Qwen 按量付费(DashScope OpenAI 兼容端),预置 base URL `https://dashscope.aliyuncs.com/compatible-mode/v1`;转售第三方模型保留厂商前缀 id(如 `kimi/kimi-k3`) |
| google | `GEMINI_API_KEY` | |
| anthropic | `ANTHROPIC_API_KEY` | |
| openai | `OPENAI_API_KEY` | |
@@ -65,9 +70,9 @@ api_key = "sk-..."
| moonshot | `MOONSHOT_API_KEY` | |
| custom | `OPENAI_API_KEY` | 任意 OpenAI 协议端点 |
网关分组(openrouter / siliconflow)经 AgentHub 的 OpenAI 客户端请求,因此凭证留空时读取的是 `OPENAI_API_KEY`,而非网关自己的变量名。
网关分组(openrouter / fireworks / siliconflow / qwen-token-plan / qwen-pay-as-you-go)经 AgentHub 的 OpenAI 客户端请求,因此凭证留空时读取的是 `OPENAI_API_KEY`,而非网关自己的变量名。
预置目录中的部分模型:deepseek-v4-pro / deepseek-v4-flash、gemini-3.1-pro-preview、claude-opus-4-8 / claude-sonnet-4-6、gpt-5.5、glm-5.2、kimi-k2.6 等(非完整清单)。
预置目录中的部分模型:deepseek-v4-pro / deepseek-v4-flash、gemini-3.1-pro-preview、claude-opus-4-8 / claude-sonnet-4-6、gpt-5.5、glm-5.2、kimi-k2.6、qwen3.8-max-preview 等(非完整清单)。
## 思考等级
+22 -8
View File
@@ -54,15 +54,16 @@ On resume, the engine takes this Trace line as the runtime config — the model,
## model_msg: complete payloads
Seven content payloads, discriminated by `payload.type`. Shared optional fields: `stop_reason` (marks an abnormal terminal state) and `signature` (a provider-fidelity field, see below):
Seven content payloads, discriminated by `payload.type`. Shared optional fields: `stop_reason` (marks an abnormal terminal state) and `fidelity` (an opaque provider-fidelity payload, see below):
```ts
type Fidelity = Record<string, unknown>; // opaque provider-fidelity payload (see below)
interface TextPayload {
type: "text";
role: "user" | "assistant";
text: string;
phase?: string | null; // segmentation marker (e.g. GPT-5 phases)
signature?: string;
fidelity?: Fidelity; // e.g. { phase } segment marker (GPT-5), { signature }
stop_reason?: StopReason;
}
@@ -70,7 +71,7 @@ interface ThinkingPayload {
type: "thinking";
role: "assistant";
thinking: string;
signature?: string; // required by some models to replay history
fidelity?: Fidelity; // required by some models to replay history
stop_reason?: StopReason;
}
@@ -79,7 +80,7 @@ interface InlineThinkingPayload {
role: "assistant";
data: string; // reasoning content in binary form
mime_type: string;
signature?: string;
fidelity?: Fidelity;
stop_reason?: StopReason;
}
@@ -89,7 +90,7 @@ interface ToolCallPayload {
name: string;
arguments: string; // arguments as a JSON string
tool_call_id: string;
signature?: string;
fidelity?: Fidelity;
stop_reason?: StopReason;
}
@@ -114,7 +115,7 @@ interface InlineDataPayload {
role: "user" | "assistant";
data: string; // other binary content
mime_type: string;
signature?: string;
fidelity?: Fidelity;
stop_reason?: StopReason;
}
```
@@ -273,7 +274,20 @@ Errors never cross an interface boundary as exceptions — they *are* messages.
## Provider-fidelity fields
Provider-specific fields such as `signature` and `phase` pass through and persist verbatim end to end — some models require them byte-for-byte when history is replayed, and any rewriting would break compatibility. This is one of the preconditions for lossless Session recovery from the Trace.
Provider-specific wire data travels in a single optional field, `fidelity` — an arbitrary JSON object the LLM client records to reproduce the original message on replay: thinking signatures, `phase` segment labels, GPT-5 encrypted reasoning, the OpenAI-compatible upstream reasoning field name:
```ts
// Claude: a thinking block closed by its signature
{ type: "thinking", thinking: "…", fidelity: { signature: "EqQBCkYIBxgCKkB…" } }
// GPT-5: encrypted reasoning (empty thinking text, fidelity only)
{ type: "thinking", thinking: "", fidelity: { id: "rs_0d3…", encrypted_content: "gAAAA…" } }
// OpenAI-compatible: the upstream field the reasoning text came from
{ type: "thinking", thinking: "…", fidelity: { reasoning_field: "reasoning_content" } }
```
The payload is opaque to PenguinHarness: it passes through and persists verbatim end to end — some models require it byte-for-byte when history is replayed, and any rewriting (or loss) would break compatibility. This is one of the preconditions for lossless Session recovery from the Trace.
## Three jobs, one protocol
+22 -8
View File
@@ -54,15 +54,16 @@ interface ToolDefinition {
## model_msg:完整消息
七种内容 payload,以 `payload.type` 判别。公共可选字段:`stop_reason`(非正常收尾时标注终态)与 `signature`(Provider 保真字段,见下文):
七种内容 payload,以 `payload.type` 判别。公共可选字段:`stop_reason`(非正常收尾时标注终态)与 `fidelity`(不透明的 Provider 保真负载,见下文):
```ts
type Fidelity = Record<string, unknown>; // 不透明的 Provider 保真负载(见下文)
interface TextPayload {
type: "text";
role: "user" | "assistant";
text: string;
phase?: string | null; // 分段标记(如 GPT-5 的 phase)
signature?: string;
fidelity?: Fidelity; // 如 { phase } 分段标记(GPT-5)、{ signature }
stop_reason?: StopReason;
}
@@ -70,7 +71,7 @@ interface ThinkingPayload {
type: "thinking";
role: "assistant";
thinking: string;
signature?: string; // 部分模型历史回放所必需
fidelity?: Fidelity; // 部分模型历史回放所必需
stop_reason?: StopReason;
}
@@ -79,7 +80,7 @@ interface InlineThinkingPayload {
role: "assistant";
data: string; // 二进制形态的思考内容
mime_type: string;
signature?: string;
fidelity?: Fidelity;
stop_reason?: StopReason;
}
@@ -89,7 +90,7 @@ interface ToolCallPayload {
name: string;
arguments: string; // 参数 JSON 字符串
tool_call_id: string;
signature?: string;
fidelity?: Fidelity;
stop_reason?: StopReason;
}
@@ -114,7 +115,7 @@ interface InlineDataPayload {
role: "user" | "assistant";
data: string; // 其他二进制内容
mime_type: string;
signature?: string;
fidelity?: Fidelity;
stop_reason?: StopReason;
}
```
@@ -272,7 +273,20 @@ type StopReason = "completed" | "failed" | "aborted" | "timeout" | "malformed";
## 保真字段
`signature` 与 `phase` 等 Provider 专有字段在整条链路上原样透传、原样存储——部分模型在历史回放时要求逐字一致,任何转写都会破坏兼容性。这是 Trace 能够无损恢复 Session 的前提之一。
Provider 专有的线上数据统一收拢在一个可选字段 `fidelity` 中——LLM 客户端为历史回放记录的任意 JSON 对象:思考签名、`phase` 分段标记、GPT-5 加密推理、OpenAI 兼容上游的推理字段名:
```ts
// Claude:由签名闭合的 thinking 块
{ type: "thinking", thinking: "…", fidelity: { signature: "EqQBCkYIBxgCKkB…" } }
// GPT-5:加密推理(thinking 文本为空,仅有 fidelity)
{ type: "thinking", thinking: "", fidelity: { id: "rs_0d3…", encrypted_content: "gAAAA…" } }
// OpenAI 兼容:思考内容来自上游哪个字段
{ type: "thinking", thinking: "…", fidelity: { reasoning_field: "reasoning_content" } }
```
该负载对 PenguinHarness 完全不透明:在整条链路上原样透传、原样存储——部分模型在历史回放时要求逐字一致,任何转写或丢失都会破坏兼容性。这是 Trace 能够无损恢复 Session 的前提之一。
## 协议的三种职责
+4 -4
View File
@@ -8,7 +8,7 @@ description: Install PenguinHarness, configure a model, and run your first Task.
One-liner for Linux / macOS:
```bash
curl -fsSL https://github.com/Prism-Shadow/penguin-harness/releases/latest/download/install.sh | sh
curl -fsSL https://penguin.ooo/install.sh | sh
```
For other options (npm, from source), see [Installation](/installation).
@@ -18,10 +18,10 @@ For other options (npm, from source), see [Installation](/installation).
PenguinHarness ships with no built-in model credentials, so configure a model first. Use the Models page in the Web UI, or the CLI:
```bash
penguin config model add --model-id deepseek-v4-pro --api-key sk-... --set-default
penguin config model add --provider deepseek --model-id deepseek-v4-pro --api-key sk-... --set-default
```
- When `--provider` is omitted, the Provider is inferred from the built-in catalog.
- A model is always referenced as a `(provider, model_id)` pair, so `--provider` and `--model-id` are both required — the Provider is never inferred from the model id. See [Models & Providers](/models) for the built-in groups.
- The API key can also come from environment variables: when a model entry has no inline api_key, AgentHub (the LLM gateway library) reads variables such as `DEEPSEEK_API_KEY`, `ANTHROPIC_API_KEY`, `OPENAI_API_KEY`, and `GEMINI_API_KEY`. A `.env` file in the working directory is loaded automatically.
## Start the Web App
@@ -30,7 +30,7 @@ penguin config model add --model-id deepseek-v4-pro --api-key sk-... --set-defau
penguin web
```
The service runs at http://127.0.0.1:7364 and opens your browser (`--no-open` to skip). First login is `admin` / `admin123` — change it right away. `penguin server` starts the same process headless.
The service runs at http://127.0.0.1:7364 and opens your browser (`--no-open` to skip). First login is `admin` / `penguin-2026` — change it right away. `penguin server` starts the same process headless.
## One-shot run
+4 -4
View File
@@ -8,7 +8,7 @@ description: 安装 PenguinHarness、配置模型并运行第一个 Task。
Linux / macOS 一键安装:
```bash
curl -fsSL https://github.com/Prism-Shadow/penguin-harness/releases/latest/download/install.sh | sh
curl -fsSL https://penguin.ooo/install.sh | sh
```
其他方式(npm、源码)见[安装](/installation)。
@@ -18,10 +18,10 @@ curl -fsSL https://github.com/Prism-Shadow/penguin-harness/releases/latest/downl
PenguinHarness 不内置任何模型凭据,使用前需要先配置一个模型。可以在 Web UI 的 Models 页面完成,也可以用 CLI:
```bash
penguin config model add --model-id deepseek-v4-pro --api-key sk-... --set-default
penguin config model add --provider deepseek --model-id deepseek-v4-pro --api-key sk-... --set-default
```
- 省略 `--provider` 时,根据内置目录自动推断 Provider。
- 模型引用始终是 `(provider, model_id)` 二元组,因此 `--provider` 与 `--model-id` 均为必填——Provider 绝不由模型 id 推断。内置分组见[模型与 Provider](/models)。
- API Key 也可以来自环境变量:当模型条目没有内联 api_key 时,LLM 网关库 AgentHub 会读取 `DEEPSEEK_API_KEY`、`ANTHROPIC_API_KEY`、`OPENAI_API_KEY`、`GEMINI_API_KEY` 等变量;工作目录下的 `.env` 会被自动加载。
## 启动 Web App
@@ -30,7 +30,7 @@ penguin config model add --model-id deepseek-v4-pro --api-key sk-... --set-defau
penguin web
```
服务运行在 http://127.0.0.1:7364 并自动打开浏览器(`--no-open` 跳过)。首次登录使用 `admin` / `admin123`,请立即修改密码。`penguin server` 启动同一进程的 headless 版本。
服务运行在 http://127.0.0.1:7364 并自动打开浏览器(`--no-open` 跳过)。首次登录使用 `admin` / `penguin-2026`,请立即修改密码。`penguin server` 启动同一进程的 headless 版本。
## 单次运行
+17 -5
View File
@@ -35,12 +35,12 @@ packages/server/src
- Cookie session: `penguin_session` (HttpOnly, SameSite=Lax), valid for 7 days with sliding renewal;
- Passwords are stored as scrypt hashes; the server keeps only the sha256 of the session token, never the plaintext;
- No open registration: the built-in admin `admin` / `admin123` is seeded at startup, and all other accounts are created by an admin;
- No open registration: the built-in admin `admin` / `penguin-2026` is seeded at startup, and all other accounts are created by an admin;
- Same-origin only — no CORS middleware is enabled.
```bash
curl -c cookies.txt -H "Content-Type: application/json" \
-d '{"userId":"admin","password":"admin123"}' \
-d '{"userId":"admin","password":"penguin-2026"}' \
http://127.0.0.1:7364/api/auth/login
```
@@ -87,6 +87,8 @@ Member writes are owner-only.
| PUT | /api/projects/:projectId/models | Full-table replace, keyed by `(provider, modelId)` |
| POST | /api/projects/:projectId/models/test | Connectivity test: `{provider, modelId, …}` → `{ok, latencyMs?, message?}` |
Every endpoint that names a model takes the complete `(provider, modelId)` pair. Nothing is inferred: a request carrying only one half is a 400, never a lookup. Where the reference itself is optional (Session creation, Schedules), omitting both halves selects the Project's default model.
### Agents
The paths below omit the `/api/projects/:projectId` prefix.
@@ -110,7 +112,7 @@ The paths below omit the `/api/projects/:projectId` prefix.
| GET / POST | /agents/:agentId/schedules | List scheduled tasks / create one (409 if the name exists) |
| GET / PUT / DELETE | /agents/:agentId/schedules/:name | Read / update / delete a single task |
Schedule writes are owner-only.
Schedule writes are owner-only. A task in new-Session mode carries `modelId` and `provider` together or not at all; the pair is checked against the Project's model table when the task is saved and again when the scheduler reconciles it.
### Session Creation and Directory Browsing
@@ -120,7 +122,7 @@ Schedule writes are owner-only.
| POST | /agents/:agentId/sessions | Create a Session: `{modelId?, provider?, workspace?, approvalMode?}` → 201 |
| GET | /dirs?path= | Server-side directory browser (backs the Workspace picker) |
On Session creation, the model defaults to the Project's default model, the Workspace defaults to an auto-created temporary directory, and the approval mode defaults to `allow-all`.
On Session creation, `modelId` and `provider` are both-or-neither: send the complete pair to pick a model, or omit both to take the Project's default model — one without the other is a 400. The Workspace defaults to an auto-created temporary directory, and the approval mode defaults to `allow-all`.
### Usage and Traces (Agent Level)
@@ -147,7 +149,7 @@ The paths below omit the `/api/sessions/:sessionId` prefix. For the storage mode
| POST | /abort | Interrupt the current Task: 202 when triggered, 204 when idle |
| POST | /compact | Trigger context compaction: 202; 409 `nothing_to_compact` when there is nothing to compact |
| GET | /files?path= | Browse the Workspace directory |
| GET | /files/content?path=&download= | Read a Workspace file (`download=1` serves it as an attachment) |
| GET | /files/content?path=&download=&preview= | Read a Workspace file (`download=1` serves it as an attachment, `preview=1` renders it in a sandbox — see below) |
| POST | /files/stat | Batch existence check: `{paths}` |
| PUT | /files/content?path= | Upload a file: `{dataBase64}`, capped at 14MB |
| GET | /traces | List this Session's Trace files |
@@ -157,6 +159,16 @@ The paths below omit the `/api/sessions/:sessionId` prefix. For the storage mode
General conventions: Sessions the user cannot access always return 404 — their existence is never leaked; only one Task or compaction runs per Session at a time, and conflicts return 409 (`task_in_progress` / `compacting`).
Workspace files may be Agent-generated, so `GET /files/content` treats them as untrusted: every response carries `X-Content-Type-Options: nosniff`, and the rest of the headers depend on the two flags (`download=1` wins over `preview=1`):
| Query | Content-Type | Content-Disposition | Content-Security-Policy |
| --- | --- | --- | --- |
| neither | `text/plain; charset=utf-8` for `.html` / `.htm` / `.svg`, the real type otherwise | `inline` | — |
| `preview=1` | the real type (`text/html`, `image/svg+xml`, …) | `inline` | `sandbox allow-scripts allow-popups allow-modals allow-forms`, sent only for `.html` / `.htm` / `.svg` |
| `download=1` | the real type | `attachment` | — |
The filename always rides along as `filename*=UTF-8''` with percent-encoding. `preview=1` is what backs "open in a new tab": the document keeps its real type and does render and run, but the sandbox deliberately omits `allow-same-origin`, so it lands in an opaque origin and can reach neither this origin's cookies nor the API — while the request itself still authenticates, because a top-level GET sends the SameSite=Lax session cookie.
Key request bodies (explicit keys):
```ts
+17 -5
View File
@@ -35,12 +35,12 @@ packages/server/src
- Cookie 会话:`penguin_session`(HttpOnly、SameSite=Lax),有效期 7 天,滑动续期;
- 密码以 scrypt 哈希存储;服务端只保存会话 Token 的 sha256,不落明文;
- 不开放注册:启动时种子化内置管理员 `admin` / `admin123`,其余账号由管理员创建;
- 不开放注册:启动时种子化内置管理员 `admin` / `penguin-2026`,其余账号由管理员创建;
- 仅限同源访问,未启用 CORS 中间件。
```bash
curl -c cookies.txt -H "Content-Type: application/json" \
-d '{"userId":"admin","password":"admin123"}' \
-d '{"userId":"admin","password":"penguin-2026"}' \
http://127.0.0.1:7364/api/auth/login
```
@@ -87,6 +87,8 @@ curl -c cookies.txt -H "Content-Type: application/json" \
| PUT | /api/projects/:projectId/models | 全表替换,条目以 `(provider, modelId)` 为键 |
| POST | /api/projects/:projectId/models/test | 连通性测试:`{provider, modelId, …}` → `{ok, latencyMs?, message?}` |
所有涉及模型的接口都要求完整的 `(provider, modelId)` 二元组,不做任何推断:只带一半的请求一律 400,绝不会退化为一次查找。模型引用本身可省略的场景(创建 Session、定时任务)省略的是整对,两半都不给即选用 Project 默认模型。
### Agent
以下路径均省略前缀 `/api/projects/:projectId`。
@@ -110,7 +112,7 @@ curl -c cookies.txt -H "Content-Type: application/json" \
| GET / POST | /agents/:agentId/schedules | 定时任务列表 / 创建(重名返回 409) |
| GET / PUT / DELETE | /agents/:agentId/schedules/:name | 读取 / 更新 / 删除单个任务 |
Schedule 写操作仅限 Owner。
Schedule 写操作仅限 Owner。新建 Session 模式的任务,`modelId` 与 `provider` 要么成对给出、要么都不给;该二元组会在任务保存时以及调度器对账时对照 Project 模型表校验。
### Session 创建与目录浏览
@@ -120,7 +122,7 @@ Schedule 写操作仅限 Owner。
| POST | /agents/:agentId/sessions | 创建 Session:`{modelId?, provider?, workspace?, approvalMode?}` → 201 |
| GET | /dirs?path= | 服务器端目录浏览(Workspace 选择器数据源) |
创建 Session 时,模型默认取 Project 默认模型,Workspace 默认自动创建临时目录,审批模式默认 `allow-all`。
创建 Session 时,`modelId` 与 `provider` 要么成对给出、要么都不给:给出完整二元组即指定模型,两个都省略则取 Project 默认模型,只给一个返回 400。Workspace 默认自动创建临时目录,审批模式默认 `allow-all`。
### 用量与 Trace(Agent 级)
@@ -147,7 +149,7 @@ Schedule 写操作仅限 Owner。
| POST | /abort | 中断当前 Task:已触发返回 202,无任务返回 204 |
| POST | /compact | 触发上下文压缩:202;无可压缩内容返回 409 `nothing_to_compact` |
| GET | /files?path= | 浏览 Workspace 目录 |
| GET | /files/content?path=&download= | 读取 Workspace 文件(`download=1` 时作为附件下载) |
| GET | /files/content?path=&download=&preview= | 读取 Workspace 文件(`download=1` 时作为附件下载,`preview=1` 以沙箱方式预览 —— 见下) |
| POST | /files/stat | 批量存在性检查:`{paths}` |
| PUT | /files/content?path= | 上传文件:`{dataBase64}`,上限 14MB |
| GET | /traces | 本 Session 的 Trace 文件列表 |
@@ -157,6 +159,16 @@ Schedule 写操作仅限 Owner。
通用约定:无权访问的 Session 一律返回 404,不泄露其存在性;每个 Session 同时只允许一个 Task 或压缩在运行,冲突时返回 409(`task_in_progress` / `compacting`)。
Workspace 文件可能由 Agent 生成,`GET /files/content` 一律按不可信内容处理:所有响应都带 `X-Content-Type-Options: nosniff`,其余响应头取决于两个开关(`download=1` 优先于 `preview=1`):
| 查询参数 | Content-Type | Content-Disposition | Content-Security-Policy |
| --- | --- | --- | --- |
| 都不带 | `.html` / `.htm` / `.svg` 降级为 `text/plain; charset=utf-8`,其余为真实类型 | `inline` | 无 |
| `preview=1` | 真实类型(`text/html`、`image/svg+xml` 等) | `inline` | `sandbox allow-scripts allow-popups allow-modals allow-forms`,仅对 `.html` / `.htm` / `.svg` 下发 |
| `download=1` | 真实类型 | `attachment` | 无 |
文件名始终以 `filename*=UTF-8''` 形式携带(百分号编码)。`preview=1` 支撑的是“在新标签页打开”:文档保留真实类型,可以正常渲染并执行脚本,但沙箱刻意不含 `allow-same-origin`,因此它落在一个不透明源里,既拿不到本源的 Cookie,也调不动 API;而请求本身仍然要鉴权 —— 顶层 GET 会带上 SameSite=Lax 的会话 Cookie。
关键请求体(明确键名):
```ts
@@ -81,7 +81,7 @@ Special case: if the latest Trace file ends with a completed compaction, that co
## Field fidelity
Provider-specific fields (such as `signature` and `phase`) are preserved verbatim in the Trace and sent back verbatim — some models require them byte-for-byte on history replay, and any rewriting would break compatibility. This is one reason the Trace stores raw OmniMessage envelopes rather than a post-processed format.
Each content message's opaque provider `fidelity` payload (thinking signatures, phase labels, encrypted reasoning, …) is preserved verbatim in the Trace and sent back verbatim — some models require it byte-for-byte on history replay, and any rewriting would break compatibility. This is one reason the Trace stores raw OmniMessage envelopes rather than a post-processed format.
## Observability
@@ -81,7 +81,7 @@ Trace 是恢复的唯一事实来源,没有独立的会话数据库需要与
## 字段保真
Provider 专有字段(如 `signature`、`phase`)在 Trace 中原样保存、原样回传——部分模型在历史回放时要求这些字段逐字一致,任何转写都会破坏兼容性。这也是 Trace 直接存储 OmniMessage 信封而非二次加工格式的原因之一。
内容消息携带的不透明 Provider 保真负载 `fidelity`(思考签名、phase 分段标记、加密推理等)在 Trace 中原样保存、原样回传——部分模型在历史回放时要求该负载逐字一致,任何转写都会破坏兼容性。这也是 Trace 直接存储 OmniMessage 信封而非二次加工格式的原因之一。
## 可观测性
+8 -7
View File
@@ -58,16 +58,17 @@ The built-in Skills, by group (the group manifest is `SKILL_GROUPS` in `packages
| Group | Skill | Purpose |
| --- | --- | --- |
| Agent Development | `agent-creation` | Turn a user requirement into a concrete agent: write the target agent's AGENTS.md and install the skills it needs |
| Office Productivity | `data-analysis` | Complete data-analysis tasks with bounded evidence inspection, explicit answer-changing decisions, native artifact handling and final output verification |
| | `firecrawl` | Web search and page scraping into clean markdown via the Firecrawl API |
| Software Development | `web-design` | Penguin visual language for generated web pages and app UIs: design tokens, components, light/dark themes and chat layouts |
| | `software-engineering` | Complete software-engineering tasks: investigate and review code, implement fixes, features and refactors with minimal scope, validate changes, and report verified outcomes |
| AI App Development | `penguin-sdk` | Build AI and RAG apps on the SDK: the createSession/run streaming loop plus a complete retrieval recipe with chunk-revealing citations |
| | `penguin-cli` | Manage model API keys, default models and per-agent Vault secrets with the penguin CLI |
| | `agenthub-models` | Call model APIs through `@prismshadow/agenthub`: streaming text, image generation, speech synthesis and embeddings |
| Agent Tuning | `agent-creation` | Turn a user requirement into a concrete agent: write the target agent's AGENTS.md and install the skills it needs |
| | `benchmark-design` | Design and calibrate a multi-Case capability Benchmark with repeated independent evaluations and a traceable baseline |
| | `agent-evaluation` | Run and score exactly one Benchmark Case run, with CLI execution, Trace provenance checks and private Rubric isolation |
| | `agent-optimization` | Improve an Agent State from direct feedback or versioned multi-Case Benchmark scores and score-linked Traces |
| Data Analysis | `data-analysis` | Complete data-analysis tasks with bounded evidence inspection, explicit answer-changing decisions, native artifact handling and final output verification |
| Penguin Development | `penguin-sdk` | Build AI apps on the SDK (the createSession/run streaming loop) |
| | `penguin-cli` | Manage model API keys, default models and per-agent Vault secrets with the penguin CLI |
| | `agenthub-models` | Call model APIs through `@prismshadow/agenthub`: streaming text, image generation, speech synthesis and embeddings |
| Web Development | `web-design` | Default visual language for generated web pages: minimal black-white-gray |
| Software Engineering | `software-engineering` | Complete software-engineering tasks: investigate and review code, implement fixes, features and refactors with minimal scope, validate changes, and report verified outcomes |
## Writing and optimizing Skills
+8 -7
View File
@@ -58,16 +58,17 @@ Skill 库以 npm 包 `@prismshadow/penguin-skills` 发布,tarball 直接携带
| 分组 | Skill | 说明 |
| --- | --- | --- |
| Agent 开发 | `agent-creation` | 把用户需求变成具体的 Agent:撰写目标 Agent 的 AGENTS.md 并安装所需 Skill |
| 办公效率 | `data-analysis` | 以有界的证据检查、显式的改答案决策、原生产物处理与最终输出校验完成数据分析任务 |
| | `firecrawl` | 经 Firecrawl API 做网络搜索与页面抓取,产出干净的 Markdown |
| 软件开发 | `web-design` | 生成网页与应用界面的 Penguin 视觉语言:设计令牌、组件配方、明暗主题与聊天布局 |
| | `software-engineering` | 完成软件工程任务:调查与审查代码,以最小改动实现修复、特性与重构,验证改动并报告经过确认的结果 |
| AI 应用开发 | `penguin-sdk` | 基于 SDK 构建 AI 与 RAG 应用:createSession/run 流式循环,外加带可溯源引用的完整检索配方 |
| | `penguin-cli` | 用 penguin CLI 管理模型 API Key、默认模型与各 Agent 的 Vault 密钥 |
| | `agenthub-models` | 经 `@prismshadow/agenthub` 调用模型 API:流式文本、图像生成、语音合成与 Embedding |
| Agent 调优 | `agent-creation` | 把用户需求变成具体的 Agent:撰写目标 Agent 的 AGENTS.md 并安装所需 Skill |
| | `benchmark-design` | 设计并校准多 Case 的能力评测 Benchmark,含重复独立评测与可追溯基线 |
| | `agent-evaluation` | 隔离执行并评分单个 Benchmark Case:CLI 执行、Trace 溯源检查、Rubric 私有隔离 |
| | `agent-optimization` | 依据直接反馈或带版本的多 Case Benchmark 分数与关联 Trace 改进 Agent State |
| 数据分析 | `data-analysis` | 以有界的证据检查、显式的改答案决策、原生产物处理与最终输出校验完成数据分析任务 |
| Penguin 开发 | `penguin-sdk` | 基于 SDK 构建 AI 应用(createSession/run 流式循环) |
| | `penguin-cli` | 用 penguin CLI 管理模型 API Key、默认模型与各 Agent 的 Vault 密钥 |
| | `agenthub-models` | 经 `@prismshadow/agenthub` 调用模型 API:流式文本、图像生成、语音合成与 Embedding |
| 网页开发 | `web-design` | 生成网页的默认视觉规范:极简黑白灰 |
| 软件工程 | `software-engineering` | 完成软件工程任务:调查与审查代码,以最小改动实现修复、特性与重构,验证改动并报告经过确认的结果 |
## 编写与优化
+1 -1
View File
@@ -23,7 +23,7 @@ penguin web
# open http://127.0.0.1:7364
```
The initial account is `admin` / `admin123`. There is no self-registration: accounts are created by an admin on the user-management page, and every new user automatically gets an independent initial Project named `<userId>-default_project`. While the initial password is still in use, a banner prompts the user to change it.
The initial account is `admin` / `penguin-2026`. There is no self-registration: accounts are created by an admin on the user-management page, and every new user automatically gets an independent initial Project named `<userId>-default_project`. While the initial password is still in use, a banner prompts the user to change it.
Logins persist for 7 days with sliding renewal; an admin password reset invalidates all of that user's login sessions.
+1 -1
View File
@@ -23,7 +23,7 @@ penguin web
# 打开 http://127.0.0.1:7364
```
初始账号为 `admin` / `admin123`。系统不开放自助注册:账号由管理员在用户管理页创建;每个新用户会自动获得一个独立的初始 Project,命名为 `<userId>-default_project`。仍在使用初始密码时,页面会以横幅提示尽快修改。
初始账号为 `admin` / `penguin-2026`。系统不开放自助注册:账号由管理员在用户管理页创建;每个新用户会自动获得一个独立的初始 Project,命名为 `<userId>-default_project`。仍在使用初始密码时,页面会以横幅提示尽快修改。
登录状态保持 7 天(滑动续期);管理员重置密码会使该用户的全部登录会话失效。
+82 -10
View File
@@ -1,16 +1,69 @@
/**
* Sticky top bar: logo + site name + "Docs" badge, then (right) a link back to the
* main site, GitHub, language/theme toggles and — on small screens — the sidebar
* toggle. The sidebar itself lives in the layout (router.tsx); this bar only flips
* its open state.
* Sticky top bar, landing-parity: logo + site name + "Docs" badge, then the SAME
* link row as the landing page's nav (section anchors on the landing home, blog,
* docs) with the same sliding hover pill — so the two sites link into each other
* seamlessly. Section/blog links are plain anchors into the landing SPA one level
* up; "Docs" routes to this site's own root. The right side keeps the language
* and theme toggles, GitHub, and — on small screens — the sidebar toggle.
*/
import { useRef } from "react";
import type { MouseEvent } from "react";
import { Link } from "react-router";
import { S } from "../lib/strings";
import { REPO_URL, SITE_URL } from "../lib/links";
import { GitHubIcon, MenuIcon, XIcon } from "./icons";
import { ThemeToggle } from "./theme-toggle";
import { LangToggle } from "./lang-toggle";
const SECTION_IDS = ["highlights", "quickstart", "benchmark", "contract", "features"] as const;
export function Nav({ menuOpen, onToggleMenu }: { menuOpen: boolean; onToggleMenu: () => void }) {
const pillRef = useRef<HTMLSpanElement | null>(null);
const pillVisible = useRef(false);
const sectionLabel: Record<(typeof SECTION_IDS)[number], string> = {
highlights: S.nav.highlights,
quickstart: S.nav.quickstart,
benchmark: S.nav.benchmark,
contract: S.nav.contract,
features: S.nav.features,
};
const linkCls =
"relative z-10 rounded-md px-2.5 py-1.5 text-sm text-gray-600 transition-colors hover:text-gray-900 dark:text-gray-400 dark:hover:text-gray-100";
// Landing-parity active style: on the docs site, the "Docs" link is the current page.
const activeLinkCls =
"relative z-10 rounded-md px-2.5 py-1.5 text-sm transition-colors bg-black text-white hover:bg-black hover:text-white dark:bg-black dark:text-white dark:ring-1 dark:ring-gray-600 dark:hover:bg-black dark:hover:text-white";
/**
* The hover pill appears IN PLACE under the first link it lands on (position
* jumps with only the fade animating), slides while moving between links, and
* fades out where it is on leave — never sweeping in from the nav's edge.
*/
const slideTo = (e: MouseEvent<HTMLElement>) => {
const el = e.currentTarget;
const pill = pillRef.current;
if (!pill) return;
if (!pillVisible.current) {
pill.style.transitionProperty = "opacity";
pill.style.left = `${el.offsetLeft}px`;
pill.style.width = `${el.offsetWidth}px`;
void pill.offsetWidth; // flush the jump before restoring the full transition
pill.style.transitionProperty = "";
pillVisible.current = true;
} else {
pill.style.left = `${el.offsetLeft}px`;
pill.style.width = `${el.offsetWidth}px`;
}
pill.style.opacity = "1";
};
const hidePill = () => {
const pill = pillRef.current;
if (pill) pill.style.opacity = "0";
pillVisible.current = false;
};
return (
<header className="sticky top-0 z-40 border-b border-gray-200 bg-white/85 backdrop-blur dark:border-gray-800 dark:bg-gray-950/85">
<div className="mx-auto flex h-14 max-w-7xl items-center gap-2 px-4 sm:px-6">
@@ -32,13 +85,32 @@ export function Nav({ menuOpen, onToggleMenu }: { menuOpen: boolean; onToggleMen
</span>
</a>
<div className="ml-auto flex items-center gap-1">
<a
href={SITE_URL}
className="hidden rounded-md px-2.5 py-1.5 text-sm text-gray-600 transition-colors hover:text-gray-900 sm:inline-block dark:text-gray-400 dark:hover:text-gray-100"
>
{S.nav.home}
<nav
className="relative ml-4 hidden items-center gap-0.5 md:flex"
aria-label="Primary"
onMouseLeave={hidePill}
>
{/* Sliding hover pill: appears in place, slides between links, fades out in place. */}
<span
ref={pillRef}
aria-hidden="true"
className="absolute top-1/2 h-8 -translate-y-1/2 rounded-md bg-gray-100 transition-[left,width,opacity] duration-200 ease-out dark:bg-gray-800"
style={{ left: 0, width: 0, opacity: 0 }}
/>
{SECTION_IDS.map((id) => (
<a key={id} href={`${SITE_URL}#${id}`} className={linkCls} onMouseEnter={slideTo}>
{sectionLabel[id]}
</a>
))}
<a href={`${SITE_URL}blog`} className={linkCls} onMouseEnter={slideTo}>
{S.nav.blog}
</a>
<Link to="/" className={activeLinkCls} onMouseEnter={slideTo} aria-current="page">
{S.nav.docs}
</Link>
</nav>
<div className="ml-auto flex items-center gap-1">
<LangToggle />
<ThemeToggle />
<a
+7 -1
View File
@@ -6,7 +6,13 @@ export const en: Strings = {
docsBadge: "Docs",
nav: {
home: "Website",
highlights: "Highlights",
quickstart: "Quick start",
benchmark: "Benchmark",
contract: "CONTRACT.md",
features: "Features",
blog: "Blog",
docs: "Docs",
github: "GitHub",
openMenu: "Open navigation",
closeMenu: "Close navigation",
+9 -1
View File
@@ -10,7 +10,15 @@ export const zh = {
docsBadge: "Docs",
nav: {
home: "产品主页",
// Landing-parity labels: the top bar mirrors the landing page's nav exactly
// (same items, same order) so the two sites link into each other seamlessly.
highlights: "特色",
quickstart: "快速开始",
benchmark: "评测",
contract: "CONTRACT.md",
features: "功能",
blog: "博客",
docs: "文档",
github: "GitHub",
openMenu: "打开目录",
closeMenu: "关闭目录",
@@ -0,0 +1,84 @@
---
title: "Partnering with the AMD Developer Program: $50 in free Fireworks credits, wired into PenguinHarness"
date: 2026-07-20
category: news
excerpt: We are delighted to partner with the AMD AI Developer Program to bring you free Fireworks coupon codes — apply for the $50 credits with this guide, then set them up in PenguinHarness in three steps.
---
We are delighted to partner with the **AMD AI Developer Program** to bring you free Fireworks redemption codes: join the program, pass the review, and you receive a Coupon Code redeemable for **$50 in Fireworks AI Credits**. PenguinHarness ships a built-in Fireworks AI gateway group — OpenAI protocol, preset base URL, five preset models — so the credits are usable the moment they land.
> Page content, credit amounts, and expiration dates are subject to change. Please refer to the application page and approval email for the latest information.
## Step 1: Join the AMD AI Developer Program
Visit the official page: <https://developer.amd.com/ai-developer-program/>. In the **Join the AMD AI Developer Program** section:
- No AMD ADP account yet: click **Create Account** and enter your personal information;
- Already have one: click **Log In** and follow the on-screen instructions.
![The AMD AI Developer Program join page](https://github.com/user-attachments/assets/47a3055b-9a95-40a1-80c6-e3bac7a9ac49)
## Step 2: Request Cloud Credits under Member Perks
After registering and signing in:
1. Click **Member Perks** in the top navigation bar;
2. Find **Cloud Credit Options**;
3. Click **Request Cloud Credits** at the bottom of the page.
![Cloud Credit Options under Member Perks](https://github.com/user-attachments/assets/773f8cf1-d72f-4aa3-83b8-f31fc7c9ed9e)
## Step 3: Complete the application form
Fill in the requested personal information:
- Under **Product Needed**, select **Fireworks AI**;
- In the **Profile** section, provide at least one public profile for account verification: LinkedIn, GitHub, a portfolio, a company or school profile, WeChat, WeCom, and similar all work.
![Application form: select Fireworks AI under Product Needed](https://github.com/user-attachments/assets/12a49136-0956-4f29-9d47-f0e473615075)
Complete the other required fields marked with `*`, double-check your email, identity, product selection, and profile link, then submit.
## Step 4: Wait for the review
AMD verifies your account and application, usually within **2–3 business days** — actual time varies with application volume, information completeness, and holidays.
## Step 5: Receive and redeem the coupon code
Once approved, AMD emails a unique **Coupon Code** redeemable for **$50 in Fireworks AI Credits** to your application address. Keep it secure — do not share it publicly, forward it, or commit it to a repository.
![The Coupon Code in the approval email](https://github.com/user-attachments/assets/c53129c7-2c87-4510-a4e3-31f52598dc24)
Redeem it and create an API key:
1. Open <https://fireworks.ai/> and sign in;
2. Click **Redeem Promo** and enter the Coupon Code from the email to redeem the $50 credits;
3. Click **Create API Key** to generate your Fireworks API key.
![Redeeming and creating an API key in the Fireworks console](https://github.com/user-attachments/assets/051a1e69-db7f-4867-b899-89981df15142)
## Set it up in PenguinHarness
With the API key in hand, three steps:
**1. Install and launch**
```bash
curl -fsSL https://penguin.ooo/install.sh | sh
penguin web # opens http://127.0.0.1:7364 (first login: admin / penguin-2026)
```
**2. Configure a Fireworks model**
Open the Models page and find the **Fireworks AI** group, then use its bulk key button to paste the API key you just created. The group presets five models — GLM 5.2, Kimi K2.7 Code, DeepSeek V4 Pro, MiniMax M3, and DeepSeek V4 Flash — with base URLs and pricing pre-filled; set any of them as the default. You can also hit the group's speed-test button to measure real TTFT and TPS before choosing.
**3. Start working**
Head back to Chat and hand the Agent its first task — e.g. "Analyze data.csv and summarize quarterly sales".
## References
- [AMD AI Developer Program](https://developer.amd.com/ai-developer-program/)
- [Official AMD Cloud Credits application video tutorial](https://www.youtube.com/watch?v=masSW53JkTY)
- Application steps and screenshots are adapted from WhatGhost's guides ([中文](https://github.com/WhatGhost/whatghost_Notebooks/blob/main/other/AMD_AI_Developer_Program_Credits_%E7%94%B3%E8%AF%B7%E6%8C%87%E5%8D%97.md) / [English](https://github.com/WhatGhost/whatghost_Notebooks/blob/main/other/AMD_AI_Developer_Program_Credits_Application_Guide_EN.md)) — thanks to the original author
- [PenguinHarness model configuration docs](https://penguin.ooo/docs/models)
@@ -0,0 +1,84 @@
---
title: 携手 AMD 开发者计划:免费领取 $50 Fireworks 额度,直连 PenguinHarness
date: 2026-07-20
category: news
excerpt: 我们很高兴与 AMD AI Developer Program 合作,为大家带来 Fireworks 的免费兑换码——按本文申请 $50 Credits,再在 PenguinHarness 里三步用起来。
---
我们很高兴与 **AMD AI Developer Program**(AMD 开发者计划)合作,为大家带来 Fireworks 的免费兑换码:加入计划并通过审核,即可获得可兑换 **$50 Fireworks AI Credits** 的 Coupon Code。PenguinHarness 内置 Fireworks AI 网关分组——OpenAI 协议、预置 base URL 与五个模型,额度到手即刻可用。
> 页面内容、Credits 金额与有效期可能调整,请以申请时的页面及审批邮件为准。
## 第一步:加入 AMD AI Developer Program
访问官方入口:<https://developer.amd.com/ai-developer-program/>,在 **Join the AMD AI Developer Program** 区域:
- 没有 AMD ADP 账户:点击 **Create Account**,填写个人信息创建账户;
- 已有账户:点击 **Log In**,按页面提示登录。
![AMD AI Developer Program 加入页面](https://github.com/user-attachments/assets/47a3055b-9a95-40a1-80c6-e3bac7a9ac49)
## 第二步:在 Member Perks 申请 Cloud Credits
注册并登录后:
1. 点击顶部导航栏中的 **Member Perks**;
2. 找到 **Cloud Credit Options**;
3. 点击底部的 **Request Cloud Credits**。
![Member Perks 中的 Cloud Credit Options](https://github.com/user-attachments/assets/773f8cf1-d72f-4aa3-83b8-f31fc7c9ed9e)
## 第三步:填写申请表
进入申请表后填写个人信息:
- **Product Needed** 处选择 **Fireworks AI**;
- **Profile** 处提供至少一个公开资料用于账户验证:LinkedIn、GitHub、Portfolio、公司 / 学校主页、WeChat、WeCom 等均可。
![申请表:Product Needed 选择 Fireworks AI](https://github.com/user-attachments/assets/12a49136-0956-4f29-9d47-f0e473615075)
填完页面中其他带 `*` 的必填项,检查邮箱、身份、产品选项与公开资料链接无误后提交。
## 第四步:等待审核
AMD 会验证账户与申请资料,通常需要 **2–3 个工作日**;实际时间可能因申请量、资料完整度或节假日而变化。
## 第五步:收到兑换码并兑换
审核通过后,AMD 会向申请邮箱发送包含唯一 **Coupon Code** 的邮件,可兑换 **$50 Fireworks AI Credits**。请妥善保存,不要公开、转发或提交到代码仓库。
![审批邮件中的 Coupon Code](https://github.com/user-attachments/assets/c53129c7-2c87-4510-a4e3-31f52598dc24)
兑换与创建 API key:
1. 打开 <https://fireworks.ai/> 并登录;
2. 点击 **Redeem Promo**,输入邮件中的 Coupon Code,兑换 $50 Credits;
3. 点击 **Create API Key**,生成 Fireworks API key。
![在 Fireworks 控制台兑换并创建 API key](https://github.com/user-attachments/assets/051a1e69-db7f-4867-b899-89981df15142)
## 在 PenguinHarness 中用起来
拿到 API key 后,三步接入:
**1. 安装并启动**
```bash
curl -fsSL https://penguin.ooo/install.sh | sh
penguin web # 打开 http://127.0.0.1:7364(首次登录:admin / penguin-2026)
```
**2. 配置 Fireworks 模型**
进入「模型仓库」页,找到 **Fireworks AI** 分组,点击「统一配置 key」粘贴刚创建的 API key。分组预置了五个模型——GLM 5.2、Kimi K2.7 Code、DeepSeek V4 Pro、MiniMax M3、DeepSeek V4 Flash——base URL 与价格已填好,任选一个设为默认即可;也可以点组头的「测速」,实测各模型的 TTFT 与 TPS 再决定。
**3. 开始使用**
回到对话页,把第一个任务交给 Agent——例如「分析 data.csv,输出各季度销售额汇总」。
## 参考链接
- [AMD AI Developer Program](https://developer.amd.com/ai-developer-program/)
- [AMD 官方 Cloud Credits 申请视频演示](https://www.youtube.com/watch?v=masSW53JkTY)
- 申请步骤与截图整理自 WhatGhost 的申请指南([中文](https://github.com/WhatGhost/whatghost_Notebooks/blob/main/other/AMD_AI_Developer_Program_Credits_%E7%94%B3%E8%AF%B7%E6%8C%87%E5%8D%97.md) / [English](https://github.com/WhatGhost/whatghost_Notebooks/blob/main/other/AMD_AI_Developer_Program_Credits_Application_Guide_EN.md)),感谢原作者
- [PenguinHarness 模型配置文档](https://penguin.ooo/docs/models)
@@ -2,70 +2,114 @@
title: "Introducing PenguinHarness: agents that build agents"
date: 2026-07-17
category: news
excerpt: The first open-source harness with recursive self-improvement is here — lightweight, efficient and secure infrastructure covering everything from automatic agent construction to continuous self-evolution.
excerpt: We proved agents can self-evolve in our GDPevo Benchmark — now we are bringing that capability to everyone. The first open-source harness with recursive self-improvement covers everything from one-sentence agent construction to continuous self-evolution.
---
Today we are releasing **PenguinHarness** — an open-source harness built for constructing and evolving agents. Its purpose fits in one line:
Today we are releasing **PenguinHarness** — an open-source harness built for constructing and evolving agents: a zero-code Harness CLI and Web UI, connected to 1000+ models. The story it tells fits in one line:
> Efficient Self-Improving Harness for Everyone.
> With LangChain, you build agents by hand — at 1× speed. With PenguinHarness, agents build agents — at 100×.
## From GDPevo to PenguinHarness: why we built this
Before PenguinHarness, our team published the [GDPevo Benchmark](https://prism-shadow.github.io/GDPevo/). In GDPevo we systematically verified one thing: **agents can self-evolve** — an Agent can score its own performance, find where the points were lost, rewrite its own prompts and Skills, and climb version after version.
With the capability proven, the question became: how does everyone get to use it? Self-evolution should not stay a curve in a paper — it should be infrastructure that works out of the box on every developer's desk. **Bringing an efficient self-improving harness to everyone is why we built PenguinHarness** — and it is right there in the name: Efficient Self-Improving Harness for Everyone.
## Why PenguinHarness
Over the past year the way agent applications are built has been converging fast: what really decides quality is not a heavyweight framework but a simple, reliable, observable harness. PenguinHarness is rebuilt from the ground up — no dependency on any agent framework, a fully open-source self-developed kernel — and it brings three things to the open-source world first:
Three reasons, in deliberate order — from task quality, to how agents get built, to how they keep improving.
- **Simplest Is the Best**: a deliberately minimal toolset over clean low-level interfaces — fewer tool calls, fewer Tokens, complex tasks done efficiently.
- **Harness for Building Agents**: with the PenguinHarness SDK, an Agent builds complete Agent applications for you, autonomously, from scratch.
- **Harness for Recursive Self-Improvement**: with PenguinHarness Skills, an Agent evaluates and optimizes itself, improving recursively over time.
### 1. Better on complex tasks, at lower cost
For the latter two, PenguinHarness is the first open-source implementation in the industry.
A deliberately minimal toolset over clean low-level interfaces: fewer tool calls, fewer Tokens, deeply tuned for open models like DeepSeek. Each harness runs the model it is normally paired with — the comparison is between the products as people actually use them — head-to-head on two suites:
## Same model, equal or better quality, lower cost
![Benchmark: PenguinHarness leads the data-analysis suite and ties OpenAI Codex on coding, at a small fraction of both rivals' cost](/blog-assets/benchmark-light.svg)
All runs use the same DeepSeek V4 Pro model, head-to-head against Claude Code and OpenAI Codex on two suites (per-run means below).
Complex data analysis (15 tasks, single run; PenguinHarness and Codex at thinking xhigh, Claude Code at max):
Complex data analysis (15 tasks, single run):
| Framework | Model | Accuracy (%) | Tokens (M) | Cost ($) |
| -------------- | --------------- | -----------: | ---------: | -------: |
| PenguinHarness | DeepSeek V4 Pro | 66.67 | 18.04 | 0.55 |
| Claude Code | Claude Opus 4.8 | 53.33 | 22.20 | 38.48 |
| OpenAI Codex | GPT-5.5 | 53.33 | 13.72 | 19.41 |
| Framework | Model | Accuracy (%) | Tokens (M) | Cost ($) |
| --- | --- | ---: | ---: | ---: |
| PenguinHarness | DeepSeek V4 Pro | 66.7 | 18.04 | 0.552 |
| Claude Code | DeepSeek V4 Pro | 66.7 | 21.17 | 0.641 |
| OpenAI Codex | DeepSeek V4 Pro | 46.7 | 13.36 | 0.427 |
Coding tasks (40 tasks × 2 runs; accuracy is over all 80 outcomes):
Coding tasks (40 tasks × 2 runs averaged, thinking high, 30 min per-case timeout, CNY pricing converted at $1 = ¥7):
| Framework | Model | Accuracy (%) | Tokens (M) | Cost ($) |
| -------------- | --------------- | -----------: | ---------: | -------: |
| PenguinHarness | DeepSeek V4 Pro | 71.25 | 200.00 | 3.81 |
| Claude Code | Claude Opus 4.8 | 86.25 | 151.61 | 146.97 |
| OpenAI Codex | GPT-5.5 | 71.25 | 251.20 | 220.08 |
| Framework | Model | Accuracy (%) | Tokens (M) | Cost ($) |
| --- | --- | ---: | ---: | ---: |
| PenguinHarness | DeepSeek V4 Pro | 50.00 | 2.10 | 0.041 |
| Claude Code | DeepSeek V4 Pro | 48.75 | 2.00 | 0.048 |
| OpenAI Codex | DeepSeek V4 Pro | 42.50 | 2.65 | 0.043 |
Tokens and cost are suite totals, not per-run means. On data analysis we take the highest accuracy of the three — 66.67% against 53.33% for both — while spending 1/35 of what Codex spent and 1/70 of Claude Code's. On coding we tie Codex at 71.25% and trail Claude Code's 86.25%, but the whole suite cost us $3.81 against their $220.08 and $146.97: comparable work, one to two orders of magnitude apart on the bill.
On the data-analysis suite PenguinHarness ties Claude Code on accuracy and clearly beats OpenAI Codex while using 14.8% fewer Tokens at 13.8% lower cost; on the coding suite it scores highest of the three at the lowest per-run cost.
### 2. One sentence, and an Agent builds your Agent app
Type one sentence, and an Agent builds the complete Agent application for you — scaffold, code, and run instructions, end to end:
```text
Collect the docs from https://github.com/ericbuess/claude-code-docs and build a RAG app that answers Claude Code questions as a configuration expert, citing its sources.
```
And this is the finished product — a docs expert with retrieval, cited sources that link to the original files, and example questions built in:
![The generated RAG app: a Claude Code docs expert answering with cited, clickable sources and example questions](/blog-assets/rag-app-en-light.webp)
**And generating this entire RAG app burned just $0.02 (¥0.2) of tokens — on DeepSeek V4 Pro.**
### 3. Self-evolution: it gets stronger with use
With PenguinHarness Skills, an Agent evaluates and optimizes itself: the Optimizer orchestrates multiple Evaluators to score in parallel, uses the scores and run traces to find where points were lost, and upgrades the Agent from version N to N+1 — with a snapshot before every round, and every request replayable in the Trace view. A self-evolution demo video is coming soon.
## Evolution within bounds, security first
The biggest worry about self-improvement is losing control. PenguinHarness answers with a contract — CONTRACT.md:
The biggest worry about self-evolution is losing control. PenguinHarness answers with a contract (CONTRACT.md):
- Evolution is strictly confined to Workspace and Skills; the harness core security boundary is never modified.
- Tool calls run only after approval, and every approval is audited.
- Risky changes snapshot first — every step of evolution can be rolled back.
- Fully open source and locally deployable: data never leaves your machine, meeting enterprise security requirements.
- Evolution is strictly confined to Workspace and Skills — the harness core security boundary is never modified;
- Tool calls require approval first, and every approval leaves an audit record;
- Risky changes are preceded by version snapshots, so any round of evolution can be rolled back;
- Fully open source and locally deployed — data never leaves your machine, meeting enterprise data-security requirements.
## Get started now
## Supported models
Install with one command (Linux / macOS, x64 / arm64, bundled Node runtime):
| Model | Providers |
| ---------------- | -------------------------------------------------------------------------------- |
| DeepSeek V4 | DeepSeek, OpenRouter, Fireworks AI, SiliconFlow, Qwen Token Plan |
| Kimi K3 | Moonshot AI, OpenRouter, Qwen Pay-As-You-Go |
| GLM 5.2 | Z.AI, OpenRouter, Fireworks AI, SiliconFlow, Qwen Token Plan, Qwen Pay-As-You-Go |
| Hunyuan 3 | OpenRouter |
| Qwen 3.8 Max | Qwen Token Plan (preview) |
| GPT 5.5 | OpenAI, OpenRouter |
| Gemini 3.5 Flash | Google Gemini, OpenRouter |
| Claude Opus 4.8 | Anthropic, OpenRouter |
Any OpenAI-protocol endpoint is supported: pick a preset above, or point a custom endpoint at any of the 1000+ online and local models.
## How to use it
Install with one command (Linux / macOS, x64 / arm64, bundled Node runtime), then launch the Web UI:
```bash
curl -fsSL https://github.com/Prism-Shadow/penguin-harness/releases/latest/download/install.sh | sh
curl -fsSL https://penguin.ooo/install.sh | sh
penguin web # opens http://127.0.0.1:7364 (first login: admin / penguin-2026)
```
Configure a model (DeepSeek as an example), run your first task, or open the desktop-grade web interface with `penguin web`:
Open the Models page, paste an API key under the DeepSeek or OpenRouter group and set it as default; then head back to Chat and hand the Agent its first task — e.g. "Analyze data.csv and summarize quarterly sales".
```bash
penguin config model add --model-id deepseek-v4-pro --api-key sk-your-key --set-default
penguin run --approve allow-all --message "Analyze data.csv and summarize quarterly sales"
penguin web
```
## What's next
PenguinHarness supports 1000+ online and local models and multi-agent collaborative evolution, and runs on as little as a single CPU. Through continuous evolution it makes complex AI development ever simpler — a more efficient, more reliable, lower-hallucination and lower-cost Agent productivity engine.
- Public release of the benchmark suite;
- A desktop app;
- Windows support;
- More to come.
Follow us on [GitHub](https://github.com/Prism-Shadow/penguin-harness) and open your first issue.
## Join the community and build with us
A self-improving harness needs a community that improves with it. Come discuss, request, and contribute — your first Issue is the best way to start:
- [Discord](https://discord.gg/eFHKqqcU3D): chat with us and other developers in real time;
- [X (Twitter)](https://x.com/code_hiyouga): follow the latest updates;
- [WeChat group](https://github.com/Prism-Shadow/penguin-harness-community/blob/main/wechat/group.jpg): Chinese community discussions;
- [GitHub](https://github.com/Prism-Shadow/penguin-harness): stars, Issues, and PRs all welcome.
Self-evolving agent infrastructure, for everyone — starting today.
@@ -2,44 +2,64 @@
title: PenguinHarness 正式发布:让 Agent 为你构建 Agent
date: 2026-07-17
category: news
excerpt: 首个支持递归自我进化的开源 Harness 正式发布——以轻量、高效、安全的方式,提供从 Agent 自动构建到持续自我进化的完整基础设施。
excerpt: 我们在 GDPevo Benchmark 中验证了 Agent 自我进化的能力,现在把它带给所有人——首个支持递归自我进化的开源 Harness 正式发布,从一句话构建 Agent 到持续自我进化,一套基础设施全部覆盖。
---
今天,我们正式发布 **PenguinHarness**——一个为构建与进化 Agent 而生的开源 Harness。它的主旨只有一句话:
今天,我们正式发布 **PenguinHarness**——一个为构建与进化 Agent 而生的开源 Harness:零代码的 Harness CLI 与 Web UI,连接 1000+ 模型。它要讲的故事只有一句话:
> Efficient Self-Improving Harness for Everyone.
> 使用 LangChain,以 1 倍速度人工构建 Agent;使用 PenguinHarness,以 100 倍速度用 Agent 构建 Agent。
## 从 GDPevo 到 PenguinHarness:我们的初心
在 PenguinHarness 之前,我们团队发布了 [GDPevo Benchmark](https://prism-shadow.github.io/GDPevo/)。在 GDPevo 中,我们系统性地验证了一件事:**Agent 可以自我进化**——让 Agent 评估自己的表现、定位失分原因、改写自己的提示词与技能,分数随版本一路上升。
能力验证了,问题就变成了:怎么让每个人都用上它?自我进化不该只是论文里的曲线,而应该是每个开发者桌面上开箱即用的基础设施。**让每个人都能使用 Efficient Self-Improving Harness,这就是我们构建 PenguinHarness 的初心**——它也因此得名:Efficient Self-Improving Harness for Everyone.
## 为什么是 PenguinHarness
过去一年里,Agent 应用的开发范式在快速收敛:真正决定效果的不是庞大的框架,而是一个简洁、可靠、可观测的 Harness。PenguinHarness 从底层重构,不依赖任何 Agent 框架,开源自研 Harness 内核,并率先把三件事带入开源世界:
三个递进的理由——从任务效果,到构建方式,再到进化能力。
- **Simplest Is the Best**:坚持最小化工具集与简洁的底层接口,以更少的工具调用与 Token 消耗,高效完成复杂任务。
- **Harness for Building Agents**:通过 PenguinHarness SDK,让 Agent 从零自主完成 Agent 应用的构建。
- **Harness for Recursive Self-Improvement**:通过 PenguinHarness Skills,Agent 以自我评估与自我优化实现递归式自我提升。
### 1. 复杂任务表现更好,成本更低
后两项能力,PenguinHarness 是业内首个开源实现。
刻意精简的工具集配合干净的底层接口:更少的工具调用、更少的 Token,对 DeepSeek 等开放模型深度适配。每个产品搭配它常用的模型——比的是大家实际会怎么用——在两套题库上正面对比:
## 同一模型,同级效果,更低消耗
![Benchmark:PenguinHarness 在数据分析题库准确率最高、编程题库与 OpenAI Codex 持平,成本仅为两者的零头](/blog-assets/benchmark-light.svg)
全部使用同一 DeepSeek V4 Pro 模型,与 Claude Code、OpenAI Codex 在两套题库上正面对比(表中为单次运行均值)。
复杂数据分析(15 题,单次运行;PenguinHarness 与 Codex 为 thinking xhigh,Claude Code 为 max):
复杂数据分析(15 题,单次运行):
| 实验框架 | 模型名称 | 准确率(%) | Token 用量(M) | 成本($) |
| -------------- | --------------- | ----------: | --------------: | --------: |
| PenguinHarness | DeepSeek V4 Pro | 66.67 | 18.04 | 0.55 |
| Claude Code | Claude Opus 4.8 | 53.33 | 22.20 | 38.48 |
| OpenAI Codex | GPT-5.5 | 53.33 | 13.72 | 19.41 |
| 实验框架 | 模型名称 | 准确率(%) | Token 用量(M) | 成本($) |
| --- | --- | ---: | ---: | ---: |
| PenguinHarness | DeepSeek V4 Pro | 66.7 | 18.04 | 0.552 |
| Claude Code | DeepSeek V4 Pro | 66.7 | 21.17 | 0.641 |
| OpenAI Codex | DeepSeek V4 Pro | 46.7 | 13.36 | 0.427 |
代码任务(40 题 × 2 次运行,准确率取全部 80 次结果):
代码任务(40 题 × 2 runs 取均值,thinking high、单题 30 分钟超时,人民币计价按 $1 = ¥7 折算):
| 实验框架 | 模型名称 | 准确率(%) | Token 用量(M) | 成本($) |
| -------------- | --------------- | ----------: | --------------: | --------: |
| PenguinHarness | DeepSeek V4 Pro | 71.25 | 200.00 | 3.81 |
| Claude Code | Claude Opus 4.8 | 86.25 | 151.61 | 146.97 |
| OpenAI Codex | GPT-5.5 | 71.25 | 251.20 | 220.08 |
| 实验框架 | 模型名称 | 准确率(%) | Token 用量(M) | 成本($) |
| --- | --- | ---: | ---: | ---: |
| PenguinHarness | DeepSeek V4 Pro | 50.00 | 2.10 | 0.041 |
| Claude Code | DeepSeek V4 Pro | 48.75 | 2.00 | 0.048 |
| OpenAI Codex | DeepSeek V4 Pro | 42.50 | 2.65 | 0.043 |
Token 与成本均为全套题目合计,不是单次均值。数据分析套件我们准确率最高——66.67% 对另两者的 53.33%——花的钱是 Codex 的 1/35、Claude Code 的 1/70;代码套件与 Codex 同为 71.25%、低于 Claude Code 的 86.25%,但整套题目只花了 $3.81,对方分别是 $220.08 与 $146.97:活干得差不多,账单差出一到两个数量级。
数据分析套件与 Claude Code 准确率持平、显著超过 OpenAI Codex,同时 Token 消耗少 14.8%、成本低 13.8%;代码套件三者中准确率最高、单次成本最低。
### 2. 一句话,让 Agent 构建 Agent 应用
输入一句话,Agent 为你构建完整的 Agent 应用——脚手架、代码、运行说明,一步到位:
```text
收集 https://github.com/ericbuess/claude-code-docs 的文档,做一个化身 Claude Code 配置专家、回答带来源引用的 RAG 问答应用。
```
这是做出来的成品——一个文档专家:检索增强、引用可点击直达原文、内置示例问题:
![生成的 RAG 应用成品:Claude Code 配置专家,回答带可点击的来源引用与示例问题](/blog-assets/rag-app-zh-light.webp)
**而生成整个 RAG 应用,仅消耗了 0.2 元($0.02)的 token——使用 DeepSeek V4 Pro 模型。**
### 3. 自进化,越用越强
借助 PenguinHarness 技能库,Agent 自己评估、自己优化:Optimizer 组织多个 Evaluator 并行打分,依据分数与运行轨迹定位失分原因,把 Agent 从版本 N 优化到版本 N+1——每轮之前自动快照,每个请求都可在轨迹观测中回放。自进化演示视频即将上线。
## 进化有界,安全先行
@@ -50,22 +70,46 @@ excerpt: 首个支持递归自我进化的开源 Harness 正式发布——以
- 风险修改之前先留版本快照,任何一次进化都可回退;
- 完全开源、本地部署,数据不出域,满足企业级数据安全。
## 现在就可以开始
## 支持的模型
一行命令安装(Linux / macOS,x64 / arm64,内嵌 Node 运行时):
| 模型 | 可用供应商 |
| ---------------- | -------------------------------------------------------------------------------- |
| DeepSeek V4 | DeepSeek, OpenRouter, Fireworks AI, SiliconFlow, Qwen Token Plan |
| Kimi K3 | Moonshot AI, OpenRouter, Qwen Pay-As-You-Go |
| GLM 5.2 | Z.AI, OpenRouter, Fireworks AI, SiliconFlow, Qwen Token Plan, Qwen Pay-As-You-Go |
| Hunyuan 3 | OpenRouter |
| Qwen 3.8 Max | Qwen Token Plan(预览) |
| GPT 5.5 | OpenAI, OpenRouter |
| Gemini 3.5 Flash | Google Gemini, OpenRouter |
| Claude Opus 4.8 | Anthropic, OpenRouter |
只要是 OpenAI 协议的端点都可以接入:从上表选择预置,或用自定义端点连接 1000+ 在线与本地模型。
## 如何使用
一行命令安装(Linux / macOS,x64 / arm64,内嵌 Node 运行时),然后启动 Web 界面:
```bash
curl -fsSL https://github.com/Prism-Shadow/penguin-harness/releases/latest/download/install.sh | sh
curl -fsSL https://penguin.ooo/install.sh | sh
penguin web # 打开 http://127.0.0.1:7364(首次登录:admin / penguin-2026)
```
配置模型(以 DeepSeek 为例)后即可运行第一个任务,或用 `penguin web` 打开桌面级 Web 界面:
进入「模型仓库」页,在 DeepSeek 或 OpenRouter 分组里粘贴 API key 并设为默认;回到对话页,把第一个任务交给 Agent——例如「分析 data.csv,输出各季度销售额汇总」。
```bash
penguin config model add --model-id deepseek-v4-pro --api-key sk-your-key --set-default
penguin run --approve allow-all --message "分析 data.csv,输出各季度销售额汇总"
penguin web
```
## 发展计划
PenguinHarness 支持 1000 多种在线与本地模型、多智能体协作进化,最低单 CPU 即可运行。通过不断进化,它会让复杂的 AI 开发越来越简单——为你提供更高效、更可靠、更低幻觉、更低成本的 Agent 生产力引擎。
- Benchmark 套件正式发布;
- 推出桌面端(Desktop)应用;
- 支持 Windows 系统;
- 更多规划,敬请期待。
欢迎在 [GitHub](https://github.com/Prism-Shadow/penguin-harness) 上关注我们,提出你的第一个 Issue。
## 加入社区,一起共建
自我进化的 Harness,也需要一个不断进化的社区。欢迎加入讨论、提出需求、贡献代码——你的第一个 Issue 就是最好的开始:
- [Discord](https://discord.gg/eFHKqqcU3D):与我们和其他开发者实时交流;
- [X(Twitter)](https://x.com/code_hiyouga):关注最新动态;
- [微信群](https://github.com/Prism-Shadow/penguin-harness-community/blob/main/wechat/group.jpg):中文社区讨论;
- [GitHub](https://github.com/Prism-Shadow/penguin-harness):Star、Issue 与 PR 都欢迎。
让每个人都用上会自我进化的 Agent 基础设施——从今天开始。
@@ -0,0 +1,48 @@
<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 920 368" width="920" height="368" role="img" aria-label="Benchmark: PenguinHarness vs Claude Code vs OpenAI Codex on two suites — comparable accuracy at a small fraction of the cost">
<text x="148" y="36" font-size="12" font-weight="400" fill="#898781" text-anchor="start" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">Accuracy · suite total · higher is better</text>
<text x="596" y="36" font-size="12" font-weight="400" fill="#898781" text-anchor="start" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">Total cost (USD) · lower is better</text>
<text x="24" y="68" font-size="13" font-weight="600" fill="#1f2328" text-anchor="start" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">Data analysis — 15 tasks, single run</text>
<rect x="147" y="78" width="1" height="92" fill="#c3c2b7"/>
<text x="136" y="94.5" font-size="12" font-weight="600" fill="#1f2328" text-anchor="end" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">PenguinHarness</text>
<path d="M149 82 h172.0 a4 4 0 0 1 4 4 v8 a4 4 0 0 1 -4 4 h-172.0 z" fill="#2a78d6"/>
<text x="337" y="94.5" font-size="12" font-weight="600" fill="#1f2328" text-anchor="start" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">66.67%</text>
<text x="136" y="124.5" font-size="12" font-weight="400" fill="#52514e" text-anchor="end" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">Claude Code</text>
<path d="M149 112 h136.8 a4 4 0 0 1 4 4 v8 a4 4 0 0 1 -4 4 h-136.8 z" fill="#898781"/>
<text x="301.7841607919604" y="124.5" font-size="12" font-weight="400" fill="#1f2328" text-anchor="start" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">53.33%</text>
<text x="136" y="154.5" font-size="12" font-weight="400" fill="#52514e" text-anchor="end" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">OpenAI Codex</text>
<path d="M149 142 h136.8 a4 4 0 0 1 4 4 v8 a4 4 0 0 1 -4 4 h-136.8 z" fill="#898781"/>
<text x="301.7841607919604" y="154.5" font-size="12" font-weight="400" fill="#1f2328" text-anchor="start" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">53.33%</text>
<rect x="595" y="78" width="1" height="92" fill="#c3c2b7"/>
<text x="584" y="94.5" font-size="12" font-weight="600" fill="#1f2328" text-anchor="end" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">PenguinHarness</text>
<path d="M597 82 h10.0 a4 4 0 0 1 4 4 v8 a4 4 0 0 1 -4 4 h-10.0 z" fill="#2a78d6"/>
<text x="623" y="94.5" font-size="12" font-weight="600" fill="#1f2328" text-anchor="start" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">$0.55</text>
<text x="584" y="124.5" font-size="12" font-weight="400" fill="#52514e" text-anchor="end" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">Claude Code</text>
<path d="M597 112 h172.0 a4 4 0 0 1 4 4 v8 a4 4 0 0 1 -4 4 h-172.0 z" fill="#898781"/>
<text x="785" y="124.5" font-size="12" font-weight="400" fill="#1f2328" text-anchor="start" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">$38.48</text>
<text x="584" y="154.5" font-size="12" font-weight="400" fill="#52514e" text-anchor="end" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">OpenAI Codex</text>
<path d="M597 142 h84.8 a4 4 0 0 1 4 4 v8 a4 4 0 0 1 -4 4 h-84.8 z" fill="#898781"/>
<text x="697.7945961503353" y="154.5" font-size="12" font-weight="400" fill="#1f2328" text-anchor="start" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">$19.41</text>
<text x="24" y="210" font-size="13" font-weight="600" fill="#1f2328" text-anchor="start" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">Coding — 40 tasks × 2 runs</text>
<rect x="147" y="220" width="1" height="92" fill="#c3c2b7"/>
<text x="136" y="236.5" font-size="12" font-weight="600" fill="#1f2328" text-anchor="end" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">PenguinHarness</text>
<path d="M149 224 h141.4 a4 4 0 0 1 4 4 v8 a4 4 0 0 1 -4 4 h-141.4 z" fill="#2a78d6"/>
<text x="306.3913043478261" y="236.5" font-size="12" font-weight="600" fill="#1f2328" text-anchor="start" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">71.25%</text>
<text x="136" y="266.5" font-size="12" font-weight="400" fill="#52514e" text-anchor="end" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">Claude Code</text>
<path d="M149 254 h172.0 a4 4 0 0 1 4 4 v8 a4 4 0 0 1 -4 4 h-172.0 z" fill="#898781"/>
<text x="337" y="266.5" font-size="12" font-weight="400" fill="#1f2328" text-anchor="start" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">86.25%</text>
<text x="136" y="296.5" font-size="12" font-weight="400" fill="#52514e" text-anchor="end" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">OpenAI Codex</text>
<path d="M149 284 h141.4 a4 4 0 0 1 4 4 v8 a4 4 0 0 1 -4 4 h-141.4 z" fill="#898781"/>
<text x="306.3913043478261" y="296.5" font-size="12" font-weight="400" fill="#1f2328" text-anchor="start" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">71.25%</text>
<rect x="595" y="220" width="1" height="92" fill="#c3c2b7"/>
<text x="584" y="236.5" font-size="12" font-weight="600" fill="#1f2328" text-anchor="end" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">PenguinHarness</text>
<path d="M597 224 h10.0 a4 4 0 0 1 4 4 v8 a4 4 0 0 1 -4 4 h-10.0 z" fill="#2a78d6"/>
<text x="623" y="236.5" font-size="12" font-weight="600" fill="#1f2328" text-anchor="start" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">$3.81</text>
<text x="584" y="266.5" font-size="12" font-weight="400" fill="#52514e" text-anchor="end" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">Claude Code</text>
<path d="M597 254 h113.5 a4 4 0 0 1 4 4 v8 a4 4 0 0 1 -4 4 h-113.5 z" fill="#898781"/>
<text x="726.5332606324973" y="266.5" font-size="12" font-weight="400" fill="#1f2328" text-anchor="start" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">$146.97</text>
<text x="584" y="296.5" font-size="12" font-weight="400" fill="#52514e" text-anchor="end" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">OpenAI Codex</text>
<path d="M597 284 h172.0 a4 4 0 0 1 4 4 v8 a4 4 0 0 1 -4 4 h-172.0 z" fill="#898781"/>
<text x="785" y="296.5" font-size="12" font-weight="400" fill="#1f2328" text-anchor="start" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">$220.08</text>
<text x="24" y="338" font-size="12" font-weight="400" fill="#898781" text-anchor="start" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">Each harness runs the model it is normally paired with: PenguinHarness on DeepSeek V4 Pro, Claude Code on Claude Opus 4.8,</text>
<text x="24" y="356" font-size="12" font-weight="400" fill="#898781" text-anchor="start" font-family="system-ui, -apple-system, 'Segoe UI', sans-serif">OpenAI Codex on GPT-5.5. Accuracy, Tokens and cost are suite totals at official pricing.</text>
</svg>

After

Width:  |  Height:  |  Size: 6.9 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 76 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 71 KiB

+21
View File
@@ -0,0 +1,21 @@
#!/bin/sh
# https://penguin.ooo/install.sh - PenguinHarness installer entry point.
#
# GitHub Pages cannot serve HTTP redirects, so this thin forwarder IS the
# stable install URL: it fetches the real installer attached to the latest
# GitHub release and runs it, forwarding every argument it was given. Usage:
#
# curl -fsSL https://penguin.ooo/install.sh | sh
# curl -fsSL https://penguin.ooo/install.sh | sh -s -- --universal
#
set -eu
# Download to a file first, then run it: piping straight into `sh` would execute
# a truncated download line by line, and the real installer removes the old
# bin/lib/web/node before moving the new ones in — a cut connection mid-way
# would leave no install at all.
TMP="$(mktemp)"
trap 'rm -f "$TMP"' EXIT
curl -fsSL "https://github.com/Prism-Shadow/penguin-harness/releases/latest/download/install.sh" -o "$TMP"
rc=0
sh "$TMP" "$@" || rc=$?
exit "$rc"
@@ -0,0 +1,62 @@
/**
* Render the landing Cases tab's "penguin sled game" finished-product shot.
*
* Same pattern as capture-readme-demo.mjs: penguin-game-mockup.html is a static,
* dependency-free mockup of the example game's play screen, captured per language
* (zh / en) and theme (light = polar day / dark = polar night) into
* packages/landing/src/assets as game-<lang>-<theme>.webp.
*
* Prereqs: Playwright's chromium only (no server, no build). Run:
* `node scripts/capture-game-mockup.mjs [--png-dir <dir>]`.
*/
import { mkdirSync, writeFileSync } from "node:fs";
import path from "node:path";
import { fileURLToPath, pathToFileURL } from "node:url";
import { chromium } from "@playwright/test";
const HERE = path.dirname(fileURLToPath(import.meta.url));
const OUT_DIR = path.resolve(HERE, "../src/assets");
const PAGE = pathToFileURL(path.join(HERE, "penguin-game-mockup.html")).href;
const pngDirArg = process.argv.indexOf("--png-dir");
const PNG_DIR = pngDirArg >= 0 ? process.argv[pngDirArg + 1] : null;
mkdirSync(OUT_DIR, { recursive: true });
if (PNG_DIR) mkdirSync(PNG_DIR, { recursive: true });
const browser = await chromium.launch();
const encoderPage = await browser.newPage();
async function saveWebp(pngBuffer, fileName) {
const dataUrl = await encoderPage.evaluate(async (b64) => {
const img = new Image();
img.src = `data:image/png;base64,${b64}`;
await img.decode();
const canvas = document.createElement("canvas");
canvas.width = img.width;
canvas.height = img.height;
canvas.getContext("2d").drawImage(img, 0, 0);
return canvas.toDataURL("image/webp", 0.9);
}, pngBuffer.toString("base64"));
writeFileSync(path.join(OUT_DIR, fileName), Buffer.from(dataUrl.split(",")[1], "base64"));
console.log(`[game] ${fileName}`);
}
for (const lang of ["en", "zh"]) {
for (const theme of ["light", "dark"]) {
const context = await browser.newContext({
viewport: { width: 1280, height: 760 },
deviceScaleFactor: 1.5,
locale: lang === "zh" ? "zh-CN" : "en-US",
});
const page = await context.newPage();
await page.goto(`${PAGE}?lang=${lang}&theme=${theme}`);
await page.waitForTimeout(300);
const png = await page.screenshot();
if (PNG_DIR) writeFileSync(path.join(PNG_DIR, `game-${lang}-${theme}.png`), png);
await saveWebp(png, `game-${lang}-${theme}.webp`);
await context.close();
}
}
await browser.close();
console.log("[game] done");
@@ -0,0 +1,65 @@
/**
* Render the README "build an Agent app in one sentence" finished-product shot.
*
* The image shows the RESULT of the example — a Claude Code docs-expert RAG app —
* rather than the build conversation: rag-app-mockup.html is a static, dependency-free
* mockup of the generated app's UI (per the web-design skill's Penguin visual language),
* captured per language (zh / en) and theme (light / dark) into assets/readme/ at the
* repo root as rag-app-<lang>-<theme>.webp (README.md uses en, README.zh.md uses zh).
*
* Prereqs: Playwright's chromium only (no server, no build). Run:
* `node scripts/capture-readme-demo.mjs [--png-dir <dir>]` (--png-dir also saves
* lossless PNG copies, handy for reviewing the shots).
*/
import { mkdirSync, writeFileSync } from "node:fs";
import path from "node:path";
import { fileURLToPath, pathToFileURL } from "node:url";
import { chromium } from "@playwright/test";
const HERE = path.dirname(fileURLToPath(import.meta.url));
const ROOT = path.resolve(HERE, "../../..");
const OUT_DIR = path.resolve(ROOT, "assets/readme");
const PAGE = pathToFileURL(path.join(HERE, "rag-app-mockup.html")).href;
const pngDirArg = process.argv.indexOf("--png-dir");
const PNG_DIR = pngDirArg >= 0 ? process.argv[pngDirArg + 1] : null;
mkdirSync(OUT_DIR, { recursive: true });
if (PNG_DIR) mkdirSync(PNG_DIR, { recursive: true });
const browser = await chromium.launch();
const encoderPage = await browser.newPage();
async function saveWebp(pngBuffer, fileName) {
const dataUrl = await encoderPage.evaluate(async (b64) => {
const img = new Image();
img.src = `data:image/png;base64,${b64}`;
await img.decode();
const canvas = document.createElement("canvas");
canvas.width = img.width;
canvas.height = img.height;
canvas.getContext("2d").drawImage(img, 0, 0);
return canvas.toDataURL("image/webp", 0.9);
}, pngBuffer.toString("base64"));
writeFileSync(path.join(OUT_DIR, fileName), Buffer.from(dataUrl.split(",")[1], "base64"));
console.log(`[demo] ${fileName}`);
}
for (const lang of ["en", "zh"]) {
for (const theme of ["light", "dark"]) {
const context = await browser.newContext({
viewport: { width: 1280, height: 760 },
deviceScaleFactor: 1.5,
locale: lang === "zh" ? "zh-CN" : "en-US",
});
const page = await context.newPage();
await page.goto(`${PAGE}?lang=${lang}&theme=${theme}`);
await page.waitForTimeout(300);
const png = await page.screenshot();
if (PNG_DIR) writeFileSync(path.join(PNG_DIR, `rag-app-${lang}-${theme}.png`), png);
await saveWebp(png, `rag-app-${lang}-${theme}.webp`);
await context.close();
}
}
await browser.close();
console.log("[demo] done");
+132 -58
View File
@@ -34,96 +34,172 @@ const MOCK = `http://127.0.0.1:${MOCK_PORT}`;
// Commands are shared across languages (code is code) and really execute.
// ---------------------------------------------------------------------------
const CMD_SCAFFOLD = `mkdir -p csv-analyst/src && cat > csv-analyst/package.json <<'EOF'
const CMD_COLLECT = `git clone --depth 1 https://github.com/ericbuess/claude-code-docs \\
claude-code-expert/corpus/claude-code-docs
find claude-code-expert/corpus -type f ! -name '*.md' -delete
rm -rf claude-code-expert/corpus/claude-code-docs/.git
ls claude-code-expert/corpus/claude-code-docs | head -6`;
const CMD_APP = `mkdir -p claude-code-expert/src claude-code-expert/public
cat > claude-code-expert/package.json <<'EOF'
{
"name": "csv-analyst",
"name": "claude-code-expert",
"private": true,
"type": "module",
"scripts": { "start": "tsx src/agent.ts" },
"dependencies": { "@prismshadow/penguin-core": "^0.1.0" }
"scripts": { "start": "tsx src/rag.ts" },
"dependencies": { "@prismshadow/penguin-core": "^0.1.0", "tsx": "^4" }
}
EOF
ls -R csv-analyst`;
cat > claude-code-expert/src/rag.ts <<'EOF'
// BM25 retrieval over corpus/ + a Session that answers with clickable [n] citations.
import fs from "node:fs";
import http from "node:http";
import path from "node:path";
import { createAgent, isModelMessage, userText } from "@prismshadow/penguin-core";
const CMD_ENTRY = `cat > csv-analyst/src/agent.ts <<'EOF'
import { createAgent, isCompleteModelMessage, userText } from "@prismshadow/penguin-core";
const walk = (d) =>
fs.readdirSync(d, { withFileTypes: true }).flatMap((e) =>
e.isDirectory() ? walk(path.join(d, e.name)) : [path.join(d, e.name)]);
const chunks = walk("corpus").flatMap((f) =>
fs.readFileSync(f, "utf8").split(/\\n(?=#{1,3} )/).map((text) => ({ source: f, text })));
const agent = await createAgent({ agentId: "csv_analyst" });
const session = await agent.createSession({ workspaceDir: process.cwd() });
for await (const out of session.run([userText("Analyze data.csv and write summary.md")], {
approve: async () => "allow",
})) {
if (isCompleteModelMessage(out) && out.payload.type === "text") {
console.log(out.payload.text);
const tok = (s) => s.toLowerCase().match(/[a-z0-9]+|[一-鿿]/g) ?? [];
const docs = chunks.map((c) => tok(c.text));
const avg = docs.reduce((n, d) => n + d.length, 0) / Math.max(docs.length, 1);
const df = new Map();
for (const d of docs) for (const w of new Set(d)) df.set(w, (df.get(w) ?? 0) + 1);
// BM25 (k1=1.2, b=0.75): idf weights rare terms, term frequency saturates, long chunks are penalized.
const score = (q, i) => {
const d = docs[i];
let s = 0;
for (const w of new Set(tok(q))) {
const f = d.filter((x) => x === w).length;
if (!f) continue;
const n = df.get(w) ?? 0;
s += Math.log(1 + (docs.length - n + 0.5) / (n + 0.5)) * (f * 2.2) / (f + 1.2 * (0.25 + 0.75 * d.length / avg));
}
}
return s;
};
const agent = await createAgent({ root: "penguin_data" });
http.createServer(async (req, res) => {
if (req.method !== "POST") { res.end(fs.readFileSync("public/index.html")); return; }
let body = "";
for await (const p of req) body += p;
const { question } = JSON.parse(body);
const hits = chunks.map((c, i) => [score(question, i), c]).filter(([s]) => s > 0)
.sort((a, b) => b[0] - a[0]).slice(0, 6).map(([, c]) => c);
res.writeHead(200, { "content-type": "text/event-stream" });
const ctx = hits.map((c, i) => "[" + (i + 1) + "] " + c.source + "\\n" + c.text).join("\\n\\n");
const session = await agent.createSession({ workspaceDir: process.cwd() });
for await (const m of session.run([userText(ctx + "\\n\\nQ: " + question)], {
approve: async () => "deny",
})) {
if (isModelMessage(m) && m.payload.type === "partial_text" && m.payload.event_type === "delta")
res.write("data: " + JSON.stringify({ delta: m.payload.text }) + "\\n\\n");
}
res.write("data: " + JSON.stringify({ sources: hits.map((c) => c.source) }) + "\\n\\n");
session.dispose();
res.end();
}).listen(4630);
EOF
wc -l csv-analyst/src/agent.ts`;
cat > claude-code-expert/public/index.html <<'EOF'
<!doctype html>
<meta charset="utf-8" />
<title>Claude Code docs expert</title>
<body style="max-width:640px;margin:3rem auto;font-family:system-ui;line-height:1.5">
<h3>Claude Code docs expert</h3>
<div id="log"></div>
<input id="q" style="width:100%;padding:.5rem" placeholder="Ask about Claude Code…" autofocus />
<script>
q.addEventListener("keydown", async (e) => {
if (e.key !== "Enter" || !q.value.trim()) return;
const question = q.value; q.value = "";
const p = document.createElement("p"); log.append(p);
const res = await fetch("/api/ask", { method: "POST", body: JSON.stringify({ question }) });
const reader = res.body.getReader(), dec = new TextDecoder();
for (let r; !(r = await reader.read()).done; )
for (const chunk of dec.decode(r.value).split("\\n\\n")) {
if (!chunk.startsWith("data:")) continue;
const d = JSON.parse(chunk.slice(5));
if (d.delta) p.textContent += d.delta;
if (d.sources) p.textContent += " Sources: " + d.sources.join(", ");
}
});
</script>
</body>
EOF
wc -l claude-code-expert/src/rag.ts`;
const TREE = `\`\`\`text
csv-analyst/
claude-code-expert/
├── package.json
└── src/
└── agent.ts
├── corpus/claude-code-docs/ # the collected docs
├── src/rag.ts
└── public/index.html
\`\`\``;
/** Per-language script: user prompt marker -> turns + session title. */
const SCRIPTS = {
zh: {
marker: "数据分析 Agent 应用",
prompt: "用 PenguinHarness SDK 创建一个数据分析 Agent 应用:读取 CSV 并输出汇总报告",
title: "构建数据分析 Agent 应用",
marker: "配置专家",
prompt:
"收集 https://github.com/ericbuess/claude-code-docs 的文档,做一个化身 Claude Code 配置专家、回答带来源引用的 RAG 问答应用。",
title: "构建 Claude Code 文档专家",
turns: [
{
thinking:
"需求是基于 penguin-core 的数据分析 Agent 应用。先创建项目骨架:package.json 与源码目录。",
text: "我来创建应用骨架:",
cmd: CMD_SCAFFOLD,
"一句话需求,penguin-sdk 技能给出了完整配方:先采集语料,再建 BM25 索引,回答经 Session 流式生成并带 [n] 引用直达原文。",
text: "先把文档采集进语料库:",
cmd: CMD_COLLECT,
},
{
thinking:
"骨架已建好。入口代码用 createAgent + createSession,把 CSV 分析任务交给 session.run。",
text: "骨架就绪,写入 Agent 入口代码:创建 Session,把 CSV 分析任务交给 session.run 并流式输出。",
cmd: CMD_ENTRY,
"语料就绪。写应用:本地 BM25 检索(零 embedding 依赖)+ 引用回答 + Penguin 风格聊天界面,引用链接到 /corpus 原文。",
text: "语料就绪,写检索与引用回答的应用入口:",
cmd: CMD_APP,
},
{
text: `数据分析 Agent 应用已创建完成:
text: `Claude Code 文档专家已就绪:
${TREE}
- 入口 \`csv-analyst/src/agent.ts\`:创建 Agent 与 Session,任务经 \`session.run\` 流式执行,工具调用逐个审批;
- 运行方式:\`cd csv-analyst && npm install && npm start\`;
- 建议下一步:在评估中心为它配一组 CSV 任务 Benchmark,交给 Optimizer 持续优化。`,
- 检索:本地 BM25 索引全部文档片段,中文提问同样支持;
- 回答:每次提问经 Session 流式生成,引用 [1][2] 可点击直达 \`corpus/\` 原文;
- 界面:Penguin 风格聊天页,空态内置示例问题;
- 运行:\`cd claude-code-expert && npm install && npm start\`,浏览器打开 http://localhost:4630。`,
},
],
},
en: {
marker: "data-analysis Agent app",
marker: "configuration expert",
prompt:
"Use the PenguinHarness SDK to create a data-analysis Agent app that reads CSV files and writes a summary report",
title: "Build a data-analysis Agent app",
"Collect the docs from https://github.com/ericbuess/claude-code-docs and build a RAG app that answers Claude Code questions as a configuration expert, citing its sources.",
// Stay under core's TITLE_MAX_CHARS (30): a longer title gets hard-clipped
// mid-word in the shots (e.g. "…docs exper").
title: "Build Claude Code docs expert",
turns: [
{
thinking:
"They want a data-analysis Agent app on penguin-core. Start with the project skeleton: package.json plus the source directory.",
text: "Let me scaffold the app first:",
cmd: CMD_SCAFFOLD,
"One sentence is enough — the penguin-sdk skill has the full recipe: collect the corpus, build a BM25 index, answer through a Session with [n] citations linking to the originals.",
text: "Collecting the docs into the corpus first:",
cmd: CMD_COLLECT,
},
{
thinking:
"Skeleton is in place. The entry uses createAgent + createSession and hands the CSV task to session.run.",
text: "Skeleton ready — now the Agent entry point: create a Session and hand the CSV analysis task to session.run, streaming the output.",
cmd: CMD_ENTRY,
"Corpus in place. Now the app: local BM25 retrieval (no embedding credential), cited answers, and a Penguin-style chat UI with citations linking to /corpus originals.",
text: "Corpus ready — now the retrieval and cited-answer entry:",
cmd: CMD_APP,
},
{
text: `The data-analysis Agent app is ready:
text: `The Claude Code docs expert is ready:
${TREE}
- Entry \`csv-analyst/src/agent.ts\`: creates the Agent and a Session; the task runs through \`session.run\` with per-tool approval;
- Run it with \`cd csv-analyst && npm install && npm start\`;
- Suggested next step: give it a CSV Benchmark suite in the evaluation center and let an Optimizer keep improving it.`,
- Retrieval: a local BM25 index over every doc chunk, Chinese questions included;
- Answers: each question streams through a Session, with [1][2] citations that link straight to the \`corpus/\` originals;
- UI: a Penguin-style chat page with example questions in the empty state;
- Run: \`cd claude-code-expert && npm install && npm start\`, then open http://localhost:4630.`,
},
],
},
@@ -466,7 +542,7 @@ try {
await waitFor(`${BASE}/`);
console.log(`[shots] server ready on ${BASE}`);
const admin = await login("admin", "admin123");
const admin = await login("admin", "penguin-2026");
const browser = await chromium.launch();
// WebP encoder: Chromium re-encodes the PNG screenshot buffer via canvas, which
@@ -549,17 +625,15 @@ try {
await page.waitForTimeout(2000);
await saveWebp(await page.screenshot(), `chat-${lang}-${theme}.webp`);
// Trace view: select the session in the list (deep-link selection is unreliable
// right after a fresh navigation, so click explicitly — sidebar shows the same
// title first in DOM order, hence .last()).
await page.goto(`${BASE}/traces?sessionId=${sessionId}`);
await page.waitForTimeout(1500);
await page
.getByText(script.title)
.last()
.click()
.catch(() => {});
await page.waitForTimeout(2500);
// Trace view: the product's canonical deep link carries BOTH agentId and
// sessionId (?sessionId= alone is ignored by TracesPage's focus wiring), and
// auto-selects the Session once its trace list loads. Waiting for the
// execution timeline's exec_command lanes guarantees every language captures
// the same opened trace — stats + a timeline with tool calls — never the
// empty "select a Session" state.
await page.goto(`${BASE}/traces?agentId=default_agent&sessionId=${sessionId}`);
await page.getByText("exec_command").first().waitFor({ timeout: 20000 });
await page.waitForTimeout(2000);
await saveWebp(await page.screenshot(), `traces-${lang}-${theme}.webp`);
// Evaluation center: open the pre-provisioned example Benchmark scoreboard.
@@ -0,0 +1,333 @@
<!doctype html>
<!--
Static, dependency-free mockup of the "penguin sled" example game's play screen —
the FINISHED PRODUCT of the draft-screen game example, captured by
capture-game-mockup.mjs per language (?lang=zh|en) and theme (?theme=light|dark;
light = polar day, dark = polar night with an aurora). A cute cartoon scene:
an Antarctic penguin on a wooden sled mid-jump over snow-capped rocks, with a
score/distance/speed HUD and a "Space to jump" hint.
-->
<html lang="en">
<head>
<meta charset="utf-8" />
<title>Penguin Sled Run</title>
<style>
html,
body {
margin: 0;
height: 100%;
font-family:
ui-sans-serif,
system-ui,
-apple-system,
"Segoe UI",
Roboto,
"PingFang SC",
"Microsoft YaHei",
sans-serif;
}
.stage {
position: relative;
width: 1280px;
height: 760px;
overflow: hidden;
background: linear-gradient(#bfdbfe 0%, #e0f2fe 55%, #f8fafc 100%);
}
body.dark .stage {
background: linear-gradient(#0b1220 0%, #16233b 60%, #1e293b 100%);
}
.hud {
position: absolute;
display: flex;
align-items: center;
gap: 8px;
padding: 10px 16px;
border-radius: 14px;
background: rgba(255, 255, 255, 0.82);
border: 1px solid rgba(15, 23, 42, 0.08);
color: #0f172a;
font-weight: 700;
box-shadow: 0 4px 14px rgba(15, 23, 42, 0.08);
}
body.dark .hud {
background: rgba(15, 23, 42, 0.72);
border-color: rgba(148, 163, 184, 0.25);
color: #e2e8f0;
box-shadow: 0 4px 14px rgba(0, 0, 0, 0.35);
}
.hud .label {
font-size: 13px;
font-weight: 600;
opacity: 0.62;
letter-spacing: 0.04em;
}
.hud .value {
font-size: 24px;
font-variant-numeric: tabular-nums;
}
#hud-score {
top: 22px;
left: 24px;
}
#hud-dist {
top: 22px;
left: 218px;
}
#hud-speed {
top: 22px;
right: 24px;
color: #1d4ed8;
}
body.dark #hud-speed {
color: #93c5fd;
}
.hint {
position: absolute;
left: 50%;
bottom: 26px;
transform: translateX(-50%);
display: flex;
align-items: center;
gap: 10px;
padding: 10px 18px;
border-radius: 999px;
font-size: 15px;
font-weight: 600;
background: rgba(255, 255, 255, 0.85);
border: 1px solid rgba(15, 23, 42, 0.08);
color: #334155;
}
body.dark .hint {
background: rgba(15, 23, 42, 0.75);
border-color: rgba(148, 163, 184, 0.25);
color: #cbd5e1;
}
.hint kbd {
padding: 3px 12px;
border-radius: 8px;
border: 1px solid #94a3b8;
border-bottom-width: 3px;
background: #fff;
font: inherit;
color: #0f172a;
}
body.dark .hint kbd {
background: #1e293b;
border-color: #64748b;
color: #e2e8f0;
}
.title {
position: absolute;
top: 26px;
left: 50%;
transform: translateX(-50%);
font-size: 17px;
font-weight: 800;
letter-spacing: 0.12em;
color: rgba(15, 23, 42, 0.55);
}
body.dark .title {
color: rgba(226, 232, 240, 0.65);
}
.day-only {
display: initial;
}
.night-only {
display: none;
}
body.dark .day-only {
display: none;
}
body.dark .night-only {
display: initial;
}
</style>
</head>
<body>
<div class="stage">
<svg viewBox="0 0 1280 760" width="1280" height="760" aria-hidden="true">
<!-- Polar day: sun; polar night: moon, stars and an aurora ribbon. -->
<g class="day-only">
<circle cx="1084" cy="150" r="58" fill="#fde68a" />
<circle cx="1084" cy="150" r="44" fill="#fcd34d" />
</g>
<g class="night-only">
<path
d="M60 60 C 320 190, 680 30, 1020 150 C 1120 185, 1220 140, 1268 100 L 1268 20 L 60 20 Z"
fill="#34d399"
opacity="0.16"
/>
<path
d="M100 100 C 360 220, 700 80, 1040 190 C 1130 218, 1210 180, 1260 150 L 1260 60 C 1000 160, 420 40, 100 100 Z"
fill="#6ee7b7"
opacity="0.12"
/>
<circle cx="1084" cy="140" r="46" fill="#e2e8f0" opacity="0.92" />
<circle cx="1066" cy="128" r="12" fill="#cbd5e1" opacity="0.8" />
<circle cx="1102" cy="156" r="8" fill="#cbd5e1" opacity="0.8" />
<g fill="#e2e8f0">
<circle cx="180" cy="120" r="3" />
<circle cx="330" cy="70" r="2.4" />
<circle cx="520" cy="130" r="2.8" />
<circle cx="760" cy="64" r="2.4" />
<circle cx="900" cy="120" r="3" />
<circle cx="240" cy="220" r="2" />
<circle cx="640" cy="200" r="2.2" />
</g>
</g>
<!-- Distant snow hills. -->
<path
d="M0 430 Q 200 330 420 420 T 860 410 T 1280 400 L 1280 760 L 0 760 Z"
class="hill-far"
fill="#e2e8f0"
/>
<path
d="M0 500 Q 260 400 560 490 T 1280 470 L 1280 760 L 0 760 Z"
class="hill-near"
fill="#f1f5f9"
/>
<!-- Ice track. -->
<rect x="0" y="560" width="1280" height="200" fill="#ffffff" />
<rect x="0" y="560" width="1280" height="12" fill="#bae6fd" opacity="0.7" />
<path
d="M0 640 H 300 M420 640 H 700 M820 640 H 1100"
stroke="#dbeafe"
stroke-width="6"
stroke-linecap="round"
/>
<!-- Rocks ahead on the track (snow-capped). -->
<g>
<path d="M760 560 q 26 -52 62 -2 q 20 30 -8 30 l -60 0 q -14 -6 6 -28 Z" fill="#94a3b8" />
<path d="M770 548 q 18 -26 40 -4 q -20 -10 -40 4 Z" fill="#f8fafc" />
<path
d="M980 560 q 34 -70 84 -4 q 24 36 -10 36 l -80 0 q -18 -8 6 -32 Z"
fill="#64748b"
/>
<path d="M994 540 q 26 -34 54 -6 q -28 -14 -54 6 Z" fill="#f8fafc" />
<path
d="M1180 560 q 22 -44 52 -2 q 16 26 -6 26 l -52 0 q -12 -6 6 -24 Z"
fill="#94a3b8"
/>
</g>
<!-- The star: a cute penguin on a wooden sled, mid-jump over the first rock.
Drawn upright and compact so every part stays connected to the body. -->
<g transform="translate(430 316) rotate(-8)">
<!-- motion lines -->
<path
d="M-158 118 h 70 M-178 156 h 58 M-150 194 h 66"
stroke="#93c5fd"
stroke-width="8"
stroke-linecap="round"
opacity="0.7"
/>
<!-- snow spray behind the sled -->
<g fill="#e0f2fe">
<circle cx="-92" cy="214" r="9" />
<circle cx="-116" cy="230" r="6" />
<circle cx="-70" cy="232" r="5" />
</g>
<!-- sled: two wooden slats + a curled steel runner -->
<path
d="M-92 208 q -22 0 -20 -22 q 2 -14 20 -14 h 250 q 26 2 26 22 q 0 14 -24 14 Z"
fill="none"
stroke="#475569"
stroke-width="9"
stroke-linecap="round"
/>
<rect x="-70" y="176" width="242" height="24" rx="12" fill="#b45309" />
<rect x="-58" y="158" width="220" height="22" rx="11" fill="#d97706" />
<!-- penguin, centered on the sled -->
<g transform="translate(52 60)">
<!-- body (black egg) -->
<path
d="M0 -96 C 62 -96 82 -34 82 20 C 82 78 44 108 0 108 C -44 108 -82 78 -82 20 C -82 -34 -62 -96 0 -96 Z"
fill="#111827"
/>
<!-- white belly, tucked inside the body outline -->
<path
d="M0 -52 C 40 -52 56 -10 56 30 C 56 66 30 88 0 88 C -30 88 -56 66 -56 30 C -56 -10 -40 -52 0 -52 Z"
fill="#f8fafc"
/>
<!-- near wing, hugging the body -->
<path d="M-66 -20 C -92 4 -92 52 -70 74 C -58 60 -58 16 -60 -8 Z" fill="#0b1220" />
<!-- far wing, raised a little (whee!) -->
<path d="M62 -30 C 92 -44 104 -16 96 12 C 82 6 70 -8 60 -18 Z" fill="#0b1220" />
<!-- cheeks -->
<circle cx="-30" cy="-14" r="11" fill="#fda4af" opacity="0.8" />
<circle cx="30" cy="-14" r="11" fill="#fda4af" opacity="0.8" />
<!-- eyes -->
<circle cx="-20" cy="-40" r="9" fill="#f8fafc" />
<circle cx="20" cy="-40" r="9" fill="#f8fafc" />
<circle cx="-18" cy="-38" r="4.4" fill="#0b1220" />
<circle cx="22" cy="-38" r="4.4" fill="#0b1220" />
<circle cx="-16" cy="-40" r="1.5" fill="#f8fafc" />
<circle cx="24" cy="-40" r="1.5" fill="#f8fafc" />
<!-- beak -->
<path d="M-11 -22 L 11 -22 L 0 -8 Z" fill="#f59e0b" />
<!-- scarf around the neck (two bands, one short tail) -->
<path d="M-40 4 Q 0 22 40 4 L 40 20 Q 0 38 -40 20 Z" fill="#ef4444" />
<path d="M28 14 q 26 6 30 26 q -18 6 -34 -6 Z" fill="#f87171" />
<!-- feet resting on the front slat -->
<path d="M-26 104 q -14 12 4 20 q 20 4 26 -10 Z" fill="#fb923c" />
<path d="M26 104 q -6 14 12 18 q 20 2 22 -12 Z" fill="#fb923c" />
</g>
</g>
</svg>
<div class="title" id="title"></div>
<div class="hud" id="hud-score">
<span class="label" id="score-label"></span><span class="value">1280</span>
</div>
<div class="hud" id="hud-dist">
<span class="label" id="dist-label"></span><span class="value">512 m</span>
</div>
<div class="hud" id="hud-speed"><span class="value">×2.4</span></div>
<div class="hint"><kbd id="key"></kbd><span id="hint"></span></div>
</div>
<script>
const params = new URLSearchParams(location.search);
const lang = params.get("lang") === "zh" ? "zh" : "en";
if (params.get("theme") === "dark") document.body.classList.add("dark");
const T = {
zh: {
title: "企 鹅 雪 橇 越 野",
score: "得分",
dist: "里程",
key: "空格",
hint: "起跳,跃过石头!",
},
en: {
title: "P E N G U I N · S L E D · R U N",
score: "SCORE",
dist: "DISTANCE",
key: "Space",
hint: "jump the rocks!",
},
}[lang];
document.documentElement.lang = lang === "zh" ? "zh-CN" : "en";
document.getElementById("title").textContent = T.title;
document.getElementById("score-label").textContent = T.score;
document.getElementById("dist-label").textContent = T.dist;
document.getElementById("key").textContent = T.key;
document.getElementById("hint").textContent = T.hint;
const dark = document.body.classList.contains("dark");
if (dark) {
for (const [sel, fill] of [
[".hill-far", "#243b55"],
[".hill-near", "#2c476b"],
]) {
document.querySelector(sel).setAttribute("fill", fill);
}
for (const rect of document.querySelectorAll("svg rect")) {
if (rect.getAttribute("fill") === "#ffffff") rect.setAttribute("fill", "#dbeafe");
}
}
</script>
</body>
</html>
@@ -0,0 +1,439 @@
<!doctype html>
<html lang="en">
<!--
Finished-product mockup of the "one sentence → RAG app" example (a Claude Code docs
expert), rendered by capture-readme-demo.mjs into assets/readme/rag-app-<lang>-<theme>.webp.
Pure static page — no build, no server; ?lang=zh|en&theme=light|dark picks the variant.
Styling follows the web-design skill's Penguin visual language: solid backgrounds,
1px hairlines, system fonts, one brand-blue accent, pure-black dark mode.
-->
<head>
<meta charset="utf-8" />
<title>Claude Code Docs Expert</title>
<style>
:root {
color-scheme: light;
--brand-50: #e8f0fe;
--brand-100: #d2e3fc;
--brand-300: #8ab4f8;
--brand-500: #4285f4;
--brand-600: #1a73e8;
--brand-700: #0b57d0;
--bg: #ffffff;
--surface: #ffffff;
--surface-2: #f9fafb;
--border: #e5e7eb;
--control-border: #d1d5db;
--fg: #111827;
--fg-muted: #4b5563;
--fg-faint: #6b7280;
--accent-bg: #111827;
--accent-fg: #ffffff;
--bubble: #f3f4f6;
--code-bg: #f9fafb;
--chip-bg: #e8f0fe;
--chip-fg: #0b57d0;
--chip-border: #d2e3fc;
--dots: rgb(26 115 232 / 0.14);
}
.dark {
color-scheme: dark;
--bg: #000000;
--surface: #0d0d0d;
--surface-2: #0d0d0d;
--border: #1f1f1f;
--control-border: #303030;
--fg: #f3f4f6;
--fg-muted: #9ca3af;
--fg-faint: #6b7280;
--accent-bg: #f3f4f6;
--accent-fg: #111827;
--bubble: #1f1f1f;
--code-bg: #0d0d0d;
--chip-bg: #041e49;
--chip-fg: #8ab4f8;
--chip-border: #062e6f;
--dots: rgb(138 180 248 / 0.16);
}
* {
box-sizing: border-box;
margin: 0;
}
body {
font-family:
ui-sans-serif,
system-ui,
-apple-system,
"Segoe UI",
Roboto,
"PingFang SC",
"Microsoft YaHei",
sans-serif;
-webkit-font-smoothing: antialiased;
background: var(--bg);
color: var(--fg);
height: 100vh;
display: flex;
flex-direction: column;
overflow: hidden;
}
code,
pre {
font-family: ui-monospace, SFMono-Regular, Menlo, Consolas, monospace;
}
nav {
height: 56px;
flex: none;
display: flex;
align-items: center;
gap: 10px;
padding: 0 24px;
border-bottom: 1px solid var(--border);
background: color-mix(in srgb, var(--bg) 85%, transparent);
}
.logo {
width: 28px;
height: 28px;
border-radius: 8px;
background: var(--accent-bg);
color: var(--accent-fg);
display: grid;
place-items: center;
flex: none;
}
.logo svg {
width: 16px;
height: 16px;
}
.app-name {
font-size: 15px;
font-weight: 600;
letter-spacing: -0.01em;
}
.nav-chip {
font-size: 11px;
font-weight: 500;
color: var(--chip-fg);
background: var(--chip-bg);
border: 1px solid var(--chip-border);
border-radius: 9999px;
padding: 2px 10px;
}
.nav-right {
margin-left: auto;
display: flex;
align-items: center;
gap: 6px;
font-size: 12px;
color: var(--fg-faint);
}
.dot {
width: 6px;
height: 6px;
border-radius: 9999px;
background: #10b981;
}
main {
flex: 1;
min-height: 0;
display: flex;
flex-direction: column;
width: 100%;
max-width: 48rem;
margin: 0 auto;
padding: 0 24px;
}
/* Decorative dot grid fading out below the nav — the one flourish the design language
allows. The ::after overlay fades to the page background, so the tail disappears
cleanly in dark mode too (light dots on black stay visible far longer under a mask). */
.dots-band {
position: absolute;
inset: 56px 0 auto 0;
height: 150px;
background-image: radial-gradient(var(--dots) 1px, transparent 1px);
background-size: 22px 22px;
pointer-events: none;
}
.dots-band::after {
content: "";
position: absolute;
inset: 0;
background: linear-gradient(
to bottom,
transparent 0%,
color-mix(in srgb, var(--bg) 55%, transparent) 55%,
var(--bg) 100%
);
}
.chat {
flex: 1;
min-height: 0;
overflow: hidden;
padding-top: 28px;
display: flex;
flex-direction: column;
gap: 20px;
position: relative;
}
.msg-user {
align-self: flex-end;
max-width: 85%;
background: var(--bubble);
border-radius: 12px;
padding: 9px 14px;
font-size: 14px;
line-height: 1.55;
}
.msg-assistant {
font-size: 14px;
line-height: 1.65;
color: var(--fg);
display: flex;
flex-direction: column;
gap: 10px;
}
.msg-assistant p {
max-width: 65ch;
}
.cite {
color: var(--brand-600);
font-size: 0.8em;
vertical-align: super;
font-weight: 600;
text-decoration: none;
}
.dark .cite {
color: var(--brand-300);
}
pre.block {
background: var(--code-bg);
border: 1px solid var(--border);
border-radius: 8px;
padding: 12px 14px;
font-size: 13px;
line-height: 1.85;
overflow-x: auto;
}
pre.block .c {
color: var(--fg-faint);
}
.sources {
display: flex;
align-items: center;
flex-wrap: wrap;
gap: 8px;
margin-top: 2px;
}
.sources-label {
font-size: 11px;
font-weight: 600;
text-transform: uppercase;
letter-spacing: 0.05em;
color: var(--fg-faint);
}
.source-chip {
display: inline-flex;
align-items: center;
gap: 6px;
font-size: 12px;
color: var(--chip-fg);
background: var(--chip-bg);
border: 1px solid var(--chip-border);
border-radius: 9999px;
padding: 3px 11px;
text-decoration: none;
}
.source-chip .n {
font-weight: 600;
}
.source-chip .path {
font-family: ui-monospace, SFMono-Regular, Menlo, Consolas, monospace;
font-size: 11px;
}
.bottom {
flex: none;
padding: 14px 0 20px;
display: flex;
flex-direction: column;
gap: 10px;
}
.suggests {
display: flex;
flex-wrap: wrap;
gap: 8px;
}
.suggest {
font-size: 12px;
color: var(--fg-muted);
border: 1px solid var(--border);
background: var(--surface);
border-radius: 9999px;
padding: 5px 12px;
}
.composer {
border: 1px solid var(--control-border);
background: var(--surface);
border-radius: 12px;
padding: 12px 12px 10px 14px;
display: flex;
align-items: flex-end;
gap: 10px;
}
.composer .ph {
flex: 1;
font-size: 14px;
color: var(--fg-faint);
padding: 4px 0 6px;
}
.send {
width: 32px;
height: 32px;
border-radius: 8px;
border: none;
background: var(--accent-bg);
color: var(--accent-fg);
display: grid;
place-items: center;
flex: none;
}
.hint {
font-size: 11px;
color: var(--fg-faint);
text-align: center;
}
</style>
</head>
<body>
<div class="dots-band"></div>
<nav>
<div class="logo" aria-hidden>
<svg viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="1.8">
<path d="M4 19.5V6a2 2 0 0 1 2-2h12a2 2 0 0 1 2 2v13.5" stroke="none" />
<path
d="M12 3c3.5 0 5.5 2.6 5.5 6.2 0 4.4 1.6 6.3 2.5 7.1-1.6 2.9-4.6 4.7-8 4.7s-6.4-1.8-8-4.7c.9-.8 2.5-2.7 2.5-7.1C6.5 5.6 8.5 3 12 3z"
/>
<circle cx="9.8" cy="8.6" r="0.6" fill="currentColor" stroke="none" />
<circle cx="14.2" cy="8.6" r="0.6" fill="currentColor" stroke="none" />
<path d="M10.6 11h2.8l-1.4 1.6z" fill="currentColor" stroke="none" />
</svg>
</div>
<span class="app-name" data-t="appName"></span>
<span class="nav-chip">RAG · claude-code-docs</span>
<div class="nav-right"><span class="dot"></span><span data-t="corpus"></span></div>
</nav>
<main>
<div class="chat">
<div class="msg-user" data-t="question"></div>
<div class="msg-assistant">
<p>
<span data-t="a1"></span>
<a class="cite" href="#">[1]</a><span data-t="a1b"></span>
</p>
<pre class="block"><span class="c" data-t="codeC1"></span>
claude mcp add github -- npx -y @modelcontextprotocol/server-github
<span class="c" data-t="codeC2"></span>
claude mcp list</pre>
<p>
<span data-t="a2"></span>
<a class="cite" href="#">[2]</a><span data-t="a2b"></span>
</p>
<div class="sources">
<span class="sources-label" data-t="sources"></span>
<a class="source-chip" href="#"
><span class="n">[1]</span><span class="path">docs/mcp.md</span
><span data-t="s1"></span
></a>
<a class="source-chip" href="#"
><span class="n">[2]</span><span class="path">docs/settings.md</span
><span data-t="s2"></span
></a>
</div>
</div>
</div>
<div class="bottom">
<div class="suggests">
<span class="suggest" data-t="q1"></span>
<span class="suggest" data-t="q2"></span>
<span class="suggest" data-t="q3"></span>
</div>
<div class="composer">
<div class="ph" data-t="placeholder"></div>
<button class="send" aria-hidden>
<svg
width="15"
height="15"
viewBox="0 0 24 24"
fill="none"
stroke="currentColor"
stroke-width="2"
stroke-linecap="round"
stroke-linejoin="round"
>
<path d="M12 19V5M6 11l6-6 6 6" />
</svg>
</button>
</div>
<div class="hint" data-t="hint"></div>
</div>
</main>
<script>
const T = {
en: {
appName: "Claude Code Docs Expert",
corpus: "674 chunks · BM25",
question: "How do I configure an MCP server for Claude Code?",
a1: "Add the server with the CLI, then verify it from inside Claude Code ",
a1b: ":",
codeC1: "# add a server (local scope by default)",
codeC2: "# verify",
a2: "Project-scoped servers live in .mcp.json at the project root and are shared with your team; user scope makes a server available across all your projects ",
a2b: ". Inside a session, /mcp shows each server's connection status.",
sources: "Sources",
s1: "— Add an MCP server",
s2: "— Settings & scopes",
q1: "What does CLAUDE.md do?",
q2: "How do I create a custom slash command?",
q3: "Where do hooks go?",
placeholder: "Ask about Claude Code…",
hint: "Answers cite the collected docs — click a source to open the original file.",
},
zh: {
appName: "Claude Code 配置专家",
corpus: "674 个片段 · BM25",
question: "如何为 Claude Code 配置 MCP 服务器?",
a1: "用 CLI 添加服务器,然后在 Claude Code 里验证 ",
a1b: ":",
codeC1: "# 添加服务器(默认 local 作用域)",
codeC2: "# 验证",
a2: "项目级服务器保存在项目根目录的 .mcp.json,可随仓库共享给团队;user 作用域则让服务器在你的所有项目中可用 ",
a2b: "。会话内可用 /mcp 查看各服务器的连接状态。",
sources: "来源",
s1: "— 添加 MCP 服务器",
s2: "— 设置与作用域",
q1: "CLAUDE.md 是做什么的?",
q2: "怎么创建自定义斜杠命令?",
q3: "Hooks 配置放在哪里?",
placeholder: "向 Claude Code 专家提问…",
hint: "回答引用采集的文档——点击来源即可打开原文。",
},
};
const params = new URLSearchParams(location.search);
const lang = params.get("lang") === "zh" ? "zh" : "en";
const theme = params.get("theme") === "dark" ? "dark" : "light";
document.documentElement.lang = lang === "zh" ? "zh" : "en";
if (theme === "dark") document.documentElement.classList.add("dark");
for (const el of document.querySelectorAll("[data-t]")) {
el.textContent = T[lang][el.dataset.t];
}
document.title = T[lang].appName;
</script>
</body>
</html>
@@ -0,0 +1,179 @@
/**
* Renders the README / blog benchmark chart from the same data the landing page uses
* (src/lib/benchmark-data.ts), so the static SVGs cannot drift from the site when the
* numbers are refreshed. Emits:
* assets/readme/benchmark-light.svg
* assets/readme/benchmark-dark.svg
* packages/landing/public/blog-assets/benchmark-light.svg
*
* Two panels per suite (accuracy, cost), horizontal bars scaled linearly from zero — the
* cost spread is ~70x, so the PenguinHarness bar is by far the shortest. That is the
* result, not a defect; a zoomed baseline would flatter everyone else. Bars are floored at
* MIN_BAR, though: at true scale the cost bar comes out ~2.5px, which reads as "no bar" and
* loses the series entirely. The floor keeps it visible and unmistakably smallest, and the
* exact figure is printed beside every bar, so nothing is overstated.
*
* Run: node packages/landing/scripts/render-benchmark-svg.mjs
*/
import { readFileSync, writeFileSync } from "node:fs";
import path from "node:path";
import { fileURLToPath } from "node:url";
const HERE = path.dirname(fileURLToPath(import.meta.url));
const LANDING = path.resolve(HERE, "..");
const REPO = path.resolve(LANDING, "..", "..");
/**
* The data module is TypeScript; rather than add a build step, read the two exported
* arrays out of the source. Any shape change here fails loudly instead of silently
* rendering a stale chart.
*/
function loadBench() {
const src = readFileSync(path.join(LANDING, "src/lib/benchmark-data.ts"), "utf8");
const consts = Object.fromEntries(
[...src.matchAll(/^const (\w+) = "([^"]+)";$/gm)].map((m) => [m[1], m[2]]),
);
const pick = (name) => {
const block = src.match(
new RegExp(`export const ${name}: BenchResult\\[\\] = \\[([\\s\\S]*?)\\n\\];`),
);
if (!block) throw new Error(`could not find ${name} in benchmark-data.ts`);
const rows = [...block[1].matchAll(/\{([\s\S]*?)\n {2}\}/g)].map((m) => {
const body = m[1];
const field = (k) => body.match(new RegExp(`${k}:\\s*([^,\\n]+)`))?.[1]?.trim();
const str = (k) => {
const v = field(k);
if (!v) throw new Error(`missing ${k}`);
return v.startsWith('"') ? v.slice(1, -1) : (consts[v] ?? v);
};
return {
framework: str("framework"),
model: str("model"),
accuracyPct: Number(field("accuracyPct")),
tokensM: Number(field("tokensM")),
costUsd: Number(field("costUsd")),
emphasized: /emphasized:\s*true/.test(body),
};
});
if (rows.length !== 3) throw new Error(`${name}: expected 3 rows, got ${rows.length}`);
for (const r of rows) {
if (!Number.isFinite(r.accuracyPct) || !Number.isFinite(r.costUsd)) {
throw new Error(`${name}: ${r.framework} has a non-numeric field`);
}
}
return rows;
};
return { DATA_BENCH: pick("DATA_BENCH"), CODE_BENCH: pick("CODE_BENCH") };
}
const THEMES = {
light: {
strong: "#1f2328",
mid: "#52514e",
muted: "#898781",
rule: "#c3c2b7",
brand: "#2a78d6",
bar: "#898781",
},
dark: {
strong: "#f0f3f6",
mid: "#c3c2b7",
muted: "#898781",
rule: "#383835",
brand: "#3987e5",
bar: "#6b6a64",
},
};
const FONT = "system-ui, -apple-system, 'Segoe UI', sans-serif";
const W = 920;
const H = 368;
const BAR_H = 16;
const ROW_H = 30;
const MAX_BAR = 176;
/** Shortest a bar may render, so a near-zero value still reads as a bar (~8% of full). */
const MIN_BAR = 14;
const esc = (s) => s.replace(/&/g, "&amp;").replace(/</g, "&lt;").replace(/>/g, "&gt;");
function text(x, y, s, { size = 12, weight = 400, fill, anchor = "start" }) {
return `<text x="${x}" y="${y}" font-size="${size}" font-weight="${weight}" fill="${fill}" text-anchor="${anchor}" font-family="${FONT}">${esc(s)}</text>`;
}
/** Rounded-end horizontal bar starting at the axis. */
function bar(x, y, w, fill) {
const r = Math.min(4, w / 2);
return `<path d="M${x} ${y} h${(w - r).toFixed(1)} a${r} ${r} 0 0 1 ${r} ${r} v${BAR_H - 2 * r} a${r} ${r} 0 0 1 -${r} ${r} h-${(w - r).toFixed(1)} z" fill="${fill}"/>`;
}
/** One measure for one suite: label column, axis rule, three bars, value labels. */
function panel(t, x0, yTop, rows, value, format) {
const axis = x0 + 124;
const max = Math.max(...rows.map(value));
const out = [
`<rect x="${axis - 1}" y="${yTop}" width="1" height="${ROW_H * rows.length + 2}" fill="${t.rule}"/>`,
];
rows.forEach((row, i) => {
const y = yTop + 4 + i * ROW_H;
const w = Math.max(MIN_BAR, (value(row) / max) * MAX_BAR);
const weight = row.emphasized ? 600 : 400;
out.push(
text(axis - 12, y + 12.5, row.framework, {
fill: row.emphasized ? t.strong : t.mid,
weight,
anchor: "end",
}),
bar(axis + 1, y, w, row.emphasized ? t.brand : t.bar),
text(axis + w + 13, y + 12.5, format(value(row)), { fill: t.strong, weight }),
);
});
return out.join("\n");
}
function suite(t, yTop, title, rows) {
return [
text(24, yTop, title, { size: 13, weight: 600, fill: t.strong }),
panel(
t,
24,
yTop + 10,
rows,
(r) => r.accuracyPct,
(v) => `${v.toFixed(2)}%`,
),
panel(
t,
472,
yTop + 10,
rows,
(r) => r.costUsd,
(v) => `$${v.toFixed(2)}`,
),
].join("\n");
}
function render(theme, { DATA_BENCH, CODE_BENCH }) {
const t = THEMES[theme];
const alt =
"Benchmark: PenguinHarness vs Claude Code vs OpenAI Codex on two suites — comparable accuracy at a small fraction of the cost";
return `<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 ${W} ${H}" width="${W}" height="${H}" role="img" aria-label="${esc(alt)}">
${text(148, 36, "Accuracy · suite total · higher is better", { fill: t.muted })}
${text(596, 36, "Total cost (USD) · lower is better", { fill: t.muted })}
${suite(t, 68, "Data analysis — 15 tasks, single run", DATA_BENCH)}
${suite(t, 210, "Coding — 40 tasks × 2 runs", CODE_BENCH)}
${text(24, 338, "Each harness runs the model it is normally paired with: PenguinHarness on DeepSeek V4 Pro, Claude Code on Claude Opus 4.8,", { fill: t.muted })}
${text(24, 356, "OpenAI Codex on GPT-5.5. Accuracy, Tokens and cost are suite totals at official pricing.", { fill: t.muted })}
</svg>
`;
}
const bench = loadBench();
const targets = [
["light", path.join(REPO, "assets/readme/benchmark-light.svg")],
["dark", path.join(REPO, "assets/readme/benchmark-dark.svg")],
["light", path.join(LANDING, "public/blog-assets/benchmark-light.svg")],
];
for (const [theme, file] of targets) {
writeFileSync(file, render(theme, bench));
console.log(`wrote ${path.relative(REPO, file)}`);
}
Binary file not shown.

After

Width:  |  Height:  |  Size: 37 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 36 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 40 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 39 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 39 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 39 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 70 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 76 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 64 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 71 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 65 KiB

After

Width:  |  Height:  |  Size: 65 KiB

Some files were not shown because too many files have changed in this diff Show More