Initialize repository with harness code and assets

Initial import of all source code, config, and README assets: the
packages workspace (cli, core, server, web, docs, landing, skills),
build scripts, tooling config, and CI workflows.

Includes the data-layout revision made on this branch: the local data
root defaults to ~/.penguin/data (PENGUIN_HOME still overrides; the
installer keeps its binaries in ~/.penguin), and every Agent lives
under <project>/agents/<agent>/ — path helpers, the three
agent-enumeration scans, the system prompt, built-in Skills, tests
and docs all follow the new layout.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018ihk8iQuo3kv2aPjAYEPuR
This commit is contained in:
Yaowei Zheng
2026-07-19 14:06:53 +08:00
committed by GitHub
parent 056bed7aeb
commit 45bfae6e94
543 changed files with 92949 additions and 0 deletions
+123
View File
@@ -0,0 +1,123 @@
---
title: The Agent Loop
description: The context_engine's master flow diagram and a stage-by-stage breakdown — approvals, concurrent tool execution, interrupt carry-over, automatic reconnect and compaction.
---
The SDK's single execution entry point is `session.run(newMessages, opts?)`: input is the list of new OmniMessages (the Prompt); the return value is an async generator that streams [OmniMessage](/omni-message). One `run` drives one complete Task, until the model produces a final answer with no tool calls.
This page shows the context_engine's overall flow first, then breaks down each stage; the message-level observable timeline and ordering guarantees are on [Message Flow & Ordering](/message-flow). Source: `packages/core/src/engine/context-engine.ts`.
## The loop at a glance
```text
session.run(newMessages, { approve, signal })
│ carry-over from a previous interrupt? → prepend to this run's input
▼
┌── turn loop (≤ max_turns, default 100) ───────────────────────┐
│ │
│ request_begin │
│ LLM.streamGenerate(newMessages) │
│ ├─ streams partial_* fragments + complete msgs │
│ ├─ for each complete tool_call: │
│ │ approve(toolCall) ──deny──► synthetic aborted output │
│ │ │allow (approvals sequential; │
│ │ ▼ decision audited) │
│ │ Environment.executeTool ──► runs concurrently, │
│ │ output streams back │
│ └─ LLMOutcome: │
│ timeout / malformed ──► reconnect within the turn │
│ (≤2, with <turn_retried>; tools not rerun) │
│ token_usage + request_end (at LLM-stream end; not waiting │
│ for tools) │
│ │
│ tool outputs reordered to original call order ──► next turn │
│ no tool_call this turn? ──► Task ends, run returns │
│ compaction trigger (context/turns)? ──► summarize/discard │
│ + Trace rotation │
└───────────────────────────────────────────────────────────────┘
signal fires (any point) ──► emit abort + build carry-over ──► run returns
```
Every message and event flows to two destinations at once: streamed live to the Human, and written to the [Trace](/sessions-and-traces).
## Inputs and outputs
```ts
const agent = await createAgent({ agentId: "default_agent" });
const session = await agent.createSession({ workspaceDir: process.cwd() });
for await (const output of session.run([userText("Clean up the CSV files under data/")], {
approve: async (toolCall) => "allow",
signal: abortController.signal,
})) {
// output: partial_* fragments, complete model_msg, event_msg
}
```
```ts
interface RunOptions {
signal?: AbortSignal; // interrupt (e.g. Ctrl-C)
approve?: ApproveFn; // per-tool approval; denies everything when omitted (conservative default)
}
```
## Lifecycle of a turn
A Task consists of consecutive Requests (turns). Each turn:
1. emits `request_begin`;
2. the LLM streams back: `partial_*` fragments followed by complete messages;
3. every complete `tool_call` triggers exactly one `approve` callback; the decision is recorded as an `approval_decision` event;
4. approved calls run **concurrently** in the Environment (approvals themselves are one at a time); outputs stream out in completion order;
5. when the LLM stream ends, its final `token_usage` is emitted and `request_end(status)` follows at once — **without waiting for tools**: still-running tools may emit output after `request_end`;
6. once the whole batch is terminal, tool results are **reordered to the original call order** and become the next turn's input — the next Request never fires before that.
The Task ends when a turn produces no `tool_call`. A denial produces a synthetic `aborted` tool output ("Tool call denied by user.") that the model reacts to.
## Interruption and carry-over
When `signal` fires, the engine emits an `abort` event and returns immediately, while constructing carry-over content for the next `run`:
- **Case A — the model's output had completed** (the turn's `tool_call`s were committed): finished tool results are re-sent as structured `tool_call_output`s; unfinished calls get an `[interrupted: tool aborted by user]` placeholder, keeping `tool_call`/output pairing strictly intact;
- **Case B — the model's output was incomplete**: the whole turn is flattened into one `<turn_aborted>` user text carrying whatever partial output existed.
Carry-over enters the model context only — it is never written to the Trace, which records only what actually happened.
## Automatic reconnect
Only LLM-side `timeout` (network timeouts, rate limits, 5xx) and `malformed` (truncated streams, JSON parse failures) trigger an in-run reconnect: the engine re-sends the original input plus a `<turn_retried>` block carrying the previous partial output, so tools are never re-executed. Default limit is 2 reconnects with linear backoff (base 250ms); beyond that the turn settles as `failed`. Tool errors are never retried — they are fed back to the model as `tool_call_output` and the model decides what to do next.
## Compaction
Compaction settings are filled in from `system_config.yaml` by the composition layer:
```ts
interface CompactionSettings {
maxContextLength: number; // context-token threshold (last token_usage's request.total); <=0 disables
maxSessionTurns: number; // cumulative Session turn threshold (counted across Tasks); <=0 = unlimited
mode: "summarize" | "discard";
prompt: string; // the Prompt used by summarize compaction
}
```
Three triggers (`compaction_begin.reason`):
| reason | Condition |
| --- | --- |
| `context` | last turn's `token_usage.request.total` ≥ `maxContextLength` (default 128000) |
| `turns` | Session turn count ≥ `maxSessionTurns` (default -1 = unlimited) |
| `manual` | the user runs `/compact` or calls `session.compact()` |
Two modes: `summarize` (default) appends the compaction Prompt to the old context, extracts the `<summary>`, wraps it as a `<context_summary>` user text and continues in a **fresh model context**; `discard` simply drops the old context. Compaction rotates the [Trace file](/sessions-and-traces) (`_002`, `_003`, …) — one Trace file always equals one complete model context. `compactability()` probes feasibility before `session.compact()` (`ok | unsupported | empty | just_compacted`).
## Concurrency model
- Within a turn: approvals are sequential, execution is concurrent, and the next turn's input keeps the original order;
- within a Session: only one Task or one compaction runs at a time (the Server rejects concurrent requests with 409);
- a [Subagent](/tools) is an independent Session with its own Trace and loop; its messages are forwarded to the parent tagged with `origin`.
## Side channels
- **Session titles**: `session.generateTitle()` is a one-shot out-of-band LLM call (no tools, no system Prompt) that never enters history or Trace;
- **Usage accounting**: each turn's `token_usage` events are persisted row by row by the Server — the raw data behind the cost statistics.
+120
View File
@@ -0,0 +1,120 @@
---
title: Agent 运行循环
description: context_engine 的总体流程图与逐环节拆解——审批、并发工具执行、中断补发、自动重连与上下文压缩。
---
SDK 的唯一执行入口是 `session.run(newMessages, opts?)`:输入本次新增的 OmniMessage 列表(Prompt),返回一个异步生成器,流式产出 [OmniMessage](/omni-message)。一次 `run` 自动跑完一个完整的 Task,直到模型给出不含工具调用的最终答复。
本页先给出 context_engine 的总体流程,再逐环节拆解;逐条消息级的可见时序与顺序保证见[消息流转与时序](/message-flow)。源码:`packages/core/src/engine/context-engine.ts`。
## 总体流程
```text
session.run(newMessages, { approve, signal })
│ 存在上次中断的补发内容?→ 前置到本轮输入
▼
┌── 轮循环(≤ max_turns,默认 100)──────────────────────────────┐
│ │
│ request_begin │
│ LLM.streamGenerate(newMessages) │
│ ├─ 流式产出 partial_* 分片 + 完整消息(thinking/text/…) │
│ ├─ 每个完整 tool_call: │
│ │ approve(toolCall) ──deny──► 合成 aborted 输出 │
│ │ │allow (审批逐个;写审计事件) │
│ │ ▼ │
│ │ Environment.executeTool ──► 并发执行,输出流式回传 │
│ └─ LLMOutcome: │
│ timeout / malformed ──► 同轮自动重连(≤2 次, │
│ 附 <turn_retried>,工具不重跑) │
│ token_usage + request_end(LLM 流结束即产出,不等工具) │
│ │
│ 工具输出按原始调用顺序重排 ──► 作为下一轮输入 │
│ 本轮无 tool_call?──► Task 结束,run 返回 │
│ 压缩触发(context/turns)?──► summarize/discard + Trace 轮转 │
└───────────────────────────────────────────────────────────────┘
signal 中断(任意时刻)──► 产出 abort 事件 + 构造补发内容 ──► run 返回
```
全程的每条消息与事件同时流向两个去处:实时输出给 Human,以及写入 [Trace](/sessions-and-traces)。
## 输入与输出
```ts
const agent = await createAgent({ agentId: "default_agent" });
const session = await agent.createSession({ workspaceDir: process.cwd() });
for await (const output of session.run([userText("整理 data/ 下的 CSV 文件")], {
approve: async (toolCall) => "allow",
signal: abortController.signal,
})) {
// output: partial_* 分片、完整 model_msg、event_msg
}
```
```ts
interface RunOptions {
signal?: AbortSignal; // 中断信号(如 Ctrl-C)
approve?: ApproveFn; // 逐工具审批;未注入时默认全部拒绝(保守策略)
}
```
## 一轮(Turn)的生命周期
Task 由若干连续的 Request(轮)组成,每轮:
1. 产出 `request_begin`;
2. LLM 流式返回:`partial_*` 分片与完整消息依次产出;
3. 每个完整的 `tool_call` 恰好触发一次 `approve` 回调,决策以 `approval_decision` 事件记录;
4. 通过审批的工具交给 Environment **并发执行**(审批本身逐个进行),输出按完成顺序流出;
5. LLM 流结束时,先产出其最后一条 `token_usage`,随即产出 `request_end(status)`——**不等待工具**,仍在执行的工具输出可出现在 `request_end` 之后;
6. 整批工具全部到达终态后,工具结果**按原始调用顺序**重排,作为下一轮输入——在此之前不会发起下一次 Request。
某轮不再产生 `tool_call` 时,Task 结束。拒绝(deny)会生成一条合成的 `aborted` 工具输出(内容为 `Tool call denied by user.`),模型据此继续。
## 中断与补发(carry-over)
`signal` 触发中断后,引擎产出 `abort` 事件并立即返回,同时为下一次 `run` 构造补发内容:
- **场景 A:模型输出已完成**(该轮 `tool_call` 已提交)——已完成的工具结果按结构化 `tool_call_output` 补发;未执行完的调用补上 `[interrupted: tool aborted by user]` 占位,保证 `tool_call` 与输出严格配对;
- **场景 B:模型输出未完成**——整轮压平为一段 `<turn_aborted>` 用户文本,携带已产生的部分输出。
补发内容只进入模型上下文,不写入 Trace——Trace 永远只记录真实发生的消息。
## 自动重连
只有 LLM 侧的 `timeout`(网络超时、限流、5xx)与 `malformed`(流截断、JSON 解析失败)会触发引擎内自动重连:同一次 `run` 内重发原始输入,并附加 `<turn_retried>` 块携带上一次的部分输出,避免工具重复执行。默认最多重连 2 次,线性退避(基数 250ms);超限后该轮以 `failed` 收场。工具错误从不重试——它们作为 `tool_call_output` 反馈给模型,由模型决定下一步。
## 上下文压缩(Compaction)
压缩配置由组装层从 `system_config.yaml` 填充默认值:
```ts
interface CompactionSettings {
maxContextLength: number; // 上下文 Token 阈值(取最近一次 token_usage 的 request.total);<=0 关闭
maxSessionTurns: number; // Session 累计轮数阈值(跨 Task 计数);<=0 不限制
mode: "summarize" | "discard";
prompt: string; // summarize 模式使用的压缩 Prompt
}
```
三种触发方式(`compaction_begin.reason`):
| reason | 触发条件 |
| --- | --- |
| `context` | 上一轮 `token_usage.request.total` ≥ `maxContextLength`(默认 128000) |
| `turns` | Session 轮数 ≥ `maxSessionTurns`(默认 -1,即不限) |
| `manual` | 用户执行 `/compact` 或调用 `session.compact()` |
两种模式:`summarize`(默认)向旧上下文追加压缩 Prompt,提取 `<summary>` 后包装为 `<context_summary>` 用户文本,在**全新的模型上下文**中继续;`discard` 直接丢弃旧上下文。压缩时 [Trace 文件随之轮转](/sessions-and-traces)(`_002`、`_003`……),一个 Trace 文件恒等于一个完整模型上下文。`session.compact()` 前可用 `compactability()` 探询可行性(`ok | unsupported | empty | just_compacted`)。
## 并发模型
- 同一轮内:审批逐个、执行并发、下一轮输入按原始顺序;
- 同一 Session:同时只有一个 Task 或一次压缩在运行(Server 侧以 409 拒绝并发请求);
- [Subagent](/tools) 是独立 Session,拥有自己的 Trace 与运行循环,消息以 `origin` 标记转发给父级。
## 相关旁路
- **Session 标题**:`session.generateTitle()` 走独立的一次性 LLM 调用(无工具、无系统 Prompt),不进入历史与 Trace;
- **用量落账**:每轮的 `token_usage` 事件被 Server 逐条入库,构成成本统计的原始数据。
+135
View File
@@ -0,0 +1,135 @@
---
title: Architecture
description: How the three-interface boundary, the context_engine and OmniMessage organize the SDK, CLI, Server and Web App into one system.
---
PenguinHarness is a pnpm monorepo whose center is the execution engine in `@prismshadow/penguin-core`; the CLI, the Server and the Web App are just different "Human implementations" of that same engine.
## Layers
```text
┌─────────────┐ ┌─────────────────────────────┐
│ CLI │ │ Web App (React SPA) │
│ (penguin) │ │ ↑ OmniMessage over SSE │
│ │ │ Server (Hono + SQLite) │
└──────┬──────┘ └──────────────┬──────────────┘
│ session.run(...) │ ← Human boundary
┌──────┴────────────────────────┴──────────────┐
│ core: context_engine (ReAct loop) │
│ ├── LLMInterface ──→ AgentHub ──→ models │
│ ├── EnvironmentInterface ──→ builtin tools│
│ ├── Agent State (editable files) │
│ └── Trace (append-only JSONL) │
└──────────────────────────────────────────────┘
```
| Package | Role |
| --- | --- |
| `packages/core` | SDK and engine: context_engine, OmniMessage, LLM/Environment interfaces, State and Trace |
| `packages/cli` | Terminal Human implementation: REPL and one-shot runs, embeds core in-process |
| `packages/server` | Web Human implementation: HTTP for input and approvals, SSE for the output stream |
| `packages/web` | Rendering SPA: streams by the OmniMessage protocol, contains no engine logic |
| `packages/skills` | The built-in skill library (a set of `SKILL.md` files) |
## Division of responsibilities
To place a design in a layer, ask where its **source of truth** lives. The four layers split as:
| Layer | Owns | Does not own |
| --- | --- | --- |
| SDK (`core`) | Protocol and execution — everything that makes messages flow | persisted user state, multi-user, any rendering |
| Server | The resident process and the multi-user runtime | engine logic (fully delegated to the SDK) |
| File layer (`~/.penguin/data`) | Everything editable and everything recorded | any computation |
| CLI / Web | Rendering and interaction | business state |
Item by item (design → owner → carrying file or module):
| Design | Owner | Carried by |
| --- | --- | --- |
| The OmniMessage protocol, message parsing and partial aggregation | SDK | `core/src/omnimessage/` — see [The OmniMessage Protocol](/omni-message) |
| The ReAct loop, carry-over, reconnect, compaction | SDK | `core/src/engine/context-engine.ts` — see [The Agent Loop](/agent-loop) |
| The approval mechanism (one decision per tool_call) | SDK | `ApproveFn` (`core/src/interfaces.ts`); the concrete mode is injected by CLI/Server |
| Tool execution and centralized close-out | SDK | `core/src/environment/` — see [Tools & Approval](/tools) |
| Model access (provider protocol adaptation) | SDK → AgentHub | `core/src/llm/` + `@prismshadow/agenthub` — see [Models & Providers](/models) |
| Trace writing and Session-recovery logic | SDK | `core/src/trace/` (the records themselves live in the file layer) |
| Subagent spawning and message forwarding | SDK | the `run_subagent` tool + the injected `SubagentRunner` |
| Multi-user auth and Project authorization | Server | `server/src/auth/`, `server/src/services/project-service.ts` |
| Session indexing, per-Session mutex, SSE forwarding | Server | `server/src/runtime/` — see [Server API](/server-api) |
| Scheduled tasks (execution) | Server | `server/src/runtime/scheduler.ts`; the task definitions live in the file layer at `agent_state/schedule/*.toml` |
| Approval-mode persistence and manual decisions | Server | `server/src/runtime/approvals.ts` + SQLite |
| Usage persistence and cost statistics | Server | `server/src/runtime/usage-recorder.ts`, `services/usage-service.ts` |
| Agent behavior definition (prompts, runtime params) | File layer | `agent_state/system_config.yaml`, `AGENTS.md` — see the [Configuration Reference](/configuration) |
| Skills | File layer | `agent_state/skills/<name>/SKILL.md` — see [Skills](/skills) |
| Secrets | File layer | Vault: `agent_state/.vault.toml`; model credentials: `.project_config.toml` (both 0600) |
| The model table and the default model | File layer | `<project>/.project_config.toml` |
| Run history (the sole source of truth for recovery) | File layer | `traces/<date>/<session>_<index>.jsonl` — see [Sessions & Traces](/sessions-and-traces) |
| Benchmark cases and scores | File layer | `benchmarks/<id>/` — see [Self-Improvement](/self-improvement) |
| Snapshots | File layer | `snapshots/v<version>.tar.gz`; the export/import service lives in the Server |
| Streaming rendering, approval UI, charts | CLI / Web | `cli/src`, `web/src` (pure rendering, no engine logic) |
The one-line rule: **what is editable or recorded lives in files; what makes messages flow lives in the SDK; what needs a resident process and multiple users lives in the Server; the rest is rendering.** The Server's SQLite stores only indexes and aggregates — it never competes with the file layer as a source of truth.
## Source layout
How each package is organized (single-purpose files, split by layer; every file's header comment is its design note):
```text
packages/
├── core/src
│ ├── agent.ts / session.ts # the createAgent composition layer and Session (run / compact / generateTitle)
│ ├── session-title.ts # one-shot title generation (out-of-band LLM call, never in Trace)
│ ├── engine/context-engine.ts # ReAct loop orchestration: turn lifecycle, approvals, carry-over, reconnect, compaction
│ ├── omnimessage/ # types.ts protocol types · builders.ts constructors · aggregate.ts partial aggregation
│ ├── llm/ # generative-model.ts AgentHub adapter · tool-call-ids.ts id uniqueness
│ ├── environment/ # environment.ts execution close-out · tools/ registry, 6 builtin tools, background sessions
│ ├── state/ # paths · default-config · project-config · model-catalog
│ │ # agent-state (Skill install, prompt assembly) · agent-vault · builtin-agents
│ ├── trace/ # writer.ts append-only JSONL · resume.ts replay-based recovery
│ └── internal/ # date and Session helpers
├── cli/src # commander entry + run / chat / config / serve commands and approval prompts
├── server/src # app assembly · db (node:sqlite) · auth · http/routes · runtime · services
├── web/src # api client · state · lib/omni stream rendering · components · feature pages
├── skills/ # loader + the skills/<name>/SKILL.md library
├── landing/ # the product landing page (with the blog)
└── docs/ # this documentation site
```
The internals of server and web are detailed on [Server API](/server-api) and the [Web App Guide](/web-app).
## The three-interface boundary
The context_engine is the heart of the system and does exactly two things: it maintains the linear message history, and it orchestrates the event flow between three interfaces. It speaks only [OmniMessage](/omni-message) and performs no protocol conversion:
- **Human** — the user-side boundary. It is deliberately not an interface class: the SDK's single entry point `session.run(newMessages, { approve, signal })` *is* the Human boundary. Input is a list of new OmniMessages plus an approval callback; output is streamed OmniMessages. The CLI and the Server are its two shipped implementations.
- **LLM** — the model-side interface (`LLMInterface`). Translates OmniMessage to requests against the AgentHub model gateway and streamed events back into OmniMessage. All provider protocol adaptation happens inside AgentHub; core never imports a vendor SDK.
- **Environment** — the tool-execution interface (`EnvironmentInterface`). Runs approved tool calls and streams results back.
Why this boundary matters: the kernel contains no provider, tool or UI specifics, so each side swaps by configuration (local shell today, other sandboxes tomorrow; CLI, Web, or programmatic callers) without touching the core. See [Core Interfaces](/interfaces) for the signatures.
## Data flow of one Task
1. Human hands a Prompt (a list of OmniMessages) to `session.run`;
2. the engine issues a Request: the LLMInterface streams `partial_*` fragments and complete messages;
3. every complete `tool_call` triggers one `approve` decision; approved calls run concurrently in the Environment;
4. tool outputs are re-fed in their original order as the next Request's input;
5. the Task ends when a turn produces no `tool_call` (the final answer).
Every message and event flows to two destinations at once: streamed live to the Human, and appended to the [Trace](/sessions-and-traces). Loop details (interruption, reconnect, compaction) are on [The Agent Loop](/agent-loop).
## The state layer
Below the engine sits a purely file-based state layer rooted at `~/.penguin/data` (override with `PENGUIN_HOME`), organized as `<project>/agents/<agent>/`:
- **Agent State** — the `agent_state/` directory: `system_config.yaml`, `AGENTS.md`, Skills, Vault. An Agent's entire behavior is editable files.
- **Project config** — `.project_config.toml`: the model table and credentials; model identity is always the `(provider, model_id)` pair.
- **Trace** — the `traces/` directory: append-only JSONL, the single source of truth for Session recovery.
The Server keeps an additional SQLite index (users, authorization, usage stats) but never duplicates the file layer's facts — the CLI, SDK and Web share one data directory and can be mixed freely.
## Key design decisions
- **One protocol, three jobs**: OmniMessage is simultaneously the SDK's external interface, the Trace on-disk format and the engine's internal currency — what streams, what is stored and what the model sees are the same thing.
- **Errors converge into messages**: the LLM and Environment never throw into the engine; results carry a five-value `stop_reason` (`completed | failed | aborted | timeout | malformed`), and only LLM-side `timeout / malformed` trigger an in-run reconnect.
- **A thin model layer**: core defines only `LLMInterface`; provider adaptation lives entirely in AgentHub (`@prismshadow/agenthub`), which is what makes any OpenAI-compatible endpoint reachable. See [Models & Providers](/models).
Source entry points: `packages/core/src/engine/context-engine.ts`, `packages/core/src/interfaces.ts`.
+135
View File
@@ -0,0 +1,135 @@
---
title: 架构总览
description: 三接口边界、context_engine 与 OmniMessage 如何把 SDK、CLI、Server、Web 组织成一个系统。
---
PenguinHarness 是一个 pnpm monorepo,核心是 `@prismshadow/penguin-core` 中的执行引擎;CLI、Server 与 Web App 都只是这同一个引擎的不同「Human 实现」。
## 分层结构
```text
┌─────────────┐ ┌─────────────────────────────┐
│ CLI │ │ Web App (React SPA) │
│ (penguin) │ │ ↑ OmniMessage over SSE │
│ │ │ Server (Hono + SQLite) │
└──────┬──────┘ └──────────────┬──────────────┘
│ session.run(...) │ ← Human 边界
┌──────┴────────────────────────┴──────────────┐
│ core: context_engine(ReAct 循环) │
│ ├── LLMInterface ──→ AgentHub ──→ 各模型 │
│ ├── EnvironmentInterface ──→ 内置工具 │
│ ├── Agent State(可编辑文件) │
│ └── Trace(追加式 JSONL) │
└──────────────────────────────────────────────┘
```
| 包 | 角色 |
| --- | --- |
| `packages/core` | SDK 与执行引擎:context_engine、OmniMessage、LLM/Environment 接口、State 与 Trace |
| `packages/cli` | 终端 Human 实现:REPL 与单次运行,直接内嵌 core |
| `packages/server` | Web Human 实现:HTTP 承接输入与审批,SSE 推送输出流 |
| `packages/web` | 渲染层 SPA:按 OmniMessage 协议流式渲染,不含业务引擎 |
| `packages/skills` | 内置技能库(`SKILL.md` 文件集合) |
## 职责划分
判断一个设计归属哪一层,只看它的**事实来源**在哪里。四层的分工:
| 层 | 承担 | 不承担 |
| --- | --- | --- |
| SDK(`core`) | 协议与执行:让消息流动起来的一切 | 持久化用户态、多用户、任何渲染 |
| Server | 常驻进程与多用户运行时 | 引擎逻辑(全部委托给 SDK) |
| 文件层(`~/.penguin/data`) | 一切可编辑的定义与一切被记录的历史 | 任何计算 |
| CLI / Web | 渲染与交互 | 业务状态 |
逐项对应(设计 → 归属 → 承载文件或模块):
| 设计 | 归属 | 承载 |
| --- | --- | --- |
| OmniMessage 协议、消息解析与分片聚合 | SDK | `core/src/omnimessage/`,见 [OmniMessage 协议](/omni-message) |
| ReAct 循环、补发、重连、压缩 | SDK | `core/src/engine/context-engine.ts`,见 [Agent 运行循环](/agent-loop) |
| 审批机制(每个 tool_call 一次决策) | SDK | `ApproveFn`(`core/src/interfaces.ts`);具体模式由 CLI/Server 注入 |
| 工具执行与统一收尾 | SDK | `core/src/environment/`,见[工具与审批](/tools) |
| 模型接入(Provider 协议适配) | SDK → AgentHub | `core/src/llm/` + `@prismshadow/agenthub`,见[模型与 Provider](/models) |
| Trace 写入与 Session 恢复逻辑 | SDK | `core/src/trace/`(记录本体在文件层) |
| Subagent 派生与消息回流 | SDK | `run_subagent` 工具 + `SubagentRunner` 注入 |
| 多用户认证与 Project 授权 | Server | `server/src/auth/`、`server/src/services/project-service.ts` |
| Session 索引、并发互斥、SSE 转发 | Server | `server/src/runtime/`,见 [Server API](/server-api) |
| 定时任务(Schedule 执行) | Server | `server/src/runtime/scheduler.ts`;任务定义在文件层 `agent_state/schedule/*.toml` |
| 审批模式持久化与人工决策 | Server | `server/src/runtime/approvals.ts` + SQLite |
| 用量落库与成本统计 | Server | `server/src/runtime/usage-recorder.ts`、`services/usage-service.ts` |
| Agent 行为定义(Prompt、运行参数) | 文件层 | `agent_state/system_config.yaml`、`AGENTS.md`,见[配置参考](/configuration) |
| Skill | 文件层 | `agent_state/skills/<name>/SKILL.md`,见[技能系统](/skills) |
| 密钥 | 文件层 | Vault:`agent_state/.vault.toml`;模型凭据:`.project_config.toml`(均 0600) |
| 模型表与默认模型 | 文件层 | `<project>/.project_config.toml` |
| 运行历史(恢复的唯一事实来源) | 文件层 | `traces/<date>/<session>_<index>.jsonl`,见 [Session 与 Trace](/sessions-and-traces) |
| Benchmark 题库与评分 | 文件层 | `benchmarks/<id>/`,见[自我进化](/self-improvement) |
| 快照 | 文件层 | `snapshots/v<version>.tar.gz`;导入导出服务在 Server |
| 流式渲染、审批 UI、统计图表 | CLI / Web | `cli/src`、`web/src`(纯渲染,不含引擎逻辑) |
一句话判定:**能编辑的与被记录的在文件层;让消息流动起来的在 SDK;需要常驻进程与多用户的在 Server;其余是渲染。**Server 的 SQLite 只存索引与聚合,从不与文件层争当事实来源。
## 源码结构
各包的目录设计(职责单一、按层拆分;文件头注释即该文件的设计说明):
```text
packages/
├── core/src
│ ├── agent.ts / session.ts # createAgent 组装层与 Session(run / compact / generateTitle)
│ ├── session-title.ts # 一次性标题生成(旁路 LLM 调用,不入 Trace)
│ ├── engine/context-engine.ts # ReAct 循环编排:轮生命周期、审批、补发、重连、压缩
│ ├── omnimessage/ # types.ts 协议类型 · builders.ts 构造函数 · aggregate.ts 分片聚合
│ ├── llm/ # generative-model.ts AgentHub 适配 · tool-call-ids.ts id 唯一化
│ ├── environment/ # environment.ts 执行与收尾 · tools/ 注册表、6 个内置工具、后台会话
│ ├── state/ # paths · default-config · project-config · model-catalog
│ │ # agent-state(Skill 安装、提示词装配)· agent-vault · builtin-agents
│ ├── trace/ # writer.ts 追加式 JSONL · resume.ts 回放恢复
│ └── internal/ # 日期与 Session 辅助
├── cli/src # commander 入口 + run / chat / config / serve 命令与审批交互
├── server/src # app 组装 · db(node:sqlite)· auth · http/routes · runtime · services
├── web/src # api 客户端 · state · lib/omni 流渲染 · components · features 各页面
├── skills/ # 加载器 + skills/<name>/SKILL.md 技能库
├── landing/ # 产品落地页(含博客)
└── docs/ # 本文档站
```
server 与 web 的内部结构分别见 [Server API](/server-api) 与 [Web App 指南](/web-app)。
## 三接口边界
context_engine 是整个系统的核心,它只做两件事:维护线性消息历史,以及在三个接口之间编排事件流。它只认识 [OmniMessage](/omni-message),不做任何协议转换:
- **Human**——用户侧边界。它不是一个接口类:SDK 的唯一入口 `session.run(newMessages, { approve, signal })` 就是 Human 边界本身。输入是新增的 OmniMessage 列表与审批回调,输出是流式 OmniMessage。CLI 与 Server 是它的两种实现形态。
- **LLM**——模型侧接口(`LLMInterface`)。把 OmniMessage 翻译为模型网关 AgentHub 的请求,把流式事件翻译回 OmniMessage。所有 Provider 协议适配都在 AgentHub 内完成,core 不直接依赖任何模型厂商 SDK。
- **Environment**——工具执行接口(`EnvironmentInterface`)。执行通过审批的工具调用,把结果以流式 OmniMessage 送回。
这一边界设计的意义:引擎内核不含任何 Provider、工具或 UI 细节,三侧实现均可按配置替换(本地 shell、其他执行沙箱;CLI、Web、程序化调用),而互不影响。接口签名详见[接口契约](/interfaces)。
## 一个 Task 的数据流
1. Human 把 Prompt(OmniMessage 列表)交给 `session.run`;
2. 引擎发起一次 Request:经 LLMInterface 流式产出 `partial_*` 与完整消息;
3. 每个完整的 `tool_call` 触发一次 `approve` 审批;通过后交 Environment 并发执行;
4. 工具输出按原始顺序回填,进入下一轮 Request;
5. 某轮不再产生 `tool_call`(最终答复)时 Task 结束。
全程的每条消息与事件同时流向两个去处:实时输出给 Human,以及追加写入 [Trace](/sessions-and-traces)。运行循环的细节(中断、重连、压缩)见 [Agent 运行循环](/agent-loop)。
## 状态层
引擎之下是纯文件的状态层,数据根目录为 `~/.penguin/data`(`PENGUIN_HOME` 可改),按 `<project>/agents/<agent>/` 组织:
- **Agent State**——`agent_state/` 目录:`system_config.yaml`、`AGENTS.md`、Skills、Vault。Agent 的全部行为定义都是可编辑文件。
- **Project 配置**——`.project_config.toml`:模型表与凭据,模型身份恒为 `(provider, model_id)` 二元组。
- **Trace**——`traces/` 目录:追加式 JSONL,恢复 Session 的唯一事实来源。
Server 额外维护一个 SQLite 索引库(用户、授权、用量统计),但从不复制文件层的事实——CLI、SDK 与 Web 共用同一份数据目录,可以混用。
## 关键设计决策
- **一个协议,三种职责**:OmniMessage 同时是 SDK 对外接口、Trace 落盘格式与引擎内部通货——「流出去的」「存下来的」「模型看到的」是同一种东西。
- **错误收敛为消息**:LLM 与 Environment 从不向引擎抛异常;结果携带五值 `stop_reason`(`completed | failed | aborted | timeout | malformed`),仅 LLM 侧的 `timeout / malformed` 触发引擎内重连。
- **薄模型层**:core 只定义 `LLMInterface`,Provider 适配全部下沉到 AgentHub(`@prismshadow/agenthub`),因此支持任意 OpenAI 兼容端点,见[模型与 Provider](/models)。
源码入口:`packages/core/src/engine/context-engine.ts`、`packages/core/src/interfaces.ts`。
+145
View File
@@ -0,0 +1,145 @@
---
title: CLI Reference
description: Complete reference for the penguin command, its subcommands, and options.
---
The CLI ships as the npm package `@prismshadow/penguin-cli`; the command is `penguin`. Running bare `penguin` prints help; `-v, --version` prints the version. A `.env` file in the working directory is loaded automatically on startup.
## Global conventions
- Model references: a model's identity is always the `(provider, model_id)` pair. `--model-id` takes the upstream model id and pairs with `--provider`. When `run` / `chat` omit `--provider`, the `--model-id` matches only if it is globally unique in the configuration; ambiguity is an error.
- Data root: `--root <dir>` overrides the data root directory. Priority: `--root` > the `PENGUIN_HOME` env var > `~/.penguin/data`.
## penguin run
Send a single message, execute one Task, then exit. If the Task aborted, the exit code is non-zero, so scripts / CI can check it.
```bash
penguin run -m "Summarize the code structure of this directory"
```
| Option | Description |
| --- | --- |
| `-m, --message <message>` | Required; the message to send |
| `--model-id <id>` | Model to use; defaults to the Project's default model |
| `--provider <group>` | Provider group of the model |
| `--project-id <id>` | Project to use |
| `--agent-id <id>` | Agent to use |
| `--workspace <path>` | Workspace directory; defaults to the current directory and must exist |
| `--approve <mode>` | Approval mode, see below |
## penguin chat
Interactive REPL; each input line starts a Task. Takes the same options as `run` (minus `-m, --message`), plus:
| Option | Description |
| --- | --- |
| `--resume [sessionId]` | Resume a Session; without an id, resumes the Agent's latest Session |
With `--resume`, the Workspace and model are locked by the original Session and cannot be overridden via `--workspace` / `--model-id` / `--provider`. On exit, a copy-pastable `penguin chat --resume <sessionId>` command is printed.
In-REPL commands:
| Input | Behavior |
| --- | --- |
| `/compact` | Proactively compact the current context |
| `/exit`, `/quit` | Quit |
Ctrl-C is state-dependent:
| State | Behavior |
| --- | --- |
| Awaiting tool approval | Deny that tool call |
| Task running | Abort the current Task and return to input |
| Input buffer non-empty | Clear the current input |
| Idle with empty buffer | Show an exit confirmation (y/N) |
## Approval modes (--approve)
| Mode | Behavior |
| --- | --- |
| `allow-all` | Auto-approve every tool call (default) |
| `deny-all` | Auto-reject every tool call |
| `read-only` | Auto-approve read-only tools; prompt for the rest |
| `always-ask` | Prompt for every tool call |
At an interactive prompt, `y` / `yes` approves and `n` / `no` denies; a bare Enter defaults to approve.
## penguin config
Manages a Project's model configuration, per-Agent vault environment variables, and the UI language. Except for `lang`, all subcommands below accept `--project-id <id>` (defaults to the default Project) and `--root <dir>`.
### model add
Add or update a model entry:
```bash
penguin config model add --model-id deepseek-v4-pro --api-key sk-... --set-default
```
| Option | Description |
| --- | --- |
| `--model-id <id>` | Required; the upstream model id |
| `--provider <group>` | Provider group; inferred from the built-in catalog when omitted |
| `--api-key <key>` | API key, stored inline in the Project's hidden `.project_config.toml` |
| `--base-url <url>` | Custom endpoint base URL |
| `--context-window <n>` | Context window size |
| `--client-type <type>` | Client protocol type |
| `--vision` / `--no-vision` | Mark vision input as supported / unsupported |
| `--price-cache-read <n>` | Cache-read price |
| `--price-cache-write <n>` | Cache-write price |
| `--price-output <n>` | Output price |
| `--set-default` | Also set as the default model |
### model default / model vision / model list
```bash
penguin config model default --model-id <id> --provider <group>
penguin config model vision --model-id <id> --provider <group>
penguin config model list
```
- `model default` sets the Project's default model; `model vision` sets the vision proxy model. Both require `--model-id` and `--provider`, and the reference must already exist in the model list.
- `model list` lists configured models; the default model is marked with `*`.
### vault
Per-Agent environment variable store, written to `agent_state/.vault.toml`. Values are injected into tool subprocess environments only — never into the model context.
```bash
penguin config vault set --key GITHUB_TOKEN --value ghp_xxx
penguin config vault list
penguin config vault remove --key GITHUB_TOKEN
```
| Subcommand | Options |
| --- | --- |
| `vault set` | `--key <name>` (required), `--value <value>` (required), `[--agent-id <id>]` |
| `vault list` | `[--agent-id <id>]` |
| `vault remove` | `--key <name>` (required), `[--agent-id <id>]` |
### lang
```bash
penguin config lang en
```
Sets the CLI UI language (`en` or `zh`) by writing `PENGUIN_LANG` into the shell startup file.
## penguin server / penguin web
Two entry points into the same service process: `server` runs headless; `web` additionally waits for readiness, prints the URL, and opens the browser.
```bash
penguin web
```
| Option | Description |
| --- | --- |
| `--port <port>` | Listen port, default 7364 |
| `--host <host>` | Listen host, default 127.0.0.1 |
| `--no-open` | `web` only: do not open the browser |
Port / host priority: command-line option > the `PORT` / `HOST` env vars (including `.env`) > defaults.
See also: [Configuration Reference](/configuration), [Models & Providers](/models).
+145
View File
@@ -0,0 +1,145 @@
---
title: CLI 参考
description: penguin 命令的子命令与选项完整参考。
---
CLI 由 npm 包 `@prismshadow/penguin-cli` 提供,命令为 `penguin`。不带子命令执行 `penguin` 时打印帮助;`-v, --version` 打印版本号。启动时自动加载工作目录下的 `.env`。
## 全局约定
- 模型引用:模型身份始终是 `(provider, model_id)` 二元组。`--model-id` 填上游模型 id,与 `--provider` 组成配对引用。`run` / `chat` 省略 `--provider` 时,仅当该 `--model-id` 在配置中全局唯一才会匹配,存在歧义则报错。
- 数据根目录:`--root <dir>` 覆盖数据根目录,优先级为 `--root` > 环境变量 `PENGUIN_HOME` > `~/.penguin/data`。
## penguin run
发送单条消息执行一个 Task,结束后退出;Task 被中止时以非零码退出,便于脚本 / CI 判断。
```bash
penguin run -m "总结当前目录的代码结构"
```
| 选项 | 说明 |
| --- | --- |
| `-m, --message <message>` | 必填,要发送的消息 |
| `--model-id <id>` | 指定模型,缺省使用 Project 默认模型 |
| `--provider <group>` | 模型所属 Provider 分组 |
| `--project-id <id>` | 指定 Project |
| `--agent-id <id>` | 指定 Agent |
| `--workspace <path>` | Workspace 目录,默认当前目录,必须已存在 |
| `--approve <mode>` | 审批模式,见下文 |
## penguin chat
交互式 REPL,每输入一行发起一个 Task。选项与 `run` 相同(除 `-m, --message` 外),另加:
| 选项 | 说明 |
| --- | --- |
| `--resume [sessionId]` | 恢复指定 Session;省略 id 时恢复该 Agent 最近的 Session |
使用 `--resume` 时,Workspace 与模型由原 Session 锁定,不可再用 `--workspace` / `--model-id` / `--provider` 覆盖。退出时会打印可直接复制的 `penguin chat --resume <sessionId>` 命令。
REPL 内命令:
| 输入 | 行为 |
| --- | --- |
| `/compact` | 主动压缩当前上下文 |
| `/exit`、`/quit` | 退出 |
Ctrl-C 的行为依状态而定:
| 状态 | 行为 |
| --- | --- |
| 等待工具审批 | 拒绝该次工具调用 |
| Task 运行中 | 中断当前 Task,返回输入 |
| 输入缓冲非空 | 清空当前输入 |
| 空闲且缓冲为空 | 显示退出确认(y/N) |
## 审批模式(--approve)
| 模式 | 行为 |
| --- | --- |
| `allow-all` | 自动批准所有工具调用(默认) |
| `deny-all` | 自动拒绝所有工具调用 |
| `read-only` | 自动批准只读工具,其余逐个询问 |
| `always-ask` | 每次工具调用都询问 |
交互询问时输入 `y` / `yes` 批准、`n` / `no` 拒绝;直接回车默认为批准。
## penguin config
管理 Project 的模型配置、Agent 级 vault 环境变量与界面语言。除 `lang` 外,以下子命令均支持 `--project-id <id>`(缺省为默认 Project)与 `--root <dir>`。
### model add
新增或更新模型条目:
```bash
penguin config model add --model-id deepseek-v4-pro --api-key sk-... --set-default
```
| 选项 | 说明 |
| --- | --- |
| `--model-id <id>` | 必填,上游模型 id |
| `--provider <group>` | Provider 分组,缺省时根据内置目录推断 |
| `--api-key <key>` | API Key,内联存入 Project 隐藏文件 `.project_config.toml` |
| `--base-url <url>` | 自定义接口地址 |
| `--context-window <n>` | 上下文窗口大小 |
| `--client-type <type>` | 客户端协议类型 |
| `--vision` / `--no-vision` | 标记是否支持视觉输入 |
| `--price-cache-read <n>` | 缓存读价格 |
| `--price-cache-write <n>` | 缓存写价格 |
| `--price-output <n>` | 输出价格 |
| `--set-default` | 同时设为默认模型 |
### model default / model vision / model list
```bash
penguin config model default --model-id <id> --provider <group>
penguin config model vision --model-id <id> --provider <group>
penguin config model list
```
- `model default` 设置 Project 默认模型;`model vision` 设置视觉代理模型。两者的 `--model-id` 与 `--provider` 均为必填,且引用必须已存在于模型列表。
- `model list` 列出已配置模型,默认模型以 `*` 标记。
### vault
按 Agent 存储环境变量,写入 `agent_state/.vault.toml`;值只注入工具子进程的环境变量,绝不进入模型上下文。
```bash
penguin config vault set --key GITHUB_TOKEN --value ghp_xxx
penguin config vault list
penguin config vault remove --key GITHUB_TOKEN
```
| 子命令 | 选项 |
| --- | --- |
| `vault set` | `--key <name>`(必填)、`--value <value>`(必填)、`[--agent-id <id>]` |
| `vault list` | `[--agent-id <id>]` |
| `vault remove` | `--key <name>`(必填)、`[--agent-id <id>]` |
### lang
```bash
penguin config lang zh
```
设置 CLI 界面语言(`en` 或 `zh`),将 `PENGUIN_LANG` 写入 shell 启动文件。
## penguin server / penguin web
两者是同一服务进程的两个入口:`server` 为 headless 模式;`web` 额外等待服务就绪、打印 URL 并打开浏览器。
```bash
penguin web
```
| 选项 | 说明 |
| --- | --- |
| `--port <port>` | 监听端口,默认 7364 |
| `--host <host>` | 监听地址,默认 127.0.0.1 |
| `--no-open` | 仅 `web`:不自动打开浏览器 |
端口 / 地址优先级:命令行选项 > 环境变量 `PORT` / `HOST`(含 `.env`)> 默认值。
相关文档:[配置参考](/configuration)、[模型与 Provider](/models)。
+186
View File
@@ -0,0 +1,186 @@
---
title: Configuration Reference
description: Complete field reference for environment variables, Project config, Agent config, the Vault, and Schedules.
---
PenguinHarness configuration has three layers: environment variables shape the deployment, the Project config manages models and credentials, and the Agent config defines a single Agent's behavior. Each Agent additionally has two kinds of state files: the Vault (private environment variables) and Schedules (timed tasks).
## Environment variables
The CLI and the server automatically load a `.env` file from the working directory on startup.
| Variable | Description | Default |
| --- | --- | --- |
| `PENGUIN_HOME` | Data root directory | `~/.penguin/data` |
| `PORT` | Web service listen port | `7364` |
| `HOST` | Web service listen address | `127.0.0.1` |
| `PENGUIN_WEB_DB` | Server SQLite database path | `<root>/web.db` |
| `PENGUIN_WEB_DIST` | Front-end static assets directory | the npm server package falls back to its bundled web-dist |
| `PENGUIN_LANG` | CLI language (`en` / `zh`), set via `penguin config lang` | `en` |
### Provider credential variables
When a model entry has no inline `api_key`, the AgentHub gateway falls back to the provider's environment variable; the `*_BASE_URL` variants override the base URL the same way:
| Provider | API key | Base URL |
| --- | --- | --- |
| deepseek | `DEEPSEEK_API_KEY` | `DEEPSEEK_BASE_URL` |
| anthropic | `ANTHROPIC_API_KEY` | `ANTHROPIC_BASE_URL` |
| openai, openrouter, siliconflow, custom | `OPENAI_API_KEY` | `OPENAI_BASE_URL` |
| google | `GEMINI_API_KEY` | `GEMINI_BASE_URL` |
| zhipu | `ZAI_API_KEY` | `ZAI_BASE_URL` |
| moonshot | `MOONSHOT_API_KEY` | `MOONSHOT_BASE_URL` |
The openrouter, siliconflow, and custom groups speak the OpenAI-compatible protocol, hence the shared `OPENAI_*` variables. Provider groups and the built-in model catalog are covered in [Models & Providers](/models).
## Project config
`<root>/<project>/.project_config.toml` is the Project's single config file: a hidden file written with mode 0600, with credentials inlined on the model entries. Model identity is always the `(provider, model_id)` pair — string concatenation is forbidden everywhere.
| Field | Description |
| --- | --- |
| `name` | Project display name (the id is shown when unset) |
| `default_model` | Paired reference `{ provider, model_id }` to the default model; must point to an entry in `models` |
| `vision_model` | The vision model that reads images on behalf of text-only models (used by `describe_image`); a paired reference |
| `[[models]]` | The list of available model entries |
Model entry (`[[models]]`) fields:
| Field | Description |
| --- | --- |
| `provider` | Provider group; together with `model_id` forms the entry's unique key |
| `model_id` | Upstream request id, sent to AgentHub unchanged |
| `context_window` | Context window size |
| `client_type` | AgentHub client protocol; inferred from `model_id` by default — third-party OpenAI-compatible models should set `openai` |
| `display_name` | Display name; persisted only when it differs from the built-in catalog |
| `vision` | Whether image input is supported; defaults to supported |
| `pricing` | Three price buckets `cache_read` / `cache_write` / `output`, in USD per million Tokens (`unit = "usd_per_mtok"`) |
| `api_key` | Inline credential; when empty, falls back to the provider environment variable |
| `base_url` | Custom base URL; preset for gateway models |
| `created_at` | Write timestamp of `api_key` (ISO 8601; a display field maintained by the interface layer) |
```toml
default_model = { provider = "deepseek", model_id = "deepseek-v4-pro" }
[[models]]
provider = "deepseek"
model_id = "deepseek-v4-pro"
context_window = 1000000
vision = false
api_key = "sk-..."
[models.pricing]
unit = "usd_per_mtok"
cache_read = 0.003571
cache_write = 0.428571
output = 0.857143
```
`pricing.unit` is currently always `usd_per_mtok` (USD per million tokens); the three buckets map onto `token_usage`'s three counters.
Edit this file via the CLI (`penguin config model …`) or the Web Models page — never by hand while the service is running, and never by the model itself, which has no right to read or write it.
## Agent config
`agent_state/system_config.yaml` defines a single Agent's behavior (YAML; comments are preserved when edited via the Web UI):
| Field | Default | Description |
| --- | --- | --- |
| `name` | — | Agent display name (falls back to the id) |
| `description` | — | Agent description |
| `version` | `1` | Agent State version (a natural number), incremented on each successful optimization |
| `system_prompt` | built-in template | Required; the only template with placeholder substitution |
| `max_turns` | `100` | Maximum LLM turns per Task |
| `model.max_tokens` | `32000` | Output Token limit per Request |
| `model.thinking_level` | `medium` | `none` / `low` / `medium` / `high` / `xhigh` |
| `model.timeoutMs` | `120000` | Per-Request timeout (milliseconds) |
| `compaction.max_context_length` | `128000` | Context Token threshold that triggers compaction |
| `compaction.max_session_turns` | `-1` | Cumulative Session turn threshold (`-1` = unlimited) |
| `compaction.mode` | `summarize` | `summarize` / `discard` |
| `compaction.prompt` | built-in template | Prompt used for summarize compaction |
| `tools.builtin` | full default toolset when omitted | Tool entries: `name` / `description` / `parameters` / `permission` (`r` or `rw`) / `forModel` / `timeoutMs` / `maxOutputLength`; once written it replaces the default list wholesale |
| `tools.mcpServers` | `[]` | MCP Server configuration (`name` + `config`); reserved for the MCP adapter layer |
Tool permissions and approval semantics are covered in [Tools & Approval](/tools).
A partial-override example (edit the file the init step generated). Note that this file is **not deep-merged with the defaults**: a key you write out takes effect wholesale, and only omitted keys fall back to the defaults above at their use sites; `system_prompt` is required (loading refuses without it), so keep the full generated template when editing other fields:
```yaml
name: default_agent
description: General-purpose agent
version: 3
# Required: keep the full generated default template ({{AGENTS_MD}} and friends; elided here).
system_prompt: |
…
max_turns: 100
model:
max_tokens: 32000
thinking_level: medium
timeoutMs: 120000
compaction:
max_context_length: 128000
max_session_turns: -1
mode: summarize
# Omitting the whole tools section = the full default toolset. Writing tools.builtin
# REPLACES the default list wholesale: carry the complete definition (including the
# parameters JSON Schema) for every tool you keep — see Tools & Approval.
```
### System prompt placeholders
`system_prompt` is the only template with placeholder substitution. Available placeholders:
| Placeholder | Injected content |
| --- | --- |
| `{{AGENTS_MD}}` | Full text of `AGENTS.md` |
| `{{VAULT_KEYS}}` | List of Vault key names (names only) |
| `{{SKILL_METADATA}}` | Metadata of installed Skills |
| `{{PLATFORM}}` | Runtime platform |
| `{{OS_VERSION}}` | Operating system version |
| `{{DATE}}` | Current date |
| `{{CWD}}` | Workspace path |
| `{{AGENT_ID}}` | Agent id |
| `{{PROJECT_DIR}}` | Project directory |
| `{{SESSION_ID}}` | Session id |
`agent_state/AGENTS.md` is the developer-editable instruction file, injected via `{{AGENTS_MD}}` and empty by default — it is also the file an optimizer edits most (see [Self-Improvement](/self-improvement)).
## Vault
`agent_state/.vault.toml` is the Agent-level environment-variable vault: a hidden file written with mode 0600.
- Key names must match `^[A-Za-z_][A-Za-z0-9_]*$` (shell environment variable naming rules);
- Values are injected only into tool subprocess environments and never enter the model context or the Trace;
- Only key names are disclosed in the system prompt via `{{VAULT_KEYS}}`;
- Managed via `penguin config vault set/list/remove` or the Web Vault tab.
## Schedules
Each file `agent_state/schedule/<name>.toml` describes one scheduled task (the filename is its identity) that sends a preset Prompt to the Agent on a cadence. Schedules execute only while the Web service (the server runtime) is running, and are managed in the Web Agent settings → Schedule tab.
| Field | Required | Description |
| --- | --- | --- |
| `prompt` | yes | The Prompt sent on each trigger |
| `enabled` | no | Enabled switch; defaults to `false` |
| `start_at` | yes | First trigger time (ISO 8601) |
| `period` | no | Cadence such as `30m` / `12h` / `7d`, minimum 5 minutes; omitted means a one-shot task |
| `end_at` | no | End time; must be later than `start_at` |
| `session_id` | no | Bind to an existing Session; mutually exclusive with the three fields below |
| `workspace` | no | Workspace for new-Session mode |
| `provider` / `model_id` | no | Paired model reference for new-Session mode |
```toml
prompt = "Check yesterday's builds and summarize the failures"
enabled = true
start_at = 2026-08-01T09:00:00Z
period = "12h"
```
## Design principle
An Agent's behavior lives entirely in editable files on disk — prompts, Skills, and configuration are data, not code. That is what makes Agents improvable by Agents: an optimizer edits exactly the same files you edit by hand. See [Self-Improvement](/self-improvement) and the [CLI Reference](/cli).
+186
View File
@@ -0,0 +1,186 @@
---
title: 配置参考
description: 环境变量、Project 配置、Agent 配置、Vault 与定时任务的完整字段参考。
---
PenguinHarness 的配置分三层:环境变量决定部署形态,Project 配置管理模型与凭证,Agent 配置定义单个 Agent 的行为。此外每个 Agent 还有 Vault(私有环境变量)与 Schedule(定时任务)两类状态文件。
## 环境变量
CLI 与服务端启动时会自动加载工作目录下的 `.env` 文件。
| 变量 | 说明 | 缺省值 |
| --- | --- | --- |
| `PENGUIN_HOME` | 数据根目录 | `~/.penguin/data` |
| `PORT` | Web 服务监听端口 | `7364` |
| `HOST` | Web 服务监听地址 | `127.0.0.1` |
| `PENGUIN_WEB_DB` | 服务端 SQLite 数据库路径 | `<root>/web.db` |
| `PENGUIN_WEB_DIST` | 前端静态资源目录 | npm 安装的服务端包回退到内置 web-dist |
| `PENGUIN_LANG` | CLI 语言(`en` / `zh`),用 `penguin config lang` 设置 | `en` |
### Provider 凭证环境变量
当模型条目未内联 `api_key` 时,AgentHub 网关按 Provider 回退读取对应环境变量;`*_BASE_URL` 变体同理覆盖 Base URL:
| Provider | API Key | Base URL |
| --- | --- | --- |
| deepseek | `DEEPSEEK_API_KEY` | `DEEPSEEK_BASE_URL` |
| anthropic | `ANTHROPIC_API_KEY` | `ANTHROPIC_BASE_URL` |
| openai、openrouter、siliconflow、custom | `OPENAI_API_KEY` | `OPENAI_BASE_URL` |
| google | `GEMINI_API_KEY` | `GEMINI_BASE_URL` |
| zhipu | `ZAI_API_KEY` | `ZAI_BASE_URL` |
| moonshot | `MOONSHOT_API_KEY` | `MOONSHOT_BASE_URL` |
openrouter、siliconflow 与 custom 分组走 OpenAI 兼容协议,因此复用 `OPENAI_*` 变量。Provider 分组与内置模型目录见[模型与 Provider](/models)。
## Project 配置
`<root>/<project>/.project_config.toml` 是 Project 唯一的配置文件:隐藏文件,落盘权限 0600,凭证内联在模型条目上。模型身份始终是 `(provider, model_id)` 成对引用,禁止任何形式的字符串拼接。
| 字段 | 说明 |
| --- | --- |
| `name` | Project 展示名(缺省显示 id) |
| `default_model` | 缺省模型的成对引用 `{ provider, model_id }`,必须指向 `models` 中的条目 |
| `vision_model` | 代读图片的视觉模型(供纯文本模型的 `describe_image` 使用),成对引用 |
| `[[models]]` | 可用模型条目列表 |
模型条目(`[[models]]`)字段:
| 字段 | 说明 |
| --- | --- |
| `provider` | Provider 分组;与 `model_id` 共同构成条目唯一键 |
| `model_id` | 上游请求 id,原样发送给 AgentHub |
| `context_window` | 上下文窗口大小 |
| `client_type` | AgentHub 客户端协议;缺省由 `model_id` 推断,OpenAI 兼容的第三方模型应设为 `openai` |
| `display_name` | 展示名;仅在与内置目录不同时持久化 |
| `vision` | 是否支持图片输入;缺省视为支持 |
| `pricing` | 三档价格 `cache_read` / `cache_write` / `output`,单位 USD 每百万 Token(`unit = "usd_per_mtok"`) |
| `api_key` | 内联凭证;留空回退到 Provider 环境变量 |
| `base_url` | 自定义 Base URL;网关模型预置 |
| `created_at` | `api_key` 写入时间(ISO 8601,界面维护的展示字段) |
```toml
default_model = { provider = "deepseek", model_id = "deepseek-v4-pro" }
[[models]]
provider = "deepseek"
model_id = "deepseek-v4-pro"
context_window = 1000000
vision = false
api_key = "sk-..."
[models.pricing]
unit = "usd_per_mtok"
cache_read = 0.003571
cache_write = 0.428571
output = 0.857143
```
`pricing.unit` 目前固定为 `usd_per_mtok`(USD 每百万 Token);三档对应 `token_usage` 的三个计数桶。
该文件通过 CLI `penguin config model …` 或 Web 的 Models 页面修改——服务运行期间不要手工编辑,模型本身则永远无权读写它。
## Agent 配置
`agent_state/system_config.yaml` 定义单个 Agent 的行为(YAML;经 Web UI 编辑时保留注释):
| 字段 | 缺省值 | 说明 |
| --- | --- | --- |
| `name` | — | Agent 展示名(缺省回退到 id) |
| `description` | — | Agent 描述 |
| `version` | `1` | Agent State 版本号(自然数),每次成功优化自增 |
| `system_prompt` | 内置模板 | 必填;唯一进行占位符替换的模板 |
| `max_turns` | `100` | 单个 Task 的最大 LLM 轮数 |
| `model.max_tokens` | `32000` | 单次输出 Token 上限 |
| `model.thinking_level` | `medium` | `none` / `low` / `medium` / `high` / `xhigh` |
| `model.timeoutMs` | `120000` | 单次 Request 超时(毫秒) |
| `compaction.max_context_length` | `128000` | 触发压缩的上下文 Token 阈值 |
| `compaction.max_session_turns` | `-1` | Session 累计轮数阈值(`-1` 不限制) |
| `compaction.mode` | `summarize` | `summarize` / `discard` |
| `compaction.prompt` | 内置模板 | summarize 压缩使用的 Prompt |
| `tools.builtin` | 缺省时为完整默认工具集 | 工具条目:`name` / `description` / `parameters` / `permission`(`r` 或 `rw`)/ `forModel` / `timeoutMs` / `maxOutputLength`;一旦写出即整体替换默认列表 |
| `tools.mcpServers` | `[]` | MCP Server 配置(`name` + `config`),预留给 MCP 适配层 |
工具权限与审批语义见[工具与审批](/tools)。
局部调整示例(在初始化生成的文件基础上修改)。注意本文件**不与默认值做 deep merge**:写出的字段整体生效,省略的字段才在使用处回退表中缺省值;`system_prompt` 是必填字段(缺失会拒绝加载),编辑其他字段时应保留初始化写入的完整模板:
```yaml
name: default_agent
description: General-purpose agent
version: 3
# 必填:保留初始化生成的完整默认模板(含 {{AGENTS_MD}} 等占位符,此处从略)。
system_prompt: |
…
max_turns: 100
model:
max_tokens: 32000
thinking_level: medium
timeoutMs: 120000
compaction:
max_context_length: 128000
max_session_turns: -1
mode: summarize
# tools 整段省略 = 使用完整默认工具集。一旦写出 tools.builtin,将**整体替换**
# 默认列表:必须为每个要保留的工具携带完整定义(含 parameters JSON Schema),
# 参见「工具与审批」页。
```
### 系统提示词占位符
`system_prompt` 是唯一进行占位符替换的模板,可用占位符:
| 占位符 | 注入内容 |
| --- | --- |
| `{{AGENTS_MD}}` | `AGENTS.md` 的全文 |
| `{{VAULT_KEYS}}` | Vault 的键名列表(仅键名) |
| `{{SKILL_METADATA}}` | 已安装 Skill 的元数据 |
| `{{PLATFORM}}` | 运行平台 |
| `{{OS_VERSION}}` | 操作系统版本 |
| `{{DATE}}` | 当前日期 |
| `{{CWD}}` | Workspace 路径 |
| `{{AGENT_ID}}` | Agent id |
| `{{PROJECT_DIR}}` | Project 目录 |
| `{{SESSION_ID}}` | Session id |
`agent_state/AGENTS.md` 是开发者可编辑的指令文件,经 `{{AGENTS_MD}}` 注入系统提示词,缺省为空——它也是优化器最常改动的文件(见[自我进化](/self-improvement))。
## Vault
`agent_state/.vault.toml` 是 Agent 级的环境变量保险库:隐藏文件,落盘权限 0600。
- 键名须匹配 `^[A-Za-z_][A-Za-z0-9_]*$`(shell 环境变量命名规则);
- 值只注入工具子进程的环境变量,永远不进入模型上下文与 Trace;
- 系统提示词中经 `{{VAULT_KEYS}}` 只披露键名;
- 通过 CLI `penguin config vault set/list/remove` 或 Web 的 Vault 标签页管理。
## 定时任务
`agent_state/schedule/<name>.toml` 每个文件描述一个定时任务(文件名即任务标识),按节律向 Agent 发送预设 Prompt。定时任务仅在 Web 服务(server 运行时)运行期间执行,在 Web 的 Agent 设置 → Schedule 标签页管理。
| 字段 | 必填 | 说明 |
| --- | --- | --- |
| `prompt` | 是 | 触发时发送的 Prompt |
| `enabled` | 否 | 是否启用,缺省 `false` |
| `start_at` | 是 | 首次触发时刻(ISO 8601) |
| `period` | 否 | 周期,形如 `30m` / `12h` / `7d`,下限 5 分钟;缺省为一次性任务 |
| `end_at` | 否 | 结束时刻,须晚于 `start_at` |
| `session_id` | 否 | 绑定既有 Session;与下列三项互斥 |
| `workspace` | 否 | 新建 Session 模式的 Workspace |
| `provider` / `model_id` | 否 | 新建 Session 模式的模型成对引用 |
```toml
prompt = "检查昨日构建结果并汇总失败原因"
enabled = true
start_at = 2026-08-01T09:00:00Z
period = "12h"
```
## 设计原则
Agent 的行为完整地存放于磁盘上的可编辑文件——提示词、Skill、配置都是数据而非代码。正因如此,Agent 才能被 Agent 改进:优化器编辑的与你手工编辑的是同一批文件。参见[自我进化](/self-improvement)与 [CLI 参考](/cli)。
+79
View File
@@ -0,0 +1,79 @@
---
title: Installation
description: Install PenguinHarness via the install script, npm, or from source.
---
## Requirements
- Linux / macOS (x64 or arm64): the install script ships platform tarballs with an official Node.js runtime bundled — no local Node needed.
- Other platforms, or installing via npm / from source: system Node.js >= 24.
## Script install (recommended)
On Linux / macOS:
```bash
curl -fsSL https://github.com/Prism-Shadow/penguin-harness/releases/latest/download/install.sh | sh
```
The script downloads the matching `penguin-{linux,darwin}-{x64,arm64}.tar.gz`, which bundles an official Node.js runtime. Other platforms do **not** fall back automatically: the script exits and asks you to install Node.js >= 24 and re-run with `--universal`, which selects the runtime-less `penguin-universal.tar.gz`.
Verify the install:
```bash
penguin -v
```
### Install location and options
| Item | Details |
| --- | --- |
| Install dir | `~/.penguin` by default; override with the `PENGUIN_INSTALL_DIR` env var |
| Command entry | A symlink `~/.local/bin/penguin` is created (the script warns if `~/.local/bin` is not on PATH) |
| Version pin | `PENGUIN_VERSION=vX.Y.Z` env var, or the `--version vX.Y.Z` script flag; defaults to the latest Release |
| Integrity check | Downloads are sha256-verified when the Release ships checksum assets |
| Upgrade | Re-run the install script; files are swapped atomically |
Script flags are passed as `curl ... | sh -s -- --universal`.
### Data directory
The data directory defaults to `~/.penguin/data` — under the install home `~/.penguin`, but never modified by install or upgrade — and is overridable with the `PENGUIN_HOME` env var. Model configuration, Session records, and other data are preserved across upgrades.
## npm install
Requires system Node.js >= 24:
```bash
npm install -g @prismshadow/penguin-cli
```
The npm package is `@prismshadow/penguin-cli`; the installed command is `penguin`. Web UI assets ship inside the `@prismshadow/penguin-server` package, so this single install yields a working `penguin web`.
## From source
Requires Node.js >= 24 and pnpm:
```bash
git clone https://github.com/Prism-Shadow/penguin-harness.git
cd penguin-harness
pnpm install && pnpm build
```
After the build, run `pnpm penguin <args>` inside the repo as the dev runner, or use the globally linked `penguin` command.
## Published npm packages
| Package | Description |
| --- | --- |
| `@prismshadow/penguin-cli` | Command-line tool providing the `penguin` command |
| `@prismshadow/penguin-core` | SDK for creating Agents and Sessions programmatically |
| `@prismshadow/penguin-server` | Web service, including the Web UI assets |
| `@prismshadow/penguin-skills` | Skill collection |
All packages are published under the Apache-2.0 license.
## Next steps
- [Quickstart](/quickstart): configure a model and run your first Task.
- [CLI Reference](/cli): the full list of commands and options.
+79
View File
@@ -0,0 +1,79 @@
---
title: 安装
description: 通过安装脚本、npm 或源码安装 PenguinHarness。
---
## 系统要求
- Linux / macOS(x64 或 arm64):安装脚本提供内置官方 Node.js 运行时的平台压缩包,解压即用,无需本机安装 Node。
- 其他平台,或通过 npm / 源码安装:需要系统 Node.js >= 24。
## 脚本安装(推荐)
在 Linux / macOS 上执行:
```bash
curl -fsSL https://github.com/Prism-Shadow/penguin-harness/releases/latest/download/install.sh | sh
```
脚本按平台下载 `penguin-{linux,darwin}-{x64,arm64}.tar.gz`,其中捆绑了官方 Node.js 运行时。其他平台**不会自动回退**:脚本会退出并提示先安装 Node.js >= 24、再携带 `--universal` 重新执行,改用不含运行时的 `penguin-universal.tar.gz`。
安装完成后验证:
```bash
penguin -v
```
### 安装位置与选项
| 项目 | 说明 |
| --- | --- |
| 安装目录 | 默认 `~/.penguin`,可用环境变量 `PENGUIN_INSTALL_DIR` 覆盖 |
| 命令入口 | 创建符号链接 `~/.local/bin/penguin`(若 `~/.local/bin` 不在 PATH 上,脚本会给出提示) |
| 版本固定 | 环境变量 `PENGUIN_VERSION=vX.Y.Z`,或脚本参数 `--version vX.Y.Z`;默认安装最新 Release |
| 完整性校验 | Release 提供 checksum 资产时自动进行 sha256 校验 |
| 升级 | 重新执行安装脚本即可,文件原子替换 |
脚本参数通过 `curl ... | sh -s -- --universal` 的形式传入。
### 数据目录
数据目录默认位于 `~/.penguin/data`(在安装主目录 `~/.penguin` 之下,但安装与升级都不会改动它),可用环境变量 `PENGUIN_HOME` 覆盖。模型配置、Session 记录等在升级后均会保留。
## npm 安装
需要系统 Node.js >= 24:
```bash
npm install -g @prismshadow/penguin-cli
```
npm 包名为 `@prismshadow/penguin-cli`,安装后的命令是 `penguin`。Web UI 静态资源随 `@prismshadow/penguin-server` 包发布,因此仅执行上述命令即可直接使用 `penguin web`。
## 源码安装
需要 Node.js >= 24 与 pnpm:
```bash
git clone https://github.com/Prism-Shadow/penguin-harness.git
cd penguin-harness
pnpm install && pnpm build
```
构建完成后,在仓库内用 `pnpm penguin <args>` 作为开发入口运行,或使用全局链接的 `penguin` 命令。
## 已发布的 npm 包
| 包 | 说明 |
| --- | --- |
| `@prismshadow/penguin-cli` | 命令行工具,提供 `penguin` 命令 |
| `@prismshadow/penguin-core` | SDK,程序化创建 Agent 与 Session |
| `@prismshadow/penguin-server` | Web 服务,含 Web UI 静态资源 |
| `@prismshadow/penguin-skills` | Skill 集合 |
全部包以 Apache-2.0 协议发布。
## 下一步
- [快速开始](/quickstart):配置模型并运行第一个 Task。
- [CLI 参考](/cli):完整的命令与选项列表。
+241
View File
@@ -0,0 +1,241 @@
---
title: Core Interfaces
description: A top-down tour of the contracts — full LLMInterface and EnvironmentInterface signatures, inner types field by field, and every swappable seam.
---
The context_engine depends on three interfaces: Human, LLM and Environment. All protocol conversion happens inside the implementations — the engine sees only [OmniMessage](/omni-message). This page goes top-down: the two big interface signatures and the Human boundary first, then each interface's inner types layer by layer. All types are exported by `@prismshadow/penguin-core`; source: `packages/core/src/interfaces.ts`.
## Overview
```text
Human (a boundary, not a class)
session.run(newMessages, { approve, signal })
│ ▲
▼ │ streamed OmniMessage
context_engine
│ │
LLMInterface │ │ EnvironmentInterface
▼ ▼
GenerativeModel Environment
└─ AgentHub gateway └─ BuiltinTool registry (exec_command …)
```
| Interface | Contract | Built-in implementation |
| --- | --- | --- |
| Human | `session.run`'s inputs and streamed output | CLI, Server (SSE) |
| LLM | `LLMInterface.streamGenerate` | `GenerativeModel` (over AgentHub) |
| Environment | `EnvironmentInterface.executeTool` et al. | `Environment` + the builtin tool registry |
Two iron rules run through every interface: **never throw into the engine** (errors converge into messages/returns carrying a `stop_reason`), and **the streaming discipline** (`start → delta → stop`, complete message immediately after).
## LLMInterface
The complete model-side contract is a single method:
```ts
interface LLMInterface {
streamGenerate(parameters: GenerativeModelParameters): AsyncGenerator<OmniMessage, LLMOutcome>;
}
interface GenerativeModelParameters {
newMessages: OmniMessage[]; // only this turn's new messages (the impl owns history; mixed roles rejected)
signal?: AbortSignal;
}
```
The generator yields `partial_*` fragments and complete messages, emits Token usage as `token_usage` events, and reports the terminal state via its **return value** (not a yielded message).
### LLMOutcome semantics
```ts
interface LLMOutcome {
status: StopReason; // completed | timeout | malformed | aborted | failed
message?: string; // display text when failed
}
```
| status | Meaning | Engine reaction |
| --- | --- | --- |
| `completed` | finished normally (token_usage already emitted) | proceed |
| `timeout` | timeout / lost connection | auto-reconnect within the run |
| `malformed` | response parse failure | auto-reconnect within the run |
| `aborted` | user interrupt | stop, hand back to the user |
| `failed` | non-retryable (auth/params, …) | stop, hand back to the user |
Implementation constraints: never throw; no internal retries — reconnecting is the engine's job (see [The Agent Loop](/agent-loop)).
### GenerativeModelConfig
The built-in implementation's init config, field by field:
```ts
interface GenerativeModelConfig {
modelId: string;
apiKey?: string;
baseUrl?: string;
clientType?: string; // AgentHub client protocol (openai / …); inferred from modelId when omitted
tools: ToolDefinition[];
systemPrompt?: string; // fully assembled system prompt, placeholders substituted
contextWindow?: number;
maxTokens?: number;
thinkingLevel?: ThinkingLevelName; // "none" | "low" | "medium" | "high" | "xhigh"
requestTimeoutMs?: number; // per-Request timeout, default 120000; <=0 disables
toolCallIds?: ToolCallIdAllocator; // Session-level tool_call_id registry (pass the same instance across compaction)
}
```
### The built-in implementation: GenerativeModel
`GenerativeModel` (`packages/core/src/llm/generative-model.ts`) grounds the contract on the `AutoLLMClient` of the `@prismshadow/agenthub` model gateway:
- the gateway maintains conversation history **statefully**, receiving only new messages each turn; resuming a Session replays committed history through a one-time `setHistory`;
- an internal `EventTranslator` translates gateway stream events into `partial_*` fragments plus complete messages, preserving the `signature` / `phase` fidelity fields; complete messages settle in thinking → text → tool_call order;
- `ToolCallIdAllocator` disambiguates providers that use the function name as the call id (append `#n` inbound, strip outbound), scoped to the whole Session;
- provider differences (tool-call formats, reasoning content, streaming events) are absorbed entirely inside the gateway — see [Models & Providers](/models).
## EnvironmentInterface
The complete tool-execution contract:
```ts
interface EnvironmentInterface {
listTools(): Promise<ToolDefinition[]>;
executeTool(request: ToolExecutionRequest): AsyncGenerator<OmniMessage>;
toolPermission(name: string): "r" | "rw" | undefined; // for frontend approval-mode decisions
dispose?(): void; // release runtime resources; idempotent
}
```
`executeTool` yields `partial_tool_call_output` fragments and ends with exactly one complete `tool_call_output`; `origin`-tagged nested messages (e.g. forwarded by `run_subagent`) pass through unchanged. Rendering is explicitly not this interface's concern — streaming rendering belongs to the CLI / Web front ends.
### ToolExecutionRequest and EnvironmentConfig
```ts
interface ToolExecutionRequest {
toolCall: OmniMessage<ToolCallPayload>; // an approved call
signal?: AbortSignal;
approve?: ApproveFn; // forwarded to tools that spawn child Sessions (approval inheritance)
}
interface EnvironmentConfig {
workspaceDir: string;
toolConfig: ToolConfig; // { customTools: ToolDefinitionConfig[]; mcpServers: MCPServerConfig[] }
services?: EnvironmentServices; // runtime services injected into individual tools
vault?: Record<string, string>; // Vault env vars, injected into exec_command / input_command subprocesses
}
interface EnvironmentServices {
subagentRunner?: SubagentRunner; // needed by run_subagent
visionDescriber?: VisionDescriberService; // needed by describe_image on text-only models
commandSessions?: CommandSessionManager; // long-running command session registry (built by Environment)
subagentSessions?: SubagentSessionManager;// background subagent session registry (likewise)
}
interface MCPServerConfig {
name: string;
config: Record<string, unknown>;
}
```
### The inner tool contract: BuiltinTool
Inside the Environment, an individual tool follows a deliberately narrower contract ("loose tool, strict framework"):
```ts
interface BuiltinTool {
name: string;
definition: ToolDefinitionConfig;
execute(
args: Record<string, unknown>,
ctx: ToolExecutionContext, // { workspaceDir, toolCallId, signal?, approve? }
): AsyncGenerator<OmniMessage, ToolResult | void>;
}
interface ToolDefinitionConfig {
name: string;
description: string;
parameters?: Record<string, unknown>; // JSON Schema
permission?: "r" | "rw";
forModel?: "vision" | "text-only"; // assembled per session-model class
timeoutMs?: number; // default 120000; <=0 disables
maxOutputLength?: number; // default 16000, head-kept truncation; <=0 disables
}
```
A tool emits only content deltas; framing, timeouts, truncation, `stop_reason` priority and errors-to-messages are all handled centrally by the Environment — it is close to impossible for a tool author to break the protocol. Extension is registration: add one `name → factory` entry to `BUILTIN_TOOL_FACTORIES` (`packages/core/src/environment/tools/registry.ts`). Per-tool parameters and behavior: [Tools & Approval](/tools).
## The Human boundary
Human is deliberately not an interface class. The SDK caller *is* the Human:
```ts
const session = await agent.createSession({ workspaceDir, modelId });
session.run(
newMessages: OmniMessage[], // input: the Prompt
opts?: RunOptions,
): AsyncGenerator<OmniMessage>; // output: streamed OmniMessage
interface RunOptions {
signal?: AbortSignal; // interrupt (e.g. Ctrl-C)
approve?: ApproveFn; // per-tool approval; denies everything when omitted
}
```
The CLI wires terminal I/O onto this boundary; the Server wires HTTP requests and SSE channels onto it. Any programmatic caller that connects becomes a new Human implementation — nothing to register.
## ApproveFn
```ts
type ApprovalDecision = "allow" | "deny";
type ApproveFn = (toolCall: OmniMessage<ToolCallPayload>) => Promise<ApprovalDecision>;
```
Constraints: called exactly once per complete `tool_call`; a throwing callback counts as `deny`; when none is injected the engine denies everything (conservative default). A Subagent inherits its parent's approval callback (invoked with an `origin` tag), so the approval policy spans the whole delegation tree.
## Subagent interfaces
Subagent creation is injected at the `createAgent` composition layer, so the Environment never back-depends on the layers above it:
```ts
interface SubagentRunner {
// Precheck errors (depth limit, unknown agent) are thrown — Environment collapses them to failed
spawn(input: {
agentId?: string; // defaults to the current Agent (self-spawn)
modelId?: string; // defaults to the Project default model
}): Promise<SubagentHandle>;
}
interface SubagentHandle {
sessionId: string; // the child Session id: the origin hop; subagent_id derives from its tail
run(input: {
prompt: string;
signal?: AbortSignal;
approve?: ApproveFn; // the parent's approval callback — forwarding is inheritance
}): AsyncGenerator<OmniMessage>;
dispose(): void; // release the child Session's runtime resources; idempotent
}
```
Spawning and running are separate, so the same child Session can accept a follow-up Prompt after a turn ends (a long-running Subagent, driven via `input_subagent`). Child Sessions run in the same Workspace with their own Trace; nesting depth is currently capped at 1.
## VisionDescriberService
The image proxy-reading service for text-only models (needed by `describe_image`):
```ts
interface VisionDescriberService {
modelId: string | null; // null when the Project has no vision_model — the tool ends with a failed explanation
createLLM?: () => LLMInterface; // one-shot LLM for the vision model (no tools, no system prompt)
}
```
## Extension seams
| To … | Do … |
| --- | --- |
| Swap or customize model access | implement `LLMInterface` (or just set `client_type` for OpenAI-compatible endpoints) |
| Swap the execution sandbox | implement `EnvironmentInterface` |
| Add a tool | implement `BuiltinTool` + register a factory; or declare it under `tools.builtin` in `system_config.yaml` |
| Customize approval policy | inject an `ApproveFn` (the CLI/Web modes are wrappers over it) |
| Change an Agent's behavior | edit its Agent State: `system_config.yaml`, `AGENTS.md`, Skills — see the [Configuration Reference](/configuration) |
+241
View File
@@ -0,0 +1,241 @@
---
title: 接口契约
description: 自顶向下的接口全览:LLMInterface 与 EnvironmentInterface 的完整签名、内层类型逐字段定义,以及每一处可替换的扩展点。
---
context_engine 依赖三个接口:Human、LLM、Environment。协议转换全部发生在接口实现内部——引擎只见 [OmniMessage](/omni-message)。本页自顶向下:先给出两大接口的完整签名与 Human 边界,再逐层展开每个接口的内部类型。类型全部由 `@prismshadow/penguin-core` 导出,源码见 `packages/core/src/interfaces.ts`。
## 总览
```text
Human(边界,非接口类)
session.run(newMessages, { approve, signal })
│ ▲
▼ │ 流式 OmniMessage
context_engine
│ │
LLMInterface │ │ EnvironmentInterface
▼ ▼
GenerativeModel Environment
└─ AgentHub 网关 └─ BuiltinTool 注册表(exec_command …)
```
| 接口 | 契约 | 内置实现 |
| --- | --- | --- |
| Human | `session.run` 的入参与流式出参 | CLI、Server(SSE) |
| LLM | `LLMInterface.streamGenerate` | `GenerativeModel`(基于 AgentHub) |
| Environment | `EnvironmentInterface.executeTool` 等 | `Environment` + 内置工具注册表 |
两条铁律贯穿所有接口:**从不向引擎抛异常**(错误收敛为带 `stop_reason` 的消息/返回值),**流式纪律**(`start → delta → stop`,随后立即产出完整消息)。
## LLMInterface
模型侧的完整契约只有一个方法:
```ts
interface LLMInterface {
streamGenerate(parameters: GenerativeModelParameters): AsyncGenerator<OmniMessage, LLMOutcome>;
}
interface GenerativeModelParameters {
newMessages: OmniMessage[]; // 仅本轮新增消息(实现自行维护历史,多 role 不接受)
signal?: AbortSignal;
}
```
生成器逐条产出 `partial_*` 分片与完整消息,Token 用量以 `token_usage` 事件产出;终态经**返回值**(而非产出消息)给出。
### LLMOutcome 语义
```ts
interface LLMOutcome {
status: StopReason; // completed | timeout | malformed | aborted | failed
message?: string; // failed 时的展示文案
}
```
| status | 含义 | 引擎的反应 |
| --- | --- | --- |
| `completed` | 正常完成(已产出 token_usage) | 继续下一步 |
| `timeout` | 超时/断连 | 同一 run 内自动重连 |
| `malformed` | 响应解析失败 | 同一 run 内自动重连 |
| `aborted` | 用户中断 | 停止交还用户 |
| `failed` | 鉴权/参数等不可重试错误 | 停止交还用户 |
实现约束:从不抛异常;不做内部重试(重连是引擎的职责,见 [Agent 运行循环](/agent-loop))。
### GenerativeModelConfig
内置实现的初始化配置,逐字段:
```ts
interface GenerativeModelConfig {
modelId: string;
apiKey?: string;
baseUrl?: string;
clientType?: string; // AgentHub 客户端协议(openai / …);缺省按 modelId 推断
tools: ToolDefinition[];
systemPrompt?: string; // 占位符替换完成后的完整系统提示词
contextWindow?: number;
maxTokens?: number;
thinkingLevel?: ThinkingLevelName; // "none" | "low" | "medium" | "high" | "xhigh"
requestTimeoutMs?: number; // 单次 Request 超时,默认 120000;<=0 关闭
toolCallIds?: ToolCallIdAllocator; // Session 级 tool_call_id 唯一性登记表(压缩重建时传同一实例)
}
```
### 内置实现:GenerativeModel
`GenerativeModel`(`packages/core/src/llm/generative-model.ts`)把契约落到模型网关 `@prismshadow/agenthub` 的 `AutoLLMClient` 上:
- 网关**有状态**地维护会话历史,每轮只接收新消息;恢复 Session 时经一次性的 `setHistory` 重放已提交历史;
- 内部的 `EventTranslator` 把网关流式事件翻译为 `partial_*` 分片 + 完整消息,保留 `signature` / `phase` 保真字段,完整消息按 thinking → text → tool_call 顺序落盘;
- `ToolCallIdAllocator` 处理个别 Provider 用函数名充当调用 id 的情况(入站追加 `#n`、出站剥离),作用域覆盖整个 Session;
- Provider 协议差异(工具调用格式、思考内容、流式事件)全部在网关内抹平,见[模型与 Provider](/models)。
## EnvironmentInterface
工具执行侧的完整契约:
```ts
interface EnvironmentInterface {
listTools(): Promise<ToolDefinition[]>;
executeTool(request: ToolExecutionRequest): AsyncGenerator<OmniMessage>;
toolPermission(name: string): "r" | "rw" | undefined; // 供前端审批模式判定
dispose?(): void; // 释放运行时资源,幂等
}
```
`executeTool` 逐条产出 `partial_tool_call_output`,并以恰好一条完整 `tool_call_output` 收尾;带 `origin` 的嵌套消息(如 `run_subagent` 转发的子 Session 消息)原样透传。渲染不是本接口的职责——流式渲染由 CLI / Web 前端完成。
### ToolExecutionRequest 与 EnvironmentConfig
```ts
interface ToolExecutionRequest {
toolCall: OmniMessage<ToolCallPayload>; // 已通过审批的调用
signal?: AbortSignal;
approve?: ApproveFn; // 转发给需要派生子 Session 的工具,实现审批继承
}
interface EnvironmentConfig {
workspaceDir: string;
toolConfig: ToolConfig; // { customTools: ToolDefinitionConfig[]; mcpServers: MCPServerConfig[] }
services?: EnvironmentServices; // 注入给个别工具的运行时服务
vault?: Record<string, string>; // Vault 环境变量,注入 exec_command / input_command 子进程
}
interface EnvironmentServices {
subagentRunner?: SubagentRunner; // run_subagent 所需
visionDescriber?: VisionDescriberService; // text-only 模型的 describe_image 所需
commandSessions?: CommandSessionManager; // 长驻命令会话登记表(Environment 内部构造)
subagentSessions?: SubagentSessionManager;// 后台 Subagent 会话登记表(同上)
}
interface MCPServerConfig {
name: string;
config: Record<string, unknown>;
}
```
### 内层工具契约:BuiltinTool
Environment 之内,单个工具遵循更窄的契约(「松工具、紧框架」):
```ts
interface BuiltinTool {
name: string;
definition: ToolDefinitionConfig;
execute(
args: Record<string, unknown>,
ctx: ToolExecutionContext, // { workspaceDir, toolCallId, signal?, approve? }
): AsyncGenerator<OmniMessage, ToolResult | void>;
}
interface ToolDefinitionConfig {
name: string;
description: string;
parameters?: Record<string, unknown>; // JSON Schema
permission?: "r" | "rw";
forModel?: "vision" | "text-only"; // 按 Session 模型类别装配
timeoutMs?: number; // 默认 120000;<=0 关闭
maxOutputLength?: number; // 默认 16000,头部保留截断;<=0 关闭
}
```
工具只产出内容增量;封帧、超时、截断、`stop_reason` 优先级、错误转消息全部由 Environment 统一处理——工具作者几乎不可能写出破坏协议的工具。注册即扩展:向 `BUILTIN_TOOL_FACTORIES`(`packages/core/src/environment/tools/registry.ts`)添加一个 `名称 → 工厂` 条目即可。逐工具的参数与行为见[工具与审批](/tools)。
## Human 边界
Human 刻意不设计为接口类。SDK 的调用方就是 Human:
```ts
const session = await agent.createSession({ workspaceDir, modelId });
session.run(
newMessages: OmniMessage[], // 输入:Prompt
opts?: RunOptions,
): AsyncGenerator<OmniMessage>; // 输出:流式 OmniMessage
interface RunOptions {
signal?: AbortSignal; // 中断信号(如 Ctrl-C)
approve?: ApproveFn; // 逐工具审批;未注入时默认全部拒绝
}
```
CLI 把终端输入输出接到这个边界上;Server 把 HTTP 请求与 SSE 通道接上来。任何程序化调用方接上来就是一种新的 Human 实现,无需注册。
## ApproveFn
```ts
type ApprovalDecision = "allow" | "deny";
type ApproveFn = (toolCall: OmniMessage<ToolCallPayload>) => Promise<ApprovalDecision>;
```
约束:每个完整 `tool_call` 恰好被调用一次;回调抛出异常按 `deny` 处理;未注入时引擎默认全部拒绝(保守策略)。Subagent 继承父级的审批回调(调用时带 `origin` 标记),审批策略天然贯穿整个委托树。
## Subagent 接口
Subagent 的创建能力在 `createAgent` 组装层注入,避免 Environment 反向依赖上层:
```ts
interface SubagentRunner {
// 深度超限、目标 Agent 不存在等前置错误以抛出表达(由 Environment 收敛为 failed)
spawn(input: {
agentId?: string; // 缺省复用当前 Agent(自派生)
modelId?: string; // 缺省用 Project 默认模型
}): Promise<SubagentHandle>;
}
interface SubagentHandle {
sessionId: string; // 子 Session id:消息 origin 的一跳,subagent_id 由其尾部派生
run(input: {
prompt: string;
signal?: AbortSignal;
approve?: ApproveFn; // 父级审批回调,转发即继承
}): AsyncGenerator<OmniMessage>;
dispose(): void; // 释放子 Session 运行时资源,幂等
}
```
派生(spawn)与运行(run)分离,同一子 Session 可以在一轮结束后接受追加 Prompt 继续运行(长驻 Subagent,经 `input_subagent` 驱动)。子 Session 在同一 Workspace 中运行、拥有独立 Trace;嵌套深度当前限制为 1。
## VisionDescriberService
text-only 模型的图像代读服务(`describe_image` 所需):
```ts
interface VisionDescriberService {
modelId: string | null; // Project 未配置 vision_model 时为 null,工具以 failed 说明收尾
createLLM?: () => LLMInterface; // 构造该视觉模型的一次性 LLM(无工具、无系统提示词)
}
```
## 扩展点一览
| 想要 | 做法 |
| --- | --- |
| 更换/自定义模型接入 | 实现 `LLMInterface`(或仅配置 `client_type` 走 OpenAI 兼容协议) |
| 更换执行沙箱 | 实现 `EnvironmentInterface` |
| 新增工具 | 实现 `BuiltinTool` + 注册工厂;或在 `system_config.yaml` 的 `tools.builtin` 中声明 |
| 定制审批策略 | 注入 `ApproveFn`(CLI/Web 的四种模式即其封装) |
| 改变 Agent 行为 | 编辑 Agent State:`system_config.yaml`、`AGENTS.md`、Skills,见[配置参考](/configuration) |
+49
View File
@@ -0,0 +1,49 @@
---
title: Introduction
description: What PenguinHarness is, what ships in the box, and the design tenets behind it.
---
PenguinHarness is an open-source AI Agent harness — a complete TypeScript stack built for constructing and evolving agents. It deploys fully locally (your data never leaves the machine), runs on as little as a single CPU, and reaches 1000+ online and local models through one unified model gateway.
In one line: **Efficient Self-Improving Harness for Everyone.**
## The three pillars
PenguinHarness is organized around three radiating concepts — the message protocol, the SDK, and the skill library — each carrying one pillar:
| Pillar | Meaning |
| --- | --- |
| **Simplest Is the Best** | A deliberately minimal toolset over clean low-level interfaces: fewer tool calls, fewer Tokens, complex tasks done efficiently. |
| **Harness for Building Agents** | With the PenguinHarness SDK, an Agent builds complete Agent applications for you — autonomously, from scratch. |
| **Harness for Recursive Self-Improvement** | With PenguinHarness Skills, an Agent evaluates and optimizes itself, improving recursively over time. |
## What ships in the box
One install gives you four layers that share a single data directory and a single message protocol:
| Component | Package | Description |
| --- | --- | --- |
| SDK | `@prismshadow/penguin-core` | The core engine: ReAct loop, the [OmniMessage protocol](/omni-message), the LLM and Environment [interface contracts](/interfaces), Agent State and Trace. |
| CLI | `@prismshadow/penguin-cli` | The `penguin` command: interactive REPL, one-shot task runs, model and Vault configuration. |
| Server | `@prismshadow/penguin-server` | The Web backend: HTTP [API and SSE streaming](/server-api), multi-user auth, Project authorization, usage statistics. |
| Web App | `@prismshadow/penguin-web` | The browser UI: multi-session chat, Agent management, skill library, model configuration, Trace observability and the evaluation center. |
## Design tenets
These principles run through every component; the design pages keep coming back to them:
- **A minimal toolset**: the shell is the universal interface — file reads, writes and edits all go through `exec_command`. See [Tools & Approval](/tools).
- **Agents are editable data**: prompts, Skills and config are editable files on disk, not hardcoded constants — what you can see, an Agent can improve. See the [Configuration Reference](/configuration).
- **Everything observable**: every request, tool call and approval decision is appended to the [Trace](/sessions-and-traces); a Session restores fully from it.
- **Errors converge into messages**: model and tool failures never throw — they become messages the model can react to. See [The Agent Loop](/agent-loop).
- **Streaming first**: text streams token by token; tool calls and results appear live.
- **Model ↔ Agent decoupling**: an Agent never binds to a model; you pick one per Session. See [Models & Providers](/models).
## A note on naming
The unified message protocol is called **OmniMessage** in technical writing (marketing materials also call it Penguin Message). This documentation uses OmniMessage throughout.
## Next steps
- [Install](/installation) PenguinHarness, then run your first Task with the [Quickstart](/quickstart).
- Start the design docs at the [Architecture](/architecture) overview to see how the pieces fit together.
+49
View File
@@ -0,0 +1,49 @@
---
title: 产品介绍
description: PenguinHarness 是什么,它由哪些部分组成,以及它的设计信条。
---
PenguinHarness 是一个开源的 AI Agent Harness——为「构建 Agent」与「进化 Agent」而生的一整套 TypeScript 基础设施。它完全本地部署,数据不出机器,最低一颗 CPU 即可运行;通过统一的模型网关可接入 1000+ 在线与本地模型。
一句话概括:**Efficient Self-Improving Harness for Everyone.**
## 三大支柱
PenguinHarness 的能力围绕三个递进的概念展开——消息协议、SDK、技能库,分别支撑三个支柱:
| 支柱 | 含义 |
| --- | --- |
| **Simplest Is the Best** | 在干净的底层接口之上刻意保持极简的工具集:更少的工具调用、更少的 Token,高效完成复杂任务。 |
| **Harness for Building Agents** | 基于 PenguinHarness SDK,由一个 Agent 从零开始为你自主构建完整的 Agent 应用。 |
| **Harness for Recursive Self-Improvement** | 基于 PenguinHarness Skills,Agent 评估并优化自己,随时间递归进化。 |
## 产品组成
一次安装即获得完整的四层交付物,它们共享同一套数据目录与同一个消息协议:
| 组件 | 包名 | 说明 |
| --- | --- | --- |
| SDK | `@prismshadow/penguin-core` | 核心引擎:ReAct 循环、[OmniMessage 协议](/omni-message)、LLM 与 Environment [接口契约](/interfaces)、Agent State 与 Trace。 |
| CLI | `@prismshadow/penguin-cli` | 命令行 `penguin`:交互式 REPL、单次任务运行、模型与 Vault 配置。 |
| Server | `@prismshadow/penguin-server` | Web 服务端:HTTP [API 与 SSE 流式通道](/server-api)、多用户认证、Project 授权、用量统计。 |
| Web App | `@prismshadow/penguin-web` | 浏览器界面:多 Session 对话、Agent 管理、技能库、模型配置、Trace 观测与评估中心。 |
## 设计信条
这些原则贯穿所有组件,后续每一页设计文档都会反复引用:
- **极简工具集**:shell 是通用接口,文件读写与命令执行统一经 `exec_command` 完成,见[工具与审批](/tools)。
- **Agent 是可编辑的数据**:Prompt、Skill、配置都是磁盘上的可编辑文件,而非硬编码——你能看到的,Agent 就能改进,见[配置参考](/configuration)。
- **全量可观测**:每一次请求、工具调用与审批决策都以追加方式写入 [Trace](/sessions-and-traces),Session 可从 Trace 完整恢复。
- **错误收敛为消息**:模型与工具的错误不抛异常,而是变成模型可以继续处理的消息,见 [Agent 运行循环](/agent-loop)。
- **流式优先**:文本逐 Token 流出,工具调用与结果实时可见。
- **模型与 Agent 解耦**:Agent 不绑定模型,每个 Session 创建时自由选择,见[模型与 Provider](/models)。
## 命名说明
统一消息协议在技术文档中称为 **OmniMessage**(产品宣传中也叫 Penguin Message)。本文档站一律使用 OmniMessage。
## 下一步
- [安装](/installation) PenguinHarness,然后跟随[快速开始](/quickstart)跑通第一个 Task。
- 从[架构总览](/architecture)进入设计文档,理解各组件如何协作。
+117
View File
@@ -0,0 +1,117 @@
---
title: Message Flow & Ordering
description: How messages travel between Human, engine, LLM, Environment and Trace — every ordering guarantee and non-guarantee, and why stream order differs from context order.
---
[The OmniMessage Protocol](/omni-message) defines what messages *are*; this page explains how they *move* and in what order they become visible: the delivery paths, the merge mechanism, the observable timeline within a turn, which orderings are guaranteed, which are not, and why "order on the stream" and "order in the model context" are two different things. Source of truth: `packages/core/src/engine/context-engine.ts`.
## Delivery paths within a turn
Five actors: Human (the SDK caller), engine (context_engine), LLM, Environment, Trace. Within one turn:
```text
Human ──run(newMessages)──► engine
engine ──write Prompt──────────────────► Trace
engine ──request_begin──► Human and Trace
engine ──streamGenerate(new messages)──► LLM
┌──────────── LLM streams partial_* and complete messages ────────┐
│ engine forwards each: simultaneously ──► Human (yield) │
│ and ──► Trace (write) │
└──────────────────────────────────────────────────────────────────┘
complete tool_call ──► engine: await approve(tc) (one at a time)
engine ──approval_decision──► Human and Trace
allow ──► Environment.executeTool (concurrent, never blocks the LLM stream)
Environment ──partial_tool_call_output──► Human, and (complete) ──► Trace
LLM stream ends: token_usage is its last message, request_end follows at once
still-running tools keep streaming output (possibly after request_end)
all outputs settled ──► reordered to original call order as the next turn's LLM input
```
Key point: **every message is written to the Trace at the same moment it enters the output stream**, so stream order and Trace order agree (the Trace merely skips partials and `origin`-tagged messages — see [Sessions & Traces](/sessions-and-traces)).
## The merge point: MergeQueue
A turn has several concurrent producers: the driver task consuming the LLM stream, plus N concurrently executing tools. All of them push into one merge queue, and a **single consumer** (the `run` generator) yields messages one at a time in **arrival order**; the turn ends only when every producer has finished and the queue is drained.
This one mechanism fixes three basic properties of message delivery:
1. the consumer sees a single **totally ordered** stream — no client-side multiplexing needed;
2. messages from different producers interleave by arrival time — tool outputs arrive in **completion order**, unrelated to call order;
3. order *within* one producer is preserved (the LLM stream is internally ordered; a single tool's fragments are ordered).
## The observable order within a turn
A turn with two tool calls, as the consumer observes it (annotated):
```text
1 event request_begin
2 partial partial_thinking(start → delta… → stop)
3 complete thinking ← the complete message right after stop
4 partial partial_text(start → delta… → stop)
5 complete text
6 partial partial_tool_call A(start → delta… → stop)
7 complete tool_call A
8 event approval_decision(allow, A) ← approvals are sequential; A starts executing
9 partial partial_tool_call B(…) ← the LLM stream continues, not waiting for A
10 complete tool_call B
11 event approval_decision(allow, B)
12 partial partial_tool_call_output B(…) ← B produces output first: completion order
13 complete tool_call_output B
14 event token_usage ← the LLM stream's last message
15 event request_end(completed) ← emitted when the LLM stream ends, not waiting for tools
16 partial partial_tool_call_output A(…) ← late output lands after request_end
17 complete tool_call_output A
(A and B settled → re-fed in A, B original order → next request_begin)
```
If a `tool_call` is denied, line 8 carries `deny` and a synthetic `aborted` `tool_call_output` ("Tool call denied by user.") follows immediately — nothing is dispatched.
## Guarantees and non-guarantees
**Guaranteed:**
| Guarantee | Meaning |
| --- | --- |
| Streaming discipline | every segment goes strictly `start → delta* → stop`, complete message right after; concatenated deltas ≡ the complete message |
| Approval position | `approval_decision` comes after its `tool_call` and before any output of that tool |
| Pairing | every committed `tool_call` gets exactly one complete `tool_call_output` (a denial gets the synthetic one) |
| LLM stream tail | `token_usage` is the LLM stream's last message, `request_end` follows immediately |
| Commit criterion | `request_end.status === "completed"` ⇔ the turn was committed by the gateway (replay keeps or drops on this) |
| Stream order = Trace order | written as streamed; the Trace only filters partials and `origin` messages |
| Transport ordering | SSE delivers per channel with monotonic ids; reconnects replay from `Last-Event-ID` or get `resync_required` — see [Server API](/server-api) |
**Not guaranteed (renderers must not rely on these):**
| Non-guarantee | Meaning |
| --- | --- |
| Tool-output order | arrival is completion order; fragments of different tools interleave — attribute by `tool_call_id` |
| `request_end` ≠ end of turn | still-running tools may emit output after `request_end` and before the next `request_begin` |
| Event/content spacing | later LLM-stream messages may land between an `approval_decision` and that tool's first output |
## Stream order vs context order
The same batch of tool outputs exists in two orders, serving two different consumers:
- **stream order (completion order)** — for the Human: whoever finishes first is visible first, for real-time rendering;
- **context order (original call order)** — for the model: before entering the next turn's input, outputs are reordered to the original `tool_call` order, matching provider pairing rules.
Therefore **a renderer must never reconstruct the context from arrival order** — hang each output onto its call via `tool_call_id`; the engine owns context ordering.
## Edge-case timelines
| Case | Observable order on the stream |
| --- | --- |
| User interrupt | (messages produced so far) → the `abort` event — the last message before `run` returns; carry-over goes to the model context only, never streamed, never written to Trace |
| Automatic reconnect | `request_end(timeout \| malformed)` → a fresh `request_begin`; the `<turn_retried>` block is model-visible only |
| Compaction | `compaction_begin` → the compaction request runs against the old context (its streamed output is **not** forwarded, only written to Trace) → that request's `token_usage` → `compaction_end(status)` |
| max_turns reached | a length notice → the run ends; unsubmitted input is kept as carry-over |
| The Prompt itself | written to Trace, not echoed back onto the stream (the caller already has it) |
| session_meta | never emitted on the main Session's stream (it lives in the Trace and the history API); a Subagent child stream's **first** message is the child's `session_meta` |
## Across Sessions: the origin chain
A child Session spawned by `run_subagent` has its own complete stream. When forwarded to the parent, each child message gets one child-Session-id hop prepended to `origin`, and it interleaves with the parent's own messages **by arrival time**; renderers route by `origin` into the nested card. Child messages are not written to the parent Trace — the parent keeps only the `subagent` pointer event, while the child's stream order is recorded in its own Trace.
## Transport ordering (SSE)
The Server pushes this exact output stream verbatim (single-line JSON) onto the per-Session SSE channel: monotonically increasing event ids, a bounded replay buffer for reconnects, `resync_required` when the replay window is gone. Event order: on reconnect the replayed gap (or `resync_required`) comes first, then the authoritative `task_state` snapshot and pending approvals; a fresh connection skips replay, so `task_state` is its first event. Details — including the bundled Web App's connect-first + dedup consumption pattern — are on the [Server API](/server-api) page.
+116
View File
@@ -0,0 +1,116 @@
---
title: 消息流转与时序
description: 消息在 Human、engine、LLM、Environment 与 Trace 之间的传递机制,每一处顺序保证与非保证,以及流序与上下文序的区别。
---
[OmniMessage 协议](/omni-message)定义了消息**是什么**,本页讲清消息**怎么传、以什么顺序可见**:传递路径、汇流机制、一轮内的可见时序、哪些顺序有保证、哪些没有,以及"流上的顺序"与"模型上下文的顺序"为何是两回事。源码依据:`packages/core/src/engine/context-engine.ts`。
## 一轮的传递路径
五个参与者:Human(SDK 调用方)、engine(context_engine)、LLM、Environment、Trace。一轮之内:
```text
Human ──run(newMessages)──► engine
engine ──写 Prompt──────────────────────► Trace
engine ──request_begin──► Human 与 Trace
engine ──streamGenerate(新消息)──► LLM
┌───────────────── LLM 流式返回 partial_* 与完整消息 ────────┐
│ engine 逐条转发:每条同时 ──► Human(yield)与 ──► Trace(写) │
└─────────────────────────────────────────────────────────────┘
完整 tool_call ──► engine:await approve(tc)(逐个)
engine ──approval_decision──► Human 与 Trace
allow ──► Environment.executeTool(并发,不阻塞 LLM 流)
Environment ──partial_tool_call_output──► Human 与(完整时)Trace
LLM 流结束:最后一条 token_usage,随即 request_end ──► Human 与 Trace
仍在执行的工具继续流出输出(可晚于 request_end)
全部输出齐 ──► 按原始调用顺序重排,作为下一轮 LLM 输入
```
要点:**每条消息在进入输出流的同时写入 Trace**,因此流序与 Trace 序一致(Trace 跳过分片与带 `origin` 的消息,见 [Session 与 Trace](/sessions-and-traces))。
## 单一汇流点:MergeQueue
一轮内存在多个并发生产者:消费 LLM 流的驱动任务,加上 N 个并发执行的工具。它们全部 push 进同一个合并队列,由**单一消费者**(`run` 生成器)按**到达顺序**逐条 yield;生产者全部完成且队列排空,这一轮才结束。
这一机制决定了消息传递的三条基本性质:
1. 消费方看到的是一条**全序**的消息流,不需要自己做多路归并;
2. 不同生产者的消息按到达时刻交错——工具输出的先后是**完成顺序**,与调用顺序无关;
3. 同一生产者内部的顺序被保留(LLM 流内部有序;单个工具的分片有序)。
## 一轮内的可见顺序
一个带两次工具调用的轮,消费方按序观察到(标注示例):
```text
1 event request_begin
2 partial partial_thinking(start → delta… → stop)
3 complete thinking ← stop 后立即跟完整消息
4 partial partial_text(start → delta… → stop)
5 complete text
6 partial partial_tool_call A(start → delta… → stop)
7 complete tool_call A
8 event approval_decision(allow, A) ← 审批逐个,决策即产出;A 开始并发执行
9 partial partial_tool_call B(…) ← LLM 流继续,不等 A
10 complete tool_call B
11 event approval_decision(allow, B)
12 partial partial_tool_call_output B(…) ← B 先有输出:完成顺序,非调用顺序
13 complete tool_call_output B
14 event token_usage ← LLM 流的最后一条
15 event request_end(completed) ← LLM 流结束即产出,不等工具
16 partial partial_tool_call_output A(…) ← 迟到输出出现在 request_end 之后
17 complete tool_call_output A
(A、B 输出齐 → 按 A、B 原始顺序进入下一轮输入 → 下一个 request_begin)
```
若某条 `tool_call` 被拒绝,第 8 行的决策为 `deny`,随即产出一条合成的 `aborted` `tool_call_output`(内容 `Tool call denied by user.`),不派发执行。
## 顺序保证与非保证
**有保证:**
| 保证 | 说明 |
| --- | --- |
| 分片纪律 | 每段严格 `start → delta* → stop`,完整消息紧随其后;全部 delta 拼接 ≡ 完整消息 |
| 审批位次 | `approval_decision` 在其 `tool_call` 之后、该工具任何输出之前 |
| 配对完整 | 每个已提交的 `tool_call` 恰好对应一条完整 `tool_call_output`(拒绝为合成输出) |
| LLM 流收尾 | `token_usage` 是 LLM 流的最后一条,`request_end` 紧随其后 |
| 提交判据 | `request_end.status === "completed"` ⇔ 该轮已被网关提交(回放据此取舍) |
| 流序 = Trace 序 | 逐条"边流边写";Trace 只是滤掉分片与 `origin` 消息 |
| 传输有序 | SSE 按通道单调 id 投递,断线按 `Last-Event-ID` 补发或 `resync_required`,见 [Server API](/server-api) |
**无保证(渲染层不得依赖):**
| 非保证 | 说明 |
| --- | --- |
| 工具输出顺序 | 到达顺序是完成顺序;多工具的分片可交错,须按 `tool_call_id` 归属 |
| `request_end` ≠ 轮结束 | 仍在执行的工具输出可出现在 `request_end` 之后、下一个 `request_begin` 之前 |
| 事件与内容的相对间隔 | `approval_decision` 与首条工具输出之间可能插入 LLM 流的后续消息 |
## 流序与上下文序
同一批工具输出存在两种顺序,服务两个不同的消费者:
- **流序(完成顺序)**——面向 Human:谁先完成谁先可见,保证实时性;
- **上下文序(原始调用顺序)**——面向模型:进入下一轮输入前按 `tool_call` 的原始顺序重排,保证与 Provider 的配对约定一致。
因此**渲染层不得用到达顺序重建上下文**——按 `tool_call_id` 把输出挂回对应调用即可;上下文顺序由 engine 负责。
## 边界情形的时序
| 情形 | 流上可见的顺序 |
| --- | --- |
| 用户中断 | (已产出的消息)→ `abort` 事件——`run` 返回前的最后一条;补发内容只进模型上下文,不上流、不进 Trace |
| 自动重连 | `request_end(timeout \| malformed)` → 新的 `request_begin`;`<turn_retried>` 块仅模型可见 |
| 上下文压缩 | `compaction_begin` → 压缩请求在旧上下文中执行(其流式输出**不上行**,只写 Trace)→ 该请求的 `token_usage` → `compaction_end(status)` |
| 达到 max_turns | 长度提示消息 → 结束;未提交的输入按补发保留 |
| Prompt 本身 | 写入 Trace,但不回流(输入方已有) |
| session_meta | 主 Session 的输出流不产出它(存在于 Trace 与历史接口中);Subagent 子流的**第一条**是子 Session 的 `session_meta` |
## 跨 Session:origin 链
`run_subagent` 派生的子 Session 有自己的完整消息流。转发给父级时,每条子消息的 `origin` 前插一跳子 Session id,与父级本地消息**按到达时刻交错**;渲染层按 `origin` 归入对应子会话卡片。子消息不写父 Trace——父 Trace 只保留 `subagent` 指针事件,子 Session 的流序记录在它自己的 Trace 里。
## 传输层顺序(SSE)
Server 把上述输出流原样(单行 JSON)推入 per-Session SSE 通道:事件 id 单调递增,有界缓冲支持断线补发,重放窗口失效时以 `resync_required` 通知客户端重拉历史。事件次序:重连时补发的缺口(或 `resync_required`)在前,随后才是权威的 `task_state` 快照与未决审批;全新连接不重放缓冲,首条即为 `task_state` 快照。细节见 [Server API](/server-api) 的流式接口一节;自带 Web App 的"连接先行 + 去重"消费模式亦在该页。
+88
View File
@@ -0,0 +1,88 @@
---
title: Models & Providers
description: Model access through the single AgentHub gateway, (provider, model_id) identity, the per-Project model table, credentials and thinking levels.
---
## One gateway
All model access goes through one gateway library: `@prismshadow/agenthub` (AutoLLMClient). Core defines only a thin `LLMInterface` (see [Interfaces](/interfaces)); per-provider protocol adaptation happens inside AgentHub, so 1000+ online and local models are reachable, including any OpenAI-compatible endpoint. The protocol translation lives in `packages/core/src/llm/generative-model.ts`.
## Model identity
A model's identity is always the `(provider, model_id)` pair: `provider` is a config group name, `model_id` the upstream request id sent to AgentHub unchanged. The two are independent fields — concatenating them into one string is forbidden anywhere in the pipeline.
## The per-Project model table
Each Project's available models are recorded in the hidden `.project_config.toml`, maintained via the CLI (`penguin config model add / default / list`, see [CLI Reference](/cli)) or the Web UI — never hand-edited. `ModelEntry` fields:
| Field | Meaning |
| --- | --- |
| `provider` | Config group name; paired with `model_id` it forms the unique key |
| `model_id` | Upstream request id |
| `context_window` | Context window |
| `client_type` | Protocol hint (e.g. `openai`); inferred by AgentHub from the model id when omitted |
| `display_name` | Display name |
| `vision` | Whether image input is supported, default true |
| `pricing` | Three price buckets (unit `usd_per_mtok`, USD per million tokens): `cache_read` / `cache_write` / `output` |
| `api_key` / `base_url` | Inlined credentials, both optional; when blank, AgentHub falls back to environment variables |
A fresh Project defaults to deepseek-v4-pro. A `vision_model` entry can additionally designate the proxy model that `describe_image` uses for text-only session models (see [Tools & Approval](/tools)); it is unset by default.
File shape (illustrative):
```toml
default_model = { provider = "deepseek", model_id = "deepseek-v4-pro" }
vision_model = { provider = "google", model_id = "gemini-3.1-pro-preview" }
[[models]]
provider = "deepseek"
model_id = "deepseek-v4-pro"
context_window = 1000000
[[models]]
provider = "custom"
model_id = "my-model"
client_type = "openai"
base_url = "https://llm.example.com/v1"
api_key = "sk-..."
```
For a model tagged `vision = false` (e.g. the DeepSeek series), images from conversation input are saved to the Session scratchpad and handed over as a file path spliced into the text, and the image-reading tool switches to `describe_image`.
## Built-in provider groups
Built-in groups and their env-var fallbacks (catalog source: `packages/core/src/state/model-catalog.ts`); each group also has a `_BASE_URL` variant (e.g. `ANTHROPIC_BASE_URL`):
| Provider | API key env var | Notes |
| --- | --- | --- |
| deepseek | `DEEPSEEK_API_KEY` | Group of the default model |
| openrouter | `OPENAI_API_KEY` | OpenAI-compatible gateway, preset base URL `https://openrouter.ai/api/v1` |
| siliconflow | `OPENAI_API_KEY` | OpenAI-compatible gateway, preset base URL `https://api.siliconflow.cn/v1` |
| google | `GEMINI_API_KEY` | |
| anthropic | `ANTHROPIC_API_KEY` | |
| openai | `OPENAI_API_KEY` | |
| zhipu | `ZAI_API_KEY` | |
| moonshot | `MOONSHOT_API_KEY` | |
| custom | `OPENAI_API_KEY` | Any OpenAI-protocol endpoint |
The gateway groups (openrouter / siliconflow) go through AgentHub's OpenAI client, so with blank credentials they read `OPENAI_API_KEY` — not a gateway-specific variable.
Some models in the preset catalog: deepseek-v4-pro / deepseek-v4-flash, gemini-3.1-pro-preview, claude-opus-4-8 / claude-sonnet-4-6, gpt-5.5, glm-5.2, kimi-k2.6 (not exhaustive).
## Thinking levels
Five levels: `none | low | medium | high | xhigh`, configured per Agent as `model.thinking_level` in `system_config.yaml`, default medium. See [Configuration](/configuration).
## Models decoupled from Agents
An Agent never binds a model: the model is chosen when a Session is created and stays locked for that Session; the same Agent can run different Sessions on different models. The three `pricing` buckets feed the usage/cost center's per-Token accounting.
Credential handling:
- an inline `api_key` is stored in the hidden Project config file with mode 0600;
- the Web UI masks it on display;
- blank credentials fall back to the provider's environment variables.
## Connectivity test
The Web Models page offers a per-model connectivity test (owner only).
+88
View File
@@ -0,0 +1,88 @@
---
title: 模型与 Provider
description: 经 AgentHub 单一网关接入模型,以 (provider, model_id) 成对标识,Project 级模型表、凭证与思考等级配置。
---
## 单一网关
所有模型访问都经由一个网关库:`@prismshadow/agenthub`(AutoLLMClient)。core 只定义一层很薄的 `LLMInterface`(见 [接口契约](/interfaces)),各 Provider 的协议适配全部由 AgentHub 完成,因此可以接入 1000+ 在线或本地模型,包括任意 OpenAI 兼容端点。协议翻译实现在 `packages/core/src/llm/generative-model.ts`。
## 模型标识
模型身份永远是 `(provider, model_id)` 成对表示:`provider` 是配置分组名,`model_id` 是原样发给上游的请求 id。二者是两个独立字段,任何环节都不允许拼接成一个字符串。
## Project 模型表
每个 Project 的可用模型记录在隐藏文件 `.project_config.toml` 中,由 CLI(`penguin config model add / default / list`,见 [CLI 参考](/cli))或 Web 界面维护,不手工编辑。`ModelEntry` 字段:
| 字段 | 说明 |
| --- | --- |
| `provider` | 配置分组名,与 `model_id` 成对构成唯一键 |
| `model_id` | 上游请求 id |
| `context_window` | 上下文窗口 |
| `client_type` | 协议提示(如 `openai`);缺省由 AgentHub 按 model id 推断 |
| `display_name` | 显示名 |
| `vision` | 是否支持图像输入,默认 true |
| `pricing` | 三档价格(单位 `usd_per_mtok`,USD 每百万 Token):`cache_read` / `cache_write` / `output` |
| `api_key` / `base_url` | 内联凭证,可留空;留空时 AgentHub 回退读环境变量 |
新建 Project 的默认模型是 deepseek-v4-pro。另可配置一条 `vision_model`,作为 text-only 模型使用 `describe_image` 时的代读模型(见 [工具与审批](/tools));默认不配置。
文件形态(示意):
```toml
default_model = { provider = "deepseek", model_id = "deepseek-v4-pro" }
vision_model = { provider = "google", model_id = "gemini-3.1-pro-preview" }
[[models]]
provider = "deepseek"
model_id = "deepseek-v4-pro"
context_window = 1000000
[[models]]
provider = "custom"
model_id = "my-model"
client_type = "openai"
base_url = "https://llm.example.com/v1"
api_key = "sk-..."
```
对标注 `vision = false` 的模型(如 DeepSeek 系列):对话输入中的图片会保存到 Session scratchpad,以文件路径形式拼入文本;读图工具切换为 `describe_image`。
## 内置 Provider 分组
内置分组及其环境变量回退(目录源:`packages/core/src/state/model-catalog.ts`);每个分组同时存在 `_BASE_URL` 变体(如 `ANTHROPIC_BASE_URL`):
| Provider | API Key 环境变量 | 说明 |
| --- | --- | --- |
| deepseek | `DEEPSEEK_API_KEY` | 默认模型所在分组 |
| openrouter | `OPENAI_API_KEY` | OpenAI 兼容网关,预置 base URL `https://openrouter.ai/api/v1` |
| siliconflow | `OPENAI_API_KEY` | OpenAI 兼容网关,预置 base URL `https://api.siliconflow.cn/v1` |
| google | `GEMINI_API_KEY` | |
| anthropic | `ANTHROPIC_API_KEY` | |
| openai | `OPENAI_API_KEY` | |
| zhipu | `ZAI_API_KEY` | |
| moonshot | `MOONSHOT_API_KEY` | |
| custom | `OPENAI_API_KEY` | 任意 OpenAI 协议端点 |
网关分组(openrouter / siliconflow)经 AgentHub 的 OpenAI 客户端请求,因此凭证留空时读取的是 `OPENAI_API_KEY`,而非网关自己的变量名。
预置目录中的部分模型:deepseek-v4-pro / deepseek-v4-flash、gemini-3.1-pro-preview、claude-opus-4-8 / claude-sonnet-4-6、gpt-5.5、glm-5.2、kimi-k2.6 等(非完整清单)。
## 思考等级
思考等级共五档:`none | low | medium | high | xhigh`,按 Agent 在 `system_config.yaml` 的 `model.thinking_level` 配置,默认 medium。见 [配置参考](/configuration)。
## 模型与 Agent 解耦
Agent 从不绑定模型:模型在创建 Session 时选定,并在该 Session 内锁定不变;同一个 Agent 可以在不同 Session 用不同模型运行。`pricing` 三档价格供用量/成本中心按 Token 计费。
凭证处理:
- 内联 `api_key` 存放在权限 0600 的隐藏 Project 配置文件中;
- Web 界面展示时打码;
- 凭证留空时回退到对应 Provider 的环境变量。
## 连通性测试
Web 的模型页为每个模型提供连通性测试(仅 owner 可用)。
+297
View File
@@ -0,0 +1,297 @@
---
title: The OmniMessage Protocol
description: One envelope, three message types, a five-value stop_reason — the unified protocol behind the SDK, the Trace and SSE, field by field.
---
OmniMessage is PenguinHarness's unified message protocol: the SDK yields it, the Trace stores it line by line, and the Server pushes it verbatim over SSE. What streams, what is stored and what the model sees are one structure — there is no second format between front end, back end and storage.
This page goes top-down: the envelope and the three message types first, then every payload field by field, then the protocol-wide semantics (streaming discipline, stop_reason, origin, fidelity fields). Type source: `packages/core/src/omnimessage/types.ts`.
## The envelope
Every message shares one envelope; only the `payload` varies:
```ts
interface OmniMessage<P extends OmniPayload = OmniPayload> {
timestamp: string; // ISO 8601 UTC
type: "session_meta" | "model_msg" | "event_msg";
payload: P;
origin?: string[]; // child-Session chain (outer→inner); absent = main Session
}
```
What each message type carries:
| type | Meaning | Volume |
| --- | --- | --- |
| `session_meta` | The full runtime configuration of one model context | exactly one per context |
| `model_msg` | Content inside the model context (text, thinking, tool calls and results) | the bulk |
| `event_msg` | Runtime events outside the context (approvals, usage, compaction, aborts) | alongside |
## session_meta
```ts
interface SessionMetaPayload {
session_id: string;
provider: string; // one half of the model-identity pair
model_id: string; // the upstream request id sent to AgentHub
model_context_window: number | string;
system_prompt: string; // fully assembled, placeholders substituted
tools: ToolDefinition[]; // the complete tool schema sent to the model
thinking_level: string; // "default" when unconfigured
agent_state: string; // absolute path of the Agent State
workspace: string; // absolute path of the Workspace
}
interface ToolDefinition {
name: string;
description: string;
parameters?: Record<string, unknown>; // JSON Schema
}
```
On resume, the engine takes this Trace line as the runtime config — the model, system prompt and Workspace are immutable for the Session's lifetime. See [Sessions & Traces](/sessions-and-traces).
## model_msg: complete payloads
Seven content payloads, discriminated by `payload.type`. Shared optional fields: `stop_reason` (marks an abnormal terminal state) and `signature` (a provider-fidelity field, see below):
```ts
interface TextPayload {
type: "text";
role: "user" | "assistant";
text: string;
phase?: string | null; // segmentation marker (e.g. GPT-5 phases)
signature?: string;
stop_reason?: StopReason;
}
interface ThinkingPayload {
type: "thinking";
role: "assistant";
thinking: string;
signature?: string; // required by some models to replay history
stop_reason?: StopReason;
}
interface InlineThinkingPayload {
type: "inline_thinking";
role: "assistant";
data: string; // reasoning content in binary form
mime_type: string;
signature?: string;
stop_reason?: StopReason;
}
interface ToolCallPayload {
type: "tool_call";
role: "assistant";
name: string;
arguments: string; // arguments as a JSON string
tool_call_id: string;
signature?: string;
stop_reason?: StopReason;
}
interface ToolCallOutputPayload {
type: "tool_call_output";
role: "user";
output: string;
images?: string[]; // data:<mime>;base64,… URLs (e.g. read_image results)
tool_call_id: string;
stop_reason?: StopReason;
}
interface ImageUrlPayload {
type: "image_url";
role: "user";
image_url: string; // web URL or base64 data URL
stop_reason?: StopReason;
}
interface InlineDataPayload {
type: "inline_data";
role: "user" | "assistant";
data: string; // other binary content
mime_type: string;
signature?: string;
stop_reason?: StopReason;
}
```
`tool_call` and `tool_call_output` pair strictly via `tool_call_id`; a turn's calls form one batch, and outputs are re-fed in the original call order (see [The Agent Loop](/agent-loop)).
## model_msg: streaming partials
Four `partial_*` payloads mirror their complete counterparts, carrying an `event_type` phase marker:
```ts
type StreamEventType = "start" | "delta" | "stop";
interface PartialTextPayload {
type: "partial_text";
role: "assistant";
event_type: StreamEventType;
text: string; // the text added by this fragment
stop_reason?: StopReason;
}
interface PartialThinkingPayload {
type: "partial_thinking";
role: "assistant";
event_type: StreamEventType;
thinking: string;
stop_reason?: StopReason;
}
interface PartialToolCallPayload {
type: "partial_tool_call";
role: "assistant";
event_type: StreamEventType;
name: string;
arguments: string; // incremental fragment of the arguments JSON
tool_call_id: string;
stop_reason?: StopReason;
}
interface PartialToolCallOutputPayload {
type: "partial_tool_call_output";
role: "user";
event_type: StreamEventType;
output: string;
images?: string[]; // images are not incremental — one delta carries the whole set
tool_call_id: string;
stop_reason?: StopReason;
}
```
### The streaming discipline
Every streamed segment follows one timing rule, with the complete message immediately after the `stop`:
```text
partial_text(start) → partial_text(delta) → … → partial_text(stop) → text (complete)
└── concatenation of all deltas ≡ the complete message ──┘
(truncation applies to both alike)
```
Renderers can therefore paint deltas incrementally and swap in the complete message in place; the Trace records only complete messages, never fragments. Interface implementations close their structures internally and never leak an unclosed fragment upward. `PartialAggregator` (`aggregate.ts`) ships a ready-made aggregator.
## event_msg
Eight event payloads, all listed field by field:
```ts
interface RequestBeginPayload {
type: "request_begin";
}
interface RequestEndPayload {
type: "request_end";
status: StopReason; // "completed" is the mechanical commit criterion for replay
}
interface ApprovalDecisionPayload {
type: "approval_decision";
decision: "allow" | "deny";
tool_call_id: string; // pairs with the approved tool_call — the audit record
}
interface TokenUsagePayload {
type: "token_usage";
session: TokenCounts; // Session cumulative
request: TokenCounts; // this Request
}
interface TokenCounts {
cache_read: number;
cache_write: number;
output: number;
total: number;
}
type CompactionReason = "context" | "turns" | "manual";
type CompactionMode = "summarize" | "discard";
interface CompactionBeginPayload {
type: "compaction_begin";
reason: CompactionReason;
mode: CompactionMode;
context: number; // context tokens at trigger time
turns: number; // cumulative turns at trigger time
}
interface CompactionEndPayload {
type: "compaction_end";
reason: CompactionReason;
mode: CompactionMode;
status: StopReason;
}
interface AbortPayload {
type: "abort";
reason?: string | null;
}
interface SubagentPayload {
type: "subagent";
session_id: string; // pointer in the parent Trace to a direct child Session
}
```
## stop_reason
A five-value enum used across messages and interface results (`LLMOutcome.status` uses the same set — see [Core Interfaces](/interfaces)):
```ts
type StopReason = "completed" | "failed" | "aborted" | "timeout" | "malformed";
```
| Value | Meaning | Engine reaction |
| --- | --- | --- |
| `completed` | finished normally | continue |
| `aborted` | user interrupt | stop, hand back to the user |
| `timeout` | LLM timeout / lost connection | LLM side only: auto-reconnect within the run |
| `malformed` | parse failure / truncated stream | LLM side only: auto-reconnect within the run |
| `failed` | other non-retryable error | stop, hand back to the user |
Errors never cross an interface boundary as exceptions — they *are* messages. See [The Agent Loop](/agent-loop).
## origin: the Subagent chain
`origin` serves Subagents: when a child Session's messages are forwarded to the parent, each hop prepends one child Session id (outer→inner), and renderers route messages into the right nested card by the chain:
```ts
// message from the main Session: no origin
{ timestamp: "…", type: "model_msg", payload: { type: "text", … } }
// message from a one-level Subagent: origin = [child Session id]
{ timestamp: "…", type: "model_msg", origin: ["session-2026-07-18-…-a1b2c3d4"], payload: { … } }
```
`origin`-tagged messages are not written to the parent Trace — the child Session has its own Trace, and the parent keeps only the `subagent` pointer event.
## Provider-fidelity fields
Provider-specific fields such as `signature` and `phase` pass through and persist verbatim end to end — some models require them byte-for-byte when history is replayed, and any rewriting would break compatibility. This is one of the preconditions for lossless Session recovery from the Trace.
## Three jobs, one protocol
| Surface | Subset used |
| --- | --- |
| SDK boundary (`session.run` output) | complete `model_msg` + streaming `partial_*` + all `event_msg` |
| Trace on disk | `session_meta` + complete `model_msg` + all `event_msg` (no partials, no `origin`-tagged messages) |
| Server SSE stream | same as the SDK boundary, verbatim single-line JSON — see [Server API](/server-api) |
How messages travel along these surfaces — and every ordering guarantee — is covered on [Message Flow & Ordering](/message-flow).
## Builders and guards
`@prismshadow/penguin-core` exports all types, a builder per message kind (`builders.ts`: `userText`, `assistantText`, `toolCall`, `toolCallOutput`, `partialText`, `tokenUsage`, `withOrigin`, `emptyTokenCounts`, `addTokenCounts`, …) and runtime guards (`isCompleteModelMessage`, `isPartialPayload`, `isModelMessage`, `isEventMessage`, `isSessionMeta`):
```ts
import { userText, isCompleteModelMessage } from "@prismshadow/penguin-core";
const prompt = userText("List the files in the current directory");
// { timestamp: "…", type: "model_msg", payload: { type: "text", role: "user", text: "…" } }
```
+296
View File
@@ -0,0 +1,296 @@
---
title: OmniMessage 协议
description: 一个信封、三类消息、五值 stop_reason——贯穿 SDK、Trace 与 SSE 的统一消息协议,逐字段定义。
---
OmniMessage 是 PenguinHarness 的统一消息协议:SDK 对外产出它,Trace 逐行存储它,Server 经 SSE 原样推送它。「流出去的」「存下来的」「模型看到的」是同一种结构,前后端与存储之间不存在第二套格式。
本页自顶向下:先定义信封与三类消息,再逐字段展开每一种 payload,最后是贯穿全协议的语义(流式纪律、stop_reason、origin、保真字段)。类型源码:`packages/core/src/omnimessage/types.ts`。
## 信封
所有消息共享同一个信封,仅 `payload` 不同:
```ts
interface OmniMessage<P extends OmniPayload = OmniPayload> {
timestamp: string; // ISO 8601 UTC
type: "session_meta" | "model_msg" | "event_msg";
payload: P;
origin?: string[]; // 子 Session 链(由外到内);缺省表示主 Session
}
```
三类消息的分工:
| type | 含义 | 数量级 |
| --- | --- | --- |
| `session_meta` | 一个模型上下文的完整运行配置 | 每个上下文恰好一条 |
| `model_msg` | 模型上下文中的内容消息(文本、思考、工具调用与结果) | 主体 |
| `event_msg` | 上下文之外的运行事件(审批、用量、压缩、中断) | 伴随 |
## session_meta
```ts
interface SessionMetaPayload {
session_id: string;
provider: string; // 模型身份二元组之一
model_id: string; // 发给 AgentHub 的上游请求 id
model_context_window: number | string;
system_prompt: string; // 占位符替换完成后的完整系统提示词
tools: ToolDefinition[]; // 发给模型的完整工具 schema
thinking_level: string; // 未配置时为 "default"
agent_state: string; // Agent State 绝对路径
workspace: string; // Workspace 绝对路径
}
interface ToolDefinition {
name: string;
description: string;
parameters?: Record<string, unknown>; // JSON Schema
}
```
恢复 Session 时,引擎直接以 Trace 中的这条消息为运行时配置——模型、系统提示词、Workspace 在 Session 生命周期内不可变,见 [Session 与 Trace](/sessions-and-traces)。
## model_msg:完整消息
七种内容 payload,以 `payload.type` 判别。公共可选字段:`stop_reason`(非正常收尾时标注终态)与 `signature`(Provider 保真字段,见下文):
```ts
interface TextPayload {
type: "text";
role: "user" | "assistant";
text: string;
phase?: string | null; // 分段标记(如 GPT-5 的 phase)
signature?: string;
stop_reason?: StopReason;
}
interface ThinkingPayload {
type: "thinking";
role: "assistant";
thinking: string;
signature?: string; // 部分模型历史回放所必需
stop_reason?: StopReason;
}
interface InlineThinkingPayload {
type: "inline_thinking";
role: "assistant";
data: string; // 二进制形态的思考内容
mime_type: string;
signature?: string;
stop_reason?: StopReason;
}
interface ToolCallPayload {
type: "tool_call";
role: "assistant";
name: string;
arguments: string; // 参数 JSON 字符串
tool_call_id: string;
signature?: string;
stop_reason?: StopReason;
}
interface ToolCallOutputPayload {
type: "tool_call_output";
role: "user";
output: string;
images?: string[]; // data:<mime>;base64,… 列表(如 read_image 的结果)
tool_call_id: string;
stop_reason?: StopReason;
}
interface ImageUrlPayload {
type: "image_url";
role: "user";
image_url: string; // 网络 URL 或 base64 data URL
stop_reason?: StopReason;
}
interface InlineDataPayload {
type: "inline_data";
role: "user" | "assistant";
data: string; // 其他二进制内容
mime_type: string;
signature?: string;
stop_reason?: StopReason;
}
```
`tool_call` 与 `tool_call_output` 通过 `tool_call_id` 严格配对;一轮内的多个调用是一个批次,输出按原始调用顺序回填(见 [Agent 运行循环](/agent-loop))。
## model_msg:流式分片
四种 `partial_*` payload 与完整消息一一对应,携带 `event_type` 标记分片阶段:
```ts
type StreamEventType = "start" | "delta" | "stop";
interface PartialTextPayload {
type: "partial_text";
role: "assistant";
event_type: StreamEventType;
text: string; // 本条分片新增的文本
stop_reason?: StopReason;
}
interface PartialThinkingPayload {
type: "partial_thinking";
role: "assistant";
event_type: StreamEventType;
thinking: string;
stop_reason?: StopReason;
}
interface PartialToolCallPayload {
type: "partial_tool_call";
role: "assistant";
event_type: StreamEventType;
name: string;
arguments: string; // 参数 JSON 的增量片段
tool_call_id: string;
stop_reason?: StopReason;
}
interface PartialToolCallOutputPayload {
type: "partial_tool_call_output";
role: "user";
event_type: StreamEventType;
output: string;
images?: string[]; // 图像不增量,由单条 delta 整体携带
tool_call_id: string;
stop_reason?: StopReason;
}
```
### 流式纪律
每段流式内容严格遵守同一时序,`stop` 之后立即跟随对应的完整消息:
```text
partial_text(start) → partial_text(delta) → … → partial_text(stop) → text(完整)
└── 全部 delta 拼接 ≡ 完整消息内容(截断也两侧同步) ──┘
```
因此渲染层可以先增量渲染、收到完整消息后原地替换;Trace 只记录完整消息,不存分片。接口实现方在内部把结构闭合完毕,永远不向上层泄漏未闭合的分片。`PartialAggregator`(`aggregate.ts`)提供现成的分片聚合实现。
## event_msg
八种事件 payload,全部逐字段列出:
```ts
interface RequestBeginPayload {
type: "request_begin";
}
interface RequestEndPayload {
type: "request_end";
status: StopReason; // completed 是回放判定「该轮已提交」的机械标准
}
interface ApprovalDecisionPayload {
type: "approval_decision";
decision: "allow" | "deny";
tool_call_id: string; // 与被审批的 tool_call 配对,构成审计记录
}
interface TokenUsagePayload {
type: "token_usage";
session: TokenCounts; // Session 累计
request: TokenCounts; // 本次 Request
}
interface TokenCounts {
cache_read: number;
cache_write: number;
output: number;
total: number;
}
type CompactionReason = "context" | "turns" | "manual";
type CompactionMode = "summarize" | "discard";
interface CompactionBeginPayload {
type: "compaction_begin";
reason: CompactionReason;
mode: CompactionMode;
context: number; // 触发时的上下文 Token 数
turns: number; // 触发时的累计轮数
}
interface CompactionEndPayload {
type: "compaction_end";
reason: CompactionReason;
mode: CompactionMode;
status: StopReason;
}
interface AbortPayload {
type: "abort";
reason?: string | null;
}
interface SubagentPayload {
type: "subagent";
session_id: string; // 父 Trace 中指向直接子 Session 的指针
}
```
## stop_reason
五值枚举,贯穿消息与接口返回(`LLMOutcome.status` 使用同一集合,见[接口契约](/interfaces)):
```ts
type StopReason = "completed" | "failed" | "aborted" | "timeout" | "malformed";
```
| 值 | 语义 | 引擎的反应 |
| --- | --- | --- |
| `completed` | 正常完成 | 继续 |
| `aborted` | 用户中断 | 停止并交还用户 |
| `timeout` | LLM 超时/断连 | 仅 LLM 侧:同一 run 内自动重连 |
| `malformed` | 响应解析失败/流截断 | 仅 LLM 侧:同一 run 内自动重连 |
| `failed` | 其他不可重试错误 | 停止并交还用户 |
错误从不以异常形式穿过接口边界——它们就是消息,见 [Agent 运行循环](/agent-loop)。
## origin:子 Session 链
`origin` 服务于 Subagent:子 Session 的消息转发给父级时,每经过一层就在数组前端添加一个子 Session id(由外到内),渲染层据此把消息归入对应的子会话卡片:
```ts
// 主 Session 的消息:无 origin
{ timestamp: "…", type: "model_msg", payload: { type: "text", … } }
// 一层 Subagent 的消息:origin = [子 Session id]
{ timestamp: "…", type: "model_msg", origin: ["session-2026-07-18-…-a1b2c3d4"], payload: { … } }
```
带 `origin` 的消息不写入父 Trace——子 Session 拥有自己的 Trace,父 Trace 只保留 `subagent` 指针事件。
## 保真字段
`signature` 与 `phase` 等 Provider 专有字段在整条链路上原样透传、原样存储——部分模型在历史回放时要求逐字一致,任何转写都会破坏兼容性。这是 Trace 能够无损恢复 Session 的前提之一。
## 协议的三种职责
| 场景 | 使用的子集 |
| --- | --- |
| SDK 边界(`session.run` 输出) | 完整 `model_msg` + 流式 `partial_*` + 全部 `event_msg` |
| Trace 落盘 | `session_meta` + 完整 `model_msg` + 全部 `event_msg`(不存分片与 `origin` 消息) |
| Server SSE 推送 | 与 SDK 边界一致,原样单行 JSON,见 [Server API](/server-api) |
消息沿这些通道传递的机制与顺序保证,见[消息流转与时序](/message-flow)。
## 构造与判别
`@prismshadow/penguin-core` 导出全部类型、每种消息的构造函数(`builders.ts`:`userText`、`assistantText`、`toolCall`、`toolCallOutput`、`partialText`、`tokenUsage`、`withOrigin`、`emptyTokenCounts`、`addTokenCounts` 等)与运行时判别函数(`isCompleteModelMessage`、`isPartialPayload`、`isModelMessage`、`isEventMessage`、`isSessionMeta`):
```ts
import { userText, isCompleteModelMessage } from "@prismshadow/penguin-core";
const prompt = userText("列出当前目录的文件");
// { timestamp: "…", type: "model_msg", payload: { type: "text", role: "user", text: "…" } }
```
+76
View File
@@ -0,0 +1,76 @@
---
title: Quickstart
description: Install PenguinHarness, configure a model, and run your first Task.
---
## Install
One-liner for Linux / macOS:
```bash
curl -fsSL https://github.com/Prism-Shadow/penguin-harness/releases/latest/download/install.sh | sh
```
For other options (npm, from source), see [Installation](/installation).
## Configure a model
PenguinHarness ships with no built-in model credentials, so configure a model first. Use the Models page in the Web UI, or the CLI:
```bash
penguin config model add --model-id deepseek-v4-pro --api-key sk-... --set-default
```
- When `--provider` is omitted, the Provider is inferred from the built-in catalog.
- The API key can also come from environment variables: when a model entry has no inline api_key, AgentHub (the LLM gateway library) reads variables such as `DEEPSEEK_API_KEY`, `ANTHROPIC_API_KEY`, `OPENAI_API_KEY`, and `GEMINI_API_KEY`. A `.env` file in the working directory is loaded automatically.
## Start the Web App
```bash
penguin web
```
The service runs at http://127.0.0.1:7364 and opens your browser (`--no-open` to skip). First login is `admin` / `admin123` — change it right away. `penguin server` starts the same process headless.
## One-shot run
```bash
penguin run -m "Create hello.txt containing Hello, Penguin"
```
The Workspace defaults to the current directory; pass `--workspace /path` to change it. The target directory must already exist.
## Interactive chat
```bash
penguin chat
```
- Each input line starts a Task.
- `/compact` compacts the context; `/exit` or `/quit` quits; Ctrl-C interrupts the running Task.
- On exit it prints a `penguin chat --resume <sessionId>` hint for resuming this Session; `--resume` without an id resumes the Agent's latest Session.
## SDK hello
After installing `@prismshadow/penguin-core`:
```ts
import { createAgent, isCompleteModelMessage, userText } from "@prismshadow/penguin-core";
const agent = await createAgent({ agentId: "default_agent" });
const session = await agent.createSession({ workspaceDir: process.cwd() });
for await (const output of session.run([userText("Create hello.txt containing hi")], {
approve: async () => "allow",
})) {
if (isCompleteModelMessage(output) && output.payload.type === "text") {
console.log(output.payload.text);
}
}
```
## Next steps
- [Web App Guide](/web-app): use PenguinHarness from the browser.
- [CLI Reference](/cli): the full list of commands and options.
- [Architecture Overview](/architecture): how the pieces fit together.
+76
View File
@@ -0,0 +1,76 @@
---
title: 快速开始
description: 安装 PenguinHarness、配置模型并运行第一个 Task。
---
## 安装
Linux / macOS 一键安装:
```bash
curl -fsSL https://github.com/Prism-Shadow/penguin-harness/releases/latest/download/install.sh | sh
```
其他方式(npm、源码)见[安装](/installation)。
## 配置模型
PenguinHarness 不内置任何模型凭据,使用前需要先配置一个模型。可以在 Web UI 的 Models 页面完成,也可以用 CLI:
```bash
penguin config model add --model-id deepseek-v4-pro --api-key sk-... --set-default
```
- 省略 `--provider` 时,根据内置目录自动推断 Provider。
- API Key 也可以来自环境变量:当模型条目没有内联 api_key 时,LLM 网关库 AgentHub 会读取 `DEEPSEEK_API_KEY`、`ANTHROPIC_API_KEY`、`OPENAI_API_KEY`、`GEMINI_API_KEY` 等变量;工作目录下的 `.env` 会被自动加载。
## 启动 Web App
```bash
penguin web
```
服务运行在 http://127.0.0.1:7364 并自动打开浏览器(`--no-open` 跳过)。首次登录使用 `admin` / `admin123`,请立即修改密码。`penguin server` 启动同一进程的 headless 版本。
## 单次运行
```bash
penguin run -m "创建 hello.txt,内容为 Hello, Penguin"
```
Workspace 默认为当前目录,可用 `--workspace /path` 指定;目标目录必须已存在。
## 交互式对话
```bash
penguin chat
```
- 每输入一行即发起一个 Task。
- `/compact` 压缩上下文;`/exit` 或 `/quit` 退出;Ctrl-C 中断正在运行的 Task。
- 退出时会打印 `penguin chat --resume <sessionId>` 提示,用于恢复本次 Session;`--resume` 不带 id 时恢复该 Agent 最近的 Session。
## SDK 示例
安装 `@prismshadow/penguin-core` 后:
```ts
import { createAgent, isCompleteModelMessage, userText } from "@prismshadow/penguin-core";
const agent = await createAgent({ agentId: "default_agent" });
const session = await agent.createSession({ workspaceDir: process.cwd() });
for await (const output of session.run([userText("Create hello.txt containing hi")], {
approve: async () => "allow",
})) {
if (isCompleteModelMessage(output) && output.payload.type === "text") {
console.log(output.payload.text);
}
}
```
## 下一步
- [Web App 指南](/web-app):在浏览器中使用 PenguinHarness。
- [CLI 参考](/cli):完整命令与选项。
- [架构总览](/architecture):了解整体设计。
@@ -0,0 +1,73 @@
---
title: Self-Improvement
description: The Skill-orchestrated Benchmark and optimization loop: score, improve, snapshot, roll back.
---
Self-improvement in PenguinHarness is not carried by special-purpose engine code — it is carried by Skills orchestrating the ordinary Agent machinery: evaluations are ordinary Sessions, optimization is ordinary file editing, and orchestration uses the built-in `run_subagent` tool. The direct payoff is that the whole process shares the same observability and recovery machinery as everyday runs.
## Three roles
| Role | Responsibility |
| --- | --- |
| Target Agent | The Agent being improved; runs evaluation tasks only inside its own Workspace |
| Evaluator | Runs and scores one Benchmark Case run |
| Optimizer | Drives the whole optimization loop |
The roles are defined by Skills, not hardcoded: the Evaluator follows the `agent-evaluation` Skill, the Optimizer follows the `agent-optimization` Skill. This applies the design principle stated in the [Configuration Reference](/configuration) — an Agent's behavior is editable files on disk, which is what makes Agents improvable by Agents.
## The loop
1. `benchmark-design` builds a multi-Case capability Benchmark: repeated independent runs, with a traceable baseline calibrated first;
2. The Optimizer orchestrates Evaluators in parallel via the `run_subagent` tool, covering the Case × runs matrix;
3. Scores plus their linked Traces show where points were lost;
4. The Optimizer edits the Target Agent's editable state — `AGENTS.md`, Skills, config — to produce version N+1;
5. A Snapshot is taken before each round; the candidate version is kept only if the total score strictly improves, otherwise rolled back.
Benchmark optimization mode requires a complete baseline series in the scoreboard — without a calibrated baseline there is no improvement to compare against. Besides this loop, `agent-optimization` also supports a one-shot feedback mode: a concrete correction is applied directly as edits to the Target Agent's state, without going through the evaluation loop.
## Benchmark storage
Benchmarks are stored per Agent under `benchmarks/<id>/`:
```text
benchmarks/<id>/
├── benchmark_config.toml # Benchmark configuration (e.g. runs per Case)
├── <case-id>/
│ ├── statement/ # the task given to the Target Agent
│ └── rubric/ # private scoring rubric, isolated from the Target Agent
└── scoreboard.yaml # evaluation records (v2 format)
```
The separation of `rubric/` from `statement/` is deliberate: the Target Agent sees only the task statement and never touches the scoring rubric.
Each evaluation record in `scoreboard.yaml` (v2 format) is timestamped and carries:
- the paired model reference `(provider, model_id)` used for the round;
- `summary_title` and `summary` (the round's conclusion and the hypothesis for the next one);
- total score, cost, and duration — Case-level metrics are the average over its runs, evaluation-level metrics are the sum over its Cases;
- per-Case run details, each run recording `score`, `cost`, `duration_ms`, and `session_id`.
The built-in `default_agent` ships with an example Benchmark (`packages/core/src/state/example-benchmark.ts`) so the evaluation pages have data out of the box; the whole directory can be deleted or replaced at any time.
## Snapshots and versions
Before each optimization round, the Agent State is packed into `snapshots/v<version>.tar.gz` (excluding the Vault — secrets never enter a snapshot). The `version` in `system_config.yaml` increments on successful optimization. The Web UI supports exporting and importing snapshots; importing a version not higher than the current one requires explicit confirmation.
## Auditable end to end
- Every Evaluator run is an ordinary Session with a full Trace;
- Scoreboard records link back to those Sessions via `session_id`; see [Sessions & Traces](/sessions-and-traces);
- The Web evaluation pages are read-only views of these files; see the [Web App Guide](/web-app).
Scores are not black-box output: every number can be traced back to the run that produced it.
## Related Skills
| Skill | Purpose |
| --- | --- |
| `agent-creation` | Turn a requirement into a working Agent: write its `AGENTS.md`, install the Skills it needs |
| `benchmark-design` | Design and calibrate a multi-Case capability Benchmark |
| `agent-evaluation` | Run and score one isolated Benchmark Case run |
| `agent-optimization` | Improve an Agent from feedback or Benchmark results |
How Skills are organized and installed is covered in the [Skill System](/skills).
@@ -0,0 +1,73 @@
---
title: 自我进化
description: 由 Skill 编排的 Benchmark 评测与优化闭环:评分、改进、Snapshot 与回滚。
---
PenguinHarness 中的自我进化不依赖专用引擎代码,而是由 Skill 编排普通的 Agent 机制完成:评测是普通的 Session,优化是普通的文件编辑,编排靠内置的 `run_subagent` 工具。这样做的直接收益是——整个过程与日常运行共用同一套可观测性与恢复机制。
## 三个角色
| 角色 | 职责 |
| --- | --- |
| Target Agent | 被改进的 Agent,只在自己的 Workspace 里执行评测任务 |
| Evaluator | 执行并评分一次 Benchmark Case 运行 |
| Optimizer | 驱动整个优化循环 |
角色由 Skill 定义而非硬编码:Evaluator 遵循 `agent-evaluation` Skill,Optimizer 遵循 `agent-optimization` Skill。这正是[配置参考](/configuration)所述设计原则的应用——Agent 的行为是磁盘上的可编辑文件,所以 Agent 可以被 Agent 改进。
## 优化循环
1. `benchmark-design` 构建多 Case 的能力 Benchmark:重复独立运行,先校准出可追溯的基线;
2. Optimizer 通过 `run_subagent` 工具并行编排 Evaluator,覆盖 Case × 运行次数矩阵;
3. 得分与其关联的 Trace 共同指出失分位置;
4. Optimizer 编辑 Target Agent 的可编辑状态——`AGENTS.md`、Skills、配置——产出版本 N+1;
5. 每轮开始前先打 Snapshot;总分严格提升才保留候选版本,否则回滚。
Benchmark 优化模式要求 scoreboard 中已有完整的基线序列——没有校准过的基线,就没有可比较的提升。除此之外 `agent-optimization` 还支持一次性反馈模式:把一条具体的纠正意见直接落实为对 Target Agent 状态的编辑,不经过评测循环。
## Benchmark 存储
Benchmark 按 Agent 存放在 `benchmarks/<id>/` 下:
```text
benchmarks/<id>/
├── benchmark_config.toml # Benchmark 配置(如每个 Case 的运行次数 runs)
├── <case-id>/
│ ├── statement/ # 交给 Target Agent 的任务描述
│ └── rubric/ # 私有评分标准,对 Target Agent 隔离
└── scoreboard.yaml # 评测记录(v2 格式)
```
`rubric/` 与 `statement/` 的隔离是刻意设计:Target Agent 只能看到题面,永远接触不到评分标准。
`scoreboard.yaml`(v2 格式)中的每条评测记录带时间戳,并记录:
- 本轮使用的模型成对引用 `(provider, model_id)`;
- `summary_title` 与 `summary`(本轮结论与下一轮假设);
- 总分、成本与耗时——Case 级指标是各次运行的平均值,评测级指标是各 Case 的加和;
- 每个 Case 的逐次运行明细,每次运行含 `score`、`cost`、`duration_ms` 与 `session_id`。
内置的 `default_agent` 预置了一个示例 Benchmark(`packages/core/src/state/example-benchmark.ts`),评测页面开箱即有数据;整个目录可随时删除或替换。
## Snapshot 与版本
每轮优化前,Agent State 被打包为 `snapshots/v<version>.tar.gz`(Vault 除外——密钥永不进入快照)。`system_config.yaml` 的 `version` 在优化成功后自增。Web UI 支持导出与导入快照,导入版本不高于当前版本时需要显式确认。
## 全程可审计
- 每次 Evaluator 运行都是一个普通的 Session,留有完整 Trace;
- scoreboard 记录通过 `session_id` 链接回这些 Session,见 [Session 与 Trace](/sessions-and-traces);
- Web 的评测页面是这些文件的只读视图,见 [Web App 指南](/web-app)。
分数不是黑盒输出:任何一个数字都可以回溯到产生它的那次运行。
## 相关 Skill
| Skill | 用途 |
| --- | --- |
| `agent-creation` | 把需求变成可用的 Agent:撰写其 `AGENTS.md`、安装所需 Skill |
| `benchmark-design` | 设计并校准多 Case 的能力 Benchmark |
| `agent-evaluation` | 隔离执行并评分一次 Benchmark Case 运行 |
| `agent-optimization` | 根据反馈或 Benchmark 结果改进 Agent |
Skill 的组织与安装方式见[技能系统](/skills)。
+236
View File
@@ -0,0 +1,236 @@
---
title: Server API
description: HTTP API reference — authentication, routes, the SSE streaming protocol, and DTO type imports.
---
The PenguinHarness server exposes a same-origin HTTP API used by the bundled Web App and by any other HTTP client. This page is the reference: authentication, route tables, and the SSE streaming protocol. For starting the server, see the [Quickstart](/quickstart).
## Overview
- Stack: Hono + @hono/node-server, requires Node >= 24;
- Storage: SQLite (built-in `node:sqlite`, WAL mode) holds only indexes and aggregates — users, auth sessions, Project authorization, Agent / Session indexes, usage, UI preferences, error records, and Schedule state; all Agent, Trace, and Workspace data stays as files under `~/.penguin/data`, shared with the CLI / SDK — see the [Configuration Reference](/configuration);
- Binding: defaults to `127.0.0.1:7364`, adjustable via the `PORT` / `HOST` environment variables;
- Request bodies: writes accept JSON only (Content-Type check, one of the CSRF defenses), capped at 20MB;
- Errors share a single shape:
```text
{ "error": { "code": "<machine-readable code>", "message": "<user-facing text>" } }
```
## Source layout
```text
packages/server/src
├── index.ts / config.ts / app.ts # startup entry · env config · Hono assembly (createApp binds no port — testable)
├── api/types.ts # the outward DTO contract (type-only import via the "./api" subpath)
├── auth/ # scrypt passwords, admin seeding, cookie sessions, auth middleware
├── db/ # node:sqlite connection, schema SQL, one repo per table
├── http/ # error bodies, request validation, SSE adapter, routes/ all route groups
├── runtime/ # session-manager (runtime driving) · channel (SSE ring buffer)
│ # approvals · usage-recorder · scheduler · title-generator
└── services/ # authorization rules, TOML/YAML config IO, Session/Trace/usage/snapshot services
```
## Authentication
- Cookie session: `penguin_session` (HttpOnly, SameSite=Lax), valid for 7 days with sliding renewal;
- Passwords are stored as scrypt hashes; the server keeps only the sha256 of the session token, never the plaintext;
- No open registration: the built-in admin `admin` / `admin123` is seeded at startup, and all other accounts are created by an admin;
- Same-origin only — no CORS middleware is enabled.
```bash
curl -c cookies.txt -H "Content-Type: application/json" \
-d '{"userId":"admin","password":"admin123"}' \
http://127.0.0.1:7364/api/auth/login
```
## Route Reference
### Auth and Account
| Method | Path | Description |
| --- | --- | --- |
| POST | /api/auth/login | Log in: `{userId, password}` → `{user}` |
| POST | /api/auth/logout | Log out, returns 204 |
| GET | /api/me | Current user info |
| PUT | /api/me/password | Change password: `{oldPassword, newPassword}` |
| GET | /api/me/prefs | Read UI preferences |
| PUT | /api/me/prefs | Write UI preferences (shallow merge) |
### User Administration (admin only)
| Method | Path | Description |
| --- | --- | --- |
| GET | /api/admin/users | List users |
| POST | /api/admin/users | Create a user: `{userId, password}` |
| POST | /api/admin/users/:userId/password | Reset a password (invalidates all of that user's login sessions) |
| DELETE | /api/admin/users/:userId | Delete a user |
### Projects and Members
| Method | Path | Description |
| --- | --- | --- |
| GET | /api/projects | Projects visible to the current user |
| POST | /api/projects | Create a Project |
| DELETE | /api/projects/:projectId | Delete a Project |
| GET | /api/projects/:projectId/members | List members |
| POST | /api/projects/:projectId/members | Add a member: `{userId}` |
| DELETE | /api/projects/:projectId/members/:userId | Remove a member |
Member writes are owner-only.
### Models
| Method | Path | Description |
| --- | --- | --- |
| GET | /api/projects/:projectId/models | List models (api_key masked) |
| PUT | /api/projects/:projectId/models | Full-table replace, keyed by `(provider, modelId)` |
| POST | /api/projects/:projectId/models/test | Connectivity test: `{provider, modelId, …}` → `{ok, latencyMs?, message?}` |
### Agents
The paths below omit the `/api/projects/:projectId` prefix.
| Method | Path | Description |
| --- | --- | --- |
| GET / POST | /agents | List / create Agents |
| DELETE | /agents/:agentId | Delete an Agent |
| GET / PUT | /agents/:agentId/config | Read / write config (AGENTS.md + system_config.yaml; PUT preserves YAML comments) |
| GET / PUT | /agents/:agentId/vault | Vault environment variables (values masked; PUT is a full replace) |
| GET | /agents/:agentId/export | Export the Agent State snapshot (tar.gz download) |
| POST | /agents/:agentId/import | Import a snapshot: `{dataBase64, confirm?}`; 409 on version conflict without confirm |
| GET / POST | /agents/:agentId/skills | List / install installed Skills |
| DELETE | /agents/:agentId/skills/:name | Uninstall a Skill |
| GET | /agents/:agentId/benchmarks | Benchmark scoring data (read-only) |
### Schedules
| Method | Path | Description |
| --- | --- | --- |
| GET / POST | /agents/:agentId/schedules | List scheduled tasks / create one (409 if the name exists) |
| GET / PUT / DELETE | /agents/:agentId/schedules/:name | Read / update / delete a single task |
Schedule writes are owner-only.
### Session Creation and Directory Browsing
| Method | Path | Description |
| --- | --- | --- |
| GET | /agents/:agentId/sessions | List Sessions (including run state) |
| POST | /agents/:agentId/sessions | Create a Session: `{modelId?, provider?, workspace?, approvalMode?}` → 201 |
| GET | /dirs?path= | Server-side directory browser (backs the Workspace picker) |
On Session creation, the model defaults to the Project's default model, the Workspace defaults to an auto-created temporary directory, and the approval mode defaults to `allow-all`.
### Usage and Traces (Agent Level)
| Method | Path | Description |
| --- | --- | --- |
| GET | /usage | Usage statistics; query parameters `from`, `to`, `groupBy`, `agentId`, `provider`, `modelId` |
| GET | /agents/:agentId/traces | Date → Session drill-down structure of Trace files |
| GET | /agents/:agentId/traces/:sessionId/:index | Read Trace events (`offset` / `limit` pagination) |
| GET | /agents/:agentId/traces/:sessionId/:index/analysis | Trace performance analysis |
### Session-Level Endpoints
The paths below omit the `/api/sessions/:sessionId` prefix. For the storage model behind Sessions and Traces, see [Sessions and Traces](/sessions-and-traces).
| Method | Path | Description |
| --- | --- | --- |
| GET | / | Session info |
| PATCH | / | Update: `{approvalMode?, archived?, title?}` |
| DELETE | / | Delete the Session (along with its Traces and scratch files) |
| GET | /messages | Full OmniMessage history |
| GET | /stream | SSE event stream (next section) |
| POST | /tasks | Start a Task: `{input: TaskInputPart[]}` → 202 |
| POST | /approvals/:toolCallId | Approval decision: `{decision}` is `allow` or `deny` → 204 |
| POST | /abort | Interrupt the current Task: 202 when triggered, 204 when idle |
| POST | /compact | Trigger context compaction: 202; 409 `nothing_to_compact` when there is nothing to compact |
| GET | /files?path= | Browse the Workspace directory |
| GET | /files/content?path=&download= | Read a Workspace file (`download=1` serves it as an attachment) |
| POST | /files/stat | Batch existence check: `{paths}` |
| PUT | /files/content?path= | Upload a file: `{dataBase64}`, capped at 14MB |
| GET | /traces | List this Session's Trace files |
| GET | /traces/:index | Read Trace events (paginated) |
| GET | /traces/:index/analysis | Trace performance analysis |
| GET | /scratchpad/:fileName | Read a session scratch file (e.g. input images) |
General conventions: Sessions the user cannot access always return 404 — their existence is never leaked; only one Task or compaction runs per Session at a time, and conflicts return 409 (`task_in_progress` / `compacting`).
Key request bodies (explicit keys):
```ts
// POST /api/sessions/:sessionId/tasks — start a Task
interface TaskCreateRequest {
input: TaskInputPart[];
}
type TaskInputPart =
| { type: "text"; text: string }
| { type: "image_url"; imageUrl: string }; // pasted images arrive as data URLs
// POST /api/sessions/:sessionId/approvals/:toolCallId
interface ApprovalDecisionRequest {
decision: "allow" | "deny";
}
```
## Streaming (SSE)
Real-time delivery uses Server-Sent Events, not WebSocket, on two channels (the ordering semantics of what the channels carry are on [Message Flow & Ordering](/message-flow)):
| Channel | Path | Contents |
| --- | --- | --- |
| Per Session | GET /api/sessions/:sessionId/stream | The Session's message stream and run events |
| Per user | GET /api/events | `hello` handshake and cross-Session notifications (schedule_fired / schedule_queued / session_created) |
### Wire Format
Default (unnamed) SSE events carry raw OmniMessage envelopes as single-line JSON — the same protocol the SDK yields and the Trace stores, see the [OmniMessage Protocol](/omni-message). Events named `server_event` carry the ServerEvent union:
```ts
export type ServerEvent =
| { type: "approval_request"; toolCall: OmniMessage<ToolCallPayload>; origin?: string[] }
| { type: "task_state"; state: "idle" | "running" | "compacting" }
| { type: "session_title"; sessionId: string; title: string }
| { type: "resync_required" }
| { type: "hello" }
| { type: "session_created"; projectId: string; agentId: string; sessionId: string; source: SessionSource }
| { type: "schedule_fired"; projectId: string; agentId: string; name: string; sessionId: string }
| { type: "schedule_queued"; projectId: string; agentId: string; name: string; sessionId: string };
```
| Event | Fired when |
| --- | --- |
| approval_request | A tool call escalated to human approval: every call under always-ask, plus rw / unknown-permission calls under read-only; pending approvals are resent on reconnect |
| task_state | The Session's run state flips (idle / running / compacting) |
| session_title | The model-generated title after the first turn has been persisted |
| resync_required | The Last-Event-ID was evicted from the buffer; the client must refetch history |
| hello | Handshake on the user channel |
| session_created | A new Session was registered (e.g. a subagent session) |
| schedule_fired | A scheduled task fired and was delivered |
| schedule_queued | The target Session is running; this firing was queued |
### Delivery Guarantees
- Event ids are monotonic per channel, shaped `<epoch>-<seq>`;
- Each channel keeps a bounded replay buffer (most recent 1000 events or 2MB);
- Reconnecting with `Last-Event-ID` replays the gap on a buffer hit; on a miss the server first sends `resync_required`, and the client refetches `/messages` before continuing;
- A heartbeat comment line is written every 20 seconds;
- Event order: on a reconnect carrying `Last-Event-ID`, **the replayed gap (or `resync_required`) arrives first**, then the initial events — the authoritative `task_state` snapshot and still-pending approval_requests — then the live stream. A fresh connection (no `Last-Event-ID`) skips replay, so its first event is the `task_state` snapshot.
### Recommended Client Pattern
The order the bundled Web App uses:
1. Connect `/stream` first and buffer incoming events;
2. GET `/messages` for the full history;
3. Replay the buffer, deduplicating the overlap;
4. Go live.
## Type Imports
All DTO types are importable type-only from the server package's `@prismshadow/penguin-server/api` subpath:
```ts
import type { ServerEvent, SessionInfo } from "@prismshadow/penguin-server/api";
```
+236
View File
@@ -0,0 +1,236 @@
---
title: Server API
description: HTTP API 参考:认证机制、路由列表、SSE 流式协议与 DTO 类型导入。
---
PenguinHarness Server 提供一套同源 HTTP API,自带的 Web App 与其他 HTTP 客户端都通过它访问。本文是接口参考:认证机制、路由列表与 SSE 流式协议。服务启动方式见[快速开始](/quickstart)。
## 总览
- 技术栈:Hono + @hono/node-server,要求 Node >= 24;
- 存储:SQLite(内置 `node:sqlite`,WAL 模式)仅存放索引与聚合数据——用户、登录会话、Project 授权、Agent / Session 索引、用量、UI 偏好、错误记录与 Schedule 状态;Agent、Trace 与 Workspace 数据全部以文件形式存放在 `~/.penguin/data` 下,与 CLI / SDK 共享,见[配置参考](/configuration);
- 监听:默认 `127.0.0.1:7364`,可用环境变量 `PORT` / `HOST` 调整;
- 请求体:写请求仅接受 JSON(Content-Type 校验,CSRF 防线之一),上限 20MB;
- 错误响应统一为:
```text
{ "error": { "code": "<机器可读错误码>", "message": "<提示文案>" } }
```
## 目录结构
```text
packages/server/src
├── index.ts / config.ts / app.ts # 启动入口 · 环境变量配置 · Hono 组装(createApp 不绑端口,便于测试)
├── api/types.ts # 对外 DTO 契约(经 "./api" 子路径供前端 type-only 引用)
├── auth/ # scrypt 密码、admin 种子、cookie 会话、认证中间件
├── db/ # node:sqlite 连接、建表 SQL、每表一个 repo
├── http/ # 错误体、请求校验、SSE 适配、routes/ 全部路由
├── runtime/ # session-manager(运行时驱动)· channel(SSE 环形缓冲)
│ # approvals · usage-recorder · scheduler · title-generator
└── services/ # 授权规则、TOML/YAML 配置读写、Session/Trace/用量/快照服务
```
## 认证
- Cookie 会话:`penguin_session`(HttpOnly、SameSite=Lax),有效期 7 天,滑动续期;
- 密码以 scrypt 哈希存储;服务端只保存会话 Token 的 sha256,不落明文;
- 不开放注册:启动时种子化内置管理员 `admin` / `admin123`,其余账号由管理员创建;
- 仅限同源访问,未启用 CORS 中间件。
```bash
curl -c cookies.txt -H "Content-Type: application/json" \
-d '{"userId":"admin","password":"admin123"}' \
http://127.0.0.1:7364/api/auth/login
```
## 路由参考
### 认证与账户
| 方法 | 路径 | 说明 |
| --- | --- | --- |
| POST | /api/auth/login | 登录:`{userId, password}` → `{user}` |
| POST | /api/auth/logout | 退出登录,返回 204 |
| GET | /api/me | 当前用户信息 |
| PUT | /api/me/password | 修改密码:`{oldPassword, newPassword}` |
| GET | /api/me/prefs | 读取 UI 偏好 |
| PUT | /api/me/prefs | 写入 UI 偏好(浅合并) |
### 用户管理(仅管理员)
| 方法 | 路径 | 说明 |
| --- | --- | --- |
| GET | /api/admin/users | 用户列表 |
| POST | /api/admin/users | 创建用户:`{userId, password}` |
| POST | /api/admin/users/:userId/password | 重置密码(该用户全部登录会话失效) |
| DELETE | /api/admin/users/:userId | 删除用户 |
### Project 与成员
| 方法 | 路径 | 说明 |
| --- | --- | --- |
| GET | /api/projects | 当前用户可见的 Project 列表 |
| POST | /api/projects | 创建 Project |
| DELETE | /api/projects/:projectId | 删除 Project |
| GET | /api/projects/:projectId/members | 成员列表 |
| POST | /api/projects/:projectId/members | 添加成员:`{userId}` |
| DELETE | /api/projects/:projectId/members/:userId | 移除成员 |
成员写操作仅限 Owner。
### 模型
| 方法 | 路径 | 说明 |
| --- | --- | --- |
| GET | /api/projects/:projectId/models | 模型列表(api_key 掩码显示) |
| PUT | /api/projects/:projectId/models | 全表替换,条目以 `(provider, modelId)` 为键 |
| POST | /api/projects/:projectId/models/test | 连通性测试:`{provider, modelId, …}` → `{ok, latencyMs?, message?}` |
### Agent
以下路径均省略前缀 `/api/projects/:projectId`。
| 方法 | 路径 | 说明 |
| --- | --- | --- |
| GET / POST | /agents | Agent 列表 / 创建 |
| DELETE | /agents/:agentId | 删除 Agent |
| GET / PUT | /agents/:agentId/config | 读写配置(AGENTS.md + system_config.yaml,PUT 保留 YAML 注释) |
| GET / PUT | /agents/:agentId/vault | Vault 环境变量(值掩码显示;PUT 全表替换) |
| GET | /agents/:agentId/export | 导出 Agent State 快照(tar.gz 下载) |
| POST | /agents/:agentId/import | 导入快照:`{dataBase64, confirm?}`;版本冲突且未确认时返回 409 |
| GET / POST | /agents/:agentId/skills | 已安装 Skill 列表 / 安装 |
| DELETE | /agents/:agentId/skills/:name | 卸载 Skill |
| GET | /agents/:agentId/benchmarks | Benchmark 评分数据(只读) |
### Schedule
| 方法 | 路径 | 说明 |
| --- | --- | --- |
| GET / POST | /agents/:agentId/schedules | 定时任务列表 / 创建(重名返回 409) |
| GET / PUT / DELETE | /agents/:agentId/schedules/:name | 读取 / 更新 / 删除单个任务 |
Schedule 写操作仅限 Owner。
### Session 创建与目录浏览
| 方法 | 路径 | 说明 |
| --- | --- | --- |
| GET | /agents/:agentId/sessions | Session 列表(含运行状态) |
| POST | /agents/:agentId/sessions | 创建 Session:`{modelId?, provider?, workspace?, approvalMode?}` → 201 |
| GET | /dirs?path= | 服务器端目录浏览(Workspace 选择器数据源) |
创建 Session 时,模型默认取 Project 默认模型,Workspace 默认自动创建临时目录,审批模式默认 `allow-all`。
### 用量与 Trace(Agent 级)
| 方法 | 路径 | 说明 |
| --- | --- | --- |
| GET | /usage | 用量统计,查询参数 `from`、`to`、`groupBy`、`agentId`、`provider`、`modelId` |
| GET | /agents/:agentId/traces | Trace 文件的日期 → Session 下钻结构 |
| GET | /agents/:agentId/traces/:sessionId/:index | 读取 Trace 事件(`offset` / `limit` 分页) |
| GET | /agents/:agentId/traces/:sessionId/:index/analysis | Trace 性能分析结果 |
### Session 级接口
以下路径均省略前缀 `/api/sessions/:sessionId`。Trace 与 Session 的存储模型见 [Session 与 Trace](/sessions-and-traces)。
| 方法 | 路径 | 说明 |
| --- | --- | --- |
| GET | / | Session 信息 |
| PATCH | / | 更新:`{approvalMode?, archived?, title?}` |
| DELETE | / | 删除 Session(连同 Trace 与暂存文件) |
| GET | /messages | 完整 OmniMessage 历史 |
| GET | /stream | SSE 事件流(见下节) |
| POST | /tasks | 发起 Task:`{input: TaskInputPart[]}` → 202 |
| POST | /approvals/:toolCallId | 审批决定:`{decision}` 取 `allow` 或 `deny` → 204 |
| POST | /abort | 中断当前 Task:已触发返回 202,无任务返回 204 |
| POST | /compact | 触发上下文压缩:202;无可压缩内容返回 409 `nothing_to_compact` |
| GET | /files?path= | 浏览 Workspace 目录 |
| GET | /files/content?path=&download= | 读取 Workspace 文件(`download=1` 时作为附件下载) |
| POST | /files/stat | 批量存在性检查:`{paths}` |
| PUT | /files/content?path= | 上传文件:`{dataBase64}`,上限 14MB |
| GET | /traces | 本 Session 的 Trace 文件列表 |
| GET | /traces/:index | 读取 Trace 事件(分页) |
| GET | /traces/:index/analysis | Trace 性能分析结果 |
| GET | /scratchpad/:fileName | 读取会话暂存文件(如输入图片) |
通用约定:无权访问的 Session 一律返回 404,不泄露其存在性;每个 Session 同时只允许一个 Task 或压缩在运行,冲突时返回 409(`task_in_progress` / `compacting`)。
关键请求体(明确键名):
```ts
// POST /api/sessions/:sessionId/tasks —— 发起一个 Task
interface TaskCreateRequest {
input: TaskInputPart[];
}
type TaskInputPart =
| { type: "text"; text: string }
| { type: "image_url"; imageUrl: string }; // 粘贴图片以 data URL 上送
// POST /api/sessions/:sessionId/approvals/:toolCallId
interface ApprovalDecisionRequest {
decision: "allow" | "deny";
}
```
## 流式接口(SSE)
实时通道采用 Server-Sent Events 而非 WebSocket,共两条(通道内承载的消息顺序语义见[消息流转与时序](/message-flow)):
| 通道 | 路径 | 内容 |
| --- | --- | --- |
| Session 级 | GET /api/sessions/:sessionId/stream | 该 Session 的消息流与运行事件 |
| 用户级 | GET /api/events | `hello` 握手与跨 Session 通知(schedule_fired / schedule_queued / session_created) |
### 传输格式
默认(未命名)SSE 事件承载原始 OmniMessage 信封(单行 JSON)——与 SDK 产出、Trace 落盘是同一套协议,见 [OmniMessage 协议](/omni-message);命名为 `server_event` 的事件承载 ServerEvent 联合类型:
```ts
export type ServerEvent =
| { type: "approval_request"; toolCall: OmniMessage<ToolCallPayload>; origin?: string[] }
| { type: "task_state"; state: "idle" | "running" | "compacting" }
| { type: "session_title"; sessionId: string; title: string }
| { type: "resync_required" }
| { type: "hello" }
| { type: "session_created"; projectId: string; agentId: string; sessionId: string; source: SessionSource }
| { type: "schedule_fired"; projectId: string; agentId: string; name: string; sessionId: string }
| { type: "schedule_queued"; projectId: string; agentId: string; name: string; sessionId: string };
```
| 事件 | 触发时机 |
| --- | --- |
| approval_request | 工具调用升级为人工审批时发出:always-ask 下的所有调用,以及 read-only 下 rw / 未知权限的调用;重连时未决审批会重发 |
| task_state | Session 运行状态翻转(idle / running / compacting) |
| session_title | 首轮后模型生成的标题已持久化 |
| resync_required | Last-Event-ID 已被缓冲区淘汰,客户端须重新拉取历史 |
| hello | 用户通道连接握手 |
| session_created | 新 Session 注册(如子 Agent 会话) |
| schedule_fired | 定时任务已触发并发送 |
| schedule_queued | 目标 Session 正在运行,本次触发已排队 |
### 投递保证
- 事件 id 按通道单调递增,形如 `<epoch>-<seq>`;
- 每通道维护有界重放缓冲(最近 1000 条事件或 2MB);
- 携带 `Last-Event-ID` 重连时,命中缓冲则补发缺口;未命中则先发 `resync_required`,客户端重新拉取 `/messages` 后继续消费;
- 每 20 秒写一条心跳注释行;
- 事件次序:带 `Last-Event-ID` 重连时,**补发的缺口(或 `resync_required`)最先送达**,随后才是初始事件——权威的 `task_state` 快照与未决的 approval_request,再进入实时流;全新连接(无 `Last-Event-ID`)不重放缓冲,首个事件即为 `task_state` 快照。
### 推荐客户端模式
自带 Web App 的接入顺序:
1. 先连接 `/stream` 并缓冲收到的事件;
2. 再 GET `/messages` 拉取完整历史;
3. 回放缓冲区并对重叠消息去重;
4. 转入实时消费。
## 类型导入
全部 DTO 类型可从服务端包的子路径 `@prismshadow/penguin-server/api` 以 type-only 方式导入:
```ts
import type { ServerEvent, SessionInfo } from "@prismshadow/penguin-server/api";
```
@@ -0,0 +1,88 @@
---
title: Sessions & Traces
description: The six-level run model, local data directory layout, Trace file design, and Session recovery.
---
All PenguinHarness runtime data lives on the local file system: configuration is editable files, history is append-only Traces. This page defines each level of the run model and explains how the Trace serves as history, recovery source, and statistics source at once.
## Run model
Six levels: Project → Agent → Workspace → Session → Task → Request.
| Concept | Definition |
| --- | --- |
| Project | Top-level unit organizing Agents; owns the model and credential configuration; in the multi-user Web setup, users and Projects are many-to-many |
| Agent | The executing subject; has exactly one Agent State (a persistent directory); one Agent can serve many Workspaces |
| Workspace | The working directory of one run — the only file scope the model sees; an explicit `workspaceDir` must already exist, otherwise a temp Workspace `workspaces/tmp-<8hex>` is created |
| Session | A continuous conversation under one (Agent, Workspace); model and Workspace are locked at Session creation; ids look like `session-YYYY-MM-DD-HH-mm-ss-<8hex>` |
| Task | One execution goal started by one Prompt; consists of one or more consecutive Requests |
| Request | One LLM API call: context and tool definitions in, streamed output out |
See the [Architecture](/architecture) page for how the levels cooperate, and the [Agent Loop](/agent-loop) for how Requests advance within a Task.
## Data layout
The data root is the `PENGUIN_HOME` environment variable, defaulting to `~/.penguin/data`. The layout is defined in one place, `packages/core/src/state/paths.ts`:
```text
<root>/<project>/
├── .project_config.toml # Project-level models & credentials (hidden file, 0600)
└── agents/
└── <agent>/
├── agent_state/ # system_config.yaml, AGENTS.md, .vault.toml,
│ # tools/, memory/, skills/, schedule/
├── traces/
│ └── <yyyy-mm-dd>/<sessionId>_<index3>.jsonl
├── scratchpad/ # temp files, one subdirectory per Session id (e.g. pasted images)
├── workspaces/ # temp Workspaces (tmp-<8hex>)
├── benchmarks/ # capability Benchmark cases and scores
└── snapshots/ # Agent State version snapshots
```
See the [Configuration Reference](/configuration) for the fields of each config file.
## Trace design
A Trace is an append-only JSON Lines file; each line is one OmniMessage envelope (see the [OmniMessage Protocol](/omni-message)). History is only ever appended, never modified in place.
- One Trace file corresponds to one complete model context. When compaction produces a new context segment, the writer rotates to a new file — `_002`, `_003`, … — with an incrementing index.
- Recorded: `session_meta`, complete `model_msg`, and all `event_msg`.
- Not recorded: streaming `partial_*` fragments (the producer appends the complete message once the segment ends), and nested messages tagged with `origin` — a subagent's messages go to the child Session's own Trace, while the parent Trace keeps a single `subagent` pointer event at the spawn site recording the child Session id.
- `request_begin` and `request_end(status)` come in pairs delimiting one Request; replay uses `request_end.status === "completed"` as the commit criterion for that turn.
See `packages/core/src/trace/writer.ts` for the implementation.
The head of a Trace (illustrative; one OmniMessage envelope per line):
```jsonl
{"timestamp":"2026-07-18T03:10:22.531Z","type":"session_meta","payload":{"session_id":"session-2026-07-18-11-10-22-3f8a1c2d","provider":"deepseek","model_id":"deepseek-v4-pro","model_context_window":1000000,"system_prompt":"…","tools":[…],"thinking_level":"medium","agent_state":"/home/u/.penguin/data/default_project/agents/default_agent/agent_state","workspace":"/home/u/work"}}
{"timestamp":"…","type":"event_msg","payload":{"type":"request_begin"}}
{"timestamp":"…","type":"model_msg","payload":{"type":"text","role":"user","text":"Create hello.txt"}}
{"timestamp":"…","type":"model_msg","payload":{"type":"tool_call","role":"assistant","name":"exec_command","arguments":"{\"cmd\":\"printf hi > hello.txt\"}","tool_call_id":"call_0"}}
{"timestamp":"…","type":"event_msg","payload":{"type":"approval_decision","decision":"allow","tool_call_id":"call_0"}}
{"timestamp":"…","type":"model_msg","payload":{"type":"tool_call_output","role":"user","output":"[no output]","tool_call_id":"call_0"}}
{"timestamp":"…","type":"event_msg","payload":{"type":"request_end","status":"completed"}}
{"timestamp":"…","type":"event_msg","payload":{"type":"token_usage","session":{…},"request":{…}}}
```
## Session recovery
The Trace is the single source of truth for recovery — there is no separate session database to keep in sync. `resumeSession` works as follows:
1. Locate the highest-index Trace file of the Session;
2. Read the runtime configuration from its `session_meta` — model, system prompt, Workspace — all immutable for the lifetime of the Session;
3. Replay the committed history into a fresh LLM context;
4. Reconstruct the carry-over (undelivered tool outputs, interruption markers) plus turn and Token counters;
5. Continue appending to the same Trace file.
Recovery requires that the Workspace and the model still exist. What recovery guarantees is structural legality: only committed turns are replayed, with `tool_call` / `tool_call_output` pairing intact; incomplete model output (thinking, text) is allowed to be lost. A truncated last line left by an abnormal process exit is tolerated and ignored. See `packages/core/src/trace/resume.ts`.
Special case: if the latest Trace file ends with a completed compaction, that context is closed as a whole — resume starts from an empty context; in summarize mode the `<context_summary>` is reconstructed and prepended to the first input after resume.
## Field fidelity
Provider-specific fields (such as `signature` and `phase`) are preserved verbatim in the Trace and sent back verbatim — some models require them byte-for-byte on history replay, and any rewriting would break compatibility. This is one reason the Trace stores raw OmniMessage envelopes rather than a post-processed format.
## Observability
Every approval decision (`approval_decision`), abort (`abort`), compaction (`compaction_begin` / `compaction_end`), and `token_usage` lands in the Trace as an event. The Web Trace view and the usage/cost statistics are both derived from this same data — there is no second source of truth; see the [Web App Guide](/web-app). The approval mechanism itself is covered in [Tools & Approval](/tools).
@@ -0,0 +1,88 @@
---
title: Session 与 Trace
description: 运行模型的六层结构、本地数据目录布局、Trace 文件设计与 Session 恢复机制。
---
PenguinHarness 的全部运行数据都落在本地文件系统:配置是可编辑文件,历史是 append-only 的 Trace。本页定义运行模型的各层概念,并说明 Trace 如何同时充当历史记录、恢复依据与统计来源。
## 运行模型
六层结构:Project → Agent → Workspace → Session → Task → Request。
| 概念 | 定义 |
| --- | --- |
| Project | 组织 Agent 的顶层单位,持有模型与凭证配置;Web 多用户部署中,用户与 Project 是多对多关系 |
| Agent | 执行主体,恰好拥有一份 Agent State(持久化目录);一个 Agent 可服务多个 Workspace |
| Workspace | 一次运行的工作目录,是模型可见的唯一文件范围;显式指定的 `workspaceDir` 必须已存在,未指定时自动创建临时 Workspace `workspaces/tmp-<8hex>` |
| Session | 同一(Agent、Workspace)下的一段连续对话;模型与 Workspace 在 Session 创建时锁定,id 形如 `session-YYYY-MM-DD-HH-mm-ss-<8hex>` |
| Task | 由一条 Prompt 发起的一个执行目标,由一个或多个连续的 Request 组成 |
| Request | 一次 LLM API 调用:上下文与工具定义送入,流式输出返回 |
各层组件如何协作见[架构总览](/architecture);Task 内部 Request 如何推进见 [Agent 运行循环](/agent-loop)。
## 数据目录
数据根目录取环境变量 `PENGUIN_HOME`,缺省 `~/.penguin/data`。目录布局由 `packages/core/src/state/paths.ts` 统一定义:
```text
<root>/<project>/
├── .project_config.toml # Project 级模型与凭证(隐藏文件,0600)
└── agents/
└── <agent>/
├── agent_state/ # system_config.yaml、AGENTS.md、.vault.toml、
│ # tools/、memory/、skills/、schedule/
├── traces/
│ └── <yyyy-mm-dd>/<sessionId>_<index3>.jsonl
├── scratchpad/ # 临时文件,按 Session id 建子目录(如粘贴的图片)
├── workspaces/ # 临时 Workspace(tmp-<8hex>)
├── benchmarks/ # 能力评测题库与得分
└── snapshots/ # Agent State 版本快照
```
配置文件的字段详见[配置参考](/configuration)。
## Trace 设计
Trace 是 append-only 的 JSON Lines 文件,每行一个 OmniMessage 信封(协议见 [OmniMessage 协议](/omni-message))。历史事件只追加、从不原地修改。
- 一个 Trace 文件对应一份完整的模型上下文。上下文压缩产生新的上下文段时,写入器轮转到 `_002`、`_003`……新文件,索引递增。
- 记录的消息:`session_meta`、完整的 `model_msg`、全部 `event_msg`。
- 不记录的消息:流式 `partial_*` 分片(片段结束后由生产方补写完整消息);带 `origin` 标记的嵌套消息——子 Agent 的消息写入子 Session 自己的 Trace,父 Trace 只在派生位置保留一个 `subagent` 指针事件,记录子 Session id。
- `request_begin` 与 `request_end(status)` 成对出现,界定一轮 Request;回放以 `request_end.status === "completed"` 作为该轮已提交的判据。
实现见 `packages/core/src/trace/writer.ts`。
一条 Trace 的开头(示意,每行一个 OmniMessage 信封):
```jsonl
{"timestamp":"2026-07-18T03:10:22.531Z","type":"session_meta","payload":{"session_id":"session-2026-07-18-11-10-22-3f8a1c2d","provider":"deepseek","model_id":"deepseek-v4-pro","model_context_window":1000000,"system_prompt":"…","tools":[…],"thinking_level":"medium","agent_state":"/home/u/.penguin/data/default_project/agents/default_agent/agent_state","workspace":"/home/u/work"}}
{"timestamp":"…","type":"event_msg","payload":{"type":"request_begin"}}
{"timestamp":"…","type":"model_msg","payload":{"type":"text","role":"user","text":"创建 hello.txt"}}
{"timestamp":"…","type":"model_msg","payload":{"type":"tool_call","role":"assistant","name":"exec_command","arguments":"{\"cmd\":\"printf hi > hello.txt\"}","tool_call_id":"call_0"}}
{"timestamp":"…","type":"event_msg","payload":{"type":"approval_decision","decision":"allow","tool_call_id":"call_0"}}
{"timestamp":"…","type":"model_msg","payload":{"type":"tool_call_output","role":"user","output":"[no output]","tool_call_id":"call_0"}}
{"timestamp":"…","type":"event_msg","payload":{"type":"request_end","status":"completed"}}
{"timestamp":"…","type":"event_msg","payload":{"type":"token_usage","session":{…},"request":{…}}}
```
## Session 恢复
Trace 是恢复的唯一事实来源,没有独立的会话数据库需要与之对齐。`resumeSession` 的流程:
1. 定位该 Session 索引最大的 Trace 文件;
2. 从文件内的 `session_meta` 读取运行配置——模型、系统提示词、Workspace,三者在 Session 生命周期内不可变;
3. 将已提交的历史回放进一份全新的 LLM 上下文;
4. 重建 carry-over(未送达的工具输出、中断标记)与轮数、Token 计数器;
5. 继续追加写入同一个 Trace 文件。
恢复的前提是 Workspace 与模型仍然存在。恢复保证的是结构合法性:只回放已提交的轮次,`tool_call` 与 `tool_call_output` 配对完整;未完成的模型输出(thinking、文本)允许丢失。异常退出留下的截断末行会被容忍并忽略。实现见 `packages/core/src/trace/resume.ts`。
特殊情形:若最新 Trace 文件以一次完成的压缩收尾,则该上下文已整体关闭——恢复从空上下文开始;summarize 模式下会重建 `<context_summary>` 摘要,前置到恢复后第一轮输入中。
## 字段保真
Provider 专有字段(如 `signature`、`phase`)在 Trace 中原样保存、原样回传——部分模型在历史回放时要求这些字段逐字一致,任何转写都会破坏兼容性。这也是 Trace 直接存储 OmniMessage 信封而非二次加工格式的原因之一。
## 可观测性
每一次审批决策(`approval_decision`)、中断(`abort`)、压缩(`compaction_begin` / `compaction_end`)与 `token_usage` 都作为事件落入 Trace。Web 的 Trace 视图与用量、成本统计均由这同一份数据派生,不存在第二事实来源;见 [Web App 指南](/web-app)。审批机制本身见[工具与审批](/tools)。
+76
View File
@@ -0,0 +1,76 @@
---
title: Skills
description: Skills package reusable instructions as directories with a SKILL.md — metadata up front, body on demand, editable by the Agent itself.
---
## Anatomy of a Skill
A Skill is a directory containing a `SKILL.md`, optionally with a custom `icon.svg`. The directory name is the authoritative skill name and must match `^[A-Za-z0-9_-]+$`; a `name` in the frontmatter is overridden by it.
Frontmatter fields:
| Field | Meaning |
| --- | --- |
| `name` | Skill name, matching the directory name |
| `description` | English one-liner injected into the system prompt |
| `short_description` / `short_description_zh` | UI labels for compact spots such as cards; not injected into the prompt |
| `version` | Natural-number version, default 1 |
| `updated` | Update date |
```md
---
name: my-skill
description: One-line English description injected into the system prompt.
short_description: Short UI label.
short_description_zh: 简短的中文标签。
version: 1
updated: 2026-07-17
---
# My Skill
Concrete steps, boundaries and acceptance criteria...
```
Parsing is tolerant: only `key: value` scalar lines inside the first `---` block are recognized; a `version` that is not a natural number falls back to 1, and a missing `updated` defaults to empty.
## Progressive loading
Skills follow an "index first, body on demand" design: the system prompt injects only each installed Skill's metadata (name + description) through the `{{SKILL_METADATA}}` placeholder, and instructs the model to read the matching `SKILL.md` in full via the shell before following it. There is no dedicated skill tool — reading the body is just one `exec_command` call (see [Tools & Approval](/tools)).
Chat can also pin skills explicitly: the message then starts with a `<use_skills>` block listing the skill names.
If a message only names a skill without a concrete task, the model is instructed to ask what is needed before starting.
## Installation and storage
Installed Skills live under `agent_state/skills/<name>/` inside the Agent State. The files are the source of truth: every read goes straight to disk with no cache, which makes Skills naturally editable.
- The built-in Agent `default_agent` gets the whole library installed at initialization;
- other Agents install on demand — through the Web UI's Skill library page, or via the SDK;
- installing writes the library `SKILL.md` verbatim (frontmatter included) and copies any `icon.svg` alongside it.
The library ships as the npm package `@prismshadow/penguin-skills`, carrying the raw `skills/` directory in the tarball; at runtime the package's `skills/<name>/SKILL.md` files are likewise the source of truth for library content.
## Built-in library
The built-in Skills, by group (the group manifest is `SKILL_GROUPS` in `packages/skills/src/index.ts`; the library directory is the source of truth as Skills are added):
| Group | Skill | Purpose |
| --- | --- | --- |
| Agent Development | `agent-creation` | Turn a user requirement into a concrete agent: write the target agent's AGENTS.md and install the skills it needs |
| | `benchmark-design` | Design and calibrate a multi-Case capability Benchmark with repeated independent evaluations and a traceable baseline |
| | `agent-evaluation` | Run and score exactly one Benchmark Case run, with CLI execution, Trace provenance checks and private Rubric isolation |
| | `agent-optimization` | Improve an Agent State from direct feedback or versioned multi-Case Benchmark scores and score-linked Traces |
| Data Analysis | `data-analysis` | Complete data-analysis tasks with bounded evidence inspection, explicit answer-changing decisions, native artifact handling and final output verification |
| Penguin Development | `penguin-sdk` | Build AI apps on the SDK (the createSession/run streaming loop) |
| | `penguin-cli` | Manage model API keys, default models and per-agent Vault secrets with the penguin CLI |
| | `agenthub-models` | Call model APIs through `@prismshadow/agenthub`: streaming text, image generation, speech synthesis and embeddings |
| Web Development | `web-design` | Default visual language for generated web pages: minimal black-white-gray |
| Software Engineering | `software-engineering` | Complete software-engineering tasks: investigate and review code, implement fixes, features and refactors with minimal scope, validate changes, and report verified outcomes |
## Writing and optimizing Skills
- Manual install: create a directory under `agent_state/skills/<name>/` and write a `SKILL.md`; the system scans `skills/` when assembling the system prompt and injects the metadata. A directory without a `SKILL.md` does not count as a Skill.
- Uninstalling deletes the whole `skills/<name>/` directory and is idempotent.
- An Agent can rewrite its own SKILL.md as part of a task — combined with Benchmark evaluation and optimization this closes the improvement loop, see [Self-Improvement](/self-improvement).
+76
View File
@@ -0,0 +1,76 @@
---
title: 技能系统
description: Skill 以目录加 SKILL.md 承载可复用指令,元数据先行、正文按需读取,并可由 Agent 自行编辑优化。
---
## Skill 的形态
一个 Skill 就是一个目录:内含一份 `SKILL.md`,可选附带一个 `icon.svg` 自定义图标。目录名即权威的 Skill 名,须匹配 `^[A-Za-z0-9_-]+$`;frontmatter 中的 `name` 以目录名为准。
frontmatter 字段:
| 字段 | 说明 |
| --- | --- |
| `name` | Skill 名,与目录名一致 |
| `description` | 英文单行描述,注入系统 Prompt |
| `short_description` / `short_description_zh` | UI 短标签(卡片等紧凑位置用),不注入 Prompt |
| `version` | 自然数版本号,默认 1 |
| `updated` | 更新日期 |
```md
---
name: my-skill
description: One-line English description injected into the system prompt.
short_description: Short UI label.
short_description_zh: 简短的中文标签。
version: 1
updated: 2026-07-17
---
# My Skill
具体的步骤、边界与验收标准……
```
解析是容错的:只识别首个 `---` 块内的 `key: value` 标量行;`version` 不是自然数时回退为 1,`updated` 缺省为空。
## 渐进式加载
Skill 采用「先索引、后正文」的设计:系统 Prompt 经 `{{SKILL_METADATA}}` 占位符只注入每个已安装 Skill 的元数据(name + description),并指示模型在任务匹配某个 Skill 时,先用 Shell 完整读取对应的 `SKILL.md`,再遵循执行。系统不设专门的 Skill 工具,读取正文就是一次 `exec_command` 调用(见 [工具与审批](/tools))。
对话中也可以显式指定 Skill:此时消息以 `<use_skills>` 块开头,列出要使用的 Skill 名。
若消息只点名 Skill 而没有给出具体任务,模型会先询问需求再开始。
## 安装与存放
已安装的 Skill 位于 Agent State 的 `agent_state/skills/<name>/`。文件即事实源:每次读取直接读文件、不设缓存,因此 Skill 天然可编辑。
- 内置 Agent `default_agent` 在初始化时安装完整 Skill 库;
- 其他 Agent 按需安装:经 Web 界面的 Skill 库页,或经 SDK;
- 安装即把库里的 `SKILL.md` 原样写入(含 frontmatter),目录内的 `icon.svg` 一并拷贝。
Skill 库以 npm 包 `@prismshadow/penguin-skills` 发布,tarball 直接携带原始 `skills/` 目录;运行时库内容的事实源同样是包内的 `skills/<name>/SKILL.md` 文件。
## 内置 Skill 库
内置 Skill 按分组列出如下(分组清单见 `packages/skills/src/index.ts` 的 `SKILL_GROUPS`,新增 Skill 时以库目录为准):
| 分组 | Skill | 说明 |
| --- | --- | --- |
| Agent 开发 | `agent-creation` | 把用户需求变成具体的 Agent:撰写目标 Agent 的 AGENTS.md 并安装所需 Skill |
| | `benchmark-design` | 设计并校准多 Case 的能力评测 Benchmark,含重复独立评测与可追溯基线 |
| | `agent-evaluation` | 隔离执行并评分单个 Benchmark Case:CLI 执行、Trace 溯源检查、Rubric 私有隔离 |
| | `agent-optimization` | 依据直接反馈或带版本的多 Case Benchmark 分数与关联 Trace 改进 Agent State |
| 数据分析 | `data-analysis` | 以有界的证据检查、显式的改答案决策、原生产物处理与最终输出校验完成数据分析任务 |
| Penguin 开发 | `penguin-sdk` | 基于 SDK 构建 AI 应用(createSession/run 流式循环) |
| | `penguin-cli` | 用 penguin CLI 管理模型 API Key、默认模型与各 Agent 的 Vault 密钥 |
| | `agenthub-models` | 经 `@prismshadow/agenthub` 调用模型 API:流式文本、图像生成、语音合成与 Embedding |
| 网页开发 | `web-design` | 生成网页的默认视觉规范:极简黑白灰 |
| 软件工程 | `software-engineering` | 完成软件工程任务:调查与审查代码,以最小改动实现修复、特性与重构,验证改动并报告经过确认的结果 |
## 编写与优化
- 手工安装:在 `agent_state/skills/<name>/` 下建目录并写入 `SKILL.md` 即可,系统组装系统 Prompt 时扫描 `skills/` 注入元数据;没有 `SKILL.md` 的目录不计为 Skill。
- 卸载即删除整个 `skills/<name>/` 目录,操作幂等。
- Agent 可以在任务中直接改写自己的 SKILL.md——配合 Benchmark 评测与优化形成闭环,见 [自我进化](/self-improvement)。
+189
View File
@@ -0,0 +1,189 @@
---
title: Tools & Approval
description: The deliberately minimal built-in toolset, its execution contract with centralized close-out, and per-call approval audited in the Trace.
---
## Design
PenguinHarness ships a deliberately minimal built-in toolset: the shell is the universal interface, and reading, writing and editing files all go through `exec_command` — there are no separate file tools. Fewer tools mean fewer schema tokens and fewer wrong calls.
## Execution contract
Every built-in tool implements the same `BuiltinTool` interface (`packages/core/src/environment/tools/types.ts`):
```ts
interface BuiltinTool {
name: string;
definition: ToolDefinitionConfig;
execute(
args: Record<string, unknown>,
ctx: ToolExecutionContext,
): AsyncGenerator<OmniMessage, ToolResult | void>;
}
interface ToolExecutionContext {
workspaceDir: string;
toolCallId: string;
signal?: AbortSignal;
approve?: ApproveFn; // forwarded to tools that spawn child Sessions (approval inheritance)
}
interface ToolResult {
stopReason?: StopReason; // the tool's self-reported terminal state (lowest priority, see below)
note?: string; // terminal marker appended outside the truncation window (e.g. exit code)
images?: string[]; // data-URL images, appended after the text output
}
```
A tool only yields incremental `partial_tool_call_output` deltas; the Environment handles the close-out centrally:
- streaming framing (start / stop) and `tool_call_id` threading;
- timeout merging, and head-kept truncation once output exceeds `maxOutputLength` (default 16000 characters);
- stop_reason priority: user interrupt > timeout > tool throw > tool self-report;
- never-empty output (`[no output]` is substituted when a tool produced nothing);
- `note` (e.g. the exit code) and images are appended outside the truncation window, so the terminal marker survives even when long output is cut.
Tools and the Environment never throw into the engine: errors collapse into `tool_call_output` messages the model can read and react to. See the [OmniMessage Protocol](/omni-message) for message structure.
## Configuration fields
Each tool is described by one `ToolDefinitionConfig`:
| Field | Meaning |
| --- | --- |
| `name` | Tool name, matching the model's `tool_call.name` |
| `description` | Tool description handed to the model |
| `parameters` | JSON Schema of the arguments |
| `permission` | `"r"` read-only / `"rw"` read-write |
| `forModel` | `"vision"` / `"text-only"`: selected by the Session model's class; omitted = available to all models |
| `timeoutMs` | Per-call timeout (ms), default 120000; `<=0` disables |
| `maxOutputLength` | Output length cap (characters); `<=0` disables |
## Built-in tools
There are 6 built-in tools (assembled via `packages/core/src/environment/tools/registry.ts`):
| Tool | Permission | Timeout (ms) | Purpose |
| --- | --- | --- | --- |
| `exec_command` | rw | 120000 | Run a shell command in the Workspace via `bash -lc`, streaming stdout/stderr |
| `input_command` | rw | 130000 | Drive a running command by `process_id`: write stdin, send Ctrl-C, poll output |
| `run_subagent` | rw | 600000 | Delegate a self-contained subtask to a child Agent in the same Workspace |
| `input_subagent` | rw | 600000 | Poll a background subagent, or send a follow-up prompt once it is idle |
| `read_image` | r | 60000 | Read an image and return it as image content (vision models) |
| `describe_image` | r | 90000 | Have the configured `vision_model` read the image and answer in text (text-only models) |
### Command sessions
`exec_command` waits in the foreground first; if the command outruns `yield_time_ms` it moves to the background and the call returns the output so far plus a `process_id`, driven from then on by `input_command`:
```text
exec_command(cmd)
├─ finishes within the foreground window (yield_time_ms, default 60000)
│ ──► full output + exit code
└─ still running ──► backgrounds, returns output so far + process_id
│
input_command(process_id[, chars]) ──► write stdin / send Ctrl-C / poll
└─ loop until the command exits
```
Both tools' arguments (explicit keys):
```ts
// exec_command
{
cmd: string; // required: the shell command to run
workdir?: string; // working directory; defaults to the Workspace root, relative paths resolve against it
yield_time_ms?: number; // foreground wait; default 60000, minimum 250, capped below the tool timeout
}
// input_command
{
process_id: string; // required: the command-session id returned by exec_command
chars?: string; // characters for stdin; send "\u0003" alone to deliver Ctrl-C; empty = poll only
yield_time_ms?: number; // wait; defaults 250 for writes, 5000 for empty polls
}
```
### Subagents
`run_subagent` hands a subtask you can fully specify in one prompt to a child Agent, with the same two-phase shape: after the foreground window (default 300000ms) it moves to the background with a `subagent_id`, driven by `input_subagent` for polling or follow-up prompts; the child's pending approvals surface while the poll waits.
```ts
// run_subagent
{
prompt: string; // required: the complete subtask (all context + the exact final output expected)
agent_id?: string; // the child Agent; defaults to the current Agent
model_id?: string; // the child Session's model; defaults to the Project default model
yield_time_ms?: number; // foreground wait; default 300000
}
// input_subagent
{
subagent_id: string; // required: the background Subagent id returned by run_subagent
prompt?: string; // follow-up task, accepted only while the child Session is idle; empty = poll only
yield_time_ms?: number; // wait; defaults 300000 with a prompt, 10000 for empty polls
}
```
- Depth is capped at 1: a subagent cannot spawn another subagent.
- The child Session inherits the parent Agent's approval callback, so the approval mode follows the parent.
- The child Session gets its own Trace, linked from the parent by a `subagent` pointer event; child messages stream back into the parent flow tagged with `origin`. See [Sessions & Traces](/sessions-and-traces).
### Image tools
`read_image` and `describe_image` are mutually exclusive, selected by the Session model's vision flag. Both accept an http(s) URL or a Workspace path and support png/jpeg/gif/webp up to 5MB. Text-only models get `describe_image`: the image plus a prompt are forwarded to the Project's configured `vision_model`, whose text answer becomes the tool output. See [Models & Providers](/models).
```ts
// read_image (vision models)
{
source: string; // required: an http(s) URL, or a file path inside the Workspace
}
// describe_image (text-only models)
{
source: string; // required: as above
prompt?: string; // what to ask about the image; defaults to a detailed description
}
```
### Background session caps
| Session type | Cap | Eviction |
| --- | --- | --- |
| Command sessions | 64 | When full, exited sessions are evicted first, then idle ones by LRU |
| Subagent sessions | 8 | Only completed ones are evicted; running subagents never — with no room, spawning is rejected |
## Approval
Every complete `tool_call` triggers exactly one approval decision:
```ts
type ApproveFn = (toolCall: OmniMessage<ToolCallPayload>) => Promise<"allow" | "deny">;
```
| Surface | Behavior |
| --- | --- |
| SDK | Pass `approve` per `session.run`; with none injected the engine denies by default (conservative — nothing gets approved unattended) |
| CLI | `--approve` takes four modes: allow-all (default) / deny-all / read-only / always-ask; read-only auto-approves `permission: "r"` tools and defers the rest to a human |
| Web / Server | The same four modes, set per Session; the mode is re-read from the DB on every decision, so changes take effect immediately; manual decisions arrive via the API |
A deny produces a synthetic aborted `tool_call_output` (`Tool call denied by user.`) for the model to react to. Every decision is written to the Trace as an `approval_decision` event, forming a complete audit record. Approval happens in the tool-execution phase of the [Agent Loop](/agent-loop).
## Custom tools & MCP
The `tools.builtin` array in `system_config.yaml` declares the toolset with entries of the same `ToolDefinitionConfig` shape. The semantics are **wholesale replacement, not merging**: omit the section entirely to keep the full default toolset; once written, the default list is replaced and every tool you keep must carry its complete definition (including the `parameters` JSON Schema — a tool's schema comes entirely from config). `tools.mcpServers` carries MCP server configs (name + config) — enumerating concrete MCP tools is reserved for a later adapter layer and not yet wired. See [Configuration](/configuration).
```yaml
tools:
# Writing builtin replaces the default toolset wholesale (this example deliberately
# keeps a minimal single-tool set).
builtin:
- name: exec_command
description: Run a shell command in the workspace.
permission: rw
timeoutMs: 120000
maxOutputLength: 16000
# parameters: the complete JSON Schema is required (see the default definition
# in packages/core/src/state/default-config.ts); elided here.
mcpServers: []
```
+187
View File
@@ -0,0 +1,187 @@
---
title: 工具与审批
description: 极简内置工具集的设计与执行契约、Environment 统一收尾规则,以及逐调用审批与 Trace 审计。
---
## 设计取向
PenguinHarness 刻意维持一个极小的内置工具集:Shell 是通用接口,文件的读取、写入、编辑全部经由 `exec_command` 完成,不设专门的文件工具。工具越少,注入的 schema 越少,Token 开销越小,模型误调用的概率也越低。
## 执行契约
所有内置工具实现同一个 `BuiltinTool` 接口(`packages/core/src/environment/tools/types.ts`):
```ts
interface BuiltinTool {
name: string;
definition: ToolDefinitionConfig;
execute(
args: Record<string, unknown>,
ctx: ToolExecutionContext,
): AsyncGenerator<OmniMessage, ToolResult | void>;
}
interface ToolExecutionContext {
workspaceDir: string;
toolCallId: string;
signal?: AbortSignal;
approve?: ApproveFn; // 供需要派生子 Session 的工具转发(审批继承)
}
interface ToolResult {
stopReason?: StopReason; // 工具自报终态(优先级最低,见下)
note?: string; // 追加在截断范围之外的终止标记(如退出码)
images?: string[]; // data URL 图像,附加在文本输出之后
}
```
工具本身只需 yield 增量的 `partial_tool_call_output`,收尾由 Environment 集中处理:
- 流式分帧(start / stop)与 `tool_call_id` 贯穿;
- 超时归并;输出超过 `maxOutputLength`(默认 16000 字符)时截断,保留开头;
- stop_reason 按优先级归并:用户中断 > 超时 > 工具抛错 > 工具自报;
- 输出永不为空:没有任何输出时补 `[no output]`;
- `note`(如退出码)与图像附加在截断范围之外,长输出被截断时终止标记不会丢失。
工具与 Environment 从不向引擎抛异常:错误一律折叠为 `tool_call_output` 消息,交给模型阅读并调整下一步。消息结构见 [OmniMessage 协议](/omni-message)。
## 配置字段
每个工具由一条 `ToolDefinitionConfig` 描述:
| 字段 | 说明 |
| --- | --- |
| `name` | 工具名,对应模型产出的 `tool_call.name` |
| `description` | 提供给模型的工具说明 |
| `parameters` | 参数 JSON Schema |
| `permission` | `"r"` 只读 / `"rw"` 读写 |
| `forModel` | `"vision"` / `"text-only"`:按 Session 模型类别装配;缺省对所有模型可用 |
| `timeoutMs` | 单次调用超时(ms),默认 120000;`<=0` 关闭 |
| `maxOutputLength` | 输出长度上限(字符);`<=0` 关闭 |
## 内置工具
共 6 个内置工具(装配入口 `packages/core/src/environment/tools/registry.ts`):
| 工具 | 权限 | 超时(ms) | 用途 |
| --- | --- | --- | --- |
| `exec_command` | rw | 120000 | 在 Workspace 内以 `bash -lc` 运行命令,流式返回 stdout/stderr |
| `input_command` | rw | 130000 | 按 `process_id` 驱动运行中的命令:写 stdin、发 Ctrl-C、轮询输出 |
| `run_subagent` | rw | 600000 | 把自包含子任务委派给同 Workspace 的子 Agent |
| `input_subagent` | rw | 600000 | 轮询后台 Subagent,或在其空闲时追加后续 Prompt |
| `read_image` | r | 60000 | 读取图片并作为图像内容返回(vision 模型) |
| `describe_image` | r | 90000 | 由 `vision_model` 代读图片并返回文字回答(text-only 模型) |
### 命令会话
`exec_command` 先在前台等待;命令超过 `yield_time_ms` 仍未结束时转入后台,返回已有输出和一个 `process_id`,之后用 `input_command` 驱动:
```text
exec_command(cmd)
├─ 前台窗口(yield_time_ms,默认 60000)内结束 ──► 完整输出 + 退出码
└─ 未结束 ──► 转入后台,返回已有输出 + process_id
│
input_command(process_id[, chars]) ──► 写 stdin / 发 Ctrl-C / 轮询
└─ 循环驱动,直至命令退出
```
两个工具的参数(明确键名):
```ts
// exec_command
{
cmd: string; // 必填:要执行的 shell 命令
workdir?: string; // 工作目录;缺省为 Workspace 根,相对路径按其解析
yield_time_ms?: number; // 前台等待时长;默认 60000,最小 250,上限受工具超时约束
}
// input_command
{
process_id: string; // 必填:exec_command 返回的命令会话 id
chars?: string; // 写入 stdin 的字符;单独发送 "\u0003" 传递 Ctrl-C;缺省仅轮询
yield_time_ms?: number; // 等待时长;有写入默认 250,空轮询默认 5000
}
```
### Subagent
`run_subagent` 把一段能一次说清的子任务交给子 Agent 执行,同样是两段式:前台窗口(默认 300000ms)过后转入后台并返回 `subagent_id`,由 `input_subagent` 轮询或追加 Prompt;子 Agent 的待审批项会在轮询等待期间浮出。
```ts
// run_subagent
{
prompt: string; // 必填:完整的子任务(含全部上下文与期望的最终产出)
agent_id?: string; // 子 Agent;缺省复用当前 Agent
model_id?: string; // 子 Session 模型;缺省用 Project 默认模型
yield_time_ms?: number; // 前台等待时长;默认 300000
}
// input_subagent
{
subagent_id: string; // 必填:run_subagent 返回的后台 Subagent id
prompt?: string; // 追加任务,仅在子 Session 空闲时接受;缺省仅轮询
yield_time_ms?: number; // 等待时长;有追加默认 300000,空轮询默认 10000
}
```
- 深度上限为 1:Subagent 不能再派生 Subagent。
- 子 Session 继承父 Agent 的审批回调,审批模式随父生效。
- 子 Session 拥有独立 Trace,父 Trace 以 `subagent` 指针事件链接;子消息带 `origin` 标记回流到父级消息流。见 [Session 与 Trace](/sessions-and-traces)。
### 图像工具
`read_image` 与 `describe_image` 互斥,按 Session 模型的 vision 标记二选一装配。两者都接受 http(s) URL 或 Workspace 路径,支持 png/jpeg/gif/webp,不超过 5MB。text-only 模型走 `describe_image`:图片连同提问转交 Project 配置的 `vision_model`,其文字回答即工具输出。见 [模型与 Provider](/models)。
```ts
// read_image(vision 模型)
{
source: string; // 必填:http(s) URL,或 Workspace 内的文件路径
}
// describe_image(text-only 模型)
{
source: string; // 必填:同上
prompt?: string; // 要对图片提出的问题;缺省为详细描述
}
```
### 后台会话上限
| 会话类型 | 上限 | 淘汰策略 |
| --- | --- | --- |
| 命令会话 | 64 | 满时优先淘汰已退出者,否则对空闲会话按 LRU 淘汰 |
| Subagent 会话 | 8 | 只淘汰已完成者;运行中的从不淘汰,无空位则拒绝派生 |
## 审批
每个完整的 `tool_call` 触发且只触发一次审批决策:
```ts
type ApproveFn = (toolCall: OmniMessage<ToolCallPayload>) => Promise<"allow" | "deny">;
```
| 使用面 | 行为 |
| --- | --- |
| SDK | 每次 `session.run` 传入 `approve` 回调;未注入时引擎默认全部拒绝(保守策略,避免无人值守下误放行) |
| CLI | `--approve` 四种模式:allow-all(默认)/ deny-all / read-only / always-ask;read-only 自动放行 `permission: "r"` 的工具,其余转人工 |
| Web / Server | 同样四种模式,按 Session 设置;每次决策前从数据库重读,改模式立即生效;人工决策经 API 送达 |
deny 会合成一条 aborted 的 `tool_call_output`(内容为 `Tool call denied by user.`),模型据此调整策略。每次决策都以 `approval_decision` 事件写入 Trace,构成完整的审计记录。审批发生在 [Agent 运行循环](/agent-loop) 的工具执行阶段。
## 自定义与 MCP
`system_config.yaml` 的 `tools.builtin` 数组以 `ToolDefinitionConfig` 同构条目声明工具集。注意语义是**整体替换而非合并**:整段省略时使用完整默认工具集;一旦写出,默认列表即被替换,要保留的每个工具都必须携带完整定义(含 `parameters` JSON Schema——工具的参数 schema 完全来自配置)。`tools.mcpServers` 承载 MCP Server 配置(name + config)——具体 MCP 工具的枚举由后续适配层接管,当前仅保留配置位。见 [配置参考](/configuration)。
```yaml
tools:
# 写出 builtin 即整体替换默认工具集(此例刻意只保留一个最小工具集)。
builtin:
- name: exec_command
description: Run a shell command in the workspace.
permission: rw
timeoutMs: 120000
maxOutputLength: 16000
# parameters: 必须携带完整 JSON Schema(默认定义见
# packages/core/src/state/default-config.ts),此处从略。
mcpServers: []
```
+106
View File
@@ -0,0 +1,106 @@
---
title: Web App Guide
description: A page-by-page guide to the Web App — login, chat, Agent management, models, usage, and traces.
---
PenguinHarness ships with a ready-to-use Web App: multi-user login, streaming chat, Agent configuration, model and usage management all happen in the browser. This guide walks through the app page by page. For installation and first launch, see the [Quickstart](/quickstart).
## Source layout
```text
packages/web/src
├── api/ # fetch wrapper · one function per API (DTOs type-only from @prismshadow/penguin-server/api) · SSE wrapper
├── state/ # auth / project / sessions / theme / locale contexts
├── lib/omni/ # OmniMessage stream → view-model reducer; connect-first + dedup stream controller
├── components/ # ui primitives (modal / drawer / select …) and the app layout
└── features/ # chat / agents / skills / models / usage / traces / benchmark / admin pages
```
## Startup and Login
```bash
penguin web
# open http://127.0.0.1:7364
```
The initial account is `admin` / `admin123`. There is no self-registration: accounts are created by an admin on the user-management page, and every new user automatically gets an independent initial Project named `<userId>-default_project`. While the initial password is still in use, a banner prompts the user to change it.
Logins persist for 7 days with sliding renewal; an admin password reset invalidates all of that user's login sessions.
The interface language (中文 / English / system) and theme (light / dark / system) can be switched at any time.
## Chat (/chat)
### Creating a Conversation
A new conversation starts as a draft: pick the Agent, the Workspace (via a server-side directory browser), the approval mode, and the model before sending the first message. The Session is created on first send, and from then on its model and Workspace are locked.
There are four approval modes: `allow-all`, `deny-all`, `read-only` (only read-only tools pass), and `always-ask`. See [Tools and Approvals](/tools).
### Streaming Rendering
- Model text renders token by token; thinking blocks are collapsible;
- Tool cards expand to show arguments and output, with a live timer while running;
- Subagents appear as nested cards; context compaction shows a banner;
- After each Task, a stats line shows tokens, TPS, elapsed time, and cost.
### Input and Shortcuts
- Enter sends, Shift+Enter inserts a newline, and images can be pasted;
- Typing `/` opens the slash menu: trigger context compaction (`/compact`) or toggle installed Skills — chosen Skills are sent along with the message in a `<use_skills>` block;
- Typing `@` mentions another Agent to hand the conversation over to it;
- When human approval is required, tool calls show inline allow/deny buttons in the message stream; the approval mode can be changed mid-Session.
### Files Panel
The files panel browses the Workspace tree, previews files (Markdown / HTML rendered), uploads files (≤ 14MB each), and downloads them.
## Agent Management (/agents)
The list page creates and deletes Agents; clicking through opens the `/agents/:agentId` settings page, organized into tabs:
| Tab | Contents |
| --- | --- |
| Overview | Basic info, plus export / import of Agent State snapshots |
| Prompt | AGENTS.md and system_prompt |
| Runtime | Runtime parameters such as max_turns, model.*, compaction.* |
| Tools | Built-in tool table and MCP server JSON configuration |
| Vault | Environment-variable entries with masked values |
| Schedule | Scheduled tasks (TOML-defined): create, edit, toggle, delete |
Scheduled tasks fire on a fixed period (minimum 5 minutes) and run only while the service is running.
## Skill Library (/skills)
Browse the Skill library by group, install Skills onto an Agent, or quick-invoke one into a chat draft.
## Model Configuration (/models)
A per-Project model table grouped by provider. Models can be added and edited: identity is the `(provider, model_id)` pair, credentials are masked, and context window, pricing, and the vision flag are configurable. You can set the default model and the vision model (which reads images on behalf of session models without image input), and run a connectivity test on any entry. Only Project owners can edit. For concepts, see [Models and Providers](/models).
## Usage (/usage)
- Filters: Agent, model, date range;
- Summary cards: today / last 7 days / cumulative;
- Charts: per-Agent share, per-model success rates, daily Token and cost trends;
- A server error panel summarizing recent server-side error records.
## Trace Browser (/traces)
Drill down Agent → date → Session → Trace file. Per-turn cards show a context-occupancy donut and a cache breakdown, alongside a lane-based execution timeline and the full event list. For the storage model, see [Sessions and Traces](/sessions-and-traces).
## Benchmark (/benchmark)
Read-only scoreboards per Benchmark: switch the metric (score / cost / duration), drill into each Case's runs, and jump to the linked Session and Trace. Works together with the [Self-Improvement](/self-improvement) workflow.
## User Administration (/admin/users)
Admin only: list and create users, reset passwords, and delete users (the built-in admin cannot be deleted).
## Projects and Members
The sidebar provides a Project switcher and supports creating new Projects. Members have two roles, owner and member: owners manage membership and exclusively edit models, Vault, and Schedules, as well as perform deletions.
## Production Deployment
The server hosts the built SPA itself (same origin, SPA fallback), so a single `penguin web` or `penguin server` process is all production needs. The npm package bundles the frontend build; to serve a custom static directory, override it with `PENGUIN_WEB_DIST` — see the [Configuration Reference](/configuration).
+106
View File
@@ -0,0 +1,106 @@
---
title: Web App 指南
description: 按页面组织的 Web App 使用指南:登录、Chat、Agent 管理、模型、用量与 Trace。
---
PenguinHarness 自带一个开箱即用的 Web App:多用户登录、流式对话、Agent 配置、模型与用量管理都在浏览器中完成。本文按页面组织,逐一介绍各页面的功能与操作。安装与首次启动见[快速开始](/quickstart)。
## 目录结构
```text
packages/web/src
├── api/ # fetch 封装 · 每个 API 一个函数(DTO type-only 来自 @prismshadow/penguin-server/api)· SSE 封装
├── state/ # auth / project / sessions / theme / locale 五个 context
├── lib/omni/ # OmniMessage 流 → 渲染视图模型 reducer;连接先行 + 去重的流控制器
├── components/ # ui 原语(modal / drawer / select …)与应用布局
└── features/ # chat / agents / skills / models / usage / traces / benchmark / admin 各页面
```
## 启动与登录
```bash
penguin web
# 打开 http://127.0.0.1:7364
```
初始账号为 `admin` / `admin123`。系统不开放自助注册:账号由管理员在用户管理页创建;每个新用户会自动获得一个独立的初始 Project,命名为 `<userId>-default_project`。仍在使用初始密码时,页面会以横幅提示尽快修改。
登录状态保持 7 天(滑动续期);管理员重置密码会使该用户的全部登录会话失效。
界面语言(中文 / English / 跟随系统)与主题(浅色 / 深色 / 跟随系统)可随时切换。
## Chat 页面(/chat)
### 新建会话
新会话从草稿开始:先选择 Agent、Workspace(服务器端目录浏览器选取)、审批模式与模型,再发送第一条消息。Session 在首次发送时才真正创建,此后该会话的模型与 Workspace 即被锁定。
审批模式共四种:`allow-all`(全部放行)、`deny-all`(全部拒绝)、`read-only`(仅放行只读工具)、`always-ask`(每次询问),详见[工具与审批](/tools)。
### 流式渲染
- 模型文本逐 Token 渲染,思考块可折叠;
- 工具卡片可展开查看参数与输出,执行中显示实时计时;
- 子 Agent 以嵌套卡片呈现;上下文压缩以横幅提示;
- 每个 Task 结束后显示统计行:Token 用量、TPS、耗时与费用。
### 输入与快捷操作
- Enter 发送,Shift+Enter 换行,支持粘贴图片;
- 输入 `/` 打开快捷菜单:触发上下文压缩(`/compact`),或勾选已安装的 Skill——所选 Skill 会以 `<use_skills>` 块随消息发送;
- 输入 `@` 提及其他 Agent,将会话交接给它;
- 需要人工审批时,工具调用在消息流中内联显示“允许 / 拒绝”按钮;审批模式在会话中途可随时调整。
### 文件面板
文件面板可浏览 Workspace 目录树、预览文件(Markdown / HTML 渲染显示)、上传文件(单个 ≤ 14MB)与下载文件。
## Agent 管理(/agents)
列表页支持创建与删除 Agent;点击进入 `/agents/:agentId` 设置页,按标签页组织:
| 标签页 | 内容 |
| --- | --- |
| Overview | 基本信息,以及 Agent State 快照的导出 / 导入 |
| Prompt | AGENTS.md 与 system_prompt |
| Runtime | max_turns、model.*、compaction.* 等运行参数 |
| Tools | 内置工具表格与 MCP Server 的 JSON 配置 |
| Vault | 环境变量条目,值以掩码显示 |
| Schedule | 定时任务(TOML 定义):创建、编辑、启停、删除 |
定时任务按固定周期触发(最短 5 分钟),且仅在服务运行期间执行。
## Skill 库(/skills)
按分组浏览 Skill 库,可将 Skill 安装到指定 Agent,或一键带入 Chat 草稿快速调用。
## 模型配置(/models)
按 Provider 分组展示当前 Project 的模型表格。支持添加与编辑模型:以 `(provider, model_id)` 为唯一标识,凭据以掩码显示,可配置上下文窗口、定价与视觉(vision)标记;可设置默认模型与视觉模型(在会话模型不支持图片输入时代为读图),并对任一模型做连通性测试。仅 Project Owner 可编辑,概念说明见[模型与 Provider](/models)。
## 用量统计(/usage)
- 筛选条件:Agent、模型、日期范围;
- 概览卡片:今日 / 近 7 天 / 累计用量;
- 图表:各 Agent 占比、各模型成功率、每日 Token 与费用趋势;
- 服务端错误面板:汇总最近的服务端错误记录。
## Trace 浏览(/traces)
按 Agent → 日期 → Session → Trace 文件逐级下钻。每回合卡片展示上下文占用环形图与缓存构成,并提供泳道式执行时间线与完整事件列表。Trace 的存储模型见 [Session 与 Trace](/sessions-and-traces)。
## Benchmark(/benchmark)
只读展示各 Benchmark 的评分板,可切换指标(得分 / 费用 / 耗时),下钻查看每个 Case 的多次运行结果,并跳转到关联的 Session 与 Trace。配合[自我进化](/self-improvement)工作流使用。
## 用户管理(/admin/users)
仅管理员可见:列出与创建用户、重置密码、删除用户(内置 admin 不可删除)。
## Project 与成员
侧边栏提供 Project 切换器,并支持创建新 Project。成员分为 Owner 与 Member 两种角色:Owner 负责成员管理,并独占模型、Vault、Schedule 的编辑以及各类删除操作。
## 生产部署
服务端自身托管构建好的 SPA(同源、SPA fallback),生产环境只需运行 `penguin web` 或 `penguin server` 一个进程。npm 安装包已内置前端产物;如需自定义静态目录,可用 `PENGUIN_WEB_DIST` 覆盖,见[配置参考](/configuration)。