diff --git a/CHANGELOG.md b/CHANGELOG.md new file mode 100644 index 0000000..83dc026 --- /dev/null +++ b/CHANGELOG.md @@ -0,0 +1,6 @@ +# Changelog + +One brief line per release. Per-release detail lives in [`changelog//`](changelog/). + +- **0.0.2** — unreleased. ([details](changelog/0.0.2/README.md)) +- **0.0.1** — 2026-07-19. First tagged release; changelog history starts after this tag. diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md new file mode 100644 index 0000000..27d800f --- /dev/null +++ b/CONTRIBUTING.md @@ -0,0 +1,100 @@ +# Contributing to PenguinHarness + +Thanks for helping build PenguinHarness! This guide covers the workspace setup, daily +commands, quality gates, and the repo's working rules. + +## Prerequisites + +- Node >= 24 +- pnpm 10 (`corepack enable` or `npm install -g pnpm`) + +## Setup and daily commands + +```bash +pnpm install +pnpm build # build first: core's exports point at dist/ + +pnpm dev # backend + web app together (prefixed logs, deps built once) +pnpm dev:server # backend at 127.0.0.1:7364 +pnpm dev:web # web app (Vite) at 127.0.0.1:7365, /api proxied +pnpm dev:docs # docs site (Vite) at 127.0.0.1:7367 +pnpm dev:landing # landing page (Vite) at 127.0.0.1:7366 + +BASE_PATH=/ pnpm build:site # assemble landing + docs exactly like the Pages deploy +``` + +Every dev command runs `scripts/dev-prebuild.mjs` first, which (behind a lock that +serializes concurrent invocations) **keeps `pnpm install` current automatically** — a +fresh clone or a pulled lockfile change installs before starting, and an up-to-date tree +pays nothing (the lockfile hash is stamped) — then prebuilds the workspace deps (skills, +core) with back-to-back builds deduped: starting `dev:server` and `dev:web` at the same +time (or just `pnpm dev`) installs and builds exactly once. `dev:docs` / `dev:landing` +run the install check only (`--install-only`). + +Copy `.env.example` to `.env` for model credentials in development. + +## Repo layout + +A pnpm monorepo (TypeScript, Node >= 24). One install ships four layers that share a +single data directory (`~/.penguin/data`) and a single message protocol (OmniMessage): + +| Package | Name | Role | +| ------------------------------------ | ----------------------------- | ------------------------------------------------------------------------------------------------------- | +| [`packages/core`](packages/core) | `@prismshadow/penguin-core` | SDK & engine: ReAct loop, OmniMessage protocol, LLM/Environment interface contracts, Agent State, Trace | +| [`packages/cli`](packages/cli) | `@prismshadow/penguin-cli` | The `penguin` command: REPL, one-shot runs, model & vault config, service launcher | +| [`packages/server`](packages/server) | `@prismshadow/penguin-server` | Web backend: HTTP API + SSE streaming, multi-user auth, Project authorization, usage stats | +| [`packages/web`](packages/web) | `@prismshadow/penguin-web` | Web App: multi-session chat, Agent/skill/model management, Trace observability, evaluation center | +| [`packages/skills`](packages/skills) | `@prismshadow/penguin-skills` | Built-in skill library (agent creation, benchmarking, evaluation, optimization, …) | +| [`packages/landing`](packages/landing) | — | Product landing page (this repo's website) | +| [`packages/docs`](packages/docs) | — | Documentation site (bilingual, deployed under `/docs/`) | + +Responsibilities split by source of truth: the **SDK** owns protocol and execution +(message parsing, the agent loop, tools), the **Server** owns the multi-user runtime +(auth, SSE streaming, scheduled tasks), and the **file layer** under `~/.penguin/data` +owns everything editable and recorded (prompts, Skills, secrets, Traces). The full map +is in [Architecture → Division of responsibilities](https://penguin.ooo/docs/architecture). + +## Quality gates + +CI runs all of these on every PR — run them locally before pushing: + +```bash +pnpm format:check # prettier +pnpm typecheck +pnpm test # unit suites for every package +``` + +End-to-end suites (optional locally, slower): + +```bash +npx playwright install chromium # once +pnpm --filter @prismshadow/penguin-web test:e2e # browser e2e against a mock LLM +pnpm test:e2e # core live-model e2e, needs DEEPSEEK_API_KEY +``` + +## Working rules + +- **English is the repository's working language** — code, comments, error/log messages, + test names and fixtures, package metadata, and developer docs. Chinese appears only + where it is the content itself: zh i18n catalogs and fields (`strings.ts` dictionaries, + CLI `i18n.ts`, `titleZh`, `short_description_zh`), `*.zh.md` documents, and test + literals that assert zh i18n output or exercise CJK-specific behavior. +- **Every change ships with a changelog entry**: add + `changelog//YYYY-MM-DD-.md` under the next unreleased version + (released versions' folders are frozen) — an H1 title, a one-sentence summary + paragraph, then details — and add a one-line link for it to that version's index, + `changelog//README.md`. The layout is documented in + [`changelog/README.md`](changelog/README.md). Related changes may share one entry + file (extending its details) instead of opening a new file per small change. +- README assets under `assets/readme/` are generated — the benchmark charts from the + landing benchmark data, and the demo screenshots via + `node packages/landing/scripts/capture-readme-demo.mjs` (build first; needs Playwright + chromium). Regenerate rather than hand-editing. + +## Pull requests + +- Branch from `main`; keep PRs focused on one topic. +- Make sure CI is green (build, format, typecheck, tests) and describe user-visible + changes in the PR body. +- New user-facing behavior should come with tests, and with docs updates when it changes + documented behavior (README, docs site). diff --git a/README.md b/README.md index 29ff03f..2c2de29 100644 --- a/README.md +++ b/README.md @@ -4,12 +4,9 @@

PenguinHarness

-

Efficient Self-Improving Harness for Everyone

+

With LangChain, you build agents by hand — at 1× speed.
With PenguinHarness, agents build agents — at 100×.

-

- Open-source, local-first infrastructure that builds AI agents for you — - from automatic agent construction to recursive self-improvement. -

+

A zero-code Harness CLI and Web UI, connected to 1000+ models.

CI @@ -19,58 +16,105 @@

- English | 简体中文 · - Website · - Docs · - Blog + Website + Docs + Blog

- - - PenguinHarness Web App — multi-session chat with live streaming tool calls - + Discord + X (Twitter) + WeChat

---- +

English | 简体中文

## Why PenguinHarness -- **Simplest Is the Best** — a deliberately minimal toolset over clean low-level interfaces: fewer tool calls, fewer tokens, complex tasks done efficiently. -- **Harness for Building Agents** — with the PenguinHarness SDK, an Agent builds complete Agent applications for you, autonomously, from scratch. -- **Harness for Recursive Self-Improvement** — with PenguinHarness Skills, an Agent evaluates and optimizes itself: benchmark, find the lost points, ship version N+1, snapshot before every round. -- **Local-first and lightweight** — 100% open source, runs on a single CPU, your data never leaves the machine. 1000+ online and local models reachable through one gateway. -- **Everything observable** — every request, tool call and approval decision lands in an append-only Trace; any Session can be resumed from it. +Three reasons, in deliberate order — from task quality, to how agents get built, to how they keep improving. -## Quickstart +### 1. 🏆 Comparable quality, one to two orders of magnitude cheaper -Install with one command (Linux / macOS, x64 / arm64, bundled Node runtime): +A deliberately minimal toolset over clean low-level interfaces: fewer tool calls, fewer tokens — deeply tuned for open models like DeepSeek. Each harness on the model it is normally paired with, same tasks, head-to-head: -```bash -curl -fsSL https://github.com/Prism-Shadow/penguin-harness/releases/latest/download/install.sh | sh +

+ + + Benchmark: PenguinHarness leads the data-analysis suite and ties OpenAI Codex on coding, at a small fraction of both rivals' cost + +

+ +**Best accuracy on data analysis — at 1/70 of Claude Code's cost.** + +### 2. ⚡ One sentence, and an Agent builds your Agent app + +Type one sentence, and an Agent builds the complete Agent application for you — scaffold, code, and run instructions, end to end: + +```text +Collect the docs from https://github.com/ericbuess/claude-code-docs and build a RAG app that answers Claude Code questions as a configuration expert, citing its sources. ``` -Or via npm (requires Node >= 24; the command it installs is `penguin`): +And this is the finished product — a docs expert with retrieval, cited sources that link to the original files, and example questions built in: + +https://github.com/user-attachments/assets/9b7033e8-f08a-4c3f-bd33-547896664e6e + +**And generating this entire RAG app burned just $0.02 (¥0.2) of tokens — on DeepSeek V4 Pro.** + +### 3. 🧬 Self-evolution: it gets stronger with use + +With PenguinHarness Skills, an Agent evaluates and optimizes itself: run the benchmark, find the lost points, ship version N+1 — with a snapshot before every round, and every request observable in the Trace view. + +https://github.com/user-attachments/assets/922d13a6-5ffc-4685-9a39-352f02f9afc0 + +## Supported Models + +| Model | Providers | +| ---------------- | -------------------------------------------------------------------------------- | +| DeepSeek V4 | DeepSeek, OpenRouter, Fireworks AI, SiliconFlow, Qwen Token Plan | +| Kimi K3 | OpenRouter, Qwen Pay-As-You-Go | +| Kimi K2.6 | Moonshot AI | +| GLM 5.2 | Z.AI, OpenRouter, Fireworks AI, SiliconFlow, Qwen Token Plan, Qwen Pay-As-You-Go | +| Hunyuan 3 | OpenRouter | +| Qwen 3.8 Max | Qwen Token Plan (preview) | +| GPT 5.5 | OpenAI, OpenRouter | +| Gemini 3.5 Flash | Google Gemini, OpenRouter | +| Claude Opus 4.8 | Anthropic, OpenRouter | + +Any OpenAI-protocol endpoint is supported: pick a preset above, or point a custom endpoint at any of the 1000+ online and local models. + +## Requirements + +| Requirement | Supported | +| ------------ | -------------------------------------------------------------------------- | +| OS | Linux, macOS | +| Architecture | x64, arm64 | +| Runtime | bundled by the one-line installer (npm installs need Node >= 24) | +| Model | an API key for at least one model | + +## Installation + +### 🌐 Web App — for humans + +🚀 Install and launch the full experience (multi-session chat, Agent/skill/model management, usage stats, Trace observability, evaluation center): ```bash -npm install -g @prismshadow/penguin-cli +curl -fsSL https://penguin.ooo/install.sh | sh +penguin web # start the service and open http://127.0.0.1:7364 (first login: admin / penguin-2026) ``` -Then launch the Web App — or stay in the terminal: +📦 Or via npm: `npm install -g @prismshadow/penguin-cli`. Configure models on the in-app Models page, then chat. + +### 🤖 CLI & SDK — for agents + +The same engine, scriptable — made to be driven by agents (and agents building agents): ```bash -penguin web # start the service and open http://127.0.0.1:7364 (first login: admin / admin123) -penguin server # same service, headless - -# configure a model once (or use the in-app Models page) -penguin config model add --model-id deepseek-v4-pro --api-key sk-... --set-default - +penguin config model add --provider deepseek --model-id deepseek-v4-pro --api-key sk-... --set-default penguin run -m "Create hello.txt containing Hello, Penguin" # one-shot task penguin chat # interactive REPL (/compact, /exit, Ctrl-C to interrupt) +penguin server # headless service (same API the Web App uses) ``` -Using the SDK directly: - ```ts import { createAgent, isCompleteModelMessage, userText } from "@prismshadow/penguin-core"; @@ -86,45 +130,38 @@ for await (const output of session.run([userText("Create hello.txt containing hi } ``` -## What's inside +## Roadmap -A pnpm monorepo (TypeScript, Node >= 24). One install ships four layers that share a single data directory (`~/.penguin/data`) and a single message protocol (OmniMessage): - -| Package | Name | Role | -| --- | --- | --- | -| [`packages/core`](packages/core) | `@prismshadow/penguin-core` | SDK & engine: ReAct loop, OmniMessage protocol, LLM/Environment interface contracts, Agent State, Trace | -| [`packages/cli`](packages/cli) | `@prismshadow/penguin-cli` | The `penguin` command: REPL, one-shot runs, model & vault config, service launcher | -| [`packages/server`](packages/server) | `@prismshadow/penguin-server` | Web backend: HTTP API + SSE streaming, multi-user auth, Project authorization, usage stats | -| [`packages/web`](packages/web) | `@prismshadow/penguin-web` | Web App: multi-session chat, Agent/skill/model management, Trace observability, evaluation center | -| [`packages/skills`](packages/skills) | `@prismshadow/penguin-skills` | Built-in skill library (agent creation, benchmarking, evaluation, optimization, …) | -| [`packages/landing`](packages/landing) | — | Product landing page (this repo's website) | -| [`packages/docs`](packages/docs) | — | Documentation site (bilingual, deployed under `/docs/`) | - -Responsibilities split by source of truth: the **SDK** owns protocol and execution (message parsing, the agent loop, tools), the **Server** owns the multi-user runtime (auth, SSE streaming, scheduled tasks), and the **file layer** under `~/.penguin/data` owns everything editable and recorded (prompts, Skills, secrets, Traces). The full design-by-design map is in [Architecture → Division of responsibilities](https://penguin.ooo/docs/architecture). - -## Documentation - -The docs site covers both usage and design: [Introduction](https://penguin.ooo/docs/) · [Quickstart](https://penguin.ooo/docs/quickstart) · [Architecture](https://penguin.ooo/docs/architecture) · [The OmniMessage Protocol](https://penguin.ooo/docs/omni-message) · [Core Interfaces](https://penguin.ooo/docs/interfaces) · [The Agent Loop](https://penguin.ooo/docs/agent-loop) · [CLI Reference](https://penguin.ooo/docs/cli) · [Server API](https://penguin.ooo/docs/server-api) · [Configuration](https://penguin.ooo/docs/configuration) - -Every doc page has a "Copy Markdown" button, so you can paste it straight into a model context. +- [ ] Public release of the benchmark suite +- [ ] Desktop app +- [ ] Windows support +- More to come… ## Development ```bash -pnpm install -pnpm build # build first: core's exports point at dist/ -pnpm typecheck -pnpm test - -pnpm dev:server # backend at 127.0.0.1:7364 -pnpm dev:web # web app (Vite) at 127.0.0.1:7365, /api proxied -pnpm dev:docs # docs site (Vite) at 127.0.0.1:7367 - -BASE_PATH=/ pnpm build:site # assemble landing + docs exactly like the Pages deploy +pnpm install && pnpm build # build first: core's exports point at dist/ +pnpm dev # backend + web app together (prefixed logs, deps built once) ``` -Copy `.env.example` to `.env` for model credentials in development. E2E tests run against a live model (`pnpm test:e2e`, needs `DEEPSEEK_API_KEY`). +See [CONTRIBUTING.md](CONTRIBUTING.md) for the full workspace guide: dev commands, quality gates, repo layout, and the changelog rule. + +## Citation + +If you use PenguinHarness in your research, please cite: + +```bibtex +@software{penguinharness2026, + author = {{PrismShadow Team}}, + title = {PenguinHarness: Efficient Self-Improving Harness for Everyone}, + year = {2026}, + url = {https://github.com/Prism-Shadow/penguin-harness}, + license = {Apache-2.0} +} +``` ## License [Apache-2.0](LICENSE) © 2026 Prism Shadow + +Built with ❤️ by [Yaowei Zheng](https://github.com/hiyouga) (author of [LlamaFactory](https://github.com/hiyouga/LlamaFactory)), the [PrismShadow AI Team](https://github.com/Prism-Shadow), and [Fable 5](https://www.anthropic.com/news/claude-fable-5-mythos-5). diff --git a/README.zh.md b/README.zh.md index 5a3bf2e..913fe44 100644 --- a/README.zh.md +++ b/README.zh.md @@ -4,11 +4,9 @@

PenguinHarness

-

Efficient Self-Improving Harness for Everyone

+

使用 LangChain,以 1 倍速度人工构建 Agent;
使用 PenguinHarness,以 100 倍速度用 Agent 构建 Agent。

-

- 开源、本地优先的 AI Agent 基础设施——从自动构建 Agent 到递归自我进化。 -

+

零代码 Harness CLI 与 Web UI,连接 1000+ 模型。

CI @@ -18,66 +16,113 @@

- English | 简体中文 · - 官网 · - 文档 · - 博客 + 官网 + 文档 + 博客

- - - PenguinHarness Web App——多 Session 对话与实时流式工具调用 - + Discord + X(Twitter) + 微信群

---- +

English | 简体中文

## 为什么选择 PenguinHarness -- **Simplest Is the Best**——在干净的底层接口之上刻意保持极简的工具集:更少的工具调用、更少的 Token,高效完成复杂任务。 -- **Harness for Building Agents**——基于 PenguinHarness SDK,由一个 Agent 从零开始为你自主构建完整的 Agent 应用。 -- **Harness for Recursive Self-Improvement**——基于 PenguinHarness Skills,Agent 评估并优化自己:跑 Benchmark、找失分点、产出 N+1 版本,每轮之前先做快照。 -- **本地优先且轻量**——100% 开源,一颗 CPU 即可运行,数据不出机器;经统一网关可接入 1000+ 在线与本地模型。 -- **全量可观测**——每次请求、工具调用与审批决策都以追加方式写入 Trace,任何 Session 均可从 Trace 恢复。 +三个递进的理由——从任务效果,到构建方式,再到进化能力。 -## 快速开始 +### 1. 🏆 效果同级,成本低一到两个数量级 -一行命令安装(Linux / macOS,x64 / arm64,内嵌 Node 运行时,解压即用): +刻意精简的工具集配合干净的底层接口:更少的工具调用、更少的 Token,对 DeepSeek 等开放模型深度适配。各自搭配常用模型、同一批任务,正面对比: -```bash -curl -fsSL https://github.com/Prism-Shadow/penguin-harness/releases/latest/download/install.sh | sh +

+ + + Benchmark:PenguinHarness 在数据分析题库准确率最高、编程题库与 OpenAI Codex 持平,成本仅为两者的零头 + +

+ +**数据分析准确率最高——成本只有 Claude Code 的 1/70。** + +### 2. ⚡ 一句话,让 Agent 构建 Agent 应用 + +输入一句话,Agent 为你构建完整的 Agent 应用——脚手架、代码、运行说明,一步到位: + +```text +收集 https://github.com/ericbuess/claude-code-docs 的文档,做一个化身 Claude Code 配置专家、回答带来源引用的 RAG 问答应用。 ``` -或经 npm 安装(需系统 Node >= 24,安装后的命令为 `penguin`): +这是做出来的成品——一个文档专家:检索增强、引用可点击直达原文、内置示例问题: + +https://github.com/user-attachments/assets/604eb626-0a5d-4a62-87e3-14ebade1cd5f + +**而生成整个 RAG 应用,仅消耗了 0.2 元($0.02)的 token——使用 DeepSeek V4 Pro 模型。** + +### 3. 🧬 自进化,越用越强 + +借助 PenguinHarness 技能库,Agent 自己评估、自己优化:跑 Benchmark、找失分点、发布 N+1 版——每轮之前自动快照,每个请求都可在轨迹观测中回放。 + +https://github.com/user-attachments/assets/aec49ae9-b743-467b-b247-37bedfeaa36e + +## 支持的模型 + +| 模型 | 可用供应商 | +| ---------------- | -------------------------------------------------------------------------------- | +| DeepSeek V4 | DeepSeek, OpenRouter, Fireworks AI, SiliconFlow, Qwen Token Plan | +| Kimi K3 | OpenRouter, Qwen Pay-As-You-Go | +| Kimi K2.6 | Moonshot AI | +| GLM 5.2 | Z.AI, OpenRouter, Fireworks AI, SiliconFlow, Qwen Token Plan, Qwen Pay-As-You-Go | +| Hunyuan 3 | OpenRouter | +| Qwen 3.8 Max | Qwen Token Plan(预览) | +| GPT 5.5 | OpenAI, OpenRouter | +| Gemini 3.5 Flash | Google Gemini, OpenRouter | +| Claude Opus 4.8 | Anthropic, OpenRouter | + +只要是 OpenAI 协议的端点都可以接入:从上表选择预置,或用自定义端点连接 1000+ 在线与本地模型。 + +## 系统需求 + +| 需求项 | 支持情况 | +| -------- | ------------------------------------------------- | +| 操作系统 | Linux、macOS | +| 架构 | x64、arm64 | +| 运行时 | 一行安装器自带(经 npm 安装需 Node >= 24) | +| 模型 | 至少一个模型的 API key | + +## 安装 + +### 🌐 Web 应用——面向人 + +🚀 一行安装,启动完整体验(多会话对话、Agent / 技能 / 模型管理、用量统计、轨迹观测、评估中心): ```bash -npm install -g @prismshadow/penguin-cli +curl -fsSL https://penguin.ooo/install.sh | sh +penguin web # 启动服务并打开 http://127.0.0.1:7364(首次登录:admin / penguin-2026) ``` -然后启动 Web App,或直接留在终端: +📦 或经 npm 安装:`npm install -g @prismshadow/penguin-cli`。在应用内模型页配置模型后即可对话。 + +### 🤖 CLI 与 SDK——面向 Agent + +同一引擎、可脚本化——为被 Agent 驱动而生(以及让 Agent 构建 Agent): ```bash -penguin web # 启动服务并打开 http://127.0.0.1:7364(初始账号 admin / admin123) -penguin server # 同一服务,无头运行 - -# 先配置一次模型(也可在 Web 的模型页完成) -penguin config model add --model-id deepseek-v4-pro --api-key sk-... --set-default - -penguin run -m "创建 hello.txt,内容为 Hello, Penguin" # 单次任务 -penguin chat # 交互式 REPL(/compact、/exit,Ctrl-C 中断) +penguin config model add --provider deepseek --model-id deepseek-v4-pro --api-key sk-... --set-default +penguin run -m "Create hello.txt containing Hello, Penguin" # 单次任务 +penguin chat # 交互式 REPL(/compact、/exit、Ctrl-C 中断) +penguin server # 无界面服务(与 Web 应用同一套 API) ``` -直接使用 SDK: - ```ts import { createAgent, isCompleteModelMessage, userText } from "@prismshadow/penguin-core"; const agent = await createAgent({ agentId: "default_agent" }); const session = await agent.createSession({ workspaceDir: process.cwd() }); -for await (const output of session.run([userText("创建 hello.txt 并写入 hi")], { - approve: async () => "allow", // 逐个工具审批 +for await (const output of session.run([userText("Create hello.txt containing hi")], { + approve: async () => "allow", // 按工具调用逐个审批 })) { if (isCompleteModelMessage(output) && output.payload.type === "text") { console.log(output.payload.text); @@ -85,45 +130,38 @@ for await (const output of session.run([userText("创建 hello.txt 并写入 hi" } ``` -## 仓库结构 +## 路线图 -pnpm monorepo(TypeScript,Node >= 24)。一次安装交付四层组件,共享同一数据目录(`~/.penguin/data`)与同一消息协议(OmniMessage): +- [ ] Benchmark 套件正式发布 +- [ ] 桌面端应用 +- [ ] Windows 系统支持 +- 更多规划,敬请期待…… -| 目录 | 包名 | 职责 | -| --- | --- | --- | -| [`packages/core`](packages/core) | `@prismshadow/penguin-core` | SDK 与引擎:ReAct 循环、OmniMessage 协议、LLM/Environment 接口契约、Agent State、Trace | -| [`packages/cli`](packages/cli) | `@prismshadow/penguin-cli` | `penguin` 命令:REPL、单次运行、模型与 Vault 配置、服务启动 | -| [`packages/server`](packages/server) | `@prismshadow/penguin-server` | Web 服务端:HTTP API + SSE 流式、多用户认证、Project 授权、用量统计 | -| [`packages/web`](packages/web) | `@prismshadow/penguin-web` | Web App:多 Session 对话、Agent/技能/模型管理、Trace 观测、评估中心 | -| [`packages/skills`](packages/skills) | `@prismshadow/penguin-skills` | 内置技能库(Agent 创建、Benchmark 设计、评估、优化等) | -| [`packages/landing`](packages/landing) | — | 产品落地页(本仓库官网) | -| [`packages/docs`](packages/docs) | — | 文档站(双语,部署于 `/docs/` 路径) | - -职责按事实来源划分:**SDK** 负责协议与执行(消息解析、运行循环、工具),**Server** 负责多用户运行时(认证、SSE 流式、定时任务),`~/.penguin/data` 下的**文件层**承载一切可编辑与被记录的状态(Prompt、Skill、密钥、Trace)。逐项对应表见[架构总览 → 职责划分](https://penguin.ooo/docs/architecture)。 - -## 文档 - -文档站覆盖使用与设计两个层面:[产品介绍](https://penguin.ooo/docs/) · [快速开始](https://penguin.ooo/docs/quickstart) · [架构总览](https://penguin.ooo/docs/architecture) · [OmniMessage 协议](https://penguin.ooo/docs/omni-message) · [接口契约](https://penguin.ooo/docs/interfaces) · [Agent 运行循环](https://penguin.ooo/docs/agent-loop) · [CLI 参考](https://penguin.ooo/docs/cli) · [Server API](https://penguin.ooo/docs/server-api) · [配置参考](https://penguin.ooo/docs/configuration) - -每页文档都带「复制 Markdown」按钮,可直接粘贴进模型上下文。 - -## 本地开发 +## 参与开发 ```bash -pnpm install -pnpm build # 先构建:core 的导出指向 dist/ -pnpm typecheck -pnpm test - -pnpm dev:server # 服务端 127.0.0.1:7364 -pnpm dev:web # Web App(Vite)127.0.0.1:7365,/api 代理到服务端 -pnpm dev:docs # 文档站(Vite)127.0.0.1:7367 - -BASE_PATH=/ pnpm build:site # 按 Pages 部署的方式组装 落地页 + 文档 +pnpm install && pnpm build # 先构建:core 的导出指向 dist/ +pnpm dev # 服务端 + Web 一起启动(带前缀日志,依赖只构建一次) ``` -开发态模型凭据可复制 `.env.example` 为 `.env` 填写。E2E 测试走真实模型(`pnpm test:e2e`,需要 `DEEPSEEK_API_KEY`)。 +完整工作区指南见 [CONTRIBUTING.md](CONTRIBUTING.md):开发命令、质量门禁、仓库结构与 changelog 规则。 -## 许可证 +## 引用 + +如果 PenguinHarness 对你的研究有帮助,请引用: + +```bibtex +@software{penguinharness2026, + author = {{PrismShadow Team}}, + title = {PenguinHarness: Efficient Self-Improving Harness for Everyone}, + year = {2026}, + url = {https://github.com/Prism-Shadow/penguin-harness}, + license = {Apache-2.0} +} +``` + +## 协议 [Apache-2.0](LICENSE) © 2026 Prism Shadow + +由 [LlamaFactory](https://github.com/hiyouga/LlamaFactory) 作者 [Yaowei Zheng](https://github.com/hiyouga)、[PrismShadow AI Team](https://github.com/Prism-Shadow) 与 [Fable 5](https://www.anthropic.com/news/claude-fable-5-mythos-5) 共同用 ❤️ 构建。 diff --git a/assets/readme/benchmark-dark.svg b/assets/readme/benchmark-dark.svg new file mode 100644 index 0000000..f927211 --- /dev/null +++ b/assets/readme/benchmark-dark.svg @@ -0,0 +1,48 @@ + +Accuracy · suite total · higher is better +Total cost (USD) · lower is better +Data analysis — 15 tasks, single run + +PenguinHarness + +66.67% +Claude Code + +53.33% +OpenAI Codex + +53.33% + +PenguinHarness + +$0.55 +Claude Code + +$38.48 +OpenAI Codex + +$19.41 +Coding — 40 tasks × 2 runs + +PenguinHarness + +71.25% +Claude Code + +86.25% +OpenAI Codex + +71.25% + +PenguinHarness + +$3.81 +Claude Code + +$146.97 +OpenAI Codex + +$220.08 +Each harness runs the model it is normally paired with: PenguinHarness on DeepSeek V4 Pro, Claude Code on Claude Opus 4.8, +OpenAI Codex on GPT-5.5. Accuracy, Tokens and cost are suite totals at official pricing. + diff --git a/assets/readme/benchmark-light.svg b/assets/readme/benchmark-light.svg new file mode 100644 index 0000000..31c480c --- /dev/null +++ b/assets/readme/benchmark-light.svg @@ -0,0 +1,48 @@ + +Accuracy · suite total · higher is better +Total cost (USD) · lower is better +Data analysis — 15 tasks, single run + +PenguinHarness + +66.67% +Claude Code + +53.33% +OpenAI Codex + +53.33% + +PenguinHarness + +$0.55 +Claude Code + +$38.48 +OpenAI Codex + +$19.41 +Coding — 40 tasks × 2 runs + +PenguinHarness + +71.25% +Claude Code + +86.25% +OpenAI Codex + +71.25% + +PenguinHarness + +$3.81 +Claude Code + +$146.97 +OpenAI Codex + +$220.08 +Each harness runs the model it is normally paired with: PenguinHarness on DeepSeek V4 Pro, Claude Code on Claude Opus 4.8, +OpenAI Codex on GPT-5.5. Accuracy, Tokens and cost are suite totals at official pricing. + diff --git a/assets/readme/rag-app-en-dark.webp b/assets/readme/rag-app-en-dark.webp new file mode 100644 index 0000000..52392a7 Binary files /dev/null and b/assets/readme/rag-app-en-dark.webp differ diff --git a/assets/readme/rag-app-en-light.webp b/assets/readme/rag-app-en-light.webp new file mode 100644 index 0000000..0a2115b Binary files /dev/null and b/assets/readme/rag-app-en-light.webp differ diff --git a/assets/readme/rag-app-zh-dark.webp b/assets/readme/rag-app-zh-dark.webp new file mode 100644 index 0000000..8a698e0 Binary files /dev/null and b/assets/readme/rag-app-zh-dark.webp differ diff --git a/assets/readme/rag-app-zh-light.webp b/assets/readme/rag-app-zh-light.webp new file mode 100644 index 0000000..b9dba57 Binary files /dev/null and b/assets/readme/rag-app-zh-light.webp differ diff --git a/changelog/0.0.2/2026-07-20-landing-site.md b/changelog/0.0.2/2026-07-20-landing-site.md new file mode 100644 index 0000000..d362c70 --- /dev/null +++ b/changelog/0.0.2/2026-07-20-landing-site.md @@ -0,0 +1,84 @@ +# Landing site + +The penguin.ooo landing page: story, structure, animations, navigation, and the install domain. + +## Landing, README, and blog retell one story, with penguin.ooo/install.sh + +The marketing surfaces now tell the same 1x/100x story, persuade through one numbered "Why" section, and install from the site's own domain. + +- **penguin.ooo/install.sh** — the landing site ships a thin `public/install.sh` that forwards to the latest GitHub release installer (GitHub Pages cannot serve real redirects), so `curl -fsSL https://penguin.ooo/install.sh | sh` works everywhere; the hero command, README (en/zh), docs installation/quickstart (en/zh), and blog posts all switch to it. +- **Landing** — the hero replaces the rotating-word headline with the two-line story ("With LangChain, you build agents by hand — at 1x; with PenguinHarness, agents build agents — at 100x", the 100x fragment in brand color) plus the "zero-code Harness CLI and Web UI, 1000+ models" subtitle. One numbered Why section (1 benchmark suites + DeepSeek tuning, 2 one-sentence build demo with the real RAG shot, 3 self-evolution loop with a "demo video coming soon" pill) absorbs the former Pillars / Showcase / Benchmark / SelfImprove sections; UI screenshots leave the home page; nav anchors become why / quickstart / contract / features (reason 1 keeps the #benchmark anchor). The blog already shares the landing Nav, theme, and language settings. +- **README (en/zh)** — subtitle now says "Harness CLI and Web UI"; website/docs/blog/community links render as shields.io badges; the top screenshot and the Changelog-Blog-Docs section are gone; the three reasons sit numbered with emoji under "Why PenguinHarness"; the models table keeps only a one-line "any OpenAI-protocol endpoint works" note; requirements become a table; install sections gain emoji; the roadmap adds a desktop app and Windows support; the credits link LlamaFactory, the PrismShadow AI Team, and Fable 5. +- **Blog** — the introduction post (en/zh) leads with the same story line and phrasing. + +## Landing back to the classic structure, with tabbed features/cases and a docs-matching nav + +The previous landing overhaul went too far; the classic section set returns, with the new material folded in as tabs and side sections instead of replacements. + +- **Restored as they were**: the rotating hero headline ("Efficient Self-Improving Harness for Developers/Enterprises"), the three pillars, the self-improvement loop (now with a "demo video coming soon" pill), Quickstart, the standalone Benchmark section, CONTRACT.md, and Security. +- **LangChain comparison moved, not deleted**: the 1x/100x story lives in its own compact section after the pillars — two cards (LangChain, hand-built, 1x, de-emphasized vs PenguinHarness, agents building agents, 100x, brand-emphasized) with the story sentence as the subtitle. +- **Features become switchable tabs**: all nine feature descriptions remain, each as a tab with its icon; the three views with real captures — multi-session chat, Trace view, Agent evaluation — show their locale/theme-matched screenshot in the active panel, bringing the Trace and benchmark shots back onto the page. +- **Use cases become tabs too**: a Cases section with a tab bar holds the RAG case only for now (prompt + captured result); future cases append as tabs. +- **Community section added at the end**: Discord / X / WeChat group / GitHub as outbound cards after the CTA. +- **Docs nav now mirrors the landing nav exactly** — same link row (Highlights / Quick start / Benchmark / CONTRACT.md / Features / Blog / Docs) with the same sliding hover pill, anchoring into the landing one level up, so the two sites link seamlessly; the old standalone "Website" link is absorbed by the row. +- **Launch blog post back to its original skeleton**: the "Why PenguinHarness" three-pillar bullets, the benchmark section wording, and the closing "start right now" paragraph return, with the GDPevo origin story, images, models table, roadmap, and community call-to-action kept as additions rather than replacements. + +## Landing trace screenshots: same opened timeline in English and Chinese + +The English trace screenshots showed the empty "select a Session" state while the Chinese ones showed a full trace; all four now capture the same opened trace with stats and an execution timeline containing tool calls. + +- The capture script navigated to `/traces?sessionId=...` — but the traces page only honors the session deep link when `agentId=` is present (the product's own links always carry both), so selection relied on a fragile title click that silently failed for English via `.catch()`. The script now uses the canonical `?agentId=default_agent&sessionId=...` deep link and waits for the timeline's `exec_command` lanes before shooting, so an empty capture fails loudly instead of shipping. +- The scripted English session title was 31 chars and core clips titles at `TITLE_MAX_CHARS = 30`, producing "…Agent ap" in the shots; the mock title is now "Build a data-analysis Agent" (27 chars). +- Regenerated `traces-{en,zh}-{light,dark}.webp`; chat and benchmark shots are unchanged. + +## Landing polish: sliding announcements, construction animation, trend value labels, feature split + +A polish pass over the landing page's motion and grouping, plus web-first usage guidance in the blog. + +- **Announcement bar** — light brand-tinted background; both entries now link to blog posts (the new-models entry to the launch post, the credits entry to the AMD post); switching SLIDES horizontally (auto-advance, hover pause) with dot indicators instead of arrows; the trailing arrow icon is gone. +- **LangChain comparison** — a looping "construction" animation on a shared cycle: the LangChain card lays one gray block at a time and never tops out before the cycle resets, while the PenguinHarness card raises a whole brand-blue skyline in under two seconds and holds it. Reduced-motion shows both skylines complete. +- **Self-improvement trends** — the three outcome charts are now driven by one rAF clock, and a value label rides each curve's moving head (score climbing, cost and time falling), changing as the line draws; reduced-motion shows the finished line with its final value. +- **Features split by capture** — only the three features with real screenshots (multi-session chat, trace view, agent evaluation) form the tab bar; the remaining six sit below as the classic card grid, closed with an "and more…" card. +- **Blog usage goes web-only** — the launch post's getting-started and the AMD post's setup now guide through `penguin web` and the Models page only; the CLI config/run alternatives are removed to funnel users into the Web UI. + +## Announcement carousel: one-direction slide, arrow kept, dots removed + +The announcement bar keeps the small arrow marking each entry as a click-through link, drops the dot switch buttons on the right, and auto-rotates by sliding in ONE direction only: a clone of the first slide follows the last, and once the clone is fully in view the track snaps back without animation — the bar never visibly slides backwards. Hovering still pauses the rotation. + +## Feature tabs drop the description card + +The three screenshot-backed feature tabs no longer render an icon card above the capture; the feature's one-line description now sits as centered text directly above the image (the tab chip already carries the title and icon). The no-capture features keep their card grid below. + +## Trend value labels carry units with animated decimals + +The self-improvement loop's three outcome labels now read as real measurements while they ride the curve head: score as a percentage with one decimal (79.0%), cost in dollars with two decimals ($0.25), and time in seconds with one decimal (83.0s) — the decimal part animates continuously along with the line. + +## Fix the nav hover highlight sweeping in from the edge + +The landing and docs navs' sliding hover pill animated from its hidden state at the nav's left edge, so the first hover sent a gray pill flying across the whole link row (and it slid back on leave). + +The pill now appears IN PLACE under the first link it lands on — the position jumps with only the fade animating — slides while moving between links, and fades out where it is when the pointer leaves. Applied identically to the landing nav and the landing-parity docs nav. + +## Merge main to restore the persistent nav active highlight + +The branch predated main's PR #6 ("correct navigation state and anchor scrolling"), so it was missing the nav's persistent active highlight — the black chip that marks the current section or route — and would have reverted that fix on merge. Merging origin/main brings it back and reconciles it with the branch's nav work. + +- Landing nav: section links route through `/#id` again with `getActiveNavItem` tracking (new `lib/nav-state.ts` + tests), the active link keeps its black chip with `aria-current`, and the branch's in-place hover pill behavior is preserved on top. +- Also restored from #6: anchor targets scroll below the sticky header via `.section-anchor` (ids moved onto the section's inner div), footer/hero/CTA anchors as router links, and the docs brand logo returning to the landing site. +- Docs nav (landing-parity) additionally marks its own "Docs" link as the current page with the same black chip. + +## Nav highlight follows the live scroll position + +The nav's black current-item chip previously updated only from the URL hash (i.e. on click); while scrolling, it stayed stale. A scroll-spy hook now measures the five section anchors on every scroll frame (rAF-throttled, document-level capture so any scrolling container works) and lights up the LAST section whose top has crossed the activation line under the sticky header — null above the first section, and in-between sections keep the previous anchor lit. On the home page the highlight is fully live; other routes keep route-based state (Blog). + +## The Cases tab shows the finished RAG app with a localized prompt + +The RAG case now displays the condensed claude-code-docs configuration-expert prompt from the strings catalog (localized zh/en) and the finished-product app screenshot matched to the visitor's locale and theme, replacing the shared English prompt and the PenguinHarness chat capture. + +## The penguin sled game joins the cases, with a $0.02 cost hook on the RAG demo + +The draft-screen game example becomes a cute Antarctic penguin sledding game, lands on the landing page as a second case after RAG, and every RAG demo now leads with how little it cost to generate. + +- The draft-screen example card (zh/en) swaps the 2D motocross runner for a penguin sledding game: Space to jump the rocks on the ice, speed and difficulty ramping up, live scoring, one-click restart, cute cartoon look — the full detailed prompt still gets submitted as-is. +- The landing Cases section gains a "Penguin sled game" tab after the RAG one, showing the finished play screen per language and theme (light = polar day, dark = polar night with an aurora) from the new dependency-free `penguin-game-mockup.html` + `capture-game-mockup.mjs` pipeline. +- README (en/zh), the launch blog post (en/zh), and the landing RAG case now carry an emphasized hook under the finished-product shot: generating the entire RAG app burned just $0.02 (¥0.2) of tokens on DeepSeek V4 Pro. diff --git a/changelog/0.0.2/2026-07-20-models-and-credentials.md b/changelog/0.0.2/2026-07-20-models-and-credentials.md new file mode 100644 index 0000000..324982d --- /dev/null +++ b/changelog/0.0.2/2026-07-20-models-and-credentials.md @@ -0,0 +1,259 @@ +# Model catalog, Models page, and credential handling + +Preset provider groups, catalog entries and ordering, and the Models page features built around them. + +## Add the Qwen Token Plan provider group to the model catalog + +The built-in catalog gains a Qwen Token Plan subscription gateway group (OpenAI-compatible, +preset base URL) with five models — qwen3.8-max-preview, qwen3.7-max, qwen3.7-plus, glm-5.2, +and deepseek-v4-pro — plus a custom provider logo. + +## Details + +- New provider `qwen-token-plan` ("Qwen Token Plan"), placed with the gateway cluster after + SiliconFlow: OpenAI-compatible endpoint preset to + `https://token-plan.cn-beijing.maas.aliyuncs.com/compatible-mode/v1`, API-key page + `https://platform.qianwenai.com/pricing/token-plan`, model-id docs page + `https://platform.qianwenai.com/docs/token-plan/personal/token-plan-personal-overview`; + env fallback is `OPENAI_API_KEY`/`OPENAI_BASE_URL` like the other gateways. +- Five catalog entries with `client_type: openai` and the inlined base URL. Vision flags per + the plan's supported-model table (qwen3.8-max-preview and qwen3.7-plus see images; the + rest do not). Pricing and context windows come from each model's page at + `www.qianwenai.com/models/` (official CNY list prices; limited-time promotions are not + stored): qwen3.7-max ¥2.4/¥12/¥36, qwen3.7-plus ¥0.4/¥2/¥8, glm-5.2 ¥2/¥8/¥28, + deepseek-v4-pro ¥1/¥12/¥24 (cache-hit/input/output per M tokens), windows 1M (glm-5.2: + 1.04M). qwen3.8-max-preview is preview-only with a quota-multiplier promotion and no + per-token list price, so it alone carries no pricing (costs read as 0, same as unpriced + user models); the pricing invariant in the catalog tests is scoped to that one entry. +- Catalog invariant updated: bare model ids may now repeat across providers (the gateway + resells vendor models under their upstream ids, e.g. `glm-5.2` / `deepseek-v4-pro`); + uniqueness is the `(provider, model_id)` pair, matching the catalog's sole lookup key. +- New provider logo: the official Qwen wordmark (icon + lettering) from the brand SVG, + gradient fills flattened to currentColor monochrome, coordinates rounded to 2dp. +- Docs (models.en/zh) provider table and gateway notes updated. + +## Add the Qwen Pay-As-You-Go provider group + +A pay-per-token gateway group below Qwen Token Plan (DashScope's OpenAI-compatible +endpoint) with four preset models — qwen3.7-max, qwen3.7-plus, and the resold +vendor-prefixed kimi/kimi-k3 and ZHIPU/GLM-5.2 — priced from each model's official page. + +## Details + +- New provider `qwen-pay-as-you-go` ("Qwen Pay-As-You-Go") right after Qwen Token Plan in + the gateway cluster: preset base URL `https://dashscope.aliyuncs.com/compatible-mode/v1`, + API-key link `https://platform.qianwenai.com/docs/api-reference/preparation/api-key`, + models page `https://www.qianwenai.com/models`; `OPENAI_*` env fallback like the other + gateways. +- Four entries (`client_type: openai` + inlined endpoint), official CNY list prices and + specs from each model's page: kimi/kimi-k3 ¥2/¥20/¥100 (1.04M, vision), qwen3.7-max + ¥2.4/¥12/¥36 (1M), ZHIPU/GLM-5.2 ¥2/¥8/¥28 (1.04M), qwen3.7-plus ¥0.4/¥2/¥8 (1M, + vision). Resold third-party models keep their vendor-prefixed upstream ids. +- The group shares the Qwen emblem (glyph extracted into a shared constant), and + `modelHomepageUrl` URL-encodes the slash-prefixed ids + (`.../models/ZHIPU%2FGLM-5.2`). Docs (models.en/zh) provider tables and gateway notes + updated. + +## Add the Fireworks AI provider group + +An OpenAI-protocol gateway group below OpenRouter with five preset models (GLM-5.2, +Kimi K2.7 Code, DeepSeek V4 Pro, MiniMax M3, DeepSeek V4 Flash), priced from each model's +Fireworks page. + +## Details + +- New provider `fireworks` ("Fireworks AI") right after OpenRouter in the gateway cluster: + preset base URL `https://api.fireworks.ai/inference/v1`, API-key page + `https://app.fireworks.ai/settings/users/api-keys`, models page + `https://app.fireworks.ai/models`; `OPENAI_*` env fallback like the other gateways. +- Five entries (`client_type: openai` + inlined endpoint) with standard-serverless USD + pricing (cached input / uncached input / output per M tokens) and specs from each page: + glm-5p2 $0.14/$1.40/$4.40 (1M), kimi-k2p7-code $0.19/$0.95/$4.00 (262K, vision), + deepseek-v4-pro $0.15/$1.74/$3.48 (1M), minimax-m3 $0.06/$0.30/$1.20 (512K, vision), + deepseek-v4-flash $0.03/$0.14/$0.28 (1M). API model ids use Fireworks' full + `accounts/fireworks/models/` form (sent verbatim). +- `modelHomepageUrl` maps the `accounts//models/` id to the model page + (`app.fireworks.ai/models//`), falling back to the models listing for + nonconforming user-added ids. A simplified starburst glyph approximates the brand mark + (same approach as Z.AI). Docs (models.en/zh) provider tables and gateway notes updated. + +## Expand the OpenRouter catalog with twelve models + +The OpenRouter gateway group grows from 4 to 16 entries, adding the current flagship and +free tiers with pricing, context windows, and vision flags taken from each model's +OpenRouter page. + +## Details + +- Added (ordered by output price): anthropic/claude-fable-5 ($10/$50), openai/gpt-5.6-sol + ($5/$30), openai/gpt-5.5 ($5/$30), anthropic/claude-opus-4.8 ($5/$25), + anthropic/claude-opus-4.7 ($5/$25), moonshotai/kimi-k3 ($3/$15), openai/gpt-5.6-terra + ($2.50/$15), anthropic/claude-sonnet-5 ($2/$10), z-ai/glm-5.2 ($0.93/$3), + deepseek/deepseek-v4-pro ($0.435/$0.87), deepseek/deepseek-v4-flash ($0.09/$0.18), and + nvidia/nemotron-3-ultra-550b-a55b:free (input/output per M tokens; all 1M context). +- Vision per the pages: the Claude models, GPT-5.5, and Kimi K3 accept image input; the + rest do not. +- None of these pages list cache pricing, so cache_read carries the standard input price + (no discount). The :free tier stores a genuine $0 price — not "unknown" — so costs + correctly compute to 0; the catalog pricing invariant gains a free-tier case. + +## Add Grok 4.5 to the OpenRouter catalog + +`x-ai/grok-4.5` joins the OpenRouter group: $2 input / $6 output per M tokens (no cache +price listed, so cache_read carries the input price), 500K context, vision-capable — +inserted at its dictionary position. + +## Add Gemini 3.5 Flash to the OpenRouter catalog + +`google/gemini-3.5-flash` joins the OpenRouter group: $1.50 input / $9 output per M tokens +(no cache price listed, so cache_read carries the input price), 1M context, vision-capable +— at its dictionary position; the README model table's Gemini row now lists OpenRouter too. + +## Order catalog models by dictionary, newer versions first + +Within each provider group, catalog entries are now in dictionary order by model id, except +that newer versions of the same series come first (gpt-5.6-* before gpt-5.5, +claude-opus-4.8 before 4.7, glm-5.2 before glm-5) — precomputed by hand in the catalog +literal, with no runtime sorting anywhere. + +## Details + +- Every provider section of MODEL_CATALOG is hand-reordered: dictionary order + (case-insensitive) across families and tiers; within a version series, the newest version + block leads, tiers inside a version staying alphabetical. Section comments and the + exact-order test assertions are updated to match. +- The order flows everywhere in-group order is preserved: new Projects' preset config, the + models page cards, and the chat model dropdown (orderModelsLikeLibrary). Existing Project + configs keep their stored order — the sync-presets merge deliberately preserves local + positions. + +## Catalog data: official Fireworks logo, two SiliconFlow models, per-model vendor pages + +The Fireworks group now wears the official burst mark, SiliconFlow gains +moonshotai/Kimi-K2.7-Code and deepseek-ai/DeepSeek-V4-Flash, and Z.AI / Moonshot models +link to their per-model docs pages. + +## Details + +- Fireworks logo: the official three-stroke burst mark (viewBox 0 0 638 315, currentColor) + replaces the interim starburst approximation. +- SiliconFlow entries (official CNY pricing, dictionary position): + deepseek-ai/DeepSeek-V4-Flash ¥0.02/¥1/¥2 (1M) and moonshotai/Kimi-K2.7-Code + ¥1.3/¥6.5/¥27 (262K, vision). +- modelHomepageUrl: zhipu -> `docs.z.ai/guides/llm/`; moonshot -> + `platform.kimi.com/docs/pricing/chat-k` (kimi-k2.6 -> chat-k26), + nonconforming ids falling back to the group's models page. + +## Add a sync-presets button to the Models page + +A small owner-only button next to the Models page search box merges the built-in catalog +into the Project's model table: catalog entries missing locally are added, entries present +on both sides are reset to the catalog's fields, locally added models and API keys stay +untouched. + +## Details + +- Union semantics (`catalog-sync.ts`, pure and unit-tested): keyed by the + `(provider, model_id)` pair. Catalog-only entries are appended (gateway base URLs preset); + intersecting entries take the catalog's context window, pricing (including removal when + the catalog carries none, e.g. the unpriced preview model), protocol, base URL, and vision + flag — the catalog wins wherever the two differ; local-only models (including + user-defined groups) are kept verbatim and in place. +- Credentials are structurally untouched: merged rows submit no `apiKey` (the PUT + full-table replace keeps stored keys when the field is absent), and existing rows keep + their credential state; a user base-URL override on a preset model is reset to the + catalog's (the API-key carve-out is the only one). +- Feedback via toasts: "Presets synced: N added, M updated", or "already up to date" + without a PUT when nothing differs. Strings added to both locales. +- The Qwen Token Plan provider logo is trimmed to the official emblem only (the wordmark + lettering dropped), on a square viewBox. + +## Model test no longer fails on thinking-only responses + +Testing a reasoning-heavy model (e.g. qwen3.8-max-preview) failed with "OpenaiClient +returned no content other than thinking (finish_reason=\"length\")": the connectivity +probe's tiny output cap was burned entirely on thinking. The probe now counts a streamed +thinking-only ending as reachable — the endpoint, credential, and model id all +demonstrably work. + +## Details + +- The probe deliberately sends one "ping" with `maxTokens: 16` and thinking disabled + (single-digit token cost by design). Reasoning models behind OpenAI-compatible endpoints + can ignore the disabled thinking level, hit `finish_reason=length` with no text, and + AgentHub 0.4 raises `EmptyResponseError` — collapsed to a `malformed` outcome, which the + probe previously reported as a test failure. +- `testModel` now tracks whether genuine model content (thinking or text, partial or + complete) was streamed, and a `malformed` ending after streamed content passes the test; + timeouts, auth/parameter failures, and malformed endings with nothing received still + fail. The logic lives in two pure functions (`isProbeContent` / `probeVerdict`) with unit + tests, including the exact qwen3.8-max-preview case. + +## Group speed test on the Models page + +Each model group header gains an owner-only speed-test action: after a quota warning it +probes the group's models one at a time, measuring time-to-first-token and output rate, and +writes tone-colored badges (green / yellow / red) onto each card; the model-homepage link +moves from the card corner into the config dialog. + +## Details + +- Server: the model-test endpoint gains a `speed` flag — the probe's output cap rises from + 16 to 64 tokens so a real streaming window exists, and the response now carries `ttftMs` + (request start -> first streamed content) and `tps` (output tokens over the streaming + window, from the completed stream's usage report; thinking-only endings carry TTFT but no + rate). The plain connectivity test is unchanged. +- Web: a gauge button on each group header (owner-only) opens a confirmation dialog warning + that one real request per model will consume API quota; on confirm the group is tested + **strictly sequentially** (concurrent probes trip provider rate limits), each result + landing on its card as it finishes. Badges: clock icon + ms for TTFT (green < 1s, yellow + <= 3s, red beyond), zap icon + tok/s for TPS (green >= 40, yellow >= 15, red below); + failures show a red "test failed" with the reason on hover. Thresholds live in a pure, + unit-tested helper; results are session-scoped. +- The model-homepage link moves off the card corner into the config dialog next to the + "get model ids" link (the card stays a single clickable surface; the freed corner hosts + the speed badges). +- Refinements: the group-header actions (add model / bulk API key / speed test) are all + icon + text buttons; the speed badges live on the card's meta line in their own + non-shrinking slot (the numbers never crowd or wrap the title row); the probe prompt + discourages reasoning and ends with an empty `` block so reasoning models + skip their thinking phase instead of burning the probe budget on it. + +## Model page refinements: draft follows the new default, ordered dropdown, GPT vision, homepage links + +Four refinements: changing the Project's default model now resets the stored draft's model +selection to follow the new default; the chat model dropdown lists models in the same order +as the model library page; all GPT models are marked vision-capable; and model cards link to +each model's homepage. + +## Details + +- Draft follows the default: when saving a default-model change on the Models page, the + current user's stored draft for that Project drops its `modelRef` — the draft chat then + resolves the (new) default live instead of pinning the old model forever. +- Dropdown order: a new `orderModelsLikeLibrary` helper (unit-tested) flattens the library + grouping — built-in provider groups in MODEL_PROVIDERS order, user-defined groups after, + custom last, in-group order preserved — and the chat model dropdown now uses it. +- GPT vision: `openai/gpt-5.6-sol` and `openai/gpt-5.6-terra` are flipped to + vision-capable — GPT models are uniformly multimodal (OpenAI product-line policy) even + where the gateway page omits the modality. +- Homepage links: a new `modelHomepageUrl` helper (unit-tested) — OpenRouter and Qwen Token + Plan have stable per-model URL patterns (working for user-added ids in those groups too; + the unpaged Token Plan preview model falls back to the plan overview), direct vendors link + to their model docs page, custom/user-defined groups have none. Model cards show the link + as a corner external-link icon (a sibling of the clickable card, since interactive + elements must not nest). + +## penguin-sdk and agenthub-models keep model keys project-local + +Both skills now spell out where model API keys belong: in the project under the working directory, never in the user's global `~/.penguin`. + +- penguin-sdk (v6) and agenthub-models (v3) instruct configuring keys with the penguin CLI into the app's own data root under CWD (`penguin config model add --root …`), or relying on vault-injected environment variables; reading, copying or falling back to model keys stored in the global `~/.penguin` directory is explicitly forbidden — that config belongs to the person running Penguin, not to the app being built. +- The no-key path stays as before: stop and ask the user to open the agent's settings via the gear icon on its card and update the key vault. + +## Vault edits take effect on the next task; global keys don't count as usable + +A vault save now invalidates the Agent's cached Session runtimes so the next task runs with the new values, and the AI-app skills stop counting keys from the global `~/.penguin` as usable. + +- Server: `PUT /agents/:agentId/vault` bumps the Agent's config generation in the session manager; every runtime built before the update is discarded on its next idle access and re-resumed via the loader (resume re-reads `agent_state/.vault.toml`; history is preserved through the Trace). A task already in flight keeps the values it started with and rebuilds on the first access after it finishes. Unit and integration tests cover idle/busy entries and the HTTP wiring; the configuration docs (en/zh) and the web Vault tab hint document the new semantics. +- penguin-sdk (v9) and agenthub-models (v6): only two sources count as a usable credential — a vault-injected environment variable, or a key configured in the app's own data root (`penguin config model list --root `). Keys in the global `~/.penguin` (what a bare `penguin config model list` without `--root` reads — the CLI defaults to the global root) or any other `.penguin` directory never count and must never be used or copied; when no counted key is usable, stop immediately and ask the user to configure one instead of building or retrying. diff --git a/changelog/0.0.2/2026-07-20-readme-blog-and-docs.md b/changelog/0.0.2/2026-07-20-readme-blog-and-docs.md new file mode 100644 index 0000000..cc85760 --- /dev/null +++ b/changelog/0.0.2/2026-07-20-readme-blog-and-docs.md @@ -0,0 +1,74 @@ +# READMEs, blog, and docs site + +## READMEs + +The repository READMEs (en/zh). + +### Restructure the README around the product story + +The README now leads with the agents-build-agents pitch and community links, then three +feature showcases (benchmark chart, one-sentence RAG demo, self-evolution), followed by +changelog/blog/docs, supported models, human-first installation, a roadmap, CONTRIBUTING, +a citation, and credits. + +### Details + +- New narrative header: "With LangChain, you build agents by hand — at 1x speed. With + PenguinHarness, agents build agents — at 100x." with the subtitle "A zero-code CLI and + Web UI, connected to 1000+ models," plus community links (Discord / X / WeChat). +- Feature 1 "Simple and Efficient": light/dark benchmark bar charts generated from the + landing benchmark data (accuracy and cost per run vs Claude Code and OpenAI Codex, all + driven by DeepSeek V4 Pro), committed as `assets/readme/benchmark-{light,dark}.svg`. +- Feature 2 "Build an Agent in One Sentence": the RAG one-sentence prompt plus a real + product screenshot captured by the new `packages/landing/scripts/capture-readme-demo.mjs` + (same real-server + mock-LLM pipeline as the landing shots), committed as + `assets/readme/rag-demo-{light,dark}.webp`. +- Feature 3 "Self-Evolution": copy describing the evaluate-optimize-snapshot loop with an + HTML-comment placeholder for the upcoming demo video. +- New sections: Changelog / Blog / Docs links, a supported-models table (DeepSeek V4, + Kimi K3, GLM 5.2, Hunyuan 3, Qwen 3.8 Max, GPT 5.5, Gemini 3.5 Flash, Claude Opus 4.8 + with their providers, plus the 1000+-via-gateways note), Requirements and Installation + split into "Web App — for humans" and "CLI & SDK — for agents", a Roadmap (benchmark + suite release), a BibTeX citation ({PrismShadow Team}), and the license/credits footer. +- New `CONTRIBUTING.md` absorbs the developer content: dev commands, repo layout table, + quality gates, the English-only and changelog working rules, and the README-asset + regeneration notes; the README's Development section now points there. +- `README.zh.md` mirrors the new structure in Chinese. + +### Refresh the README model table against the current catalog + +The supported-models table (the same eight models) becomes two columns — model on the +left, the comma-separated providers it's available from on the right (per today's catalog) +— and the note below now names all five OpenAI-compatible gateways. + +### Details + +- Availability per the catalog: DeepSeek V4 in five groups, GLM 5.2 in six, Kimi K3 via + OpenRouter and Qwen Pay-As-You-Go, Qwen 3.8 Max as the Token Plan preview, GPT 5.5 and + Claude Opus 4.8 native + OpenRouter, Hunyuan 3 via OpenRouter, Gemini 3.5 Flash native. +- The gateway note lists OpenRouter, Fireworks AI, SiliconFlow, Qwen Token Plan, and + Qwen Pay-As-You-Go. README.zh.md mirrors. + +### Split badge rows and showcase the finished RAG app + +The site badges (Website / Docs / Blog) and the community badges (Discord / X / WeChat) now sit on separate lines. The one-sentence example becomes the condensed claude-code-docs configuration-expert prompt (the chat page's example task carries the full version), and the demo image shows the FINISHED PRODUCT — the generated docs-expert app with cited clickable sources and example questions (`assets/readme/rag-app--.webp`, per-language shots; the zh README shows the Chinese prompt and shots) — instead of the PenguinHarness chat UI. The mockup renderer (`rag-app-mockup.html` + rewritten `capture-readme-demo.mjs`, no server needed) comes along so the assets stay regenerable. + +## Blog and docs site + +Blog posts and the docs site. + +### Announcement bar, AMD Fireworks-credits blog post, and the GDPevo launch story + +The site gains a rotating announcement bar, a new campaign post, and a launch post that finally tells the whole story. + +- **Announcement bar** — a switchable bar above the nav (auto-rotates every 6s, paused on hover, prev/next chevrons): entry 1 announces Kimi K3 and Qwen 3.8 Max availability (links to the models docs), entry 2 the $50 Fireworks credits campaign (links to the new post). Bilingual, on every page, scrolls away with the page. +- **New blog post `fireworks-credits-amd` (en/zh)** — announces the AMD AI Developer Program partnership bringing free Fireworks redemption codes: step-by-step application (join ADP, Member Perks, form with Fireworks AI selected, review, coupon email, redeem + API key, screenshots adapted from WhatGhost's guides with credit), then a three-step PenguinHarness setup (install, Fireworks group bulk-key + presets + speed test, run). +- **Launch post rewrite (en/zh)** — now opens with the GDPevo origin story: self-evolution was validated in the team's GDPevo Benchmark (linked), and bringing it to everyone is why PenguinHarness exists. The rest mirrors the README: three numbered reasons with the benchmark chart and RAG demo images (served from the site's own /blog-assets/), the security contract, the models table with the any-OpenAI-protocol note, install/usage steps, the roadmap (benchmark suite, desktop app, Windows), and a closing community call-to-action (Discord / X / WeChat / GitHub). + +### The launch post settles on the numbered three-reasons structure + +The launch blog post (en/zh) keeps the GDPevo origin story and the numbered "Why PenguinHarness" structure — ### 1 better on complex tasks at lower cost (benchmark chart + tables), ### 2 one-sentence Agent-builds-your-app (prompt + demo shot), ### 3 self-evolution — followed by the security contract, the models table, a Web-only "How to use it" (install + penguin web + Models page; no CLI commands), the roadmap, and the community call-to-action. + +### The launch post shows the finished RAG app with the condensed prompt + +The one-sentence build section now uses the condensed claude-code-docs configuration-expert prompt (Chinese in the zh post) and the finished-product screenshot of the generated docs-expert app (per-language, served from /blog-assets/), replacing the PenguinHarness chat capture. diff --git a/changelog/0.0.2/2026-07-20-skills.md b/changelog/0.0.2/2026-07-20-skills.md new file mode 100644 index 0000000..ac1a64d --- /dev/null +++ b/changelog/0.0.2/2026-07-20-skills.md @@ -0,0 +1,30 @@ +# Skill library and the AI-app skills + +## One-shot RAG app flow: skills rewrite and a draft-screen example task + +PR #11's squash — briefly lost to a force-push race during the branch rebase — is restored: the skills library rewrite that makes one-sentence RAG apps one-shot, a draft-screen example task card, and the finished-product showcase pipeline. + +- penguin-sdk rewritten around a complete RAG recipe (corpus collection, heading chunking, local BM25 retrieval, per-request Session SSE answers with citations, run-and-verify checklist); web-design gains the Penguin visual language and chat/RAG layout recipes; agent-creation covers skill bundles; agenthub-dev retires and a firecrawl skill joins. +- The Web draft screen gains example task cards (full prompts submitted as-is), with chat stream rendering refinements (markdown module, work summaries, stream follow) and colorful skill icons. +- The landing capture pipeline drives the docs-expert conversation and renders the finished-app mockup; README/Cases assets and prompts match the earlier finished-product switch. + +## Skill library regrouped by audience, with square card actions + +The library's six ad-hoc groups become four audience-oriented ones, and the card actions move into a tidy column. + +- New groups, in order: Office Productivity / 办公效率 (`data-analysis`, `firecrawl`), Software Development / 软件开发 (`web-design`, `software-engineering`), AI App Development / AI 应用开发 (`penguin-sdk`, `penguin-cli`, `agenthub-models`), and Agent Tuning / Agent 调优 (`agent-creation`, `benchmark-design`, `agent-evaluation`, `agent-optimization`). Docs tables, skills/server tests, and the skills e2e spec follow. +- Skill card actions (update nudge / quick invoke / manage installs) are now equal squares stacked in one column, vertically centered at the card's right edge; the metadata line moves under the header. + +## Skills: pin the CLI --root to the app dir and fast-stop when no key + +Three skills gain two hard rules for AI-app development: always target the app's own data root with the penguin CLI, and stop asking for help the moment no API key is usable. + +- penguin-sdk (v7), agenthub-models (v4), penguin-cli (v2): when building an AI app, `penguin config ...` must always pass `--root ` pointing at the app's data directory inside the current working directory (the same path given to `createAgent({ root })`, e.g. `./penguin_data`); running without `--root` writes to the global `~/.penguin/data`, which belongs to the person running Penguin, not the app. +- penguin-sdk (v7) and agenthub-models (v4): when no API key is usable, stop immediately and ask the user for help instead of looping on tool calls — re-running `env`, re-checking the vault, or retrying the build wastes turns and money; one clear check, then hand back to the user (open the agent's settings via the gear icon and add a key to the key vault). + +## System prompt exposes Provider and Model ID; skills state the key-first premise + +The Agent system prompt's Environment section now carries the session model reference, and the AI-app skills spell out the key-and-root prerequisites up front. + +- The `# Environment` section gains `Provider: {{PROVIDER}}` and `Model ID: {{MODEL_ID}}` placeholders (the session model's provider group and upstream id, filled from the resolved model entry) right before `Session ID`, and the fields are reordered to `Project Dir → Agent ID → CWD → Provider → Model ID → Session ID`. `assembleSystemPrompt` / `SessionEnvironmentValues` / `sessionEnvironment` thread the two new values through; the configuration docs' placeholder table (en/zh) follows. +- penguin-sdk (v8) and agenthub-models (v5) open with the prerequisite: for AI-app development, have the user add the model API key in **this agent's key vault** *before* building, and keep the app's Penguin data root **inside the CWD workspace** (`--root ./penguin_data`), never `~/.penguin`; model ids can come from the penguin CLI catalog. diff --git a/changelog/0.0.2/2026-07-20-tooling.md b/changelog/0.0.2/2026-07-20-tooling.md new file mode 100644 index 0000000..9d2ed87 --- /dev/null +++ b/changelog/0.0.2/2026-07-20-tooling.md @@ -0,0 +1,156 @@ +# Tooling, language, and infrastructure + +Repository language, dev commands, the changelog layout itself, and dependency upgrades. + +## Make English the repository working language + +Translated all residual non-i18n Chinese (comments, error/log messages, test titles and +fixtures, package metadata, e2e mock content) to English; Chinese remains only in i18n +catalogs, zh documents, and CJK-purpose test fixtures. + +## Details + +- Package metadata: `package.json` descriptions (root/server/web/landing) and the `link:cli` + fallback echo. +- Source comments: all remaining Chinese comments, including the SQL comment set in + `packages/server/src/db/schema.ts`. +- Hardcoded non-i18n user-facing strings (user-visible language change; error codes, + placeholders, and limits untouched): server HTTP/validation/409 messages, schedule TOML + validation errors, server logs, core SDK errors, CLI fallback errors, web provider + invariants. +- Tests: all Chinese describe/it titles and language-irrelevant fixtures translated; + assertions realigned to the new English source strings. +- e2e: mock LLM conversation content and cross-file branch markers translated consistently + (the substring relationships the mock's `includes()` dispatch relies on are preserved); + zh-locale UI assertions kept; `e2e/README.md` translated. +- Kept as required Chinese: zh i18n catalogs and fields (`strings.ts` dictionaries, CLI + `i18n.ts`, `titleZh`, `short_description_zh`, the zh language-name label, `*.zh.md` docs) and + test literals that assert zh i18n output or exercise CJK-specific behavior (display width, + CJK id validation, title-language consistency, zh-CN e2e UI assertions). +- Known gap: web `formatTaskStats` still hardcodes zh stat-line fragments; the proper fix is + routing them through the locale catalogs. + +PR: [#5](https://github.com/Prism-Shadow/penguin-harness/pull/5) (merged 2026-07-20) + +## Reorganize the changelog by release version + +The top level now holds only version folders; each version's README carries the +one-line-per-change index, and detail files live inside the folder (mirrors agenthub PR +#162). + +## Details + +- `changelog//README.md` is the release summary: one line per change with its + title and one-sentence summary, linking the detail file relatively; `0.0.2/README.md` + absorbs the section previously kept in the top-level index. +- `changelog/README.md` now documents only the layout and the entry conventions (no + per-file listing). +- The working rules (AGENTS.md, CONTRIBUTING.md) point index updates at the version + folder's README instead of the top-level file. + +## Serialize the dev prebuild to fix concurrent dev:server / dev:web clobbering + +dev:server and dev:web now share a lock-serialized, deduplicated prebuild of skills and core, +so launching both at the same time no longer corrupts dist/ (previously two tsup builds with +clean:true raced in the same output directories). + +## Details + +- New `scripts/dev-prebuild.mjs`: takes an exclusive on-disk lock (atomic mkdir under + `node_modules/`) around `pnpm --filter skills --filter core build`; concurrent invocations + wait instead of clobbering. Locks left by crashed runs are stolen when the holder PID is + dead or the lock is older than 10 minutes. +- A 5-second success stamp collapses duplicate builds: starting dev:server and dev:web + simultaneously builds once (the waiter skips), and dev:server's inner re-invocation is a + no-op. The window is deliberately tiny so an edit-then-restart cycle always rebuilds — + the "never start on stale deps" behavior is unchanged. +- Root `dev:server` now just delegates to the server package's `dev` script (which prebuilds + via the shared script), removing the historical double build of core; `dev:web` prebuilds + via the script and then starts Vite. +- `pnpm dev` run standalone from `packages/server` now also builds skills (it previously + built only core, even though the server imports skills at runtime). +- Verified: concurrent invocations wait/skip correctly and release the lock; a simultaneous + dev:server + dev:web start performs a single build with both servers coming up cleanly. + +## Add a combined pnpm dev and a dev:landing shortcut + +`pnpm dev` now starts the backend and the web app together with prefixed logs (workspace deps +built once via the shared prebuild lock), and `pnpm dev:landing` serves the landing page dev +server from the repo root. + +## Details + +- `pnpm dev` runs `concurrently -n server,web "pnpm dev:server" "pnpm dev:web"`. Merging the + two was previously unsafe: each command prebuilt skills/core with tsup `clean: true` into + the same dist/ directories and the parallel builds clobbered each other. The lock-serialized + prebuild (see the serialize-dev-prebuild entry) removed that race, and its success stamp + collapses the two prebuilds into a single build on a combined start. +- `pnpm dev:landing` delegates to the landing package's Vite dev server (port 7366, completing + the 7364/7365/7367 dev-port family). The landing package has no workspace deps, so no + prebuild is involved. +- `concurrently` added as a root devDependency; the dev-command lists in README.md and + README.zh.md now cover `dev` and `dev:landing`. +- Verified: a combined start performs exactly one skills+core build (the other prebuild + waits and skips), with the server on 127.0.0.1:7364 and Vite on localhost:7365; the + landing dev server responds on localhost:7366. Note Vite binds localhost (IPv6 ::1) — + use `localhost`, not `127.0.0.1`, when probing the Vite ports with curl. + +## Dev commands keep pnpm install current automatically + +Forgetting `pnpm install` before a dev command no longer breaks the start: every dev +command's prestep checks install freshness (lockfile-hash stamp) and runs `pnpm install` +itself when node_modules is missing or the lockfile changed. + +## Details + +- `scripts/dev-prebuild.mjs` gains an install-freshness step inside its existing lock: the + pnpm-lock.yaml content hash is stamped after a successful install, so the usual dev start + pays nothing; a fresh clone or a pulled lockfile change triggers `pnpm install` + automatically (concurrent dev commands still install/build exactly once). +- The lock and stamps move from `node_modules/` to the OS temp directory keyed by the repo + path — they must exist before the first install does (the old location crashed on a fresh + clone), and per-checkout isolation comes from the key. +- `dev:docs` / `dev:landing` have no workspace deps to build but still need current + installs: they now run the prestep with `--install-only`. + +## Upgrade AgentHub to 0.4.0 and adopt the opaque fidelity payload + +@prismshadow/agenthub 0.3.3 -> 0.4.0 (agenthub PR #159): content items replace the item-level +`signature`/`phase` fields with one opaque `fidelity` object, and OmniMessage now carries it +verbatim end to end — Trace, replay, and resume included; the `agenthub-dev` skill joins the +built-in library. + +## Details + +- OmniMessage complete payloads (text, thinking, inline_data, inline_thinking, tool_call) + replace `signature?: string` / `phase?: string | null` with `fidelity?: Record` — an opaque wire-fidelity payload written to the Trace as-is and passed back + verbatim on replay (Claude thinking signatures, GPT-5 encrypted reasoning `{id, + encrypted_content}` and `{phase}` markers, the OpenAI-compatible `{reasoning_field}` name). + Builders take the object directly; an empty object is treated as absent. +- GenerativeModel's streaming translator mirrors AgentHub baseClient aggregation: a thinking + block is closed by its fidelity payload and a run of equal fidelity is one block (the + OpenAI-compatible clients stamp every thinking delta with the same `{reasoning_field}`, + which must not split blocks — this carries agenthub's reasoning-field replay fix through + PenguinHarness so multi-turn conversations against strict OpenAI-compatible upstreams + survive); a text segment splits on a differing `fidelity.phase` and closes on + `fidelity.signature`, merging fidelity keys. Text-phase stickiness across segments is gone + (mirrors baseClient). +- The `agenthub-dev` skill (AgentHub's own model-support development workflow) is installed + into the built-in library under the Penguin Development group, completed to the library + contract (version/updated frontmatter, short descriptions, a "Before you start" section, + and a custom icon). +- Malformed-classification fix for 0.4.x: agenthub now surfaces truncated streamed tool-call + arguments as its own `ToolCallArgumentParseError` (previously a raw `SyntaxError`) and + thinking-only completions as `EmptyResponseError`; `isMalformedJsonParseError` recognizes + both (instanceof + name fallback + cause chain), so these still end as `malformed` and the + engine reconnects and retries instead of failing the turn (caught by the malformed e2e). +- Docs (omni-message, interfaces, sessions-and-traces; en + zh) and the design specs updated + to the fidelity semantics. +- Traces written before this change carried `signature`/`phase` on payloads; per the + pre-release no-migration policy they are not converted (old fields are ignored on replay — + resume of such Sessions loses provider fidelity; delete and recreate if needed). + +## The changelog folder merges related entries + +One file per change had grown to 35 fragment files for this release alone; related changes are now merged into six date-slug entry files (models-and-catalog, web-app, landing-site, blog-and-docs, readme, tooling), each keeping the standard entry shape — H1 title, one-sentence summary, then per-change details — with the version README listing one line per entry file. The layout convention itself is unchanged; the branch was also rebased onto main to linearize away the merge commit with the final tree kept byte-identical. diff --git a/changelog/0.0.2/2026-07-20-web-app.md b/changelog/0.0.2/2026-07-20-web-app.md new file mode 100644 index 0000000..54dce4a --- /dev/null +++ b/changelog/0.0.2/2026-07-20-web-app.md @@ -0,0 +1,99 @@ +# Web App + +Chat input, session titles, and the skill library in the Web App. + +## Chat input: positional slash, skill chips above the input, wider skill menu + +The `/` command menu now opens from any caret position (like `@` mentions) and running a +command removes just the token; selected skills display as chips above the input next to +the agent chip; chip remove buttons recolor on hover instead of washing a background; the +skill dropdown widens for readable descriptions while staying inside phone screens. + +## Details + +- Positional slash (`slash-token.ts`, pure + unit-tested): a `/` at the start of the text + or after whitespace opens the command menu from wherever the caret is (paths/URLs never + trigger it), mirroring the existing positional `@` mention matching; running a command + removes only the `start..end` token and keeps the rest of the text. Send-time `@` + semantics are unchanged — only a leading `@` hands off. +- Selected skills render as chips above the input in the same row as the `@` handoff chip + (skill icon + monospace name + remove x), appearing as you pick them from the dropdown; + the toolbar count badge stays in sync. +- Chip remove buttons (agent and skill) lose the hover background — the x recolors instead. +- The skill dropdown widens (26rem, clamped to the viewport) so descriptions stay readable + on desktop without overflowing phones. +- The models-page group speed button collapses to icon-only below the sm breakpoint (three + labeled actions don't fit a 390px header — caught by the layout e2e). + +## Start title generation after 1000 chars of body text + +Session titles start generating as soon as ~1000 characters of main-session body text have +streamed, instead of waiting for the whole Task to finish — long answers no longer overrun +the title material — and the title prompt now suppresses chain-of-thought. + +## Details + +- Server: the output relay fires the title generator mid-run once EARLY_TITLE_BODY_CHARS + (1000) of main-session complete body text have streamed (sub-session text doesn't count); + the generator self-guards (NULL title, single flight), and the Task-completion trigger + stays as the short-answer fallback. Covered by a gated mid-run test proving the early + fire happens while the run is still in flight. +- Core: the captured assistant-side title material is capped at 1000 chars (the user side + keeps 2000) — a title only needs the opening of the answer. +- The title prompt adds an explicit "answer immediately — do not think aloud or produce + chain-of-thought" rule and ends with an empty `` block, which many + reasoning models treat as an already-closed thinking phase (same trick as the model + probe), keeping the one-off request's budget on the title itself. + +## Skill library: update reminder, borderless groups, accent icons + +The skill library reminds you when an Agent's installed copy of a skill is older than the +library's version — an accent rotate button on the card updates every outdated Agent in one +click, and the manage-installs dialog marks outdated rows with an Update button. Group +sections lose their border, and card icons sit on a theme-accent background. + +## Details + +- The installed-skills snapshot now tracks each installed copy's `version` (the read API + already returns it — the installed SKILL.md frontmatter is the source). A pure + `outdatedAgentIds` helper (unit-tested) flags Agents whose copy is strictly older than + the library's; not-installed Agents and locally *newer* copies never trigger it. +- Card footer: when any Agent is outdated, an accent rotate button appears + ("有新版本:更新 N 个 Agent 的安装" / "Update available"); clicking reinstalls the current + library copy on every outdated Agent (install-again-is-update semantics) with one batch + success toast; partial failures keep the succeeded Agents and toast the first error. +- Manage-installs dialog: outdated rows show an accent "更新"/"Update" button next to + "已安装"/"Installed", updating just that Agent. +- Styling: group sections are borderless (header background alone carries the grouping); + card icons use the theme accent (`--accent-bg`/`--accent-fg`, following the theme-color + setting) instead of the gray bordered tile. + +## Titles strip skill-invocation markers; Agent cards show the installed skill count + +Two Web App additions: generated titles no longer leak the `` block, and Agent cards display how many Skills are installed. + +- Session titles strip machine markers before generation and in the fallback path: `stripConversationMarkers` (core, exported) removes `…` and the handoff / scheduled-task marker blocks from the material sent to the model, from the model output (via `sanitizeTitle`), and from the server's first-line fallback — so a skill invocation's marker can never become the title. Ordinary angle-bracket text (e.g. `
`) is left untouched. +- The Agent card's stats line gains an installed-skill count (open-book icon): the agents API/service compute `skillCount` from the `agent_state/skills//SKILL.md` directories, alongside the existing session / tool / vault-key / schedule counts. + +## Slash menu stays on screen; skill card buttons go horizontal and light + +Two Web App polish fixes: the slash command menu could grow past the top of the viewport, and the skill card actions change arrangement again. + +- The slash menu (which lists /compact plus every installed skill) now caps its height at min(20rem, 40vh) with internal scrolling, so its top edge never leaves the screen; the active row keeps itself scrolled into view for keyboard navigation. +- Skill card actions return to a single horizontal row (still equal squares, vertically centered at the card's right edge), and all three buttons now wear the light secondary background. + +## Password field polish, model-key visibility, and borderless skill cards + +Several small UX fixes across the password fields, the model key input, and the skill library cards, plus a redrawn penguin game shot. + +- The change-password dialog's current-password field now shows a hint naming the built-in admin's default initial password (penguin-2026), so a user who forgot it can still get in. +- The password show/hide toggle is removed from the tab order (tabIndex -1): Tab now moves between fields instead of landing on the reveal button. +- The model API key input gains the same show/hide toggle (it reuses PasswordInput), so a pasted key can be verified before saving. +- Skill library cards drop their border; hover now tints the whole card with a light gray background instead. +- The penguin sled game mockup is redrawn so the penguin stays a single connected cartoon shape (it had looked fragmented), and the game example is emphasized as 2D with a smoother, gentler difficulty ramp in the draft-screen card, its prompt, and the landing tab label. + +## The admin initial password becomes penguin-2026 + +`admin123` sits in every breach corpus, so Chrome flags the first login as a compromised password; the seeded admin now starts as `penguin-2026` — unflagged, brand-related, all lowercase plus a hyphen so it stays easy to type. + +The value swaps everywhere it appears: the server seed (`ADMIN_INITIAL_PASSWORD`), the release installer's first-login hint, READMEs, docs (quickstart / web-app / server-api), blog posts, landing quickstart copy, the screenshot capture scripts, and the e2e auth helper. The change-it-soon banner semantics are unchanged. diff --git a/changelog/0.0.2/2026-07-21-model-reference-and-review-fixes.md b/changelog/0.0.2/2026-07-21-model-reference-and-review-fixes.md new file mode 100644 index 0000000..c596379 --- /dev/null +++ b/changelog/0.0.2/2026-07-21-model-reference-and-review-fixes.md @@ -0,0 +1,17 @@ +# Explicit model reference pairs, and a batch of review fixes + +The provider group is no longer resolved automatically anywhere — a model is referenced by the complete `(provider, model_id)` pair or not at all — and a code review of the release branch turned up a further set of correctness, accessibility and documentation defects, all fixed here. + +- **The provider is never inferred, guessed, or defaulted (core, CLI, server, web).** Two mechanisms used to pick a group on the user's behalf, and both are gone. Catalog inference (`inferProviderForUpstream`) took the first bare-id match, and gateways resell vendor models under their upstream ids — `glm-5.2`, `qwen3.7-max`, `qwen3.7-plus` and `deepseek-v4-pro` each belong to two groups now — so `penguin config model add --model-id glm-5.2 --api-key ` wrote a Z.AI credential onto the Qwen Token Plan entry and sent it to Alibaba's endpoint (on the previous release the same command resolved to `zhipu`). Config resolution (`resolveModelRef`) separately matched a bare `model_id` against the Project's configured models whenever it happened to be unique. Both are deleted, along with `inferProviderForUpstream` and `providersForUpstream`. The rule everywhere is now **both halves or neither**: `--provider` is required on `config model add`; on `run` / `chat` the pair as a whole stays optional (omit both for the Project default) but supplying one half is an error; `addModel` and `Agent.createSession` require the pair at the type level; `POST /sessions` and the schedule routes return 400 for a half reference, and a schedule file persisted with only a `modelId` is reported invalid instead of silently binding to a guessed group. `run_subagent` gains a `provider` argument beside `model_id` and refuses half a reference, so a delegating model cannot land a subagent on the wrong vendor either. The trace viewer stops attributing cost to a first-match guess for pre-tag traces, showing no cost instead. CLI help, the CLI/models/server-api/quickstart/configuration/interfaces docs (en+zh), the READMEs and every skill that scripts the CLI were updated to pass the pair. +- **Escape no longer wipes the composer (Web App).** The slash menu's Escape handler cleared the whole input — correct back when the menu only opened on a draft starting with `/`, but the menu became positional this release, so dismissing it from a `/token` mid-draft destroyed everything typed so far, unrecoverably (the textarea is controlled, so undo could not bring it back). Escape now only dismisses the menu, mirroring `@` mentions via a `slashDismissed` offset. Separately, `removeSlashToken` collapsed *every* run of two or more spaces in the surviving draft rather than the join seam, silently reflowing pasted code, YAML, aligned tables and indentation; it now touches the seam alone. +- **The penguin.ooo installer forwards its arguments.** `https://penguin.ooo/install.sh` piped the real installer straight into `sh`, so `--universal` and `--version` were silently discarded — leaving unsupported-architecture users with advice ("re-run with `--universal`") that could not work through the documented URL, and turning version pinning into "latest". The forwarder now downloads to a temp file, runs it with `"$@"`, and propagates the exit code; this also removes the partial-execution hazard of piping a truncated download into a script that deletes the old install before moving the new one in. +- **Speed test reports throughput or nothing.** Speed mode raised the output cap to 64 tokens but reused the connectivity prompt, which demands a single word — so the rate was computed over a 1-3 token sample dominated by the closing usage chunk's round trip, and the same model graded green or yellow on 30ms of jitter. Speed mode now has its own prompt that keeps generating into the cap, and the new pure `probeTps` drops the rate below meaningful sample floors (16 tokens, 100ms), leaving TTFT alone rather than a fabricated number. +- **Announcement bar (landing).** Rotation pauses on keyboard focus as well as hover, so the focused link is no longer left inside an `aria-hidden` slide translated out of view (an ARIA violation and a WCAG 2.2.2 failure); under `prefers-reduced-motion` the carousel no longer strands itself on the clone slide. +- **Streaming fidelity rules are pinned by tests.** The two rules added with the AgentHub 0.4.0 payload — a `fidelity.signature` closes the open text segment, and fidelity keys merge rather than replace across deltas — had no coverage; both mutations passed the whole suite. Tests now fail on either, guarding against a replay handing the provider a signature covering text it never signed, and against GPT-5 phase segmentation being lost on resume. +- **New benchmark results, and a comparison that changed shape.** Both suites are refreshed, and the design of the comparison changed with them: instead of running all three harnesses on one shared DeepSeek V4 Pro, each now runs the model it is normally paired with — PenguinHarness on DeepSeek V4 Pro, Claude Code on Claude Opus 4.8, OpenAI Codex on GPT-5.5 — so the model column is part of the result rather than a constant, and Tokens/cost are suite totals rather than per-run means. Data analysis (15 tasks, single run): 66.67% for PenguinHarness against 53.33% for both rivals, at $0.55 versus $38.48 and $19.41. Coding (40 tasks × 2 runs, accuracy over all 80 outcomes): 71.25%, level with Codex and behind Claude Code's 86.25%, at $3.81 versus $220.08 and $146.97. The copy moved with the data — the section no longer claims "same model" or "highest accuracy of the three", both of which the new results contradict, and leads on cost-per-quality instead; the 35×/70×/58×/39× multiples in that copy are now asserted against the data by unit test so they cannot drift on the next refresh. The README and blog charts are generated from the same source by `render-benchmark-svg.mjs`. Bars stay linear and zero-based, but carry a minimum length: at true scale the cost panels put PenguinHarness at ~2px, which reads as an empty slot rather than as the winning series, so the floor keeps it visible while still plainly the shortest bar — with the exact figure printed beside every one. The README's chart caption is now a single-line conclusion rather than a paragraph of run settings, matching the RAG cost hook. +- **The demo videos reach the landing page.** The self-improvement section's "demo video coming soon" pill is replaced by the recording itself, and the RAG case tab plays the finished app instead of only picturing it — both following the READMEs, both locale-matched (`evo_zh`/`evo_en`, `rag_zh`/`rag_en`). The files stay out of this repo: at ~9 MB each against a ~17 MB history, committing four would triple what every contributor clones for assets only the marketing site shows, so they are served from the sibling `penguin-harness-community` repo via `raw.githubusercontent`, which answers with `accept-ranges: bytes` (seeking works) and `access-control-allow-origin: *`. Its `application/octet-stream` type does not block playback — `nosniff` is enforced for scripts and styles, not media — and this was verified end to end in Chromium rather than assumed. Every embed pairs a poster with `preload="none"`, so a page view downloads one ~37 KB still and no video byte until someone presses play (asserted by watching the network on load: zero mp4 requests). The standalone embed deliberately drops the BrowserFrame chrome the screenshots use, because the self-improvement recording is a narrated deck rather than the product's own UI; the RAG case keeps the frame, because there it really is the app. +- **The draft page is genuinely centred, and the upward menus stop clipping.** The brand block sat inside the upper flex spacer, so the gap above it was shorter than the gap below the example cards by exactly the brand's own height — the whole page rode high, and the slash menu, which opens upward from the composer, ran into the top of the scroll area. The brand now lives in the same centred block as the composer, pills and example tasks, making the two gaps identical (measured: 134/134 at 1000px tall, 84/84 at 900, 34/34 at 800). The slash and `@` menus additionally cap their height to the room actually measured between the composer and the nearest clipping ancestor instead of a fixed `40vh`, which could not know that distance: previously the menu was cut by ~50px at 900px tall and still 5px at 800px, and now it fits at every height down to 620px, shrinking and scrolling internally instead of being clipped. +- **Web polish.** Speed-test badges reset when the active Project changes (results are keyed by `(provider, model_id)`, so they used to leak across Projects); the work-group header ticks and settles on the same span, so its duration never jumps backwards; the markdown renderer's component map has a stable identity, so streaming no longer remounts closed code blocks and drops text selection or Copy state. +- **Sources are diffable again.** Four files embedded a literal NUL byte as a composite-key separator instead of the `\0` escape, which made git treat them as binary — `catalog-sync.ts`, the model-catalog and models tests, and the scheduler produced no reviewable diff, no `git grep` output, and no textual merge. The runtime value is unchanged. +- **Docs and metadata.** The workspace `files/content` preview mode and its response headers are documented in the server API reference (en/zh); the skills README lists the groups that actually exist; the skill-library page's docstrings describe the per-skill palette and two-column grid it really renders; `penguin-cli`'s `updated` date matches its version bump; the penguin-sdk RAG recipe stops installing the entire skill library into the app it generates; the README model table no longer lists Kimi K3 under Moonshot, which does not serve it; and a Chinese comment in the catalog's pricing block is translated. +- **A test that did not test.** The early-title case asserting that sub-session output does not count toward the 1000-character trigger stayed green with the origin guard removed; it now fails without it. diff --git a/changelog/0.0.2/README.md b/changelog/0.0.2/README.md new file mode 100644 index 0000000..cd6ebb9 --- /dev/null +++ b/changelog/0.0.2/README.md @@ -0,0 +1,17 @@ +# Version 0.0.2 + +Unreleased. + +- [2026-07-21] A model is always referenced by an explicit `(provider, model_id)` pair — the provider is never inferred, guessed, or defaulted — together with the fixes from a full review of the release branch, refreshed benchmark results, and the demo videos on the landing page. ([details](2026-07-21-model-reference-and-review-fixes.md)) + +- [2026-07-20] Model catalog: preset provider groups, catalog entries and ordering, the Models page features built around them, and where model credentials are allowed to live. ([details](2026-07-20-models-and-credentials.md)) + +- [2026-07-20] Web App: chat input and slash menu, session titles, the skill library cards, password fields, and the seeded admin password. ([details](2026-07-20-web-app.md)) + +- [2026-07-20] Landing site: the story it tells, section structure, animations, navigation, the install domain, and the case gallery. ([details](2026-07-20-landing-site.md)) + +- [2026-07-20] Skill library regrouped by audience, and the AI-app skills that turn one sentence into a running RAG app. ([details](2026-07-20-skills.md)) + +- [2026-07-20] The repository READMEs, the blog posts, and the docs site. ([details](2026-07-20-readme-blog-and-docs.md)) + +- [2026-07-20] Repository language, dev commands, the changelog layout itself, and dependency upgrades. ([details](2026-07-20-tooling.md)) diff --git a/changelog/README.md b/changelog/README.md new file mode 100644 index 0000000..f1fb32a --- /dev/null +++ b/changelog/README.md @@ -0,0 +1,10 @@ +# Changelog Details + +The root [../CHANGELOG.md](../CHANGELOG.md) keeps one brief line per release. Each release owns a folder here: + +- `/README.md` — the release summary: one brief line per change, each linking its detail file, e.g. `- [YYYY-MM-DD] Brief description. ([details](YYYY-MM-DD-short-slug.md))` +- `/YYYY-MM-DD-short-slug.md` — one detail file per entry, named by the entry date: an H1 title, then what changed and why, using `##` sections once the entry covers more than one thing. +- Entries are grouped by the surface they change — Models, Web App, landing site, skills, docs, tooling — rather than one file per commit. A related change extends the existing file and its summary line instead of opening a new one. +- Changes not yet released go into the upcoming version's folder; during release preparation, rename the folder if the number changed and add the release line to the root file. Released folders are frozen. + +Written in English. History starts after the v0.0.1 release (2026-07-19); earlier changes are not backfilled. diff --git a/install.sh b/install.sh index 33400b8..56aa9be 100644 --- a/install.sh +++ b/install.sh @@ -149,5 +149,5 @@ echo "PenguinHarness $installed_version installed to $INSTALL_DIR" echo "" echo "Get started:" echo " penguin --help # all commands" -echo " penguin web # start the Web UI at http://127.0.0.1:7364 (initial login: admin / admin123)" +echo " penguin web # start the Web UI at http://127.0.0.1:7364 (initial login: admin / penguin-2026)" echo " penguin server # headless server (PORT / HOST to override)" diff --git a/package.json b/package.json index 3595e42..f4961cd 100644 --- a/package.json +++ b/package.json @@ -16,15 +16,18 @@ "build": "pnpm -r build && pnpm link:cli", "link:cli": "pnpm --dir packages/cli link --global || echo '[link:cli] pnpm global directory not configured; skipping the global penguin link (run pnpm setup once first)'", "penguin": "tsx packages/cli/src/index.ts", - "dev:server": "pnpm --filter @prismshadow/penguin-skills --filter @prismshadow/penguin-core build && pnpm --filter @prismshadow/penguin-server dev", - "dev:web": "pnpm --filter @prismshadow/penguin-skills --filter @prismshadow/penguin-core build && pnpm --filter @prismshadow/penguin-web dev", - "dev:docs": "pnpm --filter @prismshadow/penguin-docs dev", + "dev": "concurrently -n server,web -c cyan,magenta \"pnpm dev:server\" \"pnpm dev:web\"", + "dev:server": "pnpm --filter @prismshadow/penguin-server dev", + "dev:web": "node scripts/dev-prebuild.mjs && pnpm --filter @prismshadow/penguin-web dev", + "dev:docs": "node scripts/dev-prebuild.mjs --install-only && pnpm --filter @prismshadow/penguin-docs dev", + "dev:landing": "node scripts/dev-prebuild.mjs --install-only && pnpm --filter @prismshadow/penguin-landing dev", "build:site": "node scripts/build-site.mjs" }, "devDependencies": { "@prismshadow/penguin-cli": "workspace:*", "@prismshadow/penguin-core": "workspace:*", "@types/node": "^24.0.0", + "concurrently": "^10.0.3", "prettier": "^3.9.4", "tsx": "^4.20.0", "typescript": "^5.6.0", diff --git a/packages/cli/package.json b/packages/cli/package.json index 9db412e..401b8dc 100644 --- a/packages/cli/package.json +++ b/packages/cli/package.json @@ -22,7 +22,7 @@ "penguin": "tsx src/index.ts" }, "dependencies": { - "@prismshadow/agenthub": "^0.3.3", + "@prismshadow/agenthub": "^0.4.0", "@prismshadow/penguin-core": "workspace:*", "@prismshadow/penguin-server": "workspace:*", "@prismshadow/penguin-skills": "workspace:*", diff --git a/packages/cli/src/commands/chat.ts b/packages/cli/src/commands/chat.ts index 5873ae6..cae59a3 100644 --- a/packages/cli/src/commands/chat.ts +++ b/packages/cli/src/commands/chat.ts @@ -1,12 +1,14 @@ /** * `penguin chat` — interactive REPL. * - * penguin chat [--model-id ] [--provider ] [--project-id ] [--agent-id ] + * penguin chat [--model-id --provider ] [--project-id ] [--agent-id ] * [--workspace ] [--approve ] * * Each line of input starts one conversation turn; `/compact` proactively compacts the * context (reason=manual); `/exit` or `/quit` exits. - * Uses the current directory when no Workspace is specified. + * Uses the current directory when no Workspace is specified. A model reference is always an + * explicit `(provider, model_id)` pair, so `--model-id` and `--provider` must be given + * together; giving neither uses the Project's default model. * * Multi-line input: trailing `\` continues the line; when the terminal supports bracketed * paste, a multi-line paste is treated as a single message (sent on Enter). @@ -65,6 +67,17 @@ export function registerChatCommand(program: Command, t: Messages): void { .option("--approve ", t.common.approve) .option("--resume [sessionId]", t.chat.resume) .action(async (opts) => { + // The model reference is a pair: commander can only require each option on its own, + // so the "both or neither" rule is enforced here. Giving neither is the normal case + // and falls back to the Project's default model. Skipped under --resume, which + // rejects both options outright further down with a more specific message. + // Usage errors go to stderr (as in `run` and `config model add`), unlike this file's + // informational messages, which the REPL writes to stdout. + if (opts.resume === undefined && Boolean(opts.modelId) !== Boolean(opts.provider)) { + process.stderr.write(`${t.error(t.modelRefIncomplete())}\n`); + process.exitCode = 1; + return; + } const mode = resolveApprovalMode(opts.approve, t); const out = process.stdout; diff --git a/packages/cli/src/commands/config.ts b/packages/cli/src/commands/config.ts index 187cc78..64394bf 100644 --- a/packages/cli/src/commands/config.ts +++ b/packages/cli/src/commands/config.ts @@ -2,7 +2,7 @@ * `penguin config` — manages a Project's model credentials, default model, model list, * Agent-level vault environment variables, and UI language. * - * penguin config model add --model-id [--provider ] [--api-key ] [--context-window ] [--set-default] [--root ] + * penguin config model add --model-id --provider [--api-key ] [--context-window ] [--set-default] [--root ] * penguin config model default --model-id --provider [--root ] * penguin config model vision --model-id --provider [--root ] * penguin config model list [--root ] @@ -13,12 +13,12 @@ * * `--model-id` always takes the **upstream id** (the request id sent to AgentHub verbatim), * which together with `--provider` forms a `(provider, model_id)` paired reference — - * **no string concatenation is ever performed**. For `model add`, --provider defaults to - * an inference from the built-in catalog (falling back to custom when inference fails); - * a new entry's client_type defaults according to the group's semantics (not set for - * first-party vendors; openai for custom / self-hosted groups / gateways, with the - * gateway's endpoint base URL pre-filled). For `model default` / `model vision`, - * --provider is **required**; core validation raises an error when the reference is not + * **no string concatenation is ever performed**. `--provider` is **required** on all three + * model subcommands: the group is never guessed, so `--api-key` can never land on a vendor + * the user did not name. For `model add`, a new entry's client_type defaults according to + * the group's semantics (not set for first-party vendors; openai for custom / self-hosted + * groups / gateways, with the gateway's endpoint base URL pre-filled). For `model default` + * / `model vision`, core validation raises an error when the reference is not * found in models. `--root` specifies the data root directory (priority: option > * PENGUIN_HOME > ~/.penguin/data). The UI language is controlled by the PENGUIN_LANG * environment variable; `config lang` writes it into the shell startup file and restarts @@ -39,7 +39,6 @@ import { catalogEntryFor, formatModelRef, getModel, - inferProviderForUpstream, loadAgentVault, loadProjectConfig, providerInfo, @@ -117,7 +116,7 @@ export function registerConfigCommand(program: Command, t: Messages): void { .command("add") .description(t.config.addDesc) .requiredOption("--model-id ", t.config.addModelId) - .option("--provider ", t.config.addProvider) + .requiredOption("--provider ", t.config.addProvider) .option("--api-key ", t.config.addApiKey) .option("--base-url ", t.config.addBaseUrl) .option("--context-window ", t.config.addContextWindow, parseIntArg) @@ -133,11 +132,11 @@ export function registerConfigCommand(program: Command, t: Messages): void { .option("--root ", t.common.root) .action(async (opts) => { const root = resolveRootOption(opts.root); - // --model-id takes the upstream id, paired with --provider as a reference - // (--provider defaults to catalog-based inference, falling back to custom); no - // concatenation is performed. + // --model-id takes the upstream id, paired with the required --provider as a + // reference; the group is never guessed, so --api-key can only ever land on the + // vendor the user named. No concatenation is performed. const modelId: string = opts.modelId; - const provider: string = opts.provider ?? inferProviderForUpstream(modelId); + const provider: string = opts.provider; const ref: ModelRef = { provider, model_id: modelId }; const before = await loadProjectConfig(root, opts.projectId); const existed = getModel(before, ref) !== undefined; diff --git a/packages/cli/src/commands/run.ts b/packages/cli/src/commands/run.ts index 1a8f9d9..5617aee 100644 --- a/packages/cli/src/commands/run.ts +++ b/packages/cli/src/commands/run.ts @@ -1,14 +1,14 @@ /** * `penguin run` — send a single Task in one shot. * - * penguin run -m [--model-id ] [--provider ] [--workspace ] + * penguin run -m [--model-id --provider ] [--workspace ] * [--project-id ] [--agent-id ] * [--approve ] * * Uses the current directory when Workspace is unspecified; uses the Project's default model - * when model is unspecified. `--provider` is optional: when omitted, `--model-id` is resolved - * via resolveModelRef semantics (only matches when the exact value is globally unique in the - * config; ambiguity is an error). Defaults to interactive per-call approval; `--approve` + * when model is unspecified. A model reference is always an explicit `(provider, model_id)` + * pair, so `--model-id` and `--provider` must be given together — giving only one of them is + * an error, never a lookup. Defaults to interactive per-call approval; `--approve` * selects the permission mode. * Docs: /docs/cli § "penguin run". */ @@ -31,6 +31,14 @@ export function registerRunCommand(program: Command, t: Messages): void { .option("--workspace ", t.common.workspace) .option("--approve ", t.common.approve) .action(async (opts) => { + // The model reference is a pair: commander can only require each option on its own, + // so the "both or neither" rule is enforced here. Giving neither is the normal case + // and falls back to the Project's default model. + if (Boolean(opts.modelId) !== Boolean(opts.provider)) { + process.stderr.write(`${t.error(t.modelRefIncomplete())}\n`); + process.exitCode = 1; + return; + } const mode = resolveApprovalMode(opts.approve, t); const agent = await createAgent({ diff --git a/packages/cli/src/i18n.ts b/packages/cli/src/i18n.ts index c4dd048..094b6f8 100644 --- a/packages/cli/src/i18n.ts +++ b/packages/cli/src/i18n.ts @@ -23,7 +23,7 @@ export interface Messages { projectId: string; agentId: string; modelId: string; - /** run/chat's --provider: pairs with --model-id; when omitted, resolved by unique match (ambiguity is an error). */ + /** run/chat's --provider: must be given together with --model-id (the group is never inferred). */ provider: string; /** Data root directory option (priority: --root > PENGUIN_HOME > ~/.penguin/data). */ root: string; @@ -113,6 +113,8 @@ export interface Messages { approveModeInvalid(value: string): string; /** Render label for an approval decision (frontend renders the approval_decision event; one label each for allow/deny). */ approvalDecision(decision: "allow" | "deny"): string; + /** run/chat given only one of --model-id / --provider: a model reference is always an explicit pair, never a lookup. */ + modelRefIncomplete(): string; /** --resume is mutually exclusive with --workspace/--model-id (neither can change once the Session is created). */ resumeNoOverride(): string; /** --resume given without a session id, and the current Agent has no Session at all. */ @@ -156,7 +158,7 @@ const en: Messages = { agentId: "Agent id", modelId: "Model to use (upstream model id; defaults to the Project default model)", provider: - "Provider of --model-id; when omitted, the model id must match exactly one configured entry (ambiguity is an error)", + "Provider group of --model-id; required whenever --model-id is given (the group is never inferred)", root: "Data root directory (overrides PENGUIN_HOME and ~/.penguin/data)", workspace: "Workspace directory; must already exist (defaults to the current directory)", approve: @@ -168,11 +170,11 @@ const en: Messages = { addDesc: "Add or update a model, optionally writing a credential", addModelId: "Upstream model id sent to AgentHub as-is (e.g. claude-sonnet-4-6)", addProvider: - "Provider group stored alongside model_id; inferred from the builtin catalog when omitted, else custom", + "Provider group stored alongside model_id; required, never inferred (use custom for anything without a vendor group)", addApiKey: "API key, stored inline in the Project's hidden .project_config.toml", addBaseUrl: "Custom base URL", addContextWindow: "Context window size (tokens)", - addClientType: "AgentHub client type (e.g. openai); inferred from model id when omitted", + addClientType: "AgentHub client type (e.g. openai); defaults by provider group when omitted", addVision: "Mark the model as supporting image input (vision)", addNoVision: "Mark the model as NOT supporting image input; omit both to keep current", addPriceCacheRead: "Price per 1M tokens: cache read (USD)", @@ -235,6 +237,8 @@ const en: Messages = { approveModeInvalid: (value) => `Invalid approval mode "${value}". Use allow-all, deny-all, read-only, or always-ask.`, approvalDecision: (decision) => (decision === "allow" ? "✓ [approved]" : "× [denied]"), + modelRefIncomplete: () => + "--model-id and --provider must be given together: a model reference is always an explicit (provider, model_id) pair. Omit both to use the Project default model.", resumeNoOverride: () => "--resume does not accept --workspace, --model-id or --provider: they follow the original Session and cannot change.", resumeNoSession: () => "No session to resume: this agent has no recorded sessions yet.", @@ -268,7 +272,7 @@ const zh: Messages = { projectId: "Project id", agentId: "Agent id", modelId: "本次使用的模型(上游模型 id;默认 Project 默认模型)", - provider: "--model-id 的 provider 分组;省略时 model id 须在配置中精确唯一命中(歧义报错)", + provider: "--model-id 的 provider 分组;给出 --model-id 时必须一并给出(分组不作任何推断)", root: "数据根目录(优先于 PENGUIN_HOME 与 ~/.penguin/data)", workspace: "Workspace 目录,须为已存在目录(默认当前目录)", approve: @@ -279,11 +283,11 @@ const zh: Messages = { modelDesc: "管理模型 credential 与默认模型", addDesc: "新增或更新一个模型,并可写入 credential", addModelId: "上游模型 id(如 claude-sonnet-4-6,原样发给 AgentHub)", - addProvider: "与 model_id 分列存储的 provider 分组;缺省按内置目录推断,推断不出为 custom", + addProvider: "与 model_id 分列存储的 provider 分组;必填,不作推断(无厂商分组时填 custom)", addApiKey: "API key,内联存入 Project 的隐藏文件 .project_config.toml", addBaseUrl: "自定义 base url", addContextWindow: "上下文窗口大小(token 数)", - addClientType: "AgentHub 客户端协议(如 openai);缺省由 model id 推断", + addClientType: "AgentHub 客户端协议(如 openai);缺省按 provider 分组的语义取值", addVision: "标注该模型支持图片输入(视觉)", addNoVision: "标注该模型不支持图片输入;两者都不给则保留原值", addPriceCacheRead: "每百万 token 价格:缓存读取(USD)", @@ -345,6 +349,8 @@ const zh: Messages = { approveModeInvalid: (value) => `无效的审批模式 "${value}"。请使用 allow-all、deny-all、read-only 或 always-ask。`, approvalDecision: (decision) => (decision === "allow" ? "✓ [已批准]" : "× [已拒绝]"), + modelRefIncomplete: () => + "--model-id 与 --provider 必须成对给出:模型引用始终是显式的 (provider, model_id) 组合。两者都不给则使用 Project 默认模型。", resumeNoOverride: () => "--resume 不接受 --workspace、--model-id 与 --provider:均沿用原 Session,创建后不可更换。", resumeNoSession: () => "没有可恢复的 Session:当前 Agent 还没有任何会话记录。", diff --git a/packages/cli/test/config-model.test.ts b/packages/cli/test/config-model.test.ts index 0f084c7..b3f13ee 100644 --- a/packages/cli/test/config-model.test.ts +++ b/packages/cli/test/config-model.test.ts @@ -1,10 +1,10 @@ /** * Integration tests for `penguin config model add|default|vision|list` (run through * commander's parseAsync for the full command path): --model-id always takes the - * upstream id, paired with --provider to form a (provider, model_id) reference (add's - * --provider defaults to catalog-based inference, falling back to custom when - * inference fails; default / vision require --provider and raise an error when the - * reference isn't found in models — no string concatenation is ever performed); --root + * upstream id, paired with --provider to form a (provider, model_id) reference + * (--provider is required on all three subcommands — the group is never inferred — and + * default / vision raise an error when the reference isn't found in models; no string + * concatenation is ever performed); --root * specifies the data root directory (takes priority over PENGUIN_HOME); persisted to a * single hidden .project_config.toml (mode 0600, credentials inline, provider and * model_id as separate columns); list displays provider and model_id as separate @@ -84,13 +84,15 @@ describe("penguin config model add/list (--root plus provider / model_id stored "add", "--model-id", "my-own-model", + "--provider", + "custom", "--api-key", "sk-root-secret-1", "--root", tmpRoot, ]); expect(add.code).toBe(0); - // Catalog inference fails -> falls back to the custom group (provider is a separate field, never concatenated into the id). + // The named group is stored as a separate field, never concatenated into the id. expect(add.out).toContain("Added model (provider=custom, model_id=my-own-model)."); const file = projectConfigPath(tmpRoot, DEFAULT_PROJECT_ID); @@ -119,11 +121,13 @@ describe("penguin config model add/list (--root plus provider / model_id stored expect(list.out).not.toContain("sk-root-secret-1"); }); - it("built-in catalog infers the grouping: an upstream id matching the catalog lands under that provider; --set-default writes a pair reference", async () => { + it("naming an existing pair updates that preset entry in place; --set-default writes a pair reference", async () => { const add = await runModel([ "add", "--model-id", "claude-sonnet-4-6", + "--provider", + "anthropic", "--set-default", "--root", tmpRoot, @@ -168,8 +172,16 @@ describe("penguin config model add/list (--root plus provider / model_id stored }); it("client_type defaults by grouping semantics (PRN-021): custom / self-hosted / gateway get openai, first-party providers get none", async () => { - // custom (catalog inference fails) and self-hosted groups (--provider not a catalog value): default to client_type=openai. - await runModel(["add", "--model-id", "my-openai-proxy", "--root", tmpRoot]); + // The custom group and self-hosted groups (--provider not a catalog value): default to client_type=openai. + await runModel([ + "add", + "--model-id", + "my-openai-proxy", + "--provider", + "custom", + "--root", + tmpRoot, + ]); await runModel(["add", "--model-id", "in-house-1", "--provider", "mylab", "--root", tmpRoot]); // A non-catalog id under a first-party vendor group: client_type is not set (AgentHub auto-routes by upstream id). await runModel([ @@ -241,13 +253,29 @@ describe("penguin config model add/list (--root plus provider / model_id stored }); }); -describe("model default/vision: --provider is required, (provider, model_id) pair reference", () => { +describe("model add/default/vision: --provider is required, (provider, model_id) pair reference", () => { it("missing --provider: commander usage error, nonzero exit code", async () => { const bad = await runModel(["default", "--model-id", "deepseek-v4-flash", "--root", tmpRoot]); expect(bad.code).not.toBe(0); expect(bad.err).toContain("--provider"); }); + it("add without --provider is a usage error too: the group is never inferred, so no config is written", async () => { + const bad = await runModel([ + "add", + "--model-id", + "claude-sonnet-4-6", + "--api-key", + "sk-never-stored", + "--root", + tmpRoot, + ]); + expect(bad.code).not.toBe(0); + expect(bad.err).toContain("--provider"); + // The credential must not have landed on a guessed vendor: nothing was persisted at all. + await expect(fs.access(projectConfigPath(tmpRoot, DEFAULT_PROJECT_ID))).rejects.toThrow(); + }); + it("dangling reference: the pair is not in models; the error carries the pair reference and a model list hint", async () => { const bad = await runModel([ "default", diff --git a/packages/cli/test/model-ref-pairing.test.ts b/packages/cli/test/model-ref-pairing.test.ts new file mode 100644 index 0000000..809a9a8 --- /dev/null +++ b/packages/cli/test/model-ref-pairing.test.ts @@ -0,0 +1,106 @@ +/** + * `penguin run` / `penguin chat`: a model reference is always an explicit + * (provider, model_id) pair. Commander can only mark each option required on its own, so + * the "both or neither" rule is enforced inside the action — supplying exactly one of + * --model-id / --provider is a usage error (never a lookup against the configured models), + * while supplying neither falls back to the Project's default model. + */ +import fs from "node:fs/promises"; +import os from "node:os"; +import path from "node:path"; +import { afterEach, beforeEach, describe, expect, it, vi } from "vitest"; +import { Command } from "commander"; +import { registerChatCommand } from "../src/commands/chat.js"; +import { registerRunCommand } from "../src/commands/run.js"; +import { getMessages } from "../src/i18n.js"; + +// The --resume case below reaches createAgent, which initializes Agent state on disk; point +// the data root at a throwaway directory so no test ever writes to the real ~/.penguin/data. +let tmpHome: string; +let prevHome: string | undefined; + +beforeEach(async () => { + prevHome = process.env.PENGUIN_HOME; + tmpHome = await fs.mkdtemp(path.join(os.tmpdir(), "penguin-cli-pairing-")); + process.env.PENGUIN_HOME = tmpHome; +}); + +afterEach(async () => { + if (prevHome === undefined) delete process.env.PENGUIN_HOME; + else process.env.PENGUIN_HOME = prevHome; + await fs.rm(tmpHome, { recursive: true, force: true }); +}); + +/** Runs one command, capturing stdout / stderr and the exit code (without exiting the process). */ +async function runCommand( + register: (program: Command, t: ReturnType) => void, + args: string[], +): Promise<{ out: string; err: string; code: number }> { + const program = new Command(); + program.exitOverride(); + register(program, getMessages("en")); + const out: string[] = []; + const err: string[] = []; + const outSpy = vi.spyOn(process.stdout, "write").mockImplementation((chunk) => { + out.push(String(chunk)); + return true; + }); + const errSpy = vi.spyOn(process.stderr, "write").mockImplementation((chunk) => { + err.push(String(chunk)); + return true; + }); + const prevExitCode = process.exitCode; + process.exitCode = undefined; + try { + await program.parseAsync(["node", "penguin", ...args]); + return { out: out.join(""), err: err.join(""), code: Number(process.exitCode ?? 0) }; + } catch (e) { + const exitCode = (e as { exitCode?: number }).exitCode; + return { out: out.join(""), err: err.join(""), code: exitCode || 1 }; + } finally { + outSpy.mockRestore(); + errSpy.mockRestore(); + process.exitCode = prevExitCode; + } +} + +describe("run: --model-id and --provider must be given together", () => { + it("--model-id without --provider: error on stderr, exit code 1", async () => { + const bad = await runCommand(registerRunCommand, [ + "run", + "-m", + "hi", + "--model-id", + "deepseek-v4-flash", + ]); + expect(bad.code).toBe(1); + expect(bad.err).toContain("--model-id and --provider must be given together"); + }); + + it("--provider without --model-id: same error (the pair is never half-specified)", async () => { + const bad = await runCommand(registerRunCommand, ["run", "-m", "hi", "--provider", "deepseek"]); + expect(bad.code).toBe(1); + expect(bad.err).toContain("--model-id and --provider must be given together"); + }); +}); + +describe("chat: --model-id and --provider must be given together", () => { + it("--model-id without --provider: error on stderr, exit code 1", async () => { + const bad = await runCommand(registerChatCommand, ["chat", "--model-id", "deepseek-v4-flash"]); + expect(bad.code).toBe(1); + expect(bad.err).toContain("--model-id and --provider must be given together"); + }); + + it("--resume plus a lone --model-id keeps the more specific resume error", async () => { + const bad = await runCommand(registerChatCommand, [ + "chat", + "--resume", + "sess-1", + "--model-id", + "deepseek-v4-flash", + ]); + expect(bad.code).toBe(1); + expect(bad.out).toContain("--resume does not accept"); + expect(`${bad.out}${bad.err}`).not.toContain("must be given together"); + }); +}); diff --git a/packages/core/package.json b/packages/core/package.json index 593540e..abd4c28 100644 --- a/packages/core/package.json +++ b/packages/core/package.json @@ -36,7 +36,7 @@ "build": "tsup" }, "dependencies": { - "@prismshadow/agenthub": "^0.3.3", + "@prismshadow/agenthub": "^0.4.0", "@prismshadow/penguin-skills": "workspace:*", "smol-toml": "^1.3.0", "yaml": "^2.5.0" diff --git a/packages/core/src/agent.ts b/packages/core/src/agent.ts index 5613d64..63db090 100644 --- a/packages/core/src/agent.ts +++ b/packages/core/src/agent.ts @@ -79,12 +79,14 @@ export interface CreateAgentOptions { export interface CreateSessionOptions { /** Workspace for this run; if unspecified, a temporary Workspace is created under the Agent directory. */ workspaceDir?: string; - /** Model used for this Session (upstream model_id); if unspecified, uses the Project's default Model. */ + /** + * Model used for this Session (upstream model_id); must be given together with `provider`. + * Omit both to use the Project's default Model. + */ modelId?: string; /** - * Provider grouping for `modelId` (a paired reference); if omitted, resolved via - * `resolveModelRef` semantics — `model_id` only resolves if it is a globally unique - * exact match in the config; zero or multiple matches produce a clear error. + * Provider grouping for `modelId`: the two together form the paired reference. It is never + * inferred, so it's "both or neither" — either half alone is an error, not a lookup. */ provider?: string; /** Explicit credentials; if unspecified, falls back to credentials in the Project config, then to AgentHub reading environment variables. */ @@ -134,18 +136,20 @@ export class Agent { */ async createSession(opts: CreateSessionOptions = {}): Promise { // Model is validated first (before creating the Workspace, so failure leaves no - // temp directory behind): the reference must resolve to an entry in the Project - // config (the (provider, model_id) pair is the unique key); a reference - // outside the config throws immediately rather than passing silently — otherwise - // credentials, pricing, and the context window would all be unavailable. - if (opts.modelId === undefined && opts.provider !== undefined) { + // temp directory behind): the reference must be the complete (provider, model_id) + // pair — the config's unique key — and must name an entry in the Project config; a + // reference outside the config throws immediately rather than passing silently, + // otherwise credentials, pricing, and the context window would all be unavailable. + // Half a reference is always an error: the missing half is never inferred, since a + // guessed provider would send the entry's credential to a vendor nobody named. + if ((opts.modelId === undefined) !== (opts.provider === undefined)) { throw new Error( - "provider was specified without modelId: a model reference must be given as a pair (provider cannot be used alone).", + "A model reference must be given as a (provider, model_id) pair: both must be specified, or neither (to use the Project's default model).", ); } let ref: ModelRef; - if (opts.modelId !== undefined) { - // The only entry point for resolving an "omitted provider" reference (resolveModelRef): three branches — unique match / zero matches / ambiguous. + if (opts.modelId !== undefined && opts.provider !== undefined) { + // Pair validation only (resolveModelRef): the pair either names a configured entry or errors. ref = resolveModelRef(this.projectConfig, opts.modelId, opts.provider); } else if (this.projectConfig.default_model) { ref = this.projectConfig.default_model; @@ -211,6 +215,8 @@ export class Agent { sessionEnvironment(workspaceDir, sessionId, { agentId: this.state.agentId, projectDir: projectDir(this.state.root, this.state.projectId), + provider: modelEntry.provider, + modelId: modelEntry.model_id, }), Object.keys(vault), installedSkills, @@ -445,9 +451,12 @@ export class Agent { const { workspaceDir, modelEntry, apiKey, baseUrl, systemPrompt, subagentDepth, vault } = args; // Child-Agent runner: injected into the run_subagent tool so it doesn't need to // depend on Agent/Session (breaking a circular dependency). The model can - // optionally choose agentId (omitted = call the current Agent) and modelId - // (omitted = Project default). Precheck errors (depth limit exceeded / agent - // doesn't exist) are expressed as throws, which the Environment collapses to failed. + // optionally choose agentId (omitted = call the current Agent) and the child + // Session's model (omitted = Project default); the model reference is forwarded to + // createSession as-is and must be a complete (provider, model_id) pair — half a + // reference is rejected there rather than being guessed here. Precheck errors (depth + // limit exceeded / agent doesn't exist) are expressed as throws, which the + // Environment collapses to failed. // Docs: /docs/interfaces § "Subagent interfaces" const parentAgent = this; const { root, projectId, agentId: parentAgentId } = this.state; @@ -455,7 +464,7 @@ export class Agent { // Spawn and run are separate: the same child Session can run for multiple turns // (continuing via input_subagent appending a prompt); resource cleanup is // consolidated in handle.dispose (called by the managing ManagedSubagentSession). - async spawn({ agentId, modelId }) { + async spawn({ agentId, modelId, provider }) { if (subagentDepth >= MAX_SUBAGENT_DEPTH) { throw new Error( `subagent depth limit ${MAX_SUBAGENT_DEPTH} reached; not spawning another subagent`, @@ -475,9 +484,12 @@ export class Agent { agentId !== undefined && agentId !== parentAgentId ? await createAgent({ root, projectId, agentId }) : parentAgent; + // The model reference is forwarded as a whole: createSession rejects half a pair, so a + // caller that named only one side gets the same error the CLI and HTTP layers give. const childSession = await childAgent.createSession({ workspaceDir, ...(modelId !== undefined ? { modelId } : {}), + ...(provider !== undefined ? { provider } : {}), subagentDepth: subagentDepth + 1, }); // All child-session messages are tagged with an origin (the child Session id, diff --git a/packages/core/src/environment/tools/run-subagent.ts b/packages/core/src/environment/tools/run-subagent.ts index 9c27cf9..328c296 100644 --- a/packages/core/src/environment/tools/run-subagent.ts +++ b/packages/core/src/environment/tools/run-subagent.ts @@ -82,6 +82,15 @@ export function createSubagentTool( } const agentId = typeof args.agent_id === "string" ? args.agent_id : undefined; const modelId = typeof args.model_id === "string" ? args.model_id : undefined; + const provider = typeof args.provider === "string" ? args.provider : undefined; + // A model is referenced by the complete (provider, model_id) pair — never half of one. + // Caught here rather than in createSession so the model is told which half it left out. + if ((modelId === undefined) !== (provider === undefined)) { + yield* fail( + "[run_subagent error: `model_id` and `provider` must be given together (a model reference is the pair), or both omitted to use the Project default model]", + ); + return { stopReason: "failed" }; + } const yieldMs = clampYield( args.yield_time_ms, DEFAULT_SUBAGENT_YIELD_MS, @@ -106,6 +115,7 @@ export function createSubagentTool( const handle = await runner.spawn({ ...(agentId !== undefined ? { agentId } : {}), ...(modelId !== undefined ? { modelId } : {}), + ...(provider !== undefined ? { provider } : {}), }); session = new ManagedSubagentSession(handle); } catch (err) { diff --git a/packages/core/src/index.ts b/packages/core/src/index.ts index 28f40d4..fdbd62f 100644 --- a/packages/core/src/index.ts +++ b/packages/core/src/index.ts @@ -9,7 +9,8 @@ * * ```ts * const agent = await createAgent({ agentId: "default_agent" }); - * const session = await agent.createSession({ workspaceDir, modelId }); + * // A model reference is always the (provider, model_id) pair; omit both for the Project default. + * const session = await agent.createSession({ workspaceDir, provider, modelId }); * for await (const output of session.run([userText("...")])) { ... } * ``` */ @@ -36,7 +37,12 @@ export type { } from "./engine/context-engine.js"; export { Session } from "./session.js"; export type { SessionConfig } from "./session.js"; -export { buildTitlePrompt, generateTitleWithLLM, sanitizeTitle } from "./session-title.js"; +export { + buildTitlePrompt, + generateTitleWithLLM, + sanitizeTitle, + stripConversationMarkers, +} from "./session-title.js"; export type { SessionTitleResult } from "./session-title.js"; export { Agent, createAgent } from "./agent.js"; export type { CreateAgentOptions, CreateSessionOptions, ResumeSessionOptions } from "./agent.js"; diff --git a/packages/core/src/interfaces.ts b/packages/core/src/interfaces.ts index c8c463c..fab5ea1 100644 --- a/packages/core/src/interfaces.ts +++ b/packages/core/src/interfaces.ts @@ -200,8 +200,14 @@ export interface SubagentRunner { spawn(input: { /** The child Agent's agentId; if omitted, reuses the current Agent (self-invocation). */ agentId?: string; - /** The Model used by the child Session; if omitted, uses the Project's default Model. */ + /** + * Upstream model id for the child Session, paired with `provider` — a model reference is + * always the complete pair. Omit both to use the Project's default Model; supplying one + * half without the other is rejected. + */ modelId?: string; + /** Provider group for `modelId`; required whenever `modelId` is given. */ + provider?: string; }): Promise; } diff --git a/packages/core/src/internal/session-support.ts b/packages/core/src/internal/session-support.ts index 9e6484e..2b5bbe2 100644 --- a/packages/core/src/internal/session-support.ts +++ b/packages/core/src/internal/session-support.ts @@ -35,7 +35,7 @@ export function formatSessionId(date: Date = new Date()): string { export function sessionEnvironment( workspaceDir: string, sessionId: string, - ids: { agentId: string; projectDir: string }, + ids: { agentId: string; projectDir: string; provider: string; modelId: string }, date = new Date(), ): SessionEnvironment { return { @@ -43,6 +43,8 @@ export function sessionEnvironment( cwd: workspaceDir, agentId: ids.agentId, projectDir: ids.projectDir, + provider: ids.provider, + modelId: ids.modelId, platform: process.platform, osVersion: getOsVersion(), date: formatLocalDate(date), diff --git a/packages/core/src/llm/generative-model.ts b/packages/core/src/llm/generative-model.ts index 87f7a0c..58ba081 100644 --- a/packages/core/src/llm/generative-model.ts +++ b/packages/core/src/llm/generative-model.ts @@ -21,7 +21,12 @@ * `context_engine` only consumes OmniMessage; all Uni* protocol details are encapsulated here. * Docs: /docs/interfaces § "The built-in implementation: GenerativeModel". */ -import { AutoLLMClient, ThinkingLevel } from "@prismshadow/agenthub"; +import { + AutoLLMClient, + EmptyResponseError, + ThinkingLevel, + ToolCallArgumentParseError, +} from "@prismshadow/agenthub"; import type { ContentItem, FinishReason, @@ -45,6 +50,7 @@ import { } from "../omnimessage/index.js"; import type { CompleteModelPayload, + Fidelity, OmniMessage, StopReason, TokenCounts, @@ -75,21 +81,43 @@ function parseToolArguments(raw: string): Record { return parsed !== null && typeof parsed === "object" ? (parsed as Record) : {}; } +/** + * Whether a fidelity payload carries at least one key (mirrors AgentHub's baseClient helper — + * an absent and an empty fidelity are equivalent). + */ +function hasFidelity(fidelity?: Fidelity): boolean { + return fidelity != null && Object.keys(fidelity).length > 0; +} + +/** + * Compare two fidelity payloads by value (mirrors AgentHub's baseClient helper). Fidelity + * objects are built with a stable key order by each AgentHub client, so JSON serialization is + * a faithful equality check. + */ +function fidelityEquals(a?: Fidelity, b?: Fidelity): boolean { + return JSON.stringify(a ?? {}) === JSON.stringify(b ?? {}); +} + +/** Spread helper: attach `fidelity` only when it carries at least one key. */ +function fidelityProp(fidelity?: Fidelity): { fidelity?: Fidelity } { + return hasFidelity(fidelity) ? { fidelity } : {}; +} + /** * Maps a complete OmniMessage payload to an AgentHub `ContentItem`. * Only complete model_msg payloads are supported; `partial_*` is an output-only protocol. */ function payloadToContentItem(payload: CompleteModelPayload): ContentItem { - // Provider-fidelity fields (signature / phase) are restored verbatim — some models require - // them when history is replayed back (e.g. Claude thinking signatures, GPT-5 encrypted - // reasoning and phase segmentation); losing them would break Session recovery. + // The provider-fidelity payload is opaque and restored verbatim — some models require it + // when history is replayed back (e.g. Claude thinking signatures, GPT-5 encrypted reasoning + // and phase segmentation, the OpenAI-compatible reasoning field name); losing it would break + // Session recovery. switch (payload.type) { case "text": return { type: "text", text: payload.text, - ...(payload.phase != null ? { phase: payload.phase } : {}), - ...(payload.signature !== undefined ? { signature: payload.signature } : {}), + ...fidelityProp(payload.fidelity), }; case "image_url": return { type: "image_url", image_url: payload.image_url }; @@ -98,20 +126,20 @@ function payloadToContentItem(payload: CompleteModelPayload): ContentItem { type: "inline_data", data: Buffer.from(payload.data, "base64"), mime_type: payload.mime_type, - ...(payload.signature !== undefined ? { signature: payload.signature } : {}), + ...fidelityProp(payload.fidelity), }; case "inline_thinking": return { type: "inline_thinking", data: Buffer.from(payload.data, "base64"), mime_type: payload.mime_type, - ...(payload.signature !== undefined ? { signature: payload.signature } : {}), + ...fidelityProp(payload.fidelity), }; case "thinking": return { type: "thinking", thinking: payload.thinking, - ...(payload.signature !== undefined ? { signature: payload.signature } : {}), + ...fidelityProp(payload.fidelity), }; case "tool_call": return { @@ -121,7 +149,7 @@ function payloadToContentItem(payload: CompleteModelPayload): ContentItem { arguments: parseToolArguments(payload.arguments), // On the way back, strip the uniqueness suffix to restore the provider's original id (see tool-call-ids.ts). tool_call_id: stripToolCallIdSuffix(payload.tool_call_id), - ...(payload.signature !== undefined ? { signature: payload.signature } : {}), + ...fidelityProp(payload.fidelity), }; case "tool_call_output": return { @@ -240,8 +268,8 @@ interface ToolCallAccumulator { toolCallId: string; /** The original tool_call_id reported by the provider (the attribution key for inbound events). */ providerKey: string; - /** Provider-fidelity field: signature (kept verbatim, produced alongside the complete tool_call). */ - signature: string | undefined; + /** Provider-fidelity payload (kept verbatim, produced alongside the complete tool_call). */ + fidelity: Fidelity | undefined; /** Whether this tool_call's complete message has already been emitted eagerly in `pushEvent` (avoids duplicate emission in finish). */ emitted: boolean; } @@ -274,12 +302,12 @@ export class EventTranslator { // Buffers needed for the complete message. private textBuffer = ""; private thinkingBuffer = ""; - // Provider-fidelity fields: the thinking block's signature (a signature - // marks the end of a block) and the current text segment's phase (sticky across segments; - // a differing phase marker starts a new segment) and signature. - private thinkingSignature: string | undefined; - private textPhase: string | null = null; - private textSignature: string | undefined; + // Provider-fidelity payloads of the currently open segments. Segmentation mirrors AgentHub's + // baseClient aggregation: a thinking block is closed by its fidelity payload (a run of equal + // fidelity is one block); a text segment is closed by a `fidelity.signature` and split by a + // differing `fidelity.phase`. + private thinkingFidelity: Fidelity | undefined; + private textFidelity: Fidelity | undefined; /** Provider id keys saved in order of appearance, so complete tool_calls are emitted in a stable order. */ private toolOrder: string[] = []; /** provider's original tool_call_id → the accumulator for the **latest** call under that id. */ @@ -306,17 +334,26 @@ export class EventTranslator { for (const item of event.content_items) { switch (item.type) { case "text": { - // Provider fidelity: a phase marker can arrive as an increment with **empty text** - // (e.g. GPT-5's segment markers). A phase differing from the current segment's phase - // starts a new segment — providers split by phase when replaying history, so mixing - // segments would break fidelity. Phase is sticky across segments (a subsequent segment - // with the same phase isn't re-marked). - if (item.phase != null && item.phase !== this.textPhase) { + // Mirrors AgentHub baseClient text aggregation. A `fidelity.phase` marker can arrive + // as an increment with **empty text** (e.g. GPT-5's segment markers): a phase + // differing from the current segment's phase starts a new segment — providers split + // by phase when replaying history, so mixing segments would break fidelity. A + // `fidelity.signature` closes the segment: any further content starts a new one. + // On merge, fidelity keys accumulate ({...current, ...incoming}). + const curPhase = (this.textFidelity as { phase?: unknown } | undefined)?.phase ?? null; + const inPhase = (item.fidelity as { phase?: unknown } | undefined)?.phase ?? null; + const signatureClosed = + (this.textFidelity as { signature?: unknown } | undefined)?.signature != null; + if ( + (signatureClosed && (item.text || hasFidelity(item.fidelity))) || + (inPhase != null && inPhase !== curPhase) + ) { yield* this.flushThinking("completed"); yield* this.flushText("completed"); - this.textPhase = item.phase; } - if (item.signature) this.textSignature = item.signature; + if (hasFidelity(item.fidelity)) { + this.textFidelity = { ...this.textFidelity, ...item.fidelity }; + } if (!item.text) break; // Type boundary: before a text segment starts, flush any unclosed thinking // segment, so the complete-message order matches generation order (thinking → text). @@ -332,16 +369,22 @@ export class EventTranslator { break; } case "thinking": { - // Provider fidelity: a thinking block ends with a signature (Claude's signature_delta - // is empty text + signature; redacted blocks carry sentinel text + signature; GPT-5 - // encrypted reasoning is empty text + signature). If thinking content/a new signature - // arrives after a signature is already set, that's a new block — close the current - // segment first, so each block's signature stays independently faithful and blocks - // don't bleed into each other when history is replayed. - if (this.thinkingSignature !== undefined && (item.thinking || item.signature)) { + // Mirrors AgentHub baseClient thinking aggregation: a thinking block is closed by its + // fidelity payload (Claude's signature_delta is empty text + fidelity{signature}; + // redacted blocks carry sentinel text + fidelity; GPT-5 encrypted reasoning is empty + // text + fidelity{id, encrypted_content}), and **a run of equal fidelity is one + // block** — OpenAI-compatible clients stamp every delta with the same + // fidelity{reasoning_field}, which must not split blocks. Content arriving after a + // different fidelity is set starts a new block, so each block's fidelity stays + // independently faithful when history is replayed. + if ( + hasFidelity(this.thinkingFidelity) && + !fidelityEquals(this.thinkingFidelity, item.fidelity) && + (item.thinking || hasFidelity(item.fidelity)) + ) { yield* this.flushThinking("completed"); } - if (item.signature) this.thinkingSignature = item.signature; + if (hasFidelity(item.fidelity)) this.thinkingFidelity = item.fidelity; if (!item.thinking) break; // Type boundary: before a thinking segment starts, flush any unclosed text // segment, so the complete-message order matches generation order (text → thinking). @@ -365,7 +408,7 @@ export class EventTranslator { if (item.tool_call_id) this.activeToolCallId = item.tool_call_id; const acc = this.ensureTool(providerKey, item.name); if (item.name) acc.name = item.name; - if (item.signature) acc.signature = item.signature; + if (hasFidelity(item.fidelity)) acc.fidelity = item.fidelity; // Externally always use the uniqueness-resolved id (with a `#n` suffix on provider id collisions), matching the complete tool_call. if (!this.toolStarted.has(acc.toolCallId)) { // Type boundary: before a new tool_call starts, flush any unclosed thinking/text @@ -418,7 +461,7 @@ export class EventTranslator { if (!acc || acc.emitted) acc = this.createTool(item.tool_call_id, item.name); acc.name = item.name; acc.completeArgs = JSON.stringify(item.arguments ?? {}); - if (item.signature) acc.signature = item.signature; + if (hasFidelity(item.fidelity)) acc.fidelity = item.fidelity; yield* this.emitCompleteTool(acc); break; } @@ -507,7 +550,7 @@ export class EventTranslator { arguments: acc.completeArgs ?? acc.argsBuffer, toolCallId: acc.toolCallId, stopReason, - ...(acc.signature !== undefined ? { signature: acc.signature } : {}), + ...(acc.fidelity !== undefined ? { fidelity: acc.fidelity } : {}), }); // activeToolCallId holds the provider id key (used to attribute id-less deltas); reset it by providerKey. if (this.activeToolCallId === acc.providerKey) { @@ -536,14 +579,12 @@ export class EventTranslator { yield partialThinking("stop", "", stopReason); this.thinkingStarted = false; } - // A thinking block with empty text but a signature (GPT-5 encrypted reasoning) still - // produces a complete message — the signature is required when replaying history. - if (this.thinkingBuffer || this.thinkingSignature !== undefined) { - yield thinkingMessage(this.thinkingBuffer, stopReason, { - ...(this.thinkingSignature !== undefined ? { signature: this.thinkingSignature } : {}), - }); + // A thinking block with empty text but a fidelity payload (GPT-5 encrypted reasoning) + // still produces a complete message — the fidelity is required when replaying history. + if (this.thinkingBuffer || hasFidelity(this.thinkingFidelity)) { + yield thinkingMessage(this.thinkingBuffer, stopReason, this.thinkingFidelity); this.thinkingBuffer = ""; - this.thinkingSignature = undefined; + this.thinkingFidelity = undefined; } } @@ -559,18 +600,14 @@ export class EventTranslator { yield partialText("stop", "", stopReason); this.textStarted = false; } - // A text segment with empty text but a signature (e.g. Gemini carrying a thoughtSignature on - // a text part) still produces a complete message — aligned with flushThinking, so the - // signature isn't lost or leaked into a later segment just because the buffer is empty. - if (this.textBuffer || this.textSignature !== undefined) { - yield assistantText(this.textBuffer, stopReason, { - ...(this.textPhase != null ? { phase: this.textPhase } : {}), - ...(this.textSignature !== undefined ? { signature: this.textSignature } : {}), - }); + // A text segment with empty text but a fidelity payload (e.g. Gemini carrying a + // thoughtSignature on a text part, or a GPT-5 phase marker with no text) still produces a + // complete message — aligned with flushThinking, so the fidelity isn't lost or leaked into + // a later segment just because the buffer is empty. + if (this.textBuffer || hasFidelity(this.textFidelity)) { + yield assistantText(this.textBuffer, stopReason, this.textFidelity); this.textBuffer = ""; - this.textSignature = undefined; - // textPhase is sticky across segments: a later segment with the same phase isn't - // re-marked; it's updated when a different phase marker appears. + this.textFidelity = undefined; } } @@ -601,7 +638,7 @@ export class EventTranslator { completeArgs: null, toolCallId: this.toolCallIds.allocate(providerKey), providerKey, - signature: undefined, + fidelity: undefined, emitted: false, }; this.tools.set(providerKey, acc); @@ -665,19 +702,37 @@ export function translateEvents( // --------------------------------------------------------------------------- /** - * Determines whether an error is an AgentHub / Provider response JSON parse error. + * Determines whether an error is an AgentHub / Provider "response delivered but unusable" + * parse or validation error. Two shapes (@prismshadow/agenthub 0.4.x): * - * AgentHub uses `JSON.parse` internally to parse response bodies, and a parse failure throws a - * `SyntaxError`; hence we judge directly by exception type (the `name` check also covers - * cross-realm or deserialization-reconstructed errors, and we probe down the `cause` chain for - * wrapped errors). This is not an auth/parameter failure but an incomplete LLM Request, and - * should end with `malformed` and be handed to the engine to retry. + * - A raw `SyntaxError` from `JSON.parse` on a response body; + * - AgentHub's own error classes: `ToolCallArgumentParseError` (streamed tool-call arguments + * are not valid JSON — e.g. a stream truncated mid-arguments) and `EmptyResponseError` + * (a completed response carrying thinking only, which cannot be replayed). + * + * In every case the turn was **not committed** to AgentHub history: this is not an + * auth/parameter failure but an incomplete LLM Request, and should end with `malformed` and + * be handed to the engine to reconnect and retry. Judged by exception type with a `name` + * fallback (covers cross-realm or deserialization-reconstructed errors), probing down the + * `cause` chain for wrapped errors. */ export function isMalformedJsonParseError(error: unknown): boolean { if (error == null) return false; - if (error instanceof SyntaxError) return true; + if ( + error instanceof SyntaxError || + error instanceof ToolCallArgumentParseError || + error instanceof EmptyResponseError + ) { + return true; + } const err = error as { name?: string; cause?: unknown }; - if (err.name === "SyntaxError") return true; + if ( + err.name === "SyntaxError" || + err.name === "ToolCallArgumentParseError" || + err.name === "EmptyResponseError" + ) { + return true; + } if (err.cause && err.cause !== error) { return isMalformedJsonParseError(err.cause); } @@ -692,7 +747,7 @@ export function isMalformedJsonParseError(error: unknown): boolean { * usage_metadata|finish_reason"). This is not an auth/parameter failure but an incomplete LLM * Request, and should end with `malformed` and be handed to the engine to reconnect and retry. * AgentHub doesn't provide an error type for this, so we match by message prefix - * (@prismshadow/agenthub 0.3.x), probing down the `cause` chain. + * (verified against @prismshadow/agenthub 0.4.x), probing down the `cause` chain. */ export function isIncompleteStreamError(error: unknown): boolean { if (error == null) return false; diff --git a/packages/core/src/omnimessage/builders.ts b/packages/core/src/omnimessage/builders.ts index 4311f42..e476b17 100644 --- a/packages/core/src/omnimessage/builders.ts +++ b/packages/core/src/omnimessage/builders.ts @@ -13,6 +13,7 @@ import type { CompactionMode, CompactionReason, EventMessage, + Fidelity, ImageUrlPayload, InlineDataPayload, InlineThinkingPayload, @@ -61,30 +62,28 @@ export function sessionMeta(payload: SessionMetaPayload): SessionMetaMessage { // Complete model_msg ----------------------------------------------------------- /** - * Provider fidelity fields: kept as-is and restored verbatim on replay. - * Builder convention: positional-argument-style builders carry these in a trailing `fidelity` - * object (narrowed via Pick per payload type — e.g. thinking only has signature); object-argument- - * style builders (toolCall) flatten `fidelity` fields into the parameter object alongside - * `stopReason`, mirroring the payload structure directly. + * Provider-fidelity payload (opaque, see `Fidelity` in types.ts): kept as-is and restored + * verbatim on replay. Builder convention: positional-argument-style builders take it as a + * trailing `fidelity` object; object-argument-style builders (toolCall) carry it in the + * parameter object alongside `stopReason`. An empty object is treated as absent — the + * payload field is only set when the fidelity carries at least one key. */ -export interface FidelityFields { - phase?: string | null; - signature?: string; +function fidelityProp(fidelity?: Fidelity): { fidelity?: Fidelity } { + return fidelity !== undefined && Object.keys(fidelity).length > 0 ? { fidelity } : {}; } export function textMessage( role: Role, text: string, stopReason: StopReason = "completed", - fidelity?: FidelityFields, + fidelity?: Fidelity, ): OmniMessage { return model({ type: "text", role, text, stop_reason: stopReason, - ...(fidelity?.phase != null ? { phase: fidelity.phase } : {}), - ...(fidelity?.signature !== undefined ? { signature: fidelity.signature } : {}), + ...fidelityProp(fidelity), }); } @@ -93,7 +92,7 @@ export const userText = (text: string): OmniMessage => textMessage( export const assistantText = ( text: string, stopReason: StopReason = "completed", - fidelity?: FidelityFields, + fidelity?: Fidelity, ): OmniMessage => textMessage("assistant", text, stopReason, fidelity); export function imageUrlMessage(imageUrl: string): OmniMessage { @@ -109,7 +108,7 @@ export function inlineData( role: Role, data: string, mimeType: string, - fidelity?: Pick, + fidelity?: Fidelity, ): OmniMessage { return model({ type: "inline_data", @@ -117,28 +116,28 @@ export function inlineData( data, mime_type: mimeType, stop_reason: "completed", - ...(fidelity?.signature !== undefined ? { signature: fidelity.signature } : {}), + ...fidelityProp(fidelity), }); } export function thinkingMessage( thinking: string, stopReason: StopReason = "completed", - fidelity?: Pick, + fidelity?: Fidelity, ): OmniMessage { return model({ type: "thinking", role: "assistant", thinking, stop_reason: stopReason, - ...(fidelity?.signature !== undefined ? { signature: fidelity.signature } : {}), + ...fidelityProp(fidelity), }); } export function inlineThinking( data: string, mimeType: string, - fidelity?: Pick, + fidelity?: Fidelity, ): OmniMessage { return model({ type: "inline_thinking", @@ -146,7 +145,7 @@ export function inlineThinking( data, mime_type: mimeType, stop_reason: "completed", - ...(fidelity?.signature !== undefined ? { signature: fidelity.signature } : {}), + ...fidelityProp(fidelity), }); } @@ -155,7 +154,7 @@ export function toolCall(args: { arguments: string; toolCallId: string; stopReason?: StopReason; - signature?: string; + fidelity?: Fidelity; }): OmniMessage { return model({ type: "tool_call", @@ -164,7 +163,7 @@ export function toolCall(args: { arguments: args.arguments, tool_call_id: args.toolCallId, stop_reason: args.stopReason ?? "completed", - ...(args.signature !== undefined ? { signature: args.signature } : {}), + ...fidelityProp(args.fidelity), }); } diff --git a/packages/core/src/omnimessage/types.ts b/packages/core/src/omnimessage/types.ts index 07f56e1..39c4a62 100644 --- a/packages/core/src/omnimessage/types.ts +++ b/packages/core/src/omnimessage/types.ts @@ -99,15 +99,23 @@ export interface SessionMetaPayload { // --------------------------------------------------------------------------- // Docs: /docs/omni-message § "model_msg: complete payloads" +/** + * Provider-fidelity payload (mirrors AgentHub's `Fidelity`): an arbitrary JSON-style object of + * wire-level data the LLM client records to reproduce the original message on replay — thinking + * signatures, phase labels, encrypted reasoning, the upstream reasoning field name, etc. Opaque + * to PenguinHarness: written to the Trace as-is and passed back verbatim; some models **require** + * it when history is replayed (e.g. Claude thinking signatures, GPT-5 encrypted reasoning) — + * losing it breaks Session resumption. + */ +export type Fidelity = Record; + export interface TextPayload { type: "text"; role: Role; text: string; stop_reason?: StopReason; - /** Provider fidelity field: text phase marker (e.g. GPT-5 segments by phase), kept as-is and restored verbatim. */ - phase?: string | null; - /** Provider fidelity field: signature, kept as-is and restored verbatim. */ - signature?: string; + /** Provider-fidelity payload (e.g. `phase` for GPT-5 segment markers, `signature`), kept as-is and restored verbatim. */ + fidelity?: Fidelity; } export interface ImageUrlPayload { @@ -125,8 +133,8 @@ export interface InlineDataPayload { data: string; mime_type: string; stop_reason?: StopReason; - /** Provider fidelity field: signature, kept as-is and restored verbatim. */ - signature?: string; + /** Provider-fidelity payload, kept as-is and restored verbatim. */ + fidelity?: Fidelity; } export interface ThinkingPayload { @@ -135,11 +143,12 @@ export interface ThinkingPayload { thinking: string; stop_reason?: StopReason; /** - * Provider fidelity field: thinking-block signature (Claude thinking blocks / redacted - * thinking, GPT-5 encrypted reasoning, etc. — **required** when some models replay history), - * kept as-is and restored verbatim — losing it breaks Session resumption. + * Provider-fidelity payload closing the thinking block (Claude thinking signatures / redacted + * thinking, GPT-5 encrypted reasoning, the OpenAI-compatible reasoning field name, etc. — + * **required** when some models replay history), kept as-is and restored verbatim — losing it + * breaks Session resumption. */ - signature?: string; + fidelity?: Fidelity; } export interface InlineThinkingPayload { @@ -149,8 +158,8 @@ export interface InlineThinkingPayload { data: string; mime_type: string; stop_reason?: StopReason; - /** Provider fidelity field: signature, kept as-is and restored verbatim. */ - signature?: string; + /** Provider-fidelity payload, kept as-is and restored verbatim. */ + fidelity?: Fidelity; } export interface ToolCallPayload { @@ -161,8 +170,8 @@ export interface ToolCallPayload { arguments: string; tool_call_id: string; stop_reason?: StopReason; - /** Provider fidelity field: signature, kept as-is and restored verbatim. */ - signature?: string; + /** Provider-fidelity payload, kept as-is and restored verbatim. */ + fidelity?: Fidelity; } export interface ToolCallOutputPayload { diff --git a/packages/core/src/session-title.ts b/packages/core/src/session-title.ts index f17d88c..cf99e22 100644 --- a/packages/core/src/session-title.ts +++ b/packages/core/src/session-title.ts @@ -22,6 +22,31 @@ const EXCERPT_MAX_CHARS = 2000; /** Cap on title length (fallback truncation for when the model occasionally ignores the constraint). */ const TITLE_MAX_CHARS = 30; +/** + * Special message markers that must never leak into a title. These are machine-inserted + * XML-ish blocks (a skill invocation wraps the body in `…`, a + * subagent handoff / scheduled task prepend their own blocks) — meaningful to the runtime, + * noise in a title. The list is a fixed allowlist so ordinary angle-bracket text (e.g. a + * user pasting `
`) is left untouched. + */ +const MARKER_TAGS = ["use_skills", "handoff_from", "scheduled_task"]; + +/** + * Strips machine-inserted marker blocks (see MARKER_TAGS) from conversation text so titles are + * built from the human-meaningful body only — both the material sent to the model and the + * fallback derived from the raw first message. Removes the paired `…` block (any + * inner content, across lines) and any stray unpaired `` / `` left behind. + */ +export function stripConversationMarkers(text: string): string { + let out = text; + for (const tag of MARKER_TAGS) { + out = out + .replace(new RegExp(`<${tag}>[\\s\\S]*?`, "g"), "") + .replace(new RegExp(``, "g"), ""); + } + return out.trim(); +} + export interface SessionTitleResult { /** The sanitized title; null when material is insufficient, the request fails, or the output is empty. */ title: string | null; @@ -44,6 +69,7 @@ export function buildTitlePrompt(userExcerpt: string, assistantExcerpt: string): "- Write the title in the SAME language the user is using.", "- Keep it short: at most 6 words, or ~16 characters for CJK.", "- Output ONLY the title text — no quotes, no trailing punctuation, no explanation.", + "- Answer immediately — do not think aloud or produce chain-of-thought.", "", "[User]", clip(userExcerpt), @@ -51,12 +77,15 @@ export function buildTitlePrompt(userExcerpt: string, assistantExcerpt: string): if (assistantExcerpt.trim()) { lines.push("", "[Assistant]", clip(assistantExcerpt)); } + // The trailing empty think block makes many reasoning models treat their thinking phase + // as already closed, so the one-off request spends its budget on the title itself. + lines.push("", ""); return lines.join("\n"); } -/** Sanitizes model output into a title: strips leading/trailing quotes/brackets and trailing punctuation (until stable), collapses whitespace, and truncates if too long; returns null for an empty result. */ +/** Sanitizes model output into a title: strips any leaked marker blocks and leading/trailing quotes/brackets and trailing punctuation (until stable), collapses whitespace, and truncates if too long; returns null for an empty result. */ export function sanitizeTitle(raw: string): string | null { - let t = raw.replace(/\s+/g, " ").trim(); + let t = stripConversationMarkers(raw).replace(/\s+/g, " ").trim(); // Stripping quotes can expose more punctuation underneath (or vice versa), so strip repeatedly until stable. for (let prev = ""; prev !== t;) { prev = t; @@ -81,10 +110,14 @@ export async function generateTitleWithLLM( llm: LLMInterface, args: { userText: string; assistantText: string; signal?: AbortSignal }, ): Promise { - if (!args.userText.trim()) { + // Strip machine markers before the model ever sees the material, so a skill-invocation + // block (or handoff / scheduled-task marker) can't bleed into the generated title. + const userMaterial = stripConversationMarkers(args.userText); + const assistantMaterial = stripConversationMarkers(args.assistantText); + if (!userMaterial.trim()) { return { title: null, usage: null }; } - const prompt = buildTitlePrompt(args.userText, args.assistantText); + const prompt = buildTitlePrompt(userMaterial, assistantMaterial); const gen = llm.streamGenerate({ newMessages: [userText(prompt)], ...(args.signal ? { signal: args.signal } : {}), diff --git a/packages/core/src/session.ts b/packages/core/src/session.ts index 1558a18..c30a2f9 100644 --- a/packages/core/src/session.ts +++ b/packages/core/src/session.ts @@ -65,15 +65,26 @@ export interface SessionConfig { inputImagesDir?: string; } -/** Cap on captured title material (chars per side, matching buildTitlePrompt's truncation); stops accumulating once exceeded. */ -const TITLE_MATERIAL_LIMIT = 2000; +/** + * Caps on captured title material (chars per side); accumulation stops once exceeded. The + * assistant body is capped tighter: a title only needs the opening of the answer, and hosts + * may start generating as soon as this much body text has streamed (see the Web server's + * early trigger) — a long answer would otherwise overrun the material. + */ +const TITLE_USER_MATERIAL_LIMIT = 2000; +const TITLE_ASSISTANT_MATERIAL_LIMIT = 1000; /** * Accumulates title material: the body text of complete text messages from the main session * (no origin) — thinking and tool calls naturally don't count — and stops once the cap is hit. */ -function appendTitleText(base: string, msg: OmniMessage, role: "user" | "assistant"): string { - if (base.length >= TITLE_MATERIAL_LIMIT) return base; +function appendTitleText( + base: string, + msg: OmniMessage, + role: "user" | "assistant", + limit: number, +): string { + if (base.length >= limit) return base; if (msg.origin && msg.origin.length > 0) return base; const p = msg.payload as { type?: string; role?: string; text?: string }; if (msg.type !== "model_msg" || p.type !== "text" || p.role !== role || !p.text) return base; @@ -153,12 +164,22 @@ export class Session { const capture = !this.titleMaterialFrozen; if (capture) { for (const m of newMessages) { - this.titleUserText = appendTitleText(this.titleUserText, m, "user"); + this.titleUserText = appendTitleText( + this.titleUserText, + m, + "user", + TITLE_USER_MATERIAL_LIMIT, + ); } } for await (const msg of this.engine.run(newMessages, opts)) { if (capture) { - this.titleAssistantText = appendTitleText(this.titleAssistantText, msg, "assistant"); + this.titleAssistantText = appendTitleText( + this.titleAssistantText, + msg, + "assistant", + TITLE_ASSISTANT_MATERIAL_LIMIT, + ); } yield msg; } diff --git a/packages/core/src/state/agent-state.ts b/packages/core/src/state/agent-state.ts index 450c228..42f8fa4 100644 --- a/packages/core/src/state/agent-state.ts +++ b/packages/core/src/state/agent-state.ts @@ -26,6 +26,8 @@ import { VAULT_KEYS_PLACEHOLDER, SKILL_METADATA_PLACEHOLDER, CWD_PLACEHOLDER, + PROVIDER_PLACEHOLDER, + MODEL_ID_PLACEHOLDER, DATE_PLACEHOLDER, defaultAgentsMd, defaultSystemConfig, @@ -83,6 +85,10 @@ export interface SessionEnvironmentValues { agentId: string; /** Absolute path to this Project's directory (system Prompt placeholder {{PROJECT_DIR}}; Agent State/scratchpad paths are derived from it). */ projectDir: string; + /** The session model's provider group (system Prompt placeholder {{PROVIDER}}; paired with modelId to form the model reference). */ + provider: string; + /** The session model's upstream model id (system Prompt placeholder {{MODEL_ID}}). */ + modelId: string; platform: string; osVersion: string; date: string; @@ -371,6 +377,10 @@ export function assembleSystemPrompt( .join(sessionEnvironment?.sessionId ?? "") .split(CWD_PLACEHOLDER) .join(sessionEnvironment?.cwd ?? "") + .split(PROVIDER_PLACEHOLDER) + .join(sessionEnvironment?.provider ?? "") + .split(MODEL_ID_PLACEHOLDER) + .join(sessionEnvironment?.modelId ?? "") .split(PLATFORM_PLACEHOLDER) .join(sessionEnvironment?.platform ?? "") .split(OS_VERSION_PLACEHOLDER) diff --git a/packages/core/src/state/default-config.ts b/packages/core/src/state/default-config.ts index 62a73d8..c2242e5 100644 --- a/packages/core/src/state/default-config.ts +++ b/packages/core/src/state/default-config.ts @@ -28,6 +28,8 @@ export const SESSION_ID_PLACEHOLDER = "{{SESSION_ID}}"; export const CWD_PLACEHOLDER = "{{CWD}}"; export const AGENT_ID_PLACEHOLDER = "{{AGENT_ID}}"; export const PROJECT_DIR_PLACEHOLDER = "{{PROJECT_DIR}}"; +export const PROVIDER_PLACEHOLDER = "{{PROVIDER}}"; +export const MODEL_ID_PLACEHOLDER = "{{MODEL_ID}}"; export const PLATFORM_PLACEHOLDER = "{{PLATFORM}}"; export const OS_VERSION_PLACEHOLDER = "{{OS_VERSION}}"; export const DATE_PLACEHOLDER = "{{DATE}}"; @@ -141,9 +143,11 @@ Skills are reusable instruction packages stored under /agents/ + // form) -- + { + modelId: "accounts/fireworks/models/deepseek-v4-flash", + displayName: "DeepSeek V4 Flash", + provider: "fireworks", + contextWindow: 1000000, + pricing: usd(0.03, 0.14, 0.28), + supportsVision: false, + clientType: "openai", + baseUrl: FIREWORKS_BASE_URL, + }, + { + modelId: "accounts/fireworks/models/deepseek-v4-pro", + displayName: "DeepSeek V4 Pro", + provider: "fireworks", + contextWindow: 1000000, + pricing: usd(0.15, 1.74, 3.48), + supportsVision: false, + clientType: "openai", + baseUrl: FIREWORKS_BASE_URL, + }, + { + modelId: "accounts/fireworks/models/glm-5p2", + displayName: "GLM-5.2", + provider: "fireworks", + contextWindow: 1000000, + pricing: usd(0.14, 1.4, 4.4), + supportsVision: false, + clientType: "openai", + baseUrl: FIREWORKS_BASE_URL, + }, + { + modelId: "accounts/fireworks/models/kimi-k2p7-code", + displayName: "Kimi K2.7 Code", + provider: "fireworks", + contextWindow: 262144, + pricing: usd(0.19, 0.95, 4), + supportsVision: true, + clientType: "openai", + baseUrl: FIREWORKS_BASE_URL, + }, + { + modelId: "accounts/fireworks/models/minimax-m3", + displayName: "MiniMax M3", + provider: "fireworks", + contextWindow: 524288, + pricing: usd(0.06, 0.3, 1.2), + supportsVision: true, + clientType: "openai", + baseUrl: FIREWORKS_BASE_URL, + }, + // -- SiliconFlow (gateway, official CNY pricing: cache hit / input / output) -- + { + modelId: "deepseek-ai/DeepSeek-V4-Flash", + displayName: "DeepSeek V4 Flash", + provider: "siliconflow", + contextWindow: 1000000, + pricing: cny(0.02, 1, 2), + supportsVision: false, + clientType: "openai", + baseUrl: SILICONFLOW_BASE_URL, + }, + { + modelId: "deepseek-ai/DeepSeek-V4-Pro", + displayName: "DeepSeek V4 Pro", + provider: "siliconflow", + contextWindow: 1000000, + pricing: cny(0.1, 12, 24), + supportsVision: false, + clientType: "openai", + baseUrl: SILICONFLOW_BASE_URL, + }, + { + modelId: "meituan-longcat/LongCat-2.0", + displayName: "LongCat 2.0", + provider: "siliconflow", + contextWindow: 1000000, + pricing: cny(0.1, 5, 20), + supportsVision: false, + clientType: "openai", + baseUrl: SILICONFLOW_BASE_URL, + }, + { + modelId: "moonshotai/Kimi-K2.7-Code", + displayName: "Kimi K2.7 Code", + provider: "siliconflow", + contextWindow: 262144, + pricing: cny(1.3, 6.5, 27), + supportsVision: true, + clientType: "openai", + baseUrl: SILICONFLOW_BASE_URL, + }, + { + modelId: "zai-org/GLM-5.2", + displayName: "GLM-5.2", + provider: "siliconflow", + contextWindow: 1000000, + pricing: cny(2, 8, 28), + supportsVision: false, + clientType: "openai", + baseUrl: SILICONFLOW_BASE_URL, + }, + // -- Qwen Token Plan (subscription gateway; vision flags per the plan's supported-model + // table). Pricing and context windows from each model's page at + // www.qianwenai.com/models/ (official CNY list prices; limited-time promotions such as + // the 20%/50% off discounts are not stored). qwen3.8-max-preview is preview-only with a + // quota-multiplier promotion and publishes no per-token list price nor a context window, so + // it carries no pricing and uses its family's 1M window. -- + { + modelId: "deepseek-v4-pro", + displayName: "DeepSeek V4 Pro", + provider: "qwen-token-plan", + contextWindow: 1000000, + pricing: cny(1, 12, 24), + supportsVision: false, + clientType: "openai", + baseUrl: QWEN_TOKEN_PLAN_BASE_URL, + }, + { + modelId: "glm-5.2", + displayName: "GLM-5.2", + provider: "qwen-token-plan", + contextWindow: 1048576, + pricing: cny(2, 8, 28), + supportsVision: false, + clientType: "openai", + baseUrl: QWEN_TOKEN_PLAN_BASE_URL, + }, + { + modelId: "qwen3.8-max-preview", + displayName: "Qwen 3.8 Max Preview", + provider: "qwen-token-plan", + contextWindow: 1000000, + supportsVision: true, + clientType: "openai", + baseUrl: QWEN_TOKEN_PLAN_BASE_URL, + }, + { + modelId: "qwen3.7-max", + displayName: "Qwen 3.7 Max", + provider: "qwen-token-plan", + contextWindow: 1000000, + pricing: cny(2.4, 12, 36), + supportsVision: false, + clientType: "openai", + baseUrl: QWEN_TOKEN_PLAN_BASE_URL, + }, + { + modelId: "qwen3.7-plus", + displayName: "Qwen 3.7 Plus", + provider: "qwen-token-plan", + contextWindow: 1000000, + pricing: cny(0.4, 2, 8), + supportsVision: true, + clientType: "openai", + baseUrl: QWEN_TOKEN_PLAN_BASE_URL, + }, + // -- Qwen Pay-As-You-Go (DashScope's OpenAI-compatible pay-per-token marketplace; official + // CNY list prices and specs from each model's page at www.qianwenai.com/models/ — + // resold third-party models keep their vendor-prefixed upstream ids) -- + { + modelId: "kimi/kimi-k3", + displayName: "Kimi K3", + provider: "qwen-pay-as-you-go", + contextWindow: 1048576, + pricing: cny(2, 20, 100), + supportsVision: true, + clientType: "openai", + baseUrl: QWEN_PAYG_BASE_URL, + }, + { + modelId: "qwen3.7-max", + displayName: "Qwen 3.7 Max", + provider: "qwen-pay-as-you-go", + contextWindow: 1000000, + pricing: cny(2.4, 12, 36), + supportsVision: false, + clientType: "openai", + baseUrl: QWEN_PAYG_BASE_URL, + }, + { + modelId: "qwen3.7-plus", + displayName: "Qwen 3.7 Plus", + provider: "qwen-pay-as-you-go", + contextWindow: 1000000, + pricing: cny(0.4, 2, 8), + supportsVision: true, + clientType: "openai", + baseUrl: QWEN_PAYG_BASE_URL, + }, + { + modelId: "ZHIPU/GLM-5.2", + displayName: "GLM-5.2", + provider: "qwen-pay-as-you-go", + contextWindow: 1048576, + pricing: cny(2, 8, 28), + supportsVision: false, + clientType: "openai", + baseUrl: QWEN_PAYG_BASE_URL, + }, + // -- Google Gemini (official USD pricing) -- + { + modelId: "gemini-3.5-flash", + displayName: "Gemini 3.5 Flash", + provider: "google", + contextWindow: 1048576, + pricing: usd(0.15, 1.5, 9), + supportsVision: true, + }, + { + modelId: "gemini-3.1-flash-lite", + displayName: "Gemini 3.1 Flash-Lite", + provider: "google", + contextWindow: 1048576, + pricing: usd(0.025, 0.25, 1.5), + supportsVision: true, + }, + { + // ≤200K input tier; >200K has official surcharge pricing (see file header comment). + modelId: "gemini-3.1-pro-preview", + displayName: "Gemini 3.1 Pro (Preview)", + provider: "google", + contextWindow: 1048576, + pricing: usd(0.2, 2, 12), + supportsVision: true, + }, + { + modelId: "gemini-3-flash-preview", + displayName: "Gemini 3 Flash (Preview)", + provider: "google", + contextWindow: 1048576, + pricing: usd(0.05, 0.5, 3), + supportsVision: true, + }, + // -- Anthropic (official USD pricing; cache write = 1.25 x input) -- { modelId: "claude-opus-4-8", displayName: "Claude Opus 4.8", @@ -205,7 +666,7 @@ export const MODEL_CATALOG: ModelCatalogEntry[] = [ pricing: usd(0.3, 3.75, 15), supportsVision: true, }, - // —— OpenAI —— + // -- OpenAI (official USD pricing) -- { modelId: "gpt-5.5", displayName: "GPT-5.5", @@ -256,41 +717,7 @@ export const MODEL_CATALOG: ModelCatalogEntry[] = [ pricing: usd(30, 30, 180), supportsVision: true, }, - // —— Google Gemini —— - { - // ≤200K input tier; >200K has official surcharge pricing (see file header comment). - modelId: "gemini-3.1-pro-preview", - displayName: "Gemini 3.1 Pro (Preview)", - provider: "google", - contextWindow: 1048576, - pricing: usd(0.2, 2, 12), - supportsVision: true, - }, - { - modelId: "gemini-3.5-flash", - displayName: "Gemini 3.5 Flash", - provider: "google", - contextWindow: 1048576, - pricing: usd(0.15, 1.5, 9), - supportsVision: true, - }, - { - modelId: "gemini-3-flash-preview", - displayName: "Gemini 3 Flash (Preview)", - provider: "google", - contextWindow: 1048576, - pricing: usd(0.05, 0.5, 3), - supportsVision: true, - }, - { - modelId: "gemini-3.1-flash-lite", - displayName: "Gemini 3.1 Flash-Lite", - provider: "google", - contextWindow: 1048576, - pricing: usd(0.025, 0.25, 1.5), - supportsVision: true, - }, - // —— Z.AI (GLM) —— + // -- Z.AI (GLM) -- { modelId: "glm-5.2", displayName: "GLM-5.2", @@ -332,80 +759,6 @@ export const MODEL_CATALOG: ModelCatalogEntry[] = [ pricing: cny(0.7, 4, 21), supportsVision: true, }, - // -- OpenRouter (gateway: uses OpenAI-compatible protocol, preset base URL) -- - { - modelId: "xiaomi/mimo-v2.5", - displayName: "MiMo-V2.5", - provider: "openrouter", - contextWindow: 1048576, - pricing: usd(0.0028, 0.14, 0.28), - supportsVision: true, - clientType: "openai", - baseUrl: OPENROUTER_BASE_URL, - }, - { - modelId: "tencent/hy3", - displayName: "Hy3", - provider: "openrouter", - contextWindow: 262144, - pricing: usd(0.035, 0.14, 0.58), - supportsVision: false, - clientType: "openai", - baseUrl: OPENROUTER_BASE_URL, - }, - { - // No official separate cache price published: cache_read uses the standard input price (no discount assumed). - modelId: "minimax/minimax-m3", - displayName: "MiniMax M3", - provider: "openrouter", - contextWindow: 1048576, - pricing: usd(0.06, 0.3, 1.2), - supportsVision: true, - clientType: "openai", - baseUrl: OPENROUTER_BASE_URL, - }, - { - // No official separate cache price published: cache_read uses the standard input price. - modelId: "stepfun/step-3.7-flash", - displayName: "Step 3.7 Flash", - provider: "openrouter", - contextWindow: 256000, - pricing: usd(0.04, 0.2, 1.15), - supportsVision: true, - clientType: "openai", - baseUrl: OPENROUTER_BASE_URL, - }, - // -- SiliconFlow (gateway, official CNY pricing: cache hit / input / output) -- - { - modelId: "zai-org/GLM-5.2", - displayName: "GLM-5.2", - provider: "siliconflow", - contextWindow: 1000000, - pricing: cny(2, 8, 28), - supportsVision: false, - clientType: "openai", - baseUrl: SILICONFLOW_BASE_URL, - }, - { - modelId: "deepseek-ai/DeepSeek-V4-Pro", - displayName: "DeepSeek V4 Pro", - provider: "siliconflow", - contextWindow: 1000000, - pricing: cny(0.1, 12, 24), - supportsVision: false, - clientType: "openai", - baseUrl: SILICONFLOW_BASE_URL, - }, - { - modelId: "meituan-longcat/LongCat-2.0", - displayName: "LongCat 2.0", - provider: "siliconflow", - contextWindow: 1000000, - pricing: cny(0.1, 5, 20), - supportsVision: false, - clientType: "openai", - baseUrl: SILICONFLOW_BASE_URL, - }, ]; /** Looks up a catalog entry by (provider, upstream id) pair (**the sole catalog-matching entry point**); returns undefined if not in the catalog. */ @@ -416,15 +769,6 @@ export function catalogEntryFor( return MODEL_CATALOG.find((m) => m.provider === provider && m.modelId === upstreamId); } -/** - * Infers the provider for an upstream id from the built-in catalog (used to default - * `provider` on `model add`): if it matches a catalog entry, use that entry's provider - * (upstream ids are globally unique within the catalog); otherwise custom. - */ -export function inferProviderForUpstream(upstreamId: string): string { - return MODEL_CATALOG.find((m) => m.modelId === upstreamId)?.provider ?? "custom"; -} - /** Looks up provider info by provider id; returns undefined for an unknown id. */ export function providerInfo(providerId: string): ModelProviderInfo | undefined { return MODEL_PROVIDERS.find((p) => p.id === providerId); @@ -487,3 +831,43 @@ export function presetModelEntries(): ModelEntry[] { ...(m.baseUrl !== undefined ? { base_url: m.baseUrl } : {}), })); } + +/** + * The model's own homepage/detail page for the frontend's model-card link. Gateway groups + * have a stable per-model URL pattern (works for user-added ids in those groups too); + * direct-vendor models link to the vendor's model list/docs page; the Token Plan preview + * model has no dedicated page and links to the plan's model overview; custom and + * user-defined groups have no page to vouch for. + */ +export function modelHomepageUrl(provider: string, modelId: string): string | undefined { + if (provider === "openrouter") return `https://openrouter.ai/${modelId}`; + if (provider === "qwen-token-plan") { + return modelId === "qwen3.8-max-preview" + ? providerInfo(provider)?.modelsUrl + : `https://www.qianwenai.com/models/${modelId}`; + } + if (provider === "fireworks") { + // API id "accounts//models/" -> page "app.fireworks.ai/models//"; + // nonconforming (user-added) ids fall back to the models listing. + const m = /^accounts\/([^/]+)\/models\/(.+)$/.exec(modelId); + return m + ? `https://app.fireworks.ai/models/${m[1]}/${m[2]}` + : providerInfo(provider)?.modelsUrl; + } + if (provider === "qwen-pay-as-you-go") { + return `https://www.qianwenai.com/models/${encodeURIComponent(modelId)}`; + } + if (provider === "zhipu") { + // Z.AI's per-model guide pages use the bare model id as the slug. + return `https://docs.z.ai/guides/llm/${modelId}`; + } + if (provider === "moonshot") { + // Moonshot's pricing pages: kimi-k2.6 -> chat-k26 (dot dropped); other ids fall back. + const m = /^kimi-k(\d+)\.(\d+)$/.exec(modelId); + return m + ? `https://platform.kimi.com/docs/pricing/chat-k${m[1]}${m[2]}` + : providerInfo(provider)?.modelsUrl; + } + if (provider === "custom") return undefined; + return providerInfo(provider)?.modelsUrl; +} diff --git a/packages/core/src/state/project-config.ts b/packages/core/src/state/project-config.ts index 0780070..d0f4997 100644 --- a/packages/core/src/state/project-config.ts +++ b/packages/core/src/state/project-config.ts @@ -17,11 +17,17 @@ * the unique key — string concatenation like `/` is forbidden anywhere in the * pipeline. `model_id` is the upstream request id, sent to AgentHub unchanged; `default_model` / * `vision_model` are paired `{ provider, model_id }` references (a TOML inline table). + * + * A caller always supplies the **complete pair**: `provider` is never guessed from the builtin + * catalog and never derived from whichever configured entry happens to carry the same + * `model_id`. Both halves or neither — a `model_id` without a `provider` is an error, not a + * lookup, because resolving it would silently point credentials and pricing at a vendor the + * caller never named. */ import fs from "node:fs/promises"; import path from "node:path"; import { parse as parseToml, stringify as stringifyToml } from "smol-toml"; -import { inferProviderForUpstream, presetModelEntries } from "./model-catalog.js"; +import { presetModelEntries } from "./model-catalog.js"; import { projectConfigPath } from "./paths.js"; /** Model reference: a `(provider, model_id)` pair (never string-concatenated anywhere). */ @@ -247,9 +253,11 @@ export async function saveProjectConfig( /** * Adds or updates a Model: - * - Upserts into `models`, deduplicated by the `(provider, model_id)` pair (provider may be - * omitted — the builtin catalog is used to infer the upstream id's group, falling back to - * custom if it can't be inferred); + * - Upserts into `models`, deduplicated by the `(provider, model_id)` pair — both halves are + * supplied by the caller, since the group is never guessed from the builtin catalog (a + * gateway reselling a vendor model keeps the vendor's upstream id, so a bare id names no + * single group, and guessing wrong files the caller's api_key under a vendor they never + * picked); a model outside every known group is added under `"custom"` explicitly; * - If `api_key`/`base_url` are provided, they're written inline into the entry; * - Set as the default Model (a paired reference) when `opts.setDefault` is true. * Reads the existing config (or the default), saves after the change, and returns the updated @@ -259,8 +267,8 @@ export async function addModel( root: string, projectId: string, entry: { - /** provider group; inferred from the builtin catalog when omitted (`inferProviderForUpstream`, falling back to custom if it can't be inferred). */ - provider?: string; + /** provider group (required; never inferred — pass `"custom"` for a model outside the known groups). */ + provider: string; /** Upstream model id (sent to AgentHub unchanged). */ model_id: string; context_window?: number; @@ -275,7 +283,7 @@ export async function addModel( opts?: { setDefault?: boolean }, ): Promise { const cfg = await loadProjectConfig(root, projectId); - const provider = entry.provider ?? inferProviderForUpstream(entry.model_id); + const { provider } = entry; // upsert: layers new fields on top of the existing entry; fields not explicitly provided // (e.g. context_window) keep their existing value, so a call like "just add an api_key" @@ -400,35 +408,19 @@ export function getModel(cfg: ProjectConfig, ref: ModelRef): ModelEntry | undefi } /** - * Resolves a model reference (the **single entry point for "provider omitted"**, shared by core - * and CLI/server — never set up a second one): - * - `provider` given: validated for existence by exact paired reference; - * - `provider` omitted: an exact-match lookup on `model_id` (no fuzzy matching of any kind) — - * resolvable only when **exactly one** entry matches; 0 or multiple matches always report a - * clear error (an ambiguity error lists the candidate paired references). + * Validates a `(provider, model_id)` pair against the Project config and returns it as a + * `ModelRef` (the **single validation entry point**, shared by core and CLI/server — never set + * up a second one). Both halves are required: this only ever checks that the exact pair is + * configured, it never searches for a group to attach to a bare `model_id`. A pair the config + * doesn't have throws — a reference outside the config would leave credentials, pricing, and + * the context window unavailable at request time. */ -export function resolveModelRef(cfg: ProjectConfig, modelId: string, provider?: string): ModelRef { - if (provider !== undefined) { - const ref: ModelRef = { provider, model_id: modelId }; - if (!getModel(cfg, ref)) { - throw new Error( - `Model is not in the Project config: ${formatModelRef(ref)}. Use \`penguin config model list\` to see the configured models, or \`penguin config model add\` to add one.`, - ); - } - return ref; - } - const candidates = cfg.models.filter((m) => m.model_id === modelId); - if (candidates.length === 1) { - return { provider: candidates[0]!.provider, model_id: modelId }; - } - if (candidates.length === 0) { +export function resolveModelRef(cfg: ProjectConfig, modelId: string, provider: string): ModelRef { + const ref: ModelRef = { provider, model_id: modelId }; + if (!getModel(cfg, ref)) { throw new Error( - `Model is not in the Project config: no entry has model_id ${modelId}. Use \`penguin config model list\` to see the configured models, or \`penguin config model add\` to add one.`, + `Model is not in the Project config: ${formatModelRef(ref)}. Use \`penguin config model list\` to see the configured models, or \`penguin config model add\` to add one.`, ); } - throw new Error( - `Ambiguous model reference: model_id ${modelId} matches multiple entries: ${candidates - .map((m) => formatModelRef({ provider: m.provider, model_id: m.model_id })) - .join(", ")}. Specify provider to give a paired reference.`, - ); + return ref; } diff --git a/packages/core/test/agent.test.ts b/packages/core/test/agent.test.ts index 23a5c66..d332bf5 100644 --- a/packages/core/test/agent.test.ts +++ b/packages/core/test/agent.test.ts @@ -80,14 +80,18 @@ describe("Agent.createSession workspace handling", () => { expect(path.isAbsolute(session.workspaceDir)).toBe(true); }); - it("rejects a modelId that is not in the Project config with a clear error", async () => { + it("rejects a model reference that is not in the Project config with a clear error", async () => { const agent = await createAgent(); const ws = path.join(tmpRoot, "ws-bad-model"); await fs.mkdir(ws, { recursive: true }); // A reference outside the config is not silently allowed (the unique key is provider + // model_id); the error is thrown before creating the temp Workspace. await expect( - agent.createSession({ workspaceDir: ws, modelId: "not-configured-model" }), + agent.createSession({ + workspaceDir: ws, + modelId: "not-configured-model", + provider: "custom", + }), ).rejects.toThrow(/is not in the Project config/); await expect( agent.createSession({ workspaceDir: ws, modelId: "deepseek-v4-pro", provider: "openai" }), @@ -129,59 +133,62 @@ describe("Agent.createSession model reference ((provider, model_id) pair)", () = } }); - it("resolves a unique bare model_id and accepts an explicit pair", async () => { - const agent = await createAgent(); - const ws = path.join(tmpRoot, "ws-ref-pair"); - await fs.mkdir(ws, { recursive: true }); - // Provider omitted: model_id is a globally unique exact match in the config -> resolves to that entry. - const bare = await agent.createSession({ workspaceDir: ws, modelId: "deepseek-v4-flash" }); - try { - expect(bare.provider).toBe("deepseek"); - expect(bare.modelId).toBe("deepseek-v4-flash"); - } finally { - bare.dispose(); - } - const paired = await agent.createSession({ - workspaceDir: ws, - modelId: "claude-sonnet-4-6", - provider: "anthropic", - }); - try { - expect(paired.provider).toBe("anthropic"); - expect(paired.modelId).toBe("claude-sonnet-4-6"); - } finally { - paired.dispose(); - } - }); - - it("rejects an ambiguous bare model_id and a provider without modelId", async () => { - // Two providers coexist with the same model_id: omitting provider throws an ambiguity - // error (listing the candidate pair references). + it("selects the entry named by the pair, even when a second group sells the same model_id", async () => { + // A user-run proxy resells claude-sonnet-4-6 under the same upstream id: the two entries + // coexist and the pair — not the bare id — decides which one (and therefore which + // credential and base_url) the Session runs on. await addModel(tmpRoot, DEFAULT_PROJECT_ID, { provider: "myproxy", model_id: "claude-sonnet-4-6", }); const agent = await createAgent(); - const ws = path.join(tmpRoot, "ws-ref-ambiguous"); + const ws = path.join(tmpRoot, "ws-ref-pair"); await fs.mkdir(ws, { recursive: true }); - await expect( - agent.createSession({ workspaceDir: ws, modelId: "claude-sonnet-4-6" }), - ).rejects.toThrow(/Ambiguous.*\(provider=anthropic, model_id=claude-sonnet-4-6\)/); - // Adding provider resolves it. - const session = await agent.createSession({ + const vendor = await agent.createSession({ + workspaceDir: ws, + modelId: "claude-sonnet-4-6", + provider: "anthropic", + }); + try { + expect(vendor.provider).toBe("anthropic"); + expect(vendor.modelId).toBe("claude-sonnet-4-6"); + } finally { + vendor.dispose(); + } + const proxied = await agent.createSession({ workspaceDir: ws, modelId: "claude-sonnet-4-6", provider: "myproxy", }); try { - expect(session.provider).toBe("myproxy"); + expect(proxied.provider).toBe("myproxy"); + expect(proxied.modelId).toBe("claude-sonnet-4-6"); + } finally { + proxied.dispose(); + } + }); + + it("rejects half a reference: modelId without provider, and provider without modelId", async () => { + const agent = await createAgent(); + const ws = path.join(tmpRoot, "ws-ref-half"); + await fs.mkdir(ws, { recursive: true }); + // A bare model_id is never resolved against the config, not even when exactly one entry + // carries it (deepseek-v4-flash is unique here): the group is the caller's to name. + await expect( + agent.createSession({ workspaceDir: ws, modelId: "deepseek-v4-flash" }), + ).rejects.toThrow(/must be given as a \(provider, model_id\) pair/); + // The mirror case: provider alone is not a reference either. + await expect(agent.createSession({ workspaceDir: ws, provider: "deepseek" })).rejects.toThrow( + /must be given as a \(provider, model_id\) pair/, + ); + // Neither half given is the documented "use the Project default" path, not an error. + const session = await agent.createSession({ workspaceDir: ws }); + try { + expect(session.provider).toBe("deepseek"); + expect(session.modelId).toBe("deepseek-v4-pro"); } finally { session.dispose(); } - // provider cannot be used alone (the reference must be a pair). - await expect(agent.createSession({ workspaceDir: ws, provider: "anthropic" })).rejects.toThrow( - /provider cannot be used alone/, - ); }); }); diff --git a/packages/core/test/builtin-agents.test.ts b/packages/core/test/builtin-agents.test.ts index 0efe145..a71fbff 100644 --- a/packages/core/test/builtin-agents.test.ts +++ b/packages/core/test/builtin-agents.test.ts @@ -122,6 +122,8 @@ describe("Project Dir / Agent ID placeholders", () => { cwd: "/tmp/ws", agentId: "env_agent", projectDir: "/tmp/proj", + provider: "deepseek", + modelId: "deepseek-v4-pro", platform: "linux", osVersion: "test", date: "2026-07-08", diff --git a/packages/core/test/llm.test.ts b/packages/core/test/llm.test.ts index ce52b86..9237f13 100644 --- a/packages/core/test/llm.test.ts +++ b/packages/core/test/llm.test.ts @@ -10,7 +10,11 @@ * determination. */ import { describe, expect, it } from "vitest"; -import { ThinkingLevel } from "@prismshadow/agenthub"; +import { + EmptyResponseError, + ThinkingLevel, + ToolCallArgumentParseError, +} from "@prismshadow/agenthub"; import type { UniEvent, UniMessage, UsageMetadata } from "@prismshadow/agenthub"; import type { LLMOutcome } from "../src/interfaces.js"; @@ -1105,6 +1109,37 @@ describe("isMalformedJsonParseError", () => { ); expect(isMalformedJsonParseError(new Error("socket hang up"))).toBe(false); }); + + it("detects AgentHub 0.4 parse/validation error classes (truncated tool args, thinking-only)", () => { + // A stream truncated mid-arguments surfaces as ToolCallArgumentParseError since agenthub + // 0.4 (previously a raw SyntaxError) — must stay malformed so the engine reconnects. + expect( + isMalformedJsonParseError( + new ToolCallArgumentParseError({ + client: "Claude5Client", + toolName: "exec_command", + toolCallId: "toolu_broken_1", + rawArguments: '{"cmd": "ec', + reason: "Unterminated string in JSON at position 11", + }), + ), + ).toBe(true); + // A completed thinking-only response cannot be replayed (400 on the next turn): retrying + // via malformed gives the model another chance instead of failing the turn. + expect( + isMalformedJsonParseError( + new EmptyResponseError({ client: "Claude5Client", finishReason: "stop" }), + ), + ).toBe(true); + // Also detectable via the name fallback and the cause chain. + expect( + isMalformedJsonParseError( + new Error("request failed", { + cause: new EmptyResponseError({ client: "GPT5_5Client", finishReason: null }), + }), + ), + ).toBe(true); + }); }); describe("isIncompleteStreamError", () => { @@ -1337,35 +1372,40 @@ describe("GenerativeModel.streamGenerate outcome classification (PRN-013)", () = }); }); -describe("provider fidelity fields (signature / phase)", () => { +describe("provider fidelity payloads (opaque, AgentHub 0.4 semantics)", () => { const complete = (messages: ReturnType["messages"]) => messages.filter((m) => !(m.payload as { type: string }).type.startsWith("partial_")); - it("captures the thinking signature arriving as an empty-text delta (Claude signature_delta)", () => { + it("captures the thinking fidelity arriving as an empty-text delta (Claude signature_delta)", () => { const { messages } = translateEvents([ ev({ content_items: [{ type: "thinking", thinking: "let me think" }] }), - ev({ content_items: [{ type: "thinking", thinking: "", signature: "sig-abc" }] }), + ev({ + content_items: [{ type: "thinking", thinking: "", fidelity: { signature: "sig-abc" } }], + }), ev({ content_items: [{ type: "text", text: "answer" }] }), ev({ event_type: "stop", content_items: [], finish_reason: "stop" }), ]); const thinking = complete(messages).find( (m) => (m.payload as { type: string }).type === "thinking", )!; - expect((thinking.payload as { thinking: string; signature?: string }).thinking).toBe( - "let me think", - ); - expect((thinking.payload as { signature?: string }).signature).toBe("sig-abc"); + const p = thinking.payload as { thinking: string; fidelity?: Record }; + expect(p.thinking).toBe("let me think"); + expect(p.fidelity).toEqual({ signature: "sig-abc" }); }); - it("splits adjacent thinking blocks on signature (redacted + normal keep their own signatures)", () => { + it("splits adjacent thinking blocks on differing fidelity (redacted + normal keep their own)", () => { const { messages } = translateEvents([ - // A redacted block: sentinel text + signature arrive together (Claude content_block_start). + // A redacted block: sentinel text + fidelity arrive together (Claude content_block_start). ev({ - content_items: [{ type: "thinking", thinking: "_REDACTED_THINKING", signature: "sig-red" }], + content_items: [ + { type: "thinking", thinking: "_REDACTED_THINKING", fidelity: { signature: "sig-red" } }, + ], }), // The next, ordinary thinking block. ev({ content_items: [{ type: "thinking", thinking: "visible" }] }), - ev({ content_items: [{ type: "thinking", thinking: "", signature: "sig-vis" }] }), + ev({ + content_items: [{ type: "thinking", thinking: "", fidelity: { signature: "sig-vis" } }], + }), ev({ event_type: "stop", content_items: [], finish_reason: "stop" }), ]); const thinkings = complete(messages).filter( @@ -1373,49 +1413,125 @@ describe("provider fidelity fields (signature / phase)", () => { ); expect( thinkings.map((m) => { - const p = m.payload as { thinking: string; signature?: string }; - return [p.thinking, p.signature]; + const p = m.payload as { thinking: string; fidelity?: Record }; + return [p.thinking, p.fidelity]; }), ).toEqual([ - ["_REDACTED_THINKING", "sig-red"], - ["visible", "sig-vis"], + ["_REDACTED_THINKING", { signature: "sig-red" }], + ["visible", { signature: "sig-vis" }], ]); }); - it("emits an empty-text thinking with signature (GPT-5 encrypted reasoning)", () => { + it("keeps a run of equal fidelity as one thinking block (OpenAI-compatible reasoning_field per delta)", () => { + const rf = { reasoning_field: "reasoning_content" }; const { messages } = translateEvents([ - ev({ content_items: [{ type: "thinking", thinking: "", signature: '{"id":"rs_1"}' }] }), + ev({ content_items: [{ type: "thinking", thinking: "step 1, ", fidelity: { ...rf } }] }), + ev({ content_items: [{ type: "thinking", thinking: "step 2, ", fidelity: { ...rf } }] }), + ev({ content_items: [{ type: "thinking", thinking: "done", fidelity: { ...rf } }] }), ev({ content_items: [{ type: "text", text: "answer" }] }), ev({ event_type: "stop", content_items: [], finish_reason: "stop" }), ]); - const thinking = complete(messages).find( + const thinkings = complete(messages).filter( (m) => (m.payload as { type: string }).type === "thinking", - )!; - expect((thinking.payload as { thinking: string }).thinking).toBe(""); - expect((thinking.payload as { signature?: string }).signature).toBe('{"id":"rs_1"}'); + ); + expect( + thinkings.map((m) => { + const p = m.payload as { thinking: string; fidelity?: Record }; + return [p.thinking, p.fidelity]; + }), + ).toEqual([["step 1, step 2, done", rf]]); }); - it("splits text segments on phase markers arriving as empty-text deltas (GPT-5)", () => { + it("emits an empty-text thinking with fidelity (GPT-5 encrypted reasoning) and splits on the next one", () => { const { messages } = translateEvents([ - ev({ content_items: [{ type: "text", text: "", phase: "planning" }] }), + ev({ + content_items: [ + { type: "thinking", thinking: "", fidelity: { id: "rs_1", encrypted_content: "aaa" } }, + ], + }), + ev({ + content_items: [ + { type: "thinking", thinking: "", fidelity: { id: "rs_2", encrypted_content: "bbb" } }, + ], + }), + ev({ content_items: [{ type: "text", text: "answer" }] }), + ev({ event_type: "stop", content_items: [], finish_reason: "stop" }), + ]); + const thinkings = complete(messages).filter( + (m) => (m.payload as { type: string }).type === "thinking", + ); + expect( + thinkings.map((m) => { + const p = m.payload as { thinking: string; fidelity?: Record }; + return [p.thinking, p.fidelity]; + }), + ).toEqual([ + ["", { id: "rs_1", encrypted_content: "aaa" }], + ["", { id: "rs_2", encrypted_content: "bbb" }], + ]); + }); + + it("splits text segments on fidelity.phase markers arriving as empty-text deltas (GPT-5)", () => { + const { messages } = translateEvents([ + ev({ content_items: [{ type: "text", text: "", fidelity: { phase: "planning" } }] }), ev({ content_items: [{ type: "text", text: "plan..." }] }), - ev({ content_items: [{ type: "text", text: "", phase: "answer" }] }), + ev({ content_items: [{ type: "text", text: "", fidelity: { phase: "answer" } }] }), ev({ content_items: [{ type: "text", text: "final" }] }), ev({ event_type: "stop", content_items: [], finish_reason: "stop" }), ]); const texts = complete(messages).filter((m) => (m.payload as { type: string }).type === "text"); expect( texts.map((m) => { - const p = m.payload as { text: string; phase?: string | null }; - return [p.text, p.phase]; + const p = m.payload as { text: string; fidelity?: Record }; + return [p.text, p.fidelity]; }), ).toEqual([ - ["plan...", "planning"], - ["final", "answer"], + ["plan...", { phase: "planning" }], + ["final", { phase: "answer" }], ]); }); - it("carries the tool_call signature through to the complete message", () => { + it("closes a text segment on fidelity.signature: later text becomes its own segment (the signature must not cover unsigned text)", () => { + const { messages } = translateEvents([ + // Gemini stamps a thoughtSignature on the text part it signed; whatever follows is + // unsigned. Merging them would replay a signature covering text the provider never + // signed, and the provider rejects the resumed turn. + ev({ content_items: [{ type: "text", text: "part one", fidelity: { signature: "sigA" } }] }), + ev({ content_items: [{ type: "text", text: "part two" }] }), + ev({ event_type: "stop", content_items: [], finish_reason: "stop" }), + ]); + const texts = complete(messages).filter((m) => (m.payload as { type: string }).type === "text"); + expect( + texts.map((m) => { + const p = m.payload as { text: string; fidelity?: Record }; + return [p.text, p.fidelity]; + }), + ).toEqual([ + ["part one", { signature: "sigA" }], + ["part two", undefined], + ]); + }); + + it("accumulates fidelity keys across deltas of one text segment (phase marker + trailing signature)", () => { + const { messages } = translateEvents([ + // GPT-5 opens the segment with a bare phase marker and signs it only at the end; a + // replacing (rather than merging) assignment would drop the phase and lose the + // segmentation on replay. + ev({ content_items: [{ type: "text", text: "", fidelity: { phase: "answer" } }] }), + ev({ content_items: [{ type: "text", text: "final" }] }), + ev({ content_items: [{ type: "text", text: "", fidelity: { signature: "sigB" } }] }), + ev({ event_type: "stop", content_items: [], finish_reason: "stop" }), + ]); + const texts = complete(messages).filter((m) => (m.payload as { type: string }).type === "text"); + expect( + texts.map((m) => { + const p = m.payload as { text: string; fidelity?: Record }; + return [p.text, p.fidelity]; + }), + ).toEqual([["final", { phase: "answer", signature: "sigB" }]]); + }); + + it("carries the tool_call fidelity through to the complete message", () => { const { messages } = translateEvents([ ev({ content_items: [ @@ -1424,7 +1540,7 @@ describe("provider fidelity fields (signature / phase)", () => { name: "exec_command", arguments: { cmd: "ls" }, tool_call_id: "tc1", - signature: "sig-tool", + fidelity: { signature: "sig-tool" }, }, ], }), @@ -1433,32 +1549,42 @@ describe("provider fidelity fields (signature / phase)", () => { const tc = complete(messages).find( (m) => (m.payload as { type: string }).type === "tool_call", )!; - expect((tc.payload as { signature?: string }).signature).toBe("sig-tool"); + expect((tc.payload as { fidelity?: Record }).fidelity).toEqual({ + signature: "sig-tool", + }); }); - it("round-trips fidelity fields back to UniMessage content items (setHistory path)", () => { + it("round-trips fidelity payloads back to UniMessage content items (setHistory path)", () => { const uni = mergeOmniToUniMessage([ thinkingMessage("deep", "completed", { signature: "sig-1" }), assistantText("hi", "completed", { phase: "answer", signature: "sig-2" }), - toolCall({ name: "t", arguments: "{}", toolCallId: "tc1", signature: "sig-3" }), + toolCall({ name: "t", arguments: "{}", toolCallId: "tc1", fidelity: { signature: "sig-3" } }), ]); expect(uni.content_items).toEqual([ - { type: "thinking", thinking: "deep", signature: "sig-1" }, - { type: "text", text: "hi", phase: "answer", signature: "sig-2" }, - { type: "tool_call", name: "t", arguments: {}, tool_call_id: "tc1", signature: "sig-3" }, + { type: "thinking", thinking: "deep", fidelity: { signature: "sig-1" } }, + { type: "text", text: "hi", fidelity: { phase: "answer", signature: "sig-2" } }, + { + type: "tool_call", + name: "t", + arguments: {}, + tool_call_id: "tc1", + fidelity: { signature: "sig-3" }, + }, ]); }); }); -describe("flushText signature parity (PR #39 review)", () => { - it("emits an empty-text message carrying a text signature instead of dropping it", () => { +describe("flushText fidelity parity (PR #39 review)", () => { + it("emits an empty-text message carrying a text fidelity instead of dropping it", () => { const { messages } = translateEvents([ - ev({ content_items: [{ type: "text", text: "", signature: "sig-t" }] }), + ev({ content_items: [{ type: "text", text: "", fidelity: { signature: "sig-t" } }] }), ev({ event_type: "stop", content_items: [], finish_reason: "stop" }), ]); const text = messages.find((m) => (m.payload as { type: string }).type === "text")!; expect((text.payload as { text: string }).text).toBe(""); - expect((text.payload as { signature?: string }).signature).toBe("sig-t"); + expect((text.payload as { fidelity?: Record }).fidelity).toEqual({ + signature: "sig-t", + }); }); }); diff --git a/packages/core/test/model-catalog.test.ts b/packages/core/test/model-catalog.test.ts index 2e5e82b..eeafe6a 100644 --- a/packages/core/test/model-catalog.test.ts +++ b/packages/core/test/model-catalog.test.ts @@ -6,24 +6,31 @@ import { describe, expect, it } from "vitest"; import { MODEL_CATALOG, MODEL_PROVIDERS, + modelHomepageUrl, catalogEntryFor, - inferProviderForUpstream, presetModelEntries, providerInfo, resolveModelEnv, } from "../src/state/index.js"; describe("model-catalog", () => { - it("model ids are globally unique; DeepSeek comes first (the default model's provider)", () => { + it("(provider, model_id) pairs are unique; DeepSeek comes first (the default model's provider)", () => { + // Bare model ids may repeat across providers (a gateway reselling a vendor model keeps the + // vendor's upstream id, e.g. Qwen Token Plan's glm-5.2 / deepseek-v4-pro) — uniqueness is + // the (provider, model_id) pair, matching the catalog's sole lookup key (catalogEntryFor). + const pairs = MODEL_CATALOG.map((m) => `${m.provider}\0${m.modelId}`); + expect(new Set(pairs).size).toBe(pairs.length); const ids = MODEL_CATALOG.map((m) => m.modelId); - expect(new Set(ids).size).toBe(ids.length); expect(MODEL_CATALOG[0]!.provider).toBe("deepseek"); - // Group order: DeepSeek first, followed by the OpenRouter and SiliconFlow gateways, - // then Google Gemini before Anthropic, with custom last. + // Group order: DeepSeek first, followed by the OpenRouter, SiliconFlow, and Qwen Token + // Plan gateways, then Google Gemini before Anthropic, with custom last. expect(MODEL_PROVIDERS.map((p) => p.id)).toEqual([ "deepseek", "openrouter", + "fireworks", "siliconflow", + "qwen-token-plan", + "qwen-pay-as-you-go", "google", "anthropic", "openai", @@ -61,13 +68,26 @@ describe("model-catalog", () => { } }); - it("all three price buckets are positive; context_window is a positive integer", () => { + it("price buckets are positive (preview models without a list price omit pricing); context_window is a positive integer", () => { for (const m of MODEL_CATALOG) { - expect(m.pricing, m.modelId).toBeDefined(); - expect(m.pricing!.unit).toBe("usd_per_mtok"); - expect(m.pricing!.cache_read).toBeGreaterThan(0); - expect(m.pricing!.cache_write).toBeGreaterThan(0); - expect(m.pricing!.output).toBeGreaterThan(0); + if (m.provider === "qwen-token-plan" && m.modelId === "qwen3.8-max-preview") { + // Preview-only model: the plan runs a quota-multiplier promotion and publishes no + // per-token list price, so the entry carries none and costs read as 0 (same as + // unpriced user models). + expect(m.pricing, m.modelId).toBeUndefined(); + } else if (m.modelId.endsWith(":free")) { + // Free-tier gateway model: a genuine $0 price (not "unknown"), so costs compute to 0. + expect(m.pricing, m.modelId).toBeDefined(); + expect([m.pricing!.cache_read, m.pricing!.cache_write, m.pricing!.output]).toEqual([ + 0, 0, 0, + ]); + } else { + expect(m.pricing, m.modelId).toBeDefined(); + expect(m.pricing!.unit).toBe("usd_per_mtok"); + expect(m.pricing!.cache_read).toBeGreaterThan(0); + expect(m.pricing!.cache_write).toBeGreaterThan(0); + expect(m.pricing!.output).toBeGreaterThan(0); + } expect(Number.isInteger(m.contextWindow)).toBe(true); expect(m.contextWindow!).toBeGreaterThan(0); } @@ -78,9 +98,12 @@ describe("model-catalog", () => { expect(providerInfo("nonexistent")).toBeUndefined(); }); - it("pair matching and provider-grouping inference: catalogEntryFor / inferProviderForUpstream", () => { - // catalogEntryFor is the sole catalog lookup entry point: it matches on (group, upstream id) - // pairs, so an identically named upstream id never matches across the wrong group. + it("catalogEntryFor is the sole lookup and always takes the (provider, model_id) pair", () => { + // It matches on (group, upstream id) pairs, so an identically named upstream id never + // matches across the wrong group. There is no bare-id lookup at all: a gateway reselling a + // vendor model keeps the vendor's upstream id, so a bare id names no single catalog entry + // and the catalog never offers to pick one (`glm-5.2`, `qwen3.7-max`, `qwen3.7-plus` and + // `deepseek-v4-pro` each appear under two groups). expect(catalogEntryFor("anthropic", "claude-sonnet-4-6")?.displayName).toBe( "Claude Sonnet 4.6", ); @@ -88,11 +111,11 @@ describe("model-catalog", () => { // The upstream id itself may contain / (gateway models); it is never split apart. expect(catalogEntryFor("openrouter", "xiaomi/mimo-v2.5")?.displayName).toBe("MiMo-V2.5"); expect(catalogEntryFor("custom", "my-own")).toBeUndefined(); - - // Inference for `model add` when --provider is omitted: a catalog hit yields its provider, otherwise custom. - expect(inferProviderForUpstream("deepseek-v4-pro")).toBe("deepseek"); - expect(inferProviderForUpstream("xiaomi/mimo-v2.5")).toBe("openrouter"); - expect(inferProviderForUpstream("my-own-model")).toBe("custom"); + // Each group's entry for a resold id is reached only through that group. + expect(catalogEntryFor("zhipu", "glm-5.2")?.contextWindow).toBe(1000000); + expect(catalogEntryFor("qwen-token-plan", "glm-5.2")?.contextWindow).toBe(1048576); + expect(catalogEntryFor("deepseek", "deepseek-v4-pro")?.provider).toBe("deepseek"); + expect(catalogEntryFor("qwen-token-plan", "deepseek-v4-pro")?.provider).toBe("qwen-token-plan"); }); it("presetModelEntries: provider and bare upstream model_id are separate fields; gateway models inline base_url", () => { @@ -115,42 +138,124 @@ describe("model-catalog", () => { } }); - it("gateway models (OpenRouter / SiliconFlow): openai protocol + preset base URL; env fallback is OPENAI_API_KEY", () => { + it("gateway models (OpenRouter / SiliconFlow / Qwen Token Plan): openai protocol + preset base URL; env fallback is OPENAI_API_KEY", () => { const or = MODEL_CATALOG.filter((m) => m.provider === "openrouter"); + // Dictionary order, newer versions of a series first (gpt-5.6-* before gpt-5.5, + // opus-4.8 before 4.7) — precomputed in the catalog, no runtime sorting. expect(or.map((m) => m.modelId)).toEqual([ - "xiaomi/mimo-v2.5", - "tencent/hy3", + "anthropic/claude-fable-5", + "anthropic/claude-opus-4.8", + "anthropic/claude-opus-4.7", + "anthropic/claude-sonnet-5", + "deepseek/deepseek-v4-flash", + "deepseek/deepseek-v4-pro", + "google/gemini-3.5-flash", "minimax/minimax-m3", + "moonshotai/kimi-k3", + "nvidia/nemotron-3-ultra-550b-a55b:free", + "openai/gpt-5.6-sol", + "openai/gpt-5.6-terra", + "openai/gpt-5.5", "stepfun/step-3.7-flash", + "tencent/hy3", + "x-ai/grok-4.5", + "xiaomi/mimo-v2.5", + "z-ai/glm-5.2", ]); for (const m of or) { expect(m.clientType).toBe("openai"); expect(m.baseUrl).toBe("https://openrouter.ai/api/v1"); } + const fw = MODEL_CATALOG.filter((m) => m.provider === "fireworks"); + expect(fw.map((m) => [m.modelId, m.supportsVision])).toEqual([ + ["accounts/fireworks/models/deepseek-v4-flash", false], + ["accounts/fireworks/models/deepseek-v4-pro", false], + ["accounts/fireworks/models/glm-5p2", false], + ["accounts/fireworks/models/kimi-k2p7-code", true], + ["accounts/fireworks/models/minimax-m3", true], + ]); + for (const m of fw) { + expect(m.clientType).toBe("openai"); + expect(m.baseUrl).toBe("https://api.fireworks.ai/inference/v1"); + } const sf = MODEL_CATALOG.filter((m) => m.provider === "siliconflow"); expect(sf.map((m) => m.modelId)).toEqual([ - "zai-org/GLM-5.2", + "deepseek-ai/DeepSeek-V4-Flash", "deepseek-ai/DeepSeek-V4-Pro", "meituan-longcat/LongCat-2.0", + "moonshotai/Kimi-K2.7-Code", + "zai-org/GLM-5.2", ]); for (const m of sf) { expect(m.clientType).toBe("openai"); expect(m.baseUrl).toBe("https://api.siliconflow.cn/v1"); } + const qtp = MODEL_CATALOG.filter((m) => m.provider === "qwen-token-plan"); + expect(qtp.map((m) => m.modelId)).toEqual([ + "deepseek-v4-pro", + "glm-5.2", + "qwen3.8-max-preview", + "qwen3.7-max", + "qwen3.7-plus", + ]); + for (const m of qtp) { + expect(m.clientType).toBe("openai"); + expect(m.baseUrl).toBe("https://token-plan.cn-beijing.maas.aliyuncs.com/compatible-mode/v1"); + } + // Vision flags per the plan's supported-model table: 3.8-max-preview and 3.7-plus see images. + expect(qtp.map((m) => [m.modelId, m.supportsVision])).toEqual([ + ["deepseek-v4-pro", false], + ["glm-5.2", false], + ["qwen3.8-max-preview", true], + ["qwen3.7-max", false], + ["qwen3.7-plus", true], + ]); + const qpayg = MODEL_CATALOG.filter((m) => m.provider === "qwen-pay-as-you-go"); + expect(qpayg.map((m) => [m.modelId, m.supportsVision])).toEqual([ + ["kimi/kimi-k3", true], + ["qwen3.7-max", false], + ["qwen3.7-plus", true], + ["ZHIPU/GLM-5.2", false], + ]); + for (const m of qpayg) { + expect(m.clientType).toBe("openai"); + expect(m.baseUrl).toBe("https://dashscope.aliyuncs.com/compatible-mode/v1"); + } // Routed through AgentHub's OpenAI client -> when the credential is left blank it reads OPENAI_API_KEY (not the provider's own env var name). - for (const id of ["openrouter", "siliconflow", "custom"]) { + for (const id of [ + "openrouter", + "fireworks", + "siliconflow", + "qwen-token-plan", + "qwen-pay-as-you-go", + "custom", + ]) { expect(providerInfo(id)!.envKey).toBe("OPENAI_API_KEY"); expect(providerInfo(id)!.envBaseUrlKey).toBe("OPENAI_BASE_URL"); } - // gatewayBaseUrl (prefilled by group in the frontend's "add model" dialog) is only carried by the two gateway providers. + // gatewayBaseUrl (prefilled by group in the frontend's "add model" dialog) is only carried by the gateway providers. expect(providerInfo("openrouter")!.gatewayBaseUrl).toBe("https://openrouter.ai/api/v1"); expect(providerInfo("siliconflow")!.gatewayBaseUrl).toBe("https://api.siliconflow.cn/v1"); + expect(providerInfo("qwen-token-plan")!.gatewayBaseUrl).toBe( + "https://token-plan.cn-beijing.maas.aliyuncs.com/compatible-mode/v1", + ); + expect(providerInfo("qwen-pay-as-you-go")!.gatewayBaseUrl).toBe( + "https://dashscope.aliyuncs.com/compatible-mode/v1", + ); + expect(providerInfo("fireworks")!.gatewayBaseUrl).toBe("https://api.fireworks.ai/inference/v1"); + const GATEWAYS = [ + "openrouter", + "fireworks", + "siliconflow", + "qwen-token-plan", + "qwen-pay-as-you-go", + ]; for (const p of MODEL_PROVIDERS) { - if (p.id !== "openrouter" && p.id !== "siliconflow") { + if (!GATEWAYS.includes(p.id)) { expect(p.gatewayBaseUrl, p.id).toBeUndefined(); } } - const gateway = [...or, ...sf]; + const gateway = [...or, ...fw, ...sf, ...qtp, ...qpayg]; // Pricing (USD): MiMo v2.5 and Hy3. const mimo = MODEL_CATALOG.find((m) => m.modelId === "xiaomi/mimo-v2.5")!.pricing!; expect([mimo.cache_read, mimo.cache_write, mimo.output]).toEqual([0.0028, 0.14, 0.28]); @@ -209,4 +314,47 @@ describe("resolveModelEnv (PRN-021: env fallback resolved by AgentHub routing ru } } }); + + it("modelHomepageUrl: gateway per-model pages, vendor docs fallback, none for custom groups", () => { + // Gateway URL patterns work for user-added ids in those groups too (not catalog-gated). + expect(modelHomepageUrl("openrouter", "anthropic/claude-fable-5")).toBe( + "https://openrouter.ai/anthropic/claude-fable-5", + ); + expect(modelHomepageUrl("openrouter", "someone/new-model")).toBe( + "https://openrouter.ai/someone/new-model", + ); + expect(modelHomepageUrl("qwen-token-plan", "qwen3.7-plus")).toBe( + "https://www.qianwenai.com/models/qwen3.7-plus", + ); + // Fireworks maps the accounts//models/ API id to its page path; other ids + // fall back to the models listing. + expect(modelHomepageUrl("fireworks", "accounts/fireworks/models/glm-5p2")).toBe( + "https://app.fireworks.ai/models/fireworks/glm-5p2", + ); + expect(modelHomepageUrl("fireworks", "my-own-id")).toBe("https://app.fireworks.ai/models"); + // Pay-as-you-go resells third-party models under slash-prefixed ids: the id is URL-encoded. + expect(modelHomepageUrl("qwen-pay-as-you-go", "ZHIPU/GLM-5.2")).toBe( + "https://www.qianwenai.com/models/ZHIPU%2FGLM-5.2", + ); + // The preview model has no dedicated page: falls back to the plan's model overview. + expect(modelHomepageUrl("qwen-token-plan", "qwen3.8-max-preview")).toBe( + providerInfo("qwen-token-plan")!.modelsUrl, + ); + // Direct vendors link to the vendor's model docs page. + expect(modelHomepageUrl("deepseek", "deepseek-v4-pro")).toBe( + "https://api-docs.deepseek.com/quick_start/pricing", + ); + // Z.AI and Moonshot have per-model pages (Moonshot drops the dot: kimi-k2.6 -> chat-k26). + expect(modelHomepageUrl("zhipu", "glm-5.2")).toBe("https://docs.z.ai/guides/llm/glm-5.2"); + expect(modelHomepageUrl("moonshot", "kimi-k2.6")).toBe( + "https://platform.kimi.com/docs/pricing/chat-k26", + ); + expect(modelHomepageUrl("moonshot", "kimi-k2.5")).toBe( + "https://platform.kimi.com/docs/pricing/chat-k25", + ); + expect(modelHomepageUrl("moonshot", "my-own")).toBe("https://platform.kimi.com/docs/pricing"); + // Custom and user-defined groups have no page to vouch for. + expect(modelHomepageUrl("custom", "my-model")).toBeUndefined(); + expect(modelHomepageUrl("my-own-gateway", "x")).toBeUndefined(); + }); }); diff --git a/packages/core/test/session-title.test.ts b/packages/core/test/session-title.test.ts index 960da86..2ef18aa 100644 --- a/packages/core/test/session-title.test.ts +++ b/packages/core/test/session-title.test.ts @@ -10,6 +10,7 @@ import { generateTitleWithLLM, sanitizeTitle, Session, + stripConversationMarkers, thinkingMessage, tokenUsage, userText, @@ -119,6 +120,24 @@ describe("session-title", () => { expect(sanitizeTitle("『标题』!")).toBe("标题"); expect(sanitizeTitle(" \n ")).toBeNull(); expect(sanitizeTitle("x".repeat(50))).toHaveLength(30); + // A leaked block is stripped from the model output. + expect(sanitizeTitle("\nskills: web-design\n\n构建落地页")).toBe( + "构建落地页", + ); + }); + + it("stripConversationMarkers: removes machine marker blocks, keeps the human body", () => { + // The skill-invocation block that wraps a first user message must not reach the title. + expect( + stripConversationMarkers( + "\nskills: penguin-sdk, web-design\n\n做一个 RAG 应用", + ), + ).toBe("做一个 RAG 应用"); + // Handoff and scheduled-task markers are stripped too; ordinary angle-bracket text stays. + expect(stripConversationMarkers("data_analyst继续分析")).toBe( + "继续分析", + ); + expect(stripConversationMarkers("render a
element")).toBe("render a
element"); }); it("Session.generateTitle: sends via createBareLLM; returns null when no factory is provided", async () => { @@ -162,6 +181,10 @@ describe("session-title", () => { // Material = the first Task's user text + model text (thinking does not count), matching // buildTitlePrompt's shape. expect(seen[0]).toBe(buildTitlePrompt("user question", "answer body")); + // Anti-CoT shape: an explicit no-thinking rule, and the prompt ends with an empty think + // block so reasoning models treat their thinking phase as already closed. + expect(seen[0]).toContain("do not think aloud"); + expect(seen[0]!.endsWith("")).toBe(true); // No request is sent when no material has been collected (run was never called). const idle = new Session({ diff --git a/packages/core/test/state.test.ts b/packages/core/test/state.test.ts index 540a9d3..bf5f6cf 100644 --- a/packages/core/test/state.test.ts +++ b/packages/core/test/state.test.ts @@ -41,6 +41,7 @@ import { skillsDir, systemConfigPath, toolsDir, + type ModelRef, type ProjectConfig, } from "../src/state/index.js"; import { sessionEnvironment } from "../src/internal/session-support.js"; @@ -288,6 +289,8 @@ describe("assembleSystemPrompt", () => { sessionEnvironment("/tmp/penguin-ws", "session-test-1", { agentId: DEFAULT_AGENT_ID, projectDir: "/tmp/proj", + provider: "deepseek", + modelId: "deepseek-v4-pro", }), ); expect(prompt).toContain("AGENTS.md"); @@ -341,6 +344,8 @@ describe("assembleSystemPrompt", () => { cwd: "/tmp/ws", agentId: "agent-x", projectDir: "/tmp/proj", + provider: "deepseek", + modelId: "deepseek-v4-pro", platform: "darwin", osVersion: "Darwin 25.0.0", date: "2026-06-30", @@ -393,6 +398,8 @@ describe("assembleSystemPrompt", () => { cwd: "/tmp/ws", agentId: "agent-x", projectDir: "/tmp/proj", + provider: "deepseek", + modelId: "deepseek-v4-pro", platform: "darwin", osVersion: "Darwin 25.0.0", date: "2026-06-30", @@ -466,7 +473,7 @@ describe("assembleSystemPrompt", () => { const env = sessionEnvironment( "/tmp/penguin-ws", "session-test-1", - { agentId: "agent-x", projectDir: "/tmp/proj" }, + { agentId: "agent-x", projectDir: "/tmp/proj", provider: "openai", modelId: "gpt-5.5" }, new Date("2026-06-30T00:00:00"), ); const prompt = assembleSystemPrompt(state, env); @@ -476,15 +483,19 @@ describe("assembleSystemPrompt", () => { expect(prompt).toContain("CWD: /tmp/penguin-ws"); expect(prompt).toContain("Agent ID: agent-x"); expect(prompt).toContain("Project Dir: /tmp/proj"); + expect(prompt).toContain("Provider: openai"); + expect(prompt).toContain("Model ID: gpt-5.5"); expect(prompt).toContain("Platform:"); expect(prompt).toContain("OS Version:"); expect(prompt).toContain("Date: 2026-06-30"); expect(prompt.indexOf("Platform:")).toBeLessThan(prompt.indexOf("OS Version:")); expect(prompt.indexOf("OS Version:")).toBeLessThan(prompt.indexOf("Date:")); - expect(prompt.indexOf("Date:")).toBeLessThan(prompt.indexOf("CWD:")); - expect(prompt.indexOf("CWD:")).toBeLessThan(prompt.indexOf("Agent ID:")); - expect(prompt.indexOf("Agent ID:")).toBeLessThan(prompt.indexOf("Project Dir:")); - expect(prompt.indexOf("Project Dir:")).toBeLessThan(prompt.indexOf("Session ID:")); + expect(prompt.indexOf("Date:")).toBeLessThan(prompt.indexOf("Project Dir:")); + expect(prompt.indexOf("Project Dir:")).toBeLessThan(prompt.indexOf("Agent ID:")); + expect(prompt.indexOf("Agent ID:")).toBeLessThan(prompt.indexOf("CWD:")); + expect(prompt.indexOf("CWD:")).toBeLessThan(prompt.indexOf("Provider:")); + expect(prompt.indexOf("Provider:")).toBeLessThan(prompt.indexOf("Model ID:")); + expect(prompt.indexOf("Model ID:")).toBeLessThan(prompt.indexOf("Session ID:")); }); }); @@ -531,15 +542,41 @@ describe("project-config round trip", () => { expect(getModel(loaded, { provider: "custom", model_id: "unknown-model" })).toBeUndefined(); }); - it("infers the provider from the builtin catalog when addModel omits it", async () => { - await addModel(tmpRoot, DEFAULT_PROJECT_ID, { model_id: "claude-sonnet-4-6" }); - await addModel(tmpRoot, DEFAULT_PROJECT_ID, { model_id: "my-own-model" }); + it("addModel files the entry under the provider it was given, never one of its own choosing", async () => { + // provider is a required field: nothing is inferred from the builtin catalog, so a model + // outside the known groups is filed under custom only because the caller said so. glm-5.2 + // is sold by both the Qwen Token Plan gateway and Zhipu — the caller names which one, and + // the entry (with its api_key) lands in exactly that group. + await addModel(tmpRoot, DEFAULT_PROJECT_ID, { + provider: "anthropic", + model_id: "claude-sonnet-4-6", + }); + await addModel(tmpRoot, DEFAULT_PROJECT_ID, { provider: "custom", model_id: "my-own-model" }); + await addModel(tmpRoot, DEFAULT_PROJECT_ID, { + provider: "zhipu", + model_id: "glm-5.2", + api_key: "sk-zhipu", + }); const loaded = await loadProjectConfig(tmpRoot, DEFAULT_PROJECT_ID); - // A catalog hit gets its provider; otherwise it falls back to custom. expect( getModel(loaded, { provider: "anthropic", model_id: "claude-sonnet-4-6" }), ).toBeDefined(); expect(getModel(loaded, { provider: "custom", model_id: "my-own-model" })).toBeDefined(); + expect(getModel(loaded, { provider: "zhipu", model_id: "glm-5.2" })?.api_key).toBe("sk-zhipu"); + // The key never leaks into the other group that resells the same bare id. + expect( + getModel(loaded, { provider: "qwen-token-plan", model_id: "glm-5.2" })?.api_key, + ).toBeUndefined(); + }); + + it("addModel requires an explicit provider: a bare model_id does not type-check", () => { + // Compile-time contract (asserted by `pnpm typecheck`, which includes this file): with the + // catalog inference gone there is nothing to fall back to, so the entry's provider field is + // required rather than optional. vitest only checks that the call expression exists. + const bare = { model_id: "glm-5.2" }; + // @ts-expect-error provider is required: a model reference is always a (provider, model_id) pair. + const call = (): Promise => addModel(tmpRoot, DEFAULT_PROJECT_ID, bare); + expect(call).toBeTypeOf("function"); }); it("upserts by the (provider, model_id) pair; same model_id under two providers co-exists", async () => { @@ -706,7 +743,10 @@ describe("project-config round trip", () => { (c) => c.provider === entry.provider && c.modelId === entry.model_id, )!; expect(entry.vision).toBe(cat.supportsVision ? undefined : false); - expect(entry.pricing?.unit).toBe("usd_per_mtok"); + // A catalog entry without a list price (the Token Plan preview model) presets no + // pricing; every other catalog entry stores USD pricing. + if (cat.pricing === undefined) expect(entry.pricing).toBeUndefined(); + else expect(entry.pricing?.unit).toBe("usd_per_mtok"); // A model that auto-routes leaves client_type unset; a gateway model (OpenRouter) // explicitly sets it to openai. expect(entry.client_type).toBe(cat.clientType); @@ -800,7 +840,7 @@ describe("project-config round trip", () => { }); }); -describe('resolveModelRef (the single "provider omitted" resolution entry point)', () => { +describe("resolveModelRef (validates a (provider, model_id) pair against the config)", () => { const cfg: ProjectConfig = { models: [ { provider: "deepseek", model_id: "deepseek-v4-pro" }, @@ -809,37 +849,44 @@ describe('resolveModelRef (the single "provider omitted" resolution entry point) ], }; - it("resolves a globally unique model_id without provider (exact match only)", () => { - expect(resolveModelRef(cfg, "deepseek-v4-pro")).toEqual({ + it("returns the pair when it names a configured entry", () => { + expect(resolveModelRef(cfg, "deepseek-v4-pro", "deepseek")).toEqual({ provider: "deepseek", model_id: "deepseek-v4-pro", }); - // Exact-match lookup, no fuzzy/prefix matching. - expect(() => resolveModelRef(cfg, "deepseek-v4")).toThrow(/is not in the Project config/); - }); - - it("validates the exact pair when provider is given", () => { + // The same bare model_id under two groups is never ambiguous: each pair names its own entry. expect(resolveModelRef(cfg, "shared-id", "siliconflow")).toEqual({ provider: "siliconflow", model_id: "shared-id", }); - // A wrong provider grouping likewise fails: the error includes the pair reference. + expect(resolveModelRef(cfg, "shared-id", "openrouter")).toEqual({ + provider: "openrouter", + model_id: "shared-id", + }); + }); + + it("throws when the pair is not configured; the error carries the pair reference", () => { + // Wrong group for a configured model_id: no fallback to "the entry that happens to have + // this id" — a pair the config doesn't have simply isn't a model. expect(() => resolveModelRef(cfg, "shared-id", "openai")).toThrow( - /\(provider=openai, model_id=shared-id\)/, + /is not in the Project config.*\(provider=openai, model_id=shared-id\)/, + ); + // Unknown model_id, and exact matching only (no fuzzy/prefix matching). + expect(() => resolveModelRef(cfg, "no-such-model", "deepseek")).toThrow( + /\(provider=deepseek, model_id=no-such-model\)/, + ); + expect(() => resolveModelRef(cfg, "deepseek-v4", "deepseek")).toThrow( + /is not in the Project config/, ); }); - it("errors clearly on zero matches", () => { - expect(() => resolveModelRef(cfg, "no-such-model")).toThrow( - /is not in the Project config.*no-such-model/, - ); - }); - - it("errors on ambiguity, listing the candidate pair references", () => { - expect(() => resolveModelRef(cfg, "shared-id")).toThrow(/Ambiguous/); - expect(() => resolveModelRef(cfg, "shared-id")).toThrow( - /\(provider=siliconflow, model_id=shared-id\).*\(provider=openrouter, model_id=shared-id\)/, - ); + it("requires provider: a bare model_id does not type-check (no resolution path left)", () => { + // The pair is enforced by the type checker — asserted by `pnpm typecheck`, which includes + // this file; the unused-directive error is the failure mode if the parameter ever goes + // optional again. vitest only checks that the call expression exists. + // @ts-expect-error provider is required: a model reference is always a (provider, model_id) pair. + const call = (): ModelRef => resolveModelRef(cfg, "deepseek-v4-pro"); + expect(call).toBeTypeOf("function"); }); }); diff --git a/packages/core/test/subagent.test.ts b/packages/core/test/subagent.test.ts index c89f486..3089025 100644 --- a/packages/core/test/subagent.test.ts +++ b/packages/core/test/subagent.test.ts @@ -56,7 +56,7 @@ type RunInput = { prompt: string; signal?: AbortSignal; approve?: ApproveFn }; /** Builds a SubagentRunner from a run implementation (spawn arguments observed via a spy). */ function runnerOf( run: (input: RunInput) => AsyncGenerator, - spawnSpy?: (input: { agentId?: string; modelId?: string }) => void, + spawnSpy?: (input: { agentId?: string; modelId?: string; provider?: string }) => void, ): SubagentRunner { return { async spawn(input) { @@ -136,7 +136,8 @@ afterEach(() => { describe("run_subagent tool (foreground)", () => { it("forwards stamped child messages and mirrors child text as its own output deltas", async () => { - const seen: Array<{ prompt?: string; agentId?: string; modelId?: string }> = []; + const seen: Array<{ prompt?: string; agentId?: string; modelId?: string; provider?: string }> = + []; const runner = runnerOf( async function* (input) { seen[0] = { ...seen[0], prompt: input.prompt }; @@ -150,7 +151,10 @@ describe("run_subagent tool (foreground)", () => { const { services } = makeServices(runner); const tool = createSubagentTool(DEF, services); const { out, result } = await collectWithReturn( - tool.execute({ prompt: "world", agent_id: "researcher", model_id: "m1" }, CTX), + tool.execute( + { prompt: "world", agent_id: "researcher", model_id: "m1", provider: "p1" }, + CTX, + ), ); // Child session messages pass through verbatim (with origin). @@ -163,7 +167,49 @@ describe("run_subagent tool (foreground)", () => { expect(result?.stopReason).toBe("completed"); // The model is free to choose the agent and model (spawn arguments); the prompt is // handed to run. - expect(seen[0]).toEqual({ prompt: "world", agentId: "researcher", modelId: "m1" }); + expect(seen[0]).toEqual({ + prompt: "world", + agentId: "researcher", + modelId: "m1", + provider: "p1", + }); + }); + + it("rejects half a model reference in either direction, and never spawns", async () => { + // A model is referenced by the (provider, model_id) pair; the tool refuses half of one + // rather than letting the session layer guess a group for a bare upstream id. + for (const args of [ + { prompt: "x", model_id: "m1" }, + { prompt: "x", provider: "p1" }, + ]) { + const spawned: unknown[] = []; + const runner = runnerOf( + async function* () {}, + (input) => spawned.push(input), + ); + const { services } = makeServices(runner); + const tool = createSubagentTool(DEF, services); + const { out, result } = await collectWithReturn(tool.execute(args, CTX)); + expect(result?.stopReason).toBe("failed"); + expect(ownDeltas(out)).toContain("must be given together"); + expect(spawned).toHaveLength(0); + } + }); + + it("uses the Project default model when neither half of the reference is given", async () => { + const seen: Array<{ modelId?: string; provider?: string }> = []; + const runner = runnerOf( + async function* () { + yield withOrigin(partialText("delta", "ok"), HOP); + }, + (input) => seen.push(input), + ); + const { services } = makeServices(runner); + const tool = createSubagentTool(DEF, services); + const { result } = await collectWithReturn(tool.execute({ prompt: "x" }, CTX)); + expect(result?.stopReason).toBe("completed"); + expect(seen[0]?.modelId).toBeUndefined(); + expect(seen[0]?.provider).toBeUndefined(); }); it("does not mirror deeper-nested (origin.length > 1) text into its own output", async () => { diff --git a/packages/docs/content/cli.en.md b/packages/docs/content/cli.en.md index 6e94071..6d2a674 100644 --- a/packages/docs/content/cli.en.md +++ b/packages/docs/content/cli.en.md @@ -7,7 +7,7 @@ The CLI ships as the npm package `@prismshadow/penguin-cli`; the command is `pen ## Global conventions -- Model references: a model's identity is always the `(provider, model_id)` pair. `--model-id` takes the upstream model id and pairs with `--provider`. When `run` / `chat` omit `--provider`, the `--model-id` matches only if it is globally unique in the configuration; ambiguity is an error. +- Model references: a model's identity is always the `(provider, model_id)` pair. `--model-id` takes the upstream model id and `--provider` the group it belongs to; the provider is never inferred, guessed, or defaulted. On `run` / `chat` the pair as a whole is optional — pass both to pick a model, or neither to use the Project's default model — but passing one without the other is an error. - Data root: `--root ` overrides the data root directory. Priority: `--root` > the `PENGUIN_HOME` env var > `~/.penguin/data`. ## penguin run @@ -21,8 +21,8 @@ penguin run -m "Summarize the code structure of this directory" | Option | Description | | --- | --- | | `-m, --message ` | Required; the message to send | -| `--model-id ` | Model to use; defaults to the Project's default model | -| `--provider ` | Provider group of the model | +| `--model-id ` | Upstream id of the model to use; requires `--provider`. Omit both to use the Project's default model | +| `--provider ` | Provider group of the model; required whenever `--model-id` is given | | `--project-id ` | Project to use | | `--agent-id ` | Agent to use | | `--workspace ` | Workspace directory; defaults to the current directory and must exist | @@ -74,13 +74,13 @@ Manages a Project's model configuration, per-Agent vault environment variables, Add or update a model entry: ```bash -penguin config model add --model-id deepseek-v4-pro --api-key sk-... --set-default +penguin config model add --provider deepseek --model-id deepseek-v4-pro --api-key sk-... --set-default ``` | Option | Description | | --- | --- | | `--model-id ` | Required; the upstream model id | -| `--provider ` | Provider group; inferred from the built-in catalog when omitted | +| `--provider ` | Required; the provider group the entry belongs to. It is never derived from the model id: gateways resell vendor models under their upstream ids, so a guessed group would write the credential onto another vendor's endpoint. Use `custom` for any endpoint outside the built-in groups. | | `--api-key ` | API key, stored inline in the Project's hidden `.project_config.toml` | | `--base-url ` | Custom endpoint base URL | | `--context-window ` | Context window size | diff --git a/packages/docs/content/cli.zh.md b/packages/docs/content/cli.zh.md index 2241400..163b9f8 100644 --- a/packages/docs/content/cli.zh.md +++ b/packages/docs/content/cli.zh.md @@ -7,7 +7,7 @@ CLI 由 npm 包 `@prismshadow/penguin-cli` 提供,命令为 `penguin`。不带 ## 全局约定 -- 模型引用:模型身份始终是 `(provider, model_id)` 二元组。`--model-id` 填上游模型 id,与 `--provider` 组成配对引用。`run` / `chat` 省略 `--provider` 时,仅当该 `--model-id` 在配置中全局唯一才会匹配,存在歧义则报错。 +- 模型引用:模型身份始终是 `(provider, model_id)` 二元组。`--model-id` 填上游模型 id,`--provider` 填其所属分组;provider 绝不推断、绝不猜测、也没有缺省值。`run` / `chat` 上这对参数整体可选——两个都给即指定模型,两个都不给则使用 Project 默认模型——但只给其中一个是错误。 - 数据根目录:`--root ` 覆盖数据根目录,优先级为 `--root` > 环境变量 `PENGUIN_HOME` > `~/.penguin/data`。 ## penguin run @@ -21,8 +21,8 @@ penguin run -m "总结当前目录的代码结构" | 选项 | 说明 | | --- | --- | | `-m, --message ` | 必填,要发送的消息 | -| `--model-id ` | 指定模型,缺省使用 Project 默认模型 | -| `--provider ` | 模型所属 Provider 分组 | +| `--model-id ` | 指定模型的上游 id,须与 `--provider` 同时给出;两者都不给时使用 Project 默认模型 | +| `--provider ` | 模型所属 Provider 分组,给出 `--model-id` 时必填 | | `--project-id ` | 指定 Project | | `--agent-id ` | 指定 Agent | | `--workspace ` | Workspace 目录,默认当前目录,必须已存在 | @@ -74,13 +74,13 @@ Ctrl-C 的行为依状态而定: 新增或更新模型条目: ```bash -penguin config model add --model-id deepseek-v4-pro --api-key sk-... --set-default +penguin config model add --provider deepseek --model-id deepseek-v4-pro --api-key sk-... --set-default ``` | 选项 | 说明 | | --- | --- | | `--model-id ` | 必填,上游模型 id | -| `--provider ` | Provider 分组,缺省时根据内置目录推断 | +| `--provider ` | 必填,条目所属的 Provider 分组。它绝不由模型 id 推导:网关会以上游 id 转售厂商模型,猜错分组会把凭据写到另一家厂商的接口上。内置分组之外的接口一律用 `custom`。 | | `--api-key ` | API Key,内联存入 Project 隐藏文件 `.project_config.toml` | | `--base-url ` | 自定义接口地址 | | `--context-window ` | 上下文窗口大小 | diff --git a/packages/docs/content/configuration.en.md b/packages/docs/content/configuration.en.md index 909b049..036c781 100644 --- a/packages/docs/content/configuration.en.md +++ b/packages/docs/content/configuration.en.md @@ -35,7 +35,7 @@ The openrouter, siliconflow, and custom groups speak the OpenAI-compatible proto ## Project config -`//.project_config.toml` is the Project's single config file: a hidden file written with mode 0600, with credentials inlined on the model entries. Model identity is always the `(provider, model_id)` pair — string concatenation is forbidden everywhere. +`//.project_config.toml` is the Project's single config file: a hidden file written with mode 0600, with credentials inlined on the model entries. Model identity is always the `(provider, model_id)` pair — string concatenation is forbidden everywhere, and every reference into this file carries both halves: the provider is never inferred from a bare `model_id`. | Field | Description | | --- | --- | @@ -143,9 +143,11 @@ compaction: | `{{PLATFORM}}` | Runtime platform | | `{{OS_VERSION}}` | Operating system version | | `{{DATE}}` | Current date | -| `{{CWD}}` | Workspace path | -| `{{AGENT_ID}}` | Agent id | | `{{PROJECT_DIR}}` | Project directory | +| `{{AGENT_ID}}` | Agent id | +| `{{CWD}}` | Workspace path | +| `{{PROVIDER}}` | Model provider group | +| `{{MODEL_ID}}` | Upstream model id | | `{{SESSION_ID}}` | Session id | `agent_state/AGENTS.md` is the developer-editable instruction file, injected via `{{AGENTS_MD}}` and empty by default — it is also the file an optimizer edits most (see [Self-Improvement](/self-improvement)). @@ -157,6 +159,7 @@ compaction: - Key names must match `^[A-Za-z_][A-Za-z0-9_]*$` (shell environment variable naming rules); - Values are injected only into tool subprocess environments and never enter the model context or the Trace; - Only key names are disclosed in the system prompt via `{{VAULT_KEYS}}`; +- Saving through the Web/API invalidates the Agent's cached Session runtimes: the next Task on any of its Sessions re-resumes and runs with the new values; a Task already in flight keeps the values it started with (a direct CLI file edit reaches a running server only when a Session is next created or resumed); - Managed via `penguin config vault set/list/remove` or the Web Vault tab. ## Schedules @@ -172,7 +175,7 @@ Each file `agent_state/schedule/.toml` describes one scheduled task (the f | `end_at` | no | End time; must be later than `start_at` | | `session_id` | no | Bind to an existing Session; mutually exclusive with the three fields below | | `workspace` | no | Workspace for new-Session mode | -| `provider` / `model_id` | no | Paired model reference for new-Session mode | +| `provider` / `model_id` | no | Paired model reference for new-Session mode; write both or neither — a lone `model_id` is rejected, and with neither the Project's default model is used | ```toml prompt = "Check yesterday's builds and summarize the failures" diff --git a/packages/docs/content/configuration.zh.md b/packages/docs/content/configuration.zh.md index f24e2d0..f16bd84 100644 --- a/packages/docs/content/configuration.zh.md +++ b/packages/docs/content/configuration.zh.md @@ -35,7 +35,7 @@ openrouter、siliconflow 与 custom 分组走 OpenAI 兼容协议,因此复用 ## Project 配置 -`//.project_config.toml` 是 Project 唯一的配置文件:隐藏文件,落盘权限 0600,凭证内联在模型条目上。模型身份始终是 `(provider, model_id)` 成对引用,禁止任何形式的字符串拼接。 +`//.project_config.toml` 是 Project 唯一的配置文件:隐藏文件,落盘权限 0600,凭证内联在模型条目上。模型身份始终是 `(provider, model_id)` 成对引用,禁止任何形式的字符串拼接;指向本文件的每一处引用都要带上两半,provider 绝不由裸 `model_id` 推断。 | 字段 | 说明 | | --- | --- | @@ -143,9 +143,11 @@ compaction: | `{{PLATFORM}}` | 运行平台 | | `{{OS_VERSION}}` | 操作系统版本 | | `{{DATE}}` | 当前日期 | -| `{{CWD}}` | Workspace 路径 | -| `{{AGENT_ID}}` | Agent id | | `{{PROJECT_DIR}}` | Project 目录 | +| `{{AGENT_ID}}` | Agent id | +| `{{CWD}}` | Workspace 路径 | +| `{{PROVIDER}}` | 模型 provider 分组 | +| `{{MODEL_ID}}` | 上游模型 id | | `{{SESSION_ID}}` | Session id | `agent_state/AGENTS.md` 是开发者可编辑的指令文件,经 `{{AGENTS_MD}}` 注入系统提示词,缺省为空——它也是优化器最常改动的文件(见[自我进化](/self-improvement))。 @@ -157,6 +159,7 @@ compaction: - 键名须匹配 `^[A-Za-z_][A-Za-z0-9_]*$`(shell 环境变量命名规则); - 值只注入工具子进程的环境变量,永远不进入模型上下文与 Trace; - 系统提示词中经 `{{VAULT_KEYS}}` 只披露键名; +- 经 Web/API 保存会使该 Agent 已缓存的 Session 运行时失效:其任意 Session 的下一个任务会重新恢复(resume)并使用新值;进行中的任务保持其启动时的值(CLI 直接改文件对运行中的 server 则要等 Session 下次创建或恢复时生效); - 通过 CLI `penguin config vault set/list/remove` 或 Web 的 Vault 标签页管理。 ## 定时任务 @@ -172,7 +175,7 @@ compaction: | `end_at` | 否 | 结束时刻,须晚于 `start_at` | | `session_id` | 否 | 绑定既有 Session;与下列三项互斥 | | `workspace` | 否 | 新建 Session 模式的 Workspace | -| `provider` / `model_id` | 否 | 新建 Session 模式的模型成对引用 | +| `provider` / `model_id` | 否 | 新建 Session 模式的模型成对引用;要写就两个都写,只写 `model_id` 会被拒绝,两个都不写则使用 Project 默认模型 | ```toml prompt = "检查昨日构建结果并汇总失败原因" diff --git a/packages/docs/content/installation.en.md b/packages/docs/content/installation.en.md index 8f2bebe..390519b 100644 --- a/packages/docs/content/installation.en.md +++ b/packages/docs/content/installation.en.md @@ -13,7 +13,7 @@ description: Install PenguinHarness via the install script, npm, or from source. On Linux / macOS: ```bash -curl -fsSL https://github.com/Prism-Shadow/penguin-harness/releases/latest/download/install.sh | sh +curl -fsSL https://penguin.ooo/install.sh | sh ``` The script downloads the matching `penguin-{linux,darwin}-{x64,arm64}.tar.gz`, which bundles an official Node.js runtime. Other platforms do **not** fall back automatically: the script exits and asks you to install Node.js >= 24 and re-run with `--universal`, which selects the runtime-less `penguin-universal.tar.gz`. @@ -34,7 +34,7 @@ penguin -v | Integrity check | Downloads are sha256-verified when the Release ships checksum assets | | Upgrade | Re-run the install script; files are swapped atomically | -Script flags are passed as `curl ... | sh -s -- --universal`. +Script flags go after `sh -s --`, e.g. `curl -fsSL https://penguin.ooo/install.sh | sh -s -- --universal`. ### Data directory diff --git a/packages/docs/content/installation.zh.md b/packages/docs/content/installation.zh.md index d7fc68b..c48c003 100644 --- a/packages/docs/content/installation.zh.md +++ b/packages/docs/content/installation.zh.md @@ -13,7 +13,7 @@ description: 通过安装脚本、npm 或源码安装 PenguinHarness。 在 Linux / macOS 上执行: ```bash -curl -fsSL https://github.com/Prism-Shadow/penguin-harness/releases/latest/download/install.sh | sh +curl -fsSL https://penguin.ooo/install.sh | sh ``` 脚本按平台下载 `penguin-{linux,darwin}-{x64,arm64}.tar.gz`,其中捆绑了官方 Node.js 运行时。其他平台**不会自动回退**:脚本会退出并提示先安装 Node.js >= 24、再携带 `--universal` 重新执行,改用不含运行时的 `penguin-universal.tar.gz`。 @@ -34,7 +34,7 @@ penguin -v | 完整性校验 | Release 提供 checksum 资产时自动进行 sha256 校验 | | 升级 | 重新执行安装脚本即可,文件原子替换 | -脚本参数通过 `curl ... | sh -s -- --universal` 的形式传入。 +脚本参数写在 `sh -s --` 之后,例如 `curl -fsSL https://penguin.ooo/install.sh | sh -s -- --universal`。 ### 数据目录 diff --git a/packages/docs/content/interfaces.en.md b/packages/docs/content/interfaces.en.md index 7ceb61c..d732428 100644 --- a/packages/docs/content/interfaces.en.md +++ b/packages/docs/content/interfaces.en.md @@ -89,7 +89,7 @@ interface GenerativeModelConfig { `GenerativeModel` (`packages/core/src/llm/generative-model.ts`) grounds the contract on the `AutoLLMClient` of the `@prismshadow/agenthub` model gateway: - the gateway maintains conversation history **statefully**, receiving only new messages each turn; resuming a Session replays committed history through a one-time `setHistory`; -- an internal `EventTranslator` translates gateway stream events into `partial_*` fragments plus complete messages, preserving the `signature` / `phase` fidelity fields; complete messages settle in thinking → text → tool_call order; +- an internal `EventTranslator` translates gateway stream events into `partial_*` fragments plus complete messages, preserving each item's opaque `fidelity` payload verbatim; segmentation mirrors the gateway's own aggregation — a thinking block is closed by its fidelity payload and a run of equal fidelity stays one block (OpenAI-compatible clients stamp every delta with the same `{ reasoning_field }`, which must not split blocks), while a text segment splits on a differing `fidelity.phase` and closes on a `fidelity.signature`, fidelity keys accumulating on merge; complete messages settle in thinking → text → tool_call order; - `ToolCallIdAllocator` disambiguates providers that use the function name as the call id (append `#n` inbound, strip outbound), scoped to the whole Session; - provider differences (tool-call formats, reasoning content, streaming events) are absorbed entirely inside the gateway — see [Models & Providers](/models). @@ -169,7 +169,7 @@ A tool emits only content deltas; framing, timeouts, truncation, `stop_reason` p Human is deliberately not an interface class. The SDK caller *is* the Human: ```ts -const session = await agent.createSession({ workspaceDir, modelId }); +const session = await agent.createSession({ workspaceDir, provider, modelId }); session.run( newMessages: OmniMessage[], // input: the Prompt diff --git a/packages/docs/content/interfaces.zh.md b/packages/docs/content/interfaces.zh.md index 03fab56..5478052 100644 --- a/packages/docs/content/interfaces.zh.md +++ b/packages/docs/content/interfaces.zh.md @@ -89,7 +89,7 @@ interface GenerativeModelConfig { `GenerativeModel`(`packages/core/src/llm/generative-model.ts`)把契约落到模型网关 `@prismshadow/agenthub` 的 `AutoLLMClient` 上: - 网关**有状态**地维护会话历史,每轮只接收新消息;恢复 Session 时经一次性的 `setHistory` 重放已提交历史; -- 内部的 `EventTranslator` 把网关流式事件翻译为 `partial_*` 分片 + 完整消息,保留 `signature` / `phase` 保真字段,完整消息按 thinking → text → tool_call 顺序落盘; +- 内部的 `EventTranslator` 把网关流式事件翻译为 `partial_*` 分片 + 完整消息,逐条原样保留不透明的 `fidelity` 保真负载;分段与网关自身的聚合一致——thinking 块由其 fidelity 负载闭合,连续相同的 fidelity 归为同一块(OpenAI 兼容客户端给每条增量盖同一个 `{ reasoning_field }`,不能因此切块),text 段遇到不同的 `fidelity.phase` 即切分、遇到 `fidelity.signature` 即闭合,合并时 fidelity 键累积;完整消息按 thinking → text → tool_call 顺序落盘; - `ToolCallIdAllocator` 处理个别 Provider 用函数名充当调用 id 的情况(入站追加 `#n`、出站剥离),作用域覆盖整个 Session; - Provider 协议差异(工具调用格式、思考内容、流式事件)全部在网关内抹平,见[模型与 Provider](/models)。 @@ -169,7 +169,7 @@ interface ToolDefinitionConfig { Human 刻意不设计为接口类。SDK 的调用方就是 Human: ```ts -const session = await agent.createSession({ workspaceDir, modelId }); +const session = await agent.createSession({ workspaceDir, provider, modelId }); session.run( newMessages: OmniMessage[], // 输入:Prompt diff --git a/packages/docs/content/models.en.md b/packages/docs/content/models.en.md index 418cfce..64ba372 100644 --- a/packages/docs/content/models.en.md +++ b/packages/docs/content/models.en.md @@ -11,6 +11,8 @@ All model access goes through one gateway library: `@prismshadow/agenthub` (Auto A model's identity is always the `(provider, model_id)` pair: `provider` is a config group name, `model_id` the upstream request id sent to AgentHub unchanged. The two are independent fields — concatenating them into one string is forbidden anywhere in the pipeline. +Every interface that names a model takes the complete pair: the CLI, the HTTP API, and the SDK all reject half a reference instead of completing it. The provider is never inferred from the model id and has no default, because gateways resell vendor models under their upstream ids — a guessed group would send the entry's credential to a vendor nobody named. Where a model reference is optional at all (`penguin run` / `chat`, Session creation, Schedules), the choice is between the whole pair and nothing: omit both halves to take the Project's default model. + ## The per-Project model table Each Project's available models are recorded in the hidden `.project_config.toml`, maintained via the CLI (`penguin config model add / default / list`, see [CLI Reference](/cli)) or the Web UI — never hand-edited. `ModelEntry` fields: @@ -57,7 +59,10 @@ Built-in groups and their env-var fallbacks (catalog source: `packages/core/src/ | --- | --- | --- | | deepseek | `DEEPSEEK_API_KEY` | Group of the default model | | openrouter | `OPENAI_API_KEY` | OpenAI-compatible gateway, preset base URL `https://openrouter.ai/api/v1` | +| fireworks | `OPENAI_API_KEY` | Fireworks AI (OpenAI-compatible), preset base URL `https://api.fireworks.ai/inference/v1`; API model ids look like `accounts/fireworks/models/` | | siliconflow | `OPENAI_API_KEY` | OpenAI-compatible gateway, preset base URL `https://api.siliconflow.cn/v1` | +| qwen-token-plan | `OPENAI_API_KEY` | Qwen Token Plan subscription gateway, preset base URL `https://token-plan.cn-beijing.maas.aliyuncs.com/compatible-mode/v1`; pricing from each model page's official list price (the preview model has only a quota-multiplier promo, no list price) | +| qwen-pay-as-you-go | `OPENAI_API_KEY` | Qwen pay-as-you-go (DashScope's OpenAI-compatible endpoint), preset base URL `https://dashscope.aliyuncs.com/compatible-mode/v1`; resold third-party models keep vendor-prefixed ids (e.g. `kimi/kimi-k3`) | | google | `GEMINI_API_KEY` | | | anthropic | `ANTHROPIC_API_KEY` | | | openai | `OPENAI_API_KEY` | | @@ -65,9 +70,9 @@ Built-in groups and their env-var fallbacks (catalog source: `packages/core/src/ | moonshot | `MOONSHOT_API_KEY` | | | custom | `OPENAI_API_KEY` | Any OpenAI-protocol endpoint | -The gateway groups (openrouter / siliconflow) go through AgentHub's OpenAI client, so with blank credentials they read `OPENAI_API_KEY` — not a gateway-specific variable. +The gateway groups (openrouter / fireworks / siliconflow / qwen-token-plan / qwen-pay-as-you-go) go through AgentHub's OpenAI client, so with blank credentials they read `OPENAI_API_KEY` — not a gateway-specific variable. -Some models in the preset catalog: deepseek-v4-pro / deepseek-v4-flash, gemini-3.1-pro-preview, claude-opus-4-8 / claude-sonnet-4-6, gpt-5.5, glm-5.2, kimi-k2.6 (not exhaustive). +Some models in the preset catalog: deepseek-v4-pro / deepseek-v4-flash, gemini-3.1-pro-preview, claude-opus-4-8 / claude-sonnet-4-6, gpt-5.5, glm-5.2, kimi-k2.6, qwen3.8-max-preview (not exhaustive). ## Thinking levels diff --git a/packages/docs/content/models.zh.md b/packages/docs/content/models.zh.md index db2c46e..ef8eacc 100644 --- a/packages/docs/content/models.zh.md +++ b/packages/docs/content/models.zh.md @@ -11,6 +11,8 @@ description: 经 AgentHub 单一网关接入模型,以 (provider, model_id) 模型身份永远是 `(provider, model_id)` 成对表示:`provider` 是配置分组名,`model_id` 是原样发给上游的请求 id。二者是两个独立字段,任何环节都不允许拼接成一个字符串。 +所有涉及模型的接口都要求完整的二元组:CLI、HTTP API 与 SDK 都会拒绝半个引用,而不会替你补全。provider 绝不由模型 id 推断,也没有缺省值——网关会以上游 id 转售厂商模型,猜出来的分组会把该条目的凭据发往无人指定的厂商。凡是模型引用本身可省略之处(`penguin run` / `chat`、创建 Session、定时任务),可选的是整对:两半都省略即使用 Project 默认模型。 + ## Project 模型表 每个 Project 的可用模型记录在隐藏文件 `.project_config.toml` 中,由 CLI(`penguin config model add / default / list`,见 [CLI 参考](/cli))或 Web 界面维护,不手工编辑。`ModelEntry` 字段: @@ -57,7 +59,10 @@ api_key = "sk-..." | --- | --- | --- | | deepseek | `DEEPSEEK_API_KEY` | 默认模型所在分组 | | openrouter | `OPENAI_API_KEY` | OpenAI 兼容网关,预置 base URL `https://openrouter.ai/api/v1` | +| fireworks | `OPENAI_API_KEY` | Fireworks AI(OpenAI 兼容),预置 base URL `https://api.fireworks.ai/inference/v1`;API 模型 id 形如 `accounts/fireworks/models/` | | siliconflow | `OPENAI_API_KEY` | OpenAI 兼容网关,预置 base URL `https://api.siliconflow.cn/v1` | +| qwen-token-plan | `OPENAI_API_KEY` | Qwen Token Plan 订阅网关,预置 base URL `https://token-plan.cn-beijing.maas.aliyuncs.com/compatible-mode/v1`;定价取各模型页官方牌价(预览模型仅配额倍率促销、无牌价) | +| qwen-pay-as-you-go | `OPENAI_API_KEY` | Qwen 按量付费(DashScope OpenAI 兼容端),预置 base URL `https://dashscope.aliyuncs.com/compatible-mode/v1`;转售第三方模型保留厂商前缀 id(如 `kimi/kimi-k3`) | | google | `GEMINI_API_KEY` | | | anthropic | `ANTHROPIC_API_KEY` | | | openai | `OPENAI_API_KEY` | | @@ -65,9 +70,9 @@ api_key = "sk-..." | moonshot | `MOONSHOT_API_KEY` | | | custom | `OPENAI_API_KEY` | 任意 OpenAI 协议端点 | -网关分组(openrouter / siliconflow)经 AgentHub 的 OpenAI 客户端请求,因此凭证留空时读取的是 `OPENAI_API_KEY`,而非网关自己的变量名。 +网关分组(openrouter / fireworks / siliconflow / qwen-token-plan / qwen-pay-as-you-go)经 AgentHub 的 OpenAI 客户端请求,因此凭证留空时读取的是 `OPENAI_API_KEY`,而非网关自己的变量名。 -预置目录中的部分模型:deepseek-v4-pro / deepseek-v4-flash、gemini-3.1-pro-preview、claude-opus-4-8 / claude-sonnet-4-6、gpt-5.5、glm-5.2、kimi-k2.6 等(非完整清单)。 +预置目录中的部分模型:deepseek-v4-pro / deepseek-v4-flash、gemini-3.1-pro-preview、claude-opus-4-8 / claude-sonnet-4-6、gpt-5.5、glm-5.2、kimi-k2.6、qwen3.8-max-preview 等(非完整清单)。 ## 思考等级 diff --git a/packages/docs/content/omni-message.en.md b/packages/docs/content/omni-message.en.md index b54265d..83d7194 100644 --- a/packages/docs/content/omni-message.en.md +++ b/packages/docs/content/omni-message.en.md @@ -54,15 +54,16 @@ On resume, the engine takes this Trace line as the runtime config — the model, ## model_msg: complete payloads -Seven content payloads, discriminated by `payload.type`. Shared optional fields: `stop_reason` (marks an abnormal terminal state) and `signature` (a provider-fidelity field, see below): +Seven content payloads, discriminated by `payload.type`. Shared optional fields: `stop_reason` (marks an abnormal terminal state) and `fidelity` (an opaque provider-fidelity payload, see below): ```ts +type Fidelity = Record; // opaque provider-fidelity payload (see below) + interface TextPayload { type: "text"; role: "user" | "assistant"; text: string; - phase?: string | null; // segmentation marker (e.g. GPT-5 phases) - signature?: string; + fidelity?: Fidelity; // e.g. { phase } segment marker (GPT-5), { signature } stop_reason?: StopReason; } @@ -70,7 +71,7 @@ interface ThinkingPayload { type: "thinking"; role: "assistant"; thinking: string; - signature?: string; // required by some models to replay history + fidelity?: Fidelity; // required by some models to replay history stop_reason?: StopReason; } @@ -79,7 +80,7 @@ interface InlineThinkingPayload { role: "assistant"; data: string; // reasoning content in binary form mime_type: string; - signature?: string; + fidelity?: Fidelity; stop_reason?: StopReason; } @@ -89,7 +90,7 @@ interface ToolCallPayload { name: string; arguments: string; // arguments as a JSON string tool_call_id: string; - signature?: string; + fidelity?: Fidelity; stop_reason?: StopReason; } @@ -114,7 +115,7 @@ interface InlineDataPayload { role: "user" | "assistant"; data: string; // other binary content mime_type: string; - signature?: string; + fidelity?: Fidelity; stop_reason?: StopReason; } ``` @@ -273,7 +274,20 @@ Errors never cross an interface boundary as exceptions — they *are* messages. ## Provider-fidelity fields -Provider-specific fields such as `signature` and `phase` pass through and persist verbatim end to end — some models require them byte-for-byte when history is replayed, and any rewriting would break compatibility. This is one of the preconditions for lossless Session recovery from the Trace. +Provider-specific wire data travels in a single optional field, `fidelity` — an arbitrary JSON object the LLM client records to reproduce the original message on replay: thinking signatures, `phase` segment labels, GPT-5 encrypted reasoning, the OpenAI-compatible upstream reasoning field name: + +```ts +// Claude: a thinking block closed by its signature +{ type: "thinking", thinking: "…", fidelity: { signature: "EqQBCkYIBxgCKkB…" } } + +// GPT-5: encrypted reasoning (empty thinking text, fidelity only) +{ type: "thinking", thinking: "", fidelity: { id: "rs_0d3…", encrypted_content: "gAAAA…" } } + +// OpenAI-compatible: the upstream field the reasoning text came from +{ type: "thinking", thinking: "…", fidelity: { reasoning_field: "reasoning_content" } } +``` + +The payload is opaque to PenguinHarness: it passes through and persists verbatim end to end — some models require it byte-for-byte when history is replayed, and any rewriting (or loss) would break compatibility. This is one of the preconditions for lossless Session recovery from the Trace. ## Three jobs, one protocol diff --git a/packages/docs/content/omni-message.zh.md b/packages/docs/content/omni-message.zh.md index 5f659c2..53e538b 100644 --- a/packages/docs/content/omni-message.zh.md +++ b/packages/docs/content/omni-message.zh.md @@ -54,15 +54,16 @@ interface ToolDefinition { ## model_msg:完整消息 -七种内容 payload,以 `payload.type` 判别。公共可选字段:`stop_reason`(非正常收尾时标注终态)与 `signature`(Provider 保真字段,见下文): +七种内容 payload,以 `payload.type` 判别。公共可选字段:`stop_reason`(非正常收尾时标注终态)与 `fidelity`(不透明的 Provider 保真负载,见下文): ```ts +type Fidelity = Record; // 不透明的 Provider 保真负载(见下文) + interface TextPayload { type: "text"; role: "user" | "assistant"; text: string; - phase?: string | null; // 分段标记(如 GPT-5 的 phase) - signature?: string; + fidelity?: Fidelity; // 如 { phase } 分段标记(GPT-5)、{ signature } stop_reason?: StopReason; } @@ -70,7 +71,7 @@ interface ThinkingPayload { type: "thinking"; role: "assistant"; thinking: string; - signature?: string; // 部分模型历史回放所必需 + fidelity?: Fidelity; // 部分模型历史回放所必需 stop_reason?: StopReason; } @@ -79,7 +80,7 @@ interface InlineThinkingPayload { role: "assistant"; data: string; // 二进制形态的思考内容 mime_type: string; - signature?: string; + fidelity?: Fidelity; stop_reason?: StopReason; } @@ -89,7 +90,7 @@ interface ToolCallPayload { name: string; arguments: string; // 参数 JSON 字符串 tool_call_id: string; - signature?: string; + fidelity?: Fidelity; stop_reason?: StopReason; } @@ -114,7 +115,7 @@ interface InlineDataPayload { role: "user" | "assistant"; data: string; // 其他二进制内容 mime_type: string; - signature?: string; + fidelity?: Fidelity; stop_reason?: StopReason; } ``` @@ -272,7 +273,20 @@ type StopReason = "completed" | "failed" | "aborted" | "timeout" | "malformed"; ## 保真字段 -`signature` 与 `phase` 等 Provider 专有字段在整条链路上原样透传、原样存储——部分模型在历史回放时要求逐字一致,任何转写都会破坏兼容性。这是 Trace 能够无损恢复 Session 的前提之一。 +Provider 专有的线上数据统一收拢在一个可选字段 `fidelity` 中——LLM 客户端为历史回放记录的任意 JSON 对象:思考签名、`phase` 分段标记、GPT-5 加密推理、OpenAI 兼容上游的推理字段名: + +```ts +// Claude:由签名闭合的 thinking 块 +{ type: "thinking", thinking: "…", fidelity: { signature: "EqQBCkYIBxgCKkB…" } } + +// GPT-5:加密推理(thinking 文本为空,仅有 fidelity) +{ type: "thinking", thinking: "", fidelity: { id: "rs_0d3…", encrypted_content: "gAAAA…" } } + +// OpenAI 兼容:思考内容来自上游哪个字段 +{ type: "thinking", thinking: "…", fidelity: { reasoning_field: "reasoning_content" } } +``` + +该负载对 PenguinHarness 完全不透明:在整条链路上原样透传、原样存储——部分模型在历史回放时要求逐字一致,任何转写或丢失都会破坏兼容性。这是 Trace 能够无损恢复 Session 的前提之一。 ## 协议的三种职责 diff --git a/packages/docs/content/quickstart.en.md b/packages/docs/content/quickstart.en.md index f5b6342..6df4b3c 100644 --- a/packages/docs/content/quickstart.en.md +++ b/packages/docs/content/quickstart.en.md @@ -8,7 +8,7 @@ description: Install PenguinHarness, configure a model, and run your first Task. One-liner for Linux / macOS: ```bash -curl -fsSL https://github.com/Prism-Shadow/penguin-harness/releases/latest/download/install.sh | sh +curl -fsSL https://penguin.ooo/install.sh | sh ``` For other options (npm, from source), see [Installation](/installation). @@ -18,10 +18,10 @@ For other options (npm, from source), see [Installation](/installation). PenguinHarness ships with no built-in model credentials, so configure a model first. Use the Models page in the Web UI, or the CLI: ```bash -penguin config model add --model-id deepseek-v4-pro --api-key sk-... --set-default +penguin config model add --provider deepseek --model-id deepseek-v4-pro --api-key sk-... --set-default ``` -- When `--provider` is omitted, the Provider is inferred from the built-in catalog. +- A model is always referenced as a `(provider, model_id)` pair, so `--provider` and `--model-id` are both required — the Provider is never inferred from the model id. See [Models & Providers](/models) for the built-in groups. - The API key can also come from environment variables: when a model entry has no inline api_key, AgentHub (the LLM gateway library) reads variables such as `DEEPSEEK_API_KEY`, `ANTHROPIC_API_KEY`, `OPENAI_API_KEY`, and `GEMINI_API_KEY`. A `.env` file in the working directory is loaded automatically. ## Start the Web App @@ -30,7 +30,7 @@ penguin config model add --model-id deepseek-v4-pro --api-key sk-... --set-defau penguin web ``` -The service runs at http://127.0.0.1:7364 and opens your browser (`--no-open` to skip). First login is `admin` / `admin123` — change it right away. `penguin server` starts the same process headless. +The service runs at http://127.0.0.1:7364 and opens your browser (`--no-open` to skip). First login is `admin` / `penguin-2026` — change it right away. `penguin server` starts the same process headless. ## One-shot run diff --git a/packages/docs/content/quickstart.zh.md b/packages/docs/content/quickstart.zh.md index c5e28c8..39b0ddd 100644 --- a/packages/docs/content/quickstart.zh.md +++ b/packages/docs/content/quickstart.zh.md @@ -8,7 +8,7 @@ description: 安装 PenguinHarness、配置模型并运行第一个 Task。 Linux / macOS 一键安装: ```bash -curl -fsSL https://github.com/Prism-Shadow/penguin-harness/releases/latest/download/install.sh | sh +curl -fsSL https://penguin.ooo/install.sh | sh ``` 其他方式(npm、源码)见[安装](/installation)。 @@ -18,10 +18,10 @@ curl -fsSL https://github.com/Prism-Shadow/penguin-harness/releases/latest/downl PenguinHarness 不内置任何模型凭据,使用前需要先配置一个模型。可以在 Web UI 的 Models 页面完成,也可以用 CLI: ```bash -penguin config model add --model-id deepseek-v4-pro --api-key sk-... --set-default +penguin config model add --provider deepseek --model-id deepseek-v4-pro --api-key sk-... --set-default ``` -- 省略 `--provider` 时,根据内置目录自动推断 Provider。 +- 模型引用始终是 `(provider, model_id)` 二元组,因此 `--provider` 与 `--model-id` 均为必填——Provider 绝不由模型 id 推断。内置分组见[模型与 Provider](/models)。 - API Key 也可以来自环境变量:当模型条目没有内联 api_key 时,LLM 网关库 AgentHub 会读取 `DEEPSEEK_API_KEY`、`ANTHROPIC_API_KEY`、`OPENAI_API_KEY`、`GEMINI_API_KEY` 等变量;工作目录下的 `.env` 会被自动加载。 ## 启动 Web App @@ -30,7 +30,7 @@ penguin config model add --model-id deepseek-v4-pro --api-key sk-... --set-defau penguin web ``` -服务运行在 http://127.0.0.1:7364 并自动打开浏览器(`--no-open` 跳过)。首次登录使用 `admin` / `admin123`,请立即修改密码。`penguin server` 启动同一进程的 headless 版本。 +服务运行在 http://127.0.0.1:7364 并自动打开浏览器(`--no-open` 跳过)。首次登录使用 `admin` / `penguin-2026`,请立即修改密码。`penguin server` 启动同一进程的 headless 版本。 ## 单次运行 diff --git a/packages/docs/content/server-api.en.md b/packages/docs/content/server-api.en.md index 7182536..485aa17 100644 --- a/packages/docs/content/server-api.en.md +++ b/packages/docs/content/server-api.en.md @@ -35,12 +35,12 @@ packages/server/src - Cookie session: `penguin_session` (HttpOnly, SameSite=Lax), valid for 7 days with sliding renewal; - Passwords are stored as scrypt hashes; the server keeps only the sha256 of the session token, never the plaintext; -- No open registration: the built-in admin `admin` / `admin123` is seeded at startup, and all other accounts are created by an admin; +- No open registration: the built-in admin `admin` / `penguin-2026` is seeded at startup, and all other accounts are created by an admin; - Same-origin only — no CORS middleware is enabled. ```bash curl -c cookies.txt -H "Content-Type: application/json" \ - -d '{"userId":"admin","password":"admin123"}' \ + -d '{"userId":"admin","password":"penguin-2026"}' \ http://127.0.0.1:7364/api/auth/login ``` @@ -87,6 +87,8 @@ Member writes are owner-only. | PUT | /api/projects/:projectId/models | Full-table replace, keyed by `(provider, modelId)` | | POST | /api/projects/:projectId/models/test | Connectivity test: `{provider, modelId, …}` → `{ok, latencyMs?, message?}` | +Every endpoint that names a model takes the complete `(provider, modelId)` pair. Nothing is inferred: a request carrying only one half is a 400, never a lookup. Where the reference itself is optional (Session creation, Schedules), omitting both halves selects the Project's default model. + ### Agents The paths below omit the `/api/projects/:projectId` prefix. @@ -110,7 +112,7 @@ The paths below omit the `/api/projects/:projectId` prefix. | GET / POST | /agents/:agentId/schedules | List scheduled tasks / create one (409 if the name exists) | | GET / PUT / DELETE | /agents/:agentId/schedules/:name | Read / update / delete a single task | -Schedule writes are owner-only. +Schedule writes are owner-only. A task in new-Session mode carries `modelId` and `provider` together or not at all; the pair is checked against the Project's model table when the task is saved and again when the scheduler reconciles it. ### Session Creation and Directory Browsing @@ -120,7 +122,7 @@ Schedule writes are owner-only. | POST | /agents/:agentId/sessions | Create a Session: `{modelId?, provider?, workspace?, approvalMode?}` → 201 | | GET | /dirs?path= | Server-side directory browser (backs the Workspace picker) | -On Session creation, the model defaults to the Project's default model, the Workspace defaults to an auto-created temporary directory, and the approval mode defaults to `allow-all`. +On Session creation, `modelId` and `provider` are both-or-neither: send the complete pair to pick a model, or omit both to take the Project's default model — one without the other is a 400. The Workspace defaults to an auto-created temporary directory, and the approval mode defaults to `allow-all`. ### Usage and Traces (Agent Level) @@ -147,7 +149,7 @@ The paths below omit the `/api/sessions/:sessionId` prefix. For the storage mode | POST | /abort | Interrupt the current Task: 202 when triggered, 204 when idle | | POST | /compact | Trigger context compaction: 202; 409 `nothing_to_compact` when there is nothing to compact | | GET | /files?path= | Browse the Workspace directory | -| GET | /files/content?path=&download= | Read a Workspace file (`download=1` serves it as an attachment) | +| GET | /files/content?path=&download=&preview= | Read a Workspace file (`download=1` serves it as an attachment, `preview=1` renders it in a sandbox — see below) | | POST | /files/stat | Batch existence check: `{paths}` | | PUT | /files/content?path= | Upload a file: `{dataBase64}`, capped at 14MB | | GET | /traces | List this Session's Trace files | @@ -157,6 +159,16 @@ The paths below omit the `/api/sessions/:sessionId` prefix. For the storage mode General conventions: Sessions the user cannot access always return 404 — their existence is never leaked; only one Task or compaction runs per Session at a time, and conflicts return 409 (`task_in_progress` / `compacting`). +Workspace files may be Agent-generated, so `GET /files/content` treats them as untrusted: every response carries `X-Content-Type-Options: nosniff`, and the rest of the headers depend on the two flags (`download=1` wins over `preview=1`): + +| Query | Content-Type | Content-Disposition | Content-Security-Policy | +| --- | --- | --- | --- | +| neither | `text/plain; charset=utf-8` for `.html` / `.htm` / `.svg`, the real type otherwise | `inline` | — | +| `preview=1` | the real type (`text/html`, `image/svg+xml`, …) | `inline` | `sandbox allow-scripts allow-popups allow-modals allow-forms`, sent only for `.html` / `.htm` / `.svg` | +| `download=1` | the real type | `attachment` | — | + +The filename always rides along as `filename*=UTF-8''` with percent-encoding. `preview=1` is what backs "open in a new tab": the document keeps its real type and does render and run, but the sandbox deliberately omits `allow-same-origin`, so it lands in an opaque origin and can reach neither this origin's cookies nor the API — while the request itself still authenticates, because a top-level GET sends the SameSite=Lax session cookie. + Key request bodies (explicit keys): ```ts diff --git a/packages/docs/content/server-api.zh.md b/packages/docs/content/server-api.zh.md index 1982c0f..e091b5c 100644 --- a/packages/docs/content/server-api.zh.md +++ b/packages/docs/content/server-api.zh.md @@ -35,12 +35,12 @@ packages/server/src - Cookie 会话:`penguin_session`(HttpOnly、SameSite=Lax),有效期 7 天,滑动续期; - 密码以 scrypt 哈希存储;服务端只保存会话 Token 的 sha256,不落明文; -- 不开放注册:启动时种子化内置管理员 `admin` / `admin123`,其余账号由管理员创建; +- 不开放注册:启动时种子化内置管理员 `admin` / `penguin-2026`,其余账号由管理员创建; - 仅限同源访问,未启用 CORS 中间件。 ```bash curl -c cookies.txt -H "Content-Type: application/json" \ - -d '{"userId":"admin","password":"admin123"}' \ + -d '{"userId":"admin","password":"penguin-2026"}' \ http://127.0.0.1:7364/api/auth/login ``` @@ -87,6 +87,8 @@ curl -c cookies.txt -H "Content-Type: application/json" \ | PUT | /api/projects/:projectId/models | 全表替换,条目以 `(provider, modelId)` 为键 | | POST | /api/projects/:projectId/models/test | 连通性测试:`{provider, modelId, …}` → `{ok, latencyMs?, message?}` | +所有涉及模型的接口都要求完整的 `(provider, modelId)` 二元组,不做任何推断:只带一半的请求一律 400,绝不会退化为一次查找。模型引用本身可省略的场景(创建 Session、定时任务)省略的是整对,两半都不给即选用 Project 默认模型。 + ### Agent 以下路径均省略前缀 `/api/projects/:projectId`。 @@ -110,7 +112,7 @@ curl -c cookies.txt -H "Content-Type: application/json" \ | GET / POST | /agents/:agentId/schedules | 定时任务列表 / 创建(重名返回 409) | | GET / PUT / DELETE | /agents/:agentId/schedules/:name | 读取 / 更新 / 删除单个任务 | -Schedule 写操作仅限 Owner。 +Schedule 写操作仅限 Owner。新建 Session 模式的任务,`modelId` 与 `provider` 要么成对给出、要么都不给;该二元组会在任务保存时以及调度器对账时对照 Project 模型表校验。 ### Session 创建与目录浏览 @@ -120,7 +122,7 @@ Schedule 写操作仅限 Owner。 | POST | /agents/:agentId/sessions | 创建 Session:`{modelId?, provider?, workspace?, approvalMode?}` → 201 | | GET | /dirs?path= | 服务器端目录浏览(Workspace 选择器数据源) | -创建 Session 时,模型默认取 Project 默认模型,Workspace 默认自动创建临时目录,审批模式默认 `allow-all`。 +创建 Session 时,`modelId` 与 `provider` 要么成对给出、要么都不给:给出完整二元组即指定模型,两个都省略则取 Project 默认模型,只给一个返回 400。Workspace 默认自动创建临时目录,审批模式默认 `allow-all`。 ### 用量与 Trace(Agent 级) @@ -147,7 +149,7 @@ Schedule 写操作仅限 Owner。 | POST | /abort | 中断当前 Task:已触发返回 202,无任务返回 204 | | POST | /compact | 触发上下文压缩:202;无可压缩内容返回 409 `nothing_to_compact` | | GET | /files?path= | 浏览 Workspace 目录 | -| GET | /files/content?path=&download= | 读取 Workspace 文件(`download=1` 时作为附件下载) | +| GET | /files/content?path=&download=&preview= | 读取 Workspace 文件(`download=1` 时作为附件下载,`preview=1` 以沙箱方式预览 —— 见下) | | POST | /files/stat | 批量存在性检查:`{paths}` | | PUT | /files/content?path= | 上传文件:`{dataBase64}`,上限 14MB | | GET | /traces | 本 Session 的 Trace 文件列表 | @@ -157,6 +159,16 @@ Schedule 写操作仅限 Owner。 通用约定:无权访问的 Session 一律返回 404,不泄露其存在性;每个 Session 同时只允许一个 Task 或压缩在运行,冲突时返回 409(`task_in_progress` / `compacting`)。 +Workspace 文件可能由 Agent 生成,`GET /files/content` 一律按不可信内容处理:所有响应都带 `X-Content-Type-Options: nosniff`,其余响应头取决于两个开关(`download=1` 优先于 `preview=1`): + +| 查询参数 | Content-Type | Content-Disposition | Content-Security-Policy | +| --- | --- | --- | --- | +| 都不带 | `.html` / `.htm` / `.svg` 降级为 `text/plain; charset=utf-8`,其余为真实类型 | `inline` | 无 | +| `preview=1` | 真实类型(`text/html`、`image/svg+xml` 等) | `inline` | `sandbox allow-scripts allow-popups allow-modals allow-forms`,仅对 `.html` / `.htm` / `.svg` 下发 | +| `download=1` | 真实类型 | `attachment` | 无 | + +文件名始终以 `filename*=UTF-8''` 形式携带(百分号编码)。`preview=1` 支撑的是“在新标签页打开”:文档保留真实类型,可以正常渲染并执行脚本,但沙箱刻意不含 `allow-same-origin`,因此它落在一个不透明源里,既拿不到本源的 Cookie,也调不动 API;而请求本身仍然要鉴权 —— 顶层 GET 会带上 SameSite=Lax 的会话 Cookie。 + 关键请求体(明确键名): ```ts diff --git a/packages/docs/content/sessions-and-traces.en.md b/packages/docs/content/sessions-and-traces.en.md index a8fad63..9f1290f 100644 --- a/packages/docs/content/sessions-and-traces.en.md +++ b/packages/docs/content/sessions-and-traces.en.md @@ -81,7 +81,7 @@ Special case: if the latest Trace file ends with a completed compaction, that co ## Field fidelity -Provider-specific fields (such as `signature` and `phase`) are preserved verbatim in the Trace and sent back verbatim — some models require them byte-for-byte on history replay, and any rewriting would break compatibility. This is one reason the Trace stores raw OmniMessage envelopes rather than a post-processed format. +Each content message's opaque provider `fidelity` payload (thinking signatures, phase labels, encrypted reasoning, …) is preserved verbatim in the Trace and sent back verbatim — some models require it byte-for-byte on history replay, and any rewriting would break compatibility. This is one reason the Trace stores raw OmniMessage envelopes rather than a post-processed format. ## Observability diff --git a/packages/docs/content/sessions-and-traces.zh.md b/packages/docs/content/sessions-and-traces.zh.md index 8890185..955e489 100644 --- a/packages/docs/content/sessions-and-traces.zh.md +++ b/packages/docs/content/sessions-and-traces.zh.md @@ -81,7 +81,7 @@ Trace 是恢复的唯一事实来源,没有独立的会话数据库需要与 ## 字段保真 -Provider 专有字段(如 `signature`、`phase`)在 Trace 中原样保存、原样回传——部分模型在历史回放时要求这些字段逐字一致,任何转写都会破坏兼容性。这也是 Trace 直接存储 OmniMessage 信封而非二次加工格式的原因之一。 +内容消息携带的不透明 Provider 保真负载 `fidelity`(思考签名、phase 分段标记、加密推理等)在 Trace 中原样保存、原样回传——部分模型在历史回放时要求该负载逐字一致,任何转写都会破坏兼容性。这也是 Trace 直接存储 OmniMessage 信封而非二次加工格式的原因之一。 ## 可观测性 diff --git a/packages/docs/content/skills.en.md b/packages/docs/content/skills.en.md index 517debf..68e5e57 100644 --- a/packages/docs/content/skills.en.md +++ b/packages/docs/content/skills.en.md @@ -58,16 +58,17 @@ The built-in Skills, by group (the group manifest is `SKILL_GROUPS` in `packages | Group | Skill | Purpose | | --- | --- | --- | -| Agent Development | `agent-creation` | Turn a user requirement into a concrete agent: write the target agent's AGENTS.md and install the skills it needs | +| Office Productivity | `data-analysis` | Complete data-analysis tasks with bounded evidence inspection, explicit answer-changing decisions, native artifact handling and final output verification | +| | `firecrawl` | Web search and page scraping into clean markdown via the Firecrawl API | +| Software Development | `web-design` | Penguin visual language for generated web pages and app UIs: design tokens, components, light/dark themes and chat layouts | +| | `software-engineering` | Complete software-engineering tasks: investigate and review code, implement fixes, features and refactors with minimal scope, validate changes, and report verified outcomes | +| AI App Development | `penguin-sdk` | Build AI and RAG apps on the SDK: the createSession/run streaming loop plus a complete retrieval recipe with chunk-revealing citations | +| | `penguin-cli` | Manage model API keys, default models and per-agent Vault secrets with the penguin CLI | +| | `agenthub-models` | Call model APIs through `@prismshadow/agenthub`: streaming text, image generation, speech synthesis and embeddings | +| Agent Tuning | `agent-creation` | Turn a user requirement into a concrete agent: write the target agent's AGENTS.md and install the skills it needs | | | `benchmark-design` | Design and calibrate a multi-Case capability Benchmark with repeated independent evaluations and a traceable baseline | | | `agent-evaluation` | Run and score exactly one Benchmark Case run, with CLI execution, Trace provenance checks and private Rubric isolation | | | `agent-optimization` | Improve an Agent State from direct feedback or versioned multi-Case Benchmark scores and score-linked Traces | -| Data Analysis | `data-analysis` | Complete data-analysis tasks with bounded evidence inspection, explicit answer-changing decisions, native artifact handling and final output verification | -| Penguin Development | `penguin-sdk` | Build AI apps on the SDK (the createSession/run streaming loop) | -| | `penguin-cli` | Manage model API keys, default models and per-agent Vault secrets with the penguin CLI | -| | `agenthub-models` | Call model APIs through `@prismshadow/agenthub`: streaming text, image generation, speech synthesis and embeddings | -| Web Development | `web-design` | Default visual language for generated web pages: minimal black-white-gray | -| Software Engineering | `software-engineering` | Complete software-engineering tasks: investigate and review code, implement fixes, features and refactors with minimal scope, validate changes, and report verified outcomes | ## Writing and optimizing Skills diff --git a/packages/docs/content/skills.zh.md b/packages/docs/content/skills.zh.md index e630604..0c82aec 100644 --- a/packages/docs/content/skills.zh.md +++ b/packages/docs/content/skills.zh.md @@ -58,16 +58,17 @@ Skill 库以 npm 包 `@prismshadow/penguin-skills` 发布,tarball 直接携带 | 分组 | Skill | 说明 | | --- | --- | --- | -| Agent 开发 | `agent-creation` | 把用户需求变成具体的 Agent:撰写目标 Agent 的 AGENTS.md 并安装所需 Skill | +| 办公效率 | `data-analysis` | 以有界的证据检查、显式的改答案决策、原生产物处理与最终输出校验完成数据分析任务 | +| | `firecrawl` | 经 Firecrawl API 做网络搜索与页面抓取,产出干净的 Markdown | +| 软件开发 | `web-design` | 生成网页与应用界面的 Penguin 视觉语言:设计令牌、组件配方、明暗主题与聊天布局 | +| | `software-engineering` | 完成软件工程任务:调查与审查代码,以最小改动实现修复、特性与重构,验证改动并报告经过确认的结果 | +| AI 应用开发 | `penguin-sdk` | 基于 SDK 构建 AI 与 RAG 应用:createSession/run 流式循环,外加带可溯源引用的完整检索配方 | +| | `penguin-cli` | 用 penguin CLI 管理模型 API Key、默认模型与各 Agent 的 Vault 密钥 | +| | `agenthub-models` | 经 `@prismshadow/agenthub` 调用模型 API:流式文本、图像生成、语音合成与 Embedding | +| Agent 调优 | `agent-creation` | 把用户需求变成具体的 Agent:撰写目标 Agent 的 AGENTS.md 并安装所需 Skill | | | `benchmark-design` | 设计并校准多 Case 的能力评测 Benchmark,含重复独立评测与可追溯基线 | | | `agent-evaluation` | 隔离执行并评分单个 Benchmark Case:CLI 执行、Trace 溯源检查、Rubric 私有隔离 | | | `agent-optimization` | 依据直接反馈或带版本的多 Case Benchmark 分数与关联 Trace 改进 Agent State | -| 数据分析 | `data-analysis` | 以有界的证据检查、显式的改答案决策、原生产物处理与最终输出校验完成数据分析任务 | -| Penguin 开发 | `penguin-sdk` | 基于 SDK 构建 AI 应用(createSession/run 流式循环) | -| | `penguin-cli` | 用 penguin CLI 管理模型 API Key、默认模型与各 Agent 的 Vault 密钥 | -| | `agenthub-models` | 经 `@prismshadow/agenthub` 调用模型 API:流式文本、图像生成、语音合成与 Embedding | -| 网页开发 | `web-design` | 生成网页的默认视觉规范:极简黑白灰 | -| 软件工程 | `software-engineering` | 完成软件工程任务:调查与审查代码,以最小改动实现修复、特性与重构,验证改动并报告经过确认的结果 | ## 编写与优化 diff --git a/packages/docs/content/web-app.en.md b/packages/docs/content/web-app.en.md index d402199..01f9572 100644 --- a/packages/docs/content/web-app.en.md +++ b/packages/docs/content/web-app.en.md @@ -23,7 +23,7 @@ penguin web # open http://127.0.0.1:7364 ``` -The initial account is `admin` / `admin123`. There is no self-registration: accounts are created by an admin on the user-management page, and every new user automatically gets an independent initial Project named `-default_project`. While the initial password is still in use, a banner prompts the user to change it. +The initial account is `admin` / `penguin-2026`. There is no self-registration: accounts are created by an admin on the user-management page, and every new user automatically gets an independent initial Project named `-default_project`. While the initial password is still in use, a banner prompts the user to change it. Logins persist for 7 days with sliding renewal; an admin password reset invalidates all of that user's login sessions. diff --git a/packages/docs/content/web-app.zh.md b/packages/docs/content/web-app.zh.md index b909439..989ace0 100644 --- a/packages/docs/content/web-app.zh.md +++ b/packages/docs/content/web-app.zh.md @@ -23,7 +23,7 @@ penguin web # 打开 http://127.0.0.1:7364 ``` -初始账号为 `admin` / `admin123`。系统不开放自助注册:账号由管理员在用户管理页创建;每个新用户会自动获得一个独立的初始 Project,命名为 `-default_project`。仍在使用初始密码时,页面会以横幅提示尽快修改。 +初始账号为 `admin` / `penguin-2026`。系统不开放自助注册:账号由管理员在用户管理页创建;每个新用户会自动获得一个独立的初始 Project,命名为 `-default_project`。仍在使用初始密码时,页面会以横幅提示尽快修改。 登录状态保持 7 天(滑动续期);管理员重置密码会使该用户的全部登录会话失效。 diff --git a/packages/docs/src/components/nav.tsx b/packages/docs/src/components/nav.tsx index 87545c9..976f819 100644 --- a/packages/docs/src/components/nav.tsx +++ b/packages/docs/src/components/nav.tsx @@ -1,16 +1,69 @@ /** - * Sticky top bar: logo + site name + "Docs" badge, then (right) a link back to the - * main site, GitHub, language/theme toggles and — on small screens — the sidebar - * toggle. The sidebar itself lives in the layout (router.tsx); this bar only flips - * its open state. + * Sticky top bar, landing-parity: logo + site name + "Docs" badge, then the SAME + * link row as the landing page's nav (section anchors on the landing home, blog, + * docs) with the same sliding hover pill — so the two sites link into each other + * seamlessly. Section/blog links are plain anchors into the landing SPA one level + * up; "Docs" routes to this site's own root. The right side keeps the language + * and theme toggles, GitHub, and — on small screens — the sidebar toggle. */ +import { useRef } from "react"; +import type { MouseEvent } from "react"; +import { Link } from "react-router"; import { S } from "../lib/strings"; import { REPO_URL, SITE_URL } from "../lib/links"; import { GitHubIcon, MenuIcon, XIcon } from "./icons"; import { ThemeToggle } from "./theme-toggle"; import { LangToggle } from "./lang-toggle"; +const SECTION_IDS = ["highlights", "quickstart", "benchmark", "contract", "features"] as const; + export function Nav({ menuOpen, onToggleMenu }: { menuOpen: boolean; onToggleMenu: () => void }) { + const pillRef = useRef(null); + const pillVisible = useRef(false); + + const sectionLabel: Record<(typeof SECTION_IDS)[number], string> = { + highlights: S.nav.highlights, + quickstart: S.nav.quickstart, + benchmark: S.nav.benchmark, + contract: S.nav.contract, + features: S.nav.features, + }; + + const linkCls = + "relative z-10 rounded-md px-2.5 py-1.5 text-sm text-gray-600 transition-colors hover:text-gray-900 dark:text-gray-400 dark:hover:text-gray-100"; + // Landing-parity active style: on the docs site, the "Docs" link is the current page. + const activeLinkCls = + "relative z-10 rounded-md px-2.5 py-1.5 text-sm transition-colors bg-black text-white hover:bg-black hover:text-white dark:bg-black dark:text-white dark:ring-1 dark:ring-gray-600 dark:hover:bg-black dark:hover:text-white"; + + /** + * The hover pill appears IN PLACE under the first link it lands on (position + * jumps with only the fade animating), slides while moving between links, and + * fades out where it is on leave — never sweeping in from the nav's edge. + */ + const slideTo = (e: MouseEvent) => { + const el = e.currentTarget; + const pill = pillRef.current; + if (!pill) return; + if (!pillVisible.current) { + pill.style.transitionProperty = "opacity"; + pill.style.left = `${el.offsetLeft}px`; + pill.style.width = `${el.offsetWidth}px`; + void pill.offsetWidth; // flush the jump before restoring the full transition + pill.style.transitionProperty = ""; + pillVisible.current = true; + } else { + pill.style.left = `${el.offsetLeft}px`; + pill.style.width = `${el.offsetWidth}px`; + } + pill.style.opacity = "1"; + }; + + const hidePill = () => { + const pill = pillRef.current; + if (pill) pill.style.opacity = "0"; + pillVisible.current = false; + }; + return (
@@ -32,13 +85,32 @@ export function Nav({ menuOpen, onToggleMenu }: { menuOpen: boolean; onToggleMen -
- - {S.nav.home} + + +
Page content, credit amounts, and expiration dates are subject to change. Please refer to the application page and approval email for the latest information. + +## Step 1: Join the AMD AI Developer Program + +Visit the official page: . In the **Join the AMD AI Developer Program** section: + +- No AMD ADP account yet: click **Create Account** and enter your personal information; +- Already have one: click **Log In** and follow the on-screen instructions. + +![The AMD AI Developer Program join page](https://github.com/user-attachments/assets/47a3055b-9a95-40a1-80c6-e3bac7a9ac49) + +## Step 2: Request Cloud Credits under Member Perks + +After registering and signing in: + +1. Click **Member Perks** in the top navigation bar; +2. Find **Cloud Credit Options**; +3. Click **Request Cloud Credits** at the bottom of the page. + +![Cloud Credit Options under Member Perks](https://github.com/user-attachments/assets/773f8cf1-d72f-4aa3-83b8-f31fc7c9ed9e) + +## Step 3: Complete the application form + +Fill in the requested personal information: + +- Under **Product Needed**, select **Fireworks AI**; +- In the **Profile** section, provide at least one public profile for account verification: LinkedIn, GitHub, a portfolio, a company or school profile, WeChat, WeCom, and similar all work. + +![Application form: select Fireworks AI under Product Needed](https://github.com/user-attachments/assets/12a49136-0956-4f29-9d47-f0e473615075) + +Complete the other required fields marked with `*`, double-check your email, identity, product selection, and profile link, then submit. + +## Step 4: Wait for the review + +AMD verifies your account and application, usually within **2–3 business days** — actual time varies with application volume, information completeness, and holidays. + +## Step 5: Receive and redeem the coupon code + +Once approved, AMD emails a unique **Coupon Code** redeemable for **$50 in Fireworks AI Credits** to your application address. Keep it secure — do not share it publicly, forward it, or commit it to a repository. + +![The Coupon Code in the approval email](https://github.com/user-attachments/assets/c53129c7-2c87-4510-a4e3-31f52598dc24) + +Redeem it and create an API key: + +1. Open and sign in; +2. Click **Redeem Promo** and enter the Coupon Code from the email to redeem the $50 credits; +3. Click **Create API Key** to generate your Fireworks API key. + +![Redeeming and creating an API key in the Fireworks console](https://github.com/user-attachments/assets/051a1e69-db7f-4867-b899-89981df15142) + +## Set it up in PenguinHarness + +With the API key in hand, three steps: + +**1. Install and launch** + +```bash +curl -fsSL https://penguin.ooo/install.sh | sh +penguin web # opens http://127.0.0.1:7364 (first login: admin / penguin-2026) +``` + +**2. Configure a Fireworks model** + +Open the Models page and find the **Fireworks AI** group, then use its bulk key button to paste the API key you just created. The group presets five models — GLM 5.2, Kimi K2.7 Code, DeepSeek V4 Pro, MiniMax M3, and DeepSeek V4 Flash — with base URLs and pricing pre-filled; set any of them as the default. You can also hit the group's speed-test button to measure real TTFT and TPS before choosing. + +**3. Start working** + +Head back to Chat and hand the Agent its first task — e.g. "Analyze data.csv and summarize quarterly sales". + +## References + +- [AMD AI Developer Program](https://developer.amd.com/ai-developer-program/) +- [Official AMD Cloud Credits application video tutorial](https://www.youtube.com/watch?v=masSW53JkTY) +- Application steps and screenshots are adapted from WhatGhost's guides ([中文](https://github.com/WhatGhost/whatghost_Notebooks/blob/main/other/AMD_AI_Developer_Program_Credits_%E7%94%B3%E8%AF%B7%E6%8C%87%E5%8D%97.md) / [English](https://github.com/WhatGhost/whatghost_Notebooks/blob/main/other/AMD_AI_Developer_Program_Credits_Application_Guide_EN.md)) — thanks to the original author +- [PenguinHarness model configuration docs](https://penguin.ooo/docs/models) diff --git a/packages/landing/content/blog/fireworks-credits-amd.zh.md b/packages/landing/content/blog/fireworks-credits-amd.zh.md new file mode 100644 index 0000000..d62e454 --- /dev/null +++ b/packages/landing/content/blog/fireworks-credits-amd.zh.md @@ -0,0 +1,84 @@ +--- +title: 携手 AMD 开发者计划:免费领取 $50 Fireworks 额度,直连 PenguinHarness +date: 2026-07-20 +category: news +excerpt: 我们很高兴与 AMD AI Developer Program 合作,为大家带来 Fireworks 的免费兑换码——按本文申请 $50 Credits,再在 PenguinHarness 里三步用起来。 +--- + +我们很高兴与 **AMD AI Developer Program**(AMD 开发者计划)合作,为大家带来 Fireworks 的免费兑换码:加入计划并通过审核,即可获得可兑换 **$50 Fireworks AI Credits** 的 Coupon Code。PenguinHarness 内置 Fireworks AI 网关分组——OpenAI 协议、预置 base URL 与五个模型,额度到手即刻可用。 + +> 页面内容、Credits 金额与有效期可能调整,请以申请时的页面及审批邮件为准。 + +## 第一步:加入 AMD AI Developer Program + +访问官方入口:,在 **Join the AMD AI Developer Program** 区域: + +- 没有 AMD ADP 账户:点击 **Create Account**,填写个人信息创建账户; +- 已有账户:点击 **Log In**,按页面提示登录。 + +![AMD AI Developer Program 加入页面](https://github.com/user-attachments/assets/47a3055b-9a95-40a1-80c6-e3bac7a9ac49) + +## 第二步:在 Member Perks 申请 Cloud Credits + +注册并登录后: + +1. 点击顶部导航栏中的 **Member Perks**; +2. 找到 **Cloud Credit Options**; +3. 点击底部的 **Request Cloud Credits**。 + +![Member Perks 中的 Cloud Credit Options](https://github.com/user-attachments/assets/773f8cf1-d72f-4aa3-83b8-f31fc7c9ed9e) + +## 第三步:填写申请表 + +进入申请表后填写个人信息: + +- **Product Needed** 处选择 **Fireworks AI**; +- **Profile** 处提供至少一个公开资料用于账户验证:LinkedIn、GitHub、Portfolio、公司 / 学校主页、WeChat、WeCom 等均可。 + +![申请表:Product Needed 选择 Fireworks AI](https://github.com/user-attachments/assets/12a49136-0956-4f29-9d47-f0e473615075) + +填完页面中其他带 `*` 的必填项,检查邮箱、身份、产品选项与公开资料链接无误后提交。 + +## 第四步:等待审核 + +AMD 会验证账户与申请资料,通常需要 **2–3 个工作日**;实际时间可能因申请量、资料完整度或节假日而变化。 + +## 第五步:收到兑换码并兑换 + +审核通过后,AMD 会向申请邮箱发送包含唯一 **Coupon Code** 的邮件,可兑换 **$50 Fireworks AI Credits**。请妥善保存,不要公开、转发或提交到代码仓库。 + +![审批邮件中的 Coupon Code](https://github.com/user-attachments/assets/c53129c7-2c87-4510-a4e3-31f52598dc24) + +兑换与创建 API key: + +1. 打开 并登录; +2. 点击 **Redeem Promo**,输入邮件中的 Coupon Code,兑换 $50 Credits; +3. 点击 **Create API Key**,生成 Fireworks API key。 + +![在 Fireworks 控制台兑换并创建 API key](https://github.com/user-attachments/assets/051a1e69-db7f-4867-b899-89981df15142) + +## 在 PenguinHarness 中用起来 + +拿到 API key 后,三步接入: + +**1. 安装并启动** + +```bash +curl -fsSL https://penguin.ooo/install.sh | sh +penguin web # 打开 http://127.0.0.1:7364(首次登录:admin / penguin-2026) +``` + +**2. 配置 Fireworks 模型** + +进入「模型仓库」页,找到 **Fireworks AI** 分组,点击「统一配置 key」粘贴刚创建的 API key。分组预置了五个模型——GLM 5.2、Kimi K2.7 Code、DeepSeek V4 Pro、MiniMax M3、DeepSeek V4 Flash——base URL 与价格已填好,任选一个设为默认即可;也可以点组头的「测速」,实测各模型的 TTFT 与 TPS 再决定。 + +**3. 开始使用** + +回到对话页,把第一个任务交给 Agent——例如「分析 data.csv,输出各季度销售额汇总」。 + +## 参考链接 + +- [AMD AI Developer Program](https://developer.amd.com/ai-developer-program/) +- [AMD 官方 Cloud Credits 申请视频演示](https://www.youtube.com/watch?v=masSW53JkTY) +- 申请步骤与截图整理自 WhatGhost 的申请指南([中文](https://github.com/WhatGhost/whatghost_Notebooks/blob/main/other/AMD_AI_Developer_Program_Credits_%E7%94%B3%E8%AF%B7%E6%8C%87%E5%8D%97.md) / [English](https://github.com/WhatGhost/whatghost_Notebooks/blob/main/other/AMD_AI_Developer_Program_Credits_Application_Guide_EN.md)),感谢原作者 +- [PenguinHarness 模型配置文档](https://penguin.ooo/docs/models) diff --git a/packages/landing/content/blog/introducing-penguinharness.en.md b/packages/landing/content/blog/introducing-penguinharness.en.md index 5816b8c..f516864 100644 --- a/packages/landing/content/blog/introducing-penguinharness.en.md +++ b/packages/landing/content/blog/introducing-penguinharness.en.md @@ -2,70 +2,114 @@ title: "Introducing PenguinHarness: agents that build agents" date: 2026-07-17 category: news -excerpt: The first open-source harness with recursive self-improvement is here — lightweight, efficient and secure infrastructure covering everything from automatic agent construction to continuous self-evolution. +excerpt: We proved agents can self-evolve in our GDPevo Benchmark — now we are bringing that capability to everyone. The first open-source harness with recursive self-improvement covers everything from one-sentence agent construction to continuous self-evolution. --- -Today we are releasing **PenguinHarness** — an open-source harness built for constructing and evolving agents. Its purpose fits in one line: +Today we are releasing **PenguinHarness** — an open-source harness built for constructing and evolving agents: a zero-code Harness CLI and Web UI, connected to 1000+ models. The story it tells fits in one line: -> Efficient Self-Improving Harness for Everyone. +> With LangChain, you build agents by hand — at 1× speed. With PenguinHarness, agents build agents — at 100×. + +## From GDPevo to PenguinHarness: why we built this + +Before PenguinHarness, our team published the [GDPevo Benchmark](https://prism-shadow.github.io/GDPevo/). In GDPevo we systematically verified one thing: **agents can self-evolve** — an Agent can score its own performance, find where the points were lost, rewrite its own prompts and Skills, and climb version after version. + +With the capability proven, the question became: how does everyone get to use it? Self-evolution should not stay a curve in a paper — it should be infrastructure that works out of the box on every developer's desk. **Bringing an efficient self-improving harness to everyone is why we built PenguinHarness** — and it is right there in the name: Efficient Self-Improving Harness for Everyone. ## Why PenguinHarness -Over the past year the way agent applications are built has been converging fast: what really decides quality is not a heavyweight framework but a simple, reliable, observable harness. PenguinHarness is rebuilt from the ground up — no dependency on any agent framework, a fully open-source self-developed kernel — and it brings three things to the open-source world first: +Three reasons, in deliberate order — from task quality, to how agents get built, to how they keep improving. -- **Simplest Is the Best**: a deliberately minimal toolset over clean low-level interfaces — fewer tool calls, fewer Tokens, complex tasks done efficiently. -- **Harness for Building Agents**: with the PenguinHarness SDK, an Agent builds complete Agent applications for you, autonomously, from scratch. -- **Harness for Recursive Self-Improvement**: with PenguinHarness Skills, an Agent evaluates and optimizes itself, improving recursively over time. +### 1. Better on complex tasks, at lower cost -For the latter two, PenguinHarness is the first open-source implementation in the industry. +A deliberately minimal toolset over clean low-level interfaces: fewer tool calls, fewer Tokens, deeply tuned for open models like DeepSeek. Each harness runs the model it is normally paired with — the comparison is between the products as people actually use them — head-to-head on two suites: -## Same model, equal or better quality, lower cost +![Benchmark: PenguinHarness leads the data-analysis suite and ties OpenAI Codex on coding, at a small fraction of both rivals' cost](/blog-assets/benchmark-light.svg) -All runs use the same DeepSeek V4 Pro model, head-to-head against Claude Code and OpenAI Codex on two suites (per-run means below). +Complex data analysis (15 tasks, single run; PenguinHarness and Codex at thinking xhigh, Claude Code at max): -Complex data analysis (15 tasks, single run): +| Framework | Model | Accuracy (%) | Tokens (M) | Cost ($) | +| -------------- | --------------- | -----------: | ---------: | -------: | +| PenguinHarness | DeepSeek V4 Pro | 66.67 | 18.04 | 0.55 | +| Claude Code | Claude Opus 4.8 | 53.33 | 22.20 | 38.48 | +| OpenAI Codex | GPT-5.5 | 53.33 | 13.72 | 19.41 | -| Framework | Model | Accuracy (%) | Tokens (M) | Cost ($) | -| --- | --- | ---: | ---: | ---: | -| PenguinHarness | DeepSeek V4 Pro | 66.7 | 18.04 | 0.552 | -| Claude Code | DeepSeek V4 Pro | 66.7 | 21.17 | 0.641 | -| OpenAI Codex | DeepSeek V4 Pro | 46.7 | 13.36 | 0.427 | +Coding tasks (40 tasks × 2 runs; accuracy is over all 80 outcomes): -Coding tasks (40 tasks × 2 runs averaged, thinking high, 30 min per-case timeout, CNY pricing converted at $1 = ¥7): +| Framework | Model | Accuracy (%) | Tokens (M) | Cost ($) | +| -------------- | --------------- | -----------: | ---------: | -------: | +| PenguinHarness | DeepSeek V4 Pro | 71.25 | 200.00 | 3.81 | +| Claude Code | Claude Opus 4.8 | 86.25 | 151.61 | 146.97 | +| OpenAI Codex | GPT-5.5 | 71.25 | 251.20 | 220.08 | -| Framework | Model | Accuracy (%) | Tokens (M) | Cost ($) | -| --- | --- | ---: | ---: | ---: | -| PenguinHarness | DeepSeek V4 Pro | 50.00 | 2.10 | 0.041 | -| Claude Code | DeepSeek V4 Pro | 48.75 | 2.00 | 0.048 | -| OpenAI Codex | DeepSeek V4 Pro | 42.50 | 2.65 | 0.043 | +Tokens and cost are suite totals, not per-run means. On data analysis we take the highest accuracy of the three — 66.67% against 53.33% for both — while spending 1/35 of what Codex spent and 1/70 of Claude Code's. On coding we tie Codex at 71.25% and trail Claude Code's 86.25%, but the whole suite cost us $3.81 against their $220.08 and $146.97: comparable work, one to two orders of magnitude apart on the bill. -On the data-analysis suite PenguinHarness ties Claude Code on accuracy and clearly beats OpenAI Codex while using 14.8% fewer Tokens at 13.8% lower cost; on the coding suite it scores highest of the three at the lowest per-run cost. +### 2. One sentence, and an Agent builds your Agent app + +Type one sentence, and an Agent builds the complete Agent application for you — scaffold, code, and run instructions, end to end: + +```text +Collect the docs from https://github.com/ericbuess/claude-code-docs and build a RAG app that answers Claude Code questions as a configuration expert, citing its sources. +``` + +And this is the finished product — a docs expert with retrieval, cited sources that link to the original files, and example questions built in: + +![The generated RAG app: a Claude Code docs expert answering with cited, clickable sources and example questions](/blog-assets/rag-app-en-light.webp) + +**And generating this entire RAG app burned just $0.02 (¥0.2) of tokens — on DeepSeek V4 Pro.** + +### 3. Self-evolution: it gets stronger with use + +With PenguinHarness Skills, an Agent evaluates and optimizes itself: the Optimizer orchestrates multiple Evaluators to score in parallel, uses the scores and run traces to find where points were lost, and upgrades the Agent from version N to N+1 — with a snapshot before every round, and every request replayable in the Trace view. A self-evolution demo video is coming soon. ## Evolution within bounds, security first -The biggest worry about self-improvement is losing control. PenguinHarness answers with a contract — CONTRACT.md: +The biggest worry about self-evolution is losing control. PenguinHarness answers with a contract (CONTRACT.md): -- Evolution is strictly confined to Workspace and Skills; the harness core security boundary is never modified. -- Tool calls run only after approval, and every approval is audited. -- Risky changes snapshot first — every step of evolution can be rolled back. -- Fully open source and locally deployable: data never leaves your machine, meeting enterprise security requirements. +- Evolution is strictly confined to Workspace and Skills — the harness core security boundary is never modified; +- Tool calls require approval first, and every approval leaves an audit record; +- Risky changes are preceded by version snapshots, so any round of evolution can be rolled back; +- Fully open source and locally deployed — data never leaves your machine, meeting enterprise data-security requirements. -## Get started now +## Supported models -Install with one command (Linux / macOS, x64 / arm64, bundled Node runtime): +| Model | Providers | +| ---------------- | -------------------------------------------------------------------------------- | +| DeepSeek V4 | DeepSeek, OpenRouter, Fireworks AI, SiliconFlow, Qwen Token Plan | +| Kimi K3 | Moonshot AI, OpenRouter, Qwen Pay-As-You-Go | +| GLM 5.2 | Z.AI, OpenRouter, Fireworks AI, SiliconFlow, Qwen Token Plan, Qwen Pay-As-You-Go | +| Hunyuan 3 | OpenRouter | +| Qwen 3.8 Max | Qwen Token Plan (preview) | +| GPT 5.5 | OpenAI, OpenRouter | +| Gemini 3.5 Flash | Google Gemini, OpenRouter | +| Claude Opus 4.8 | Anthropic, OpenRouter | + +Any OpenAI-protocol endpoint is supported: pick a preset above, or point a custom endpoint at any of the 1000+ online and local models. + +## How to use it + +Install with one command (Linux / macOS, x64 / arm64, bundled Node runtime), then launch the Web UI: ```bash -curl -fsSL https://github.com/Prism-Shadow/penguin-harness/releases/latest/download/install.sh | sh +curl -fsSL https://penguin.ooo/install.sh | sh +penguin web # opens http://127.0.0.1:7364 (first login: admin / penguin-2026) ``` -Configure a model (DeepSeek as an example), run your first task, or open the desktop-grade web interface with `penguin web`: +Open the Models page, paste an API key under the DeepSeek or OpenRouter group and set it as default; then head back to Chat and hand the Agent its first task — e.g. "Analyze data.csv and summarize quarterly sales". -```bash -penguin config model add --model-id deepseek-v4-pro --api-key sk-your-key --set-default -penguin run --approve allow-all --message "Analyze data.csv and summarize quarterly sales" -penguin web -``` +## What's next -PenguinHarness supports 1000+ online and local models and multi-agent collaborative evolution, and runs on as little as a single CPU. Through continuous evolution it makes complex AI development ever simpler — a more efficient, more reliable, lower-hallucination and lower-cost Agent productivity engine. +- Public release of the benchmark suite; +- A desktop app; +- Windows support; +- More to come. -Follow us on [GitHub](https://github.com/Prism-Shadow/penguin-harness) and open your first issue. +## Join the community and build with us + +A self-improving harness needs a community that improves with it. Come discuss, request, and contribute — your first Issue is the best way to start: + +- [Discord](https://discord.gg/eFHKqqcU3D): chat with us and other developers in real time; +- [X (Twitter)](https://x.com/code_hiyouga): follow the latest updates; +- [WeChat group](https://github.com/Prism-Shadow/penguin-harness-community/blob/main/wechat/group.jpg): Chinese community discussions; +- [GitHub](https://github.com/Prism-Shadow/penguin-harness): stars, Issues, and PRs all welcome. + +Self-evolving agent infrastructure, for everyone — starting today. diff --git a/packages/landing/content/blog/introducing-penguinharness.zh.md b/packages/landing/content/blog/introducing-penguinharness.zh.md index 6b34c4d..22864bf 100644 --- a/packages/landing/content/blog/introducing-penguinharness.zh.md +++ b/packages/landing/content/blog/introducing-penguinharness.zh.md @@ -2,44 +2,64 @@ title: PenguinHarness 正式发布:让 Agent 为你构建 Agent date: 2026-07-17 category: news -excerpt: 首个支持递归自我进化的开源 Harness 正式发布——以轻量、高效、安全的方式,提供从 Agent 自动构建到持续自我进化的完整基础设施。 +excerpt: 我们在 GDPevo Benchmark 中验证了 Agent 自我进化的能力,现在把它带给所有人——首个支持递归自我进化的开源 Harness 正式发布,从一句话构建 Agent 到持续自我进化,一套基础设施全部覆盖。 --- -今天,我们正式发布 **PenguinHarness**——一个为构建与进化 Agent 而生的开源 Harness。它的主旨只有一句话: +今天,我们正式发布 **PenguinHarness**——一个为构建与进化 Agent 而生的开源 Harness:零代码的 Harness CLI 与 Web UI,连接 1000+ 模型。它要讲的故事只有一句话: -> Efficient Self-Improving Harness for Everyone. +> 使用 LangChain,以 1 倍速度人工构建 Agent;使用 PenguinHarness,以 100 倍速度用 Agent 构建 Agent。 + +## 从 GDPevo 到 PenguinHarness:我们的初心 + +在 PenguinHarness 之前,我们团队发布了 [GDPevo Benchmark](https://prism-shadow.github.io/GDPevo/)。在 GDPevo 中,我们系统性地验证了一件事:**Agent 可以自我进化**——让 Agent 评估自己的表现、定位失分原因、改写自己的提示词与技能,分数随版本一路上升。 + +能力验证了,问题就变成了:怎么让每个人都用上它?自我进化不该只是论文里的曲线,而应该是每个开发者桌面上开箱即用的基础设施。**让每个人都能使用 Efficient Self-Improving Harness,这就是我们构建 PenguinHarness 的初心**——它也因此得名:Efficient Self-Improving Harness for Everyone. ## 为什么是 PenguinHarness -过去一年里,Agent 应用的开发范式在快速收敛:真正决定效果的不是庞大的框架,而是一个简洁、可靠、可观测的 Harness。PenguinHarness 从底层重构,不依赖任何 Agent 框架,开源自研 Harness 内核,并率先把三件事带入开源世界: +三个递进的理由——从任务效果,到构建方式,再到进化能力。 -- **Simplest Is the Best**:坚持最小化工具集与简洁的底层接口,以更少的工具调用与 Token 消耗,高效完成复杂任务。 -- **Harness for Building Agents**:通过 PenguinHarness SDK,让 Agent 从零自主完成 Agent 应用的构建。 -- **Harness for Recursive Self-Improvement**:通过 PenguinHarness Skills,Agent 以自我评估与自我优化实现递归式自我提升。 +### 1. 复杂任务表现更好,成本更低 -后两项能力,PenguinHarness 是业内首个开源实现。 +刻意精简的工具集配合干净的底层接口:更少的工具调用、更少的 Token,对 DeepSeek 等开放模型深度适配。每个产品搭配它常用的模型——比的是大家实际会怎么用——在两套题库上正面对比: -## 同一模型,同级效果,更低消耗 +![Benchmark:PenguinHarness 在数据分析题库准确率最高、编程题库与 OpenAI Codex 持平,成本仅为两者的零头](/blog-assets/benchmark-light.svg) -全部使用同一 DeepSeek V4 Pro 模型,与 Claude Code、OpenAI Codex 在两套题库上正面对比(表中为单次运行均值)。 +复杂数据分析(15 题,单次运行;PenguinHarness 与 Codex 为 thinking xhigh,Claude Code 为 max): -复杂数据分析(15 题,单次运行): +| 实验框架 | 模型名称 | 准确率(%) | Token 用量(M) | 成本($) | +| -------------- | --------------- | ----------: | --------------: | --------: | +| PenguinHarness | DeepSeek V4 Pro | 66.67 | 18.04 | 0.55 | +| Claude Code | Claude Opus 4.8 | 53.33 | 22.20 | 38.48 | +| OpenAI Codex | GPT-5.5 | 53.33 | 13.72 | 19.41 | -| 实验框架 | 模型名称 | 准确率(%) | Token 用量(M) | 成本($) | -| --- | --- | ---: | ---: | ---: | -| PenguinHarness | DeepSeek V4 Pro | 66.7 | 18.04 | 0.552 | -| Claude Code | DeepSeek V4 Pro | 66.7 | 21.17 | 0.641 | -| OpenAI Codex | DeepSeek V4 Pro | 46.7 | 13.36 | 0.427 | +代码任务(40 题 × 2 次运行,准确率取全部 80 次结果): -代码任务(40 题 × 2 runs 取均值,thinking high、单题 30 分钟超时,人民币计价按 $1 = ¥7 折算): +| 实验框架 | 模型名称 | 准确率(%) | Token 用量(M) | 成本($) | +| -------------- | --------------- | ----------: | --------------: | --------: | +| PenguinHarness | DeepSeek V4 Pro | 71.25 | 200.00 | 3.81 | +| Claude Code | Claude Opus 4.8 | 86.25 | 151.61 | 146.97 | +| OpenAI Codex | GPT-5.5 | 71.25 | 251.20 | 220.08 | -| 实验框架 | 模型名称 | 准确率(%) | Token 用量(M) | 成本($) | -| --- | --- | ---: | ---: | ---: | -| PenguinHarness | DeepSeek V4 Pro | 50.00 | 2.10 | 0.041 | -| Claude Code | DeepSeek V4 Pro | 48.75 | 2.00 | 0.048 | -| OpenAI Codex | DeepSeek V4 Pro | 42.50 | 2.65 | 0.043 | +Token 与成本均为全套题目合计,不是单次均值。数据分析套件我们准确率最高——66.67% 对另两者的 53.33%——花的钱是 Codex 的 1/35、Claude Code 的 1/70;代码套件与 Codex 同为 71.25%、低于 Claude Code 的 86.25%,但整套题目只花了 $3.81,对方分别是 $220.08 与 $146.97:活干得差不多,账单差出一到两个数量级。 -数据分析套件与 Claude Code 准确率持平、显著超过 OpenAI Codex,同时 Token 消耗少 14.8%、成本低 13.8%;代码套件三者中准确率最高、单次成本最低。 +### 2. 一句话,让 Agent 构建 Agent 应用 + +输入一句话,Agent 为你构建完整的 Agent 应用——脚手架、代码、运行说明,一步到位: + +```text +收集 https://github.com/ericbuess/claude-code-docs 的文档,做一个化身 Claude Code 配置专家、回答带来源引用的 RAG 问答应用。 +``` + +这是做出来的成品——一个文档专家:检索增强、引用可点击直达原文、内置示例问题: + +![生成的 RAG 应用成品:Claude Code 配置专家,回答带可点击的来源引用与示例问题](/blog-assets/rag-app-zh-light.webp) + +**而生成整个 RAG 应用,仅消耗了 0.2 元($0.02)的 token——使用 DeepSeek V4 Pro 模型。** + +### 3. 自进化,越用越强 + +借助 PenguinHarness 技能库,Agent 自己评估、自己优化:Optimizer 组织多个 Evaluator 并行打分,依据分数与运行轨迹定位失分原因,把 Agent 从版本 N 优化到版本 N+1——每轮之前自动快照,每个请求都可在轨迹观测中回放。自进化演示视频即将上线。 ## 进化有界,安全先行 @@ -50,22 +70,46 @@ excerpt: 首个支持递归自我进化的开源 Harness 正式发布——以 - 风险修改之前先留版本快照,任何一次进化都可回退; - 完全开源、本地部署,数据不出域,满足企业级数据安全。 -## 现在就可以开始 +## 支持的模型 -一行命令安装(Linux / macOS,x64 / arm64,内嵌 Node 运行时): +| 模型 | 可用供应商 | +| ---------------- | -------------------------------------------------------------------------------- | +| DeepSeek V4 | DeepSeek, OpenRouter, Fireworks AI, SiliconFlow, Qwen Token Plan | +| Kimi K3 | Moonshot AI, OpenRouter, Qwen Pay-As-You-Go | +| GLM 5.2 | Z.AI, OpenRouter, Fireworks AI, SiliconFlow, Qwen Token Plan, Qwen Pay-As-You-Go | +| Hunyuan 3 | OpenRouter | +| Qwen 3.8 Max | Qwen Token Plan(预览) | +| GPT 5.5 | OpenAI, OpenRouter | +| Gemini 3.5 Flash | Google Gemini, OpenRouter | +| Claude Opus 4.8 | Anthropic, OpenRouter | + +只要是 OpenAI 协议的端点都可以接入:从上表选择预置,或用自定义端点连接 1000+ 在线与本地模型。 + +## 如何使用 + +一行命令安装(Linux / macOS,x64 / arm64,内嵌 Node 运行时),然后启动 Web 界面: ```bash -curl -fsSL https://github.com/Prism-Shadow/penguin-harness/releases/latest/download/install.sh | sh +curl -fsSL https://penguin.ooo/install.sh | sh +penguin web # 打开 http://127.0.0.1:7364(首次登录:admin / penguin-2026) ``` -配置模型(以 DeepSeek 为例)后即可运行第一个任务,或用 `penguin web` 打开桌面级 Web 界面: +进入「模型仓库」页,在 DeepSeek 或 OpenRouter 分组里粘贴 API key 并设为默认;回到对话页,把第一个任务交给 Agent——例如「分析 data.csv,输出各季度销售额汇总」。 -```bash -penguin config model add --model-id deepseek-v4-pro --api-key sk-your-key --set-default -penguin run --approve allow-all --message "分析 data.csv,输出各季度销售额汇总" -penguin web -``` +## 发展计划 -PenguinHarness 支持 1000 多种在线与本地模型、多智能体协作进化,最低单 CPU 即可运行。通过不断进化,它会让复杂的 AI 开发越来越简单——为你提供更高效、更可靠、更低幻觉、更低成本的 Agent 生产力引擎。 +- Benchmark 套件正式发布; +- 推出桌面端(Desktop)应用; +- 支持 Windows 系统; +- 更多规划,敬请期待。 -欢迎在 [GitHub](https://github.com/Prism-Shadow/penguin-harness) 上关注我们,提出你的第一个 Issue。 +## 加入社区,一起共建 + +自我进化的 Harness,也需要一个不断进化的社区。欢迎加入讨论、提出需求、贡献代码——你的第一个 Issue 就是最好的开始: + +- [Discord](https://discord.gg/eFHKqqcU3D):与我们和其他开发者实时交流; +- [X(Twitter)](https://x.com/code_hiyouga):关注最新动态; +- [微信群](https://github.com/Prism-Shadow/penguin-harness-community/blob/main/wechat/group.jpg):中文社区讨论; +- [GitHub](https://github.com/Prism-Shadow/penguin-harness):Star、Issue 与 PR 都欢迎。 + +让每个人都用上会自我进化的 Agent 基础设施——从今天开始。 diff --git a/packages/landing/public/blog-assets/benchmark-light.svg b/packages/landing/public/blog-assets/benchmark-light.svg new file mode 100644 index 0000000..31c480c --- /dev/null +++ b/packages/landing/public/blog-assets/benchmark-light.svg @@ -0,0 +1,48 @@ + +Accuracy · suite total · higher is better +Total cost (USD) · lower is better +Data analysis — 15 tasks, single run + +PenguinHarness + +66.67% +Claude Code + +53.33% +OpenAI Codex + +53.33% + +PenguinHarness + +$0.55 +Claude Code + +$38.48 +OpenAI Codex + +$19.41 +Coding — 40 tasks × 2 runs + +PenguinHarness + +71.25% +Claude Code + +86.25% +OpenAI Codex + +71.25% + +PenguinHarness + +$3.81 +Claude Code + +$146.97 +OpenAI Codex + +$220.08 +Each harness runs the model it is normally paired with: PenguinHarness on DeepSeek V4 Pro, Claude Code on Claude Opus 4.8, +OpenAI Codex on GPT-5.5. Accuracy, Tokens and cost are suite totals at official pricing. + diff --git a/packages/landing/public/blog-assets/rag-app-en-light.webp b/packages/landing/public/blog-assets/rag-app-en-light.webp new file mode 100644 index 0000000..0a2115b Binary files /dev/null and b/packages/landing/public/blog-assets/rag-app-en-light.webp differ diff --git a/packages/landing/public/blog-assets/rag-app-zh-light.webp b/packages/landing/public/blog-assets/rag-app-zh-light.webp new file mode 100644 index 0000000..b9dba57 Binary files /dev/null and b/packages/landing/public/blog-assets/rag-app-zh-light.webp differ diff --git a/packages/landing/public/install.sh b/packages/landing/public/install.sh new file mode 100644 index 0000000..b2348fc --- /dev/null +++ b/packages/landing/public/install.sh @@ -0,0 +1,21 @@ +#!/bin/sh +# https://penguin.ooo/install.sh - PenguinHarness installer entry point. +# +# GitHub Pages cannot serve HTTP redirects, so this thin forwarder IS the +# stable install URL: it fetches the real installer attached to the latest +# GitHub release and runs it, forwarding every argument it was given. Usage: +# +# curl -fsSL https://penguin.ooo/install.sh | sh +# curl -fsSL https://penguin.ooo/install.sh | sh -s -- --universal +# +set -eu +# Download to a file first, then run it: piping straight into `sh` would execute +# a truncated download line by line, and the real installer removes the old +# bin/lib/web/node before moving the new ones in — a cut connection mid-way +# would leave no install at all. +TMP="$(mktemp)" +trap 'rm -f "$TMP"' EXIT +curl -fsSL "https://github.com/Prism-Shadow/penguin-harness/releases/latest/download/install.sh" -o "$TMP" +rc=0 +sh "$TMP" "$@" || rc=$? +exit "$rc" diff --git a/packages/landing/scripts/capture-game-mockup.mjs b/packages/landing/scripts/capture-game-mockup.mjs new file mode 100644 index 0000000..541d5bd --- /dev/null +++ b/packages/landing/scripts/capture-game-mockup.mjs @@ -0,0 +1,62 @@ +/** + * Render the landing Cases tab's "penguin sled game" finished-product shot. + * + * Same pattern as capture-readme-demo.mjs: penguin-game-mockup.html is a static, + * dependency-free mockup of the example game's play screen, captured per language + * (zh / en) and theme (light = polar day / dark = polar night) into + * packages/landing/src/assets as game--.webp. + * + * Prereqs: Playwright's chromium only (no server, no build). Run: + * `node scripts/capture-game-mockup.mjs [--png-dir ]`. + */ +import { mkdirSync, writeFileSync } from "node:fs"; +import path from "node:path"; +import { fileURLToPath, pathToFileURL } from "node:url"; +import { chromium } from "@playwright/test"; + +const HERE = path.dirname(fileURLToPath(import.meta.url)); +const OUT_DIR = path.resolve(HERE, "../src/assets"); +const PAGE = pathToFileURL(path.join(HERE, "penguin-game-mockup.html")).href; + +const pngDirArg = process.argv.indexOf("--png-dir"); +const PNG_DIR = pngDirArg >= 0 ? process.argv[pngDirArg + 1] : null; + +mkdirSync(OUT_DIR, { recursive: true }); +if (PNG_DIR) mkdirSync(PNG_DIR, { recursive: true }); + +const browser = await chromium.launch(); +const encoderPage = await browser.newPage(); +async function saveWebp(pngBuffer, fileName) { + const dataUrl = await encoderPage.evaluate(async (b64) => { + const img = new Image(); + img.src = `data:image/png;base64,${b64}`; + await img.decode(); + const canvas = document.createElement("canvas"); + canvas.width = img.width; + canvas.height = img.height; + canvas.getContext("2d").drawImage(img, 0, 0); + return canvas.toDataURL("image/webp", 0.9); + }, pngBuffer.toString("base64")); + writeFileSync(path.join(OUT_DIR, fileName), Buffer.from(dataUrl.split(",")[1], "base64")); + console.log(`[game] ${fileName}`); +} + +for (const lang of ["en", "zh"]) { + for (const theme of ["light", "dark"]) { + const context = await browser.newContext({ + viewport: { width: 1280, height: 760 }, + deviceScaleFactor: 1.5, + locale: lang === "zh" ? "zh-CN" : "en-US", + }); + const page = await context.newPage(); + await page.goto(`${PAGE}?lang=${lang}&theme=${theme}`); + await page.waitForTimeout(300); + const png = await page.screenshot(); + if (PNG_DIR) writeFileSync(path.join(PNG_DIR, `game-${lang}-${theme}.png`), png); + await saveWebp(png, `game-${lang}-${theme}.webp`); + await context.close(); + } +} + +await browser.close(); +console.log("[game] done"); diff --git a/packages/landing/scripts/capture-readme-demo.mjs b/packages/landing/scripts/capture-readme-demo.mjs new file mode 100644 index 0000000..859017e --- /dev/null +++ b/packages/landing/scripts/capture-readme-demo.mjs @@ -0,0 +1,65 @@ +/** + * Render the README "build an Agent app in one sentence" finished-product shot. + * + * The image shows the RESULT of the example — a Claude Code docs-expert RAG app — + * rather than the build conversation: rag-app-mockup.html is a static, dependency-free + * mockup of the generated app's UI (per the web-design skill's Penguin visual language), + * captured per language (zh / en) and theme (light / dark) into assets/readme/ at the + * repo root as rag-app--.webp (README.md uses en, README.zh.md uses zh). + * + * Prereqs: Playwright's chromium only (no server, no build). Run: + * `node scripts/capture-readme-demo.mjs [--png-dir ]` (--png-dir also saves + * lossless PNG copies, handy for reviewing the shots). + */ +import { mkdirSync, writeFileSync } from "node:fs"; +import path from "node:path"; +import { fileURLToPath, pathToFileURL } from "node:url"; +import { chromium } from "@playwright/test"; + +const HERE = path.dirname(fileURLToPath(import.meta.url)); +const ROOT = path.resolve(HERE, "../../.."); +const OUT_DIR = path.resolve(ROOT, "assets/readme"); +const PAGE = pathToFileURL(path.join(HERE, "rag-app-mockup.html")).href; + +const pngDirArg = process.argv.indexOf("--png-dir"); +const PNG_DIR = pngDirArg >= 0 ? process.argv[pngDirArg + 1] : null; + +mkdirSync(OUT_DIR, { recursive: true }); +if (PNG_DIR) mkdirSync(PNG_DIR, { recursive: true }); + +const browser = await chromium.launch(); +const encoderPage = await browser.newPage(); +async function saveWebp(pngBuffer, fileName) { + const dataUrl = await encoderPage.evaluate(async (b64) => { + const img = new Image(); + img.src = `data:image/png;base64,${b64}`; + await img.decode(); + const canvas = document.createElement("canvas"); + canvas.width = img.width; + canvas.height = img.height; + canvas.getContext("2d").drawImage(img, 0, 0); + return canvas.toDataURL("image/webp", 0.9); + }, pngBuffer.toString("base64")); + writeFileSync(path.join(OUT_DIR, fileName), Buffer.from(dataUrl.split(",")[1], "base64")); + console.log(`[demo] ${fileName}`); +} + +for (const lang of ["en", "zh"]) { + for (const theme of ["light", "dark"]) { + const context = await browser.newContext({ + viewport: { width: 1280, height: 760 }, + deviceScaleFactor: 1.5, + locale: lang === "zh" ? "zh-CN" : "en-US", + }); + const page = await context.newPage(); + await page.goto(`${PAGE}?lang=${lang}&theme=${theme}`); + await page.waitForTimeout(300); + const png = await page.screenshot(); + if (PNG_DIR) writeFileSync(path.join(PNG_DIR, `rag-app-${lang}-${theme}.png`), png); + await saveWebp(png, `rag-app-${lang}-${theme}.webp`); + await context.close(); + } +} + +await browser.close(); +console.log("[demo] done"); diff --git a/packages/landing/scripts/capture-shots.mjs b/packages/landing/scripts/capture-shots.mjs index 56559f5..314a8fe 100644 --- a/packages/landing/scripts/capture-shots.mjs +++ b/packages/landing/scripts/capture-shots.mjs @@ -34,96 +34,172 @@ const MOCK = `http://127.0.0.1:${MOCK_PORT}`; // Commands are shared across languages (code is code) and really execute. // --------------------------------------------------------------------------- -const CMD_SCAFFOLD = `mkdir -p csv-analyst/src && cat > csv-analyst/package.json <<'EOF' +const CMD_COLLECT = `git clone --depth 1 https://github.com/ericbuess/claude-code-docs \\ + claude-code-expert/corpus/claude-code-docs +find claude-code-expert/corpus -type f ! -name '*.md' -delete +rm -rf claude-code-expert/corpus/claude-code-docs/.git +ls claude-code-expert/corpus/claude-code-docs | head -6`; + +const CMD_APP = `mkdir -p claude-code-expert/src claude-code-expert/public +cat > claude-code-expert/package.json <<'EOF' { - "name": "csv-analyst", + "name": "claude-code-expert", "private": true, "type": "module", - "scripts": { "start": "tsx src/agent.ts" }, - "dependencies": { "@prismshadow/penguin-core": "^0.1.0" } + "scripts": { "start": "tsx src/rag.ts" }, + "dependencies": { "@prismshadow/penguin-core": "^0.1.0", "tsx": "^4" } } EOF -ls -R csv-analyst`; +cat > claude-code-expert/src/rag.ts <<'EOF' +// BM25 retrieval over corpus/ + a Session that answers with clickable [n] citations. +import fs from "node:fs"; +import http from "node:http"; +import path from "node:path"; +import { createAgent, isModelMessage, userText } from "@prismshadow/penguin-core"; -const CMD_ENTRY = `cat > csv-analyst/src/agent.ts <<'EOF' -import { createAgent, isCompleteModelMessage, userText } from "@prismshadow/penguin-core"; +const walk = (d) => + fs.readdirSync(d, { withFileTypes: true }).flatMap((e) => + e.isDirectory() ? walk(path.join(d, e.name)) : [path.join(d, e.name)]); +const chunks = walk("corpus").flatMap((f) => + fs.readFileSync(f, "utf8").split(/\\n(?=#{1,3} )/).map((text) => ({ source: f, text }))); -const agent = await createAgent({ agentId: "csv_analyst" }); -const session = await agent.createSession({ workspaceDir: process.cwd() }); - -for await (const out of session.run([userText("Analyze data.csv and write summary.md")], { - approve: async () => "allow", -})) { - if (isCompleteModelMessage(out) && out.payload.type === "text") { - console.log(out.payload.text); +const tok = (s) => s.toLowerCase().match(/[a-z0-9]+|[一-鿿]/g) ?? []; +const docs = chunks.map((c) => tok(c.text)); +const avg = docs.reduce((n, d) => n + d.length, 0) / Math.max(docs.length, 1); +const df = new Map(); +for (const d of docs) for (const w of new Set(d)) df.set(w, (df.get(w) ?? 0) + 1); +// BM25 (k1=1.2, b=0.75): idf weights rare terms, term frequency saturates, long chunks are penalized. +const score = (q, i) => { + const d = docs[i]; + let s = 0; + for (const w of new Set(tok(q))) { + const f = d.filter((x) => x === w).length; + if (!f) continue; + const n = df.get(w) ?? 0; + s += Math.log(1 + (docs.length - n + 0.5) / (n + 0.5)) * (f * 2.2) / (f + 1.2 * (0.25 + 0.75 * d.length / avg)); } -} + return s; +}; +const agent = await createAgent({ root: "penguin_data" }); + +http.createServer(async (req, res) => { + if (req.method !== "POST") { res.end(fs.readFileSync("public/index.html")); return; } + let body = ""; + for await (const p of req) body += p; + const { question } = JSON.parse(body); + const hits = chunks.map((c, i) => [score(question, i), c]).filter(([s]) => s > 0) + .sort((a, b) => b[0] - a[0]).slice(0, 6).map(([, c]) => c); + res.writeHead(200, { "content-type": "text/event-stream" }); + const ctx = hits.map((c, i) => "[" + (i + 1) + "] " + c.source + "\\n" + c.text).join("\\n\\n"); + const session = await agent.createSession({ workspaceDir: process.cwd() }); + for await (const m of session.run([userText(ctx + "\\n\\nQ: " + question)], { + approve: async () => "deny", + })) { + if (isModelMessage(m) && m.payload.type === "partial_text" && m.payload.event_type === "delta") + res.write("data: " + JSON.stringify({ delta: m.payload.text }) + "\\n\\n"); + } + res.write("data: " + JSON.stringify({ sources: hits.map((c) => c.source) }) + "\\n\\n"); + session.dispose(); + res.end(); +}).listen(4630); EOF -wc -l csv-analyst/src/agent.ts`; +cat > claude-code-expert/public/index.html <<'EOF' + + +Claude Code docs expert + +

Claude Code docs expert

+
+ + + +EOF +wc -l claude-code-expert/src/rag.ts`; const TREE = `\`\`\`text -csv-analyst/ +claude-code-expert/ ├── package.json -└── src/ - └── agent.ts +├── corpus/claude-code-docs/ # the collected docs +├── src/rag.ts +└── public/index.html \`\`\``; /** Per-language script: user prompt marker -> turns + session title. */ const SCRIPTS = { zh: { - marker: "数据分析 Agent 应用", - prompt: "用 PenguinHarness SDK 创建一个数据分析 Agent 应用:读取 CSV 并输出汇总报告", - title: "构建数据分析 Agent 应用", + marker: "配置专家", + prompt: + "收集 https://github.com/ericbuess/claude-code-docs 的文档,做一个化身 Claude Code 配置专家、回答带来源引用的 RAG 问答应用。", + title: "构建 Claude Code 文档专家", turns: [ { thinking: - "需求是基于 penguin-core 的数据分析 Agent 应用。先创建项目骨架:package.json 与源码目录。", - text: "我来创建应用骨架:", - cmd: CMD_SCAFFOLD, + "一句话需求,penguin-sdk 技能给出了完整配方:先采集语料,再建 BM25 索引,回答经 Session 流式生成并带 [n] 引用直达原文。", + text: "先把文档采集进语料库:", + cmd: CMD_COLLECT, }, { thinking: - "骨架已建好。入口代码用 createAgent + createSession,把 CSV 分析任务交给 session.run。", - text: "骨架就绪,写入 Agent 入口代码:创建 Session,把 CSV 分析任务交给 session.run 并流式输出。", - cmd: CMD_ENTRY, + "语料就绪。写应用:本地 BM25 检索(零 embedding 依赖)+ 引用回答 + Penguin 风格聊天界面,引用链接到 /corpus 原文。", + text: "语料就绪,写检索与引用回答的应用入口:", + cmd: CMD_APP, }, { - text: `数据分析 Agent 应用已创建完成: + text: `Claude Code 文档专家已就绪: ${TREE} -- 入口 \`csv-analyst/src/agent.ts\`:创建 Agent 与 Session,任务经 \`session.run\` 流式执行,工具调用逐个审批; -- 运行方式:\`cd csv-analyst && npm install && npm start\`; -- 建议下一步:在评估中心为它配一组 CSV 任务 Benchmark,交给 Optimizer 持续优化。`, +- 检索:本地 BM25 索引全部文档片段,中文提问同样支持; +- 回答:每次提问经 Session 流式生成,引用 [1][2] 可点击直达 \`corpus/\` 原文; +- 界面:Penguin 风格聊天页,空态内置示例问题; +- 运行:\`cd claude-code-expert && npm install && npm start\`,浏览器打开 http://localhost:4630。`, }, ], }, en: { - marker: "data-analysis Agent app", + marker: "configuration expert", prompt: - "Use the PenguinHarness SDK to create a data-analysis Agent app that reads CSV files and writes a summary report", - title: "Build a data-analysis Agent app", + "Collect the docs from https://github.com/ericbuess/claude-code-docs and build a RAG app that answers Claude Code questions as a configuration expert, citing its sources.", + // Stay under core's TITLE_MAX_CHARS (30): a longer title gets hard-clipped + // mid-word in the shots (e.g. "…docs exper"). + title: "Build Claude Code docs expert", turns: [ { thinking: - "They want a data-analysis Agent app on penguin-core. Start with the project skeleton: package.json plus the source directory.", - text: "Let me scaffold the app first:", - cmd: CMD_SCAFFOLD, + "One sentence is enough — the penguin-sdk skill has the full recipe: collect the corpus, build a BM25 index, answer through a Session with [n] citations linking to the originals.", + text: "Collecting the docs into the corpus first:", + cmd: CMD_COLLECT, }, { thinking: - "Skeleton is in place. The entry uses createAgent + createSession and hands the CSV task to session.run.", - text: "Skeleton ready — now the Agent entry point: create a Session and hand the CSV analysis task to session.run, streaming the output.", - cmd: CMD_ENTRY, + "Corpus in place. Now the app: local BM25 retrieval (no embedding credential), cited answers, and a Penguin-style chat UI with citations linking to /corpus originals.", + text: "Corpus ready — now the retrieval and cited-answer entry:", + cmd: CMD_APP, }, { - text: `The data-analysis Agent app is ready: + text: `The Claude Code docs expert is ready: ${TREE} -- Entry \`csv-analyst/src/agent.ts\`: creates the Agent and a Session; the task runs through \`session.run\` with per-tool approval; -- Run it with \`cd csv-analyst && npm install && npm start\`; -- Suggested next step: give it a CSV Benchmark suite in the evaluation center and let an Optimizer keep improving it.`, +- Retrieval: a local BM25 index over every doc chunk, Chinese questions included; +- Answers: each question streams through a Session, with [1][2] citations that link straight to the \`corpus/\` originals; +- UI: a Penguin-style chat page with example questions in the empty state; +- Run: \`cd claude-code-expert && npm install && npm start\`, then open http://localhost:4630.`, }, ], }, @@ -466,7 +542,7 @@ try { await waitFor(`${BASE}/`); console.log(`[shots] server ready on ${BASE}`); - const admin = await login("admin", "admin123"); + const admin = await login("admin", "penguin-2026"); const browser = await chromium.launch(); // WebP encoder: Chromium re-encodes the PNG screenshot buffer via canvas, which @@ -549,17 +625,15 @@ try { await page.waitForTimeout(2000); await saveWebp(await page.screenshot(), `chat-${lang}-${theme}.webp`); - // Trace view: select the session in the list (deep-link selection is unreliable - // right after a fresh navigation, so click explicitly — sidebar shows the same - // title first in DOM order, hence .last()). - await page.goto(`${BASE}/traces?sessionId=${sessionId}`); - await page.waitForTimeout(1500); - await page - .getByText(script.title) - .last() - .click() - .catch(() => {}); - await page.waitForTimeout(2500); + // Trace view: the product's canonical deep link carries BOTH agentId and + // sessionId (?sessionId= alone is ignored by TracesPage's focus wiring), and + // auto-selects the Session once its trace list loads. Waiting for the + // execution timeline's exec_command lanes guarantees every language captures + // the same opened trace — stats + a timeline with tool calls — never the + // empty "select a Session" state. + await page.goto(`${BASE}/traces?agentId=default_agent&sessionId=${sessionId}`); + await page.getByText("exec_command").first().waitFor({ timeout: 20000 }); + await page.waitForTimeout(2000); await saveWebp(await page.screenshot(), `traces-${lang}-${theme}.webp`); // Evaluation center: open the pre-provisioned example Benchmark scoreboard. diff --git a/packages/landing/scripts/penguin-game-mockup.html b/packages/landing/scripts/penguin-game-mockup.html new file mode 100644 index 0000000..125282e --- /dev/null +++ b/packages/landing/scripts/penguin-game-mockup.html @@ -0,0 +1,333 @@ + + + + + + Penguin Sled Run + + + +
+ + +
+
+ 1280 +
+
+ 512 m +
+
×2.4
+
+
+ + + + diff --git a/packages/landing/scripts/rag-app-mockup.html b/packages/landing/scripts/rag-app-mockup.html new file mode 100644 index 0000000..3728c94 --- /dev/null +++ b/packages/landing/scripts/rag-app-mockup.html @@ -0,0 +1,439 @@ + + + + + + Claude Code Docs Expert + + + +
+ + +
+
+
+
+

+ + [1] +

+

+claude mcp add github -- npx -y @modelcontextprotocol/server-github
+
+
+claude mcp list
+

+ + [2] +

+ +
+
+ +
+
+ + + +
+
+
+ +
+
+
+
+ + + + diff --git a/packages/landing/scripts/render-benchmark-svg.mjs b/packages/landing/scripts/render-benchmark-svg.mjs new file mode 100644 index 0000000..430733a --- /dev/null +++ b/packages/landing/scripts/render-benchmark-svg.mjs @@ -0,0 +1,179 @@ +/** + * Renders the README / blog benchmark chart from the same data the landing page uses + * (src/lib/benchmark-data.ts), so the static SVGs cannot drift from the site when the + * numbers are refreshed. Emits: + * assets/readme/benchmark-light.svg + * assets/readme/benchmark-dark.svg + * packages/landing/public/blog-assets/benchmark-light.svg + * + * Two panels per suite (accuracy, cost), horizontal bars scaled linearly from zero — the + * cost spread is ~70x, so the PenguinHarness bar is by far the shortest. That is the + * result, not a defect; a zoomed baseline would flatter everyone else. Bars are floored at + * MIN_BAR, though: at true scale the cost bar comes out ~2.5px, which reads as "no bar" and + * loses the series entirely. The floor keeps it visible and unmistakably smallest, and the + * exact figure is printed beside every bar, so nothing is overstated. + * + * Run: node packages/landing/scripts/render-benchmark-svg.mjs + */ +import { readFileSync, writeFileSync } from "node:fs"; +import path from "node:path"; +import { fileURLToPath } from "node:url"; + +const HERE = path.dirname(fileURLToPath(import.meta.url)); +const LANDING = path.resolve(HERE, ".."); +const REPO = path.resolve(LANDING, "..", ".."); + +/** + * The data module is TypeScript; rather than add a build step, read the two exported + * arrays out of the source. Any shape change here fails loudly instead of silently + * rendering a stale chart. + */ +function loadBench() { + const src = readFileSync(path.join(LANDING, "src/lib/benchmark-data.ts"), "utf8"); + const consts = Object.fromEntries( + [...src.matchAll(/^const (\w+) = "([^"]+)";$/gm)].map((m) => [m[1], m[2]]), + ); + const pick = (name) => { + const block = src.match( + new RegExp(`export const ${name}: BenchResult\\[\\] = \\[([\\s\\S]*?)\\n\\];`), + ); + if (!block) throw new Error(`could not find ${name} in benchmark-data.ts`); + const rows = [...block[1].matchAll(/\{([\s\S]*?)\n {2}\}/g)].map((m) => { + const body = m[1]; + const field = (k) => body.match(new RegExp(`${k}:\\s*([^,\\n]+)`))?.[1]?.trim(); + const str = (k) => { + const v = field(k); + if (!v) throw new Error(`missing ${k}`); + return v.startsWith('"') ? v.slice(1, -1) : (consts[v] ?? v); + }; + return { + framework: str("framework"), + model: str("model"), + accuracyPct: Number(field("accuracyPct")), + tokensM: Number(field("tokensM")), + costUsd: Number(field("costUsd")), + emphasized: /emphasized:\s*true/.test(body), + }; + }); + if (rows.length !== 3) throw new Error(`${name}: expected 3 rows, got ${rows.length}`); + for (const r of rows) { + if (!Number.isFinite(r.accuracyPct) || !Number.isFinite(r.costUsd)) { + throw new Error(`${name}: ${r.framework} has a non-numeric field`); + } + } + return rows; + }; + return { DATA_BENCH: pick("DATA_BENCH"), CODE_BENCH: pick("CODE_BENCH") }; +} + +const THEMES = { + light: { + strong: "#1f2328", + mid: "#52514e", + muted: "#898781", + rule: "#c3c2b7", + brand: "#2a78d6", + bar: "#898781", + }, + dark: { + strong: "#f0f3f6", + mid: "#c3c2b7", + muted: "#898781", + rule: "#383835", + brand: "#3987e5", + bar: "#6b6a64", + }, +}; + +const FONT = "system-ui, -apple-system, 'Segoe UI', sans-serif"; +const W = 920; +const H = 368; +const BAR_H = 16; +const ROW_H = 30; +const MAX_BAR = 176; +/** Shortest a bar may render, so a near-zero value still reads as a bar (~8% of full). */ +const MIN_BAR = 14; + +const esc = (s) => s.replace(/&/g, "&").replace(//g, ">"); + +function text(x, y, s, { size = 12, weight = 400, fill, anchor = "start" }) { + return `${esc(s)}`; +} + +/** Rounded-end horizontal bar starting at the axis. */ +function bar(x, y, w, fill) { + const r = Math.min(4, w / 2); + return ``; +} + +/** One measure for one suite: label column, axis rule, three bars, value labels. */ +function panel(t, x0, yTop, rows, value, format) { + const axis = x0 + 124; + const max = Math.max(...rows.map(value)); + const out = [ + ``, + ]; + rows.forEach((row, i) => { + const y = yTop + 4 + i * ROW_H; + const w = Math.max(MIN_BAR, (value(row) / max) * MAX_BAR); + const weight = row.emphasized ? 600 : 400; + out.push( + text(axis - 12, y + 12.5, row.framework, { + fill: row.emphasized ? t.strong : t.mid, + weight, + anchor: "end", + }), + bar(axis + 1, y, w, row.emphasized ? t.brand : t.bar), + text(axis + w + 13, y + 12.5, format(value(row)), { fill: t.strong, weight }), + ); + }); + return out.join("\n"); +} + +function suite(t, yTop, title, rows) { + return [ + text(24, yTop, title, { size: 13, weight: 600, fill: t.strong }), + panel( + t, + 24, + yTop + 10, + rows, + (r) => r.accuracyPct, + (v) => `${v.toFixed(2)}%`, + ), + panel( + t, + 472, + yTop + 10, + rows, + (r) => r.costUsd, + (v) => `$${v.toFixed(2)}`, + ), + ].join("\n"); +} + +function render(theme, { DATA_BENCH, CODE_BENCH }) { + const t = THEMES[theme]; + const alt = + "Benchmark: PenguinHarness vs Claude Code vs OpenAI Codex on two suites — comparable accuracy at a small fraction of the cost"; + return ` +${text(148, 36, "Accuracy · suite total · higher is better", { fill: t.muted })} +${text(596, 36, "Total cost (USD) · lower is better", { fill: t.muted })} +${suite(t, 68, "Data analysis — 15 tasks, single run", DATA_BENCH)} +${suite(t, 210, "Coding — 40 tasks × 2 runs", CODE_BENCH)} +${text(24, 338, "Each harness runs the model it is normally paired with: PenguinHarness on DeepSeek V4 Pro, Claude Code on Claude Opus 4.8,", { fill: t.muted })} +${text(24, 356, "OpenAI Codex on GPT-5.5. Accuracy, Tokens and cost are suite totals at official pricing.", { fill: t.muted })} + +`; +} + +const bench = loadBench(); +const targets = [ + ["light", path.join(REPO, "assets/readme/benchmark-light.svg")], + ["dark", path.join(REPO, "assets/readme/benchmark-dark.svg")], + ["light", path.join(LANDING, "public/blog-assets/benchmark-light.svg")], +]; +for (const [theme, file] of targets) { + writeFileSync(file, render(theme, bench)); + console.log(`wrote ${path.relative(REPO, file)}`); +} diff --git a/packages/landing/src/assets/evo-poster-en.webp b/packages/landing/src/assets/evo-poster-en.webp new file mode 100644 index 0000000..f8fdd0b Binary files /dev/null and b/packages/landing/src/assets/evo-poster-en.webp differ diff --git a/packages/landing/src/assets/evo-poster-zh.webp b/packages/landing/src/assets/evo-poster-zh.webp new file mode 100644 index 0000000..69ab3ec Binary files /dev/null and b/packages/landing/src/assets/evo-poster-zh.webp differ diff --git a/packages/landing/src/assets/game-en-dark.webp b/packages/landing/src/assets/game-en-dark.webp new file mode 100644 index 0000000..4ef7cc6 Binary files /dev/null and b/packages/landing/src/assets/game-en-dark.webp differ diff --git a/packages/landing/src/assets/game-en-light.webp b/packages/landing/src/assets/game-en-light.webp new file mode 100644 index 0000000..9421d7d Binary files /dev/null and b/packages/landing/src/assets/game-en-light.webp differ diff --git a/packages/landing/src/assets/game-zh-dark.webp b/packages/landing/src/assets/game-zh-dark.webp new file mode 100644 index 0000000..067678d Binary files /dev/null and b/packages/landing/src/assets/game-zh-dark.webp differ diff --git a/packages/landing/src/assets/game-zh-light.webp b/packages/landing/src/assets/game-zh-light.webp new file mode 100644 index 0000000..e58a0c9 Binary files /dev/null and b/packages/landing/src/assets/game-zh-light.webp differ diff --git a/packages/landing/src/assets/rag-app-en-dark.webp b/packages/landing/src/assets/rag-app-en-dark.webp new file mode 100644 index 0000000..52392a7 Binary files /dev/null and b/packages/landing/src/assets/rag-app-en-dark.webp differ diff --git a/packages/landing/src/assets/rag-app-en-light.webp b/packages/landing/src/assets/rag-app-en-light.webp new file mode 100644 index 0000000..0a2115b Binary files /dev/null and b/packages/landing/src/assets/rag-app-en-light.webp differ diff --git a/packages/landing/src/assets/rag-app-zh-dark.webp b/packages/landing/src/assets/rag-app-zh-dark.webp new file mode 100644 index 0000000..8a698e0 Binary files /dev/null and b/packages/landing/src/assets/rag-app-zh-dark.webp differ diff --git a/packages/landing/src/assets/rag-app-zh-light.webp b/packages/landing/src/assets/rag-app-zh-light.webp new file mode 100644 index 0000000..b9dba57 Binary files /dev/null and b/packages/landing/src/assets/rag-app-zh-light.webp differ diff --git a/packages/landing/src/assets/shots/benchmark-en-dark.webp b/packages/landing/src/assets/shots/benchmark-en-dark.webp index d48d05c..e6229e8 100644 Binary files a/packages/landing/src/assets/shots/benchmark-en-dark.webp and b/packages/landing/src/assets/shots/benchmark-en-dark.webp differ diff --git a/packages/landing/src/assets/shots/benchmark-en-light.webp b/packages/landing/src/assets/shots/benchmark-en-light.webp index 05116dd..d8c77e4 100644 Binary files a/packages/landing/src/assets/shots/benchmark-en-light.webp and b/packages/landing/src/assets/shots/benchmark-en-light.webp differ diff --git a/packages/landing/src/assets/shots/benchmark-zh-dark.webp b/packages/landing/src/assets/shots/benchmark-zh-dark.webp index 815428a..23e7dc4 100644 Binary files a/packages/landing/src/assets/shots/benchmark-zh-dark.webp and b/packages/landing/src/assets/shots/benchmark-zh-dark.webp differ diff --git a/packages/landing/src/assets/shots/benchmark-zh-light.webp b/packages/landing/src/assets/shots/benchmark-zh-light.webp index 70ee6f7..3731335 100644 Binary files a/packages/landing/src/assets/shots/benchmark-zh-light.webp and b/packages/landing/src/assets/shots/benchmark-zh-light.webp differ diff --git a/packages/landing/src/assets/shots/chat-en-dark.webp b/packages/landing/src/assets/shots/chat-en-dark.webp index 5e3567a..10a40dc 100644 Binary files a/packages/landing/src/assets/shots/chat-en-dark.webp and b/packages/landing/src/assets/shots/chat-en-dark.webp differ diff --git a/packages/landing/src/assets/shots/chat-en-light.webp b/packages/landing/src/assets/shots/chat-en-light.webp index 6af357b..767d90d 100644 Binary files a/packages/landing/src/assets/shots/chat-en-light.webp and b/packages/landing/src/assets/shots/chat-en-light.webp differ diff --git a/packages/landing/src/assets/shots/chat-zh-dark.webp b/packages/landing/src/assets/shots/chat-zh-dark.webp index 237e726..1eabeec 100644 Binary files a/packages/landing/src/assets/shots/chat-zh-dark.webp and b/packages/landing/src/assets/shots/chat-zh-dark.webp differ diff --git a/packages/landing/src/assets/shots/chat-zh-light.webp b/packages/landing/src/assets/shots/chat-zh-light.webp index 17b4e2a..ae3a858 100644 Binary files a/packages/landing/src/assets/shots/chat-zh-light.webp and b/packages/landing/src/assets/shots/chat-zh-light.webp differ diff --git a/packages/landing/src/assets/shots/traces-en-dark.webp b/packages/landing/src/assets/shots/traces-en-dark.webp index 7b50434..8f73512 100644 Binary files a/packages/landing/src/assets/shots/traces-en-dark.webp and b/packages/landing/src/assets/shots/traces-en-dark.webp differ diff --git a/packages/landing/src/assets/shots/traces-en-light.webp b/packages/landing/src/assets/shots/traces-en-light.webp index 1139619..dd81aed 100644 Binary files a/packages/landing/src/assets/shots/traces-en-light.webp and b/packages/landing/src/assets/shots/traces-en-light.webp differ diff --git a/packages/landing/src/assets/shots/traces-zh-dark.webp b/packages/landing/src/assets/shots/traces-zh-dark.webp index ad9cb8c..20855c2 100644 Binary files a/packages/landing/src/assets/shots/traces-zh-dark.webp and b/packages/landing/src/assets/shots/traces-zh-dark.webp differ diff --git a/packages/landing/src/assets/shots/traces-zh-light.webp b/packages/landing/src/assets/shots/traces-zh-light.webp index 7a07794..a4c6903 100644 Binary files a/packages/landing/src/assets/shots/traces-zh-light.webp and b/packages/landing/src/assets/shots/traces-zh-light.webp differ diff --git a/packages/landing/src/components/announcement-bar.tsx b/packages/landing/src/components/announcement-bar.tsx new file mode 100644 index 0000000..3765ad2 --- /dev/null +++ b/packages/landing/src/components/announcement-bar.tsx @@ -0,0 +1,122 @@ +/** + * Rotating announcement bar above the sticky nav (it scrolls away with the page). + * Light brand-tinted background; announcements auto-advance by SLIDING in one + * direction only (a clone of the first slide follows the last one, and the track + * snaps back without animation once the clone is fully in view). No manual + * switch controls; hovering or focusing the bar pauses the rotation. Every + * announcement is a link into the blog, with a small arrow marking it as + * click-through. + */ +import { useEffect, useState } from "react"; +import { Link } from "react-router"; +import { S } from "../lib/strings"; +import { ArrowRightIcon } from "./icons"; + +const ROTATE_MS = 6000; + +/** Both announcements point at blog posts (content/blog/.*.md). */ +const ITEMS = [ + { key: "models", to: "/blog/introducing-penguinharness" }, + { key: "fireworks", to: "/blog/fireworks-credits-amd" }, +] as const; + +const REDUCED_MOTION_QUERY = "(prefers-reduced-motion: reduce)"; + +function prefersReducedMotion(): boolean { + return window.matchMedia(REDUCED_MOTION_QUERY).matches; +} + +export function AnnouncementBar() { + const texts: Record<(typeof ITEMS)[number]["key"], string> = { + models: S.announcement.models, + fireworks: S.announcement.fireworks, + }; + // pos runs 0..ITEMS.length where ITEMS.length is the clone of slide 0. + const [pos, setPos] = useState(0); + const [animate, setAnimate] = useState(true); + // Focus pauses like hover does: tabIndex={-1} never blurs an already focused + // link, so advancing past it would strand document focus inside an aria-hidden + // slide translated out of the viewport, and WCAG 2.2.2 wants the rotation + // stoppable without a pointer. The two are tracked apart so the pointer + // leaving the bar cannot restart it while focus is still inside. + const [hovered, setHovered] = useState(false); + const [focused, setFocused] = useState(false); + const paused = hovered || focused; + const [reduced, setReduced] = useState(prefersReducedMotion); + + // Tracked live rather than read once, because the preference can be toggled + // while the page is open and it decides how the rotation wraps around. + useEffect(() => { + const mq = window.matchMedia(REDUCED_MOTION_QUERY); + const onChange = (e: MediaQueryListEvent) => setReduced(e.matches); + setReduced(mq.matches); + mq.addEventListener("change", onChange); + return () => mq.removeEventListener("change", onChange); + }, []); + + useEffect(() => { + if (paused) return; + // styles.css kills every transition under reduced motion, so the clone's + // transitionend never arrives and the snap-back below never runs: those + // users wrap around the real slides instead and never land on the clone. + const advance = (p: number) => { + if (reduced) return (p + 1) % ITEMS.length; + return p >= ITEMS.length ? p : p + 1; + }; + const timer = setInterval(() => setPos(advance), ROTATE_MS); + return () => clearInterval(timer); + }, [paused, reduced]); + + // After snapping back to the real first slide, re-enable the transition on the + // next frame pair so the jump itself never animates. + useEffect(() => { + if (animate) return; + const raf = requestAnimationFrame(() => requestAnimationFrame(() => setAnimate(true))); + return () => cancelAnimationFrame(raf); + }, [animate]); + + const slides = [...ITEMS, ITEMS[0]!]; + + return ( +
setHovered(true)} + onMouseLeave={() => setHovered(false)} + onFocusCapture={() => setFocused(true)} + onBlurCapture={() => setFocused(false)} + className="border-b border-brand-100 bg-brand-50 dark:border-brand-950 dark:bg-brand-950/50" + > +
+
{ + if (e.target !== e.currentTarget || e.propertyName !== "transform") return; + if (pos === ITEMS.length) { + setAnimate(false); + setPos(0); + } + }} + > + {slides.map((item, i) => ( +
+ + {texts[item.key]} +
+ ))} +
+
+
+ ); +} diff --git a/packages/landing/src/components/demo-video.tsx b/packages/landing/src/components/demo-video.tsx new file mode 100644 index 0000000..5ccf0e2 --- /dev/null +++ b/packages/landing/src/components/demo-video.tsx @@ -0,0 +1,45 @@ +/** + * Demo video embed. The file itself is hosted in the sibling community repo (see + * `demoVideoUrl` in lib/links.ts for why), so the important part here is that nothing + * is fetched until the visitor asks: `preload="none"` plus a poster means a page view + * costs one ~37 KB still, not a ~9 MB download, and the section still looks finished + * before anyone presses play. + * + * Deliberately NOT wrapped in BrowserFrame: that chrome (with its localhost address bar) + * says "this is the product's UI", which is true of a screenshot and of the case demos, + * but not of a narrated slide deck. A plain framed surface instead. + */ + +export function DemoVideo({ + src, + poster, + label, + caption, +}: { + src: string; + poster: string; + /** Accessible name — the surrounding copy is decorative to a screen reader. */ + label: string; + caption?: string; +}) { + return ( +
+
+
+ {caption && ( +
+ {caption} +
+ )} +
+ ); +} diff --git a/packages/landing/src/components/icons.tsx b/packages/landing/src/components/icons.tsx index 5da366f..306e8b2 100644 --- a/packages/landing/src/components/icons.tsx +++ b/packages/landing/src/components/icons.tsx @@ -315,6 +315,22 @@ export function ChevronDownIcon(props: IconProps) { ); } +export function ChevronLeftIcon(props: IconProps) { + return ( + + + + ); +} + +export function ChevronRightIcon(props: IconProps) { + return ( + + + + ); +} + export function ExternalLinkIcon(props: IconProps) { return ( diff --git a/packages/landing/src/components/nav.tsx b/packages/landing/src/components/nav.tsx index f8eb96f..73dbc3b 100644 --- a/packages/landing/src/components/nav.tsx +++ b/packages/landing/src/components/nav.tsx @@ -1,30 +1,34 @@ /** * Sticky top navigation: logo + section anchors + blog link + language/theme toggles + * GitHub. Desktop links share a sliding hover pill, while the selected section or route - * keeps its own active background; section links route through "/#id" so the Router - * and URL hash stay in sync while native smooth scrolling remains enabled. A disclosure - * menu covers small screens. + * keeps its own active background. On the home page the active section tracks the + * LIVE scroll position (scroll-spy), so the highlight follows as you scroll; on other + * routes it falls back to route state (e.g. Blog). Section links route through "/#id" + * so the URL hash stays in sync on click. A disclosure menu covers small screens. */ -import { useState } from "react"; +import { useRef, useState } from "react"; import type { MouseEvent } from "react"; import { Link, useLocation } from "react-router"; import { S } from "../lib/strings"; import { DOCS_URL, REPO_URL } from "../lib/links"; import { getActiveNavItem, SECTION_IDS } from "../lib/nav-state"; +import type { ActiveNavItem, SectionId } from "../lib/nav-state"; +import { useScrollSpy } from "../lib/use-scroll-spy"; import { GitHubIcon, MenuIcon, XIcon } from "./icons"; import { ThemeToggle } from "./theme-toggle"; import { LangToggle } from "./lang-toggle"; -interface Indicator { - left: number; - width: number; -} +/** Stable empty list: keeps the spy idle away from the home page. */ +const NO_IDS: readonly string[] = []; export function Nav() { const { pathname, hash } = useLocation(); - const activeItem = getActiveNavItem(pathname, hash); + const onHome = pathname === "/"; + const spied = useScrollSpy(onHome ? SECTION_IDS : NO_IDS) as SectionId | null; + const activeItem: ActiveNavItem = onHome ? spied : getActiveNavItem(pathname, hash); const [open, setOpen] = useState(false); - const [indicator, setIndicator] = useState(null); + const pillRef = useRef(null); + const pillVisible = useRef(false); const sectionLabel: Record<(typeof SECTION_IDS)[number], string> = { highlights: S.nav.highlights, @@ -45,9 +49,33 @@ export function Nav() { const deskLinkCls = (active: boolean) => `relative z-10 rounded-md px-2.5 py-1.5 text-sm transition-colors ${active ? activeLinkCls : deskInactiveLinkCls}`; + /** + * The hover pill appears IN PLACE under the first link it lands on (position + * jumps with only the fade animating), slides while moving between links, and + * fades out where it is on leave — never sweeping in from the nav's edge. + */ const slideTo = (e: MouseEvent) => { const el = e.currentTarget; - setIndicator({ left: el.offsetLeft, width: el.offsetWidth }); + const pill = pillRef.current; + if (!pill) return; + if (!pillVisible.current) { + pill.style.transitionProperty = "opacity"; + pill.style.left = `${el.offsetLeft}px`; + pill.style.width = `${el.offsetWidth}px`; + void pill.offsetWidth; // flush the jump before restoring the full transition + pill.style.transitionProperty = ""; + pillVisible.current = true; + } else { + pill.style.left = `${el.offsetLeft}px`; + pill.style.width = `${el.offsetWidth}px`; + } + pill.style.opacity = "1"; + }; + + const hidePill = () => { + const pill = pillRef.current; + if (pill) pill.style.opacity = "0"; + pillVisible.current = false; }; const desktopLinks = ( @@ -122,17 +150,14 @@ export function Nav() { diff --git a/packages/landing/src/lib/benchmark-data.ts b/packages/landing/src/lib/benchmark-data.ts index 535553c..904c573 100644 --- a/packages/landing/src/lib/benchmark-data.ts +++ b/packages/landing/src/lib/benchmark-data.ts @@ -1,11 +1,14 @@ /** - * Benchmark data, two suites driven by the same DeepSeek V4 Pro model, published in - * one unified shape: framework / model / accuracy (%) / Tokens (M) / cost ($), all - * as per-run means. Suite specifics (case count, runs, thinking level, timeout, - * pricing source) live in the footnote strings, not in the table. - * - Data analysis: 15 tasks, single run, USD pricing. - * - Coding: 40 tasks x 2 runs averaged; official CNY pricing converted at $1 = ¥7 - * (0.289 / 0.338 / 0.299 CNY per run). + * Benchmark data, two suites published in one unified shape: framework / model / + * accuracy (%) / Tokens (M) / cost ($). + * + * Each harness runs the model it is normally paired with rather than a shared one — + * the comparison is between products as people actually use them, so the model column + * is part of the result and not a constant. Token and cost figures are **suite totals** + * across every task and run, not per-run means. Suite specifics (case count, runs, + * thinking level, pricing source) live in the footnote strings, not in the table. + * - Data analysis: 15 tasks, single run. + * - Coding: 40 tasks x 2 runs (accuracy is over all 80 outcomes). */ import type { HarnessKind } from "../components/harness-logo"; @@ -20,14 +23,16 @@ export interface BenchResult { emphasized?: boolean; } -const MODEL = "DeepSeek V4 Pro"; +const DEEPSEEK = "DeepSeek V4 Pro"; +const OPUS = "Claude Opus 4.8"; +const GPT = "GPT-5.5"; export const DATA_BENCH: BenchResult[] = [ { kind: "penguin", framework: "PenguinHarness", - model: MODEL, - accuracyPct: 66.7, + model: DEEPSEEK, + accuracyPct: 66.67, tokensM: 18.037757, costUsd: 0.552406, emphasized: true, @@ -35,18 +40,18 @@ export const DATA_BENCH: BenchResult[] = [ { kind: "claude", framework: "Claude Code", - model: MODEL, - accuracyPct: 66.7, - tokensM: 21.166305, - costUsd: 0.640706, + model: OPUS, + accuracyPct: 53.33, + tokensM: 22.197759, + costUsd: 38.479975, }, { kind: "codex", framework: "OpenAI Codex", - model: MODEL, - accuracyPct: 46.7, - tokensM: 13.362259, - costUsd: 0.427011, + model: GPT, + accuracyPct: 53.33, + tokensM: 13.72473, + costUsd: 19.413715, }, ]; @@ -54,46 +59,55 @@ export const CODE_BENCH: BenchResult[] = [ { kind: "penguin", framework: "PenguinHarness", - model: MODEL, - accuracyPct: 50.0, - tokensM: 2.1, - costUsd: 0.0413, + model: DEEPSEEK, + accuracyPct: 71.25, + tokensM: 200.0, + costUsd: 3.812, emphasized: true, }, { kind: "claude", framework: "Claude Code", - model: MODEL, - accuracyPct: 48.75, - tokensM: 2.0, - costUsd: 0.0483, + model: OPUS, + accuracyPct: 86.25, + tokensM: 151.61, + costUsd: 146.97, }, { kind: "codex", framework: "OpenAI Codex", - model: MODEL, - accuracyPct: 42.5, - tokensM: 2.65, - costUsd: 0.0427, + model: GPT, + accuracyPct: 71.25, + tokensM: 251.2, + costUsd: 220.08, }, ]; -/** 66.7 -> "66.7%" (chart caps). */ +/** 66.67 -> "66.67%" (chart caps). */ export function formatPct(pct: number): string { - return `${pct.toFixed(1)}%`; + return `${pct.toFixed(2)}%`; } -/** Table accuracy at suite precision: 66.7 (1dp) vs 48.75 (2dp). */ +/** Table accuracy; both suites publish at 2dp (10/15 -> 66.67, 57/80 -> 71.25). */ export function formatAccuracy(pct: number, dp: number): string { return pct.toFixed(dp); } -/** 18.037757 -> "18.0M" / 2.65 -> "2.65M" (chart caps at suite precision). */ -export function formatTokensM(tokens: number, dp = 1): string { +/** 18.037757 -> "18.04M" / 200 -> "200.00M". */ +export function formatTokensM(tokens: number, dp = 2): string { return `${tokens.toFixed(dp)}M`; } -/** Chart cost caps at suite precision: $0.55 vs $0.041. */ +/** Chart cost caps: $0.55 vs $220.08 — one format across a three-order-of-magnitude spread. */ export function formatUsd(cost: number, dp = 2): string { return `$${cost.toFixed(dp)}`; } + +/** + * How many times more a rival spent for the same suite ($38.48 vs $0.55 -> 70). The + * headline claim is built from the data rather than typed into the copy, so it cannot + * drift when the numbers are refreshed. + */ +export function costMultiple(rival: BenchResult, ours: BenchResult): number { + return Math.round(rival.costUsd / ours.costUsd); +} diff --git a/packages/landing/src/lib/links.ts b/packages/landing/src/lib/links.ts index 384c94d..f704214 100644 --- a/packages/landing/src/lib/links.ts +++ b/packages/landing/src/lib/links.ts @@ -11,8 +11,26 @@ export const LICENSE_URL = `${REPO_URL}/blob/main/LICENSE`; */ export const DOCS_URL = `${import.meta.env.BASE_URL}docs/`; -/** One-line installer (Linux / macOS, x64 / arm64, bundled Node runtime). */ -export const INSTALL_CMD = `curl -fsSL ${REPO_URL}/releases/latest/download/install.sh | sh`; +/** + * One-line installer (Linux / macOS, x64 / arm64, bundled Node runtime). + * penguin.ooo/install.sh is this site's own public/install.sh — a thin forwarder + * to the latest GitHub release installer (Pages cannot serve real redirects). + */ +export const INSTALL_CMD = "curl -fsSL https://penguin.ooo/install.sh | sh"; + +/** + * Demo videos live in the sibling `penguin-harness-community` repo rather than in this + * one: they are ~9 MB each, and this repo's whole history is ~17 MB, so committing them + * here would triple what every contributor clones for assets only the marketing site + * shows. Served by raw.githubusercontent with `accept-ranges: bytes` (seeking works), + * `access-control-allow-origin: *` and a 5-minute cache. The `application/octet-stream` + * content type does not block playback — `nosniff` is only enforced for scripts and + * styles, and `
+ {/* The loop above is a diagram; this is the same loop actually running. Replaces the + "demo video coming soon" pill that stood here while the recording was in progress. */} +
+ +
); } diff --git a/packages/landing/src/sections/showcase.tsx b/packages/landing/src/sections/showcase.tsx deleted file mode 100644 index fe0332e..0000000 --- a/packages/landing/src/sections/showcase.tsx +++ /dev/null @@ -1,112 +0,0 @@ -/** - * Use cases: daily tasks + zero-code AI development, illustrated with real - * screenshots captured from the Web App with Playwright. Twelve variants — three - * views (chat building an Agent app / trace view / evaluation center) x two UI - * languages x two themes (WebP) — and the one shown always matches the visitor's - * active locale and theme. Each figure carries a scenario tag. - */ -import { S } from "../lib/strings"; -import { useLocale } from "../state/locale"; -import type { Locale } from "../state/locale"; -import { Section } from "../components/section"; -import { BrowserFrame } from "../components/browser-frame"; -import chatZhLight from "../assets/shots/chat-zh-light.webp"; -import chatZhDark from "../assets/shots/chat-zh-dark.webp"; -import chatEnLight from "../assets/shots/chat-en-light.webp"; -import chatEnDark from "../assets/shots/chat-en-dark.webp"; -import tracesZhLight from "../assets/shots/traces-zh-light.webp"; -import tracesZhDark from "../assets/shots/traces-zh-dark.webp"; -import tracesEnLight from "../assets/shots/traces-en-light.webp"; -import tracesEnDark from "../assets/shots/traces-en-dark.webp"; -import benchmarkZhLight from "../assets/shots/benchmark-zh-light.webp"; -import benchmarkZhDark from "../assets/shots/benchmark-zh-dark.webp"; -import benchmarkEnLight from "../assets/shots/benchmark-en-light.webp"; -import benchmarkEnDark from "../assets/shots/benchmark-en-dark.webp"; - -type ShotSet = Record; - -const SHOTS: Record<"chat" | "traces" | "benchmark", ShotSet> = { - chat: { - zh: { light: chatZhLight, dark: chatZhDark }, - en: { light: chatEnLight, dark: chatEnDark }, - }, - traces: { - zh: { light: tracesZhLight, dark: tracesZhDark }, - en: { light: tracesEnLight, dark: tracesEnDark }, - }, - benchmark: { - zh: { light: benchmarkZhLight, dark: benchmarkZhDark }, - en: { light: benchmarkEnLight, dark: benchmarkEnDark }, - }, -}; - -function ThemedShot({ set, alt }: { set: ShotSet; alt: string }) { - const { locale } = useLocale(); - const pair = set[locale]; - // Explicit intrinsic size (all shots are 1920x1200): reserves layout before the - // image loads, so in-page anchor jumps don't drift when the showcase pops in. - return ( - <> - {alt} - {alt} - - ); -} - -function Caption({ tag, text }: { tag: string; text: string }) { - return ( -
- - {tag} - - {text} -
- ); -} - -export function Showcase() { - return ( -
-
- - - - -
- -
-
- - - - -
-
- - - - -
-
-
- ); -} diff --git a/packages/landing/src/styles.css b/packages/landing/src/styles.css index 1270c71..5cb87a2 100644 --- a/packages/landing/src/styles.css +++ b/packages/landing/src/styles.css @@ -331,6 +331,31 @@ } } +/* ---------- Building-blocks loop (LangChain comparison): each block pops in at its + negative delay within the shared cycle, holds while the rest of the tower lands, + then the whole skyline fades and rebuilds. Reduced-motion shows towers complete. */ +.anim-block { + animation: block-build var(--cycle, 7.2s) linear infinite; +} +@keyframes block-build { + 0% { + opacity: 0; + transform: translateY(5px) scale(0.92); + } + 3% { + opacity: 1; + transform: none; + } + 93% { + opacity: 1; + transform: none; + } + 100% { + opacity: 0; + transform: none; + } +} + /* Printing never scrolls, so scroll-reveal must not hide content on paper/PDF. */ @media print { .reveal { diff --git a/packages/landing/test/benchmark-data.test.ts b/packages/landing/test/benchmark-data.test.ts index d10a56a..134e08f 100644 --- a/packages/landing/test/benchmark-data.test.ts +++ b/packages/landing/test/benchmark-data.test.ts @@ -2,31 +2,55 @@ import { describe, expect, it } from "vitest"; import { CODE_BENCH, DATA_BENCH, + costMultiple, formatAccuracy, formatPct, formatTokensM, formatUsd, } from "../src/lib/benchmark-data"; -describe("benchmark data (unified per-run means)", () => { +describe("benchmark data (unified suite totals)", () => { it("formats the data-analysis suite at its published precision", () => { const penguin = DATA_BENCH[0]!; - expect(formatPct(penguin.accuracyPct)).toBe("66.7%"); - expect(formatAccuracy(penguin.accuracyPct, 1)).toBe("66.7"); + expect(formatPct(penguin.accuracyPct)).toBe("66.67%"); + expect(formatAccuracy(penguin.accuracyPct, 2)).toBe("66.67"); expect(formatTokensM(penguin.tokensM, 2)).toBe("18.04M"); - expect(formatUsd(penguin.costUsd, 3)).toBe("$0.552"); + expect(formatUsd(penguin.costUsd, 2)).toBe("$0.55"); }); - it("formats the coding suite at its published precision (CNY converted at 7:1)", () => { + it("formats the coding suite at its published precision", () => { const penguin = CODE_BENCH[0]!; - expect(formatAccuracy(penguin.accuracyPct, 2)).toBe("50.00"); - expect(formatTokensM(penguin.tokensM, 2)).toBe("2.10M"); - expect(formatUsd(penguin.costUsd, 3)).toBe("$0.041"); - // 0.289 CNY / 7 -> ~0.0413 USD - expect(penguin.costUsd).toBeCloseTo(0.289 / 7, 3); + expect(formatAccuracy(penguin.accuracyPct, 2)).toBe("71.25"); + expect(formatTokensM(penguin.tokensM, 2)).toBe("200.00M"); + expect(formatUsd(penguin.costUsd, 2)).toBe("$3.81"); }); - it("uses the unified framework names with PenguinHarness as the only emphasized row", () => { + it("keeps every published accuracy on its suite's grid (n=15 and n=80)", () => { + // 10/15 -> 66.67, 8/15 -> 53.33; 57/80 -> 71.25, 69/80 -> 86.25. A number off the grid + // means a transcription slip, which the charts would render without complaint. + for (const row of DATA_BENCH) + expect(Math.round((row.accuracyPct / 100) * 15)).toBeCloseTo((row.accuracyPct / 100) * 15, 1); + for (const row of CODE_BENCH) expect(((row.accuracyPct / 100) * 80) % 1).toBeCloseTo(0, 6); + }); + + it("backs the copy's cost claims: 35x/70x on data analysis, 58x/39x on coding", () => { + const [dPenguin, dClaude, dCodex] = DATA_BENCH as [ + (typeof DATA_BENCH)[0], + (typeof DATA_BENCH)[0], + (typeof DATA_BENCH)[0], + ]; + expect(costMultiple(dCodex, dPenguin)).toBe(35); + expect(costMultiple(dClaude, dPenguin)).toBe(70); + const [cPenguin, cClaude, cCodex] = CODE_BENCH as [ + (typeof CODE_BENCH)[0], + (typeof CODE_BENCH)[0], + (typeof CODE_BENCH)[0], + ]; + expect(costMultiple(cCodex, cPenguin)).toBe(58); + expect(costMultiple(cClaude, cPenguin)).toBe(39); + }); + + it("pairs each harness with its own model and emphasizes only PenguinHarness", () => { for (const suite of [DATA_BENCH, CODE_BENCH]) { expect(suite.map((r) => r.framework)).toEqual([ "PenguinHarness", @@ -34,7 +58,7 @@ describe("benchmark data (unified per-run means)", () => { "OpenAI Codex", ]); expect(suite.filter((r) => r.emphasized).map((r) => r.framework)).toEqual(["PenguinHarness"]); - for (const row of suite) expect(row.model).toBe("DeepSeek V4 Pro"); + expect(suite.map((r) => r.model)).toEqual(["DeepSeek V4 Pro", "Claude Opus 4.8", "GPT-5.5"]); } }); }); diff --git a/packages/server/README.md b/packages/server/README.md index b5e6252..8095060 100644 --- a/packages/server/README.md +++ b/packages/server/README.md @@ -37,7 +37,7 @@ pnpm --filter @prismshadow/penguin-server start # node dist/index.js - **CSRF**: session cookie is `SameSite=Lax` and writes accept only `Content-Type: application/json`; no CSRF token yet. - **No login rate limiting**: add throttling at a reverse proxy for public deployments. -- **Built-in admin starts as `admin` / `admin123`**: change it immediately (a banner keeps reminding until you do). +- **Built-in admin starts as `admin` / `penguin-2026`**: change it immediately (a banner keeps reminding until you do). - Passwords use `node:crypto` scrypt (`scrypt$N$r$p$salt$hash`, timingSafeEqual); login sessions renew on a 7-day sliding window; the DB stores only the token's sha256. - Model credentials live in the Project's hidden 0600 config file; the API always masks them. - Behind a reverse proxy, disable response buffering for SSE paths (the server already sends `X-Accel-Buffering: no`) and forward `x-forwarded-proto` to enable Secure cookies. diff --git a/packages/server/package.json b/packages/server/package.json index 3c44f6b..e4e7261 100644 --- a/packages/server/package.json +++ b/packages/server/package.json @@ -25,7 +25,7 @@ "node": ">=24" }, "scripts": { - "dev": "pnpm --filter @prismshadow/penguin-core build && tsx watch src/index.ts", + "dev": "node ../../scripts/dev-prebuild.mjs && tsx watch src/index.ts", "start": "node --disable-warning=ExperimentalWarning dist/index.js", "typecheck": "tsc --noEmit -p tsconfig.json", "test": "vitest run --passWithNoTests", diff --git a/packages/server/src/api/types.ts b/packages/server/src/api/types.ts index 4b03b6e..47f8ac6 100644 --- a/packages/server/src/api/types.ts +++ b/packages/server/src/api/types.ts @@ -273,6 +273,8 @@ export interface ModelTestRequest { apiKey?: string; /** "Clear saved API key" is checked: the test does **not** fall back to the stored key (tests against the current draft). */ clearApiKey?: boolean; + /** Speed-test mode: raises the probe's output cap (16 -> 64 tokens) so TTFT/TPS are measurable; costs a little more quota. */ + speed?: boolean; /** * base URL (not secret; the frontend always sends the form's current value): a string * means use it, `null` means explicitly clear it (no fallback to the stored value), @@ -283,10 +285,23 @@ export interface ModelTestRequest { clientType?: string; } -/** Connectivity test result: carries round-trip latency when ok, and a reason on failure (truncated raw provider error). */ +/** + * Connectivity test result: carries round-trip latency when ok, and a reason on failure + * (truncated raw provider error). When streamed content was observed, also carries the + * time-to-first-token and, when usage was reported (completed streams), the output rate. + */ export interface ModelTestResponse { ok: boolean; latencyMs?: number; + /** Time from request start to the first streamed content (thinking or text), ms. */ + ttftMs?: number; + /** + * Output tokens per second over the streaming window (first content -> stream end), 1dp. + * Omitted unless the sample is large enough to mean anything: a reply of a few tokens is + * dominated by the final chunk's round trip, so the rate it yields tracks network jitter + * rather than the model. Callers render TTFT alone in that case. + */ + tps?: number; message?: string; } @@ -342,6 +357,8 @@ export interface AgentSummary { vaultKeyCount: number; /** Schedule count (number of .toml files under agent_state/schedule/, including invalid ones). */ scheduleCount: number; + /** Installed Skill count (number of agent_state/skills// directories with a SKILL.md). */ + skillCount: number; } export interface AgentsResponse { @@ -460,12 +477,12 @@ export interface DirListResponse { } export interface SessionCreateRequest { - /** Upstream id of the session's model (paired with provider); defaults to the Project's default Model. */ + /** Upstream id of the session's model; always sent together with provider. Omit both for the Project's default Model. */ modelId?: string; /** - * Provider group for `modelId`; when omitted, resolved via resolveModelRef semantics — - * modelId can only be resolved if it's globally unique by exact match in the config; - * 0 or multiple matches return 400. + * Provider group for `modelId`. A model reference is always a complete + * (provider, modelId) pair — the provider is never inferred, so sending one field + * without the other returns 400 instead of being resolved. */ provider?: string; /** Any existing directory on the server; defaults to auto-creating a temporary Workspace. */ @@ -930,9 +947,9 @@ export interface ScheduleItem { /** Bound target Session; defaults to creating a new Session each time. */ sessionId?: string; workspace?: string; - /** Model for new-Session mode (upstream id, paired with provider); defaults to the Project's default reference. */ + /** Model for new-Session mode (upstream id, always paired with provider); absent means the Project's default reference. */ modelId?: string; - /** Provider group for `modelId`; when omitted, resolved via resolveModelRef semantics (resolvable only on a unique match). */ + /** Provider group for `modelId`; present exactly when `modelId` is — a model reference is always a pair. */ provider?: string; status: ScheduleStatus; invalidReason?: string; @@ -959,9 +976,12 @@ export interface ScheduleUpsertRequest { endAt?: string; sessionId?: string; workspace?: string; - /** Model for new-Session mode (upstream id); defaults to the Project's default reference. */ + /** Model for new-Session mode (upstream id); always sent together with provider, omit both for the Project's default reference. */ modelId?: string; - /** Provider group for `modelId`; when omitted, validated as uniquely resolvable via resolveModelRef semantics at save/reconciliation time. */ + /** + * Provider group for `modelId`. Both fields are sent as a pair (400 otherwise); the + * pair is checked against the Project config at save/reconciliation time. + */ provider?: string; } diff --git a/packages/server/src/auth/service.ts b/packages/server/src/auth/service.ts index 279521b..d1ba279 100644 --- a/packages/server/src/auth/service.ts +++ b/packages/server/src/auth/service.ts @@ -3,7 +3,7 @@ * login / logout / password change / session validation. * * - No open registration: on startup, if there are no users at all, the built-in - * admin `admin` is seeded (initial password admin123), and it adopts + * admin `admin` is seeded (initial password penguin-2026), and it adopts * `default_project`; all other users are created by an admin via the user * backend (admin-service). * - An initial password (whether seeded or set by an admin) is flagged with @@ -22,7 +22,7 @@ export const MIN_PASSWORD_LENGTH = 8; /** Built-in admin: user_id and initial password (matches the README and login-page hint). */ export const ADMIN_USER_ID = "admin"; -export const ADMIN_INITIAL_PASSWORD = "admin123"; +export const ADMIN_INITIAL_PASSWORD = "penguin-2026"; function sha256Hex(value: string): string { return createHash("sha256").update(value).digest("hex"); diff --git a/packages/server/src/http/routes/models.ts b/packages/server/src/http/routes/models.ts index 93d199f..d197797 100644 --- a/packages/server/src/http/routes/models.ts +++ b/packages/server/src/http/routes/models.ts @@ -160,6 +160,10 @@ export function modelsRoutes(deps: AppDeps): Hono { if (typeof body.clearApiKey !== "boolean") throw badRequest("clearApiKey must be a boolean."); req.clearApiKey = body.clearApiKey; } + if (body.speed !== undefined) { + if (typeof body.speed !== "boolean") throw badRequest("speed must be a boolean."); + req.speed = body.speed; + } // null = explicit clear (test against the draft, don't fall back to the stored value); empty string is treated as null. if (body.baseUrl !== undefined) { if (body.baseUrl !== null && typeof body.baseUrl !== "string") { diff --git a/packages/server/src/http/routes/schedules.ts b/packages/server/src/http/routes/schedules.ts index 7786b20..7eb3a43 100644 --- a/packages/server/src/http/routes/schedules.ts +++ b/packages/server/src/http/routes/schedules.ts @@ -213,7 +213,8 @@ async function upsert( const raw = serializeSchedule(fields); const parsed = parseScheduleFile(name, raw); if (!parsed.ok) throw badRequest(`Invalid schedule configuration: ${parsed.error}`); - // At save time, verify the model reference resolves (resolveModelRef semantics; same rules as reconciliation) so we never persist a broken file. + // At save time, verify the (provider, modelId) pair names a configured model (same rules as reconciliation) so we never persist a broken file. + // The pairing rule itself is enforced by parseScheduleFile above, which rejects half a reference. const refError = await validateScheduleModelRef(deps.config.root, projectId, parsed.def); if (refError !== null) throw badRequest(`Invalid schedule configuration: ${refError}`); await writeScheduleFile(deps.config.root, projectId, agentId, name, raw); diff --git a/packages/server/src/http/routes/sessions.ts b/packages/server/src/http/routes/sessions.ts index 8c7f7c5..659e614 100644 --- a/packages/server/src/http/routes/sessions.ts +++ b/packages/server/src/http/routes/sessions.ts @@ -108,10 +108,13 @@ export function agentSessionsRoutes(deps: AppDeps): Hono { const body = await readJson(c); const modelId = optionalString(body, "modelId", { minLen: 1, label: "modelId" }); const provider = optionalString(body, "provider", { minLen: 1, label: "provider" }); - // Model reference is submitted as a pair: provider can't appear without modelId (core does the same validation; this catches it early). - if (provider !== undefined && modelId === undefined) { + // Model reference is submitted as a pair — both or neither. Neither half is ever + // inferred from the other, so half a reference is rejected here instead of being + // resolved (core does the same validation; this catches it early). Omitting both + // falls back to the Project's default model. + if ((modelId === undefined) !== (provider === undefined)) { throw badRequest( - "provider is specified but modelId is not: a model reference must be given as a pair.", + "modelId and provider must be given together as a model reference pair: specify both, or neither to use the Project's default model.", ); } const approvalMode = optionalEnum(body, "approvalMode", APPROVAL_MODES); @@ -361,6 +364,11 @@ export function sessionsRoutes(deps: AppDeps): Hono { const row = resolveSession(c); const rel = c.req.query("path") ?? ""; const download = c.req.query("download") === "1"; + // Sandboxed top-level preview ("open in a new tab" for html): the document keeps its REAL + // content type but carries a CSP sandbox WITHOUT allow-same-origin — it renders and runs + // fully in an opaque origin, so agent-generated markup cannot reach this origin's cookies + // or API. The request itself still authenticates (top-level GET sends the Lax cookie). + const preview = !download && c.req.query("preview") === "1"; const { data, fileName, contentType, scriptable } = await deps.workspaceFiles.read( row.workspace, rel, @@ -368,15 +376,22 @@ export function sessionsRoutes(deps: AppDeps): Hono { const disposition = download ? "attachment" : "inline"; // Same-origin XSS defense: html/svg inline previews are always returned as plain // text (Workspace files may be Agent-generated and untrusted); downloads - // (attachment) keep the real content type. Paired with nosniff to prevent MIME - // sniffing from undoing this. - const effectiveType = !download && scriptable ? "text/plain; charset=utf-8" : contentType; + // (attachment) keep the real content type, and sandboxed previews keep it under the + // CSP above. Paired with nosniff to prevent MIME sniffing from undoing this. + const effectiveType = + !download && scriptable && !preview ? "text/plain; charset=utf-8" : contentType; return new Response(new Uint8Array(data), { status: 200, headers: { "Content-Type": effectiveType, "Content-Disposition": `${disposition}; filename*=UTF-8''${encodeURIComponent(fileName)}`, "X-Content-Type-Options": "nosniff", + ...(preview && scriptable + ? { + "Content-Security-Policy": + "sandbox allow-scripts allow-popups allow-modals allow-forms", + } + : {}), }, }); }); diff --git a/packages/server/src/http/routes/vault.ts b/packages/server/src/http/routes/vault.ts index 5010f3c..d733ad0 100644 --- a/packages/server/src/http/routes/vault.ts +++ b/packages/server/src/http/routes/vault.ts @@ -48,9 +48,10 @@ export function vaultRoutes(deps: AppDeps): Hono { deps.projectService.requireProjectOwner(c.var.user.userId, projectId); const req = parseVaultUpdate(await readJson(c)); const res = await deps.agentConfigService.updateVault(projectId, agentId, req); - // Effective-value semantics: no hot update — an already-built runtime is neither - // evicted nor reloaded; the new value only applies to Sessions created or resumed - // afterward. + // Effective-value semantics: no hot swap into a Task already in flight, but every + // runtime built before this update is invalidated — the next Task on any Session + // of this Agent re-resumes and picks up the new vault values. + deps.manager.invalidateAgentRuntimes(projectId, agentId); return c.json(res); }); diff --git a/packages/server/src/index.ts b/packages/server/src/index.ts index 41d366c..c23b8b8 100644 --- a/packages/server/src/index.ts +++ b/packages/server/src/index.ts @@ -20,7 +20,7 @@ const config = resolveServerConfig(); const deps = buildAppDeps(config); const app = createApp(deps); -// Built-in admin seed (idempotent): creates admin (initial password admin123) and adopts default_project when the users table is empty. +// Built-in admin seed (idempotent): creates admin (initial password penguin-2026) and adopts default_project when the users table is empty. await deps.authService.seedAdmin(); // Schedule scheduler: startup reconciliation (missed, don't backfill) + periodic scan; only active while the server is running. diff --git a/packages/server/src/runtime/channel.ts b/packages/server/src/runtime/channel.ts index d9b2216..d9465d8 100644 --- a/packages/server/src/runtime/channel.ts +++ b/packages/server/src/runtime/channel.ts @@ -8,7 +8,7 @@ * recycled/recreated or the process restarts, so a stale Last-Event-ID always misses * and falls through to resync — this prevents a silent false-hit event loss when the * new epoch's event count happens to exceed the old id; - * - A bounded ring buffer (most recent 1000 entries or 2MB, whichever comes first, + * - A bounded ring buffer (most recent 10,000 entries or 8MB, whichever comes first, * evicting the oldest on overflow) serves replay-on-reconnect via `Last-Event-ID`; * an evicted/unknown id is handled by the caller sending `resync_required`; * - Unicast (sendTo) is used for one-off replay at subscribe time (pending approvals / @@ -37,8 +37,12 @@ export interface ChannelOptions { maxBufferBytes?: number; } -const DEFAULT_MAX_COUNT = 1000; -const DEFAULT_MAX_BYTES = 2 * 1024 * 1024; +// Sized so a mid-stream reconnect during a fast large-code reply still hits replay instead of +// resync_required: the server publishes one event per provider delta (a 240KB reply ≈ 5k events +// / 1.4MB), and 1000 events covered only ~6s of such a stream — every longer blip forced a full +// client-side history rebuild. +const DEFAULT_MAX_COUNT = 10_000; +const DEFAULT_MAX_BYTES = 8 * 1024 * 1024; /** Buffered entry: seq is stored separately so hit checks never need to parse the string id. */ interface BufferedEvent { diff --git a/packages/server/src/runtime/schedule-file.ts b/packages/server/src/runtime/schedule-file.ts index d73d6e0..7db9d53 100644 --- a/packages/server/src/runtime/schedule-file.ts +++ b/packages/server/src/runtime/schedule-file.ts @@ -35,13 +35,13 @@ export interface ScheduleDefinition { sessionId?: string; /** Workspace for new-Session mode (same semantics as manually starting a session; auto-creates a temp directory if unspecified). */ workspace?: string; - /** Model for new-Session mode (upstream id, paired with provider; defaults to the Project's default reference). */ + /** Model for new-Session mode (upstream id, always paired with provider; omit both for the Project's default reference). */ modelId?: string; /** - * Vendor grouping for `model_id` (paired reference); when omitted, resolved per - * resolveModelRef semantics — whether the reference is resolvable is validated by the - * caller against config at reconciliation/save time (this module does pure parsing - * and never touches config). + * Vendor grouping for `model_id`; always present when `model_id` is (the pairing rule + * is enforced here). Whether that pair names a configured model is validated by the + * caller against config at reconciliation/save time — this module does pure parsing + * and never touches config. */ provider?: string; } @@ -148,13 +148,8 @@ export function parseScheduleFile(name: string, raw: string): ScheduleParseResul } provider = t["provider"]; } - if (provider !== undefined && modelId === undefined) { - return { - ok: false, - error: - "provider is only used together with model_id (a model reference must be given as a pair)", - }; - } + // Target conflicts are reported first: when session_id is set the whole model reference + // is out of place, so complaining about its shape would point at the wrong fix. if ( sessionId !== undefined && (workspace !== undefined || modelId !== undefined || provider !== undefined) @@ -164,6 +159,18 @@ export function parseScheduleFile(name: string, raw: string): ScheduleParseResul error: "Pick one target: workspace and provider / model_id are only for new-Session mode", }; } + // A model reference is always a complete (provider, model_id) pair: neither half is + // ever inferred from the other, so a file carrying only one of them is invalid rather + // than something to resolve later. Omit both to run on the Project's default model. + // A file written before this rule existed (model_id alone) therefore parses as + // invalid: it is listed in invalidFiles and skipped, never scheduled. + if ((modelId === undefined) !== (provider === undefined)) { + return { + ok: false, + error: + "provider and model_id must be given together (a model reference is always a pair); omit both to use the Project's default model", + }; + } return { ok: true, diff --git a/packages/server/src/runtime/schedule-store.ts b/packages/server/src/runtime/schedule-store.ts index 46b58cd..da07ed0 100644 --- a/packages/server/src/runtime/schedule-store.ts +++ b/packages/server/src/runtime/schedule-store.ts @@ -88,19 +88,23 @@ export function serializeSchedule(fields: { } /** - * Resolvability check for a schedule's model reference (shared by save and - * reconciliation): when the definition has `model_id`, it's - * resolved against Project config per resolveModelRef semantics — omitting provider is - * only resolvable when model_id matches exactly one entry globally; zero hits or - * ambiguity means unresolvable. Returns an error message (unresolvable / config read - * failure), or null if resolvable (or no model reference at all). + * Config check for a schedule's model reference (shared by save and reconciliation): + * the reference is a complete `(provider, model_id)` pair and must name an entry in the + * Project config — nothing is inferred, so a pair that matches no entry is simply an + * error. Returns an error message (no such model / half a reference / config read + * failure), or null when the pair is configured (or there is no model reference at all). */ export async function validateScheduleModelRef( root: string, projectId: string, def: Pick, ): Promise { - if (def.modelId === undefined) return null; + if (def.modelId === undefined && def.provider === undefined) return null; + // parseScheduleFile already rejects half a reference, so this only guards callers that + // build a definition by hand — the missing half is reported, never guessed. + if (def.modelId === undefined || def.provider === undefined) { + return "provider and model_id must be given together (a model reference is always a pair); omit both to use the Project's default model"; + } try { const cfg = await loadProjectConfig(root, projectId); resolveModelRef(cfg, def.modelId, def.provider); diff --git a/packages/server/src/runtime/scheduler.ts b/packages/server/src/runtime/scheduler.ts index 1a7e75b..721304d 100644 Binary files a/packages/server/src/runtime/scheduler.ts and b/packages/server/src/runtime/scheduler.ts differ diff --git a/packages/server/src/runtime/session-manager.ts b/packages/server/src/runtime/session-manager.ts index b3d077a..7bb79a3 100644 --- a/packages/server/src/runtime/session-manager.ts +++ b/packages/server/src/runtime/session-manager.ts @@ -8,6 +8,9 @@ * createSession using the index row's workspace/modelId, yielding a new session_id * and updating the index's primary key; the Task response body always returns the * current actual id; + * - Vault effectiveness: a vault update bumps the Agent's config generation + * (invalidateAgentRuntimes); runtimes built earlier are discarded on their next + * idle access and re-resumed, so the next Task always runs with current values; * - Per-Session mutual exclusion: only one Task/compaction may be in progress at a * time; * - run/compact drive: consumes the output stream in the background, publishing each @@ -187,6 +190,12 @@ interface RuntimeEntry { abort: AbortController | null; /** The in-flight drive Promise (awaited during graceful shutdown). */ running: Promise | null; + /** + * Agent config generation this runtime was built under (see + * invalidateAgentRuntimes): once it falls behind the Agent's current generation, + * the entry is discarded on its next idle access and re-resumed via the loader. + */ + generation: number; /** Timestamp of last activity (refreshed on load / status flip / drive completion), used for idle-eviction checks. */ lastActivityMs: number; } @@ -197,6 +206,13 @@ const ENTRY_SWEEP_INTERVAL_MS = 60 * 1000; /** Cap on collected model text for title material (accumulation stops beyond this; the generator side also truncates further). */ const TITLE_EXCERPT_LIMIT = 4000; +/** + * Early title trigger: once this many characters of main-session body text have streamed, + * title generation starts right away instead of waiting for the Task to finish — the core + * Session has captured its (capped) material by then, and a long answer would only overrun + * it. Short conversations are still covered by the completion trigger in drive's finally. + */ +const EARLY_TITLE_BODY_CHARS = 1000; /** Composite Agent key (used as a Set key, avoiding projectId/agentId concatenation ambiguity). */ function agentKey(projectId: string, agentId: string): string { @@ -280,6 +296,8 @@ export class SessionManager { private readonly deletingAgents = new Set(); /** Sessions currently being deleted (guards against the entry/Trace file being rebuilt and reviving it inside the deletion race window). */ private readonly deletingSessions = new Set(); + /** Per-Agent config generation (key = agentKey), bumped by invalidateAgentRuntimes on vault updates. */ + private readonly agentGenerations = new Map(); private readonly sweepTimer: NodeJS.Timeout; constructor(private readonly deps: SessionManagerDeps) { @@ -324,10 +342,25 @@ export class SessionManager { approvals: new ApprovalRegistry(), abort: null, running: null, + generation: this.generationOf(row.projectId, row.agentId), lastActivityMs: Date.now(), }); } + /** + * After an Agent's vault is updated: bump the Agent's config generation so every + * runtime built before the update is discarded on its next idle access and + * re-resumed via the loader — resume re-reads agent_state/.vault.toml, so the next + * Task on any of this Agent's Sessions runs with the new values (history is + * preserved through the Trace). A Task already in flight is neither aborted nor + * hot-swapped: it keeps the values it started with, and its entry is rebuilt on + * the first access after it returns to idle (see ensureEntry). + */ + invalidateAgentRuntimes(projectId: string, agentId: string): void { + const key = agentKey(projectId, agentId); + this.agentGenerations.set(key, this.generationOf(projectId, agentId) + 1); + } + // —— Task / compaction drive —— /** @@ -578,10 +611,30 @@ export class SessionManager { } } + private generationOf(projectId: string, agentId: string): number { + return this.agentGenerations.get(agentKey(projectId, agentId)) ?? 0; + } + /** get-or-resume-or-heal: use directly on an active-table hit; otherwise load via the loader, updating the index's primary key on self-heal. */ private async ensureEntry(sessionId: string): Promise { const existing = this.entries.get(sessionId); - if (existing) return existing; + if (existing) { + if (existing.generation === this.generationOf(existing.projectId, existing.agentId)) { + return existing; + } + // Built before the last vault update: discard once idle and fall through to a + // fresh load (resume re-reads the vault). A busy entry is returned as-is — the + // in-flight run keeps its values and assertIdle rejects the new Task anyway; + // it is rebuilt on the first access after it finishes. + if ( + existing.status !== "idle" || + existing.running !== null || + existing.approvals.size !== 0 + ) { + return existing; + } + this.entries.delete(sessionId); + } const row = this.deps.sessions.findById(sessionId); if (!row) { throw new HttpError( @@ -590,6 +643,9 @@ export class SessionManager { "Session does not exist or you do not have access.", ); } + // Captured before the (awaited) load: a vault update racing with the load leaves + // this entry stale, so the access after next rebuilds it with the new values. + const generation = this.generationOf(row.projectId, row.agentId); const session = await this.deps.loader.load(row); // The Session/Agent was marked for deletion while loading: discard the load result, // don't rebuild the entry (avoids reviving an orphaned Trace). @@ -613,6 +669,7 @@ export class SessionManager { approvals: new ApprovalRegistry(), abort: null, running: null, + generation, lastActivityMs: Date.now(), }; this.entries.set(currentId, entry); @@ -630,6 +687,8 @@ export class SessionManager { gen: AsyncGenerator, titleSource?: { userExcerpt: string }, ): Promise { + let earlyTitleFired = false; + let mainBodyChars = 0; const ctx: UsageContext = { projectId: entry.projectId, agentId: entry.agentId, @@ -715,6 +774,27 @@ export class SessionManager { child.assistantExcerpt += (child.assistantExcerpt ? "\n" : "") + nested.text; } } + // Early title trigger (see EARLY_TITLE_BODY_CHARS): fire as soon as enough main- + // session body text has streamed; maybeGenerate self-guards (NULL title, single + // flight), so the completion trigger in finally stays as the short-answer fallback. + if (!earlyTitleFired && titleSource?.userExcerpt.trim()) { + const p = msg.payload as { type?: string; role?: string; text?: string }; + if ( + (!msg.origin || msg.origin.length === 0) && + msg.type === "model_msg" && + p.type === "text" && + p.role === "assistant" && + p.text + ) { + mainBodyChars += p.text.length; + if (mainBodyChars >= EARLY_TITLE_BODY_CHARS) { + earlyTitleFired = true; + this.deps.titles?.maybeGenerate(ctx, entry.session, { + fallbackText: titleSource.userExcerpt, + }); + } + } + } // Re-fetch the channel before every publish (matches publishEvent): the channel // may have been recycled and recreated during a long wait on approval, and // holding a stale reference would send output to an orphaned, detached channel. diff --git a/packages/server/src/runtime/title-generator.ts b/packages/server/src/runtime/title-generator.ts index 33ed640..00fbc6b 100644 --- a/packages/server/src/runtime/title-generator.ts +++ b/packages/server/src/runtime/title-generator.ts @@ -14,7 +14,12 @@ * - Silent failure (logged): the title stays NULL and naturally retries after the next * Task completes. */ -import { emptyTokenCounts, sanitizeTitle, tokenUsage } from "@prismshadow/penguin-core"; +import { + emptyTokenCounts, + sanitizeTitle, + stripConversationMarkers, + tokenUsage, +} from "@prismshadow/penguin-core"; import type { SessionsRepo } from "../db/repos/sessions.js"; import type { ChannelHub } from "./channel.js"; import type { ErrorSink } from "./error-recorder.js"; @@ -132,7 +137,11 @@ export class TitleGenerator implements TitleNotifier { /** Fallback title: take the material's first non-empty line, sanitize and truncate; if sanitizing empties it out (pure punctuation, etc.) fall back to the truncated original text; returns null if all-whitespace. */ function fallbackTitle(text: string): string | null { - const firstLine = text.split("\n").find((l) => l.trim().length > 0); + // Strip machine markers first: a skill invocation prepends a `` block, so the + // raw first non-empty line would otherwise be that marker rather than the user's request. + const firstLine = stripConversationMarkers(text) + .split("\n") + .find((l) => l.trim().length > 0); if (!firstLine) return null; // sanitizeTitle strips a pure-punctuation line down to empty — in that case keep the truncated original text, guaranteeing "a title is always obtained". return sanitizeTitle(firstLine) ?? firstLine.trim().slice(0, 30); diff --git a/packages/server/src/services/agent-service.ts b/packages/server/src/services/agent-service.ts index 9343995..018876b 100644 --- a/packages/server/src/services/agent-service.ts +++ b/packages/server/src/services/agent-service.ts @@ -10,6 +10,8 @@ * template's comments). */ import fs from "node:fs/promises"; +import type { Dirent } from "node:fs"; +import path from "node:path"; import { HttpError } from "../http/errors.js"; import { agentDir, @@ -20,6 +22,7 @@ import { isValidId, loadAgentVault, scheduleDir, + skillsDir, systemConfigPath, } from "@prismshadow/penguin-core"; import type { AgentsRepo } from "../db/repos/agents.js"; @@ -41,6 +44,8 @@ export interface AgentListItem { vaultKeyCount: number; /** Number of scheduled tasks (count of .toml files under schedule/, including invalid ones). */ scheduleCount: number; + /** Number of installed Skills (count of skills// directories that contain a SKILL.md). */ + skillCount: number; } export class AgentService { @@ -86,11 +91,12 @@ export class AgentService { ); return Promise.all( sorted.map(async (row) => { - const [meta, updatedAt, vaultKeyCount, scheduleCount] = await Promise.all([ + const [meta, updatedAt, vaultKeyCount, scheduleCount, skillCount] = await Promise.all([ this.agentConfig.readCardMeta(projectId, row.agentId), this.configUpdatedAt(projectId, row.agentId), this.vaultKeyCount(projectId, row.agentId), this.scheduleCount(projectId, row.agentId), + this.skillCount(projectId, row.agentId), ]); return { agentId: row.agentId, @@ -99,6 +105,7 @@ export class AgentService { ...(updatedAt !== undefined ? { updatedAt } : {}), vaultKeyCount, scheduleCount, + skillCount, }; }), ); @@ -123,6 +130,30 @@ export class AgentService { } } + /** Number of installed Skills: count of skills// directories containing a SKILL.md (0 if the directory doesn't exist). */ + private async skillCount(projectId: string, agentId: string): Promise { + const base = skillsDir(this.root, projectId, agentId); + let dirents: Dirent[]; + try { + dirents = await fs.readdir(base, { withFileTypes: true }); + } catch { + return 0; + } + const present = await Promise.all( + dirents + .filter((d) => d.isDirectory()) + .map(async (d) => { + try { + await fs.access(path.join(base, d.name, "SKILL.md")); + return true; + } catch { + return false; + } + }), + ); + return present.filter(Boolean).length; + } + /** Last config modification time: the later of system_config.yaml and AGENTS.md mtime; omitted if neither is readable. */ private async configUpdatedAt(projectId: string, agentId: string): Promise { const paths = [ @@ -218,6 +249,8 @@ export class AgentService { version: meta.version, vaultKeyCount: 0, scheduleCount: 0, + // Read the real count: coreCreateAgent seeds the default skill set for default_agent. + skillCount: await this.skillCount(projectId, agentId), }; } } diff --git a/packages/server/src/services/project-config-service.ts b/packages/server/src/services/project-config-service.ts index 312573a..7779e68 100644 --- a/packages/server/src/services/project-config-service.ts +++ b/packages/server/src/services/project-config-service.ts @@ -27,7 +27,7 @@ import { resolveModelEnv, userText, } from "@prismshadow/penguin-core"; -import type { ModelRef } from "@prismshadow/penguin-core"; +import type { LLMOutcome, ModelRef, OmniMessage } from "@prismshadow/penguin-core"; import type { ModelInfo, ModelPricingDto, @@ -94,6 +94,26 @@ function showRef(provider: string, modelId: string): string { return `(provider=${provider}, model_id=${modelId})`; } +/** + * Connectivity probe prompt: asks for one word, so the whole exchange fits in a + * single-digit output budget. The wording discourages reasoning and the trailing empty + * makes many reasoning models treat their thinking phase as already + * closed - keeping the probe's tiny budget on actual output instead of burning it on + * thinking. + */ +const PROBE_PROMPT = + 'ping - reply with the single word "pong" and nothing else. Do not think or explain.\n'; + +/** + * Speed probe prompt: a one-word answer can't be timed (a compliant model emits 1-3 + * tokens, and the window is then dominated by the final usage chunk's round trip), so + * speed mode asks for something long enough to run into its raised cap. Counting to 50 is + * deterministic and needs no knowledge, so every model produces the same token stream at + * its own decoding rate. Carries the same anti-thinking hint as the connectivity prompt. + */ +const SPEED_PROBE_PROMPT = + "Count from 1 to 50 as a comma-separated list, and nothing else. Do not think or explain.\n"; + export class ProjectConfigService { constructor(private readonly root: string) {} @@ -203,12 +223,19 @@ export class ProjectConfigService { * as a pair in the request body; sends one minimal request using that model's * config (optionally overridden with an unsaved apiKey / baseUrl) — no tools, no * system prompt, thinking disabled, a tiny output cap, 20s timeout — just to see - * whether it completes normally. The model id sent to AgentHub is `modelId` + * whether the endpoint answers. The model id sent to AgentHub is `modelId` * itself (the upstream id verbatim; client_type inference follows it). * + * A reasoning-heavy model may ignore the disabled thinking level and burn the + * whole tiny output cap on thinking (finish_reason=length with no text — AgentHub + * raises EmptyResponseError, collapsed to a malformed outcome): the endpoint + * demonstrably streamed model output, which is everything a connectivity test + * proves, so that case counts as ok too (see probeVerdict). + * * Never throws: the LLM layer collapses auth/parameter/network errors into an * `LLMOutcome`, which is translated here into ok / message. Consumes very few - * Tokens (single-digit output), and writes no Trace and records no usage. + * Tokens (single-digit output; speed mode spends up to its 64-token cap to have a + * window worth timing), and writes no Trace and records no usage. */ async testModel(projectId: string, req: ModelTestRequest): Promise { const raw = await this.readRaw(projectId); @@ -236,18 +263,43 @@ export class ProjectConfigService { ...(clientType ? { clientType } : {}), tools: [], thinkingLevel: "none", - maxTokens: 16, + // Speed mode pairs a raised cap with a prompt that keeps generating (see + // SPEED_PROBE_PROMPT), so the stream lasts long enough for TTFT/TPS to describe + // decoding rather than one round trip; the plain connectivity test keeps the + // single-digit-token budget. + maxTokens: req.speed ? 64 : 16, requestTimeoutMs: 20_000, }); - const gen = llm.streamGenerate({ newMessages: [userText("ping")] }); + const gen = llm.streamGenerate({ + newMessages: [userText(req.speed ? SPEED_PROBE_PROMPT : PROBE_PROMPT)], + }); + let sawContent = false; + let firstContentAt: number | null = null; + let outputTokens = 0; for (;;) { const step = await gen.next(); if (step.done) { - const outcome = step.value; - if (outcome.status === "completed") - return { ok: true, latencyMs: Date.now() - startedAt }; - const detail = "message" in outcome && outcome.message ? outcome.message : outcome.status; - return { ok: false, message: String(detail).slice(0, 300) }; + const verdict = probeVerdict(step.value, sawContent); + if (!verdict.ok) return verdict; + const res: ModelTestResponse = { ok: true, latencyMs: Date.now() - startedAt }; + if (firstContentAt !== null) { + res.ttftMs = firstContentAt - startedAt; + // Output rate over the streaming window (first content -> stream end), dropped + // when the sample is too small to mean anything (see probeTps): usage is only + // reported on completed streams, so thinking-only malformed endings carry TTFT + // but no rate, and so does a model that answers in a couple of tokens. + const tps = probeTps(outputTokens, Date.now() - firstContentAt); + if (tps !== undefined) res.tps = tps; + } + return res; + } + if (isProbeContent(step.value)) { + sawContent = true; + if (firstContentAt === null) firstContentAt = Date.now(); + } + const p = step.value.payload as { type?: string; request?: { output?: number } }; + if (p.type === "token_usage" && typeof p.request?.output === "number") { + outputTokens = p.request.output; } } } catch (err) { @@ -496,3 +548,56 @@ export class ProjectConfigService { return this.getModels(projectId); } } + +/** + * Whether a streamed message carries genuine model content (thinking or text, partial delta + * or complete backfill) — the probe's "the endpoint really answered" signal. Tool calls and + * event messages don't count (the probe declares no tools). + */ +export function isProbeContent(msg: OmniMessage): boolean { + const p = msg.payload as { type?: string; thinking?: string; text?: string }; + if (p.type === "partial_thinking" || p.type === "thinking") return Boolean(p.thinking); + if (p.type === "partial_text" || p.type === "text") return Boolean(p.text); + return false; +} + +/** + * Probe verdict from the terminal LLM outcome. `completed` always passes. A `malformed` + * ending after genuine streamed content also passes: the typical case is a reasoning-heavy + * model that ignores the disabled thinking level and burns the probe's tiny max_tokens + * entirely on thinking (finish_reason=length -> AgentHub's EmptyResponseError) — the + * endpoint, credential, and model id all demonstrably work, which is what a connectivity + * test measures. Everything else (auth/parameter failures, timeouts, malformed with nothing + * received) fails with the outcome's message. + */ +export function probeVerdict( + outcome: LLMOutcome, + sawContent: boolean, +): { ok: true } | { ok: false; message: string } { + if (outcome.status === "completed") return { ok: true }; + if (outcome.status === "malformed" && sawContent) return { ok: true }; + const detail = "message" in outcome && outcome.message ? outcome.message : outcome.status; + return { ok: false, message: String(detail).slice(0, 300) }; +} + +/** + * Sample floors below which the streaming window says nothing about decoding rate. The + * speed probe's 64-token cap clears both by a wide margin, so hitting a floor means the + * model didn't really stream (a one-word answer, or usage that never arrived). + */ +const PROBE_TPS_MIN_TOKENS = 16; +const PROBE_TPS_MIN_WINDOW_MS = 100; + +/** + * Output rate (tokens/s) over the probe's streaming window (first content -> stream end), + * rounded to 1dp; undefined when the sample is too small to be meaningful. A stream's + * closing usage chunk costs a round trip on its own, so a two-token answer measures network + * jitter and nothing else — 2 tokens in 30ms reads as 66.7 tok/s and the same model 30ms + * later reads as 33.3, which the card badges would paint green vs yellow. Callers report + * TTFT alone rather than a fabricated rate; a malformed ending carries no usage at all and + * lands here as 0 tokens. + */ +export function probeTps(outputTokens: number, windowMs: number): number | undefined { + if (outputTokens < PROBE_TPS_MIN_TOKENS || windowMs <= PROBE_TPS_MIN_WINDOW_MS) return undefined; + return Math.round((outputTokens / (windowMs / 1000)) * 10) / 10; +} diff --git a/packages/server/src/services/session-service.ts b/packages/server/src/services/session-service.ts index c8758d1..36deabf 100644 --- a/packages/server/src/services/session-service.ts +++ b/packages/server/src/services/session-service.ts @@ -7,9 +7,9 @@ * (provider, model_id) / workspace, which is backfilled into a DB row * (approval_mode defaults, createdAt is taken from the timestamp embedded in * session_id). - * Create: via core's `agent.createSession` (model reference as a provider + modelId - * pair; defaults to the Project's default reference, 400 if there is none; omitting - * provider goes through resolveModelRef for unique resolution); the new Session is + * Create: via core's `agent.createSession` (the model reference is always a complete + * (provider, modelId) pair — both or neither; omitting both falls back to the + * Project's default reference, 400 if there is none); the new Session is * added to session-manager's active table (state idle). */ import path from "node:path"; @@ -22,6 +22,7 @@ import { } from "@prismshadow/penguin-core"; import type { ApprovalMode, SessionInfo } from "../api/types.js"; import { HttpError, isMissingCredential, modelCredentialMissing } from "../http/errors.js"; +import { badRequest } from "../http/validate.js"; import type { SessionRow, SessionsRepo } from "../db/repos/sessions.js"; import type { SessionManager } from "../runtime/session-manager.js"; import type { ProjectConfigService } from "./project-config-service.js"; @@ -141,28 +142,38 @@ export class SessionService { } /** - * Create a Session: model reference `(provider, modelId)` as a pair; defaults to - * the Project's default reference (400 prompting to configure a model first if - * there is none); when provider is omitted, core's resolveModelRef performs - * unique resolution (400 on zero matches / ambiguity). `workspace` is already + * Create a Session: the model reference is a complete `(provider, modelId)` pair. + * Half a reference is a client error, never something to resolve — the missing half + * is never guessed, since a guessed provider would send an entry's credential to a + * vendor nobody named. Omitting both falls back to the Project's default reference + * (400 prompting to configure a model first if there is none). `workspace` is already * validated by the route guard. The new Session is added to the active table * (idle). */ async createSession(args: { projectId: string; agentId: string; - /** Upstream id of the session's model (paired with provider); defaults to the Project's default reference. */ + /** Upstream id of the session's model; always paired with provider. Omit both for the Project's default reference. */ modelId?: string; - /** The provider group for `modelId`; if omitted, resolveModelRef performs unique resolution. */ + /** The provider group for `modelId`; always paired with modelId, never inferred. */ provider?: string; workspace?: string; approvalMode?: ApprovalMode; /** Session source marker (schedule when triggered by a scheduled task; defaults to user-created). */ source?: "schedule"; }): Promise { - let modelId = args.modelId; - let provider = args.provider; - if (modelId === undefined) { + if ((args.modelId === undefined) !== (args.provider === undefined)) { + throw badRequest( + "modelId and provider must be given together as a (provider, modelId) pair: specify both, or neither to use the Project's default model.", + ); + } + let modelId: string; + let provider: string; + if (args.modelId !== undefined && args.provider !== undefined) { + modelId = args.modelId; + provider = args.provider; + } else { + // The guard above leaves only "both omitted" here: fall back to the Project default. const def = await this.deps.projectConfig.getDefaultModelRef(args.projectId); if (def === undefined) { throw new HttpError( @@ -183,12 +194,12 @@ export class SessionService { try { session = await agent.createSession({ modelId, - ...(provider !== undefined ? { provider } : {}), + provider, ...(args.workspace !== undefined ? { workspaceDir: args.workspace } : {}), }); } catch (err) { // A missing credential is its own category (the frontend shows localized text - // by code); other core errors (zero matches / ambiguous reference, Workspace + // by code); other core errors (the pair naming no configured entry, Workspace // not existing, etc.) are collapsed to 400 — the guard already blocks most cases. if (isMissingCredential(err)) throw modelCredentialMissing(modelId); throw new HttpError( diff --git a/packages/server/test/model-probe.test.ts b/packages/server/test/model-probe.test.ts new file mode 100644 index 0000000..64936c2 --- /dev/null +++ b/packages/server/test/model-probe.test.ts @@ -0,0 +1,99 @@ +/** + * Model connectivity-probe tests (testModel's pure parts): streamed-content detection, the + * outcome verdict — in particular the reasoning-heavy-model case where the probe's tiny + * max_tokens is burned entirely on thinking (finish_reason=length -> AgentHub + * EmptyResponseError -> malformed outcome) yet the endpoint demonstrably works — and the + * output-rate guard that keeps a too-small sample from being reported as throughput. + */ +import { describe, expect, it } from "vitest"; +import { + assistantText, + partialText, + partialThinking, + thinkingMessage, + tokenUsage, + toolCall, + emptyTokenCounts, +} from "@prismshadow/penguin-core"; +import { isProbeContent, probeTps, probeVerdict } from "../src/services/project-config-service.js"; + +describe("isProbeContent", () => { + it("counts thinking and text content, partial or complete", () => { + expect(isProbeContent(partialThinking("delta", "let me think"))).toBe(true); + expect(isProbeContent(partialText("delta", "pong"))).toBe(true); + expect(isProbeContent(thinkingMessage("burned the whole budget", "malformed"))).toBe(true); + expect(isProbeContent(assistantText("pong"))).toBe(true); + }); + + it("ignores empty segments, events, and tool calls", () => { + expect(isProbeContent(partialThinking("start"))).toBe(false); + expect(isProbeContent(partialThinking("stop", "", "malformed"))).toBe(false); + expect(isProbeContent(thinkingMessage("", "malformed", { id: "rs_1" }))).toBe(false); + expect(isProbeContent(tokenUsage(emptyTokenCounts(), emptyTokenCounts()))).toBe(false); + expect(isProbeContent(toolCall({ name: "t", arguments: "{}", toolCallId: "tc1" }))).toBe(false); + }); +}); + +describe("probeVerdict", () => { + it("passes a completed probe", () => { + expect(probeVerdict({ status: "completed" }, false)).toEqual({ ok: true }); + }); + + it("passes a thinking-only malformed ending when content was streamed (reasoning model hit the tiny cap)", () => { + const outcome = { + status: "malformed" as const, + message: 'OpenaiClient returned no content other than thinking (finish_reason="length").', + }; + expect(probeVerdict(outcome, true)).toEqual({ ok: true }); + }); + + it("fails a malformed ending with nothing received (broken response, not a working endpoint)", () => { + const verdict = probeVerdict({ status: "malformed", message: "unexpected EOF" }, false); + expect(verdict).toEqual({ ok: false, message: "unexpected EOF" }); + }); + + it("fails timeouts and errors even when content was streamed (flaky is not ok)", () => { + expect(probeVerdict({ status: "timeout" }, true)).toEqual({ ok: false, message: "timeout" }); + expect(probeVerdict({ status: "failed", message: "401 unauthorized" }, true)).toEqual({ + ok: false, + message: "401 unauthorized", + }); + expect(probeVerdict({ status: "aborted" }, true)).toEqual({ ok: false, message: "aborted" }); + }); + + it("truncates long failure messages to 300 characters", () => { + const verdict = probeVerdict({ status: "failed", message: "x".repeat(500) }, false); + expect(verdict.ok).toBe(false); + if (!verdict.ok) expect(verdict.message).toHaveLength(300); + }); +}); + +describe("probeTps", () => { + it("reports the rate for a sample that spans a real streaming window", () => { + // The speed probe's 64-token cap: a full window is what the badge thresholds grade. + expect(probeTps(64, 1000)).toBe(64); + expect(probeTps(48, 900)).toBe(53.3); + }); + + it("reports at the sample floors", () => { + expect(probeTps(16, 400)).toBe(40); + expect(probeTps(16, 101)).toBe(158.4); + }); + + it("omits the rate for a one-word answer (jitter, not throughput)", () => { + // 2 tokens in 30ms would read as 66.7 tok/s; 30ms of jitter later, as 33.3 — the caller + // still reports ttftMs, just no rate. + expect(probeTps(2, 30)).toBeUndefined(); + expect(probeTps(2, 60)).toBeUndefined(); + expect(probeTps(15, 300)).toBeUndefined(); + }); + + it("omits the rate for a window too short to time, however many tokens arrived", () => { + expect(probeTps(64, 100)).toBeUndefined(); + expect(probeTps(64, 0)).toBeUndefined(); + }); + + it("omits the rate when no usage was reported (malformed thinking-only ending)", () => { + expect(probeTps(0, 5000)).toBeUndefined(); + }); +}); diff --git a/packages/server/test/models.test.ts b/packages/server/test/models.test.ts index d40e7ea..8eca08e 100644 Binary files a/packages/server/test/models.test.ts and b/packages/server/test/models.test.ts differ diff --git a/packages/server/test/schedule-file.test.ts b/packages/server/test/schedule-file.test.ts index bd32fd6..f893755 100644 --- a/packages/server/test/schedule-file.test.ts +++ b/packages/server/test/schedule-file.test.ts @@ -49,6 +49,11 @@ describe("parseScheduleFile", () => { }); }); + it("keeps a complete model reference pair", () => { + const def = defOf(`${BASE}provider = "custom"\nmodel_id = "m1"\n`); + expect(def).toMatchObject({ provider: "custom", modelId: "m1" }); + }); + it("enabled defaults to off; omitting period means one-shot", () => { const def = defOf(`prompt = "p"\nstart_at = "2026-07-16T09:00:00Z"\n`); expect(def.enabled).toBe(false); @@ -67,6 +72,10 @@ describe("parseScheduleFile", () => { [`${BASE}session_id = "s"\nworkspace = "/tmp/w"\n`, "new-Session mode"], [`${BASE}session_id = "s"\nmodel_id = "m1"\n`, "new-Session mode"], [`${BASE}model_id = ""\n`, "model_id"], + // A model reference is always a pair: half of one is invalid in either direction + // (model_id alone is what a file written before this rule looks like). + [`${BASE}model_id = "m1"\n`, "given together"], + [`${BASE}provider = "custom"\n`, "given together"], ]; for (const [raw, hint] of cases) { const r = parseScheduleFile("x", raw); diff --git a/packages/server/test/scheduler.test.ts b/packages/server/test/scheduler.test.ts index c9361a1..d684c7b 100644 --- a/packages/server/test/scheduler.test.ts +++ b/packages/server/test/scheduler.test.ts @@ -47,9 +47,9 @@ describe("scheduler", () => { beforeEach(async () => { root = await makeTempRoot(); - // A schedule's model reference is resolved against the Project config during - // reconciliation, so provide a minimal config here: - // m-bench is globally unique, so provider can be omitted and still resolve uniquely. + // A schedule's model reference is checked against the Project config during + // reconciliation, so provide a minimal config here: the file must name the whole + // (provider, model_id) pair for it to match. await saveProjectConfig(root, P, { default_model: { provider: "custom", model_id: "m-bench" }, models: [{ provider: "custom", model_id: "m-bench" }], @@ -279,6 +279,7 @@ describe("scheduler", () => { `start_at = "${iso(T0 + MIN)}"`, `period = "5m"`, `workspace = "/tmp/ws"`, + `provider = "custom"`, `model_id = "m-bench"`, ]); await scheduler.tickOnce(); @@ -287,31 +288,50 @@ describe("scheduler", () => { nowMs = T0 + 6 * MIN; await scheduler.tickOnce(); expect(created).toHaveLength(2); - // The file only supplies model_id (no provider): passed through as-is, resolved - // into a paired reference downstream via a unique match. + // The file's model reference is passed straight through as a pair. expect(created[0]).toMatchObject({ projectId: P, agentId: A, workspace: "/tmp/ws", + provider: "custom", modelId: "m-bench", }); - expect("provider" in created[0]!).toBe(false); expect(started.map((s) => s.sessionId)).toEqual(["session-new-1", "session-new-2"]); }); - it("new-Session mode: when the file gives provider, it is passed through as a pair", async () => { - await writeFile("paired", [ - `prompt = "paired"`, + it("new-Session mode: omitting the model reference entirely leaves the Project default to the session creator", async () => { + await writeFile("default-model", [ + `prompt = "default"`, `enabled = true`, `start_at = "${iso(T0 + MIN)}"`, - `provider = "custom"`, - `model_id = "m-bench"`, ]); await scheduler.tickOnce(); nowMs = T0 + MIN; await scheduler.tickOnce(); expect(created).toHaveLength(1); - expect(created[0]).toMatchObject({ provider: "custom", modelId: "m-bench" }); + expect("provider" in created[0]!).toBe(false); + expect("modelId" in created[0]!).toBe(false); + }); + + it("a file carrying model_id without provider is invalid: never scheduled, error recorded", async () => { + // A schedule file persisted before the pairing rule can hold model_id alone. It must + // fail exactly like any other invalid file — skipped with an error recorded — rather + // than resolving the provider or throwing out of the tick loop. + await writeFile("legacy", [ + `prompt = "legacy"`, + `enabled = true`, + `start_at = "${iso(T0 + MIN)}"`, + `model_id = "m-bench"`, + ]); + await scheduler.tickOnce(); + nowMs = T0 + 2 * MIN; + await scheduler.tickOnce(); + expect(created).toHaveLength(0); + expect(started).toHaveLength(0); + expect(repo.find(P, A, "legacy")).toBeNull(); + const recorded = errors.filter((e) => e.code === "schedule_invalid_file"); + expect(recorded.length).toBeGreaterThan(0); + expect((recorded[0]?.err as Error).message).toContain("given together"); }); it("invalid files are skipped with an error recorded, without affecting other tasks", async () => { diff --git a/packages/server/test/schedules.test.ts b/packages/server/test/schedules.test.ts index 60ba576..3d9feaa 100644 --- a/packages/server/test/schedules.test.ts +++ b/packages/server/test/schedules.test.ts @@ -146,17 +146,24 @@ describe("schedules api", () => { ).toBe(404); }); - it("validation 400: period below the minimum, invalid instant, pick-one target, invalid task name", async () => { + it("validation 400: period below the minimum, invalid instant, pick-one target, half a model reference, invalid task name", async () => { const ok = { prompt: "p", enabled: false, startAt: FUTURE }; const cases = [ { name: "j1", ...ok, period: "4m" }, { name: "j2", ...ok, startAt: "someday" }, { name: "j3", ...ok, sessionId: "s", workspace: "/w" }, + // A model reference is a pair: neither half alone is accepted (nothing is inferred). + { name: "j4", ...ok, modelId: "deepseek-v4-pro" }, + { name: "j5", ...ok, provider: "deepseek" }, { name: "bad.name", ...ok }, ]; for (const body of cases) { expect((await owner.post(base, body)).status, JSON.stringify(body)).toBe(400); } + // Nothing was persisted for any of them. + const list = (await (await owner.get(base)).json()) as SchedulesResponse; + expect(list.schedules).toEqual([]); + expect(list.invalidFiles).toEqual([]); }); it("hand-edited files: invalid ones land in invalidFiles; expired one-shots are marked missed during reconciliation", async () => { @@ -175,4 +182,22 @@ describe("schedules api", () => { // For a hand-edited file, the creator falls back to the Project owner. expect(stale?.creatorUserId).toBe("owner_a"); }); + + it("a file holding model_id without provider lands in invalidFiles instead of being scheduled", async () => { + // What a schedule file persisted before the pairing rule looks like: the provider is + // never filled in for it, the file is simply reported invalid and skipped. + const dir = scheduleDir(t.root, projectId, "default_agent"); + await fs.mkdir(dir, { recursive: true }); + await fs.writeFile( + path.join(dir, "legacy.toml"), + `prompt = "p"\nenabled = true\nstart_at = "${FUTURE}"\nmodel_id = "deepseek-v4-pro"\n`, + "utf8", + ); + const list = (await (await owner.get(base)).json()) as SchedulesResponse; + expect(list.schedules).toEqual([]); + expect(list.invalidFiles.map((f) => f.name)).toEqual(["legacy"]); + expect(list.invalidFiles[0]?.error).toContain("given together"); + // Reading it by name reports the same error rather than 404 or a 500. + expect((await owner.get(`${base}/legacy`)).status).toBe(400); + }); }); diff --git a/packages/server/test/session-index.test.ts b/packages/server/test/session-index.test.ts index 278ad71..beec5a7 100644 --- a/packages/server/test/session-index.test.ts +++ b/packages/server/test/session-index.test.ts @@ -103,6 +103,21 @@ describe("session-index", () => { expect(list.sessions.map((s) => s.sessionId)).toContain(session.sessionId); }); + it("half a model reference is 400: the missing half is never inferred", async () => { + await configureModels(); + // Only modelId: even though it names the one configured model, the provider is never + // filled in for the caller — a reference is submitted as a pair or not at all. + const onlyModel = await api.post(base(), { modelId: "claude-sonnet-4-6" }); + expect(onlyModel.status).toBe(400); + const onlyProvider = await api.post(base(), { provider: "anthropic" }); + expect(onlyProvider.status).toBe(400); + // The complete pair works, and so does omitting both (Project default). + expect( + (await api.post(base(), { provider: "anthropic", modelId: "claude-sonnet-4-6" })).status, + ).toBe(201); + expect((await api.post(base(), {})).status).toBe(201); + }); + it("an explicit Workspace only needs to exist; it may live outside the Project directory", async () => { await configureModels(); const inside = path.join(t.root, projectId, "my-workdir"); diff --git a/packages/server/test/session-manager.test.ts b/packages/server/test/session-manager.test.ts index a31fdf9..a848042 100644 --- a/packages/server/test/session-manager.test.ts +++ b/packages/server/test/session-manager.test.ts @@ -2,8 +2,9 @@ * Unit tests for the Session runtime (a fake Session / Loader is injected; no * real LLM requests are made): driving and state transitions, 409 mutual * exclusion, the four approval modes and taking effect immediately on change, - * abort collapsing to deny, self-healing id swaps, and LLM / tool errors in the - * message stream being persisted (core doesn't throw, so try/catch can't catch them). + * abort collapsing to deny, self-healing id swaps, vault invalidation re-resuming + * stale runtimes, and LLM / tool errors in the message stream being persisted + * (core doesn't throw, so try/catch can't catch them). */ import { afterEach, beforeEach, describe, expect, it } from "vitest"; import type { DatabaseSync } from "node:sqlite"; @@ -475,6 +476,62 @@ describe("session-manager", () => { expect(loads).toBe(1); }); + it("invalidateAgentRuntimes: an idle entry is discarded and re-resumed on next access; other Agents unaffected", async () => { + let loads = 0; + const loader: SessionLoader = { + load: async () => { + loads++; + return approvalFakeSession("session-1"); + }, + }; + const manager = makeManager(loader); + sessions.updateApprovalMode("session-1", "allow-all"); + // Adopted after creation: records the current generation, so Tasks reuse it without loading. + manager.adopt(ROW, approvalFakeSession("session-1")); + + // Another Agent's vault update leaves this entry alone. + manager.invalidateAgentRuntimes("p1", "other_agent"); + await manager.startTask("session-1", [userText("a")]); + await waitFor(() => manager.statusOf("session-1") === "idle"); + expect(loads).toBe(0); + + // This Agent's vault update: the next Task rebuilds the runtime via the loader. + manager.invalidateAgentRuntimes("p1", "a1"); + await manager.startTask("session-1", [userText("b")]); + await waitFor(() => manager.statusOf("session-1") === "idle"); + expect(loads).toBe(1); + + // Rebuilt once only: the fresh entry is current again. + await manager.startTask("session-1", [userText("c")]); + await waitFor(() => manager.statusOf("session-1") === "idle"); + expect(loads).toBe(1); + }); + + it("invalidateAgentRuntimes mid-run: the in-flight Task keeps its runtime; the first Task after it finishes re-resumes", async () => { + let loads = 0; + const loader: SessionLoader = { + load: async () => { + loads++; + return approvalFakeSession("session-1"); + }, + }; + const manager = makeManager(loader); + await manager.startTask("session-1", [userText("go")]); // built here: load #1 + await waitFor(() => manager.pendingApprovalCount("session-1") === 1); + + manager.invalidateAgentRuntimes("p1", "a1"); // vault updated while the Task waits on approval + // The pending approval still targets the live entry: the run completes on the old runtime. + expect(manager.decideApproval("session-1", "tc-1", "allow")).toBe(true); + await waitFor(() => manager.statusOf("session-1") === "idle"); + expect(loads).toBe(1); + + // First Task after it finished: the stale entry is discarded and re-resumed with current values. + sessions.updateApprovalMode("session-1", "allow-all"); + await manager.startTask("session-1", [userText("next")]); + await waitFor(() => manager.statusOf("session-1") === "idle"); + expect(loads).toBe(2); + }); + it("sweepIdle: entries that are running / have pending approvals are not evicted", async () => { const manager = makeManager(loaderOf(approvalFakeSession("session-1"))); await manager.startTask("session-1", [userText("go")]); @@ -547,4 +604,121 @@ describe("session-manager", () => { await new Promise((r) => setTimeout(r, 10)); expect(notified.length).toBe(1); }); + + it("fires title generation early once 1000 chars of body text have streamed (mid-run, before the Task finishes)", async () => { + const notified: { ctx: UsageContext; req: TitleRequest }[] = []; + // Gate the run after the big body text so the early trigger provably happens mid-run. + let release!: () => void; + const gate = new Promise((r) => (release = r)); + const longSession: RuntimeSession = { + sessionId: "session-1", + toolPermission: () => "rw", + generateTitle: async () => ({ title: null, usage: null }), + compactability: () => "ok" as const, + async *run() { + yield assistantText("x".repeat(600)); + // Sub-session text doesn't count toward the main-session body threshold + // (asserted on its own in the next test). + yield withOrigin(assistantText("z".repeat(2000)), "session-sub"); + yield assistantText("y".repeat(600)); // crosses the 1000-char threshold + await gate; + yield assistantText("tail"); + }, + async *compact() {}, + }; + const manager = new SessionManager({ + sessions, + channels, + loader: loaderOf(longSession), + recorder: { record: async () => {} }, + titles: { + maybeGenerate: (ctx, _session, req) => notified.push({ ctx, req }), + }, + log: () => {}, + }); + + await manager.startTask("session-1", [userText("long question")]); + // The early trigger fires while the run is still in flight (gated before completion). + await waitFor(() => notified.length === 1); + expect(manager.statusOf("session-1")).toBe("running"); + expect(notified[0]!.req.fallbackText).toBe("long question"); + expect(notified[0]!.req.material).toBeUndefined(); + + // Completion still notifies as the short-answer fallback path (the generator itself dedups). + release(); + await waitFor(() => manager.statusOf("session-1") === "idle"); + await waitFor(() => notified.length === 2); + }); + + it("a subagent's output never fires the early title: only main-session body text counts toward the 1000 chars", async () => { + const notified: { ctx: UsageContext; req: TitleRequest }[] = []; + const driven: OmniMessage[] = []; + // Gate the run right after the sub-session's long output: while the parent is parked + // here the completion trigger hasn't run yet, so an empty `notified` proves the child's + // text alone never crossed the threshold (without the gate an early fire would be + // indistinguishable from the completion one). + let release!: () => void; + const gate = new Promise((r) => (release = r)); + const delegating: RuntimeSession = { + sessionId: "session-1", + toolPermission: () => "rw", + generateTitle: async () => ({ title: null, usage: null }), + compactability: () => "ok" as const, + async *run() { + yield toolCall({ + name: "run_subagent", + arguments: JSON.stringify({ prompt: "Research the background of this question" }), + toolCallId: "sub-1", + }); + const hop = "child-1"; + yield withOrigin( + sessionMeta({ + session_id: "child-1", + model_id: "m-child", + provider: "custom", + model_context_window: 1000, + system_prompt: "sys", + tools: [], + thinking_level: "default", + agent_state: "/root/p1/child_agent/agent_state", + workspace: "/tmp/w-child", + }), + hop, + ); + // Far past the threshold, but it belongs to the sub-session's own conversation. + yield withOrigin(assistantText("z".repeat(3000)), hop); + await gate; + yield assistantText("short answer"); // the parent's whole body, well under 1000 chars + }, + async *compact() {}, + }; + const manager = new SessionManager({ + sessions, + channels, + loader: loaderOf(delegating), + recorder: { + record: async (_ctx, msg) => { + driven.push(msg); + }, + }, + titles: { + maybeGenerate: (ctx, _session, req) => notified.push({ ctx, req }), + }, + log: () => {}, + }); + + await manager.startTask("session-1", [userText("delegate this")]); + // Recording happens after the early-title check, so once the sub-session's 3000 chars + // have been driven through the parent's counter has already seen everything it will + // ever see from the child — and it must still be at zero. + await waitFor(() => driven.length === 3); + expect(manager.statusOf("session-1")).toBe("running"); + expect(notified.length).toBe(0); + + // Only the completion path notifies: once for the parent, once for the sub-session's own title. + release(); + await waitFor(() => manager.statusOf("session-1") === "idle"); + await waitFor(() => notified.length === 2); + expect(notified.map((n) => n.ctx.sessionId)).toEqual(["session-1", "child-1"]); + }); }); diff --git a/packages/server/test/skills.test.ts b/packages/server/test/skills.test.ts index 6f33ede..5cecf23 100644 --- a/packages/server/test/skills.test.ts +++ b/packages/server/test/skills.test.ts @@ -57,11 +57,10 @@ describe("skills api", () => { expect(res.status).toBe(200); const body = (await res.json()) as SkillLibraryResponse; expect(body.groups.map((g) => g.id)).toEqual([ - "agent-development", - "data-analysis", - "penguin-development", - "web-development", - "software-engineering", + "office-productivity", + "software-development", + "ai-app-development", + "agent-tuning", ]); for (const group of body.groups) { expect(group.title.length).toBeGreaterThan(0); @@ -71,20 +70,22 @@ describe("skills api", () => { expect("description" in group).toBe(false); } // Members within a group follow the SKILL_GROUPS list order (as ungrouped by loadSkillGroups). - expect(body.groups[0]!.skills.map((s) => s.name)).toEqual([ - "agent-creation", - "benchmark-design", - "agent-evaluation", - "agent-optimization", + expect(body.groups[0]!.skills.map((s) => s.name)).toEqual(["data-analysis", "firecrawl"]); + expect(body.groups[1]!.skills.map((s) => s.name)).toEqual([ + "web-design", + "software-engineering", ]); - expect(body.groups[1]!.skills.map((s) => s.name)).toEqual(["data-analysis"]); expect(body.groups[2]!.skills.map((s) => s.name)).toEqual([ "penguin-sdk", "penguin-cli", "agenthub-models", ]); - expect(body.groups[3]!.skills.map((s) => s.name)).toEqual(["web-design"]); - expect(body.groups[4]!.skills.map((s) => s.name)).toEqual(["software-engineering"]); + expect(body.groups[3]!.skills.map((s) => s.name)).toEqual([ + "agent-creation", + "benchmark-design", + "agent-evaluation", + "agent-optimization", + ]); const skills = body.groups.flatMap((g) => g.skills); for (const skill of skills) { expect(skill.name.length).toBeGreaterThan(0); diff --git a/packages/server/test/title-generator.test.ts b/packages/server/test/title-generator.test.ts index 5844662..b5dc213 100644 --- a/packages/server/test/title-generator.test.ts +++ b/packages/server/test/title-generator.test.ts @@ -164,6 +164,17 @@ describe("title-generator", () => { await waitFor(() => recorded.length > 0); }); + it("fallback strips a leading block so the marker never becomes the title", async () => { + const calls = { count: 0, args: [] as unknown[] }; + const gen = makeGenerator(); + gen.maybeGenerate(CTX, fakeSession({ title: null, usage: null }, calls), { + fallbackText: "\nskills: web-design\n\n做一个落地页", + }); + await waitFor(() => sessions.findById(ROW.sessionId)?.title !== null); + // Not "" — the marker block is removed before the first line is taken. + expect(sessions.findById(ROW.sessionId)?.title).toBe("做一个落地页"); + }); + it("LLM returns null and the fallback material is blank → the title stays NULL (retryable next time)", async () => { const calls = { count: 0, args: [] as unknown[] }; const gen = makeGenerator(); diff --git a/packages/server/test/vault.test.ts b/packages/server/test/vault.test.ts index 3180854..c04ad22 100644 --- a/packages/server/test/vault.test.ts +++ b/packages/server/test/vault.test.ts @@ -3,14 +3,29 @@ * agent_state/.vault.toml): GET masks values (plaintext * is never sent), PUT is owner-only, whole-table replace semantics (omitting * value keeps the original, an absent key is deleted, a new key must supply a - * value), 400 on key/shape validation, 404 for a nonexistent Agent, and vaults - * of different Agents are independent of each other. + * value), 400 on key/shape validation, 404 for a nonexistent Agent, vaults of + * different Agents are independent of each other, and PUT invalidates the Agent's + * cached Session runtimes (the next Task re-resumes and sees the new values). */ import { afterEach, beforeEach, describe, expect, it } from "vitest"; +import { userText } from "@prismshadow/penguin-core"; import type { ProjectCreateResponse, VaultResponse } from "../src/api/types.js"; -import { apiClient, createTestApp, provisionUser } from "./helpers.js"; +import type { RuntimeSession } from "../src/runtime/session-manager.js"; +import { apiClient, createTestApp, provisionUser, waitFor } from "./helpers.js"; import type { TestApp } from "./helpers.js"; +/** Minimal fake runtime Session: one assistant reply, no approvals (keeps the loader-count test free of LLM calls). */ +function fakeRuntimeSession(sessionId: string): RuntimeSession { + return { + sessionId, + toolPermission: () => "rw", + generateTitle: async () => ({ title: null, usage: null }), + compactability: () => "ok" as const, + async *run() {}, + async *compact() {}, + }; +} + describe("vault api", () => { let t: TestApp; let owner: ReturnType; @@ -18,9 +33,19 @@ describe("vault api", () => { let outsider: ReturnType; let projectId: string; let vaultPath: string; + /** Loader call count: how many times the manager (re)built a runtime from the index. */ + let loads: number; beforeEach(async () => { - t = await createTestApp(); + loads = 0; + t = await createTestApp({ + loader: { + load: async (row) => { + loads++; + return fakeRuntimeSession(row.sessionId); + }, + }, + }); const a = await provisionUser(t.app, "owner_a"); const b = await provisionUser(t.app, "member_b"); const c = await provisionUser(t.app, "outsider_c"); @@ -112,6 +137,42 @@ describe("vault api", () => { expect(modelsBody.models[0]!.credential).toBeTruthy(); }); + it("PUT invalidates the Agent's cached Session runtimes: the next Task re-resumes and picks up the new vault", async () => { + t.deps.sessionsRepo.insert({ + sessionId: "vault-sess-1", + projectId, + agentId: "default_agent", + modelId: "m1", + provider: "custom", + workspace: t.root, + approvalMode: "allow-all", + title: null, + createdAt: new Date().toISOString(), + }); + const idle = () => t.deps.manager.statusOf("vault-sess-1") === "idle"; + + // First Task builds the runtime (load #1); the second reuses the active-table entry. + await t.deps.manager.startTask("vault-sess-1", [userText("a")]); + await waitFor(idle); + await t.deps.manager.startTask("vault-sess-1", [userText("b")]); + await waitFor(idle); + expect(loads).toBe(1); + + // Vault update via the API: the cached runtime is stale, so the next Task re-resumes + // (the loader re-reads .vault.toml — the new value reaches the next Task's environment). + const put = await owner.put(vaultPath, { entries: [{ key: "NEW_KEY", value: "v-secret-1" }] }); + expect(put.status).toBe(200); + await t.deps.manager.startTask("vault-sess-1", [userText("c")]); + await waitFor(idle); + expect(loads).toBe(2); + + // Reads don't invalidate: the rebuilt runtime is reused. + expect((await owner.get(vaultPath)).status).toBe(200); + await t.deps.manager.startTask("vault-sess-1", [userText("d")]); + await waitFor(idle); + expect(loads).toBe(2); + }); + it("Agent-level isolation: different Agents' vaults are independent; a nonexistent Agent 404", async () => { await owner.put(vaultPath, { entries: [{ key: "ONLY_DEFAULT", value: "v-default-1" }] }); diff --git a/packages/server/test/workspace-files.test.ts b/packages/server/test/workspace-files.test.ts index 7178803..d99e995 100644 --- a/packages/server/test/workspace-files.test.ts +++ b/packages/server/test/workspace-files.test.ts @@ -122,6 +122,7 @@ describe("files/stat route (batch existence check)", () => { let owner: ReturnType; let outsider: ReturnType; let sessionId: string; + let workspace: string; beforeEach(async () => { t = await createTestApp(); @@ -141,6 +142,7 @@ describe("files/stat route (batch existence check)", () => { await owner.post(`/api/projects/${projectId}/agents/default_agent/sessions`, {}) ).json()) as SessionCreateResponse; sessionId = sess.session.sessionId; + workspace = sess.session.workspace; await fs.mkdir(path.join(sess.session.workspace, "sub")); await fs.writeFile(path.join(sess.session.workspace, "a.txt"), "A"); await fs.writeFile(path.join(sess.session.workspace, "sub", "b.md"), "B"); @@ -149,6 +151,35 @@ describe("files/stat route (batch existence check)", () => { await t.cleanup(); }); + it("files/content on html: inline stays text/plain; preview=1 keeps text/html under a CSP sandbox; download keeps the real type with no CSP", async () => { + await fs.writeFile(path.join(workspace, "page.html"), ""); + const url = `/api/sessions/${sessionId}/files/content?path=page.html`; + + const inline = await owner.get(url); + expect(inline.status).toBe(200); + expect(inline.headers.get("content-type")).toContain("text/plain"); + expect(inline.headers.get("content-security-policy")).toBeNull(); + + const preview = await owner.get(`${url}&preview=1`); + expect(preview.status).toBe(200); + expect(preview.headers.get("content-type")).toContain("text/html"); + const csp = preview.headers.get("content-security-policy") ?? ""; + expect(csp).toContain("sandbox"); + expect(csp).not.toContain("allow-same-origin"); + expect(preview.headers.get("content-disposition")).toContain("inline"); + + // download wins over preview: attachment + real type, no CSP needed. + const download = await owner.get(`${url}&download=1&preview=1`); + expect(download.headers.get("content-type")).toContain("text/html"); + expect(download.headers.get("content-disposition")).toContain("attachment"); + expect(download.headers.get("content-security-policy")).toBeNull(); + + // Non-scriptable files are unaffected by preview. + const txt = await owner.get(`/api/sessions/${sessionId}/files/content?path=a.txt&preview=1`); + expect(txt.headers.get("content-type")).toContain("text/plain"); + expect(txt.headers.get("content-security-policy")).toBeNull(); + }); + it("existing files return in order, deduplicated; missing / directory / out-of-bounds all count as nonexistent, always 200", async () => { const res = await owner.post(`/api/sessions/${sessionId}/files/stat`, { paths: ["sub/b.md", "a.txt", "sub/b.md", "nope.txt", "sub", "../escape.txt", "/etc/passwd"], diff --git a/packages/skills/README.md b/packages/skills/README.md index d422fc0..aae9277 100644 --- a/packages/skills/README.md +++ b/packages/skills/README.md @@ -4,17 +4,16 @@ The PenguinHarness built-in skill library. A Skill is a directory with a `SKILL. Skills follow the "index first, body on demand" design: only their metadata is injected into an Agent's system prompt; the Agent reads the full `SKILL.md` via shell when it actually needs it. -Included skills: +Included skills, in the order of the `SKILL_GROUPS` manifest in `src/index.ts` (a skill directory missing from the manifest is still loaded, and lands in an "Other" group): | Group | Skills | | --- | --- | -| Agent Development | `agent-creation`, `benchmark-design`, `agent-evaluation`, `agent-optimization` | -| Data Analysis | `data-analysis` | -| Penguin Development | `penguin-sdk`, `penguin-cli`, `agenthub-models` | -| Web Development | `web-design` | -| Software Engineering | `software-engineering` | +| Office Productivity | `data-analysis`, `firecrawl` | +| Software Development | `web-design`, `software-engineering` | +| AI App Development | `penguin-sdk`, `penguin-cli`, `agenthub-models` | +| Agent Tuning | `agent-creation`, `benchmark-design`, `agent-evaluation`, `agent-optimization` | -The first group powers the self-improvement loop: design a Benchmark, evaluate the Target Agent, optimize it to version N+1 with a snapshot before every round. +Agent Tuning powers the self-improvement loop: create the Target Agent, design a Benchmark, evaluate it, optimize it to version N+1 with a snapshot before every round. ## Documentation diff --git a/packages/skills/skills/agent-creation/SKILL.md b/packages/skills/skills/agent-creation/SKILL.md index 7b0eb68..74f408c 100644 --- a/packages/skills/skills/agent-creation/SKILL.md +++ b/packages/skills/skills/agent-creation/SKILL.md @@ -3,8 +3,8 @@ name: agent-creation description: Turn a user requirement into a concrete agent — write the target agent's AGENTS.md and install the skills it needs. short_description: Turn a requirement into a working agent. short_description_zh: 把需求变成可用的 Agent。 -version: 1 -updated: 2026-07-17T00:00:00Z +version: 2 +updated: 2026-07-20T13:00:00Z --- # Agent Creation @@ -13,7 +13,7 @@ This skill turns a user requirement into a working agent configuration — plain ## Before you start -If the user's message only invokes this skill (e.g. "use agent-creation skill") without a concrete requirement, ask the user what agent they want and what it should do. Do not start until the requirement is clear. +If the user's message only invokes this skill (e.g. "use agent-creation skill") without a concrete requirement, ask the user what agent they want and what it should do. But when the requirement is already concrete — even a single sentence like "an expert that answers questions about X" — do **not** ask follow-up questions: derive the role and rules from that sentence, apply the defaults below, and list your assumptions in the final reply. ## Locate the target agent @@ -34,7 +34,7 @@ An agent directory contains `agent_state/` (`system_config.yaml`, `AGENTS.md`, ` - Role — what the agent is for, in one or two sentences. - Domain guidance — the concrete rules, steps and constraints derived from the user requirement. -Be concise: AGENTS.md is prompt context, not documentation. +Be concise: AGENTS.md is prompt context, not documentation. For a domain expert that answers from a knowledge base, a good AGENTS.md is a few lines: the role sentence, "answer strictly from the provided context blocks", citation rules ("cite blocks inline as [1][2]"), a refusal rule for questions the context cannot answer, and "answer in the language of the question". ## Install skills @@ -44,8 +44,8 @@ A skill is a directory `agent_state/skills//` containing a `SKILL.md --- name: description: -version: 1 -updated: +version: +updated: --- @@ -57,6 +57,12 @@ Installing is all it takes: the frontmatter metadata of every `SKILL.md` under ` Write skills yourself, or fetch existing ones from the internet with shell commands (`curl`, `git clone`) and place them under `skills/`. Anything fetched from the internet must be read in full and reviewed before installing — a skill becomes durable instructions the target agent will follow in every future session; never install one you have not read, and tell the user what it does. +Library skills can be copied from any agent that already has them (e.g. `default_agent`, which ships the whole library) — copy the entire `skills//` directory. Common bundles, so you don't under-equip the target: + +- **App builder** (builds apps or web frontends): `penguin-sdk`, `web-design`, `agenthub-models`. +- **Knowledge expert** (answers questions over a document set): usually **no** harness agent is needed — build a RAG app with the penguin-sdk skill instead, and configure the app's embedded agent (below). +- **Evaluation loop**: `benchmark-design`, `agent-evaluation`, `agent-optimization`. + ## Set name and description In the target's `agent_state/system_config.yaml`, set the top-level `name:` and `description:` fields so the agent is recognizable in lists. Edit only these two fields. @@ -71,3 +77,7 @@ cp "$PROJECT_DIR/agents/default_agent/agent_state/system_config.yaml" "$TARGET/a ``` A new agent starts with no skills — install only what it needs. Then write its AGENTS.md, name and description as above. + +## The embedded agent of an SDK app + +An app built with the penguin-sdk skill carries its own agent inside the project (`createAgent({ root })` initializes `/penguin_data/default_project/agents/default_agent/` on first run). That directory has exactly the layout described here, and everything in this skill applies to it: write the app's persona into its `agent_state/AGENTS.md` (the penguin-sdk recipe keeps the source of truth in the project's `persona.md` and copies it in during ingest), and set `name`/`description` in its `system_config.yaml` so the app is recognizable. This is how "the app becomes an expert on X": the persona lives in the embedded agent's AGENTS.md, not in application code. diff --git a/packages/skills/skills/agenthub-models/SKILL.md b/packages/skills/skills/agenthub-models/SKILL.md index 7ba0678..6890b6b 100644 --- a/packages/skills/skills/agenthub-models/SKILL.md +++ b/packages/skills/skills/agenthub-models/SKILL.md @@ -3,8 +3,8 @@ name: agenthub-models description: Call model APIs through @prismshadow/agenthub — streaming text generation, image generation, speech synthesis and embeddings with one client. short_description: Call model APIs with one AgentHub client. short_description_zh: 用一个 AgentHub 客户端调用模型 API。 -version: 1 -updated: 2026-07-17T00:00:00Z +version: 7 +updated: 2026-07-21T00:00:00Z --- # AgentHub Model APIs @@ -29,6 +29,20 @@ const client = new AutoLLMClient({ model: "", apiKey: "", baseUrl If the user's message only invokes this skill (e.g. "use agenthub-models skill") without a concrete task, ask the user what they want to build. Do not write code until the requirement is clear. +**Important prerequisite — set the key up first, then develop.** When the script is an AI app you are building for the user, have them add the model API key in **this agent's key vault** (gear icon on its card, Agents page → settings → key vault tab) *before* you start, so the credential is in your shell environment. If the app stores its own model config, keep its Penguin data root **inside the CWD workspace** (`--root ./penguin_data`), never `~/.penguin`. Model ids can come from the penguin CLI catalog and the id table below. + +Check for a usable API key before writing code — the client needs one for whichever provider you target: + +```bash +env | grep -oE "(DEEPSEEK|OPENAI|ANTHROPIC|GEMINI)_API_KEY" || echo none +``` + +Vault keys also appear in your Vault Keys section. **Only two sources count as a usable key**: a vault-injected environment variable (the check above), or — when the app stores its own model config — a key already configured in the app's own data root (`penguin config model list --root `). Keys living in the global `~/.penguin` or any other `.penguin` directory do **not** count — a bare `penguin config model list` (no `--root`) reads the global store, because the CLI defaults to the global root unless `--root` is given, so a key showing up there proves nothing for your script and must never be used or copied. + +If neither counted source yields a usable key, **stop immediately and ask the user to configure one — do not write code, and do not keep calling tools to retry**: ask them to add one in the agent's **key vault** (gear icon on the agent's card, Agents page → settings → key vault tab); vault values reach your shell environment on the next task. Re-checking the environment or the vault in a loop just wastes turns — one clear check, then hand back to the user. + +Keep model API keys **project-local**: for an app that stores its own model config, write the key into the project under the working directory with the penguin CLI, **always passing `--root ` for a directory inside the current working directory** (`penguin config model add --root ./penguin_data --provider --model-id --api-key `) — without `--root` it writes to the global `~/.penguin/data` instead. `--provider` is required alongside `--model-id`: a model entry is the `(provider, model_id)` pair and the group is never inferred (use `custom` for an endpoint outside the built-in groups). Otherwise rely on vault-injected environment variables. Never read, copy or fall back to model keys stored in the user's global `~/.penguin` directory — that config belongs to the person running Penguin, not to your script. + ## Model IDs Use exact model ids. If an id is not in the table below and the user has not given one, ask the user to confirm the exact id before writing code. diff --git a/packages/skills/skills/firecrawl/SKILL.md b/packages/skills/skills/firecrawl/SKILL.md new file mode 100644 index 0000000..7aa343a --- /dev/null +++ b/packages/skills/skills/firecrawl/SKILL.md @@ -0,0 +1,69 @@ +--- +name: firecrawl +description: Search the web and scrape pages into clean markdown with the Firecrawl API — query-based discovery, single-URL extraction including public PDFs, driven by curl with a vault-stored API key. +short_description: Web search and page scraping via Firecrawl. +short_description_zh: 用 Firecrawl 做网络搜索与页面抓取。 +version: 1 +updated: 2026-07-20T14:00:00Z +--- + +# Firecrawl + +Firecrawl turns the live web into agent-ready markdown over a plain REST API (`https://api.firecrawl.dev/v2`, `Authorization: Bearer $FIRECRAWL_API_KEY`). Two calls cover most web work: `/search` to discover pages by query, `/scrape` to extract clean content from a URL you already have. Use it whenever a task needs current web information or the content of a specific page. + +## Before you start + +If the user's message only invokes this skill (e.g. "use firecrawl skill") without a concrete task, ask what they want to search or scrape. When the task is concrete, check the credential first: + +```bash +[ -n "$FIRECRAWL_API_KEY" ] && echo ok || echo missing +``` + +If missing (also visible in your Vault Keys section), ask the user to add `FIRECRAWL_API_KEY` to this agent's **key vault** — gear icon on the agent card → settings → key vault tab; keys come from the Firecrawl dashboard (https://www.firecrawl.dev/signin). Vault values reach your shell environment on the next task. Only fall back to the keyless tier (below) when the user cannot provide a key right now. + +## Search + +```bash +curl -sS -X POST https://api.firecrawl.dev/v2/search \ + -H "Authorization: Bearer $FIRECRAWL_API_KEY" -H "content-type: application/json" \ + -d '{"query": "", "limit": 5}' \ + | jq '[.data.web[] | {url, title, description}]' +``` + +- Results live in `.data.web[]`, each with `url` / `title` / `description`; `limit` defaults to 10 (per source). +- Adding `"scrapeOptions": {"formats": ["markdown"]}` returns each result's page content inline — prefer the two-step search → scrape flow instead when only a hit or two matters; content-included search costs far more credits and context. +- Useful filters: `"sources": [{"type": "news"}]` (or `images`), `"includeDomains": ["docs.example.com"]`, `"categories": ["github"]` (or `research` / `pdf`), `"tbs"` for time-bounded queries. + +## Scrape + +```bash +curl -sS -X POST https://api.firecrawl.dev/v2/scrape \ + -H "Authorization: Bearer $FIRECRAWL_API_KEY" -H "content-type: application/json" \ + -d '{"url": ""}' \ + | jq -r '.data.markdown' > .md +``` + +- Markdown is the default format; the page's main content only (`onlyMainContent` defaults to true). Metadata sits in `.data.metadata` (`title`, `sourceURL`, `statusCode`). +- Public document URLs (PDF, DOCX, …) scrape to markdown the same way. +- JS-heavy pages that come back empty: retry with `"waitFor": 2000`. + +## Workflow and context economy + +Search first for discovery, scrape once you have the URL. Never dump full page markdown into your context or reply: pipe it to a file (as above) and read the relevant parts with shell (`grep`/`sed`), then cite `metadata.sourceURL` for every claim you take from a page. + +## No key available + +The keyless free tier only works through official Firecrawl clients and is rate-limited: + +```bash +npx -y firecrawl-cli@latest search "" +npx -y firecrawl-cli@latest scrape -o page.md +``` + +Use it as a stopgap and tell the user to add a real key to the vault — accounts unlock the full API and higher limits. + +## Errors + +- `401` — missing/invalid key: re-check the vault entry name `FIRECRAWL_API_KEY`. +- `402` / `429` — out of credits or rate-limited: report to the user; do not retry-loop. +- Anything else: `https://docs.firecrawl.dev` is the source of truth for request/response schemas. diff --git a/packages/skills/skills/firecrawl/icon.svg b/packages/skills/skills/firecrawl/icon.svg new file mode 100644 index 0000000..c1318dd --- /dev/null +++ b/packages/skills/skills/firecrawl/icon.svg @@ -0,0 +1,3 @@ + + + diff --git a/packages/skills/skills/penguin-cli/SKILL.md b/packages/skills/skills/penguin-cli/SKILL.md index c7a1bea..8d04a40 100644 --- a/packages/skills/skills/penguin-cli/SKILL.md +++ b/packages/skills/skills/penguin-cli/SKILL.md @@ -3,8 +3,8 @@ name: penguin-cli description: Manage model API keys, default models and per-agent vault secrets with the penguin CLI. short_description: Manage models and secrets with the penguin CLI. short_description_zh: 用 penguin CLI 管理模型与密钥。 -version: 1 -updated: 2026-07-17T00:00:00Z +version: 3 +updated: 2026-07-21T00:00:00Z --- # Penguin CLI @@ -17,20 +17,20 @@ If the user's message only invokes this skill (e.g. "use penguin-cli skill") wit ## Models -Add or update a model (upsert by the stored model id; re-run with more options to amend an entry): +Add or update a model (upsert by the `(provider, model_id)` pair; re-run with more options to amend an entry): ```bash -penguin config model add --model-id [--provider ] [--api-key ] [--base-url ] \ +penguin config model add --provider --model-id [--api-key ] [--base-url ] \ [--client-type ] [--context-window ] [--vision | --no-vision] \ [--price-cache-read ] [--price-cache-write ] [--price-output ] \ [--project-id ] [--root ] [--set-default] ``` -- `--model-id` takes the provider's upstream model id (what the API expects). The stored id is always `/`: `--provider` picks the provider group, and when omitted it is inferred from the built-in catalog (unrecognized ids fall back to `custom`). The upstream id is persisted automatically as the entry's request id, so nothing extra is needed for it to reach the API unchanged. +- A model is identified by the `(provider, model_id)` pair, so `--provider` and `--model-id` are **both required** — the group is never inferred from the model id, because gateways resell vendor models under their upstream ids and a wrong guess would send the key to another vendor's endpoint. `--model-id` takes the provider's upstream model id (what the API expects) and is persisted as the entry's request id, so it reaches the API unchanged; `--provider` names the group (`deepseek`, `openai`, `anthropic`, `google`, `openrouter`, `siliconflow`, … — `custom` for any other endpoint). - For any OpenAI chat-completion compatible endpoint use `--client-type openai --base-url `; omit `--client-type` to auto-route by model id. - Prices are USD per million tokens (cache read / cache write / output). - `--vision` / `--no-vision` mark whether the model accepts images; omitting both keeps the current value (default is vision-capable). -- All `penguin config model ...` and `penguin config vault ...` commands accept `--root ` to target another data root (default `PENGUIN_HOME`, then `~/.penguin/data`). +- All `penguin config model ...` and `penguin config vault ...` commands accept `--root ` to target another data root (default `PENGUIN_HOME`, then `~/.penguin/data`). **When configuring models for an AI app you are building, always pass `--root ` pointing at the app's own data directory inside the current working directory** (e.g. `--root ./penguin_data`, the same path the app gives `createAgent({ root })`); running without `--root` writes to the global `~/.penguin/data`, which belongs to the person running Penguin — not to the app. Other model commands: @@ -61,7 +61,7 @@ penguin config lang # persist the CLI language via PENGUIN_LANG in you ## Running agents -`penguin run -m "" [--model-id ] [--agent-id ] [--workspace ] [--approve ]` runs one task; `penguin chat [--resume [session_id]]` starts or resumes an interactive chat with the same options. +`penguin run -m "" [--provider --model-id ] [--agent-id ] [--workspace ] [--approve ]` runs one task; `penguin chat [--resume [session_id]]` starts or resumes an interactive chat with the same options. The model reference stays a pair here too: pass `--provider` and `--model-id` together, or neither to run on the project's default model — one without the other is rejected. ## Storage diff --git a/packages/skills/skills/penguin-sdk/SKILL.md b/packages/skills/skills/penguin-sdk/SKILL.md index 7d155d1..6120e76 100644 --- a/packages/skills/skills/penguin-sdk/SKILL.md +++ b/packages/skills/skills/penguin-sdk/SKILL.md @@ -1,10 +1,10 @@ --- name: penguin-sdk -description: Build AI apps on the Penguin Harness SDK — self-contained projects inside the Workspace, model configuration, and the createSession/run streaming loop. -short_description: Build AI apps on the Penguin Harness SDK. -short_description_zh: 基于 Penguin Harness SDK 构建 AI 应用。 -version: 1 -updated: 2026-07-17T00:00:00Z +description: Build AI apps on the Penguin Harness SDK — self-contained projects, the createSession/run streaming loop, and a complete RAG recipe that ingests documents into a knowledge base and answers with citations behind a web UI. +short_description: Build AI and RAG apps on the Penguin Harness SDK. +short_description_zh: 基于 Penguin SDK 构建 AI 与 RAG 应用。 +version: 11 +updated: 2026-07-21T00:00:00Z --- # Penguin Harness SDK @@ -19,48 +19,65 @@ To have an agent perform a task, use the `run_subagent` tool — the SDK is for ## Before you start -If the user's message only invokes this skill (e.g. "use penguin-sdk skill") without a concrete app to build, ask the user what they want to build. Do not start until the requirement is clear. +If the user's message only invokes this skill (e.g. "use penguin-sdk skill") without a concrete app to build, ask the user what they want to build. But when the request names a concrete goal — even a single sentence like "build a RAG app that answers questions about these docs" — do **not** ask follow-up questions: build it end to end with the defaults in this skill (self-contained workspace project, project default model, BM25 retrieval, web UI styled per the web-design skill) and list the assumptions you made in your final reply. ## Project location -Create the app in the current workspace directory by default (the `CWD` value from your Environment section), as a self-contained project — do not place it under `` or depend on any path outside the project folder. Point the agent data root at a directory inside the project with `createAgent({ root })`, resolved from the source file so it stays relative: +Create the app in the current workspace directory by default (the `CWD` value from your Environment section), as a self-contained project — do not place it under `` or depend on any path outside the project folder. When creating the app's agent, the data root defaults **under the working directory (CWD)** too: point `createAgent({ root })` at a directory inside the project, resolved from the source file so it stays relative: ```ts -const agent = await createAgent({ root: path.resolve(import.meta.dirname, "penguin_data") }); +const agent = await createAgent({ root: path.join(import.meta.dirname, "penguin_data") }); ``` With every reference relative to the project, the user can move or copy the folder anywhere and it still runs. +## Before you build: keys and the data root + +**Important prerequisite — set the key up first, then develop.** For AI-app development, have the user add the model API key in **this agent's key vault** (gear icon on its card, Agents page → settings → key vault tab) *before* you start building, so the credential is in your shell environment when you configure and test the app. Model ids to offer the user can come straight from the penguin CLI catalog (`penguin config model add --help`, and the agenthub-models skill's id table). + +**The app's Penguin data root must live inside the CWD workspace — never `~/.penguin`.** Point `createAgent({ root })` and every `penguin config ... --root ` at a directory under the current working directory (e.g. `./penguin_data`); the global `~/.penguin` directory belongs to the person running Penguin and must never hold the app's config or keys. + +## Check the model first + +Before writing any code, verify a usable model credential exists — a finished app that cannot answer is a failed delivery discovered too late: + +```bash +env | grep -oE "(DEEPSEEK|OPENAI|ANTHROPIC|GEMINI)_API_KEY" || echo none +``` + +Vault keys also appear in your Vault Keys section. **Only two sources count as a usable credential**: a vault-injected environment variable (the check above), or a key already configured in the app's own data root (`penguin config model list --root `). Keys living in the global `~/.penguin` or any other `.penguin` directory do **not** count — a bare `penguin config model list` (no `--root`) reads the global store, because the CLI defaults to the global root unless `--root` is given, so a key showing up there proves nothing for the app and must never be used or copied. + +If neither counted source yields a usable key, **stop immediately and ask the user to configure one — do not start building, and do not keep calling tools to retry**: ask them to open the agent's settings via the **gear icon** on its card (left side, Agents page) and add a model API key (e.g. `DEEPSEEK_API_KEY`) in the **key vault** tab — vault values reach your shell environment on the next task. Re-running `env`, re-checking the vault, or attempting the build in a loop wastes turns and money; one clear check, then hand back to the user. Build only after a credential is confirmed, or clearly agree with the user to build now and verify later. + ## Setup ```bash -npm install @prismshadow/penguin-core +npm install @prismshadow/penguin-core tsx ``` -If the package is not yet available on your npm registry (it is developed in the PenguinHarness monorepo and may not be published), develop inside a checkout of the PenguinHarness repo instead: add your app as a workspace package under `packages/` and depend on `"@prismshadow/penguin-core": "workspace:*"`, then run `pnpm install && pnpm build` at the repo root. Tell the user which route you took. +If the package is not on your npm registry (it is developed in the PenguinHarness monorepo and may not be published), develop inside a checkout of the PenguinHarness repo instead: add your app as a workspace package under `packages/`, depend on `"@prismshadow/penguin-core": "workspace:*"`, then `pnpm install && pnpm build` at the repo root. Tell the user which route you took. -A model must be configured for the app's data root. Two ways: +Configure a model for the app's data root, in this order — stop at the first that works: -1. The penguin CLI, pointed at the project-local data directory: +1. `penguin config model add --root --provider --model-id --api-key [--base-url ] [--client-type openai] --set-default` — prefer `--client-type openai --base-url ` (works with any OpenAI-compatible endpoint; exact ids in the agenthub-models skill). `--provider` is required: a model is always the `(provider, model_id)` pair and the group is never inferred from the id (`custom` for an endpoint outside the built-in groups). +2. Environment variables cover the **credential only** (`DEEPSEEK_API_KEY`, `OPENAI_API_KEY`, `ANTHROPIC_API_KEY`, …) — model selection still comes from the project config, whose preset default is `deepseek-v4-pro`. Env-only setup therefore works out of the box only with `DEEPSEEK_API_KEY`; for another vendor either run the CLI command above or pass a configured `{ provider, modelId }` pair to `createSession`. -```bash -penguin config model add --root --model-id --api-key [--base-url ] [--client-type openai] --set-default -``` +Keep model API keys **project-local**: configure them with the penguin CLI into the app's own data root under the working directory, so the project stays self-contained and movable. When building an AI app, **always pass `--root ` pointing at the app's data directory inside the current working directory** (the same path you give `createAgent({ root })`, e.g. `./penguin_data`) — never run `penguin config ...` without `--root`, or it writes to the global `~/.penguin/data` instead of the project. Never read, copy or fall back to model keys stored in the user's global `~/.penguin` directory — that config belongs to the person running Penguin, not to the app you are building. -Generally prefer the OpenAI protocol client (chat completion): `--client-type openai --base-url ` works with any OpenAI-compatible endpoint. Use exact model ids — see the agenthub-models skill for the id table. +Model config lives in one hidden file under the data root's project directory: `.project_config.toml`. It is CLI-only — never read, print or edit it. -2. Environment variables as a fallback: without a configured credential the SDK reads the provider's env vars (e.g. `OPENAI_API_KEY`, `ANTHROPIC_API_KEY`, `DEEPSEEK_API_KEY`). +If neither route yields a usable credential, do not fake the verification: finish the build, report it as **unverified**, and tell the user exactly how to unblock you — in the Penguin web app, open this agent's settings via the **gear icon** on its card (Agents page) and add a model API key (e.g. `DEEPSEEK_API_KEY`) under the **key vault** tab. Vault keys are injected into your shell environment on the next task, so once the user has added one, you can run the self-test to completion. -Model config lives in a single hidden file under the data root's project directory: `.project_config.toml` (model list, settings and per-model credentials such as `api_key` inlined in each model entry). It is CLI-only — never read, print or edit it; the CLI above manages it. +## Streaming loop -## Minimal conversational app +The raw `run()` stream mixes model, event and session-meta payloads — always narrow with the exported guards (`isModelMessage`, `isCompleteModelMessage`, `isEventMessage`) before touching `payload.type`; accessing `msg.payload.type` directly does not typecheck. ```ts import path from "node:path"; import readline from "node:readline/promises"; -import { createAgent, userText } from "@prismshadow/penguin-core"; +import { createAgent, isModelMessage, userText } from "@prismshadow/penguin-core"; -const agent = await createAgent({ root: path.resolve(import.meta.dirname, "penguin_data") }); +const agent = await createAgent({ root: path.join(import.meta.dirname, "penguin_data") }); const session = await agent.createSession({ workspaceDir: process.cwd() }); const rl = readline.createInterface({ input: process.stdin, output: process.stdout }); @@ -71,18 +88,246 @@ for (;;) { for await (const msg of session.run([userText(line)], { approve: async () => "allow", // demo only — a real app should ask its user ("deny" blocks the call) })) { - const p = msg.payload; - if (p.type === "partial_text" && p.event_type === "delta") process.stdout.write(p.text); + if (isModelMessage(msg)) { + const p = msg.payload; + if (p.type === "partial_text" && p.event_type === "delta") process.stdout.write(p.text); + } } process.stdout.write("\n"); } rl.close(); +session.dispose(); ``` -Key points: - +- `createSession({ workspaceDir, provider, modelId })` — `workspaceDir` must already exist (omit for an auto temp dir); the model reference is the `(provider, modelId)` pair, so pass both to pick a configured model or neither for the project default — passing one alone throws. +- The `approve` callback gates every tool call; **omitting it denies everything**. - An Agent's behavior is edited in its `agent_state/` files (system_config.yaml, AGENTS.md, skills/), not in code. -- `createSession({ workspaceDir, modelId })` — omit `workspaceDir` for a temporary workspace, omit `modelId` for the project default model. -- `session.run(messages, { approve, signal })` is an async generator of OmniMessages; filter the payload types you care about (`partial_text` deltas carry the streamed answer). -- The `approve` callback gates every tool call — auto-allow is for demos; a production app should prompt its user before returning `"allow"`. -- Call `session.run()` again on the same Session for the next user turn — the context carries over. +- Call `session.dispose()` when done to release background processes. + +## RAG knowledge app + +The default recipe when the user wants an app that answers questions over a document set ("docs QA", "knowledge base", "chat with our docs", "become an expert on X"). The core contributes the agent loop only — retrieval is app code. Default to **lexical BM25**: no extra dependencies, no embedding credential, works offline. (Semantic upgrade: embed chunks via `@prismshadow/agenthub` — see the agenthub-models skill — and rank by cosine; only when an embedding-capable key is configured.) + +``` +my-app/ + package.json # "type": "module"; scripts: ingest / start + persona.md # the embedded agent's role — write it per the agent-creation skill + ingest.ts # corpus/ → data/index.json; initializes penguin_data/, installs persona + rag.ts # BM25 retrieval over the chunk index + server.ts # POST /api/ask streams SSE; serves public/ + public/index.html # chat UI — build it per the web-design skill + corpus/ # collected source documents + data/index.json # generated chunk index + penguin_data/ # agent data root (generated; model config lives here) +``` + +**Collect** — clone or fetch the sources into `corpus/`, keeping only text formats: + +```bash +git clone --depth 1 corpus/ # or curl pages into corpus/ +find corpus -type f ! -regex '.*\.\(md\|mdx\|txt\|html?\)$' -delete && rm -rf corpus/*/.git +``` + +**Ingest** (`ingest.ts`) — split on markdown headings, cap chunk size, write one JSON index; also initialize `penguin_data/`, install the persona and strip the skills the embedded agent doesn't need: + +```ts +import fs from "node:fs"; +import path from "node:path"; +import { createAgent } from "@prismshadow/penguin-core"; + +const ROOT = import.meta.dirname; +const walk = (d: string): string[] => + fs.readdirSync(d, { withFileTypes: true }).flatMap((e) => + e.isDirectory() ? walk(path.join(d, e.name)) : [path.join(d, e.name)]); + +const STATE = path.join( + ROOT, "penguin_data", "default_project", "agents", "default_agent", "agent_state"); +await createAgent({ root: path.join(ROOT, "penguin_data") }); +fs.copyFileSync(path.join(ROOT, "persona.md"), path.join(STATE, "AGENTS.md")); +// A fresh default_agent is initialized with the whole built-in Skill library, and every installed +// Skill's metadata is injected into the system prompt of every /api/ask. This app only answers +// from retrieved context, so remove them: unrelated skill descriptions cost tokens on each +// question and pull the answer off-topic when one happens to match the wording of a question. +fs.rmSync(path.join(STATE, "skills"), { recursive: true, force: true }); + +const chunks: { id: number; source: string; heading: string; text: string }[] = []; +for (const f of walk(path.join(ROOT, "corpus")).filter((f) => /\.(md|mdx|txt|html?)$/i.test(f))) { + const raw = fs.readFileSync(f, "utf8"); + const text = /\.html?$/i.test(f) ? raw.replace(/<[^>]+>/g, " ") : raw; + const source = path.relative(ROOT, f); + let heading = path.basename(f); + for (const block of text.split(/^(?=#{1,3} )/m)) { + heading = block.match(/^#{1,3} (.+)/)?.[1] ?? heading; + for (let i = 0; i < block.length; i += 1500) { + const piece = block.slice(i, i + 1500).trim(); + if (piece.length > 40) chunks.push({ id: chunks.length, source, heading, text: piece }); + } + } +} +fs.mkdirSync(path.join(ROOT, "data"), { recursive: true }); +fs.writeFileSync(path.join(ROOT, "data", "index.json"), JSON.stringify(chunks)); +console.log(`indexed ${chunks.length} chunks`); +``` + +**Retrieve** (`rag.ts`) — standard BM25 (k1 = 1.2, b = 0.75); the tokenizer treats each CJK character as a token so Chinese queries work. The corpus-wide statistics (per-chunk term frequencies, document frequencies, average length) never change once the corpus is indexed, so build them **once** in `loadIndex` — a per-query rescan would make every question O(corpus): + +```ts +import fs from "node:fs"; +import path from "node:path"; + +export interface Chunk { id: number; source: string; heading: string; text: string } +export interface Index { + chunks: Chunk[]; + tf: Map[]; // per-chunk term → count + len: number[]; // per-chunk token length + df: Map; // term → number of chunks containing it + avg: number; // mean chunk length (BM25 length normalization) +} + +const tokenize = (s: string): string[] => s.toLowerCase().match(/[a-z0-9]+|[一-鿿]/g) ?? []; + +export function loadIndex(): Index { + const chunks: Chunk[] = JSON.parse( + fs.readFileSync(path.join(import.meta.dirname, "data", "index.json"), "utf8")); + const tf: Map[] = []; + const len: number[] = []; + const df = new Map(); + for (const c of chunks) { + const toks = tokenize(`${c.heading} ${c.text}`); + const m = new Map(); + for (const t of toks) m.set(t, (m.get(t) ?? 0) + 1); + for (const t of m.keys()) df.set(t, (df.get(t) ?? 0) + 1); + tf.push(m); + len.push(toks.length); + } + const avg = len.reduce((n, l) => n + l, 0) / Math.max(len.length, 1); + return { chunks, tf, len, df, avg }; +} + +export function search(index: Index, query: string, k = 6): Chunk[] { + const { chunks, tf, len, df, avg } = index; + const q = [...new Set(tokenize(query))]; + const score = (i: number): number => { + let s = 0; + for (const t of q) { + const f = tf[i]!.get(t) ?? 0; + if (f === 0) continue; + const n = df.get(t) ?? 0; + s += Math.log(1 + (chunks.length - n + 0.5) / (n + 0.5)) * + (f * 2.2) / (f + 1.2 * (0.25 + (0.75 * len[i]!) / avg)); + } + return s; + }; + return chunks.map((_, i) => [score(i), i] as const) + .filter(([s]) => s > 0).sort((a, b) => b[0] - a[0]).slice(0, k) + .map(([, i]) => chunks[i]!); +} +``` + +**Answer & serve** (`server.ts`) — one Session per request (stateless QA), retrieved chunks numbered into the prompt, deltas streamed over SSE, sources sent as the final event. A pure QA session needs no tool calls — deny every approval; a denied or tool-less turn terminates normally. Do **not** clear the toolset with `tools: { builtin: [] }`: an empty tools array is sent to the provider verbatim and some OpenAI-compatible endpoints reject it with a 400, which surfaces as a silent empty answer. Guard the request boundary — a malformed body must return 400, never reject the async handler (an unhandled rejection takes the whole server down) — and abort the run if the client disconnects mid-answer so you stop generating (and paying) for a page nobody is reading. + +```ts +import fs from "node:fs"; +import http from "node:http"; +import path from "node:path"; +import { createAgent, isModelMessage, userText } from "@prismshadow/penguin-core"; +import { loadIndex, search } from "./rag.ts"; + +const ROOT = import.meta.dirname; +const PUB = path.join(ROOT, "public"); +const agent = await createAgent({ root: path.join(ROOT, "penguin_data") }); +const index = loadIndex(); +const MIME: Record = { ".html": "text/html", ".css": "text/css", ".js": "text/javascript" }; + +http.createServer(async (req, res) => { + res.on("error", () => {}); // a client that vanishes mid-write must not throw an uncaught EPIPE + if (req.method === "POST" && req.url === "/api/ask") { + let question: string; + try { + let body = ""; + for await (const part of req) body += part; // a mid-body connection reset rejects here — caught below, never fatal + const parsed = JSON.parse(body) as { question?: unknown }; + if (typeof parsed.question !== "string" || !parsed.question.trim()) throw new Error(); + question = parsed.question; + } catch { + res.writeHead(400, { "content-type": "application/json" }); + res.end(JSON.stringify({ error: "expected a JSON body { question: string }" })); + return; + } + const hits = search(index, question); + const context = hits.map((c, i) => `[${i + 1}] ${c.source} — ${c.heading}\n${c.text}`).join("\n\n"); + const ac = new AbortController(); + res.on("close", () => ac.abort()); // client navigated away → cancel the in-flight generation + // Create the Session BEFORE committing headers: a model-config failure then returns a real + // HTTP error instead of an unhandled rejection with a 200 already on the wire. + let session; + try { + session = await agent.createSession({ workspaceDir: ROOT }); + } catch { + res.writeHead(503, { "content-type": "application/json" }); + res.end(JSON.stringify({ error: "no model configured yet — see the setup steps" })); + return; + } + res.writeHead(200, { "content-type": "text/event-stream", "cache-control": "no-cache" }); + try { + const prompt = `Answer from the context below; cite blocks inline as [1][2]. If the context is not enough, say so.\n\n${context}\n\nQuestion: ${question}`; + for await (const msg of session.run([userText(prompt)], { approve: async () => "deny", signal: ac.signal })) { + if (isModelMessage(msg)) { + const p = msg.payload; + if (p.type === "partial_text" && p.event_type === "delta" && !res.writableEnded) + res.write(`data: ${JSON.stringify({ delta: p.text })}\n\n`); + } + } + // Sources carry the matched chunk text verbatim: the UI must be able to show the exact + // block behind each [n], not just a file link. + if (!res.writableEnded) + res.write(`data: ${JSON.stringify({ sources: hits.map((c) => ({ source: c.source, heading: c.heading, url: `/${c.source}`, text: c.text })) })}\n\n`); + } catch { + // The run failed after headers were sent, or the client left: surface an error event (best effort), then clean up. + if (!res.writableEnded) res.write(`data: ${JSON.stringify({ error: "generation failed" })}\n\n`); + } finally { + session.dispose(); + if (!res.writableEnded) res.end(); + } + return; + } + const pathname = (req.url ?? "/").split("?")[0] ?? "/"; + // /corpus/* serves the source documents read-only, so citation links resolve to real files. + const inCorpus = pathname.startsWith("/corpus/"); + const base = inCorpus ? path.join(ROOT, "corpus") : PUB; + const rel = inCorpus + ? pathname.slice("/corpus/".length) + : pathname === "/" + ? "index.html" + : pathname.slice(1); + const file = path.normalize(path.join(base, rel)); + if (file.startsWith(base + path.sep) && fs.existsSync(file) && fs.statSync(file).isFile()) { + res.writeHead(200, { "content-type": MIME[path.extname(file)] ?? "text/plain" }); + res.end(fs.readFileSync(file)); + } else { + res.writeHead(404); + res.end(); + } +}).listen(Number(process.env.PORT ?? 4630), () => console.log("http://localhost:4630")); +``` + +**UI** (`public/index.html`) — a chat interface built per the web-design skill: message list, streamed assistant text appended delta by delta, the final `sources` event rendered as citation chips, an empty state inviting the first question with **3–4 example questions the corpus can actually answer** (pill chips; clicking one submits it), and a visible error state when `/api/ask` fails. Citations must satisfy both of these, never bare text: + +- **Reveal the original chunk**: clicking a citation chip (or an inline `[n]`) opens a popover/panel showing the matched chunk's `text` from the sources event **verbatim** — the numbering maps 1:1 to the context blocks in the prompt, so `[n]` always reveals exactly the block the answer drew on. +- **Link to the real document**: inside the popover, `` using the `url` field (`/corpus/`, which this server serves) — clicking the chip itself opens the popover, the document link lives within it. When the corpus was cloned from a public repository, prefer mapping the path to the canonical upstream page instead (e.g. the GitHub blob URL derived from the clone URL). + +**Persona** (`persona.md`) — the embedded agent's role, written per the agent-creation skill. Shape: one role sentence ("You are an expert on X; you answer strictly from the provided context blocks"), citation and refusal rules, answer language follows the question. + +## Verify before you hand over + +Never declare the app done without running it: + +1. `npm install` succeeds (or the workspace route builds). +2. Model configured for `penguin_data` (CLI or env var; no usable key → see Setup: ask the user to add one to this agent's key vault, and report the app as unverified for now). +3. `npm run ingest` prints `indexed N chunks` with N > 0. +4. Start `npm start` in the background, then ask a real question: + `curl -N -sS -X POST localhost:4630/api/ask -H 'content-type: application/json' -d '{"question":""}'` — expect streamed `data:` deltas ending in a `sources` event that carries `source`, `url` **and the matched chunk `text`** per hit. If nothing streams, the model call failed: re-check step 2 and the provider endpoint before touching the code. + Then `curl` one of the returned source `url`s — it must return the document, not a 404 (citation links have to resolve). +5. Open the UI (or screenshot it) to confirm the layout renders. + +Fix any failure and re-verify. Report with backtick-wrapped relative paths (`server.ts`, `public/index.html`, …), how to start the app, and the assumptions you made. diff --git a/packages/skills/skills/web-design/SKILL.md b/packages/skills/skills/web-design/SKILL.md index db57a28..4a9b087 100644 --- a/packages/skills/skills/web-design/SKILL.md +++ b/packages/skills/skills/web-design/SKILL.md @@ -1,40 +1,98 @@ --- name: web-design -description: Default visual language for generated web pages — minimal black-white-gray, square corners, hierarchy from weight and spacing; colors, radii and gradients only on explicit request. -short_description: Minimal monochrome defaults for generated web pages. -short_description_zh: 生成网页的极简黑白默认视觉规范。 -version: 1 -updated: 2026-07-17T00:00:00Z +description: Penguin visual language for generated web pages and app UIs — GitHub-style simplicity with a single blue accent, light and pure-black dark themes, design tokens, and component and chat-interface recipes. +short_description: Penguin-style visual defaults for generated web UIs. +short_description_zh: 生成网页的 Penguin 风格视觉规范。 +version: 3 +updated: 2026-07-20T15:00:00Z --- # Web Design -Default visual rules for every web page or frontend interface you generate. Apply them to any HTML/CSS you produce unless the user explicitly asks otherwise. +Default visual language for every web page or frontend you generate, distilled from the Penguin Harness landing page and web app. The idea is **GitHub-style simplicity**: solid backgrounds, 1px borders instead of shadows, system fonts, one blue accent used sparingly. Depth comes from hairline borders, not shadows or gradients; dark mode is pure black, not navy. Apply these defaults unless the user explicitly asks for another style. ## Before you start -If the user's message only invokes this skill (e.g. "use web-design skill") without a concrete page or interface to build, ask the user what they want to build. Do not start until the requirement is clear. +If the user's message only invokes this skill (e.g. "use web-design skill") without a concrete page or interface to build, ask what they want to build. When a concrete build is already requested (an app UI, a landing page, a RAG chat interface), do **not** ask about styling — apply the defaults below. -## Core rules +Non-negotiable for ANY text input that sends on Enter: **never send while an IME composition is in progress** (check `event.isComposing`, falling back to `event.keyCode === 229`, on keydown). For Chinese/Japanese/Korean input methods, that Enter only confirms the composed text — auto-sending on it fires half-typed messages. Details in the composer recipe below. -- Monochrome only: black, white and grays. No accent colors, no gradients, no decorative shadows. -- Square corners everywhere: `border-radius: 0` on buttons, cards, inputs, images and modals. -- Hierarchy comes from font weight, font size, spacing and thin light-gray borders — never from colored blocks or backgrounds. -- Generous whitespace: prefer more spacing over more dividers; let sections breathe. -- Introduce colors, rounded corners or gradients **only when the user explicitly asks for them**, and only where asked — the rest of the page stays monochrome and square. - -## Tokens - -Base every stylesheet on a small monochrome token set: +## Design tokens ```css :root { - --fg: #111111; /* primary text */ - --fg-muted: #666666; /* secondary text */ - --bg: #ffffff; /* page background */ - --bg-subtle: #f5f5f5; /* raised surfaces */ - --border: #e2e2e2; /* hairline borders (1px) */ + color-scheme: light; + /* Brand blue — the only accent family. Use sparingly: links, eyebrows, tiny dots, tints. */ + --brand-50: #e8f0fe; --brand-100: #d2e3fc; --brand-300: #8ab4f8; --brand-500: #4285f4; + --brand-600: #1a73e8; /* accent text/icons in light mode */ --brand-700: #0b57d0; /* links on white */ + /* Neutrals (Tailwind gray) */ + --gray-50: #f9fafb; --gray-100: #f3f4f6; --gray-200: #e5e7eb; --gray-300: #d1d5db; + --gray-400: #9ca3af; --gray-500: #6b7280; --gray-600: #4b5563; --gray-900: #111827; + --bg: #ffffff; --surface: #ffffff; --border: var(--gray-200); --control-border: var(--gray-300); + --fg: var(--gray-900); --fg-muted: var(--gray-600); --fg-faint: var(--gray-500); + --accent-bg: #111827; --accent-fg: #ffffff; /* primary buttons are near-black, not blue */ + --radius-control: 6px; /* buttons, inputs, chips */ --radius-card: 12px; /* cards, panels */ + --ease: cubic-bezier(0.2, 0.7, 0.3, 1); +} +.dark { + color-scheme: dark; /* pure black, no blue tint */ + --bg: #000000; --surface: #0d0d0d; --border: #1f1f1f; --control-border: #303030; + --fg: #f3f4f6; --fg-muted: #9ca3af; --fg-faint: #6b7280; + --accent-bg: #f3f4f6; --accent-fg: #111827; /* primary button inverts to light */ + --brand-600: #8ab4f8; --brand-700: #8ab4f8; /* brand text flips to the 300 tone */ } ``` -Tailwind equivalent: stick to `text-neutral-900` / `text-neutral-500` / `bg-white` / `bg-neutral-100` / `border-neutral-200` / `rounded-none`; do not use color utilities (`blue-*`, `emerald-*`, ...), `rounded-*` variants other than `rounded-none`, or `bg-gradient-*`. +- Toggle dark mode with a `dark` class on `` (persist the choice; default to `prefers-color-scheme`). +- Primary buttons are **neutral black/white**, never blue fills. Brand blue is reserved for accents: links, section eyebrows, small status dots, `--brand-50` tinted chips. +- Pills, badges and dots use `border-radius: 9999px`; everything else uses the two radii above. No gradients or elevation shadows on content (modals excepted; flat focus rings drawn with `box-shadow` are fine). + +## Typography + +System fonts only — no CDN fonts, no @font-face. The CJK entries matter (bilingual product): + +```css +body { font-family: ui-sans-serif, system-ui, -apple-system, "Segoe UI", Roboto, + "PingFang SC", "Microsoft YaHei", sans-serif; -webkit-font-smoothing: antialiased; } +code, pre, kbd { font-family: ui-monospace, SFMono-Regular, Menlo, Consolas, monospace; } +``` + +Headings are always `font-weight: 600` with `letter-spacing: -0.025em` — nothing heavier, no thin weights. Scale: page/hero title 30–36px; section title 24–30px; card title 16–18px; body 14px / line-height 1.5; captions 12px in `--fg-faint`; code 13px / line-height 1.85. Hierarchy comes from weight, size, spacing and hairlines — never from colored blocks. + +## Language + +Write the page's UI copy in the language the user's request was made in — a Chinese request gets a Chinese interface, an English request an English one — and set `` to match. Code identifiers, CSS class names and code comments stay English. + +## Components + +- **Primary button** — `background: var(--accent-bg); color: var(--accent-fg); border-radius: 6px; padding: 6px 12px; font-size: 14px; font-weight: 500;` hover: `opacity: .9`. Large CTA variant: height 44px, radius 8px, padding 0 20px. +- **Secondary button** — white/`--surface` bg, `1px solid var(--control-border)`, same paddings; hover swaps bg to `--gray-100`/dark `#1f1f1f`. +- **Card** — `border: 1px solid var(--border); border-radius: 12px; background: var(--surface); padding: 24px;` hover changes **only the border color** (one step darker) — no lift, shadow or scale. +- **Input / textarea** — `border: 1px solid var(--control-border); border-radius: 6px; padding: 8px 12px;` focus: border one step darker + `box-shadow: 0 0 0 2px rgb(156 163 175 / .3)`; no default outline stacking. +- **Pill chip** — `border-radius: 9999px; border: 1px solid var(--border); padding: 4px 10px; font-size: 12px; color: var(--fg-muted);` hover: bg `--gray-50`. Brand-tinted variant: `--brand-50` bg, `--brand-700` text. +- **Sticky nav** — `height: 56px; border-bottom: 1px solid var(--border); background: color-mix(in srgb, var(--bg) 85%, transparent); backdrop-filter: blur(8px);` logo 28px + product name 15px semibold. +- **Code block** — `--gray-50`/dark `--surface` bg card with a bordered header row (mono 12px label + copy button), body mono 13px. +- **Focus** — `:focus-visible { outline: 3px solid rgb(107 114 128 / .4); outline-offset: 2px; }`. + +## Motion + +One easing everywhere: `var(--ease)`, durations 120–280ms. Entrances rise in (`translateY(10px)` + fade, 280ms, stagger siblings by 40ms); overlays fade (120ms); hover feedback is **color-only** (`transition: color, background-color, border-color 150ms`) — never transform on hover. Always include: + +```css +@media (prefers-reduced-motion: reduce) { * { animation: none !important; transition: none !important; } } +``` + +## Chat / RAG app layout + +The default shape for a generated conversational or docs-QA app: + +- **Shell** — centered column, `max-width: 48rem`, `padding: 0 16px`; sticky nav on top with the app name; message list grows, composer pinned at the bottom. +- **Empty state** — vertically centered title + one-line subtitle in `--fg-muted`, over an optional dot-grid backdrop (`background-image: radial-gradient(rgb(26 115 232 / .14) 1px, transparent 1px); background-size: 22px 22px;` faded out with a bottom mask) — the only decorative flourish allowed. Below it, a wrapped row of 3–4 example-question pill chips the app can genuinely answer; clicking one fills and submits the composer. +- **Messages** — user messages right-aligned in a `--gray-100`/dark `#1f1f1f` rounded bubble (radius 12px, padding 8px 14px, max-width 85%); assistant messages plain on the page background, no bubble. Stream deltas into the assistant message as they arrive with a 1-character pulsing cursor; render markdown. +- **Citations** — after an answer, a wrapped row of pill chips: `[1] path — heading`, brand-tinted variant. Clicking a chip (or an inline `[n]` in the answer) opens a popover/panel showing the **verbatim original text block** the citation refers to, with a link to open the full source document; a citation that is only a label or only a link is not enough. +- **Composer** — a bordered card (radius 12px) with a borderless textarea inside and a small primary send button bottom-right; Enter sends, Shift+Enter for newline; disable while streaming. **Never send while an IME composition is in progress**: on keydown, ignore Enter when `event.isComposing` (or `event.keyCode === 229`) — for CJK input methods that Enter only confirms the composed text, and auto-sending on it fires half-typed messages. +- **States** — loading: three pulsing dots in `--fg-faint`; error: 13px `#b91c1c` text on `#fef2f2` (dark: `#f87171` on `#450a0a`) in a rounded box with a retry affordance. Never leave a silent failure. + +## Page layout (marketing / landing) + +Content width `max-width: 72rem`, gutters 16–24px; section rhythm `padding: 64–96px 0`; section header = uppercase 14px semibold brand-colored eyebrow → title → one-line subtitle in `--fg-muted`, then a card grid (`gap: 20px`, 2–3 columns, collapsing to one on mobile). Hero: centered, logo + name, headline with at most one brand-colored word, CTA pair (primary + secondary). Responsive by default: single column under 640px, tap targets ≥ 40px, no horizontal scroll. diff --git a/packages/skills/src/index.ts b/packages/skills/src/index.ts index 4ab175f..38724a4 100644 --- a/packages/skills/src/index.ts +++ b/packages/skills/src/index.ts @@ -139,34 +139,28 @@ export function loadLibrarySkills(): LibrarySkill[] { */ export const SKILL_GROUPS: SkillGroupInfo[] = [ { - id: "agent-development", - title: "Agent Development", - titleZh: "Agent 开发", - skills: ["agent-creation", "benchmark-design", "agent-evaluation", "agent-optimization"], + id: "office-productivity", + title: "Office Productivity", + titleZh: "办公效率", + skills: ["data-analysis", "firecrawl"], }, { - id: "data-analysis", - title: "Data Analysis", - titleZh: "数据分析", - skills: ["data-analysis"], + id: "software-development", + title: "Software Development", + titleZh: "软件开发", + skills: ["web-design", "software-engineering"], }, { - id: "penguin-development", - title: "Penguin Development", - titleZh: "Penguin 开发", + id: "ai-app-development", + title: "AI App Development", + titleZh: "AI 应用开发", skills: ["penguin-sdk", "penguin-cli", "agenthub-models"], }, { - id: "web-development", - title: "Web Development", - titleZh: "网页开发", - skills: ["web-design"], - }, - { - id: "software-engineering", - title: "Software Engineering", - titleZh: "软件工程", - skills: ["software-engineering"], + id: "agent-tuning", + title: "Agent Tuning", + titleZh: "Agent 调优", + skills: ["agent-creation", "benchmark-design", "agent-evaluation", "agent-optimization"], }, ]; diff --git a/packages/skills/test/skills.test.ts b/packages/skills/test/skills.test.ts index 0260b0f..b2e8933 100644 --- a/packages/skills/test/skills.test.ts +++ b/packages/skills/test/skills.test.ts @@ -40,8 +40,9 @@ describe("loadLibrarySkills", () => { expect(skill.shortDescriptionZh, skill.name).toBeTruthy(); expect(skill.shortDescription!.length, skill.name).toBeLessThan(skill.description.length); expect(skill.shortDescriptionZh!.length, skill.name).toBeLessThan(skill.description.length); - // Pre-release, version is always 1. - expect(skill.version).toBe(1); + // version is a natural number, bumped on every content change (updated moves with it). + expect(Number.isInteger(skill.version), skill.name).toBe(true); + expect(skill.version, skill.name).toBeGreaterThanOrEqual(1); expect(skill.updated).toMatch(/^\d{4}-\d{2}-\d{2}T\d{2}:\d{2}:\d{2}(?:\.\d+)?Z$/); // content is the full SKILL.md text including frontmatter (written as-is on install). expect(skill.content.startsWith("---\n")).toBe(true); @@ -88,32 +89,32 @@ describe("loadSkillGroups / groupSkills", () => { it("loads groups per SKILL_GROUPS, members complete with Chinese titles, no Other group", () => { const groups = loadSkillGroups(); expect(groups.map((g) => g.id)).toEqual([ - "agent-development", - "data-analysis", - "penguin-development", - "web-development", - "software-engineering", + "office-productivity", + "software-development", + "ai-app-development", + "agent-tuning", ]); - expect(groups[0]!.skills.map((s) => s.name)).toEqual([ - "agent-creation", - "benchmark-design", - "agent-evaluation", - "agent-optimization", - ]); - expect(groups[1]!.skills.map((s) => s.name)).toEqual(["data-analysis"]); - expect(groups[1]!.title).toBe("Data Analysis"); - expect(groups[1]!.titleZh).toBe("数据分析"); + expect(groups[0]!.skills.map((s) => s.name)).toEqual(["data-analysis", "firecrawl"]); + expect(groups[0]!.title).toBe("Office Productivity"); + expect(groups[0]!.titleZh).toBe("办公效率"); + expect(groups[1]!.skills.map((s) => s.name)).toEqual(["web-design", "software-engineering"]); + expect(groups[1]!.title).toBe("Software Development"); + expect(groups[1]!.titleZh).toBe("软件开发"); expect(groups[2]!.skills.map((s) => s.name)).toEqual([ "penguin-sdk", "penguin-cli", "agenthub-models", ]); - expect(groups[3]!.skills.map((s) => s.name)).toEqual(["web-design"]); - expect(groups[3]!.title).toBe("Web Development"); - expect(groups[3]!.titleZh).toBe("网页开发"); - expect(groups[4]!.skills.map((s) => s.name)).toEqual(["software-engineering"]); - expect(groups[4]!.title).toBe("Software Engineering"); - expect(groups[4]!.titleZh).toBe("软件工程"); + expect(groups[2]!.title).toBe("AI App Development"); + expect(groups[2]!.titleZh).toBe("AI 应用开发"); + expect(groups[3]!.skills.map((s) => s.name)).toEqual([ + "agent-creation", + "benchmark-design", + "agent-evaluation", + "agent-optimization", + ]); + expect(groups[3]!.title).toBe("Agent Tuning"); + expect(groups[3]!.titleZh).toBe("Agent 调优"); for (const group of groups) { expect(group.title).toBeTruthy(); expect(group.titleZh).toBeTruthy(); @@ -126,14 +127,13 @@ describe("loadSkillGroups / groupSkills", () => { const stray = fakeSkill("stray-skill"); const groups = groupSkills([fakeSkill("agent-creation"), stray]); expect(groups.map((g) => g.id)).toEqual([ - "agent-development", - "data-analysis", - "penguin-development", - "web-development", - "software-engineering", + "office-productivity", + "software-development", + "ai-app-development", + "agent-tuning", "other", ]); - const other = groups[5]!; + const other = groups[4]!; expect(other.title).toBe("Other"); expect(other.titleZh).toBe("其他"); expect(other.skills).toEqual([stray]); @@ -142,29 +142,29 @@ describe("loadSkillGroups / groupSkills", () => { it("groupSkills: missing members are skipped; no Other group when all are grouped", () => { const groups = groupSkills([fakeSkill("penguin-cli")]); expect(groups.map((g) => g.id)).toEqual([ - "agent-development", - "data-analysis", - "penguin-development", - "web-development", - "software-engineering", + "office-productivity", + "software-development", + "ai-app-development", + "agent-tuning", ]); expect(groups[0]!.skills).toEqual([]); expect(groups[1]!.skills).toEqual([]); expect(groups[2]!.skills.map((s) => s.name)).toEqual(["penguin-cli"]); expect(groups[3]!.skills).toEqual([]); - expect(groups[4]!.skills).toEqual([]); }); it("SKILL_GROUPS hardcodes member names (sole group info source outside library files)", () => { expect(SKILL_GROUPS.map((g) => ({ id: g.id, skills: g.skills }))).toEqual([ + { id: "office-productivity", skills: ["data-analysis", "firecrawl"] }, + { id: "software-development", skills: ["web-design", "software-engineering"] }, { - id: "agent-development", + id: "ai-app-development", + skills: ["penguin-sdk", "penguin-cli", "agenthub-models"], + }, + { + id: "agent-tuning", skills: ["agent-creation", "benchmark-design", "agent-evaluation", "agent-optimization"], }, - { id: "data-analysis", skills: ["data-analysis"] }, - { id: "penguin-development", skills: ["penguin-sdk", "penguin-cli", "agenthub-models"] }, - { id: "web-development", skills: ["web-design"] }, - { id: "software-engineering", skills: ["software-engineering"] }, ]); }); }); diff --git a/packages/web/e2e/auth.mjs b/packages/web/e2e/auth.mjs index 6cc4850..4d471a5 100644 --- a/packages/web/e2e/auth.mjs +++ b/packages/web/e2e/auth.mjs @@ -1,7 +1,7 @@ /** * e2e auth helper: with signup disabled, test users are always provisioned via * the built-in admin account, then logged in. The server seeds an admin - * (admin / admin123) on startup; a single e2e run shares one data root, and + * (admin / penguin-2026) on startup; a single e2e run shares one data root, and * provisioning is idempotent (reuses the user if it already exists) so a * single spec can be rerun on its own. */ @@ -9,7 +9,7 @@ import { request } from "@playwright/test"; const BASE = process.env.BASE_URL; export const ADMIN_ID = "admin"; -export const ADMIN_PASSWORD = "admin123"; +export const ADMIN_PASSWORD = "penguin-2026"; /** Log in: the cookie lands in the given request context (page.request is the browser context); returns user. */ export async function login(ctx, userId, password) { diff --git a/packages/web/e2e/skills.spec.mjs b/packages/web/e2e/skills.spec.mjs index b7a97a7..7449eb6 100644 --- a/packages/web/e2e/skills.spec.mjs +++ b/packages/web/e2e/skills.spec.mjs @@ -3,7 +3,7 @@ * - the sidebar nav shows "技能库" ("Skill Library"), and the page renders the library's skill * cards across group sections (a collapsible group header: group name + skill count, * **no icon**; the group name follows the UI language — when the server ships a Chinese - * group name it's "Agent 开发 / 数据分析 / Penguin 开发 / 网页开发 / 软件工程", falling back to + * group name it's "办公效率 / 软件开发 / AI 应用开发 / Agent 调优", falling back to * English by default); cards carry a custom icon (icon.svg sanitized then inlined, not the book * fallback), with metadata showing version and usage count (worded semantically, not a bare * number badge); @@ -35,11 +35,10 @@ const P = "password123"; // falling back to English (both states are asserted). // The group header is a collapsible button (group name + skill count); matched by a substring of its accessible name. const GROUPS = [ - /Agent Development|Agent 开发/, - /Data Analysis|数据分析/, - /Penguin Development|Penguin 开发/, - /Web Development|网页开发/, - /Software Engineering|软件工程/, + /Office Productivity|办公效率/, + /Software Development|软件开发/, + /AI App Development|AI 应用开发/, + /Agent Tuning|Agent 调优/, ]; const SKILLS = [ "agent-creation", diff --git a/packages/web/src/api/endpoints.ts b/packages/web/src/api/endpoints.ts index 3d69fff..44e3d3b 100644 --- a/packages/web/src/api/endpoints.ts +++ b/packages/web/src/api/endpoints.ts @@ -302,7 +302,7 @@ export const getUsage = ( to?: string; groupBy: UsageGroupBy; agentId?: string; - /** Model filters are given as a pair (only takes effect when both provider and modelId are supplied). */ + /** Model filter is always a whole pair — both fields or neither; a model is never referenced by id alone. */ provider?: string; modelId?: string; }, @@ -330,6 +330,10 @@ export const listWorkspaceFiles = (sessionId: string, path: string) => export const workspaceFileUrl = (sessionId: string, path: string, download = false): string => `/api/sessions/${sessionId}/files/content?path=${encodeURIComponent(path)}${download ? "&download=1" : ""}`; +/** Sandboxed top-level preview URL (open an html file in a new tab): real content type under a CSP sandbox, see the server route. */ +export const workspaceFilePreviewUrl = (sessionId: string, path: string): string => + `/api/sessions/${sessionId}/files/content?path=${encodeURIComponent(path)}&preview=1`; + export const uploadWorkspaceFile = (sessionId: string, path: string, dataBase64: string) => apiFetch(`/api/sessions/${sessionId}/files/content`, { method: "PUT", diff --git a/packages/web/src/components/account/change-password-dialog.tsx b/packages/web/src/components/account/change-password-dialog.tsx index c6f571e..7a96644 100644 --- a/packages/web/src/components/account/change-password-dialog.tsx +++ b/packages/web/src/components/account/change-password-dialog.tsx @@ -73,6 +73,7 @@ export function ChangePasswordDialog({ open, onClose }: { open: boolean; onClose value={oldPassword} onChange={(e) => setOldPassword(e.target.value)} autoComplete="current-password" + hint={S.account.oldPasswordHint} autoFocus /> setVisible((v) => !v)} className={`absolute inset-y-0 right-0 flex items-center justify-center text-gray-400 transition-colors duration-150 hover:text-gray-600 dark:hover:text-gray-300 ${ size === "sm" ? "w-8" : "w-10" diff --git a/packages/web/src/components/ui/provider-logo.tsx b/packages/web/src/components/ui/provider-logo.tsx index 3d79c99..58a8dd4 100644 --- a/packages/web/src/components/ui/provider-logo.tsx +++ b/packages/web/src/components/ui/provider-logo.tsx @@ -3,10 +3,12 @@ * light and dark themes). * * Anthropic / OpenAI / Google Gemini / DeepSeek / Moonshot AI / OpenRouter / - * SiliconFlow use each vendor's brand mark (for recognition purposes, not under - * trademark license); Z.AI uses a simplified geometric approximation of its - * branded glyph (not an exact reproduction of the trademark); unknown vendors and - * custom models use a generic cube. All are pure paths, no external image assets. + * SiliconFlow / Qwen Token Plan / Qwen Pay-As-You-Go / Fireworks AI use each + * vendor's brand mark (for recognition purposes, not under trademark license; + * Qwen's official gradient wordmark is flattened to currentColor monochrome); + * Z.AI uses a simplified geometric approximation of its branded glyph (not an + * exact reproduction of the trademark); unknown vendors and custom models use + * a generic cube. All are pure paths, no external image assets. */ import type { ReactNode } from "react"; @@ -18,6 +20,25 @@ interface Glyph { path: ReactNode; } +/** + * The official Qwen emblem (the wordmark's icon, lettering dropped), gradient fills + * flattened to currentColor; coordinates rounded to 2dp from the brand SVG. Shared by the + * Qwen Token Plan and Qwen Pay-As-You-Go groups. + */ +const QWEN_GLYPH: Glyph = { + viewBox: "-0.76 -0.47 28.6 28.6", + path: ( + <> + + + + + + + + ), +}; + const GLYPHS: Record = { anthropic: { path: ( @@ -60,6 +81,19 @@ const GLYPHS: Record = { ), }, + "qwen-token-plan": QWEN_GLYPH, + "qwen-pay-as-you-go": QWEN_GLYPH, + fireworks: { + // The official Fireworks AI burst mark (three strokes of the wordmark's icon). + viewBox: "0 0 638 315", + path: ( + <> + + + + + ), + }, custom: { stroke: true, path: ( diff --git a/packages/web/src/features/agents/agents-page.tsx b/packages/web/src/features/agents/agents-page.tsx index 2342541..9607ff3 100644 --- a/packages/web/src/features/agents/agents-page.tsx +++ b/packages/web/src/features/agents/agents-page.tsx @@ -53,6 +53,9 @@ const CARD_ICONS = { vaultKeys: "M15.5 7.5l3 3L22 7l-3-3M21 2l-9.6 9.6M13 15.5a5.5 5.5 0 1 1-11 0 5.5 5.5 0 0 1 11 0z", /** Schedule count (alarm clock: dial + hands + twin bells, distinct from the plain clock face used for "last modified") */ schedules: "M12 21a7 7 0 1 0 0-14 7 7 0 0 0 0 14zm0-10v3l2 1.5M5 3L2.5 5.5M19 3l2.5 2.5", + /** Installed skill count (open book, same family as the skill library) */ + skills: + "M12 6.5C10.5 5 8 4.5 4 5v12c4-.5 6.5 0 8 1.5 1.5-1.5 4-2 8-1.5V5c-4-.5-6.5 0-8 1.5zm0 0V18", /** Usage (bar chart, same as sidebar "Usage Center") */ usage: "M4 20V10m6 10V4m6 16v-7m4 7H2", /** Traces (eye line icon: observability; follows text color, no fill) */ @@ -245,6 +248,13 @@ export function AgentsPage() { {a.scheduleCount} + + + {a.skillCount} + - JSON.stringify([ref.provider ?? null, ref.modelId]); +const modelOptionValue = (ref: ModelRefDto): string => JSON.stringify([ref.provider, ref.modelId]); -const parseModelOption = (v: string): ScheduleModelRef | null => { +const parseModelOption = (v: string): ModelRefDto | null => { if (!v) return null; - const [provider, modelId] = JSON.parse(v) as [string | null, string]; - return provider === null ? { modelId } : { provider, modelId }; + const [provider, modelId] = JSON.parse(v) as [string, string]; + return { provider, modelId }; }; /** Display label for a model option: upstream id + provider name (shown side by side, not a composite id). */ -const modelOptionLabel = (ref: ScheduleModelRef): string => - ref.provider === undefined - ? ref.modelId - : `${ref.modelId} · ${providerInfo(ref.provider)?.label ?? ref.provider}`; +const modelOptionLabel = (ref: ModelRefDto): string => + `${ref.modelId} · ${providerInfo(ref.provider)?.label ?? ref.provider}`; + +/** + * Stored schedule fields → a model reference. A reference is always the complete + * (provider, model_id) pair, since provider is never inferred: the DTO types the two + * fields independently, so this guard is what keeps the form and the upsert body from + * ever assembling half a reference. A file that sets only one half is rejected by the + * server when parsed (it surfaces under invalidFiles, never as a listed row), so in + * practice this returns null only when the schedule uses the Project's default model. + */ +const itemModelRef = (item: Pick): ModelRefDto | null => + item.modelId && item.provider ? { provider: item.provider, modelId: item.modelId } : null; /** Modal form state (shared by create/edit): non-null editing means editing that task (name locked). */ interface FormState { @@ -94,8 +93,8 @@ interface FormState { target: "new" | "session"; sessionId: string; workspace: string; - /** Model for the new-Session mode (null = Project default, modelId/provider omitted). */ - model: ScheduleModelRef | null; + /** Model for the new-Session mode (null = Project default, provider and modelId both omitted). */ + model: ModelRefDto | null; } const EMPTY_FORM: FormState = { @@ -191,10 +190,7 @@ export function SchedulesTab({ agentId }: { agentId: string }) { ? { workspace: form.workspace.trim() } : {}), ...(form.target === "new" && form.model - ? { - modelId: form.model.modelId, - ...(form.model.provider !== undefined ? { provider: form.model.provider } : {}), - } + ? { modelId: form.model.modelId, provider: form.model.provider } : {}), }; setBusy(true); @@ -218,6 +214,7 @@ export function SchedulesTab({ agentId }: { agentId: string }) { setBusy(true); setError(null); setNotice(null); + const model = itemModelRef(item); try { await api.updateSchedule(projectId, agentId, item.name, { prompt: item.prompt, @@ -227,9 +224,8 @@ export function SchedulesTab({ agentId }: { agentId: string }) { ...(item.endAt !== undefined ? { endAt: item.endAt } : {}), ...(item.sessionId !== undefined ? { sessionId: item.sessionId } : {}), ...(item.workspace !== undefined ? { workspace: item.workspace } : {}), - // Model reference is resent as a pair (a hand-written file with provider omitted stays omitted). - ...(item.modelId !== undefined ? { modelId: item.modelId } : {}), - ...(item.provider !== undefined ? { provider: item.provider } : {}), + // Model reference is resent as a whole pair or not at all — never half of one. + ...(model ? { modelId: model.modelId, provider: model.provider } : {}), }); setNotice(S.common.saved); await load(); @@ -240,7 +236,7 @@ export function SchedulesTab({ agentId }: { agentId: string }) { } }; - /** Edit: prefill this row into the modal form (submits via PUT; model reference prefilled as a pair, provider not auto-filled when absent). */ + /** Edit: prefill this row into the modal form (submits via PUT; the model reference is prefilled only as a complete pair). */ const startEdit = (item: ScheduleItem) => { openForm({ editing: item.name, @@ -253,13 +249,7 @@ export function SchedulesTab({ agentId }: { agentId: string }) { target: item.sessionId ? "session" : "new", sessionId: item.sessionId ?? "", workspace: item.workspace ?? "", - model: - item.modelId !== undefined - ? { - modelId: item.modelId, - ...(item.provider !== undefined ? { provider: item.provider } : {}), - } - : null, + model: itemModelRef(item), }); }; @@ -494,9 +484,9 @@ export function SchedulesTab({ agentId }: { agentId: string }) { onChange={(e) => set({ model: parseModelOption(e.target.value) })} > - {/* When the prefilled reference is no longer in the model config (including a - hand-written file missing only provider), add an extra option so it isn't - displayed as "Project default". */} + {/* When the prefilled pair is no longer in the model config (the entry was + renamed or deleted), add an extra option so it isn't displayed as + "Project default". */} {form.model && !models.some( (m) => diff --git a/packages/web/src/features/chat/chat-input.tsx b/packages/web/src/features/chat/chat-input.tsx index bcfaa26..2364ac0 100644 --- a/packages/web/src/features/chat/chat-input.tsx +++ b/packages/web/src/features/chat/chat-input.tsx @@ -12,7 +12,9 @@ * logo + name display; * `/` opens the slash command menu (`/compact` compresses context, replacing the button; each * installed skill gets its own entry; pressing Enter on `/` toggles that skill's - * selection and clears the input, without sending); + * selection without sending). Matching is positional like `@`: a slash opens the menu from any + * caret position, running a command removes just that token, and Escape only dismisses the menu — + * the rest of the draft is never touched; * `@` opens the agent selection menu; once picked it becomes a fixed highlighted target chip * above the text body (only one allowed, picking again replaces it; removed via backspace or the * x button); only a leading `@` at the start of the text counts — typing or pasting text starting @@ -54,8 +56,9 @@ import { GlyphIcon } from "../../components/ui/glyph-icon"; import { SkillIcon } from "../skills/skill-icon-view"; import { ZoomableImage } from "../../components/ui/image-zoom"; import { ProviderLogo } from "../../components/ui/provider-logo"; -import { matchesQuery, sameModelRef } from "../models/model-grouping"; +import { matchesQuery, orderModelsLikeLibrary, sameModelRef } from "../models/model-grouping"; import { filterAgents, matchMention, splitLeadingMention } from "./agent-mentions"; +import { matchSlash, removeSlashToken } from "./slash-token"; import { BOOK_ICON, buildSkillsMessage, @@ -231,7 +234,9 @@ function ModelSelect({ const current = models.find((m) => sameModelRef(m, value)); // Display rule matches the model page's card: display name, or falls back to the upstream id (grouping is already conveyed by the provider logo). const label = current ? modelLabel(current) : (value?.modelId ?? "…"); - const filtered = models.filter((m) => matchesQuery(m, query)); + // Dropdown order mirrors the model library page: provider groups in MODEL_PROVIDERS order + // (user-defined groups after, custom last), in-group order preserved. + const filtered = orderModelsLikeLibrary(models).filter((m) => matchesQuery(m, query)); return ( {/* When the card is narrower than @md, only the icon + badge remain (title shows the full name). */} {S.chat.skillsSelect} - {/* Selected-count badge: chips no longer exist, so this is the only visible rendering of the selected state. */} + {/* Selected-count badge (the chip row above the input mirrors the selection too). */} {selected.length > 0 && ( {selected.length} @@ -616,9 +623,17 @@ export function ChatInput({ }) { const { locale } = useLocale(); const [text, setText] = useState(initialText ?? ""); + /** Live text mirror for slash-command run() closures (the commands memo deliberately doesn't depend on text). */ + const textRef = useRef(text); + textRef.current = text; const [images, setImages] = useState([]); const [busy, setBusy] = useState(false); const [slashIndex, setSlashIndex] = useState(0); + // Slash token start where Escape closed the menu (mirrors mentionDismissed: the menu stays shut for that one token). + const [slashDismissed, setSlashDismissed] = useState(null); + // Anchor for the popups that open upward, and the room actually available above them. + const anchorRef = useRef(null); + const [upwardMaxH, setUpwardMaxH] = useState(); // @ handoff target (chip, fixed at the front of the input); only one allowed, picking again replaces it directly. const [target, setTarget] = useState(null); // Selected skills (dropdown checklist, multi-select): initial value comes from draft restore (quick-invoke pre-selection), cleared on successful send. @@ -654,11 +669,15 @@ export function ChatInput({ [selectedSkills, onSkillsChange], ); + /** The slash token currently under the caret (kept in a ref so command run() closures always remove the live token). */ + const slashMatchRef = useRef>(null); const commands = useMemo(() => { - /** Clears the input after a slash command runs (setText is async: wait for the DOM value to update before measuring height). */ + /** Removes just the slash token after a command runs (the rest of the text stays; setText is async: wait for the DOM value to update before measuring height). */ const clearInput = () => { - setText(""); - onTextChange?.(""); + const match = slashMatchRef.current; + const next = match ? removeSlashToken(textRef.current, match) : ""; + setText(next); + onTextChange?.(next); requestAnimationFrame(autoGrow); }; return [ @@ -681,10 +700,14 @@ export function ChatInput({ })), ]; }, [onCompact, onTextChange, skills, locale, toggleSkill]); - const trimmed = text.trim(); + // Positional matching (like @ mentions): a slash opens the menu from any caret position; + // running a command removes just the token, leaving the rest of the text intact. Doesn't + // reopen after Escape until the caret sits on a different token. + const slashTok = !running && !compacting ? matchSlash(text, caret) : null; + slashMatchRef.current = slashTok; const slashMatches = - text.trimStart().startsWith("/") && !running && !compacting - ? commands.filter((c) => c.cmd.startsWith(trimmed) || trimmed === "/") + slashTok && slashTok.start !== slashDismissed + ? commands.filter((c) => c.cmd.startsWith(`/${slashTok.query}`)) : []; const slashOpen = slashMatches.length > 0; const activeSlash = slashMatches[Math.min(slashIndex, slashMatches.length - 1)]; @@ -696,6 +719,31 @@ export function ChatInput({ const mentionOpen = mentionMatches.length > 0; const activeMention = mentionMatches[Math.min(mentionIndex, mentionMatches.length - 1)]; + // Both menus above are drawn upward (`bottom-full`) from the composer, so their ceiling is + // whatever ancestor clips overflow — on the draft page that's the centered scroll area, whose + // top edge sits well below the viewport's. A static `40vh` cap can't know that distance and + // clipped the first rows on shorter windows, so measure the real gap when a menu opens. + useEffect(() => { + if (!slashOpen && !mentionOpen) return; + const measure = () => { + const el = anchorRef.current; + if (!el) return; + let ceiling = 0; + for (let p = el.parentElement; p; p = p.parentElement) { + if (getComputedStyle(p).overflowY !== "visible") { + ceiling = p.getBoundingClientRect().top; + break; + } + } + // Less the menu's own 6px offset from the composer, plus a little breathing room. + const room = el.getBoundingClientRect().top - ceiling - 14; + setUpwardMaxH(Math.max(96, Math.min(320, Math.round(room)))); + }; + measure(); + window.addEventListener("resize", measure); + return () => window.removeEventListener("resize", measure); + }, [slashOpen, mentionOpen]); + /** Auto-grow the textarea (caps at roughly 6 lines, scrolls internally beyond that). */ const autoGrow = () => { const el = textareaRef.current; @@ -829,10 +877,10 @@ export function ChatInput({ return; } if (e.key === "Escape") { - setText(""); - onTextChange?.(""); - // setText is async: wait for the DOM value to update before measuring height (measuring synchronously would read the pre-clear content). - requestAnimationFrame(autoGrow); + // Only closes the popup, doesn't clear the input: with positional matching the `/token` + // is part of the text body like any other word, and wiping a controlled textarea is not + // undoable with Ctrl+Z. Reopens if the user keeps typing on another token. + setSlashDismissed(slashTok?.start ?? null); return; } } @@ -853,7 +901,7 @@ export function ChatInput({ return; } if (e.key === "Escape") { - // Only closes the popup, doesn't clear the input (unlike slash: the @ prefix is part of the text body), reopens if the user keeps typing. + // Only closes the popup, doesn't clear the input (the `@token` is part of the text body), reopens if the user keeps typing. setMentionDismissed(mention?.start ?? null); return; } @@ -909,14 +957,21 @@ export function ChatInput({ }; return ( -
- {/* Slash command menu (triggered by typing /; /compact plus one entry per installed skill) */} +
+ {/* Slash command menu (triggered by typing /; /compact plus one entry per installed skill). + Height is capped to the room measured above the composer (see upwardMaxH) with internal + scrolling, so a long skill list never pushes the menu's top edge out of view; the active + row keeps itself scrolled into view. */} {slashOpen && ( -
+
{slashMatches.map((c, i) => ( - + @{target.agentId} + + + )} + {selectedSkills.map((name) => { + const meta = skills.find((sk) => sk.name === name); + return ( + + + {name} + + + ); + })}
)} @@ -1039,9 +1126,14 @@ export function ChatInput({ setCaret(caretNow); setSlashIndex(0); setMentionIndex(0); - // Closing via Escape only persists for "the same @ mention": continuing to type - // within that mention won't reopen the menu; it re-opens once the cursor is no - // longer on that mention (deleted, moved away, or replaced by a new one). + // Closing via Escape only persists for "the same token": continuing to type within + // that slash command / mention won't reopen the menu; it re-opens once the cursor is + // no longer on that token (deleted, moved away, or replaced by a new one). + setSlashDismissed((d) => { + if (d === null) return null; + const m = matchSlash(value, caretNow); + return m && m.start === d ? d : null; + }); setMentionDismissed((d) => { if (d === null) return null; const m = matchMention(value, caretNow); diff --git a/packages/web/src/features/chat/draft-view.tsx b/packages/web/src/features/chat/draft-view.tsx index fb5c6df..b7e3dd7 100644 --- a/packages/web/src/features/chat/draft-view.tsx +++ b/packages/web/src/features/chat/draft-view.tsx @@ -45,6 +45,7 @@ import { Dropdown } from "../../components/ui/dropdown"; import { PenguinLogo } from "../../components/ui/penguin-logo"; import { toastError } from "../../components/ui/toast"; import { ChatInput } from "./chat-input"; +import { buildSkillsMessage } from "./skill-use"; import { clearDraft, draftKey, loadDraft, saveDraft } from "./draft-cache"; import type { DraftCache } from "./draft-cache"; import { handoffMessage } from "./agent-mentions"; @@ -53,6 +54,18 @@ import { sameModelRef } from "../models/model-grouping"; /** Coalescing window for writing body text to the cache: keystrokes are frequent, so a short batch accumulates before persisting (option changes are still written immediately). */ const DRAFT_SAVE_DEBOUNCE_MS = 300; +/** + * Example tasks on the draft screen, in display order (game card first, LoL music player, + * then the RAG build). Copy lives in S.chat.exampleTasks[id]; skills are pinned via a + * `` block — only those the selected Agent actually has installed are included, + * so the block never references a skill the agent can't read. + */ +const EXAMPLE_TASKS: { id: "game" | "lol" | "rag"; skills: string[] }[] = [ + { id: "game", skills: ["web-design"] }, + { id: "lol", skills: ["web-design"] }, + { id: "rag", skills: ["penguin-sdk", "web-design"] }, +]; + export function DraftView({ projectId, models, @@ -146,8 +159,11 @@ export function DraftView({ // mount render would trigger ChatInput's pruning effect and wrongly clear the // quick-invoke preselection. const [agentSkills, setAgentSkills] = useState([]); + /** Whether the skills fetch for the current Agent has settled — the example task waits for it so its `` pinning doesn't silently depend on network timing. */ + const [skillsLoaded, setSkillsLoaded] = useState(false); useEffect(() => { setAgentSkills((prev) => (prev.length > 0 ? [] : prev)); + setSkillsLoaded(false); if (!agentId) return; let cancelled = false; api @@ -155,7 +171,10 @@ export function DraftView({ .then((res) => { if (!cancelled) setAgentSkills(res.skills); }) - .catch(() => undefined); + .catch(() => undefined) + .finally(() => { + if (!cancelled) setSkillsLoaded(true); + }); return () => { cancelled = true; }; @@ -244,12 +263,23 @@ export function DraftView({ setCurrentAgentId(a.agentId); }; + // One in-flight guard shared by every send entry point (composer send / example task / + // @-handoff): a second submission while one is running would create a second Session with + // its own first task and a racing navigation. The ref is the synchronous guard; the state + // drives disabled styling on the example button (the composer has its own busy state). + const sendingRef = useRef(false); + const [sending, setSending] = useState(false); + // First message sent: only now is the Session created (Agent / Workspace / Model / approval // mode are all locked in together), then the route jumps once sent; returns false on any - // failure, so the input area keeps the draft and can resend. + // failure, so the input area keeps the draft and can resend. `keepDraft` is set by sends + // that did not consume the composer text (the example task), so a typed-but-unsent draft + // survives the navigation instead of being silently discarded. const onSend = useCallback( - async (input: TaskInputPart[]): Promise => { - if (!agentId) return false; + async (input: TaskInputPart[], keepDraft = false): Promise => { + if (!agentId || sendingRef.current) return false; + sendingRef.current = true; + setSending(true); let createdId: string | null = null; try { const body: SessionCreateRequest = { approvalMode }; @@ -263,7 +293,7 @@ export function DraftView({ createdId = created.session.sessionId; const res = await api.postTask(createdId, { input }); add(created.session); - discardDraft(); + if (!keepDraft) discardDraft(); navigate(`/chat/${res.sessionId}`, { replace: true }); return true; } catch (e) { @@ -273,18 +303,45 @@ export function DraftView({ if (createdId) void api.deleteSession(createdId).catch(() => undefined); toastError(apiErrorText(e, modelRef ? { modelId: modelRef.modelId } : {})); return false; + } finally { + sendingRef.current = false; + setSending(false); } }, [projectId, agentId, approvalMode, modelRef, workspace, add, discardDraft, navigate], ); + // Example tasks: one click submits the canned prompt exactly like a hand-typed send (the + // busy id drives the clicked card's spinner; the shared in-flight guard and all failure + // handling live in onSend). keepDraft: an example never consumes the composer text, so a + // typed-but-unsent draft must survive. The selected model / Workspace / approval mode apply as-is. + const [exampleBusy, setExampleBusy] = useState<"game" | "lol" | "rag" | null>(null); + const runExample = useCallback( + async (task: (typeof EXAMPLE_TASKS)[number]) => { + if (exampleBusy !== null) return; + setExampleBusy(task.id); + try { + const names = task.skills.filter((n) => agentSkills.some((s) => s.name === n)); + await onSend( + [{ type: "text", text: buildSkillsMessage(names, S.chat.exampleTasks[task.id].prompt) }], + true, + ); + } finally { + setExampleBusy(null); + } + }, + [exampleBusy, agentSkills, onSend], + ); + // @ handoff: opens a new chat for the @-mentioned agent (approval mode carries over from the // draft's current value; model/Workspace use the creation defaults), first input = // source block + the text and images with the @ mention stripped. const selectedAgent = agents.find((a) => a.agentId === agentId) ?? null; const onHandoff = useCallback( async (target: AgentSummary, input: TaskInputPart[]): Promise => { - if (!selectedAgent) return false; + if (!selectedAgent || sendingRef.current) return false; + sendingRef.current = true; + setSending(true); const origin: TaskInputPart = { type: "text", text: handoffMessage({ @@ -308,6 +365,9 @@ export function DraftView({ apiErrorText(e, models?.defaultModel ? { modelId: models.defaultModel.modelId } : {}), ); return false; + } finally { + sendingRef.current = false; + setSending(false); } }, [projectId, selectedAgent, approvalMode, add, discardDraft, navigate, models], @@ -319,15 +379,21 @@ export function DraftView({ const vision = modelInfo?.vision !== false; return ( -
+
{/* - * Vertical layout: two symmetric flex-1 spaces above and below, so **the input card lands - * exactly on the viewport's centerline**; the brand area sits at the bottom of the upper - * space (right above the input card). When the viewport is too short, the upper space - * shrinks to the brand area's height, the lower space compresses to padding, and the whole - * page falls back to natural document-flow scrolling (without clipping the top). + * Vertical layout: everything visible — brand, input card, ownership pills, example tasks — + * lives in ONE block between two empty flex-1 spacers, so the block is centred and the free + * space above and below it is exactly equal. The brand deliberately sits inside that block + * rather than in the upper spacer: keeping it in the spacer made the upper gap shorter than + * the lower one by the brand's own height, which pushed the card up the viewport and left + * the slash menu — it opens upward, `bottom-full` — too little room, so it clipped against + * the top of this scroll container. When the viewport is too short the spacers collapse to + * nothing, the container's own py-6 keeps the content off the edges, and the page falls back + * to natural scrolling. */} -
+
+ +
{/* Large brand logo + brand name + subtitle (e2e tests identify the draft page by this heading). The asset is square-cropped and the graphic already has a bit of built-in padding, so a small margin is enough to sit visually close to the title. */} @@ -338,9 +404,7 @@ export function DraftView({

{S.chat.draftSubtitle}

-
-
+ + {/* Example tasks: one-click canned builds showing off the one-sentence → app flow, + stacked vertically in display order on every viewport. Disabled until + agents/models/skills are resolved (onSend would silently no-op without an Agent); + hover only darkens the border, per the card convention. */} +
+ {EXAMPLE_TASKS.map((task) => { + const copy = S.chat.exampleTasks[task.id]; + return ( + + ); + })} +
- {/* Lower symmetric space (compresses to padding at minimum height) */} -
+ {/* Lower symmetric space — empty, so it matches the upper one exactly */} +
); } diff --git a/packages/web/src/features/chat/live-duration.tsx b/packages/web/src/features/chat/live-duration.tsx index 116b7fc..acd397b 100644 --- a/packages/web/src/features/chat/live-duration.tsx +++ b/packages/web/src/features/chat/live-duration.tsx @@ -1,12 +1,13 @@ /** - * Live duration (inline display on running thinking/tool cards): ticks every second. sinceMs + * Live duration (inline display on running thinking/tool cards): ticks every second, showing + * whole seconds only — decimals appear only on the settled value once the item finishes. sinceMs * comes from the server-side message timestamp and may drift from the local clock; negative * values are shown as 0; a pulsing ellipsis is shown when missing. * `offsetMs` is the already-settled duration of a prior segment (e.g. a tool call's argument * generation phase), added on top of the live segment as it ticks. */ import { useEffect, useState } from "react"; -import { humanizeDuration } from "../../lib/format"; +import { humanizeDurationLive } from "../../lib/format"; export function LiveDuration({ sinceMs, offsetMs = 0 }: { sinceMs?: number; offsetMs?: number }) { const [now, setNow] = useState(() => Date.now()); @@ -15,5 +16,5 @@ export function LiveDuration({ sinceMs, offsetMs = 0 }: { sinceMs?: number; offs return () => clearInterval(id); }, []); if (sinceMs === undefined) return …; - return <>{humanizeDuration(Math.max(0, offsetMs) + Math.max(0, now - sinceMs))}; + return <>{humanizeDurationLive(Math.max(0, offsetMs) + Math.max(0, now - sinceMs))}; } diff --git a/packages/web/src/features/chat/md.tsx b/packages/web/src/features/chat/md.tsx new file mode 100644 index 0000000..08de409 --- /dev/null +++ b/packages/web/src/features/chat/md.tsx @@ -0,0 +1,76 @@ +/** + * Markdown body memoized by text identity. The stream model mutates chat items in place and + * only reassigns `text` when a delta arrives, so every settled message keeps the same string + * instance across the per-frame version bumps — the shallow prop compare then skips its entire + * micromark → mdast → React re-parse, and only the actively-streaming message re-renders. + * Without this, each animation frame during a stream re-parsed the WHOLE transcript (O(n²) + * over a long reply), which visibly froze the UI while large code blocks streamed in. + * + * Fenced code blocks render through CodeBlock (language chrome + copy button + Shiki + * highlight). `streaming` disables highlighting while deltas are still arriving — + * re-tokenizing a growing block every frame is O(n²) main-thread cost — and the settle + * re-render (new string instance, streaming=false) highlights each block exactly once. + * Inline code keeps the default rendering (`.md-body code` styling). + */ +import { isValidElement, memo } from "react"; +import type { ReactNode } from "react"; +import ReactMarkdown from "react-markdown"; +import type { Components } from "react-markdown"; +import remarkGfm from "remark-gfm"; +import { CodeBlock } from "./code-block"; + +/** Flatten a react-markdown code element's children to plain text (string or string array in practice). */ +function codeText(children: unknown): string { + if (typeof children === "string") return children; + if (Array.isArray(children)) return children.filter((c) => typeof c === "string").join(""); + return ""; +} + +/** Fenced-block adapter: unwraps the
 pair react-markdown emits into CodeBlock. */
+function MdPre({ children, streaming }: { children?: ReactNode; streaming: boolean }) {
+  if (isValidElement(children)) {
+    const props = children.props as { className?: string; children?: unknown };
+    const language = /language-([\w+-]+)/.exec(props.className ?? "")?.[1] ?? "";
+    return (
+      
+    );
+  }
+  return 
{children}
; +} + +/** + * The two `components` maps, built once at module scope instead of inline per render. + * react-markdown uses `components.pre` as the element **type**, so a fresh arrow each render is + * a new type on every commit: React unmounts and remounts every code block, dropping the user's + * text selection and resetting each block's Copy-button state — ~8 times a second while a reply + * streams, including for blocks that closed long ago. `streaming` is the only thing the adapter + * closes over, so one frozen map per value is enough; the single flip between them happens on + * the settle render, which re-parses the message anyway. + */ +const STREAMING_COMPONENTS: Components = { + pre: (props) => {props.children}, +}; +const SETTLED_COMPONENTS: Components = { + pre: (props) => {props.children}, +}; + +export const Md = memo(function Md({ + text, + streaming = false, +}: { + text: string; + streaming?: boolean; +}) { + return ( + + {text} + + ); +}); diff --git a/packages/web/src/features/chat/message-item.tsx b/packages/web/src/features/chat/message-item.tsx index bc21d9a..ce83b05 100644 --- a/packages/web/src/features/chat/message-item.tsx +++ b/packages/web/src/features/chat/message-item.tsx @@ -5,14 +5,13 @@ * stats lines. Items have a light entrance animation. */ import { useState } from "react"; -import ReactMarkdown from "react-markdown"; -import remarkGfm from "remark-gfm"; import { S } from "../../lib/strings"; import { useLocale } from "../../state/locale"; import { formatMessageTime } from "../../lib/format"; import { STAT_ICONS } from "../../lib/stat-icons"; import { splitImageAttachments } from "../../lib/attachments"; import type { ChatItem } from "../../lib/omni/stream-model"; +import { Md } from "./md"; import { GlyphIcon } from "../../components/ui/glyph-icon"; import { ZoomableImage } from "../../components/ui/image-zoom"; import { MessageFilesCard } from "./message-files-card"; @@ -158,8 +157,8 @@ export function MessageItem({ item, ctx }: { item: ChatItem; ctx: StreamRenderCo // which is more useful than copying segment by segment. return (
- {/* Re-renders the accumulated text directly while streaming (a key point of the contract implementation). */} - {item.text} + {/* Re-renders the accumulated text directly while streaming (a key point of the contract implementation); memoized so settled messages skip the re-parse, and code blocks highlight once on settle (see md.tsx). */} + {item.streaming && ▌} {item.stopReason && item.stopReason !== "completed" && ( [{item.stopReason}] diff --git a/packages/web/src/features/chat/message-stream.tsx b/packages/web/src/features/chat/message-stream.tsx index 900c933..551f1bf 100644 --- a/packages/web/src/features/chat/message-stream.tsx +++ b/packages/web/src/features/chat/message-stream.tsx @@ -5,7 +5,7 @@ * StreamRenderContext threads the pending-approval map and approval callback down to tool * cards at any nesting depth. */ -import { useEffect, useRef } from "react"; +import { useEffect, useLayoutEffect, useRef, useState } from "react"; import type { ReactNode } from "react"; import { S } from "../../lib/strings"; import type { ChatItem } from "../../lib/omni/stream-model"; @@ -144,45 +144,175 @@ export function MessageStream({ // An upward-swipe intent immediately exits auto-follow; scrolling back near the bottom resumes it — see stream-follow.ts (#75) for the exact rule. const followRef = useRef(null); const follow = (followRef.current ??= createStreamFollow()); + // Back-to-bottom button visibility (React state — the follow object mutates outside React): + // shown only when follow is off AND there is actually content below the fold. The second + // condition matters: a wheel-up flick over a list that already fits the viewport exits follow + // (stream-follow.ts keeps that rule deliberately) but must not surface a "jump to latest" + // button with nothing to jump to. Synced after every event that can change either input + // (wheel-up at the very top fires no scroll event, so syncing on scroll alone is not enough) + // and in the version effect (content growing while unstuck fires no scroll event either). + const [showJump, setShowJump] = useState(false); + /** + * Animated-return phase (back-to-bottom clicked, glide in flight). The glide is a + * self-driven rAF loop rather than native scrollTo({behavior:"smooth"}) for two + * live-tested reasons: (1) re-aiming a native smooth scroll on every stream commit resets + * its easing, capping descent below a fast stream's growth rate — the return then never + * arrives; (2) an in-flight native smooth scroll cannot be interrupted, so a wheel-up + * cancel would keep dragging the user downward for hundreds of ms. The rAF loop chases the + * LIVE bottom each frame with a velocity floor above any realistic growth rate, and + * cancelling it stops the descent within a frame. While returning, the button stays hidden + * and the stick snap must not fire (an instant snap on the next commit would kill the glide). + */ + const returningRef = useRef(false); + const returnRafRef = useRef(null); + const cancelReturn = () => { + if (returnRafRef.current !== null) { + cancelAnimationFrame(returnRafRef.current); + returnRafRef.current = null; + } + returningRef.current = false; + }; + useEffect(() => () => cancelReturn(), []); + /** Local previous scrollTop, only for detecting the user fighting the return animation upward (the follow object keeps its own). */ + const lastTopRef = useRef(null); + const syncJump = () => { + const el = scrollRef.current; + setShowJump( + !returningRef.current && + !follow.stick && + el !== null && + el.scrollHeight - el.scrollTop - el.clientHeight > 1, + ); + }; const onScroll = () => { const el = scrollRef.current; if (!el) return; + const prevTop = lastTopRef.current; + lastTopRef.current = el.scrollTop; + // Scrollbar-drag / keyboard scroll upward during the return cancels it (wheel/touch cancel in their own handlers). + if (returningRef.current && prevTop !== null && el.scrollTop < prevTop - 1) { + cancelReturn(); + } follow.scrolled({ scrollTop: el.scrollTop, scrollHeight: el.scrollHeight, clientHeight: el.clientHeight, }); + syncJump(); }; - useEffect(() => { + // Layout effect (not useEffect): the stick-to-bottom snap must land before paint, otherwise + // fast streams show the bottom edge "catching up" by the growth of each commit. Suppressed + // during the animated return — the glide owns the scroll position until it arrives. + useLayoutEffect(() => { const el = scrollRef.current; - if (el && follow.stick) el.scrollTop = el.scrollHeight; + if (el && follow.stick && !returningRef.current) el.scrollTop = el.scrollHeight; + syncJump(); + // syncJump is recreated per render; the effect intentionally keys on stream growth only. + // eslint-disable-next-line react-hooks/exhaustive-deps }, [version, follow]); + /** Back-to-bottom: glide down to the live bottom (reduced motion gets an instant jump); follow re-engages on arrival. */ + const jumpToLatest = () => { + const el = scrollRef.current; + if (!el) return; + if (window.matchMedia("(prefers-reduced-motion: reduce)").matches) { + follow.resume(); + el.scrollTop = el.scrollHeight; + syncJump(); + return; + } + cancelReturn(); + returningRef.current = true; + // Far away: teleport to three viewports above the bottom first, then glide the rest — + // bounds the animation to well under a second regardless of how far the user scrolled. + const far = el.scrollHeight - el.clientHeight - el.scrollTop; + if (far > 3 * el.clientHeight) el.scrollTop = el.scrollHeight - 4 * el.clientHeight; + let last = performance.now(); + const stepFrame = (now: number) => { + returnRafRef.current = null; + const live = scrollRef.current; + if (!live || !returningRef.current) return; + const dt = Math.min(0.05, Math.max(0.001, (now - last) / 1000)); + last = now; + const remaining = live.scrollHeight - live.clientHeight - live.scrollTop; + if (remaining <= 1) { + returningRef.current = false; + follow.resume(); + live.scrollTop = live.scrollHeight; + syncJump(); + return; + } + // Proportional ease-out toward the LIVE bottom, with a time-based velocity floor + // (3200px/s) that outruns any realistic streaming growth so the glide always lands. + const step = Math.max(remaining * Math.min(1, dt * 10), 3200 * dt); + live.scrollTop += Math.min(step, remaining); + returnRafRef.current = requestAnimationFrame(stepFrame); + }; + returnRafRef.current = requestAnimationFrame(stepFrame); + syncJump(); + }; + return ( -
follow.wheel(e.deltaY)} - onTouchStart={(e) => { - const t = e.touches[0]; - if (t) follow.touchStart(t.clientY); - }} - onTouchMove={(e) => { - const t = e.touches[0]; - if (t) follow.touchMove(t.clientY); - }} - onTouchEnd={() => follow.touchEnd()} - className="anim-fade h-full overflow-y-auto px-4 py-4 md:px-6" - > -
- {items.length === 0 ? ( - - ) : ( - - )} +
+
{ + follow.wheel(e.deltaY); + if (e.deltaY < 0) cancelReturn(); // user takes over: the glide stops within a frame + syncJump(); + }} + onTouchStart={(e) => { + const t = e.touches[0]; + if (t) follow.touchStart(t.clientY); + }} + onTouchMove={(e) => { + const t = e.touches[0]; + if (t) follow.touchMove(t.clientY); + if (returningRef.current) cancelReturn(); // touching the list mid-return = taking control + syncJump(); + }} + onTouchEnd={() => follow.touchEnd()} + className="anim-fade h-full overflow-y-auto px-4 py-4 md:px-6" + > +
+ {items.length === 0 ? ( + + ) : ( + + )} +
+ {/* Back-to-bottom (shows once the user scrolls away from content below the fold): floats + just above the composer; clicking returns to the bottom and re-enters follow, so the + view keeps tracking the live stream. */} + {showJump && ( + + )}
); } diff --git a/packages/web/src/features/chat/slash-token.ts b/packages/web/src/features/chat/slash-token.ts new file mode 100644 index 0000000..5ef085b --- /dev/null +++ b/packages/web/src/features/chat/slash-token.ts @@ -0,0 +1,42 @@ +/** + * Positional slash-command matching for the chat input (pure logic, unit-tested): + * like @ mentions, a `/` opens the command menu from ANY caret position — it must sit at + * the start of the text or be preceded by whitespace (so URLs and paths like `a/b` never + * trigger it), with only command characters between the `/` and the caret. Running a + * command removes just the `start..end` token, leaving the rest of the text intact. + */ + +/** Command characters allowed between `/` and the caret (command names and skill names: letters, digits, underscore, hyphen). */ +const CMD_PREFIX = /^[\w-]*$/; + +/** The slash token currently being typed: `start` is the index of `/`, `query` is the text between `/` and the caret (no leading slash), `end` extends over the same token to the right of the caret. */ +export interface SlashMatch { + start: number; + end: number; + query: string; +} + +/** Finds the slash command currently being typed at the caret; returns null if none. */ +export function matchSlash(text: string, caret: number): SlashMatch | null { + const before = text.slice(0, caret); + const at = before.lastIndexOf("/"); + if (at < 0) return null; + if (at > 0 && !/\s/.test(before[at - 1]!)) return null; + const query = before.slice(at + 1); + if (!CMD_PREFIX.test(query)) return null; + const rest = /^[\w-]*/.exec(text.slice(caret))![0]; + return { start: at, end: caret + rest.length, query }; +} + +/** Removes the matched token from the text (what a slash command's run leaves behind); collapses a doubled space at the seam. */ +export function removeSlashToken(text: string, match: SlashMatch): string { + const before = text.slice(0, match.start); + const after = text.slice(match.end); + // Only the seam is touched: the token used to separate its two neighbours, so dropping it would + // leave `a b` — one of those two spaces goes. The rest of the draft stays byte-for-byte (a + // global collapse would silently reflow pasted code, YAML, aligned tables and indentation). + const joined = + before.endsWith(" ") && after.startsWith(" ") ? before + after.slice(1) : before + after; + // The token was the entire content (possibly padded with whitespace): leave a truly empty input. + return /^\s*$/.test(joined) ? "" : joined; +} diff --git a/packages/web/src/features/chat/stream-follow.ts b/packages/web/src/features/chat/stream-follow.ts index 4d0b168..38a44d4 100644 --- a/packages/web/src/features/chat/stream-follow.ts +++ b/packages/web/src/features/chat/stream-follow.ts @@ -35,6 +35,8 @@ export interface StreamFollow { touchEnd(): void; /** scroll event (user scrolling and programmatic stick-to-bottom share this path): moving up exits; otherwise nearing the bottom resumes; the first event initializes from position. */ scrolled(m: ScrollMetrics): void; + /** Explicit re-entry (the back-to-bottom button): resumes follow immediately — the caller scrolls to the bottom right after, and the resulting scroll event sees a bottom position, keeping it stuck. */ + resume(): void; } export function createStreamFollow(): StreamFollow { @@ -72,5 +74,11 @@ export function createStreamFollow(): StreamFollow { } if (dist < 80) stick = true; }, + resume() { + stick = true; + // Forget the last position: the caller jumps to the bottom right after, and that large + // downward scroll must not be judged against a stale historical scrollTop. + lastTop = null; + }, }; } diff --git a/packages/web/src/features/chat/thinking-block.tsx b/packages/web/src/features/chat/thinking-block.tsx index 63102ac..91c3955 100644 --- a/packages/web/src/features/chat/thinking-block.tsx +++ b/packages/web/src/features/chat/thinking-block.tsx @@ -4,8 +4,6 @@ * style and set of running-state icons (in progress / done / failed) as tool cards. */ import { useState } from "react"; -import ReactMarkdown from "react-markdown"; -import remarkGfm from "remark-gfm"; import { S } from "../../lib/strings"; import { humanizeDuration } from "../../lib/format"; import type { ThinkingItem } from "../../lib/omni/stream-model"; @@ -13,6 +11,7 @@ import { Chevron } from "../../components/ui/chevron"; import { StatusIcon } from "../../components/ui/status-icon"; import type { RunState } from "../../components/ui/status-icon"; import { LiveDuration } from "./live-duration"; +import { Md } from "./md"; export function ThinkingBlock({ item }: { item: ThinkingItem }) { const [open, setOpen] = useState(false); @@ -47,7 +46,7 @@ export function ThinkingBlock({ item }: { item: ThinkingItem }) { {open && (
- {item.thinking} +
)}
diff --git a/packages/web/src/features/chat/use-session-stream.ts b/packages/web/src/features/chat/use-session-stream.ts index 2822332..f886952 100644 --- a/packages/web/src/features/chat/use-session-stream.ts +++ b/packages/web/src/features/chat/use-session-stream.ts @@ -26,6 +26,9 @@ import type { StreamModel } from "../../lib/omni/stream-model"; export type { PendingApproval } from "../../lib/omni/stream-controller"; +/** Minimum spacing between version commits: below one frame at 8fps, invisible as staleness, but ~8× fewer full re-parses of the growing message during a fast large-code stream. */ +const BUMP_MIN_INTERVAL_MS = 120; + export interface SessionStreamState { /** View model (updated in place; version bump triggers re-render). */ model: StreamModel; @@ -70,13 +73,33 @@ export function useSessionStream( const placeholderRef = useRef(null); if (placeholderRef.current === null) placeholderRef.current = createStreamModel(); const rafRef = useRef(null); + const throttleRef = useRef(null); + const lastBumpAtRef = useRef(0); - // Coalesce high-frequency deltas: multiple pushes within one frame trigger only a single re-render. + // Coalesce high-frequency deltas: multiple pushes within one frame trigger only a single + // re-render, and commits are additionally spaced ≥BUMP_MIN_INTERVAL_MS apart. Every commit + // re-renders (and re-parses) the currently streaming message at its full accumulated length, + // so per-frame commits during a large streamed reply go O(n²) and freeze the main thread — + // the interval caps that at a bounded, invisible staleness. Inside the interval a single + // trailing timer is armed (the flush): the final deltas always land even if no further push + // ever arrives. const bump = useCallback(() => { - if (rafRef.current !== null) return; + if (rafRef.current !== null || throttleRef.current !== null) return; + const commit = () => { + lastBumpAtRef.current = Date.now(); + setVersion((v) => v + 1); + }; + const wait = lastBumpAtRef.current + BUMP_MIN_INTERVAL_MS - Date.now(); + if (wait > 0) { + throttleRef.current = window.setTimeout(() => { + throttleRef.current = null; + commit(); + }, wait); + return; + } rafRef.current = requestAnimationFrame(() => { rafRef.current = null; - setVersion((v) => v + 1); + commit(); }); }, []); @@ -134,6 +157,10 @@ export function useSessionStream( cancelAnimationFrame(rafRef.current); rafRef.current = null; } + if (throttleRef.current !== null) { + window.clearTimeout(throttleRef.current); + throttleRef.current = null; + } }; // initialStatus only serves as the first-frame placeholder; it doesn't rebuild the connection // on parent re-renders. diff --git a/packages/web/src/features/chat/work-group.tsx b/packages/web/src/features/chat/work-group.tsx index dca4813..a85e229 100644 --- a/packages/web/src/features/chat/work-group.tsx +++ b/packages/web/src/features/chat/work-group.tsx @@ -24,6 +24,7 @@ import { Chevron } from "../../components/ui/chevron"; import { StatusIcon } from "../../components/ui/status-icon"; import { approvalKey } from "../../lib/omni/stream-model"; import type { ChatItem } from "../../lib/omni/stream-model"; +import { LiveDuration } from "./live-duration"; import { MessageItem } from "./message-item"; import type { StreamRenderContext } from "./message-stream"; import { summarizeWork } from "./work-summary"; @@ -62,8 +63,11 @@ export function WorkGroup({ /** Whether this group is the last segment of the message stream (current turn still in progress): decides the default expanded/collapsed state. */ isLast: boolean; }) { + // Whether any item is in flight right now — also the only window in which the group's span is + // still growing, which the duration display below depends on. + const itemsRunning = items.some((it) => itemActive(it, ctx)); // Last segment + Task running = the model might still call another tool → show Running (even if there's no active item right now). - const active = (isLast && ctx.taskRunning) || items.some((it) => itemActive(it, ctx)); + const active = (isLast && ctx.taskRunning) || itemsRunning; const pending = hasPendingApproval(items, ctx); const [open, setOpen] = useState(isLast); const userToggled = useRef(false); @@ -79,7 +83,7 @@ export function WorkGroup({ // A pending approval must stay actionable: expand the group body regardless of collapsed state (the approval row lives inside it). const shown = open || pending; - const { steps, durationMs } = summarizeWork(items); + const { steps, durationMs, startMs } = summarizeWork(items); return (
@@ -106,11 +110,27 @@ export function WorkGroup({ {S.chat.workGroupSteps(steps)} )} - {durationMs > 0 && ( - - {humanizeDuration(durationMs)} - - )} + {/* Both states show the same quantity: the summarizeWork span (earliest item start → + latest item end). That's the canonical definition here — it's what work-summary.ts + documents, and unlike a group-open→now wall clock it is reconstructible from the + stored timestamps, so reloading the transcript reproduces the same number. While an + item is in flight its end isn't known yet, so the tick extends the span to *now* + (whole seconds); with nothing in flight the value freezes at the computed span, with + decimals — the same "don't tick through a wait we don't count" idiom as a tool card + parked on an approval. Ticking on while the group merely stays Running (the model + thinking between steps, or streaming its answer) would climb past the span and then + snap backwards the moment the group settles. */} + {itemsRunning + ? startMs !== undefined && ( + + + + ) + : durationMs > 0 && ( + + {humanizeDuration(durationMs)} + + )} {pending && !shown && ( {S.chat.approvalWaiting} diff --git a/packages/web/src/features/chat/work-summary.ts b/packages/web/src/features/chat/work-summary.ts index 89bb5be..5ff1cd6 100644 --- a/packages/web/src/features/chat/work-summary.ts +++ b/packages/web/src/features/chat/work-summary.ts @@ -2,38 +2,54 @@ * Summary for the "Reasoning & Tools" group header (pure logic, unit-testable): step count only * counts **tool calls** (thinking doesn't count as a step). * - * Duration is computed as the **union of time intervals**: overlapping time from parallel tool - * calls is counted only once, presenting the group's wall-clock working time rather than the sum - * of each item's duration (naive per-item summation would report 15-way parallel work spanning - * 9 minutes as 99 minutes). Intervals follow the same settlement convention as each item's - * durationMs (see settleToolDuration in stream-model.ts): thinking is [startedAtMs, +durationMs]; - * a tool's durationMs = argument-generation segment + execution segment (approval wait time is - * excluded), and the two segments are not adjacent on the timeline, so they must be split back - * into two intervals using the same formula: [argStartedAtMs, +generation segment] and - * [approvalAtMs ?? callStartedAtMs, +execution segment] — computing the whole span from the - * execution start point would shift the generation segment into the execution period, producing - * false overlap with parallel tools. Gaps between intervals (waiting for the model's next step) - * are not counted; a segment missing a start point can't be checked for overlap and falls back to - * plain summation. + * Duration is the group's **real wall-clock span**: earliest segment start → latest segment end, + * including everything in between — approval waits, gaps while the model decides its next step, + * parallel work counted once. The header answers "how long did this group take", not "how much + * compute happened inside it" (per-item cards carry the fine-grained settled durations). + * Segment endpoints follow the same settlement convention as each item's durationMs (see + * settleToolDuration in stream-model.ts): thinking is [startedAtMs, +durationMs]; a tool's + * durationMs = argument-generation segment + execution segment (the approval wait between them + * is excluded from the item's own duration, but lands inside the group span), so the endpoints + * are [argStartedAtMs, +generation segment] and [approvalAtMs ?? callStartedAtMs, +execution + * segment]. Segments missing a start point can't be placed on the timeline and fall back to + * plain summation on top of the span. `startMs` is the earliest placed start — the anchor the + * group header ticks from while an item is still in flight (once nothing is running the header + * freezes at the computed span, so the displayed number never jumps backwards). */ import type { ChatItem } from "../../lib/omni/stream-model"; -export function summarizeWork(items: ChatItem[]): { steps: number; durationMs: number } { +export function summarizeWork(items: ChatItem[]): { + steps: number; + durationMs: number; + startMs?: number; +} { let steps = 0; - let durationMs = 0; - const intervals: [number, number][] = []; + let fallbackMs = 0; + let minStart: number | undefined; + let maxEnd: number | undefined; + /** Group open = the earliest start stamp of ANY item, settled or not — a still-streaming first item must already anchor the live tick. */ + const seen = (startMs: number | undefined) => { + if (startMs !== undefined && (minStart === undefined || startMs < minStart)) minStart = startMs; + }; const add = (startMs: number | undefined, spanMs: number) => { if (spanMs <= 0) return; - if (startMs === undefined) durationMs += spanMs; - else intervals.push([startMs, startMs + spanMs]); + if (startMs === undefined) { + fallbackMs += spanMs; + return; + } + seen(startMs); + const end = startMs + spanMs; + if (maxEnd === undefined || end > maxEnd) maxEnd = end; }; for (const it of items) { if (it.kind === "thinking") { + seen(it.startedAtMs); if (it.durationMs !== undefined) add(it.startedAtMs, it.durationMs); continue; } if (it.kind !== "tool_call") continue; steps += 1; + seen(it.argStartedAtMs ?? it.callStartedAtMs ?? it.approvalAtMs); if (it.durationMs === undefined) continue; const genMs = it.argStartedAtMs !== undefined && it.callStartedAtMs !== undefined @@ -42,18 +58,10 @@ export function summarizeWork(items: ChatItem[]): { steps: number; durationMs: n add(it.argStartedAtMs, genMs); add(it.approvalAtMs ?? it.callStartedAtMs, it.durationMs - genMs); } - intervals.sort((a, b) => a[0] - b[0]); - let curStart: number | null = null; - let curEnd = 0; - for (const [start, end] of intervals) { - if (curStart === null || start > curEnd) { - if (curStart !== null) durationMs += curEnd - curStart; - curStart = start; - curEnd = end; - } else if (end > curEnd) { - curEnd = end; - } - } - if (curStart !== null) durationMs += curEnd - curStart; - return { steps, durationMs }; + const spanMs = minStart !== undefined && maxEnd !== undefined ? maxEnd - minStart : 0; + return { + steps, + durationMs: spanMs + fallbackMs, + ...(minStart !== undefined ? { startMs: minStart } : {}), + }; } diff --git a/packages/web/src/features/chat/workspace-browser.tsx b/packages/web/src/features/chat/workspace-browser.tsx index 7827b73..22ed988 100644 --- a/packages/web/src/features/chat/workspace-browser.tsx +++ b/packages/web/src/features/chat/workspace-browser.tsx @@ -401,10 +401,21 @@ export function WorkspaceBrowser({ ))}
)} + {/* Ghost style, matching the toolbar's upload label (text-xs, transparent until hover) — the bordered secondary look stood out from every neighbor. */} + {/\.html?$/i.test(preview.name) && ( + + {S.files.openInNewTab} + + )}
{S.files.download} diff --git a/packages/web/src/features/models/catalog-sync.ts b/packages/web/src/features/models/catalog-sync.ts new file mode 100644 index 0000000..7de6f6e --- /dev/null +++ b/packages/web/src/features/models/catalog-sync.ts @@ -0,0 +1,74 @@ +/** + * One-click sync of the Project model table with the built-in catalog ("sync presets" next + * to the search box): union semantics — catalog entries not configured locally are added; + * entries present on both sides are reset to the catalog's fields (context window, pricing, + * protocol, base URL, vision — the catalog wins wherever the two differ, including removing + * pricing the catalog doesn't carry); locally added models (including user-defined groups) + * are kept untouched. Credentials are never touched: merged rows carry no apiKey input (the + * PUT keeps the stored key) and existing rows keep their credential display state. + */ +import { presetModelEntries } from "@prismshadow/penguin-core/model-catalog"; +import type { RowState } from "./models-page"; + +type PresetEntry = ReturnType[number]; + +/** The catalog-owned fields of a row, in RowState's string-typed form (mirrors toRow). */ +function presetFields(p: PresetEntry) { + return { + vision: p.vision !== false, + contextWindow: p.context_window !== undefined ? String(p.context_window) : "", + clientType: p.client_type ?? "", + cacheRead: p.pricing ? String(p.pricing.cache_read) : "", + cacheWrite: p.pricing ? String(p.pricing.cache_write) : "", + output: p.pricing ? String(p.pricing.output) : "", + baseUrl: p.base_url ?? "", + }; +} + +/** A brand-new row for a catalog entry not configured locally (original: null -> added on PUT). */ +function presetToRow(p: PresetEntry): RowState { + return { + provider: p.provider, + modelId: p.model_id, + original: null, + ...presetFields(p), + originalBaseUrl: "", + apiKeyInput: "", + clearApiKey: false, + }; +} + +/** + * Merges the current rows with the built-in catalog. Existing rows keep their identity, + * credential state, and list position (fields are updated in place); catalog-only entries + * are appended in catalog order. Returns the merged rows plus added/updated counts for the + * success toast (updated counts only rows whose catalog-owned fields actually changed). + */ +export function syncRowsWithCatalog( + rows: RowState[], + preset: PresetEntry[] = presetModelEntries(), +): { rows: RowState[]; added: number; updated: number } { + const key = (provider: string, modelId: string) => `${provider}\0${modelId}`; + const index = new Map(rows.map((r, i) => [key(r.provider, r.modelId), i])); + const next = [...rows]; + let added = 0; + let updated = 0; + for (const p of preset) { + const i = index.get(key(p.provider, p.model_id)); + if (i === undefined) { + next.push(presetToRow(p)); + added += 1; + continue; + } + const row = next[i]!; + const fields = presetFields(p); + const changed = (Object.keys(fields) as (keyof typeof fields)[]).some( + (k) => row[k] !== fields[k], + ); + if (changed) { + next[i] = { ...row, ...fields }; + updated += 1; + } + } + return { rows: next, added, updated }; +} diff --git a/packages/web/src/features/models/model-grouping.ts b/packages/web/src/features/models/model-grouping.ts index ae09fa8..344b5fe 100644 --- a/packages/web/src/features/models/model-grouping.ts +++ b/packages/web/src/features/models/model-grouping.ts @@ -88,3 +88,13 @@ export function groupModelRows( (g) => g.rows.length > 0 || (!searching && g.provider.id === "custom"), ); } + +/** + * Flattens the library grouping into one ordered list (the chat model dropdown uses this): + * rows ordered exactly as the model page shows them — built-in provider groups in + * MODEL_PROVIDERS order, then user-defined groups, custom last; in-group row order + * preserved. + */ +export function orderModelsLikeLibrary(rows: T[]): T[] { + return groupModelRows(rows, "").flatMap((g) => g.rows); +} diff --git a/packages/web/src/features/models/models-page.tsx b/packages/web/src/features/models/models-page.tsx index 010b7ec..a3aeb37 100644 --- a/packages/web/src/features/models/models-page.tsx +++ b/packages/web/src/features/models/models-page.tsx @@ -18,12 +18,13 @@ * the protocol follows group semantics: a first-party vendor group doesn't persist * client_type (AgentHub auto-routes by upstream id, with env fallback resolved live from the * id), while custom / user-defined groups / gateways use a fixed OpenAI protocol, and - * gateways (OpenRouter / SiliconFlow) additionally pre-fill their endpoint base URL; the - * "get model id / API key" external links sit next to the corresponding input's label - * (shown in both add and edit dialogs). The group list ends with an "add group" action - * (user-defined groups share custom's semantics; the group appears once the first model - * saves successfully — groups are carried by the model entry's provider field, not - * persisted separately). + * gateways (OpenRouter / SiliconFlow / Qwen Token Plan) additionally pre-fill their endpoint + * base URL; the "get model id / API key" external links sit next to the corresponding + * input's label (shown in both add and edit dialogs). The group list ends with an "add + * group" action (user-defined groups share custom's semantics; the group appears once the + * first model saves successfully — groups are carried by the model entry's provider field, + * not persisted separately). The header also holds an owner-only "sync presets" action next + * to the search box (union-merge with the built-in catalog, see catalog-sync.ts). * * Saving does a PUT full-table replace (models not present are deleted; an empty apiKey * means keep the existing value); only the owner can edit. @@ -42,15 +43,18 @@ import { ApiError } from "../../api/client"; import { S } from "../../lib/strings"; import { useDocumentTitle } from "../../lib/use-document-title"; import { useProject } from "../../state/project"; +import { useAuth } from "../../state/auth"; import { USD_TO_CNY, useTheme } from "../../state/theme"; import type { Currency } from "../../state/theme"; import { Button } from "../../components/ui/button"; import { Input } from "../../components/ui/input"; +import { PasswordInput } from "../../components/ui/password-input"; import { Modal } from "../../components/ui/modal"; import { Select } from "../../components/ui/select"; import { toastError, toastSuccess } from "../../components/ui/toast"; import { Badge } from "../../components/ui/badge"; import { Chevron } from "../../components/ui/chevron"; +import { GlyphIcon } from "../../components/ui/glyph-icon"; import { ProviderLogo } from "../../components/ui/provider-logo"; import { SkeletonList } from "../../components/ui/skeleton"; import { EmptyState } from "../../components/ui/empty-state"; @@ -58,11 +62,16 @@ import { formatDateTime, humanizeTokens } from "../../lib/format"; import { MODEL_PROVIDERS, catalogEntryFor, + modelHomepageUrl, providerInfo, resolveModelEnv, } from "@prismshadow/penguin-core/model-catalog"; import type { ModelProviderInfo } from "@prismshadow/penguin-core/model-catalog"; import { groupModelRows, sameModelRef, userProviderInfo } from "./model-grouping"; +import { draftKey, loadDraft, saveDraft } from "../chat/draft-cache"; +import { syncRowsWithCatalog } from "./catalog-sync"; +import { tpsTone, ttftTone } from "./speed-test"; +import type { SpeedResult, SpeedTone } from "./speed-test"; /** Display currency follows the user setting (pricing is always stored in USD/million tokens; conversion happens only for display and input). */ const CURRENCY_SYMBOL: Record = { USD: "$", CNY: "¥" }; @@ -104,6 +113,26 @@ function inputToUsd(inputStr: string, currency: Currency): string { return currency === "CNY" ? trimNum(n / USD_TO_CNY) : trimNum(n); } +/** Group-header action glyphs (24x24 line paths): add, bulk key, gauge for speed test. */ +const PLUS_ICON = "M12 5v14M5 12h14"; +const KEY_ICON = + "M21 2l-2 2m-7.61 7.61a5.5 5.5 0 1 1-7.778 7.778 5.5 5.5 0 0 1 7.777-7.777zm0 0L15.5 7.5m0 0l3 3L22 7l-3-3m-3.5 3.5L19 4"; + +/** Speed-test glyphs (24x24 line paths): gauge for the group action, clock = TTFT, zap = TPS. */ +const GAUGE_ICON = "M12 14l3.5-3.5M20.49 17A10 10 0 1 0 3.5 17"; +const CLOCK_ICON = "M12 22a10 10 0 1 0 0-20 10 10 0 0 0 0 20ZM12 7v5l3.5 2"; +const ZAP_ICON = "M13 2 3 14h9l-1 8 10-12h-9l1-8Z"; + +/** Metric tone -> text color classes for the card speed badges. */ +const TONE_CLASS: Record = { + green: "text-green-600 dark:text-green-400", + yellow: "text-amber-600 dark:text-amber-400", + red: "text-red-600 dark:text-red-400", +}; + +/** In-page key for one model's speed result. */ +const speedKey = (provider: string, modelId: string) => `${provider}\u0000${modelId}`; + /** Default context window (tokens) for custom models when left unset. */ const CUSTOM_CONTEXT_DEFAULT = 128000; @@ -299,6 +328,13 @@ export function ModelsPage() { const { currentProject } = useProject(); const projectId = currentProject?.projectId ?? null; const isOwner = currentProject?.role === "owner"; + const userId = useAuth().user?.userId ?? null; + /** Per-model speed results (in-memory, reset on every project switch; "pending" while that model's turn is running). */ + const [speedResults, setSpeedResults] = useState>(new Map()); + /** Group whose speed-test confirmation dialog is open (provider id). */ + const [speedFor, setSpeedFor] = useState(null); + /** Group currently being speed-tested (provider id); tests run strictly one model at a time. */ + const [speedRunning, setSpeedRunning] = useState(null); const [rows, setRows] = useState(null); const [defaultModel, setDefaultModel] = useState(undefined); @@ -327,6 +363,10 @@ export function ModelsPage() { if (!projectId) return; setRows(null); setLoadError(null); + // Speed results are keyed by (provider, model_id) only, so another Project's identically + // named model would inherit a timing measured against a different endpoint and key — + // drop them along with the rows they annotate whenever the active Project changes. + setSpeedResults(new Map()); try { const res = await api.getModels(projectId); setRows(res.models.map(toRow)); @@ -368,6 +408,13 @@ export function ModelsPage() { setRows(res.models.map(toRow)); setDefaultModel(res.defaultModel); setVisionModel(res.visionModel); + // Default model changed: drop the stored draft's model selection so the draft chat + // follows the new default (a stored pick would otherwise pin the old model forever). + if (userId && res.defaultModel && !sameModelRef(res.defaultModel, defaultModel)) { + const key = draftKey(userId, projectId); + const draft = loadDraft(key); + if (draft.modelRef) saveDraft(key, { ...draft, modelRef: undefined }); + } toastSuccess(successText ?? S.common.saved); return true; } catch (e) { @@ -383,6 +430,60 @@ export function ModelsPage() { const groups = useMemo(() => (rows ? groupModelRows(rows, query) : []), [rows, query]); + /** + * "Sync presets": merge the built-in catalog into the current table (union; the catalog + * wins on differing preset entries, local additions and API keys stay untouched — see + * catalog-sync.ts). No-op with a toast when everything is already up to date. + */ + const syncPresets = async () => { + if (!rows) return; + const merged = syncRowsWithCatalog(rows); + if (merged.added === 0 && merged.updated === 0) { + toastSuccess(S.models.syncUpToDate); + return; + } + await persist( + merged.rows, + defaultModel, + visionModel, + S.models.syncDone(merged.added, merged.updated), + ); + }; + + /** + * Group speed test: one real request per model, strictly sequential (concurrent probes + * trip provider rate limits), each result written to the card as it lands. The + * confirmation dialog (speedFor) has already warned about quota by the time this runs. + */ + const runSpeedTest = async (providerId: string) => { + if (!projectId || !rows) return; + const targets = rows.filter((r) => r.provider === providerId); + setSpeedRunning(providerId); + try { + for (const row of targets) { + const key = speedKey(row.provider, row.modelId); + setSpeedResults((prev) => new Map(prev).set(key, "pending")); + try { + const res = await api.testModel(projectId, { + provider: row.provider, + modelId: row.modelId, + speed: true, + }); + setSpeedResults((prev) => new Map(prev).set(key, res)); + } catch (e) { + setSpeedResults((prev) => + new Map(prev).set(key, { + ok: false, + message: e instanceof ApiError ? e.message : S.common.unknownError, + }), + ); + } + } + } finally { + setSpeedRunning(null); + } + }; + /** * "Add group" confirm: a valid name that doesn't conflict with a built-in group or an * existing provider proceeds directly to that group's add-model dialog — groups are @@ -426,9 +527,9 @@ export function ModelsPage() {

{S.models.readOnlyHint}

)}
- {/* The header only holds search (add-model entry points live in each group header); - on narrow screens (flex-wrap wraps it to its own line) the search box shrinks - flexibly, fixed width at >=sm. */} + {/* The header holds search plus the owner-only "sync presets" action (add-model + entry points live in each group header); on narrow screens (flex-wrap wraps it + to its own line) the search box shrinks flexibly, fixed width at >=sm. */}
+ {isOwner && ( + + )}
@@ -498,6 +609,7 @@ export function ModelsPage() { disabled={busy} onClick={() => setAddingTo(group.provider.id)} > + {S.models.addToGroup} @@ -514,10 +626,34 @@ export function ModelsPage() { disabled={busy} onClick={() => setGroupKeyFor(group.provider.id)} > + {S.models.groupApiKey} )} + {isOwner && ( + + )} {group.provider.apiKeyUrl && ( // Not enough room on phone width for all group-level actions: collapse // it the same way as "bulk configure key", keeping the add entry and @@ -568,6 +704,7 @@ export function ModelsPage() { currency={currency} isDefault={sameModelRef(rowRef(row), defaultModel)} isVisionModel={sameModelRef(rowRef(row), visionModel)} + speed={speedResults.get(speedKey(row.provider, row.modelId))} onOpen={() => setEditing(rowRef(row))} /> )) @@ -627,6 +764,26 @@ export function ModelsPage() { /> )} + {speedFor !== null && ( + setSpeedFor(null)}> +

+ {S.models.speedTestConfirm(rows?.filter((r) => r.provider === speedFor).length ?? 0)} +

+
+ + +
+
+ )} {addGroupOpen && ( void; }) { const priced = row.cacheRead || row.cacheWrite || row.output; @@ -747,6 +911,40 @@ function ModelCard({ : S.models.noKey, ].filter((v): v is string => v !== null); + const speedBadges = + speed === "pending" ? ( + {S.models.speedPending} + ) : speed ? ( + speed.ok ? ( + + {speed.ttftMs !== undefined && ( + + + {Math.round(speed.ttftMs)}ms + + )} + {speed.tps !== undefined && ( + + + {speed.tps} tok/s + + )} + + ) : ( + + {S.models.speedFailed} + + ) + ) : null; return ( ); @@ -1020,16 +1223,29 @@ function ModelDialog({ {S.models.modelId} - {dialogProvider?.modelsUrl && ( - - {S.models.getModelIds} ↗ - - )} + + {/* Model homepage (the model's own page; gateway groups have per-model URLs) — only for existing rows, whose identity is settled. */} + {row && modelHomepageUrl(row.provider, row.modelId) && ( + + {S.models.homepage} ↗ + + )} + {dialogProvider?.modelsUrl && ( + + {S.models.getModelIds} ↗ + + )} + )} - {/* 1) API key — the most commonly used, placed first in the field section; "get API key" link next to the label. */} -