docs(blog): 0.1.4 — Windows support, goal mode, and the agents panel (#96)

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
Yaowei Zheng
2026-07-28 02:01:53 +08:00
committed by GitHub
parent 2b708d37fe
commit f70b4299ac
5 changed files with 1286 additions and 2 deletions
@@ -0,0 +1,60 @@
---
title: "PenguinHarness 0.1.4: Windows support, goal mode, and the agents panel"
date: 2026-07-27
category: news
excerpt: 0.1.4 has one theme — the harness meets you where you work and keeps working when you step away. It now installs and runs on Windows with its own one-liner and CI; a new goal mode keeps driving a Session until the objective is actually done rather than merely replied to; and a new agents panel turns subagent fan-outs into a live call graph. Here is what shipped.
---
PenguinHarness 0.1.4 is out, and the release has one theme: the harness meets you where you work and keeps working when you step away. It now installs and runs on Windows; a new **goal mode** keeps driving a Session until an objective is actually done rather than merely replied to; and a new **agents panel** turns subagent fan-outs from nested cards into a live call graph you can watch and steer. Feature by feature:
## Windows is now a first-class install
One line in PowerShell:
```powershell
irm https://penguin.ooo/install.ps1 | iex
```
That fetches `penguin-win32-x64.zip` — the official Windows Node runtime bundled, so nothing needs to be preinstalled — verifies its SHA256, unpacks with a staged swap that never touches your `data\` directory, and puts `penguin` on your user PATH (reading and writing the registry value with its kind preserved, so an existing `%USERPROFILE%`-style entry survives the edit). The stable URL is a forwarder that downloads fully before executing, so a truncated stream cannot half-install. Prefer npm? `npm install -g @prismshadow/penguin-cli` works anywhere Node ≥ 24 does.
The interesting part was not packaging but the agent itself. Every `exec_command` used to die on Windows with `spawn bash ENOENT`, because the command session hardcoded bash. Command sessions now resolve their shell per platform: Git-Bash first when present (best compatibility with the POSIX-oriented skill ecosystem — and a `bash` that resolves into the Windows system directory is rejected, because that is the WSL launcher, a different filesystem view entirely), then `pwsh`, then `powershell`, with `PENGUIN_SHELL` as the override. The chosen shell is announced to the model through a `Shell:` line in the session environment, so it writes PowerShell syntax when PowerShell is what it has, instead of emitting bash into the void.
All of it is kept true by CI: a `ci-windows` job — full build, typecheck and tests, plus a PowerShell parse gate — now runs beside the required Ubuntu job, and getting it green surfaced and fixed real Windows findings, including a workspace-upload symlink guard that POSIX's `O_NOFOLLOW` had been providing silently. The remaining limits are documented rather than pretended away: the package is x64-only for now (ARM64 runs via emulation), Ctrl-C in `input_command` hard-kills the whole command tree instead of interrupting the foreground command, and upgrading means re-running the installer — in-place `penguin update` still refuses on Windows.
## Goal mode: loop until it is done
A normal Task ends when the model stops calling tools and replies. That is right for a request and wrong for an objective — "make the check suite green" is not done just because the model went quiet. Goal mode inverts the contract: you state an objective, and the system keeps driving Tasks on the same Session, re-injecting the objective every round, until the goal reaches a terminal state.
<img class="dark:hidden" src="/blog-assets/goal-mode-en-light.webp" alt="Goal mode mid-loop: round 3 of a check-suite objective, with the goal banner above the composer tracking the objective, the round count and tokens against the budget" width="1920" height="1350" />
<img class="hidden dark:block" src="/blog-assets/goal-mode-en-dark.webp" alt="Goal mode mid-loop in dark theme: round 3 of a check-suite objective, with the goal banner above the composer tracking the objective, the round count and tokens against the budget" width="1920" height="1350" />
Completion is claimed through a protocol, not inferred from silence. Each goal run creates a `GOAL.yaml` next to the session's `PLAN.md`; the model may change exactly one field, `status`, and only to `complete` or `blocked`. The rules injected with every round tell it to audit a completion claim against evidence — files, command output, test results — before writing it, never to shrink the objective to an easier subset, and to claim `blocked` only after the same blocker has held for three consecutive rounds, so a transient obstacle does not end the goal. An optional token budget (`500k`, `2m`) is checked between rounds; when it runs out, the model gets one wrap-up round to summarize progress and remaining work, and the goal ends as budget-limited instead of pretending success.
In the Web App, the composer's new "+" menu engages a goal chip with an inline budget field (`/goal` in the slash menu does the same). Every round renders as a regular user bubble with a "Goal · round N" notice beneath it, and a live banner above the composer tracks the objective, the round count and tokens against budget — the screenshot above is round 3 of a real loop, mid-verification. The CLI has the same power as `/goal[:<budget>] <objective>` in chat and `--goal` on `penguin run`, where only a completed goal exits 0 — an objective you can put in a script. The SDK keeps its one entry point: `session.run(input, { goal: { budget } })`.
## The agents panel: see the whole tree
`run_subagent` used to inline the child's entire conversation into the parent's message flow — cards nested inside cards, unreadable past the second child. 0.1.4 moves child conversations into a dedicated **agents panel** that docks on the right exactly like the Workspace files panel: toolbar toggle, drag to resize, a bottom sheet on phones. In the message flow, each child leaves a single bar row — avatar, resolved agent name, a spinner while it runs, and an amber dot whenever an approval is pending anywhere in its subtree, so a nested approval stays discoverable even with the panel closed.
<img class="dark:hidden" src="/blog-assets/agents-panel-en-light.webp" alt="The agents panel: a call graph with the main session as root and two named subagents with live elapsed times, above the selected child's streaming conversation" width="1920" height="1350" />
<img class="hidden dark:block" src="/blog-assets/agents-panel-en-dark.webp" alt="The agents panel in dark theme: a call graph with the main session as root and two named subagents with live elapsed times, above the selected child's streaming conversation" width="1920" height="1350" />
The top of the panel is a **call graph**: one node per participating agent — avatar, name, run-state dot and elapsed time, ticking while the child runs and frozen at the settled span once it is done, so a reloaded page shows the same durations as a live one. The main session is the root, and edges show who spawned whom. Click a node and the conversation below switches to that child, rendered by the same machinery as the main stream: its own user prompt is there, tool cards stream live, and approval buttons work from inside the panel. Visibility is task-scoped — every new Task starts with the panel closed, it auto-opens once when the task first spawns a subagent, and your manual open or close wins for the rest of the task. The graph follows the latest Task by default; click a bar row from an older turn and it pins that turn's historical spawn tree instead.
## Also in 0.1.4
The same release makes an in-progress reply survive a page refresh — the server keeps a live tail per running session, so after a reload the already-streamed prefix is back immediately and keeps growing. The chat header's cost and elapsed chips now tick live while a Task runs, joining the token count that already did. Trace files gained export and import, so a trajectory can move across deployments. And the app finally knows its own version: a "Check for updates" row with the running version inline, plus a one-click in-place update for admins. The full list is in the [v0.1.4 release notes](https://github.com/Prism-Shadow/penguin-harness/releases/tag/v0.1.4).
## Get it
```bash
# Linux / macOS
curl -fsSL https://penguin.ooo/install.sh | sh
```
```powershell
# Windows
irm https://penguin.ooo/install.ps1 | iex
```
Or `npm install -g @prismshadow/penguin-cli` with Node ≥ 24 — 0.1.4 is the version to install, since 0.1.3 carries the same feature set but never reached npm. Then run `penguin web`, add a model key on the Models page — and state an objective.
@@ -0,0 +1,60 @@
---
title: "PenguinHarness 0.1.4:Windows 支持、目标模式与智能体面板"
date: 2026-07-27
category: news
excerpt: 0.1.4 只有一个主题——harness 来到你工作的地方,并在你离开后继续工作。Windows 有了自己的一行安装命令与完整 CI;新的目标模式让 Session 循环推进,直到目标真正完成而不是「回复完了」;新的智能体面板把子智能体的并发执行变成一张实时调用图。逐项说明如下。
---
PenguinHarness 0.1.4 发布了。这个版本只有一个主题:harness 来到你工作的地方,并在你离开后继续工作。它现在可以在 Windows 上安装运行;新的**目标模式**让 Session 循环推进,直到目标真正完成而不是「回复完了」;新的**智能体面板**把子智能体从层层嵌套的卡片变成一张可观察、可操作的实时调用图。逐项来看:
## Windows 成为一等公民
PowerShell 里一行命令:
```powershell
irm https://penguin.ooo/install.ps1 | iex
```
它会下载 `penguin-win32-x64.zip`——内置官方 Windows Node 运行时,无需预装任何东西——校验 SHA256,用「先落地再替换」的方式解压安装(绝不触碰你的 `data\` 目录),并把 `penguin` 加入用户 PATH:注册表值按原始类型读写,已有的 `%USERPROFILE%` 风格条目不会被展开破坏。稳定地址是一个「完整下载后才执行」的转发脚本,截断的流不可能装出半个产品。偏好 npm 的话,`npm install -g @prismshadow/penguin-cli` 在任何 Node ≥ 24 的环境同样可用。
有意思的部分不在打包,而在智能体本身。过去每一次 `exec_command` 在 Windows 上都死于 `spawn bash ENOENT`——命令会话把 bash 写死了。现在命令会话按平台解析 shell:优先 Git-Bash(与 POSIX 取向的技能生态兼容性最好;解析到 Windows 系统目录的 `bash` 会被拒绝——那是 WSL 启动器,完全是另一个文件系统视图),其次 `pwsh`,再次 `powershell`,`PENGUIN_SHELL` 可随时覆盖。选定的 shell 会通过会话环境里新增的 `Shell:` 一行告知模型,让它在拿到 PowerShell 时写 PowerShell 语法,而不是对着空气输出 bash。
这一切由 CI 保真:新增的 `ci-windows` 任务(完整构建、类型检查、测试,外加 PowerShell 语法门禁)与必需的 Ubuntu 任务并排运行;把它跑绿的过程本身就发现并修复了数个真实的 Windows 问题,包括一个此前由 POSIX 的 `O_NOFOLLOW` 默默兜底的工作区上传符号链接漏洞。剩余的限制如实写进文档而非假装不存在:安装包目前仅有 x64(ARM64 经转译运行),`input_command` 里的 Ctrl-C 会整树杀掉命令进程而非中断前台命令,升级方式是重跑安装器——`penguin update` 在 Windows 上仍拒绝原地更新。
## 目标模式:循环到完成为止
普通 Task 在模型停止调用工具并给出回复时结束。这对「一个请求」是对的,对「一个目标」是错的——「让检查套件全绿」不会因为模型不说话了就算完成。目标模式把契约倒了过来:你给出目标,系统在同一个 Session 上持续驱动 Task,每一轮重新注入目标,直到目标进入终态。
<img class="dark:hidden" src="/blog-assets/goal-mode-zh-light.webp" alt="目标模式进行中:检查套件目标的第 3 轮,输入框上方的目标横幅实时显示目标、轮次与 token 用量对预算" width="1920" height="1350" />
<img class="hidden dark:block" src="/blog-assets/goal-mode-zh-dark.webp" alt="深色主题下的目标模式:检查套件目标的第 3 轮,输入框上方的目标横幅实时显示目标、轮次与 token 用量对预算" width="1920" height="1350" />
完成必须通过协议声明,而不是从沉默里推断。每次目标运行都会在会话的 `PLAN.md` 旁创建一个 `GOAL.yaml`;模型只允许改动其中一个字段 `status`,且只能改成 `complete` 或 `blocked`。每轮注入的工作规则要求它:声明完成前先对照证据(文件、命令输出、测试结果)逐项自查;不得把目标缩水成更容易的子集;只有同一障碍连续三轮存在才可声明 `blocked`,一次性的阻碍不会终结目标。可选的 token 预算(`500k`、`2m`)在轮与轮之间检查;预算耗尽时模型获得最后一个收尾轮——总结进展、列出剩余工作——然后目标以「预算受限」结束,而不是伪装成功。
在 Web App 里,输入框新增的「+」菜单可挂上目标条并内联填写预算(斜杠菜单里的 `/goal` 等价)。每一轮都渲染成一条普通的用户气泡,下方带「目标 · 第 N 轮」的标注;输入框上方的实时横幅跟踪目标、轮次与 token 用量对预算——上面的截图正是一次真实循环的第 3 轮验证现场。CLI 侧同样齐备:聊天里的 `/goal[:<预算>] <目标>`,以及 `penguin run` 的 `--goal`——只有真正完成的目标才以 0 退出,一个可以写进脚本的目标。SDK 保持唯一入口:`session.run(input, { goal: { budget } })`。
## 智能体面板:看清整棵树
`run_subagent` 过去把子会话的完整对话内联进父消息流——卡片套卡片,两个子智能体之后就没法读了。0.1.4 把子会话对话移进专门的**智能体面板**,像工作区文件面板一样停靠在右侧:工具栏开关、拖拽调宽、窄屏变成底部抽屉。消息流里每个子智能体只留下一条横条:头像、解析出的智能体名称、运行中的旋转指示,以及子树中任何位置有待审批时的琥珀色圆点——面板关着也不会错过嵌套审批。
<img class="dark:hidden" src="/blog-assets/agents-panel-zh-light.webp" alt="智能体面板:调用图以主会话为根、两个具名子智能体带实时用时,下方是选中子会话的流式对话" width="1920" height="1350" />
<img class="hidden dark:block" src="/blog-assets/agents-panel-zh-dark.webp" alt="深色主题下的智能体面板:调用图以主会话为根、两个具名子智能体带实时用时,下方是选中子会话的流式对话" width="1920" height="1350" />
面板顶部是一张**调用图**:每个参与的智能体一个节点——头像、名称、运行状态点与用时,运行中按墙钟实时走秒,结束后定格为实际耗时,刷新后的页面与实时页面显示同样的数字。主会话是根节点,边表示谁派生了谁。点击节点,下方对话切换到对应子会话,由与主消息流完全相同的机制渲染:它自己的用户提示词就在那里,工具卡片实时流式输出,审批按钮在面板里直接可用。可见性按任务划定作用域——每个新 Task 开始时面板关闭,任务第一次派生子智能体时自动打开一次,此后你的手动开合在该任务内始终优先。调用图默认跟随最新 Task;点击更早轮次的横条,则钉住那一轮的历史派生树。
## 0.1.4 还有
同一个版本里:进行中的回复现在能挺过页面刷新——服务端为每个运行中的会话维护实时尾部,刷新后已流出的前缀立即回来并继续增长;聊天页头部的费用与用时统计在 Task 运行期间实时走动,与此前就随请求推进的 token 计数并列;Trace 文件支持导出与导入,一条轨迹可以跨部署迁移;应用终于知道自己的版本——「检查更新」一行内联显示当前版本号,管理员可一键原地更新。完整清单见 [v0.1.4 发布说明](https://github.com/Prism-Shadow/penguin-harness/releases/tag/v0.1.4)。
## 获取
```bash
# Linux / macOS
curl -fsSL https://penguin.ooo/install.sh | sh
```
```powershell
# Windows
irm https://penguin.ooo/install.ps1 | iex
```
或 `npm install -g @prismshadow/penguin-cli`(Node ≥ 24)——请安装 0.1.4:0.1.3 的功能集与之相同,但未能发布到 npm。然后 `penguin web`,在模型页配好密钥——给它一个目标。
+2 -1
View File
@@ -11,7 +11,8 @@
"typecheck": "tsc --noEmit -p tsconfig.json",
"test": "vitest run --passWithNoTests",
"shots": "node scripts/capture-shots.mjs",
"blog-shots": "node scripts/capture-blog-shots.mjs"
"blog-shots": "node scripts/capture-blog-shots.mjs",
"blog-013-shots": "node scripts/capture-blog-013-shots.mjs"
},
"dependencies": {
"react": "^19.1.0",
File diff suppressed because it is too large Load Diff
+2 -1
View File
@@ -96,7 +96,7 @@ describe("frontmatter mapping (author / pinned / category)", () => {
it("reads the pinned flag and sorts the pinned post first", () => {
for (const locale of ["en", "zh"] as const) {
const posts = postsFor(locale);
expect(posts.length).toBe(11);
expect(posts.length).toBe(12);
// The launch post stays the single pinned post; newer posts sort under it by date.
expect(posts.filter((p) => p.pinned).map((p) => p.slug)).toEqual([
"introducing-penguinharness",
@@ -120,6 +120,7 @@ describe("frontmatter mapping (author / pinned / category)", () => {
it("filters by the news category, newest first", () => {
expect(postsFor("en", "news").map((p) => p.slug)).toEqual([
"introducing-penguinharness",
"penguinharness-0-1-4",
"free-models-in-penguin-harness",
"gemini-3-6-in-penguinharness",
"fireworks-credits-amd",