docs(changelog): 2026-08-04 batch entries (#192)

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
Yaowei Zheng
2026-08-04 23:07:16 +08:00
committed by GitHub
parent ab388b4f8c
commit d87c993dd4
8 changed files with 90 additions and 0 deletions
@@ -0,0 +1,13 @@
# Web App: skills management on the Agent settings page, deep-linked list icons
## Skills tab
Agent settings gains a Skills tab, placed right after Tools. The installed list re-reads `agent_state/skills/` on every fetch — the directory stays the single source of truth, exactly as the vault and schedule files do, so hand-installed or agent-installed skills show up without any registry. Rows carry the skill icon, name, localized short description and version/updated metadata; uninstalling asks for confirmation first.
An Import dialog opens from the tab header and leads with the recommended path: install by chatting with the agent. The source field takes more than a webpage URL — a GitHub/GitLab repo or directory, a local folder path, an install command from another ecosystem (`npx skills add …`, Claude Code / Codex plugin installs — those plugins are essentially skill files), or a bare marketplace reference — and a small classifier tailors the generated prompt per form (fetch the page / clone and locate the SKILL.md directories / read the folder / work out what the command would install and fetch that from its source rather than running it blindly). Every variant keeps the review-first clause and points at the skill-porting library skill when installed. "Open a new chat" pre-fills the composer with the generated prompt through the existing draft cache — the agent can adapt what it installs, which a blind unzip cannot. The dialog's second path uploads a skill zip: a new member-level `POST /api/projects/:p/agents/:a/skills/archive` accepts the archive as base64 JSON (the API's established upload shape), takes `SKILL.md` at the zip root or inside exactly one top-level directory, derives the name from that directory (frontmatter otherwise) under the usual name pattern, rejects path-escaping entries, and caps the payload at 200 files / 5 MB per file / 20 MB uncompressed. Uploading a name that is already installed answers 409 and the dialog offers an explicit overwrite, which replaces the directory wholesale. Unzipping uses a new `fflate` dependency in the server package.
The reverse direction ships too: each row carries an export action that downloads the installed skill as `<name>.zip` — `<name>-v<version>.zip` when the frontmatter declares an explicit version — the whole directory under a single top-level folder, served as a direct attachment like the trace download, under the same size caps, so a skill round-trips through the import endpoint unchanged. Export and uninstall are icon buttons with tooltips, following the icon-plus-tooltip rule the tool rows adopted.
## Agent list
The tools, vault-keys, schedules and skills stat icons on each agent row are now buttons that open the settings page landed on the matching tab. Settings learned a `?tab=` deep link for this: the value is validated against the live tab list (unknown values fall back to Overview) and tab switches keep the URL in sync without polluting history. The active-session count badge is gone from list rows — the sessions column already tells the story there — while the settings Overview keeps its active-session figure.
@@ -0,0 +1,9 @@
# Fix: ANSI color codes no longer leak into tool output
A nested `penguin run` driven through `exec_command` and polled with `input_command` filled Web tool cards with `[36m`/`[0m` fragments spliced into words (#102). Three layers each contributed, and each is fixed:
- **CLI**: the renderer wrote escape codes unconditionally. Color is now decided once per output stream — TTY, `NO_COLOR` unset, `TERM` not `dumb`, with a non-empty `FORCE_COLOR` overriding in either direction, matching Node's own semantics — and every renderer escape routes through that palette, so piped output is plain text.
- **Command tool environment**: the child environment always set `NO_COLOR=1` and `TERM=dumb`, but an inherited `FORCE_COLOR` silently won (Node ignores `NO_COLOR` when `FORCE_COLOR` is set). `FORCE_COLOR` and `CLICOLOR_FORCE` are now stripped — removed, not blanked, since Node reads an empty `FORCE_COLOR` as "on" — so the hardening finally holds; a vault-provided value still passes through by design.
- **Web**: tool output renders through a defensive ANSI stripper (CSI including multi-parameter SGR, OSC, two-byte escapes, and an incomplete trailing sequence cut mid-stream), applied at render time only — historical Traces display clean without their files being rewritten. The Trace event inspector deliberately keeps raw payloads: it is the raw-data view.
Regression tests cover all three layers, including the reported `FORCE_COLOR=3` + `NO_COLOR=1` + `TERM=dumb` combination and sequences split across streaming chunk boundaries.
@@ -0,0 +1,9 @@
# Models: OpenRouter gains qwen/qwen3.8-max
Prices and specs read from OpenRouter's models API on 2026-08-04.
The OpenRouter group gains `qwen/qwen3.8-max` — USD 2/6 per mtok input/output, cache reads at the published 0.25 hit price, cache writes at the genuine 2.5 (1.25× input) premium, a 1,000,000-token context window, vision — inserted ahead of `qwen/qwen3.6-35b-a3b` per the newer-versions-first ordering. The block's provenance comment now counts it among the rows carrying a real cache-write premium. The catalog already had the model through the qianwenai resale groups; this adds the OpenRouter spelling with OpenRouter's own USD prices.
The agenthub-models skill (v10) adds a Qwen 3.8 Max row recording the OpenRouter id — no other gateway in the catalog resells it, so that column stands alone.
As with every catalog addition there is no automatic migration: existing Projects keep their stored model tables and pick the row up through the models page's explicit "sync presets" action; new Projects see it out of the box.
@@ -0,0 +1,9 @@
# Web App: per-Project defaults for new chats
Project settings gains a "New chat defaults" section (owner-editable; members see it read-only; placed below Members, laid out as a compact two-column grid): the default Agent, working directory (empty means the auto temp dir), approval mode and thinking level applied when a chat is created. The workspace and model controls are the composer's own pickers — the draft view's directory browser and the chat input's model picker were extracted into shared components (`workspace-select.tsx`, `model-select.tsx`) and both original surfaces refactored onto them, so the settings dialog and the composer render one implementation. In the dialog they wear a `form` trigger variant styled with the dialog's own field tokens (the composer keeps its pill triggers byte-for-byte), their menus paint above the modal, and the modal's private Escape stack was generalized into a shared esc-layer stack so Escape closes the topmost thing first — menu, then dialog. The values live in an optional `[default_chat]` table in `.project_config.toml` — `agent_id`, `workspace`, `approval_mode`, `thinking_level` — served by a member `GET` and owner `PUT /api/projects/:p/chat-defaults`; the PUT is a declarative whole-block replace (an omitted key clears it), rejects unknown agents and invalid enum values, and read-modify-writes the TOML so models and credentials are untouched.
The draft view seeds from these beneath anything more specific: route state (e.g. "New Chat" from an agent row) beats an unsent draft, which beats the project default, which beats the previous hardcoded fallbacks. The model row deliberately adds no new key: it renders and writes the same top-level `default_model` the models page owns, through a narrow owner `PUT /api/projects/:p/models/default` that validates membership in the models table, and changing it releases any draft-pinned model exactly as the models page does (the two surfaces now share that helper).
Thinking level resolves as: the Agent's explicit `model.thinking_level`, else the project default, else `medium`. The draft picker shows the effective value, and a pick still writes the Agent's own config — the project value is only ever a fallback. One deliberate behavior change rides on this: an Agent config with no explicit level now follows the project default where it previously always meant the built-in default.
The block is additive — configs without it behave exactly as before and no migration runs. One version-skew caveat: an older penguin CLI (0.2.0 and earlier) rewriting a project config drops the block, since its loader rebuilds a known-keys literal; the current code round-trips it. The TOML renderer also learned to emit table-valued keys after scalar keys, so a later scalar append (say, a rename on a config that lacked `name`) can no longer be parsed into a preceding table.
@@ -0,0 +1,12 @@
# Skills: skill-porting — install skills from any ecosystem
New library skill `skill-porting` (agent-tuning group) teaching an agent to bring skills in from the outside world and land them correctly in `agent_state/skills/<name>/`. Its schema tables were verified by directly fetching each source on 2026-08-04 — the skill says so and treats the live JSON as authoritative:
- **Claude Code plugin marketplaces** (`anthropics/claude-plugins-official`, 278 plugins): the marketplace.json shape, all four `source` forms (relative path, sha-pinned `url`, `git-subdir`, `github`) and the five places a plugin can keep skills.
- **Codex plugins** (`openai/plugins`, 180 plugins): the `.agents/plugins/marketplace.json` entry shape with its `policy` block and `.codex-plugin/plugin.json` layout.
- **`npx skills add` / skills.sh** (`vercel-labs/skills`): spec forms, per-agent install destinations, and the in-repo discovery order — resolved by fetching from the source repo, never by blindly running the installer.
- Plain **GitHub repos/subdirectories** (sparse checkout / tarball / raw fetch variants) and **local folders**, plus the agentskills.io SKILL.md convention.
Each flow ends in the same normalization: adapt frontmatter to penguin's fields (`short_description`, `short_description_zh`, integer `version`, `updated`), port or drop commands/agents/hooks penguin has no runtime for (honestly — the skill forbids pretending a dropped capability survived), and verify the install. Safety first is mandatory: read every file before installing, refuse content that exfiltrates, phones home, or overrides safety rules, prefer pinned revisions.
The web Import dialog's generated prompt points agents at this skill when it is installed. Docs skill tables (en/zh) gained the one row the library-sync test requires.
@@ -0,0 +1,11 @@
# Web App: the Traces page scales like the session sidebar
The trace list used to fetch everything at once: every agent auto-expanded on load, each expansion walked the agent's entire trace tree server-side with a stat per file, every session row and file pill rendered, and any session the paged sidebar store hadn't loaded fell back to its raw id in mono type — long-lived Projects froze the page.
`GET /api/projects/:p/agents/:a/traces` now takes optional `offset`/`limit`. Without them the response shape is byte-identical to before; with them the server returns newest-first session groups plus a total, and only the returned page is stat-ed (sizes written back to the index).
Discovery itself no longer touches the tree per request. A SQLite-derived **trace index** (`trace_files` plus per-session facts read once at registration: origin, workspace, a first-prompt title fallback, model ref) serves every consumer. The hot path is two stats — an mtime gate on the traces root and the newest date dir; a moved gate triggers an incremental readdir of only the changed date dirs, each new session classified once with a bounded head-read; server-side writes (import, session/agent/project deletion) update the index synchronously; and every consumer's miss forces one full reconcile before it may 404, so a stale index costs one extra scan, never a lie. The index is a rebuildable cache — the on-disk Trace files stay the single source of truth, and the first request after an upgrade or restart simply rebuilds it. The same index now feeds the agents-list 30-day activity sparkline (previously a full-history walk on every agents-list request), the `/messages` subagent resolver (previously a project-wide scan), trace locate/download/delete, CLI-session discovery — adoption reads no file anymore — and session resume. Listings do zero head-reads; titles come from the stored facts: DB session title, else the registered first-prompt fallback.
The page now mirrors the sidebar's structure outright, not just its pure helpers. The sidebar's group header, folder section, "more" rows and group-mode toggle were extracted verbatim from its inner closures into shared `components/ui/group-list.tsx` — and the sidebar itself was refactored onto them, so the two surfaces render from one implementation. The traces tree gains both grouping modes (by Workspace or by Agent, on the same persisted preference as the sidebar — flipping the mode on either surface flips both) and the three lazy folders per group (subagent, schedule, archived), each paging independently. Server support is additive: paged session groups carry `category` and `workspace`, the response carries per-category and per-workspace counts, and a `category=` filter pages within one folder — all answered from the trace index, with classification exact on the first request since facts are stored at registration.
Elsewhere the earlier fixes stand: the page-size split for load-more (traces-specific steps — 20 sessions, 15 groups, larger than the sidebar's 10s since the full-height page holds more), group caps with a "more groups" row, pinned-first ordering for a deep-linked agent, only the focused (or first) agent expanded by default, capped file pills, and a memoized title lookup. CLI-origin sessions follow the user's "show CLI sessions" preference here too, defaulting to hidden; a "View trace" deep link to a CLI session still resolves and opens the detail pane regardless of the preference. Session ids no longer render anywhere on the page — an untitled session shows the standard "New chat" label, and the detail subtitle keeps date · size only. The chat page links back: its session info dropdown gains a "View trace" action deep-linking to the traces page with the agent focused and the session selected, falling back to a one-agent full fetch when the session sits beyond the first page or inside a folder.
@@ -0,0 +1,13 @@
# Web App: outline windowing, a cost stat that stays put, quieter failure and update chrome
## Conversation outline
The tick-rail minimap over the stream's left gutter now appears only once a conversation reaches 5 exchanges, and renders a sliding window of at most 20 turns either side of the active one instead of one tick per turn forever — the window recenters as reading position moves, shifts rather than shrinks at the ends, and small edge dots mark hidden ranges; global turn numbering is preserved and the toolbar dropdown fallback still lists every turn. Two overlap bugs went with it: the tick stack is height-adaptive with an overflow backstop, so very long runs can no longer spill ticks over the toolbar and composer; and the gutter-fit check, which compared against a hardcoded 768 px column, now resolves the column's real 48 rem width against the live root font size — under browser font scaling the old check kept the rail visible while the widened prose column ran beneath it.
## Cost stat
The toolbar cost chip could vanish mid-run: goal rounds reset the live task buckets while the running state blocked the session-total refetch, a page opened during an active run never fetched the accrued total at all, and the idle blip between queued follow-ups could clobber a known total with an empty response. The displayed figure is now sticky and monotone while a session runs — the last fetched total plus each finished Task's settled increment plus the open Task's live estimate, reconciled verbatim once the session is idle. The usage fetch fires on session open regardless of run state, and an empty response can no longer erase a known value.
## Quieter chrome
Tool rows no longer print stop-reason markers at all — `[failed]`, `[aborted]`, `[timeout]`, `[malformed]` and `[auth]` are gone; the status icon is the single carrier of the outcome, with the raw reason still in its tooltip and aria-label, and the Trace viewer keeping the literal per-event values. The home page's "new version" hint sheds its accent pill for plain superscript text set in the version line's own type; only the link affordance remains.
+14
View File
@@ -5,3 +5,17 @@
- [2026-08-04] Web App: the public fixed admin password `penguin-2026` is replaced by a random `penguin-<4 digits>` seed printed once at first start (`PENGUIN_SEED_ADMIN_PASSWORD` pins it for tests, policy-checked), and the login endpoint gains per-username exponential throttling (5 free failures, 1s doubling to 60s, `429 too_many_attempts`, reset on success, identical for unknown usernames) so the 4-digit space cannot be enumerated. ([details](2026-08-04-web-app.md))
- [2026-08-04] CLI: `penguin web` readiness-probe failures are no longer swallowed — the last error is classified (timeout / refused / reset / permission / DNS) and reported with actionable, localized guidance instead of a generic "not responding yet". ([details](2026-08-04-cli-web-probe-diagnostics.md))
- [2026-08-04] Skills: new `skill-porting` library skill — verified, schema-accurate flows for bringing skills in from Claude Code plugin marketplaces, Codex plugins, `npx skills add`/skills.sh, GitHub repos and local folders, normalizing them into penguin's skill format with a read-everything-first safety gate. ([details](2026-08-04-skill-porting-library-skill.md))
- [2026-08-04] Web App: skills become manageable from Agent settings — a Skills tab after Tools listing the installed set straight from disk, uninstall with confirmation, an Import dialog that leads with chat-driven install (URL + generated review-first prompt) and also accepts a skill zip through a new validated archive endpoint, and per-skill zip export that round-trips through that import; the agent list's tools/vault/schedules/skills icons deep-link to the matching settings tab via a new `?tab=` param, and the list rows drop the active-session badge. ([details](2026-08-04-agent-skills-and-list.md))
- [2026-08-04] Web App: Project settings gains a "New chat defaults" section — default Agent, working directory, approval mode and thinking level stored as an additive `[default_chat]` block in `.project_config.toml`, seeded into new drafts beneath route overrides and unsent drafts; the model row reads and writes the same `default_model` the models page owns through a narrow set-default endpoint, and thinking level resolves Agent-explicit → project default → `medium`. ([details](2026-08-04-project-chat-defaults.md))
- [2026-08-04] Web App: the Traces page becomes a second surface of the session sidebar — shared group/folder/mode-toggle components (the sidebar refactored onto them), workspace and agent grouping on the sidebar's own persisted preference, lazy subagent/schedule/archived folders, server-side session paging — and trace discovery plus the agents-list activity sparkline move off per-request filesystem walks onto a SQLite-derived, mtime-reconciled trace index (rebuildable cache; disk stays the source of truth). CLI sessions now follow the show-CLI-sessions preference (default hidden), titles fall back to the first user prompt, and raw session ids are gone from the page. ([details](2026-08-04-traces-page-scaling.md))
- [2026-08-04] Web App: chat refinements — the conversation outline rail appears from 5 exchanges and windows to ±20 turns around the active one (fixing its overlap onto the composer and, under browser font scaling, onto the prose column), the toolbar cost stat no longer blinks out mid-run, tool rows drop the stop-reason text markers entirely in favor of the status icon, and the home update hint sheds its pill for plain superscript text. ([details](2026-08-04-web-chat-refinements.md))
- [2026-08-04] Fix: ANSI color codes no longer leak into tool output — the CLI gates color on TTY/`NO_COLOR`/`TERM` with `FORCE_COLOR` override semantics matching Node, the command tool strips inherited `FORCE_COLOR`/`CLICOLOR_FORCE` from child environments so its `NO_COLOR=1` hardening wins, and the Web renders tool output through a defensive ANSI stripper covering historical Traces. ([details](2026-08-04-ansi-tool-output.md))
- [2026-08-04] Models: OpenRouter gains `qwen/qwen3.8-max` (1M context, vision, published cache-hit and cache-write prices), and the agenthub-models skill (v10) records the OpenRouter spelling — existing Projects pick it up through the models page's explicit "sync presets" action. ([details](2026-08-04-model-catalog-openrouter-qwen38-max.md))