docs(changelog): record the 2026-07-24 batch (#65)

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
Yaowei Zheng
2026-07-26 23:07:54 +08:00
committed by GitHub
parent 790f45191e
commit 658b6265bb
7 changed files with 111 additions and 1 deletions
+1
View File
@@ -6,6 +6,7 @@ The root [../CHANGELOG.md](../CHANGELOG.md) keeps one brief line per release. Wo
- `<folder>/YYYY-MM-DD-short-slug.md` — one detail file per entry, named by the entry date: an H1 title, then what changed and why, using `##` sections once the entry covers more than one thing.
- `<version>/RELEASE.md` — the announcement published verbatim as the GitHub Release body: a lead sentence, then `## Install`, `## Highlights`, `## Notable in this release`, `## Requirements`, and a closing link to this folder. Only a numbered folder has one; it is written at release, when the version is finally known.
- Entries are grouped by the surface they change — Models, Web App, landing site, skills, docs, tooling — rather than one file per commit. A related change extends the existing file and its summary line instead of opening a new one.
- A batch that carries backward-compatibility handling collects it in one `YYYY-MM-DD-backward-compatibility.md`: what old shape is tolerated, how far it reaches, whether the user must act, and when it can be removed. The surface entries reference that file instead of re-explaining it. A batch with no such handling has no such file.
- Unreleased changes go in `unreleased/`, never in a numbered folder: the next version number is not knowable while the work is being written — the batch that accumulated as `0.2.0/` shipped as 0.1.1 — so a guessed number only creates a rename to get wrong later. At release, rename `unreleased/` to the version actually shipped, swap its `# Unreleased` heading for `# Version X.Y.Z` plus a `Released on <date>.` line, write `RELEASE.md`, add the release line to the root file, and create a fresh empty `unreleased/`. Released folders are frozen.
- `RELEASE.md` must be **committed before the tag is created**: the release workflow reads it from the tag's own checkout, so a file added afterwards is never published. When it is missing, the workflow falls back to notes GitHub generates from the merged PRs rather than publishing an empty body.
@@ -12,6 +12,10 @@ The preview URL takes its port from the server's own binding rather than the bro
The link is a plain anchor with `rel="noopener noreferrer"` rather than a fetch followed by `window.open`: opening a tab after an await trips popup blockers, and a script-opened window keeps an `opener` handle back to the app — the exact reference the separate origin exists to deny. The server also binds the IPv6 loopback alongside IPv4, because `localhost` commonly resolves to `::1` first and every preview URL would otherwise refuse the connection.
The in-app rendered view now goes through the same door. It used to feed the file's text into its iframe via `srcdoc` — a document with no real base URL, so a page's relative subresources (`foo.png`, `app.js`, `style.css`) could never resolve — while "Open in new tab" already served real URLs on the preview origin. With isolation available, the Files panel's iframe now loads the same `preview-redirect` URL, so relative assets resolve and the page's storage works in-panel too, under a sandbox (`allow-scripts allow-same-origin allow-forms allow-popups allow-modals allow-downloads` — now named exactly in the docs) that stays strictly tighter than the unsandboxed new tab; without isolation, the old `srcdoc` fallback and its warning remain. The source text is fetched lazily on the first Source toggle, so big files are no longer downloaded twice and a transient fetch failure affects only the Source view.
Follow-up hardening on the isolation signal: `previewIsolated` is refetched after a UI login rather than trusted from mount time, so a deployment without a separate origin no longer sends the panel down the isolated path after logging in through the UI; and the server rejects a configured preview origin whose host equals the app's own — cookies ignore ports, so a port difference separates nothing — reporting `false` on `/api/me` and engaging the safe fallback instead of granting `allow-same-origin` to a same-origin document.
## Unified form controls, one notification rule, and deduplicated copy
The form-control layer had grown five near-identical copies of the same scaffolding. A shared `Field` wrapper now owns the label / hint / error markup, and a shared `controlBase` class string owns the border, hover, focus-ring and dark-mode treatment, so `Input`, `Textarea`, `Select`, `OptionMenu` and `PasswordInput` all read from one source instead of each carrying their own. `Select` dropped its hand-rolled open/close/position effect and now reuses the same `usePortalPanel` hook `OptionMenu` already used, so both menus flip up-or-down, clamp to the viewport and close on scroll/Esc through one code path. The close button triplicated across `Modal` / `Drawer` / `Sheet` became a single `CloseButton`, and the chevron, checkmark and plus glyphs that were inlined at a dozen call sites became shared icon components. The Agent runtime tab's thinking-level and compaction-mode dropdowns dropped their separate "Reset to default" link (a `ResettableOptionMenu` wrapper) in favour of a plain `OptionMenu` — the link only rewound an in-session pick, which closing the dialog already does.
@@ -55,3 +59,23 @@ The auto-opened "last conversation" follows one rule now too. Entering `/chat` w
Actions that write files on the server used to be a mix: deletions confirmed, a skill *update* confirmed, but a settings save (AGENTS.md / system prompt, runtime params, tool overrides), a skill *uninstall* (which deletes the installed copy, local edits included) and a workspace upload landing on an existing name all executed on one click. They now all confirm first through the shared `ConfirmModal`, which also got a face-lift: a compact card with **no title bar** — just a tinted icon badge that sets the tone at a glance (red warning triangle for deletions, neutral pencil for saves/overwrites), the message beside it (each message carries its own subject and consequence), and a small Cancel/Confirm pair. The former titles still name the dialogs for assistive tech. Every existing confirmation (delete session/agent/project/schedule/vault entry, model dialog writes, import version conflict) was moved onto the same component, so confirmations look identical everywhere. Adding a vault variable under an already-configured name also confirms now (it overwrites an unrecoverable value); quick-set controls inside the chat (approval mode, thinking level) intentionally stay one-click.
Clicking save with nothing changed used to do nothing silently (most forms skipped the request; the models dialog even reported "Saved"). Every save path now detects the no-op — the agent settings tabs, the schedule edit dialog and the model edit dialog — and shows an info toast: "No changes to save".
## The chat input steers or queues while a task runs
Sending during a run used to go nowhere — the input stayed usable, but submitting was a no-op until the task finished. While a task runs, the composer now does two real things: **Steer** (the default) delivers the draft mid-run as a `[user_steering]` message riding into the model's next request, and **Queue** posts the whole draft — images, skills and the @ handoff included, through the full normal composition path — with `queueIfBusy`, so the server holds it and auto-sends it as an ordinary follow-up task when the run ends. The mode is picked in a new More settings popover on the input toolbar — an extensible panel of setting rows, available in draft state too, replacing an inline segmented switch that overflowed on narrow screens — and the choice is remembered per user the way the sidebar grouping mode is, so a running-state send simply follows it. Stop and Send also became one button: with the composer empty it stops the run, and any content turns it into Send, which steers or queues per the remembered mode. The toolbar itself is now a left/right row whose pickers render through the shared portal layer, so a panel is never clipped by the scrolling toolbar on a narrow screen. A steer that races the run's completion falls back to that same full send path. Queued follow-ups survive abort and compaction, with a server-driven "N queued" hint near the input that survives reloads; a successful steer shows a lightweight queued hint until its message appears in the stream. Steering messages render as in-flow, user-styled chips in the conversation — they never split the Task — and the running placeholder explains the delivery.
## /model switches the conversation to another model
A `/model` slash command in an active idle session opens the model picker (the same list and grouping as the draft picker, current model marked). Picking another model continues the conversation handoff-style: a NEW session for the same agent on the picked model, reusing the source session's Workspace and approval mode, whose first message opens with a paired `[model_switch_from]` source block naming the source session, its latest trace file, the workspace and the previous model — the model reads the earlier context from the source trace when it needs it. Any text left after the command rides into that first task like a normal send, and the new session shows a "switched model" banner linking back to the source. The command consumes its token immediately — Escape and click-outside just close, like `/compact` — the run state is re-checked at pick time, and picking the current model is a no-op.
## Thinking level becomes a per-turn override in active sessions
The read-only thinking tag in an active session becomes an editable picker listing just the levels. It initializes to the Agent config's level, fetched through the existing agent-config endpoint, and while untouched it sends nothing with a task — the config fallback keeps applying, so a mid-session config edit still takes effect. Once the user picks a level it pins for the session and rides on every task, without ever writing through to the Agent settings; the draft picker's switch-becomes-default behavior is unchanged.
## Call descriptions on tool cards, and approvals that show the payload
Tool-call cards show the model-written call description as the collapsed-header subtitle, and the agent settings Tools tab gains a per-row call-description switch for each tool entry whose parameters declare the property. The pending-approval block renders the decoded file-tool payload — `old_string` / `new_string` / `content` alongside the path — in a scrollable pane, and the approval preview keeps the real command for the shell tool: the user approves the actual change, not the model's summary of it. File-tool previews shorten the path to one parent directory plus filename; the full path stays in the expanded arguments.
## Free models carry a Free badge
Free rows — the `:free` variants and the Free Models Router — now show a light-yellow "Free" badge, a new shared Badge tone deliberately distinct from the warning amber, on the Models page cards, next to the default and vision badges, and in the chat model picker. A shared `isFreeModel` helper decides: the entry must carry explicit pricing with all three buckets zero, so an unpriced or partially priced row — its costs merely unknown — never badges as free.
@@ -0,0 +1,32 @@
# Backward compatibility in this batch
What this batch keeps tolerating from data and configuration already on disk, how far each allowance reaches, whether the user has to do anything, and when it could be dropped. Other entries describe the features themselves and point here.
## Legacy angle-bracket markers stay readable
System-synthesized markers moved to the paired square-bracket form (`[summary]`, `[context_summary]`, `[use_skills]`, `[handoff_from]`, `[scheduled_task]`, `[developer_instructions]`, `[turn_aborted]`, `[turn_retried]`, `[model_switch_from]`, plus the inner transcript tags). Producers emit only the new form; every parser that can meet older material accepts both. Two sources make this unavoidable: Traces written before the change, which are re-rendered and replayed on resume, and the compaction prompt persisted in each existing agent's `system_config.yaml`, which keeps instructing the model to answer in `<summary>` tags for as long as that file is untouched. `model_switch_from` is a narrower case — its angle form only ever existed in traces produced while the merged `/model` switch ran on main before the markers landed — but it rides in the same dual-form parsers.
Nothing is required of the user. The policy now lives in one place — core's `omnimessage/markers/` module (exported as `@prismshadow/penguin-core/markers`), which owns every marker's producers and parsers — so removing the allowance later is an edit in one module rather than a sweep. Removal is indefinite all the same: the dual parsers cannot go while old Traces are still expected to open and old agent configs are still honored verbatim, so this is major-release material, not cleanup for a later minor version.
## session_meta.thinking_level is still read on resume
The thinking level became a per-request parameter and is no longer written into `session_meta`. Traces recorded earlier still carry the field, so resume reads it loosely and keeps honoring it as that session's default level — which matters most for a subagent session that inherited a level from its parent and would otherwise come back at its own agent's configured level. The legacy literal `"default"` and a missing field both fall back to the agent config, as before; rebuilt metas never write the field again, and the Traces view still renders the row when a legacy meta carries it.
Nothing is required of the user. This one is cheap to keep (a single tolerant read plus the view's fallback) and could be dropped once resuming pre-change Traces is no longer supported — again a major-release decision.
## Existing agents keep their stored config verbatim
An agent's `system_config.yaml` is loaded exactly as written, with no migration and no merge against the current defaults. A pre-existing agent therefore keeps its old system prompt — including the old marker documentation and the old `Project Dir` wording the new prompt replaces with App Data Dir — and its recorded `tools.builtin` list stays frozen, so it does **not** gain `read_file` / `edit_file` / `write_file`; the settings UI adds no rows either. This is deliberate: a stored config is the user's file.
Adopting the new defaults therefore requires the user to act, in one of two ways: run **Restore default configuration** (agent settings, Overview tab; `POST /api/projects/:projectId/agents/:agentId/config/reset`), which rewrites the file to the current defaults and preserves only the agent's name, description and version — every customization, including an edited system prompt, model and compaction settings and MCP servers, is overwritten; or hand-edit the YAML to add just the wanted entries. Other Agent State files (AGENTS.md, skills, vault) are untouched either way. Nothing here is a shim with an end date — verbatim loading is the intended behavior and stays.
## A missing tools.call_description means enabled
The per-tool `call_description` flag governs whether a tool's `description` argument is offered to the model. A missing key reads as enabled, so the four command/subagent entries in configs written before the flag existed keep taking call descriptions without anyone editing anything, and an explicit `false` filters the property out of the assembled schema without ever rewriting the stored YAML. This is a defaulting rule rather than a compatibility shim; it has no removal date, and flipping the default later would need a migration.
## What this batch does not add
The separate-origin HTML preview, the OpenRouter catalog rows and the dev data root carry no compatibility handling for stored data. Two consequences are worth knowing anyway, neither of them a tolerated old shape:
- New catalog rows reach new Projects automatically, because the preset list is written when a Project's `.project_config.toml` is created. An existing Project does not change by itself — the user picks up the new models through **Sync presets** on the Models page or `penguin config model add` from the CLI.
- Development entry points now default `PENGUIN_HOME` to `~/.penguin/dev-data`. Data written by earlier from-source runs stays where it was, under the installed root; a developer who wants to keep working against it exports `PENGUIN_HOME` explicitly, which still wins over the default.
@@ -0,0 +1,43 @@
# Models and core: file tools, mid-run steering, /model switches, and new OpenRouter models
## File tools: read_file, edit_file, write_file
The builtin toolset gains three file tools. `read_file` returns a `cat -n`-style line-numbered view with `offset`/`limit` paging (2000 lines by default), reading incrementally instead of slurping the file: an abort-aware scan with a hard 8MB cap that retains only the requested window, detects CRLF, and truncates overlong single lines. Its output cap is 64000 characters against the 16000 default of the other builtins, and the output is self-budgeted so the trailing "file has N lines total; showing X–Y" continuation note always survives truncation. It refuses binary content with a pointer to the shell and image tools, and refuses `.vault.toml` / `.project_config.toml` by basename — it is auto-approved under read-only approval, so the secret files need their own guard. `edit_file` performs exact `old_string` → `new_string` replacement: the string must occur exactly once — zero occurrences and duplicates both fail with an explanation — `replace_all` covers bulk changes, and success answers with a summary line plus a git-style unified diff: one hunk per replacement site with nearby sites merged, capped at 5 hunks on a `replace_all` storm with an "…and N more replacements" note, self-budgeted so the summary and note survive truncation. `write_file` writes a whole file, creating parent directories as needed, and reports whether it created or overwrote — an overwrite appends a real line diff against the previous content (a `+X/−Y` one-liner once the diff passes ~60 lines), a created file carries none — and the CLI colors the `+`/`-`/`@@` lines of both tools' diff outputs. Both writers land atomically (a temp file in the target directory, then a rename that preserves existing permission bits), and the approval prompt in both UIs shows the decoded edit or write payload, so the user approves the actual change. All three resolve relative paths against the Workspace and accept absolute paths under the same trust model as the shell tool, which remains the general-purpose fallback for everything else. The docs site's framing — it used to present "no file tools at all" as a design tenet — was rewritten around the nine-tool builtin set (zh + en); the landing blog's earlier posts keep their historical six-tool description. New agents get the file tools; an existing agent's frozen tool list does not, until its config is restored to the current defaults — see [backward compatibility](2026-07-24-backward-compatibility.md).
## Model-written call descriptions
The four command/subagent tools — `exec_command`, `input_command`, `run_subagent`, `input_subagent` — take an optional `description` argument: one model-written sentence, in the user's language, saying what the call is doing, shown in the UIs while it runs. The parameter lives in the four entries' schemas in `system_config.yaml` like any other tool parameter, and it is required there while the entry's `call_description` toggle — missing means enabled — is on, so a tool that offers a description always carries one and the frontends can pick a call's display form up front instead of guessing mid-stream; switching the toggle off filters the property and its `required` entry out of the assembled schema, never rewriting the stored YAML. The web Tools tab switches it per row. The file tools take no `description`: their path argument is self-describing. Every schema leads with its most user-visible argument — `description` first on these four tools, `file_path` first on the file tools — so it streams first into the CLI and web previews.
## Exact CLI call and output formats
Tool calls in the CLI render as `name <- description (payload)`, or plain `name <- payload` for a tool whose schema offers no description — the form is settled up front from the session's assembled tool list rather than guessed as arguments stream, so the plain form streams live while a described call waits for its sentence and only ever appears in the described form. Call and output lines share an identical `[tool-NNN] name` prefix, so `[tool-NNN] exec_command <- $ date` pairs with `[tool-NNN] exec_command -> output`; an output whose call was never seen keeps the bare `[tool-NNN] ->` tag. File-tool previews shorten the path to one parent directory plus filename, and the chat and run startup banner now prints the version and the Agent, Workspace and Model each on its own line.
## Steering: message a running task
A running task used to be out of reach — anything sent mid-run just waited for the next idle. `Session.steer` and `POST /api/sessions/:sessionId/steer` now deliver a message into a RUNNING task: the text is queued and handed over at the next input assembly as a standalone user message wrapped in a paired `[user_steering]` marker, sent alongside that turn's tool outputs — written to Trace as real user input, streamed at that same point, and never splitting the Task in the UIs. Tool outputs are never rewritten, and steering authority comes from the message's user role — a tool printing the marker inside its own output cannot impersonate the user. The queue drains at every input assembly, including immediately after a compaction (steering that arrives during a long summarize request is delivered right behind the summary), and is discarded only when the run exits; a turn that ends with the queue still non-empty continues the loop with the queued text through the same mechanism. The marker is documented in the default prompt's System markers section.
A busy session can also hold ordinary follow-ups: `POST /api/sessions/:sessionId/tasks` accepts `queueIfBusy`, answering 202 and queuing the input server-side; when the run ends — abort included — each queued input auto-starts as a normal task, in order, carrying the per-turn thinking level stored alongside it rather than falling back to the session default, with the queued count carried on task-state events and the session snapshot. In the CLI, a line typed while a task runs becomes steering (the acknowledgment echoes the text), the renderer holds streamed output while the user is composing a line and flushes it on submit, and a steer that loses the completion race is re-submitted as the next normal prompt.
## System markers move to paired square brackets
Every system-synthesized marker switches from the angle-bracket to the paired square-bracket form: `[turn_aborted]`, `[turn_retried]`, `[context_summary]`, `[summary]`, `[use_skills]`, `[handoff_from]`, `[scheduled_task]`, `[developer_instructions]`, plus the inner transcript tags inside synthetic blocks. Producers emit only the new form, while the parsers keep reading the legacy angle form — see [backward compatibility](2026-07-24-backward-compatibility.md). Every marker's producers and parsers moved into one core module, `omnimessage/markers/`, also reachable as the `@prismshadow/penguin-core/markers` subpath, so the tag literals and the dual-form reading rule have a single home instead of being spread across core, server and the frontends. The docs' marker literals were updated accordingly (zh + en).
## Thinking level is per-request; session_meta keeps invariants only
`thinking_level` is removed from `session_meta`: the meta now holds only per-session invariants, and anything the user can change mid-conversation is either a per-turn parameter or a new session. The level threads through as a per-request parameter instead — `GenerativeModelParameters`, `RunOptions` and `TaskCreateRequest` — with the construction-time value acting as just the default; compaction requests deliberately keep the default. Traces that recorded the field are still honored on resume — see [backward compatibility](2026-07-24-backward-compatibility.md).
## Switching models: a new session that reads the source trace
Continuing a conversation on another model injects no history into the new session: thinking payloads and provider fidelity are model-bound and cannot be replayed faithfully across models, so the earlier history-carrying fork design was dropped for the handoff pattern. The web's `/model` command opens a NEW session for the same agent on the picked model through the ordinary session-creation API — deliberately reusing the source session's Workspace, so files the conversation refers to stay reachable, and its approval mode — and posts a first task whose input opens with a paired `[model_switch_from]` source block — the source session id, the absolute path of its latest trace file, the workspace, and the previous model pair — following the square-bracket marker convention. The model reads the source trace itself when it needs the earlier context. `SessionInfo` gains an optional `tracePath` (the latest shard, populated on the single-session GET only), and the marker is stripped from generated titles.
## The prompt presents the project dir as the App Data Dir; restoring defaults is the adoption path
Models kept mistaking the Environment section's `Project Dir` for the user's working directory (which is `CWD`) — treating its contents as task input and writing task deliverables there. The default prompt now presents the same value under a name that says what it is: the Environment line reads App Data Dir (the `{{PROJECT_DIR}}` placeholder is unchanged, so existing custom prompts are unaffected), and the File system section states the semantics explicitly — PenguinHarness's application data root, holding every agent's data files plus the project-level data files; not the current task's directory, not user-provided input, never a place for task deliverables (the working folder is `CWD`). Prompt paths (Agent State, other agents' assets, scratchpad, skills) are spelled `<app_data_dir>/agents/…`, and the secrets rule points at `.project_config.toml` directly under the App Data Dir. Six of the builtin skills read the App Data Dir field from the Environment section — the eval-family's project-root derivation simplifies to the App Data Dir itself — and the web system-prompt editor's placeholder hints document `{{PROJECT_DIR}}` with the app-data-root description, in the prompt's order with the full set.
Default changes never reach existing agents by themselves, since a stored config is loaded verbatim. The adoption path is a new "Restore default configuration" operation — `POST /api/projects/:projectId/agents/:agentId/config/reset` and the matching action on the agent settings Overview tab, behind a confirm dialog — the config-side analogue of a skill update; what it preserves and what it overwrites is spelled out under [backward compatibility](2026-07-24-backward-compatibility.md).
## OpenRouter additions: three free models and Claude Opus 5
The catalog gains three $0 OpenRouter rows: `inclusionai/ling-3.0-flash:free` and `poolside/laguna-m.1:free` — free tiers of Ling 3.0 Flash (a 124B-parameter MoE) and Laguna M.1, both with a 262,144 context — and `openrouter/free`, the Free Models Router that sends each request to a random free model. The router deliberately claims no vision support, since the harness must not send images to a router whose target may be text-only, and records a conservative 128,000-token context floor, so the 75% compaction clamp fires at 96,000 tokens: with no window recorded, the clamp was silently disabled, and a long Session routed to a small-window target would hard-fail instead of compacting early. The catalog test's free-pricing invariant covers exactly the `:free` suffix and the router id, the web UI badges such rows as Free, and the models docs (zh + en) note the free variants, their zero cost, and OpenRouter's free-tier rate limits and data policy.
Requested alongside the free lineup, `anthropic/claude-opus-5` joins the OpenRouter block as a paid flagship: context 1,000,000, vision, and per-million pricing of $5 input / $25 output with the real published cache prices ($0.50 read, $6.25 write) rather than the repeat-input convention — positioned newest-of-series first, right after `claude-fable-5`.
@@ -0,0 +1,3 @@
# Sites: a free-models post, featured in the announcement bar
The blog gains a bilingual post, `free-models-in-penguin-harness`: the free lineup now in the catalog — Nemotron 3 Ultra (free), the newly added Ling 3.0 Flash and Laguna M.1 free tiers, and the `openrouter/free` Free Models Router — how to enable it (a free OpenRouter key; presets automatic on new Projects, Sync presets on the web Models page or `penguin config model add` from the CLI), and honest caveats: free-tier rate limits, prompts handled under the upstream provider's own terms, availability and quality variance, and the router's varying per-request target — PenguinHarness records a conservative window for it so long Sessions compact early instead of failing — fine for trying things out and light automation, paid models for serious work. The post is featured as the newest announcement-bar slide, at the head of the rotation.
@@ -0,0 +1,3 @@
# Tooling: the dev data root separates from the installed one
Running the server or CLI from source used to operate on the very same data as the user's real installed penguin: core's `resolveRoot()` returns `PENGUIN_HOME ?? ~/.penguin/data`, the dev entry points set nothing, and so `pnpm dev` shared agents, sessions and `web.db` with the production install. The development entry points — `pnpm dev`, `pnpm dev:server`, `pnpm penguin`, and packages/cli's own `penguin` script — now default `PENGUIN_HOME` to a separate dev data root, `~/.penguin/dev-data`, via POSIX `${PENGUIN_HOME:-…}` substitution in the scripts themselves. An explicitly exported `PENGUIN_HOME` still wins, and the product defaults — `resolveRoot()`, the server config, CLI `--root` — are unchanged: the isolation lives only in the dev scripts, so the installed penguin keeps `~/.penguin/data`. CONTRIBUTING.md's development-commands section and the installation docs (zh + en) document the dev root.
+5 -1
View File
@@ -2,5 +2,9 @@
Changes since v0.1.1. The version number is assigned at release, when this folder is renamed.
- [2026-07-24] Models and core: three builtin file tools (`read_file` / `edit_file` / `write_file`) with bounded reads, atomic writes and the shell as the general fallback, model-written call descriptions configured per tool entry, exact CLI call/output formats and a fuller startup banner, mid-run steering delivered as standalone `[user_steering]` user messages plus a queue-if-busy follow-up path, all system markers unified onto paired square brackets (the angle form stays parsed), the thinking level turned into a per-request parameter with `session_meta` reduced to invariants, model switching done handoff-style through a new session that reads the source trace, the default prompt presenting the project dir as the App Data Dir with a restore-defaults action as existing agents' adoption path, and OpenRouter catalog rows for three free models plus the paid Claude Opus 5. ([details](2026-07-24-models-and-core.md))
- [2026-07-24] Sites: a bilingual free-models blog post — the lineup, how to enable it, and honest caveats — featured as the newest announcement-bar slide. ([details](2026-07-24-sites-and-blog.md))
- [2026-07-24] Tooling: the dev entry points default `PENGUIN_HOME` to a separate `~/.penguin/dev-data` root, so running from source no longer shares agents, sessions and `web.db` with the installed penguin (an exported `PENGUIN_HOME` still wins; product defaults unchanged). ([details](2026-07-24-tooling.md))
- [2026-07-24] Backward compatibility: what this batch keeps tolerating from existing Traces and configs — the legacy angle-bracket markers, `session_meta.thinking_level` on resume, agents' verbatim stored config (and the Restore default configuration path to the new prompt and file tools), and a missing `tools.call_description` reading as enabled. ([details](2026-07-24-backward-compatibility.md))
- [2026-07-24] LLM request errors surface their underlying `cause` (e.g. `terminated: other side closed (UND_ERR_SOCKET)`) instead of a bare `terminated`, visible in the Cost Center and Traces. ([details](2026-07-24-llm-request-errors.md))
- [2026-07-22] Web App: unified the form controls onto a shared Field/portal layer (now with required-field `*` markers), moved notifications to one rule (success/info → top toast, field errors inline with a red border; error prompts localized by code), generalized write confirmations (every file-writing save / uninstall / upload-overwrite now confirms via one polished dialog with tone icons — skill updates show each agent's `v_old → v_new` — and an unchanged save toasts "no changes to save"), made the Cost Center show the full error message, deduplicated the i18n copy (dead keys removed, shared `common` labels, lowercase "agent"), and reworked the sidebar session list onto category-filtered paging (default load = active rows only; the Subagents/Scheduled/Archived folders load on first open and page independently); plus the earlier separate-origin Workspace HTML previews. ([details](2026-07-22-web-app.md))
- [2026-07-22] Web App: unified the form controls onto a shared Field/portal layer (now with required-field `*` markers), moved notifications to one rule (success/info → top toast, field errors inline with a red border; error prompts localized by code), generalized write confirmations (every file-writing save / uninstall / upload-overwrite now confirms via one polished dialog with tone icons — skill updates show each agent's `v_old → v_new` — and an unchanged save toasts "no changes to save"), made the Cost Center show the full error message, deduplicated the i18n copy (dead keys removed, shared `common` labels, lowercase "agent"), and reworked the sidebar session list onto category-filtered paging (default load = active rows only; the Subagents/Scheduled/Archived folders load on first open and page independently); plus the separate-origin Workspace HTML previews (now also the in-app rendered view's path), mid-run steering or queued follow-ups from the chat input, `/model` continuing the conversation on a picked model in a new session, a per-turn thinking-level override in active sessions, a light-yellow Free badge on free models in the Models cards and the chat picker, and tool-call cards carrying model-written call descriptions with payload-showing approvals. ([details](2026-07-22-web-app.md))