From a4e7820efe28b1d738fca27908cb67e3b01ea384 Mon Sep 17 00:00:00 2001 From: Yaowei Zheng Date: Tue, 28 Jul 2026 00:31:45 +0800 Subject: [PATCH] docs(changelog): record the 2026-07-26/27 batch (#80) Co-authored-by: Claude Fable 5 --- changelog/unreleased/2026-07-26-core.md | 11 ++++++++ changelog/unreleased/2026-07-26-web-app.md | 27 +++++++++++++++++++ .../unreleased/2026-07-26-windows-support.md | 19 +++++++++++++ .../2026-07-27-backward-compatibility.md | 15 +++++++++++ .../2026-07-27-llm-request-errors.md | 9 +++++++ changelog/unreleased/2026-07-27-tooling.md | 3 +++ changelog/unreleased/README.md | 12 +++++++++ 7 files changed, 96 insertions(+) create mode 100644 changelog/unreleased/2026-07-26-core.md create mode 100644 changelog/unreleased/2026-07-26-web-app.md create mode 100644 changelog/unreleased/2026-07-26-windows-support.md create mode 100644 changelog/unreleased/2026-07-27-backward-compatibility.md create mode 100644 changelog/unreleased/2026-07-27-llm-request-errors.md create mode 100644 changelog/unreleased/2026-07-27-tooling.md diff --git a/changelog/unreleased/2026-07-26-core.md b/changelog/unreleased/2026-07-26-core.md new file mode 100644 index 0000000..2632a03 --- /dev/null +++ b/changelog/unreleased/2026-07-26-core.md @@ -0,0 +1,11 @@ +# Core: resume ignores a legacy thinking level; empty compaction summaries no longer count as success + +Since the thinking level became a per-turn run parameter (0.1.2), `session_meta` no longer records one — but resume still honored a `thinking_level` field found in an old Trace's meta, restoring it as the resumed Session's default. That compatibility read is removed: resume now always takes the level from the Agent's current config, and a legacy field is ignored outright. The one behavioral consequence is that a resumed legacy subagent session no longer keeps the level it inherited at spawn time — it follows the resuming Agent's configuration like everything else. Spawn-time inheritance in live runs, the per-turn picker, and what meta records are all unchanged. + +## An empty compaction summary is a failure, not a result + +A compaction request that came back with only thinking or a tool call — no text — used to be judged `completed`, wrapping an empty `` and discarding the old context: the one thing compaction must never do, since the Agent loses its task state (issue #83). Three fixes close it. The compaction request keeps its configuration — tools included — **byte-identical to ordinary requests** (the same frozen config object, pinned by a test): omitting or altering the tool list would change the request prefix and blow the provider's prompt cache at exactly the moment the context is largest, multiplying the compaction's cost tens of times. A `completed` response whose extracted summary trims to empty, or that answers with a tool call instead, is rejected: any committed tool calls are first **repaired** — one synthesized failed `tool_call_output` per call, in the ordinary tool-failure shape, written to the Trace and prepended to the resent input, so the provider-side pairing stays intact (and if the compaction is abandoned, the repairs move into the carry-over so the live LLM object is never left with a dangling `tool_use`) — then the request retries, with a dedicated cap of 5 rejections separate from the shared transport-reconnect budget; on exhaustion the compaction fails via `compaction_end(failed)` with the original context and Trace file retained. And resume no longer fabricates an empty summary from a textless compaction closure: such a closure is void and the original context replays, so Traces already bitten by the bug recover their task state too. + +## Mid-task compaction stops re-sending what AgentHub already has + +The empty-summary fix left one bounded hazard behind (issue #85): when auto-compaction fires mid-task, the pending turn input rides along in the compaction request — and once an attempt is **committed** into AgentHub's history, resending that input on a retry (or after giving up) hands strict providers duplicate `tool_result`s to reject. The rule is now a clean binary on one fact: if any compaction attempt committed, the carry is consumed — retries resend only the pairing repairs plus the prompt, and after abandonment the next request continues from committed state without re-sending the turn's outputs; if nothing committed, the previous carry-over is restored verbatim (object-identical). An abort inside the compaction window merges into the pending state instead of overwriting it, and replay reconstructs the same absorbed state after a restart, so live and resumed sessions agree. diff --git a/changelog/unreleased/2026-07-26-web-app.md b/changelog/unreleased/2026-07-26-web-app.md new file mode 100644 index 0000000..ce49b32 --- /dev/null +++ b/changelog/unreleased/2026-07-26-web-app.md @@ -0,0 +1,27 @@ +# Web App: subagent panel with a live call graph, streaming that survives refresh, live header stats, one-line mobile rows, Trace import/export, and in-app version/update + +Six chat-and-observability changes land together: subagent conversations move out of the message flow into a dockable panel with a call graph, an in-progress reply now survives a page refresh, the header statistics tick live, running-state rows stop growing on phones, Trace files can be moved across deployments, and the app can tell you what version it runs and update itself. + +## Subagent conversations move to a side panel with a call graph + +A `run_subagent` call used to inline the child session's entire conversation into the parent's message flow, nested cards within cards. The child conversation now lives in a dedicated **agents panel** that docks on the right exactly like the Workspace files panel (toolbar toggle, drag-to-resize with width memory, bottom sheet on narrow viewports, and the two docks are mutually exclusive so they never crush the chat column). In the message flow each child leaves only a full-width bar row styled like the other collapsed stream rows — agent avatar, resolved agent name (from the child's own `session_meta`, falling back to the spawn arguments, the session list, then a short id), a spinner while it runs, and an amber dot whenever an approval is pending anywhere in its subtree; the dot is part of the accessible name, and the toolbar button carries the same badge, so a nested approval stays discoverable with the panel closed. The panel's top shows a **call graph**: one node per participating agent (avatar + name + run-state dot + elapsed time — ticking wall clock while it runs, frozen from message timestamps when done, so a reloaded page shows the same durations as a live one; the main session is the root), edges for who spawned whom, repolled children deduplicated. Clicking a node switches the conversation shown below — rendered by the same machinery as the main stream, so nested tool cards stream live and approval buttons work from inside the panel. The graph shows the latest Task by default and resets to it on every Session switch; clicking a row from an older turn pins that turn's Task and shows its historical spawn tree instead. Opening either docked panel (any path, toolbar included) closes the other, so the two can never crowd the chat column together. The panel's visibility is **task-scoped**: every new Task starts with it closed (an open panel from an unrelated previous task never lingers), it auto-opens once per task when that task first spawns a subagent (docked layouts only), and a manual open or close always wins for the rest of the task; entering a session starts closed too — this per-task rule is the one place the panel deliberately diverges from the files panel's persistence, and a goal-mode iteration counts as a task boundary. The child conversation also finally shows its own user side live: the subagent runner forwards the spawn prompt (and follow-up inputs) through the origin-tagged stream, so the panel renders the child's user messages identically live and after a reload. + +## An in-progress reply survives refresh and reconnect + +Refreshing mid-Task used to lose whatever was streaming — a half-written reply, thinking, or a long tool call's partial output — until the message completed, because history is read from the Trace (complete messages only) and a fresh SSE connection replays nothing. The server now keeps a **live tail** per running session: an aggregate of every open fragment, keyed by origin chain and fragment identity, updated in the same tick a message is published and capped so a runaway output keeps only its tail. The messages endpoint returns it as `live: { cursor, fragments }`, captured atomically before the trace read begins. The client seeds the fragments as synthetic `partial_* start` events at the cursor position, drops its buffered partials at or before the cursor (their content is already inside the fragments), and lets complete messages keep flowing through the existing overlap dedup — deliberately never cursor-dropping them, because a publish can precede its trace append landing on disk. The effect: after a refresh, close-and-reopen, or a resync, the already-streamed prefix is back immediately and keeps growing, for the main session and for subagents alike. The end-to-end test reloads mid-command and mid-text and asserts the final transcript is byte-identical to a never-reloaded run. + +## Header statistics update live + +The three chips in the chat header (session Tokens including subagents, cost, elapsed) effectively refreshed only when a Task finished. Elapsed now ticks once per second while a Task runs (settled cross-Task total plus the running Task's wall clock, folded together by the same pure helper the tests pin down, so the running addition can never double-count across the Task boundary). Cost shows the last server-recorded value plus a live estimate converted from the running Task's usage buckets at the current Model's pricing — subagents on other models make the estimate slightly approximate mid-Task, and the idle refetch reconciles to the recorded value; without pricing the live addition is skipped and the `*` uncosted marker keeps its meaning. Tokens already advanced per completed request and are untouched. On a brand-new session the cost chip stays hidden until the first usage lands rather than flashing a formatted zero the instant a Task starts. + +## Running-state rows stay one line on mobile + +Below 640px, running-state UI no longer wraps or grows: the reply's stats footer — whose fixed height made wrapped chips paint over the content below, the one genuine bug — stays one line, and both message footers (reply and user bubble) are now **always visible on phones**: their hover-reveal never fired under touch, a gap that predated this PR. To fit 390px without scrolling, the TPS chip hides below `sm` and the numeric chips use compact variants there (money at two significant digits so a nonzero cost never reads `$0.00`; whole seconds from 10s up); everything still sits in a hidden-scrollbar scroll row as a fallback, with the copy button pinned at the end. The header shows a pulsing dot instead of the "Running" text and an icon-only Workspace button; the work group hides the step count and collapses the awaiting-approval pill to a labeled amber dot. On tool cards the textual decision indicator is **gone at every breakpoint** — the status icon alone carries "Approved · auto/manual" or "Denied · …" as its accessible name (exactly one source of truth, and a denied card no longer also says "aborted"). Action buttons follow the opposite rule: Allow/Deny stay **textual everywhere** — indicators may be iconic, buttons must be words — and while a call awaits approval on a phone, its full command wraps into view (the one sanctioned exception to the one-line rule: you read what you approve), collapsing back after the decision. Desktop rendering is pixel-unchanged. + +## Trace files can be exported and imported + +The Traces page can now move a trajectory across deployments: any member downloads a Trace file verbatim (`application/x-ndjson` attachment), and a Project owner imports one under any agent via an upload button on the agent's row (14MB cap, mirroring the Agent snapshot import). An imported file must be valid Trace JSONL whose first record is a `session_meta` with a filename-safe `session_id` — re-validated in the service right where the path is built. A session id that already exists under the target agent is rejected outright (409): index _N_ of a session means "context segment _N_", so absorbing an unrelated file there would splice it into the existing transcript, silently become the resume point, and race a live session's writer. An accepted import is therefore always file 001 of a new session, written with `wx` so even concurrent imports cannot clobber each other, filed under its first record's local date (matching the Trace writer's convention), and oversized picks are refused client-side before any upload. The tree refreshes and selects the imported file immediately; imported sessions are browsable like any other directory-discovered ("unmanaged") session. + +## The app knows its version and can update itself + +The sidebar user menu gains a **Check for updates** row directly below Change password, with the running version inline on its right (`v0.1.3`, no product-name prefix) and a superscript "New version available" badge when the check finds one; the row's tooltip carries a human-readable `Last updated Jul 26` (`最近更新日期 7 月 26 日`) taken **offline-only** from the release-stamped `BUILD_DATE` — every future release stores its date at publish time in both the tarball and npm jobs, and no network lookup is ever made for it. Clicking the row forces a fresh check (`?force=1` presses through the server's TTL cache — user-initiated and low-volume — and the outcome is re-cached), with "you're on the latest" and failures surfacing through the shared toast. The new-chat draft page carries the same quiet line under its subtitle, with the badge as the release link. Opening the menu also triggers an update check, the server's first and only outbound internet call: the latest GitHub Release is fetched with the same request shape as `penguin update`, cached server-side (an hour on success, ten minutes on failure), strictly fail-soft — a failed lookup reports "no update known", never an error — and `PENGUIN_UPDATE_CHECK=off` disables it entirely. When a newer release exists, the avatar gains a dot and the menu a release-notes link; admins additionally get **Update now**, which re-runs the CLI's own updater (`penguin server|web` exports its entry as `PENGUIN_CLI_ENTRY`; the endpoint spawns `node update --yes` with a fixed argument list — nothing from the request reaches the command). The CLI's install-kind detection and refusals pass through verbatim, and a successful run says the one thing that still needs a human: restart the service to load the new code. diff --git a/changelog/unreleased/2026-07-26-windows-support.md b/changelog/unreleased/2026-07-26-windows-support.md new file mode 100644 index 0000000..dcd0ec2 --- /dev/null +++ b/changelog/unreleased/2026-07-26-windows-support.md @@ -0,0 +1,19 @@ +# Windows: a real install story — shell selection, install.ps1, and a win-x64 release package + +The audit found the dependency graph already clean for Windows (`node:sqlite`, pure-JS `tar`, no native modules) and `npm install -g @prismshadow/penguin-cli` already yields a working `penguin` — but every `exec_command` died with `spawn bash ENOENT`, because the command session hardcoded `spawn("bash", ["-lc", cmd])`. That one line was the product-level blocker; the rest was packaging and honesty. + +## The agent actually works: shell selection + +Command sessions now resolve their shell once per process: POSIX stays `bash -lc` bit for bit; on Windows the resolver probes PATH for Git-Bash first (best compatibility with the POSIX-oriented skill ecosystem; a `bash` that resolves into the Windows system directory is rejected — that's the WSL launcher, a different filesystem view entirely), then `pwsh`, then `powershell` (`-NoLogo -NoProfile -Command`). `PENGUIN_SHELL` overrides everywhere, with the argument shape inferred from the basename. The chosen shell is announced to the model through a new `Shell:` line in the session environment, so it writes commands in the syntax it actually has. Existing Agents whose `system_config.yaml` predates the `{{SHELL}}` placeholder get the same line through a narrow assembly-time fallback (win32 only, in-memory, no migration; see the backward-compatibility entry) — without it their models would keep emitting bash into PowerShell forever. Termination follows the platform: POSIX keeps process-group signal escalation; Windows kills the whole tree via `taskkill /pid /t /f` (no console signal delivery to piped children, so `input_command`'s Ctrl-C degrades to a hard kill — documented rather than pretended away). + +## A sandbox gap the new CI caught + +Windows has no `O_NOFOLLOW`, and the workspace upload path's `(O_NOFOLLOW ?? 0)` silently erased the final-segment symlink guard there — a practical "preset a symlink, overwrite a file outside the Workspace by upload" escape. Uploads on win32 now refuse a final-segment symlink via `lstat` (best-effort against the race; POSIX keeps the atomic open-time guarantee). + +## Install and release + +`install.ps1` at the repo root mirrors `install.sh`: `PENGUIN_VERSION` / `PENGUIN_INSTALL_DIR` knobs, SHA256 verification, a staged rename-then-delete swap that never touches `data\`, generated `penguin.cmd` / `penguin.ps1` shims (both CRLF), and a user-`Path` update that reads and writes the registry value raw with its kind preserved — the naive `[Environment]` API expands `REG_EXPAND_SZ` on read and writes back `REG_SZ`, irreversibly hard-coding a user's `%USERPROFILE%`-style entries. The stable URL `https://penguin.ooo/install.ps1` is a forwarder that downloads fully before executing — a truncated stream can't half-install. The release workflow gains `penguin-win32-x64.zip` (official Windows Node bundled, `node.exe` at the archive root, CRLF launchers) plus its checksum, and uploads `install.ps1` alongside `install.sh`. Upgrading is re-running the installer; in-place `penguin update` still refuses on Windows, and the docs say so. The one-liner is `irm https://penguin.ooo/install.ps1 | iex`, with `npm install -g` (Node ≥ 24) as the script-free alternative; the landing page shows both install rows and the README roadmap checks off Windows support. + +## Keeping it true: CI and the long tail + +A `ci-windows` job (windows-latest, full build/typecheck/test plus a PowerShell parse gate) now runs beside the required Ubuntu job; getting it green surfaced four genuine Windows findings (file-lock `EBUSY` on cleanup, a wrong timeout-test premise, the symlink gap above, path-separator test assumptions), all fixed. `.gitattributes` pins LF (with CRLF for `.ps1`) so a Windows checkout can't fail the Prettier gate or corrupt shell scripts. Cross-platform env handling replaces the POSIX-only `VAR=x` npm-script prefixes with a small runner script, `penguin config lang` refuses cleanly on win32 (pointing at `setx PENGUIN_LANG`), and the installer was functionally exercised with real PowerShell against a local fake release: fresh install, upgrade preserving `data\`, checksum-mismatch abort, and the full `irm | iex` chain. Known and documented limits: no 0600 semantics for config/vault files (NTFS ACLs apply), x64 package only (ARM64 runs via emulation), execution-policy notes for the `.ps1` shim, no SIGTERM-driven graceful shutdown, and — now stated in the install and tools docs, not just in source — Ctrl-C in `input_command` kills the whole command-session tree on Windows instead of interrupting the foreground command. Keeping the Windows CI honest also meant giving win32 test runs a longer vitest timeout and platform-gating one spawn-timing assertion; each flake it caught was a real timing sensitivity, not a product bug. diff --git a/changelog/unreleased/2026-07-27-backward-compatibility.md b/changelog/unreleased/2026-07-27-backward-compatibility.md new file mode 100644 index 0000000..2f6b0e6 --- /dev/null +++ b/changelog/unreleased/2026-07-27-backward-compatibility.md @@ -0,0 +1,15 @@ +# Backward compatibility in this batch + +Per the repo rule, every compatibility decision of the batch is recorded here once; the feature entries reference this file instead of re-telling it. + +## Legacy `thinking_level` in old Traces is ignored on resume (owner-directed) + +Old Traces recorded a `thinking_level` in their `session_meta`; since 0.1.2 the field is no longer written, and resume used to keep honoring it when present. The owner chose explicit incompatibility over continued support: resume now ignores the field entirely and always reads the Agent's current config. **Old shape tolerated:** the field may appear in old meta JSON and is skipped without error (the Trace itself stays fully readable). **Scope:** `resumeSession` only; live behavior and the per-turn picker were never affected. **User action:** none — the one visible consequence is that a resumed legacy subagent session follows the resuming Agent's configured level instead of its spawn-time inherited one. **Removal:** nothing to remove; no compat code was retained beyond the tolerant skip. + +## Existing Agents without `{{SHELL}}` get the Shell line via an assembly-time fallback + +The `Shell:` environment line ships in the default `system_prompt` template through a new `{{SHELL}}` placeholder — but `system_config.yaml` is baked at Agent creation and never auto-upgraded, so every Agent created before this release lacks the placeholder, and on Windows its model would never learn which shell `exec_command` speaks. Rather than a disk migration, prompt assembly applies a narrow fallback: **on win32 only**, when the template contains no `{{SHELL}}` and the output carries no `Shell:` line already, the line is appended to the Environment block in memory. **Old shape tolerated:** pre-`{{SHELL}}` templates (and templates that hardcode their own `Shell:` line, which win). **Scope:** prompt assembly on Windows; POSIX output is byte-identical, no file is rewritten. **User action:** none; resetting an Agent's config to defaults also picks up the placeholder and makes the fallback moot for that Agent. **Removal:** delete the fallback once pre-`{{SHELL}}` Agent configs are no longer expected in the wild (tracked by the comment at the fallback site). + +## Additive protocol fields need no handling + +The `StopReason` enum's sixth value `auth` (old readers fall through like `failed`; old Traces never contain it), `request_end`'s `message` and `retry_in_ms`, the `credentials_updated` server event, and `updatedAt` on the models response are all additive: old Traces simply lack them, old readers ignore them, and nothing interprets legacy data differently. Recorded here only to state that the check was made; there is no compat behavior to retire. diff --git a/changelog/unreleased/2026-07-27-llm-request-errors.md b/changelog/unreleased/2026-07-27-llm-request-errors.md new file mode 100644 index 0000000..5eac912 --- /dev/null +++ b/changelog/unreleased/2026-07-27-llm-request-errors.md @@ -0,0 +1,9 @@ +# Core and Web App: transport and quota errors reconnect with a visible countdown, auth failures lock the Session recoverably + +Two field failures used to abort the turn outright with `[Aborted]: llm request error: …` — a dropped connection (`terminated: other side closed (UND_ERR_SOCKET)`) and a gateway quota rejection (`403 … no active subscription (insufficient_user_quota)`). Neither is a verdict on the Session: a socket drop is just the network, and a quota error heals when the balance tops up. Both now go through the engine's in-run reconnect (the `[turn_retried]` flow that re-sends the turn's input with the already-produced content attached). + +The classifier got honest about error shapes: retryability probes the `cause` chain (Node's `fetch` wraps the real transport failure — `UND_ERR_SOCKET` and friends live on `cause`) and the parsed provider body in its OpenAI and Anthropic SDK nestings; a bare `terminated` only counts alongside transport vocabulary, so "terminated by content filter" copy does not retry. The quota carve-out stays tight — 402/403 with `insufficient_user_quota` / `insufficient_quota` codes or a message naming quota/subscription — and **authentication signals are checked first**: a definitive auth code can never be swallowed by the quota keyword heuristic into the retry path. + +Retries follow one shared exponential ladder, `min(250ms × 2^(n−1), 30s)`, up to **8 reconnects (~62s of total patience)** — the first two attempts are exactly as fast as before, so a transport blip still recovers in under a second, while the tail is long enough for a quota window to actually close. Compaction requests keep their own cap of 3 (~1.75s): a failed compaction already preserves the original context and retries on the next trigger, so failing fast beats stalling the session. While a longer wait runs, the retry line shows a **live countdown** (whole seconds, anchored on the client clock, driven by the new additive `retry_in_ms` on `request_end` — computed by the same formula as the actual sleep) with two inline controls: **Retry now** (`POST /api/sessions/:id/retry-now`, skips the remaining wait without consuming an attempt) and **Give up** (the ordinary abort). `request_end` also carries the failure detail (`message`), so the Cost center's errors panel records the real cause — the 403 body, not a generic "timeout" — even when retries follow. + +Authentication errors don't retry; they surface as a sixth `StopReason` value — `auth` on the outcome and on the streamed `request_end.status` (the abort event stays reason-only; no new fields) — and the composer locks — **recoverably**. The Session pins its model *reference* at creation, but credentials are read from the current Project config on load, so fixing the key on the Models page is the actual remedy: the models update now invalidates every cached runtime in the Project (previously only vault edits did — a fixed key could sit unused for up to the 30-minute idle sweep) and broadcasts a `credentials_updated` event, unlocking open tabs instantly. Across reloads the lock is a time gate keyed on the last main-session `request_end("auth")` timestamp — dead only while that failure is newer than the credentials' file mtime — so a fixed key stays fixed, a still-wrong key re-locks on its next failure, and any later successful request clears the state too. The notice's primary action opens the Models page, with Retry and New Session as escapes, and the stuck draft stays selectable. A subagent's auth failure never locks the parent. End-to-end proof: a quota 403 retries through a visible, decreasing countdown (with Retry-now firing early on click), and a 401 locks the composer, unlocks the moment the key is updated via the API with no click required, and stays unlocked after a reload. diff --git a/changelog/unreleased/2026-07-27-tooling.md b/changelog/unreleased/2026-07-27-tooling.md new file mode 100644 index 0000000..ca65e3b --- /dev/null +++ b/changelog/unreleased/2026-07-27-tooling.md @@ -0,0 +1,3 @@ +# Tooling: keeping the Windows CI honest + +Three follow-ups hardened the new `ci-windows` job against the runners' pathologies, each backed by evidence rather than a blind retry: the server test suite's cleanup now retries `ENOTEMPTY`/`EBUSY` removals and runs under platform-aware deadlines (one failed cleanup used to cascade into "database is not open" across later files); core's win32 test deadline moved to 120s after instrumented runs showed a single cold Git-Bash spawn taking 35.7s on a degraded runner — with an amplification experiment proving latency flows through the engine 1:1, no race — plus a platform-neutral regression test pinning that a slow tool delays a run only by its own latency; and two cross-PR combination breaks caught by main's CI (a test fake missing a newly required member, platform-naive path expectations) were fixed the hour they appeared. diff --git a/changelog/unreleased/README.md b/changelog/unreleased/README.md index 8590d05..62720ac 100644 --- a/changelog/unreleased/README.md +++ b/changelog/unreleased/README.md @@ -2,4 +2,16 @@ Changes since v0.1.2. The version number is assigned at release, when this folder is renamed. +- [2026-07-27] Core and Web App: transport drops (`UND_ERR_SOCKET`-class) and provider quota errors (`insufficient_user_quota`) now reconnect on an exponential ladder (up to 8 attempts ≈62s) with a live countdown plus Retry-now/Give-up controls, real failure details reach the Cost center, and authentication failures lock the composer recoverably — updating the key on the Models page invalidates cached runtimes and unlocks open Sessions instantly. ([details](2026-07-27-llm-request-errors.md)) + +- [2026-07-26] Web App: subagent conversations move into a docked agents panel with a live call graph of the latest Task (compact chips remain in the stream), an in-progress reply now survives refresh/reconnect via a server-kept live tail, the header statistics (Tokens/cost/elapsed) tick live while a Task runs, running-state rows stay one line on mobile with icon-only colored approval buttons, Trace files can be exported and imported from the Traces page, and the sidebar user menu shows the running version with an update reminder and an admin-run self-update. ([details](2026-07-26-web-app.md)) + +- [2026-07-26] Windows: the command session picks a real shell (Git-Bash → pwsh → powershell, `PENGUIN_SHELL` override, the choice announced to the model), a win32 symlink-upload sandbox gap is closed, and the release ships `install.ps1` (`irm https://penguin.ooo/install.ps1 | iex`) plus a `penguin-win32-x64.zip` package with a bundled Node runtime, verified by a new windows-latest CI job running the full test suite. ([details](2026-07-26-windows-support.md)) + +- [2026-07-26] Core: resuming a session now ignores a legacy Trace's recorded `thinking_level`; an empty compaction summary is no longer a success — compaction keeps its tools byte-identical (protecting the prompt-cache prefix), rejects empty or tool-calling responses with paired repair outputs and up to 5 retries, and resume replays the original context instead of a fabricated empty summary; and a committed mid-task compaction attempt now absorbs the turn's pending input, so retries and follow-ups never re-send what AgentHub already holds. ([details](2026-07-26-core.md)) + +- [2026-07-27] Tooling: Windows CI hardening — retrying test cleanups, evidence-sized deadlines, and same-hour fixes for two cross-PR combination breaks. ([details](2026-07-27-tooling.md)) + +- [2026-07-27] Backward compatibility: the batch's compat decisions in one place — legacy `thinking_level` ignored on resume (owner-directed, nothing retained), the win32 assembly-time `Shell:` line fallback for pre-`{{SHELL}}` Agent configs (with its removal condition), and the additive-only protocol fields. ([details](2026-07-27-backward-compatibility.md)) + - [2026-07-26] Goal mode: state an objective with an optional token budget and the system loops Tasks on one Session via `session.run(input, { goal })` — a `[goal]` round protocol embedding GOAL.yaml, budget wrap-up round, runaway safeguards (a cut-off round is terminal; 100-round backstop) — surfaced as the CLI's `/goal` and `run --goal`, a `goal` field on the tasks API, and the Web composer's new "+" menu. ([details](2026-07-26-goal-mode.md))