From dec7565830f2a995ffa77485c8bd278ebe1b2868 Mon Sep 17 00:00:00 2001 From: Yaowei Zheng Date: Wed, 29 Jul 2026 23:59:47 +0800 Subject: [PATCH] docs(changelog): record the 2026-07-29 composer and system-prompt batch (#123) Co-authored-by: Claude Opus 5 (1M context) --- ...omposer-attachments-and-switch-commands.md | 29 +++++++++++++++++++ .../2026-07-29-default-system-prompt.md | 27 +++++++++++++++++ .../unreleased/2026-07-29-session-elapsed.md | 17 +++++++++++ changelog/unreleased/README.md | 6 ++++ 4 files changed, 79 insertions(+) create mode 100644 changelog/unreleased/2026-07-29-composer-attachments-and-switch-commands.md create mode 100644 changelog/unreleased/2026-07-29-default-system-prompt.md create mode 100644 changelog/unreleased/2026-07-29-session-elapsed.md diff --git a/changelog/unreleased/2026-07-29-composer-attachments-and-switch-commands.md b/changelog/unreleased/2026-07-29-composer-attachments-and-switch-commands.md new file mode 100644 index 0000000..7e0a646 --- /dev/null +++ b/changelog/unreleased/2026-07-29-composer-attachments-and-switch-commands.md @@ -0,0 +1,29 @@ +# Web App: file attachments in the composer, and `/agent` in place of the `@` mention + +The composer gains a second upload entry and loses its `@` trigger. A message can now carry arbitrary files the way it already carried pasted images, and the two commands that move a conversation elsewhere — `/agent` and `/model` — stage their pick as a chip instead of acting on it, so the message that opens the new conversation is the one the user meant to write. + +## File attachments + +The "+" menu holds a paperclip entry directly below image upload: any file type, up to 20 files per message, 10MB each and 12MB decoded in total (an oversize pick is refused in the browser before it is read, so nothing is uploaded to earn the rejection; the server answers 413 `file_too_large` / `too_many_files` for anything that gets past that). Selected files show as removable chips beside the image thumbnails, in the order they were picked, and — like images — an attachments-only message with no text is sendable. Goal mode clears and blocks them, since the server takes text-only goal input. + +The request body cap those limits sit under is now measured rather than read: it compared `content-length`, which a chunked request simply omits, so a body of any size passed. It counts bytes off the stream instead. + +There is no upload endpoint: a chat that is still a draft has no Session to upload to, so the file rides the task request itself as a base64 `data:` URL, exactly like a pasted image. `TaskInputPart` gains a `{type: "file", fileName, dataUrl}` variant; the server validates it without touching the filesystem, then writes the bytes into the Session scratchpad — deleted with the Session, so cleanup stays free — and appends one `[attached file: ]` line to the message text. The model opens the file with its ordinary file tools; the bytes never enter the conversation. + +That line is appended **after** the user's text rather than prefixed as a `[tag]…[/tag]` block: `[use_skills]` and the origin blocks are parsed at the start of a message, and a second leading block would break that chain. Appending has its own rule for the same reason — `[handoff_from]` and `[model_switch_from]` are only recognized when they are the *whole* message, so a message that is nothing but one of those blocks is skipped and the lines become a message of their own. Without that, a handoff carrying attachments and no text turned its own origin block into unparseable text, and the new conversation opened with the raw marker in a user bubble instead of a banner. + +The line's spelling lives in one core module shared with the existing `[attached image: …]` producer, so the two conventions cannot drift; the Web renderer pulls both back out of the body text, images becoming pictures and files an "Attached files" notice — and, like an image line, only when the path actually resolves into the Session scratchpad, so a line typed by hand stays ordinary text instead of rendering inside system chrome. + +An attachment keeps the name it was given. Only ASCII outside `[A-Za-z0-9._-]` is rewritten — a space or shell metacharacter inside a path the model is about to paste into a command is a genuine footgun — while everything above ASCII survives, minus Unicode category C (invisible controls, and the bidi overrides behind file-name spoofing): `报告 2026.pdf` lands as `报告-2026.pdf` instead of an anonymous `file.pdf`. Windows hygiene comes with it — trailing dots and spaces dropped, device names prefixed (`con.txt` → `_con.txt`) — and the stem is capped by UTF-8 **bytes** on a character boundary, since a CJK character costs three of them and a character count would let a Chinese name blow the filesystem's real limit. + +The scratchpad read endpoint changed to match: its traversal guard was a character whitelist, which a non-ASCII name could not pass. It now runs a structural check (no separators, no control characters, not a relative marker) and confirms it by resolving the path and requiring its parent to be exactly that Session's directory — strictly stronger than the charset was, since it also rejects the shapes a character class misses, such as a Windows drive-relative `C:evil.png`. Responses from that endpoint now carry `X-Content-Type-Options: nosniff`, because arbitrary user-uploaded bytes are served from the App's own origin. + +## `/agent`, and both switches staged until send + +Typing `@` no longer opens an agent picker; the leading-`@` send-time parsing is gone with it. The same handoff is now reached through `/agent`, which consumes its slash token and opens a picker built like the `/model` one — a search box matching agent id and display name, avatar rows, keyboard navigation, Escape or a click outside to dismiss. Both pickers share one panel component rather than a copy. Like `/model`, `/agent` is offered in an active session only: a draft has no conversation to hand over, and picks its Agent in the draft page's own selector. + +Both switch commands stage rather than act. Picking an agent or a model sends nothing: it pins a chip above the text body, the user keeps typing, and Enter (or Send) performs the switch — the handoff into a new chat for that Agent, or the fork of this conversation onto the picked model. Only an empty composer falls back to the default auto-message; selected skills keep riding along in a `[use_skills]` block. The two chips are mutually exclusive with each other and with goal mode, either is dropped with its × or with Backspace at the start of the text, and both are cached with the draft text they belong to — so a session switch or a reload restores the intent whole instead of leaving the text without its chip. + +A staged model fork additionally waits for the session to be idle, and says so above the composer: it branches a new session off a Trace that a run or a compaction is still appending to, so firing it mid-run would point the new model at a truncated conversation. Every send path treats a staged model switch the way it already treated the handoff target: it alone makes the composer sendable, it keeps the single action button on Send rather than Stop, and it blocks mid-run steering — a staged switch is not a message to the agent running here. + +The `[handoff_from]` origin block told the receiving model that it had been "@-mentioned". The tag and its field lines are a persisted format that old Traces still parse through, so only the prose changed: it is now trigger-agnostic and survives the composer changing how a handoff is started. diff --git a/changelog/unreleased/2026-07-29-default-system-prompt.md b/changelog/unreleased/2026-07-29-default-system-prompt.md new file mode 100644 index 0000000..7e6c818 --- /dev/null +++ b/changelog/unreleased/2026-07-29-default-system-prompt.md @@ -0,0 +1,27 @@ +# Core: a tighter default system prompt, a reply-language rule, and a shared tooling directory + +The built-in Agent template is re-read by the model on every turn, and had grown wordy in exactly the sections that repeat most. It loses about a tenth of its words (1087 → 969) with no rule going with them, and two behaviours the prompt never stated are now in it. + +## What the trim changed + +Personality pins the reply language to the user's own — the tool schema already demanded that of every call description, so the two agreed in practice while only one of them said so. It scopes to prose: code, identifiers and commit messages keep their own conventions. The process/port constraint collapses from two sentences into one, with the same guarantees: don't kill what you didn't start (PenguinHarness's own services included) unless asked, don't take a service port, pick another free port when the one you want is busy. + +API auth handling stops being split across two sections. Constraints carried "retry at most once" and Stop rules carried a fourth rule for what to do afterwards; it is now stated once, as the special case of the existing "an error you cannot resolve" rule — retry at most once, then stop calling tools and ask the user to update the key in the Agent's vault or the model settings outside the chat. The facts that made the old wording long are intact: the secret is never pasted into the conversation, and a new key only takes effect in the next conversation, so further retries cannot succeed. + +System markers lose about half their words while keeping all four markers and their instructions. File system goes from eight bullets to seven and 262 words to 243: two merges — the workspace-relative-path convention into the `CWD` bullet, another agent's state into the description of your own — against one bullet gained for shared tooling. Suggested workflows already recommended dispatching independent subtasks in parallel; it now says out loud that this is the fast way through a large task. Its Playwright/curl line folds into the Tool use bullet, which keeps both halves the two lines used to carry separately: prefer Playwright when it is installed, otherwise `curl`. + +## Tooling installs once + +File system gains a convention, and the Skills section loses the "There is no skill tool" filler that explained an absence. It sits in File system rather than Skills because it governs any task that installs a tool, not only a skill run, and it is written as the choice the model actually faces: install into the project's own environment when it has one; otherwise keep the reusable ones — Python virtualenvs, model and package caches — under `/agents//shared_env//` and reuse them across Sessions. + +Node is called out separately, with pnpm preferred **in the project itself** — both halves matter. Its shared store keeps repeated installs from duplicating on disk, and the location is what keeps the install resolvable: `node_modules` resolves from the project upward, so a package placed under the agent directory would be invisible to code running in `CWD`. + +`shared_env/` is a prompt-level convention, not a path the code creates: `paths.ts` is untouched, and Agent State snapshots still package `agent_state/` alone, so a virtualenv can never bloat an export. The data-layout tree in the Sessions & Traces documentation lists the directory with that caveat spelled out, and the per-Agent layout in `01-PRINCIPLES.md` records it too. + +## The App Data Dir bullet stops contradicting itself + +The same section's App Data Dir bullet said the directory "is NOT the task's directory" and that deliverables must never be left there. A temporary Workspace is `/agents//workspaces/tmp-<8hex>` — inside that very tree, and `CWD` for any task that is given one — so the rule broke exactly where it was needed. It now carries the input rule alone (the App Data Dir's contents were not supplied by the user and are never task input) and says plainly that `CWD` may sit inside the tree, that one folder being the task's and the rest not. Where deliverables go was already stated positively by the scratchpad bullet below it, which holds wherever `CWD` points. + +## Existing Agents + +Nothing migrates and nothing is rewritten. An Agent always runs with its on-disk `agent_state/system_config.yaml` verbatim, so this reaches **newly created** Agents only. An existing Agent adopts it through the settings page's *Restore default configuration* action, which overwrites the whole configuration — custom system prompt, tool list, model and compaction settings, MCP Servers — keeping only `name`, `description` and `version`. diff --git a/changelog/unreleased/2026-07-29-session-elapsed.md b/changelog/unreleased/2026-07-29-session-elapsed.md new file mode 100644 index 0000000..7b6f23d --- /dev/null +++ b/changelog/unreleased/2026-07-29-session-elapsed.md @@ -0,0 +1,17 @@ +# Web App: the Session's elapsed time is taken from Trace timestamps, so a reload stops restarting it + +The elapsed chip in the chat header restarted from zero whenever a running Session was reloaded, and the figure it eventually settled on depended on the browser clock. Both came from the same root: the running Task was the one part of the number not derived from the Trace. + +## Reloading mid-run resumes the chip instead of restarting it + +The chip renders the settled cross-Task total plus the running Task's wall clock so far, and that second half ticks from a local-clock anchor. A history rebuild feeds every replayed message the same "now", so the anchor was stamped with the instant the page loaded: a Task that had already been running for five minutes was treated as having just started, and the chip dropped back to the settled total and climbed from zero again. + +The anchor is now back-dated, once the replay finishes, by the elapsed already behind the Task, so the ticking value resumes where it left off. That elapsed is measured entirely in server time and only then applied to the local clock, which keeps a client/server clock offset out of the result and lets the chip keep ticking smoothly. A live stream is unaffected: it pushes one message at a time with the real current clock, where the elapsed is still zero when the Task opens. + +It is measured against the server's own clock at read time, taken from the messages response's HTTP `Date` header, rather than against the Trace's last entry. That distinction is what makes an event still in flight count: while a tool is executing, a Request is streaming or a compaction is running, nothing has been appended to the Trace since it began, so a reload measuring only what the Trace records would show none of the time that event has already taken. The Trace's own span remains the floor — it is the fallback when no `Date` header comes back, and a cached or rewritten header can only be older than the true present, so it can under-report but never overshoot. The header carries whole seconds, which is finer than a chip that ticks in whole seconds can show. + +## A settled round reads the same live and replayed + +A round's duration is the span from its first message to its last non-compaction `request_end`. One case escaped that: a round with no `request_end` at all — interrupted before its first Request even ran — was measured with the local clock when it happened to be watched live, and with the message span when it was replayed later. The same round therefore showed one number before a refresh and a different one after, with idle-detection and mid-join latency folded into the first. It now settles to its message span on both paths, which the interrupting abort's own timestamp still bounds. + +The local clock now reaches the elapsed figure in exactly one place — animating the Task in flight — and never a settled one. The meaning of the statistic is unchanged: still the sum of each Task's wall clock. Compaction handling is unchanged too, with a mid-round compaction inside the span and a post-round one outside it, and the chat page keeps its deliberate difference from the Trace page's total. diff --git a/changelog/unreleased/README.md b/changelog/unreleased/README.md index 2e75f5c..0ce97cf 100644 --- a/changelog/unreleased/README.md +++ b/changelog/unreleased/README.md @@ -8,6 +8,12 @@ Changes since v0.1.4. The version number is assigned at release, when this folde - [2026-07-29] Core and tooling: `PORT` / `HOST` and the internal CLI plumbing no longer reach commands the Agent runs, so a dev server started by `exec_command` picks its own port instead of binding the harness's; and the development backend moves to 7368, so `pnpm dev` no longer collides with an installed `penguin web` — or, quieter and worse, proxies to it. ([details](2026-07-29-harness-env-and-dev-ports.md)) +- [2026-07-29] Web App: the composer can attach files of any type — written to the Session scratchpad and handed to the model as `[attached file: ]` lines, keeping non-ASCII names — and the `@` mention becomes an `/agent` command; both switch commands are session-only and stage their pick as a chip, cached with the text, until Enter sends. ([details](2026-07-29-composer-attachments-and-switch-commands.md)) + +- [2026-07-29] Core: the default system prompt is about a tenth shorter (1087 → 969 words), with the trim concentrated in the sections re-read every turn. It now pins replies to the user's language, states the API-key stop rule once, points at parallel subagents for large tasks, and has tooling installed once into a shared per-Agent `shared_env/` directory while project dependencies stay in the project. The App Data Dir rule also stops forbidding deliverables in a tree that a temporary Workspace lives inside. Existing Agents keep their own prompt. ([details](2026-07-29-default-system-prompt.md)) + +- [2026-07-29] Web App: the chat header's elapsed chip no longer restarts from zero when a running Session is reloaded and now counts an event still in flight, and a round interrupted before its first Request settles to the same figure whether it was watched live or replayed — the running Task is anchored to the server's own clock instead of the browser's. ([details](2026-07-29-session-elapsed.md)) + - [2026-07-28] Web App: model configuration stops inviting the browser's saved login — every form control opts out of autofill unless it declares a real credential role, and a secret field says `new-password`, the only value Chrome and Safari honor on a password box; a closed docked side panel no longer paints its 1px divider next to the open one, which read as a second, empty panel beside the Workspace files; a provider-signature message carrying no text stops drawing an empty reply bubble after a thinking segment; Benchmark costs now follow the display currency selected in settings; the tool card states a call's outcome once instead of twice; and no page can grow the document into a second scrollbar that drags the whole shell. ([details](2026-07-28-web-app.md)) - [2026-07-28] Web App and CLI: duration and byte abbreviations now carry into the next unit instead of printing `1m60s` or `1024KB` — both helpers rounded the value they displayed but chose the unit (or the minute/second split) from the raw input. ([details](2026-07-28-duration-and-byte-formatting.md))