diff --git a/changelog/unreleased/2026-07-27-input-images.md b/changelog/unreleased/2026-07-27-input-images.md new file mode 100644 index 0000000..09803e5 --- /dev/null +++ b/changelog/unreleased/2026-07-27-input-images.md @@ -0,0 +1,21 @@ +# Images reach every input: mid-run steering, goal objectives, and no dead action button + +Steering — the message you send to an agent that is already running, delivered between turns as a `[user_steering]` user message — used to be text and nothing else. That was a limitation of the channel, not of the model: the very moment an image is most useful is mid-run, when the agent has just produced the wrong layout and the fastest correction is a screenshot. Steering now carries images the way a prompt does, and an image with no caption is a complete steering message on its own. + +Core queues the images with the text — `session.steer(input)` takes an OmniMessage list, the same shape `session.run` takes a prompt in — and delivers them as ordinary user image messages directly behind the `[user_steering]` block — the same shape a prompt uses, so the LLM client, the Trace, and replay all already know what to do with them. On a model without vision they fold into `[attached image: ]` lines **inside** the block, exactly as a prompt's images fold, because the block has to remain the whole text: a line appended after the closing tag would cost the message its steering identity and every render layer would read it as a new Task. A scratchpad that cannot be written to ends the run — the picture usually arrives *because* the run is going the wrong way, so carrying on without it would spend the rest of the task heading further that way. + +The renderers follow the same grouping rule they already apply to a prompt's images — everything up to the next non-image message belongs to the message in front of it. In the web app the images land inside the steering chip instead of becoming bubbles of their own, and neither the chat stream nor the Trace timeline opens a new Task for them; an images-only prompt sent later is still a genuine new turn. + +## A goal can be stated with a picture + +"Make the page match this mockup" is a goal, and a screenshot states it better than a paragraph — but a goal objective used to reject images outright. It accepts them now, with one deliberate difference from every other input: they are folded to scratchpad paths **whatever the model's vision**. A goal objective is re-injected as the text of every round's `[goal]` block, so an image cannot ride along as an image. Sending it in round 1 alone would leave every later round pointing at something compaction has since removed — a silent failure, since the objective still *reads* correct. As a path line it survives every round and every compaction, and the model spends tokens on it only when it actually looks. Text is still required: a picture alone states no objective, so a text-less goal input is refused. + +In the chat, round 1 shows the attachments in full under its bubble and later rounds collapse them into a one-line chip that expands on click — they really are in every round's input, but a twenty-round goal should not repeat the same picture twenty times. + +The three input paths now share one fold, bound once per Session, with each path deciding for itself whether to apply it: a prompt and a steering message fold only without vision, a goal objective always. The gate used to be implicit — "was the scratchpad directory configured" stood in for "does the model lack vision" — which worked only while there was exactly one rule to encode. + +## The mid-run button stops being able to strand you + +While a task runs, the composer's single action button flips from Stop to its send face as soon as the draft has any content. Steer mode could not carry every draft, so a draft it could not carry — an image with no text, back when images were not steerable — produced a **permanently disabled** send button that had also displaced Stop: the message could not be sent, and the run could not be stopped from the composer either. The only exit was deleting the attachment. + +With images steerable, that specific draft simply steers. What steering still cannot carry — selected skills with no text or image (a `[use_skills]` block is task-level setup, not something to hand a turn already under way), or an `@` handoff target — now falls back to the **follow-up queue** for that one send: the button stays live and labeled as the queue action, the whole draft goes out as one message held server-side until the run finishes, and an `@` handoff still opens its new chat directly. Stop is what's left over: the button's face whenever neither channel will take what is in the box — an empty composer, but also a goal objective, any draft at all once the model key has been rejected, and a staged `/model` fork the running session is still blocking. "There is something in the box" and "this send has somewhere to go" are different questions, and Stop belongs to the second one. diff --git a/changelog/unreleased/README.md b/changelog/unreleased/README.md index eb88152..b445c7f 100644 --- a/changelog/unreleased/README.md +++ b/changelog/unreleased/README.md @@ -22,6 +22,8 @@ Changes since v0.1.4. The version number is assigned at release, when this folde - [2026-07-28] Web App and CLI: duration and byte abbreviations now carry into the next unit instead of printing `1m60s` or `1024KB` — both helpers rounded the value they displayed but chose the unit (or the minute/second split) from the raw input. ([details](2026-07-28-duration-and-byte-formatting.md)) +- [2026-07-27] Input images: mid-run steering carries them (an uncaptioned image is a complete steering message), a goal objective accepts them as scratchpad paths on every model since it is re-injected as text each round, and the composer's mid-run action button can no longer sit disabled with Stop displaced. ([details](2026-07-27-input-images.md)) + - [2026-07-27] Windows: the `win32-x64` package bundles MinGit under `git/`, so `exec_command` has a real bash even on a machine with no Git for Windows — the shell stops depending on what happens to be installed. A user's own Git for Windows still wins (its MSYS userland is the fuller one); the bundle is the floor, and PowerShell is now reached only by npm installs. GPLv2 obligations are recorded in a new root `THIRD-PARTY-NOTICES.md`. ([details](2026-07-27-windows-bundled-shell.md)) - [2026-07-27] Sites: the 0.1.4 release post in both languages, with a capture script for its screenshots. ([details](2026-07-27-sites-and-blog.md)) diff --git a/packages/cli/src/commands/chat.ts b/packages/cli/src/commands/chat.ts index 640ebf3..5dea8a1 100644 --- a/packages/cli/src/commands/chat.ts +++ b/packages/cli/src/commands/chat.ts @@ -248,7 +248,7 @@ export function registerChatCommand(program: Command, t: Messages): void { if (state === "running") { const text = line.trim(); if (text.length === 0) return; - if (session.steer(text)) { + if (session.steer([userText(text)])) { // Printed via the renderer while the typing hold is still engaged, so the ack // lands before the held stream output flushes underneath it. renderer.printLine(dim(t.steerQueued(text))); diff --git a/packages/core/src/agent.ts b/packages/core/src/agent.ts index c9ce393..3f96421 100644 --- a/packages/core/src/agent.ts +++ b/packages/core/src/agent.ts @@ -291,15 +291,12 @@ export class Agent { createLLM: rt.createLLM, createBareLLM: rt.createBareLLM, compaction: rt.compaction, - // Model doesn't support images: input images are written to the session scratchpad and their paths appended to the text (viewed via describe_image). - ...(modelEntry.vision === false - ? { - inputImagesDir: path.join( - scratchpadDir(this.state.root, this.state.projectId, this.state.agentId), - sessionId, - ), - } - : {}), + // Where an input image lands when it becomes a path line (see SessionConfig.imagesDir). + imagesDir: path.join( + scratchpadDir(this.state.root, this.state.projectId, this.state.agentId), + sessionId, + ), + modelHasVision: modelEntry.vision !== false, // Goal mode's control file lives in the session scratchpad; the path is fixed per // Session, so it is wired here rather than passed per-run. goalFilePath: goalFilePath( @@ -451,15 +448,12 @@ export class Agent { createLLM: rt.createLLM, createBareLLM: rt.createBareLLM, compaction: rt.compaction, - // Model doesn't support images: input images are written to the session scratchpad and their paths appended to the text (viewed via describe_image). - ...(modelEntry.vision === false - ? { - inputImagesDir: path.join( - scratchpadDir(this.state.root, this.state.projectId, this.state.agentId), - sessionId, - ), - } - : {}), + // Where an input image lands when it becomes a path line (see SessionConfig.imagesDir). + imagesDir: path.join( + scratchpadDir(this.state.root, this.state.projectId, this.state.agentId), + sessionId, + ), + modelHasVision: modelEntry.vision !== false, // Goal mode's control file lives in the session scratchpad; the path is fixed per // Session, so it is wired here rather than passed per-run. goalFilePath: goalFilePath( diff --git a/packages/core/src/engine/context-engine.ts b/packages/core/src/engine/context-engine.ts index b8f1046..21db464 100644 --- a/packages/core/src/engine/context-engine.ts +++ b/packages/core/src/engine/context-engine.ts @@ -200,8 +200,27 @@ export interface ContextEngineDeps { compaction?: CompactionSettings; /** This Session's session_meta message; written at the start of the new Trace file after compaction splits it. */ sessionMeta?: OmniMessage; + /** + * Input adapter for a session whose model has no vision: folds image messages into text + * lines appended to the input's user text. Absent = the model takes images directly. `run`'s + * Prompt is folded by the caller before it reaches the engine; this hook exists for the one + * input the engine assembles itself — steering (see `steeringMessages`). + * + * Expected to settle rather than reject: it runs mid-Task, and Session's binding already + * degrades a failure into text saying the images were dropped. + */ + foldInputImages?: (messages: OmniMessage[]) => Promise; } +const isImageMessage = (m: OmniMessage): boolean => + (m.payload as { type?: string }).type === "image_url"; + +/** Whether a message carries steering of its own — an image, or text that isn't blank. */ +const carriesSteering = (m: OmniMessage): boolean => { + const p = m.payload as { type?: string; text?: string }; + return p.type === "image_url" || (p.type === "text" && (p.text ?? "").trim().length > 0); +}; + /** Whether compaction is possible; when not `ok`, `compact()` is a no-op and yields no messages (see ContextEngine.compactability). */ export type CompactAvailability = "ok" | "unsupported" | "empty" | "just_compacted"; @@ -378,7 +397,7 @@ export class ContextEngine { * later Task would be more surprising than losing it; hosts get `steer() === false` * after that point and fall back to a normal task). */ - private steeringQueue: string[] = []; + private steeringQueue: OmniMessage[][] = []; /** Whether a `run` is currently in flight (gates `steer`; compaction does not count). */ private taskRunning = false; @@ -424,28 +443,39 @@ export class ContextEngine { * Queues a steering message for the running Task: it is delivered with the next request * input as a standalone `[user_steering]` user message — alongside that turn's tool * outputs, or alone as the continuation input when the turn produced no tool calls. - * Returns false when no Task is running (the host should then submit the text as a - * normal task instead). + * `input` is an OmniMessage list, the shape `run` takes a Prompt in: its user text becomes + * the block's body and its images ride behind that text, exactly as a Prompt carries them; + * on a model without vision they are folded into path lines at delivery (see + * deliverSteering). Returns false when no Task is running (the host should then submit the + * message as a normal task instead). + * + * An input with neither text nor images queues nothing and still returns true: `false` is + * specifically "send this as a normal task", which would be the wrong advice for an empty + * one. Every host guards against this already; the check is here so an empty + * `[user_steering]` block can't reach the model through a host that forgets. */ - steer(text: string): boolean { + steer(input: OmniMessage[]): boolean { if (!this.taskRunning) return false; - this.steeringQueue.push(text); + if (!input.some(carriesSteering)) return true; + this.steeringQueue.push(input); return true; } /** * Drains the steering queue into standalone `[user_steering]` user messages (one per - * queued text, in arrival order), yielding each to the output stream and writing it to - * Trace — steering is real user input: unlike a normal Prompt (which the render layer - * already holds locally) this text never reached the consumer, and replay attributes it - * positionally to the next turn's input like any other user message. Returns the messages - * for the caller to append to the next request input; an empty queue is a no-op. + * queued entry, in arrival order, each followed by its images), yielding every message to + * the output stream and writing it to Trace — steering is real user input: unlike a normal + * Prompt (which the render layer already holds locally) this text never reached the + * consumer, and replay attributes it positionally to the next turn's input like any other + * user message. Returns the messages for the caller to append to the next request input; + * an empty queue is a no-op. */ private async *deliverSteering(): AsyncGenerator { if (this.steeringQueue.length === 0) return []; const drained = this.steeringQueue; this.steeringQueue = []; - const messages = drained.map((text) => userText(userSteeringText(text))); + const messages: OmniMessage[] = []; + for (const input of drained) messages.push(...(await this.steeringMessages(input))); for (const msg of messages) { yield msg; await this.write(msg); @@ -453,6 +483,39 @@ export class ContextEngine { return messages; } + /** + * One queued steering input -> the messages carrying it: its user text collected into the + * `[user_steering]`-wrapped message, followed by everything else it held — the images, on a + * vision model. That is the shape a Prompt uses, so every consumer down the line — LLM + * client, Trace, replay — already knows it. + * + * When `deps.foldInputImages` is given, the input goes through it **before** the wrapping so + * the images land inside the block: `parseUserSteeringText` only recognizes a text that is + * exactly one block, and anything appended after the closing tag would cost the message its + * steering identity — every render layer would read it as a new Task. + */ + private async steeringMessages(input: OmniMessage[]): Promise { + // No images, no fold: an image-free steering message is the same message either way. + const fold = input.some(isImageMessage) ? this.deps.foldInputImages : undefined; + const messages = fold ? await fold(input) : input; + const texts: string[] = []; + const rest: OmniMessage[] = []; + for (const msg of messages) { + const p = msg.payload as { type?: string; role?: string; text?: string }; + if (p.type === "text" && p.role === "user") texts.push(p.text ?? ""); + else rest.push(msg); + } + // `foldInputImages` is public API, so a third-party adapter can return something else, and + // both ways it can break lose the picture: an image that survived the fold goes to the one + // model known to refuse it, and no text at all means the images were dropped rather than + // written down as paths. Name the contract instead of delivering a steering message that + // lost what it was sent to carry. + if (fold && (rest.some(isImageMessage) || texts.length === 0)) { + throw new Error("foldInputImages must return the input's images folded into a user text."); + } + return [userText(userSteeringText(texts.join("\n\n"))), ...rest]; + } + /** The actual Task loop behind `run` (split out so run's finally can close the steering window on every exit path). */ private async *runToCompletion( newMessages: OmniMessage[], diff --git a/packages/core/src/internal/session-support.ts b/packages/core/src/internal/session-support.ts index 4b6b62b..a041001 100644 --- a/packages/core/src/internal/session-support.ts +++ b/packages/core/src/internal/session-support.ts @@ -103,6 +103,14 @@ export async function createTempWorkspace( ); } +/** + * Stands in for a single image that could not be turned into a path line — a data URL that + * doesn't parse. Shown to the model and the user instead of silently dropping the attachment. + * A scratchpad that can't be written to is a different matter: the fold throws and the run + * ends, rather than carrying on without what the sender attached. + */ +const IMAGE_DROPPED_NOTE = "[an attached image could not be saved and was dropped]"; + /** Maps a data URL's mime type to a file extension on disk; unknown mimes use bin (the image-reading tool sniffs the magic bytes and doesn't rely on the extension). */ const MIME_TO_EXT: Record = { "image/png": "png", @@ -111,6 +119,9 @@ const MIME_TO_EXT: Record = { "image/webp": "webp", }; +/** An input image message — what both conversions below pull out of the input. */ +const isImage = (m: OmniMessage): boolean => (m.payload as { type?: string }).type === "image_url"; + /** * Input conversion for when the session model doesn't support images: image messages * in the `run` input are written to disk as files (base64 data URLs are saved to the @@ -126,8 +137,6 @@ export async function imagesToScratchpadPaths( input: OmniMessage[], dir: string, ): Promise { - const isImage = (m: OmniMessage): boolean => - (m.payload as { type?: string }).type === "image_url"; if (!input.some(isImage)) return input; const lines: string[] = []; @@ -140,7 +149,7 @@ export async function imagesToScratchpadPaths( } const match = /^data:([^;,]+);base64,(.+)$/s.exec(url); if (!match) { - lines.push("[an attached image could not be saved and was dropped]"); + lines.push(IMAGE_DROPPED_NOTE); continue; } await fs.mkdir(dir, { recursive: true }); diff --git a/packages/core/src/omnimessage/markers/steering.ts b/packages/core/src/omnimessage/markers/steering.ts index e2c8bb1..2ac18c4 100644 --- a/packages/core/src/omnimessage/markers/steering.ts +++ b/packages/core/src/omnimessage/markers/steering.ts @@ -5,7 +5,11 @@ * queues it and delivers it as a **standalone user text message** wrapped in a * `[user_steering]…[/user_steering]` block, sent with the next request input alongside that * turn's tool outputs (or as the continuation input when the turn produced no tool calls) — - * the model sees it without the agent loop being interrupted. The message is real user input: + * the model sees it without the agent loop being interrupted. Images sent with the message + * follow it as ordinary user image messages (a model without vision gets `[attached image: …]` + * lines inside the block instead), so consumers group them the way they already group a + * Prompt's images: everything up to the next non-image message belongs to the steering + * message, and none of it starts a new Task. The message is real user input: * written to Trace like any Prompt and yielded to the output stream. The marker exists so the * model and the render layers can tell it apart from a task-starting Prompt: UIs keep it inside * the running Task (no new Task segment) and render it as user speech instead of raw markers. diff --git a/packages/core/src/session.ts b/packages/core/src/session.ts index f36a6ac..2244a8b 100644 --- a/packages/core/src/session.ts +++ b/packages/core/src/session.ts @@ -59,13 +59,18 @@ export interface SessionConfig { /** Session resume: the full historical messages of the current context (for rendering, including interrupted turns and their markers), for frontend display. */ resumedHistory?: OmniMessage[]; /** - * Set when the session's model doesn't support images (the composition layer decides this via - * ModelEntry.vision): images in `run` input are saved to this directory (the session's - * scratchpad), and the path is appended to the user text instead — the model views the image - * via describe_image, and images never enter the session history directly (some providers - * return a 400 outright on image input). + * This Session's scratchpad directory: where an input image is saved when it becomes an + * `[attached image: ]` line instead of riding the request as an image. The model + * reads it back with describe_image / read_image, and the Web turns the path into a + * thumbnail again. Always set — each input path decides on its own whether to use it. */ - inputImagesDir?: string; + imagesDir: string; + /** + * Whether the session's model accepts image input (from ModelEntry.vision). Prompts and + * steering messages fold their images only when this is false; goal objectives fold either + * way, since they are re-injected as text every round (see `runGoal`). + */ + modelHasVision: boolean; /** * Absolute path of this Session's GOAL.yaml (the composition layer derives it from the * agent scratchpad — see `goalFilePath` in state/paths.ts). Goal mode @@ -126,9 +131,18 @@ export class Session { private readonly trace?: TraceSink; private readonly meta: OmniMessage; private readonly createBareLLM?: () => LLMInterface; - private readonly inputImagesDir?: string; + private readonly imagesDir: string; + private readonly modelHasVision: boolean; private readonly goalFile?: string; private metaWritten = false; + /** + * The image fold, bound to this Session's scratchpad — Session is the layer that knows both + * the directory and the model's capability, so it binds the conversion once and each input + * path calls it under its own rule (see `modelHasVision`). The body reads `imagesDir` at call + * time, so field ordering in the constructor doesn't matter. + */ + private readonly foldImages = (messages: OmniMessage[]): Promise => + imagesToScratchpadPaths(messages, this.imagesDir); /** Title material (used by `generateTitle` as the default): the user input and model body text of the first Task that contains user text. */ private titleUserText = ""; private titleAssistantText = ""; @@ -146,7 +160,8 @@ export class Session { this.metaWritten = config.metaAlreadyWritten ?? false; if (config.resumedHistory) this.resumedHistory = config.resumedHistory; if (config.createBareLLM) this.createBareLLM = config.createBareLLM; - if (config.inputImagesDir) this.inputImagesDir = config.inputImagesDir; + this.imagesDir = config.imagesDir; + this.modelHasVision = config.modelHasVision; if (config.goalFilePath) this.goalFile = config.goalFilePath; this.engine = new ContextEngine({ llm: config.llm, @@ -157,6 +172,13 @@ export class Session { ...(config.createLLM ? { createLLM: config.createLLM } : {}), ...(config.compaction ? { compaction: config.compaction } : {}), ...(config.initialEngineState ? { initialState: config.initialEngineState } : {}), + // The engine assembles one input of its own — a steering message with images — and folds + // it through the same converter `runTask` uses, failures included: a scratchpad that + // can't be written to ends the run rather than dropping the attachment and carrying on. + // The picture usually arrives BECAUSE the run is going the wrong way, so continuing + // without it spends the rest of the Task heading further that way. A vision model simply + // isn't given the function, which is all the engine needs to know about the subject. + ...(config.modelHasVision ? {} : { foldInputImages: this.foldImages }), sessionMeta: this.meta, }); } @@ -173,9 +195,10 @@ export class Session { * Docs: /docs/agent-loop § "The loop at a glance". * * With `opts.goal` present, the same call runs **goal mode**: the input's text becomes the - * objective, and the Session loops Tasks — each round's input is the `[goal]` protocol - * block followed by the text (round 1 verbatim, later rounds the objective) — until the - * goal file says stop, the budget runs out, or a round is cut off. Round inputs are + * objective (attached images fold into `[attached image: …]` lines within it, whatever the + * model's vision — see `runGoal`), and the Session loops Tasks — each round's input is the + * `[goal]` protocol block followed by the text (round 1 verbatim, later rounds the objective) + * — until the goal file says stop, the budget runs out, or a round is cut off. Round inputs are * yielded onto the stream before each round (a plain run never yields its own input), and * the final message is exactly one `goal_finished` event carrying the outcome. * Docs: /docs/goal-mode. @@ -195,11 +218,8 @@ export class Session { newMessages: OmniMessage[], opts?: RunOptions, ): AsyncGenerator { - // Model doesn't support images: input images are saved to disk first (session scratchpad), - // then the path is appended to the text before it reaches the engine/Trace. - if (this.inputImagesDir) { - newMessages = await imagesToScratchpadPaths(newMessages, this.inputImagesDir); - } + // Folded before Trace and title material, so the path lines are what gets recorded. + if (!this.modelHasVision) newMessages = await this.foldImages(newMessages); await this.ensureMetaWritten(); // Self-captures title material (the title is derived from the first-turn // conversation text): while material isn't frozen yet, collect this call's user text and @@ -231,12 +251,20 @@ export class Session { } /** - * The goal-mode branch of `run`: validates the input (text-only — the objective is - * re-injected every round, and images have no place in the protocol block), then drives - * the goal loop, running each round through the single-Task path with the same per-call - * options (approval, signal, thinking level). The loop's terminal `goal_finished` event is - * additionally written to the Trace (best-effort, like session_meta) so the goal's end - * survives with its conversation. + * The goal-mode branch of `run`: validates the input, folds any attached images into the + * objective, then drives the goal loop, running each round through the single-Task path with + * the same per-call options (approval, signal, thinking level). The loop's terminal + * `goal_finished` event is additionally written to the Trace (best-effort, like session_meta) + * so the goal's end survives with its conversation. + * + * Images fold here whether or not the model has vision, because the objective is re-injected + * as the text of every round's `[goal]` block — an image message has nowhere to sit in that. + * Sending it in round 1 alone would leave later rounds pointing at something compaction has + * since dropped, while the objective still reads correct; a path line survives every round + * and every compaction, and the model pays for the picture only when it looks. The lines go + * at the end of the text, where goal-loop's `stripLeadingMarkerBlocks` leaves them alone. + * + * Text is required and checked before the fold: a picture alone doesn't state a goal. */ private async *runGoal( newMessages: OmniMessage[], @@ -246,16 +274,27 @@ export class Session { if (!this.goalFile) { throw new Error("Goal mode is unavailable: this Session has no goal file path configured."); } - const texts: string[] = []; - for (const m of newMessages) { + const isUserText = (m: OmniMessage): boolean => { const p = m.payload as { type?: string; role?: string; text?: string }; + return m.type === "model_msg" && p.type === "text" && p.role === "user" && !!p.text?.trim(); + }; + // The objective itself. Checked before the fold, which only ever appends to a user text — + // so once this passes, the joined text below cannot come out empty. + if (!newMessages.some(isUserText)) { + throw new Error("Goal mode requires a non-empty text objective."); + } + const folded = await this.foldImages(newMessages); + const texts: string[] = []; + for (const m of folded) { + const p = m.payload as { type?: string; role?: string; text?: string }; + // The fold has already turned the images into text, so anything left that isn't user + // text was never one of the two kinds goal mode accepts. if (m.type !== "model_msg" || p.type !== "text" || p.role !== "user" || !p.text) { - throw new Error("Goal mode requires text-only user input (the objective)."); + throw new Error("Goal mode accepts user text and images only (the objective)."); } texts.push(p.text); } const text = texts.join("\n").trim(); - if (!text) throw new Error("Goal mode requires a non-empty objective."); const loop = runGoalLoop( { run: (msgs) => this.runTask(msgs, opts) }, { @@ -283,13 +322,16 @@ export class Session { * Queues a steering message for the running Task: the engine delivers it between turns as * a standalone `[user_steering]` user message — sent with the next request input alongside * that turn's tool outputs, or alone as the continuation input when the turn produced no - * tool calls — so the model sees it without the loop being interrupted. Returns false when - * no Task is running — the host should then submit the text as a normal task instead. - * Delivery is independent of approval mode; anything still queued when the run exits - * (abort included) is discarded. + * tool calls — so the model sees it without the loop being interrupted. `input` is an + * OmniMessage list, the same shape `run` takes a Prompt in: its user text becomes the + * block's body, and its images ride behind that message just as a Prompt's do; on a model + * without vision they become `[attached image: …]` path lines inside the block instead (see + * the engine's steeringMessages). Returns false when no Task is running — the host should + * then submit the message as a normal task instead. Delivery is independent of approval + * mode; anything still queued when the run exits (abort included) is discarded. */ - steer(text: string): boolean { - return this.engine.steer(text); + steer(input: OmniMessage[]): boolean { + return this.engine.steer(input); } /** diff --git a/packages/core/test/engine.test.ts b/packages/core/test/engine.test.ts index a244f4d..8e33256 100644 --- a/packages/core/test/engine.test.ts +++ b/packages/core/test/engine.test.ts @@ -7,13 +7,14 @@ * turn, until some turn produces no tool_call (Task done) or is interrupted. Approval/execution * are within-turn interactions, and execution can overlap. */ -import { mkdtemp, readFile, rm } from "node:fs/promises"; +import { mkdtemp, readFile, readdir, rm } from "node:fs/promises"; import { tmpdir } from "node:os"; import { join } from "node:path"; import { afterEach, beforeEach, describe, expect, it, vi } from "vitest"; import { assistantText, emptyTokenCounts, + imageUrlMessage, isCompleteModelMessage, partialText, partialToolCallOutput, @@ -32,8 +33,14 @@ import { Environment } from "../src/environment/index.js"; import { Writer, readTrace } from "../src/trace/index.js"; import { ContextEngine, reconnectDelayMs } from "../src/engine/context-engine.js"; import { goalRoundMessage } from "../src/goal/goal-prompts.js"; +import { parseUserSteeringText } from "../src/omnimessage/markers/index.js"; +import { imagesToScratchpadPaths } from "../src/internal/session-support.js"; import type { ApproveFn, EnvironmentInterface, ToolPermission } from "../src/interfaces.js"; +/** A real 1x1 PNG data URL: the non-vision fold actually decodes and writes it to disk. */ +const PNG_DATA_URL = + "data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAAEAAAABCAYAAAAfFcSJAAAADUlEQVR42mP8z8BQDwAEhQGAhKmMIQAAAABJRU5ErkJggg=="; + /** Deterministic fake LLM: the first turn yields a tool_call, the second yields the final reply. */ class FakeLLM implements LLMInterface { calls = 0; @@ -2275,13 +2282,13 @@ describe("ContextEngine mid-run steering ([user_steering])", () => { const engine = new ContextEngine({ llm, environment: steeringEnvironment(), trace }); // Idle: nothing running yet -> steer refuses, the host falls back to a normal task. - expect(engine.steer("too early")).toBe(false); + expect(engine.steer([userText("too early")])).toBe(false); // Queue two messages while the task runs (deterministically: from the approval callback, // i.e. after the tool_call streamed but before the tool executed). const approve: ApproveFn = async () => { - expect(engine.steer("focus on the tests")).toBe(true); - expect(engine.steer("also update the docs")).toBe(true); + expect(engine.steer([userText("focus on the tests")])).toBe(true); + expect(engine.steer([userText("also update the docs")])).toBe(true); return "allow"; }; const all = await collectRun(engine, [userText("go")], approve); @@ -2319,7 +2326,7 @@ describe("ContextEngine mid-run steering ([user_steering])", () => { expect((recordedOutputs[0]!.payload as { output: string }).output).toBe("tool result"); // Task over: the queue window is closed again. - expect(engine.steer("late")).toBe(false); + expect(engine.steer([userText("late")])).toBe(false); }); it("delivers steering left at loop end as a [user_steering] continuation turn (traced, streamed)", async () => { @@ -2333,7 +2340,7 @@ describe("ContextEngine mid-run steering ([user_steering])", () => { inputs.push(params.newMessages); if (inputs.length === 1) { yield assistantText("final answer"); - expect(engineRef!.steer("one more thing")).toBe(true); + expect(engineRef!.steer([userText("one more thing")])).toBe(true); return { status: "completed" }; } yield assistantText("handled the follow-up"); @@ -2355,6 +2362,206 @@ describe("ContextEngine mid-run steering ([user_steering])", () => { expect(steeringTexts(await readTrace(trace.currentPath()))).toEqual([wrapped]); }); + /** Image URLs of the image messages in a list, in order. */ + const steeredImages = (msgs: OmniMessage[]): string[] => + msgs + .map((m) => m.payload as { type?: string; image_url?: string }) + .filter((p) => p.type === "image_url") + .map((p) => p.image_url!); + + it("carries a steering message's images right behind its text (vision model), streamed and traced with it", async () => { + // An image with no caption is a complete steering message on its own; each entry's images + // follow that entry's text, so two steers stay distinguishable in the delivered order. + let engineRef: ContextEngine | null = null; + const inputs: OmniMessage[][] = []; + const llm: LLMInterface = { + async *streamGenerate(params): AsyncGenerator { + inputs.push(params.newMessages); + if (inputs.length === 1) { + expect(engineRef!.steer([imageUrlMessage("data:image/png;base64,AAAA")])).toBe(true); + expect( + engineRef!.steer([ + userText("this one too"), + imageUrlMessage("data:image/png;base64,BBBB"), + ]), + ).toBe(true); + yield assistantText("final answer"); + return { status: "completed" }; + } + yield assistantText("looked at both"); + return { status: "completed" }; + }, + }; + const trace = new Writer({ tracesDir: traces, sessionId: "sess_steer_img" }); + const engine = new ContextEngine({ llm, environment: steeringEnvironment(), trace }); + engineRef = engine; + + const all = await collectRun(engine, [userText("go")], allowAll); + + expect(inputs).toHaveLength(2); + expect( + inputs[1]!.map((m) => { + const p = m.payload as { type: string; text?: string; image_url?: string }; + return p.type === "image_url" ? `img:${p.image_url}` : p.text; + }), + ).toEqual([ + "[user_steering]\n\n[/user_steering]", + "img:data:image/png;base64,AAAA", + "[user_steering]\nthis one too\n[/user_steering]", + "img:data:image/png;base64,BBBB", + ]); + // Streamed and traced like the text they belong to (a plain run never yields its Prompt, + // so these are the steering images and nothing else). + const expectedImages = ["data:image/png;base64,AAAA", "data:image/png;base64,BBBB"]; + expect(steeredImages(all)).toEqual(expectedImages); + expect(steeredImages(await readTrace(trace.currentPath()))).toEqual(expectedImages); + }); + + it("without vision, a steering message's images fold into [attached image: …] lines INSIDE the block", async () => { + // The block must stay the whole text: lines appended after the closing tag would cost the + // message its steering identity (parseUserSteeringText) and read as a new Task everywhere. + const scratch = await mkdtemp(join(tmpdir(), "penguin-steer-img-")); + try { + let engineRef: ContextEngine | null = null; + const inputs: OmniMessage[][] = []; + const llm: LLMInterface = { + async *streamGenerate(params): AsyncGenerator { + inputs.push(params.newMessages); + if (inputs.length === 1) { + expect( + engineRef!.steer([userText("look at this"), imageUrlMessage(PNG_DATA_URL)]), + ).toBe(true); + // An image with no caption of its own: the whole steering message is the picture. + expect(engineRef!.steer([imageUrlMessage(PNG_DATA_URL)])).toBe(true); + yield assistantText("final answer"); + return { status: "completed" }; + } + yield assistantText("read the file"); + return { status: "completed" }; + }, + }; + const engine = new ContextEngine({ + llm, + environment: steeringEnvironment(), + foldInputImages: (messages) => imagesToScratchpadPaths(messages, scratch), + }); + engineRef = engine; + await collectRun(engine, [userText("go")], allowAll); + + // Both entries arrive as text and nothing else — every image became a line inside a block. + expect(inputs[1]!.map((m) => (m.payload as { type: string }).type)).toEqual(["text", "text"]); + const inner = inputs[1]!.map((m) => + parseUserSteeringText((m.payload as { text: string }).text), + ); + expect(inner[0]).toMatch(/^look at this\n\n\[attached image: .+\]$/); + // The caption-less one is the path line and nothing else: no blank line standing in for + // the text that was never sent. + expect(inner[1]).toMatch(/^\[attached image: .+\]$/); + // Both files really landed in the scratchpad. + expect(await readdir(scratch)).toHaveLength(2); + } finally { + await rm(scratch, { recursive: true, force: true }); + } + }); + + // Session hands the engine the same throwing fold a Prompt gets: an unwritable scratchpad + // ends the run rather than dropping the attachment and carrying on. The picture usually + // arrives BECAUSE the run is going the wrong way, so continuing without it would spend the + // rest of the Task heading further that way. + it("a steering fold that fails ends the run instead of carrying on without the images", async () => { + let engineRef: ContextEngine | null = null; + const inputs: OmniMessage[][] = []; + const llm: LLMInterface = { + async *streamGenerate(params): AsyncGenerator { + inputs.push(params.newMessages); + if (inputs.length === 1) { + expect( + engineRef!.steer([ + userText("look at this"), + imageUrlMessage("data:image/png;base64,AAAA"), + ]), + ).toBe(true); + } + yield assistantText("final answer"); + return { status: "completed" }; + }, + }; + const engine = new ContextEngine({ + llm, + environment: steeringEnvironment(), + foldInputImages: () => + Promise.reject(Object.assign(new Error("ENOSPC: no space left"), { code: "ENOSPC" })), + }); + engineRef = engine; + + await expect(collectRun(engine, [userText("go")], allowAll)).rejects.toThrow(/ENOSPC/); + // It ended at delivery: no second request went out carrying a note in place of the image. + expect(inputs).toHaveLength(1); + }); + + it("a fold returning an unreadable shape names the broken contract instead of carrying on", async () => { + // foldInputImages is public API (ContextEngineDeps is exported), so a third-party adapter + // can return the wrong thing. Sending the images on as messages would be the worse answer: + // a fold is configured precisely because the model does not take images. + let engineRef: ContextEngine | null = null; + const inputs: OmniMessage[][] = []; + const llm: LLMInterface = { + async *streamGenerate(params): AsyncGenerator { + inputs.push(params.newMessages); + if (inputs.length === 1) { + expect( + engineRef!.steer([ + userText("look at this"), + imageUrlMessage("data:image/png;base64,AAAA"), + ]), + ).toBe(true); + } + yield assistantText("final answer"); + return { status: "completed" }; + }, + }; + const engine = new ContextEngine({ + llm, + environment: steeringEnvironment(), + foldInputImages: async () => [], + }); + engineRef = engine; + + await expect(collectRun(engine, [userText("go")], allowAll)).rejects.toThrow( + /foldInputImages must return/, + ); + expect(inputs).toHaveLength(1); + }); + + it("a steering message with neither text nor images queues nothing (and asks for no fallback)", async () => { + // `false` means "no Task running — send it as a normal task", which would be the wrong + // advice for an empty message; so it returns true and simply delivers nothing. + let engineRef: ContextEngine | null = null; + const inputs: OmniMessage[][] = []; + const llm: LLMInterface = { + async *streamGenerate(params): AsyncGenerator { + inputs.push(params.newMessages); + if (inputs.length === 1) { + expect(engineRef!.steer([userText(" ")])).toBe(true); + yield assistantText("final answer"); + return { status: "completed" }; + } + yield assistantText("must not happen"); + return { status: "completed" }; + }, + }; + const engine = new ContextEngine({ llm, environment: steeringEnvironment() }); + engineRef = engine; + const all = await collectRun(engine, [userText("go")], allowAll); + + // Nothing queued -> the turn produced no tool calls and no steering, so the Task ends. + expect(inputs).toHaveLength(1); + const userTexts = all.filter( + (m) => m.type === "model_msg" && (m.payload as { role?: string }).role === "user", + ); + expect(userTexts).toHaveLength(0); + }); + it("steering queued during a mid-run compaction is delivered right after it (never swallowed)", async () => { // Turn 1 completes over the context threshold -> summarize compaction runs on the old // LLM; the user steers DURING the compaction request (the acceptance window stays open); @@ -2366,7 +2573,7 @@ describe("ContextEngine mid-run steering ([user_steering])", () => { const texts = params.newMessages.map((m) => (m.payload as { text?: string }).text ?? ""); if (texts.some((t) => t.includes("summary prompt"))) { // The compaction request: steering arrives while it streams. - expect(engineRef!.steer("switch to staging")).toBe(true); + expect(engineRef!.steer([userText("switch to staging")])).toBe(true); yield assistantText("[summary]the gist[/summary]"); yield tokenUsage(emptyTokenCounts(), { cache_read: 0, @@ -2425,7 +2632,7 @@ describe("ContextEngine mid-run steering ([user_steering])", () => { const engine = new ContextEngine({ llm, environment: steeringEnvironment() }); const ac = new AbortController(); const approve: ApproveFn = async () => { - expect(engine.steer("stale steering")).toBe(true); + expect(engine.steer([userText("stale steering")])).toBe(true); ac.abort(); return "allow"; }; @@ -2433,7 +2640,7 @@ describe("ContextEngine mid-run steering ([user_steering])", () => { expect(first.some((m) => (m.payload as { type?: string }).type === "abort")).toBe(true); // Aborted: whatever was queued is dropped with the run (documented steering contract). - expect(engine.steer("after abort")).toBe(false); + expect(engine.steer([userText("after abort")])).toBe(false); await collectRun(engine, [userText("continue")], allowAll); const followUpTexts = (llm.receivedSecondInput ?? []) .map((m) => { diff --git a/packages/core/test/goal.test.ts b/packages/core/test/goal.test.ts index 0a6ce7c..2cd742d 100644 --- a/packages/core/test/goal.test.ts +++ b/packages/core/test/goal.test.ts @@ -5,6 +5,7 @@ import { afterEach, beforeEach, describe, expect, it } from "vitest"; import { parse as parseYaml } from "yaml"; import { UNLIMITED_BUDGET, + Session, abortEvent, assistantText, buildSkillsMessage, @@ -12,6 +13,7 @@ import { emptyTokenCounts, goalFilePath, goalFinishedOf, + imageUrlMessage, isGoalRoundInput, parseGoalMessage, stripConversationMarkers, @@ -19,7 +21,15 @@ import { userText, withOrigin, } from "../src/index.js"; -import type { GoalOutcome, OmniMessage, TokenCounts } from "../src/index.js"; +import type { + EnvironmentInterface, + GoalOutcome, + LLMInterface, + LLMOutcome, + OmniMessage, + SessionMetaPayload, + TokenCounts, +} from "../src/index.js"; // The file protocol, prompt composition and the loop are internal to `session.run` (not part // of the SDK barrel); tests reach them through their modules directly. import { readGoalStatus, serializeGoalFile, writeGoalFile } from "../src/goal/goal-file.js"; @@ -389,3 +399,99 @@ describe("runGoalLoop", () => { expect(emptyTokenCounts().total).toBe(0); }); }); + +/** + * The Session-level goal entry (`run(input, { goal })`): input validation and the image fold. + * A goal objective folds its images to `[attached image: ]` lines on any model, unlike a + * Prompt — it is re-injected as the text of each round's `[goal]` block, which leaves an image + * message nowhere to sit. See Session.runGoal. + */ +describe("Session.runGoal input", () => { + const PNG_DATA_URL = + "data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAAEAAAABCAYAAAAfFcSJAAAADUlEQVR42mP8z8BQDwAEhQGAhKmMIQAAAABJRU5ErkJggg=="; + + const fakeEnvironment: EnvironmentInterface = { + listTools: async () => [], + // eslint-disable-next-line require-yield + executeTool: async function* () { + throw new Error("not used"); + }, + toolPermission: () => undefined, + }; + + /** A model that answers each round with one final text, marking the goal complete on `completeOn`. */ + function fakeLLM(completeOn: number): LLMInterface { + let round = 0; + return { + async *streamGenerate() { + round++; + if (round >= completeOn) await setStatus("complete"); + yield assistantText(`round ${round} done`); + return { status: "completed" } satisfies LLMOutcome; + }, + }; + } + + // `modelHasVision: true` throughout: the fold runs regardless, and a vision model is the case + // that would break if runGoal ever grew the `if (!this.modelHasVision)` the other paths have. + function makeSession(completeOn = 1): Session { + const meta: SessionMetaPayload = { + session_id: "session-1", + provider: "custom", + model_id: "m1", + model_context_window: 1000, + system_prompt: "sp", + tools: [], + agent_state: dir, + workspace: dir, + }; + return new Session({ + meta, + llm: fakeLLM(completeOn), + environment: fakeEnvironment, + imagesDir: path.join(dir, "scratchpad", "session-1"), + modelHasVision: true, + goalFilePath: file, + }); + } + + /** Drives the goal and returns the text of each round's injected `[goal]` input. */ + async function roundInputs(session: Session, input: OmniMessage[]): Promise { + const texts: string[] = []; + for await (const msg of session.run(input, { goal: {} })) { + if (isGoalRoundInput(msg)) texts.push((msg.payload as { text: string }).text); + } + return texts; + } + + it("folds an attached image into the objective and re-injects it every round — vision model included", async () => { + const session = makeSession(2); + const rounds = await roundInputs(session, [ + userText("Match this mockup"), + imageUrlMessage(PNG_DATA_URL), + ]); + expect(rounds).toHaveLength(2); + // The picture is on disk, and both rounds point at that same file. + const saved = await fs.readdir(path.join(dir, "scratchpad", "session-1")); + expect(saved).toHaveLength(1); + const line = `[attached image: ${path.join(dir, "scratchpad", "session-1", saved[0]!)}]`; + for (const text of rounds) expect(text).toContain(line); + // Round 2 re-injects the objective alone, which is where the line matters most: it + // survives because stripLeadingMarkerBlocks only removes leading blocks, and the fold + // appends at the end. + expect(parseGoalMessage(rounds[1]!)?.rest).toBe(`Match this mockup\n\n${line}`); + // No image message ever reaches the round input. + expect(rounds.every((t) => !t.includes("data:image"))).toBe(true); + }); + + it("rejects an objective with no text: an image alone states no goal", async () => { + const session = makeSession(); + await expect(roundInputs(session, [imageUrlMessage(PNG_DATA_URL)])).rejects.toThrow( + /non-empty text objective/, + ); + // A blank text is no better. + await expect( + roundInputs(session, [userText(" "), imageUrlMessage(PNG_DATA_URL)]), + ).rejects.toThrow(/non-empty text objective/); + }); +}); diff --git a/packages/core/test/input-images.test.ts b/packages/core/test/input-images.test.ts index 0f13d53..421ce24 100644 --- a/packages/core/test/input-images.test.ts +++ b/packages/core/test/input-images.test.ts @@ -8,11 +8,20 @@ * images above and the server's `[attached file: …]` uploads). */ import { afterEach, beforeEach, describe, expect, it } from "vitest"; -import { mkdtemp, readFile, readdir, rm } from "node:fs/promises"; +import { mkdtemp, readFile, readdir, rm, writeFile } from "node:fs/promises"; import { tmpdir } from "node:os"; import path from "node:path"; import { appendAttachmentLines, imagesToScratchpadPaths } from "../src/internal/session-support.js"; +import { Session } from "../src/index.js"; +import type { + EnvironmentInterface, + LLMInterface, + LLMOutcome, + OmniMessage, + SessionMetaPayload, +} from "../src/index.js"; import { + assistantText, buildHandoffMessage, buildModelSwitchMessage, buildScheduledMessage, @@ -160,3 +169,105 @@ describe("appendAttachmentLines placement", () => { expect(textAt(out, 1)).toBe(block); }); }); + +/** + * Session's own wiring of the fold, which is where the per-input rules live: a Prompt and a + * steering message fold only without vision, and BOTH keep the throwing conversion. Steering + * used to get a copy that swallowed a disk failure and delivered the text with the images + * replaced by a note; it does not any more, because a picture sent mid-run usually arrives + * *because* the run is going the wrong way — carrying on without it spends the rest of the + * Task heading further that way. The engine only ever awaits what Session hands it, so this + * is the layer the rule actually lives in. + */ +describe("Session input-image wiring", () => { + const fakeEnvironment: EnvironmentInterface = { + listTools: async () => [], + // eslint-disable-next-line require-yield + executeTool: async function* () { + throw new Error("not used"); + }, + toolPermission: () => undefined, + }; + + const meta = (): SessionMetaPayload => ({ + session_id: "session-1", + provider: "custom", + model_id: "m1", + model_context_window: 1000, + system_prompt: "sp", + tools: [], + agent_state: tmp, + workspace: tmp, + }); + + /** An LLM that steers once from inside its first request, then answers. */ + function steeringLLM(steer: () => void): LLMInterface { + let turn = 0; + return { + async *streamGenerate(): AsyncGenerator { + if (++turn === 1) steer(); + yield assistantText(`turn ${turn}`); + return { status: "completed" }; + }, + }; + } + + /** A path that cannot be created: `/blocker` is a regular file, so mkdir under it fails. */ + async function unwritableDir(): Promise { + const blocker = path.join(tmp, "blocker"); + await writeFile(blocker, "not a directory", "utf8"); + return path.join(blocker, "session-1"); + } + + it("a steering image that cannot be saved ends the run instead of being dropped with a note", async () => { + let session!: Session; + session = new Session({ + meta: meta(), + llm: steeringLLM(() => { + expect(session.steer([userText("look at this"), imageUrlMessage(DATA_URL)])).toBe(true); + }), + environment: fakeEnvironment, + imagesDir: await unwritableDir(), + modelHasVision: false, + }); + + const drain = async () => { + for await (const _ of session.run([userText("go")])) { + // consume + } + }; + await expect(drain()).rejects.toThrow(/ENOTDIR|ENOENT|EEXIST|not a directory/i); + }); + + it("a vision model gets no fold at all, so an unwritable scratchpad never comes up", async () => { + // The engine is handed no converter when the model takes images: the steering images ride + // as messages and nothing touches the directory, broken or not. + const inputs: OmniMessage[][] = []; + let session!: Session; + session = new Session({ + meta: meta(), + llm: { + async *streamGenerate(params): AsyncGenerator { + inputs.push(params.newMessages); + if (inputs.length === 1) { + expect(session.steer([userText("look at this"), imageUrlMessage(DATA_URL)])).toBe(true); + } + yield assistantText(`turn ${inputs.length}`); + return { status: "completed" }; + }, + }, + environment: fakeEnvironment, + imagesDir: await unwritableDir(), + modelHasVision: true, + }); + + for await (const _ of session.run([userText("go")])) { + // consume + } + expect(inputs).toHaveLength(2); + expect(inputs[1]!.map((m) => (m.payload as { type: string }).type)).toEqual([ + "text", + "image_url", + ]); + }); +}); diff --git a/packages/core/test/session-title.test.ts b/packages/core/test/session-title.test.ts index f4e3d66..a865c31 100644 --- a/packages/core/test/session-title.test.ts +++ b/packages/core/test/session-title.test.ts @@ -58,6 +58,9 @@ const META: SessionMetaPayload = { workspace: "/tmp/w", }; +/** Image-fold wiring every Session takes; these tests send no images, so it is never exercised. */ +const IMAGES = { imagesDir: "/tmp/scratchpad/session-title-1", modelHasVision: true } as const; + describe("session-title", () => { it("generateTitleWithLLM: collects model text and usage, returns the sanitized result", async () => { const seen: string[] = []; @@ -127,6 +130,7 @@ describe("session-title", () => { it("Session.generateTitle: sends via createBareLLM; returns null when no factory is provided", async () => { const withFactory = new Session({ meta: META, + ...IMAGES, llm: fakeLLM([]), environment: fakeEnvironment, createBareLLM: () => fakeLLM([assistantText("Title A")]), @@ -140,6 +144,7 @@ describe("session-title", () => { const withoutFactory = new Session({ meta: META, + ...IMAGES, llm: fakeLLM([]), environment: fakeEnvironment, }); @@ -153,6 +158,7 @@ describe("session-title", () => { const seen: string[] = []; const session = new Session({ meta: META, + ...IMAGES, llm: fakeLLM([thinkingMessage("thinking"), assistantText("answer body")]), environment: fakeEnvironment, createBareLLM: () => fakeLLM([assistantText("Title B")], { status: "completed" }, seen), @@ -173,6 +179,7 @@ describe("session-title", () => { // No request is sent when no material has been collected (run was never called). const idle = new Session({ meta: META, + ...IMAGES, llm: fakeLLM([]), environment: fakeEnvironment, createBareLLM: () => fakeLLM([assistantText("must not be produced")]), diff --git a/packages/docs/content/agent-loop.en.md b/packages/docs/content/agent-loop.en.md index 1d3ff34..ffdb842 100644 --- a/packages/docs/content/agent-loop.en.md +++ b/packages/docs/content/agent-loop.en.md @@ -87,7 +87,19 @@ Carry-over enters the model context only — it is never written to the Trace, w ## Mid-run steering -While a Task is running, the host can queue a user message with `session.steer(text)` without interrupting the loop: at the next input assembly the engine delivers it as a **standalone user text message** wrapped in `[user_steering]…[/user_steering]`, sent alongside that turn's tool outputs (or alone as the continuation input when the turn produced no tool calls — the Task keeps going instead of ending). Steering is real user input: written to Trace like any Prompt, yielded to the output stream, and replayed as ordinary turn input on resume; tool outputs are never rewritten. The queue is drained at **every** input assembly — including right after a mid-run compaction, so steering that arrives during the compaction request is delivered, never swallowed. `steer` returns `false` when no Task is running (hosts then submit a normal task); the queue is discarded only when the run exits (abort included). +While a Task is running, the host can queue a user message with `session.steer(input)` — an OmniMessage list, the same shape `run` takes a Prompt in — without interrupting the loop: at the next input assembly the engine delivers it as a **standalone user text message** wrapped in `[user_steering]…[/user_steering]`, sent alongside that turn's tool outputs (or alone as the continuation input when the turn produced no tool calls — the Task keeps going instead of ending). The input's user text becomes the block's body; its images follow it as ordinary user image messages, so an image with no caption is a complete steering message; on a model without vision they fold into `[attached image: ]` lines **inside** the block instead, exactly as a Prompt's images do (the block must stay the whole text, or the message would lose its steering identity and read as a new Task). Steering is real user input: written to Trace like any Prompt, yielded to the output stream, and replayed as ordinary turn input on resume; tool outputs are never rewritten. The queue is drained at **every** input assembly — including right after a mid-run compaction, so steering that arrives during the compaction request is delivered, never swallowed. `steer` returns `false` when no Task is running (hosts then submit a normal task); the queue is discarded only when the run exits (abort included). + +## Input images + +An input image either rides the request as an image message or becomes an `[attached image: ]` line pointing at a file in the session scratchpad — the model then views it with `read_image` / `describe_image`, and the Web restores the thumbnail from the path. The conversion is one function bound once per Session (it is the only layer that knows both the scratchpad and the model's capability), and **each input path decides for itself whether to apply it**: + +| Input | Folds when | Applied at | +| --- | --- | --- | +| Prompt (`run`) | the model has no vision | run entry, before Trace and title material | +| Steering (`steer`) | the model has no vision | delivery, at the turn boundary — queuing must stay synchronous, and a queue discarded on abort would otherwise leave orphan files | +| Goal objective | **always** | before the objective is extracted, so the path lines survive every round's re-injection | + +Goal mode is the exception because its objective is re-injected as text every round: see [Goal mode](/docs/goal-mode). ## Automatic reconnect diff --git a/packages/docs/content/agent-loop.zh.md b/packages/docs/content/agent-loop.zh.md index cb44c8a..542ac2d 100644 --- a/packages/docs/content/agent-loop.zh.md +++ b/packages/docs/content/agent-loop.zh.md @@ -84,7 +84,19 @@ Task 由若干连续的 Request(轮)组成,每轮: ## 运行中插话(Steering) -Task 运行期间,宿主可通过 `session.steer(text)` 排队一条用户消息而不打断循环:引擎在下一次输入组装时把它作为**独立的用户文本消息**送出,内容包裹在 `[user_steering]…[/user_steering]` 中,与该轮工具输出一起进入下一次请求(该轮没有工具调用时则单独作为继续输入,Task 不会就此结束)。插话是真实的用户输入:像 Prompt 一样写入 Trace、推送到输出流,恢复重放时按普通轮次输入处理;工具输出本身从不被改写。队列在**每次**输入组装时排空——包括运行中压缩完成后的那次,压缩请求期间到达的插话不会被吞掉。没有 Task 运行时 `steer` 返回 `false`(宿主转为发起普通 Task);仅在运行退出(含中断)时丢弃队列。 +Task 运行期间,宿主可通过 `session.steer(input)` 排队一条用户消息而不打断循环(`input` 是 OmniMessage 列表,与 `run` 接收 Prompt 的形状一致):引擎在下一次输入组装时把它作为**独立的用户文本消息**送出,内容包裹在 `[user_steering]…[/user_steering]` 中,与该轮工具输出一起进入下一次请求(该轮没有工具调用时则单独作为继续输入,Task 不会就此结束)。输入中的用户文本成为标记块的正文,图片紧跟其后,作为普通用户图片消息送出,因此一张没有配文的图片本身就是一条完整的插话;模型不支持视觉时,图片改为折叠成 `[attached image: ]` 路径行写在标记块**内部**,与 Prompt 的图片走同一条路(标记块必须仍是整条文本,否则这条消息会丢掉插话身份、被当成新 Task)。插话是真实的用户输入:像 Prompt 一样写入 Trace、推送到输出流,恢复重放时按普通轮次输入处理;工具输出本身从不被改写。队列在**每次**输入组装时排空——包括运行中压缩完成后的那次,压缩请求期间到达的插话不会被吞掉。没有 Task 运行时 `steer` 返回 `false`(宿主转为发起普通 Task);仅在运行退出(含中断)时丢弃队列。 + +## 输入图片 + +一张输入图片要么以图片消息的形态跟着请求走,要么变成一行 `[attached image: <路径>]`,指向会话 scratchpad 里的文件——模型再用 `read_image` / `describe_image` 去看,Web 则从路径还原出缩略图。这个转换是每个 Session 绑定一次的同一个函数(Session 是唯一同时知道 scratchpad 目录和模型能力的层),而**是否折叠由各输入路径自己决定**: + +| 输入 | 何时折叠 | 折叠时机 | +| --- | --- | --- | +| Prompt(`run`) | 模型不支持图片 | `run` 入口,早于写 Trace 和取标题素材 | +| 插话(`steer`) | 模型不支持图片 | 投递时,即 turn 边界——入队必须保持同步,而中断时被丢弃的队列若已折叠会留下没人读的孤儿文件 | +| 目标(goal) | **总是** | 抽取 objective 之前,这样路径行才能活过每一轮的重新注入 | + +目标模式是唯一的例外,因为它的目标每轮都作为文本重新注入:见[目标模式](/docs/goal-mode)。 ## 自动重连 diff --git a/packages/docs/content/goal-mode.en.md b/packages/docs/content/goal-mode.en.md index 9e1a230..290ae60 100644 --- a/packages/docs/content/goal-mode.en.md +++ b/packages/docs/content/goal-mode.en.md @@ -44,6 +44,12 @@ Each round's user message is a `[goal]` protocol block followed by a plain body - `blocked` → the loop stops; what the model needs from you is in its final reply. The injected rules require the **same blocking condition to persist for three consecutive rounds** before the model may claim `blocked`, so a transient obstacle doesn't end the goal. - `active` → budget permitting, the next round fires. +### Images in an objective + +An objective may carry attached images — "make the page match this mockup" is a goal, and a screenshot states it better than a paragraph. They are always saved to the session scratchpad and referenced from the objective as `[attached image: ]` lines, **whatever the model's vision**: the objective is re-injected as the text of every round's block, so an image cannot ride along as an image. Sending it in round 1 alone would leave every later round pointing at something compaction has since removed, while the objective still reads correct. As a path it survives every round and every compaction, and the model spends tokens on it only when it actually looks (`read_image`, or `describe_image` without vision). An image cannot stand in for the text — a picture alone states no objective, so a text-less goal input is rejected. + +The chat page shows the attachments in full under round 1's bubble and collapses them into a one-line chip on later rounds (click to expand): they are part of every round's input, but a twenty-round goal shouldn't repeat the same picture twenty times. + A round that ends in an abort (user stop, LLM failure) ends the whole goal without re-firing — on-disk state stays `active`, so the workspace and goal file remain a clean resume point. In the Web App the regular stop button aborts the entire loop; in the CLI, Ctrl-C does. The same applies to a round the engine cut off at the per-Task turn cap (`max_turns`): the model never got to write the goal file, so the loop ends as `aborted` instead of re-firing the same cutoff forever. ## Token budget diff --git a/packages/docs/content/goal-mode.zh.md b/packages/docs/content/goal-mode.zh.md index 7f84666..d90f899 100644 --- a/packages/docs/content/goal-mode.zh.md +++ b/packages/docs/content/goal-mode.zh.md @@ -44,6 +44,12 @@ status: active - `blocked` → 循环停止;模型缺什么写在它最后一条回复里。注入规则要求**同一阻塞条件持续三个连续轮次**后才允许声明 `blocked`,临时性障碍不会终结目标。 - `active` → 预算允许则进入下一轮。 +### 目标里的图片 + +目标可以附带图片——「把页面改成这张设计稿的样子」本身就是一个目标,而一张截图比一段描述说得清楚。图片一律写入会话 scratchpad,在目标文本里以 `[attached image: <路径>]` 行引用,**与模型是否支持视觉无关**:目标每轮都作为协议块的文本被重新注入,图片没法以图片的形态跟着走。只在第一轮发,后续每一轮的目标就指向了一个早已被压缩掉的东西,而目标文本读起来却毫无破绽。作为路径,它跨越每一轮、每一次压缩都稳定存在,而模型只在真的要看的时候才付出 token(有视觉用 `read_image`,没有则 `describe_image`)。图片不能替代文字——一张图说明不了目标,所以没有文字的目标输入会被拒绝。 + +聊天页在第一轮气泡下方完整展示附图,后续轮次收成一行 chip(点击展开):它确实在每一轮的输入里,但二十轮的目标不该把同一张图重复二十次。 + 某一轮以中断结束(用户停止、LLM 故障)时整个目标随之结束、不再续推——磁盘上的状态保持 `active`,工作区与目标文件就是干净的断点。Web App 中常规停止按钮即中止整个循环;CLI 中是 Ctrl-C。被单 Task 轮次上限(`max_turns`)掐断的轮同理:模型没来得及写目标文件,循环以 `aborted` 结束,而不是永远重演同一次掐断。 ## Token 预算 diff --git a/packages/docs/content/server-api.en.md b/packages/docs/content/server-api.en.md index 908e2c5..f7eef10 100644 --- a/packages/docs/content/server-api.en.md +++ b/packages/docs/content/server-api.en.md @@ -161,8 +161,8 @@ The paths below omit the `/api/sessions/:sessionId` prefix. For the storage mode | DELETE | / | Delete the Session (along with its Traces and scratch files) | | GET | /messages | Full OmniMessage history; while a Task runs the response also carries `live` (the in-progress stream tail, see below) | | GET | /stream | SSE event stream (next section) | -| POST | /tasks | Start a Task: `{input: TaskInputPart[], thinkingLevel?, queueIfBusy?}` → 202. With `queueIfBusy`, a busy session holds the input as a follow-up (`queued: true`) and auto-starts it as an ordinary next task once idle; `task_state` events report the queued count. `file` input parts are written to the Session scratchpad and handed to the model as `[attached file: ]` lines (see the request body below) | -| POST | /steer | Mid-run steering: `{text}` queues a message for the running Task (delivered between turns as a standalone `[user_steering]` user message) → 202; 409 `not_running` when no Task is in progress | +| POST | /tasks | Start a Task: `{input: TaskInputPart[], thinkingLevel?, queueIfBusy?}` → 202. With `queueIfBusy`, a busy session holds the input as a follow-up (`queued: true`) and auto-starts it as an ordinary next task once idle; `task_state` events report the queued count. `file` input parts are written to the Session scratchpad and handed to the model as `[attached file: ]` lines (see the request body below). With `goal: {budget?}` the input starts a goal loop instead: it must carry non-empty text (an image alone states no objective), any images it carries fold into the objective as scratchpad path lines whatever the model's vision, and `file` parts are refused — nothing folds them into a re-injected objective — see [Goal mode](/docs/goal-mode) | +| POST | /steer | Mid-run steering: `{text, images?}` queues a message for the running Task (delivered between turns as a standalone `[user_steering]` user message, with its images right behind it) → 202; either field can carry the message on its own, but a request with neither is a 400; 409 `not_running` when no Task is in progress | | POST | /approvals/:toolCallId | Approval decision: `{decision}` is `allow` or `deny` → 204 | | POST | /abort | Interrupt the current Task: 202 when triggered, 204 when idle | | POST | /retry-now | "Retry now" on the reconnect countdown: skips the in-progress backoff wait, firing the next retry immediately (attempt counter unchanged) → 200 `{skipped}` — `skipped:false` is the benign "no wait in progress" case, never an error | diff --git a/packages/docs/content/server-api.zh.md b/packages/docs/content/server-api.zh.md index 7ff12a2..f8a0eb3 100644 --- a/packages/docs/content/server-api.zh.md +++ b/packages/docs/content/server-api.zh.md @@ -161,8 +161,8 @@ Trace 下载对任意成员开放;导入仅限 owner(同 Agent 快照导入 | DELETE | / | 删除 Session(连同 Trace 与暂存文件) | | GET | /messages | 完整 OmniMessage 历史;Task 运行期间响应额外携带 `live`(进行中的流式尾部,见下) | | GET | /stream | SSE 事件流(见下节) | -| POST | /tasks | 发起 Task:`{input: TaskInputPart[], thinkingLevel?, queueIfBusy?}` → 202。带 `queueIfBusy` 时,运行中的 Session 会把输入暂存为跟进消息(`queued: true`),空闲后按序自动作为普通 Task 发出;`task_state` 事件携带排队数。`file` 输入部分会写入该 Session 的 scratchpad,并以 `[attached file: ]` 行交给模型(见下方请求体) | -| POST | /steer | 运行中插话:`{text}` 为运行中的 Task 排队一条消息(作为独立的 `[user_steering]` 用户消息随下一轮送达)→ 202;无 Task 运行返回 409 `not_running` | +| POST | /tasks | 发起 Task:`{input: TaskInputPart[], thinkingLevel?, queueIfBusy?}` → 202。带 `queueIfBusy` 时,运行中的 Session 会把输入暂存为跟进消息(`queued: true`),空闲后按序自动作为普通 Task 发出;`task_state` 事件携带排队数。`file` 类型的输入会写入 Session scratchpad,以 `[attached file: <路径>]` 行交给模型(见下方请求体)。带 `goal: {budget?}` 时该输入转为发起目标循环:必须含非空文字(一张图说明不了目标),随行的图片一律折叠成 scratchpad 路径行写入目标文本、与模型是否支持视觉无关,而 `file` 会被拒绝——没有东西能把它折进每轮重注入的目标里——见[目标模式](/docs/goal-mode) | +| POST | /steer | 运行中插话:`{text, images?}` 为运行中的 Task 排队一条消息(作为独立的 `[user_steering]` 用户消息随下一轮送达,图片紧随其后)→ 202;两个字段任一非空即可成消息,都为空则 400;无 Task 运行返回 409 `not_running` | | POST | /approvals/:toolCallId | 审批决定:`{decision}` 取 `allow` 或 `deny` → 204 | | POST | /abort | 中断当前 Task:已触发返回 202,无任务返回 204 | | POST | /retry-now | 重连倒计时上的「立即重试」:跳过进行中的退避等待、立刻发起下一次重试(重试计数不变)→ 200 `{skipped}`——`skipped:false` 表示当前没有等待可跳过(良性空操作,非错误) | diff --git a/packages/docs/content/web-app.en.md b/packages/docs/content/web-app.en.md index 6047436..1a53757 100644 --- a/packages/docs/content/web-app.en.md +++ b/packages/docs/content/web-app.en.md @@ -49,7 +49,8 @@ There are four approval modes: `allow-all`, `deny-all`, `read-only` (only read-o - Enter sends, Shift+Enter inserts a newline, and images can be pasted; - The "+" menu holds the input add-ons: **image upload**, **file attachment** and goal mode. An attachment can be any type (up to 20 at a time, ≤ 10MB each and 12MB in total; an oversize pick is refused before it is read, so nothing is uploaded to earn the rejection); selected files show as removable chips above the text body in the order they were picked, and a message with attachments and no text is sendable. On send the files are written into the Session's scratchpad — deleted with the Session — and the message gains an `[attached file: ]` line per file, which the conversation renders as an "Attached files" notice: the bytes never enter the conversation, the model opens each file by path with its ordinary file tools; - Typing `/` opens the slash menu: trigger context compaction (`/compact`), hand the conversation over to another Agent (`/agent`), switch the model (`/model`) — both switch commands appear in an active session only, since a draft has nothing to switch and picks its Agent and model up front — or toggle installed Skills — chosen Skills are sent along with the message in a `[use_skills]` block; -- While a Task is running the input stays live and the toolbar keeps a single action button: an empty composer shows **Stop**, and typing turns it into **Send**, whose behavior follows the **mid-run send mode** from the toolbar's More-settings popover (a compact extensible settings panel, also available in draft state; the choice is remembered): **Steer** (default) delivers the text mid-run as a `[user_steering]` user message with the next turn, **Queue** holds the whole message server-side as a follow-up and auto-sends it as an ordinary new message when the run finishes (an "N queued" hint shows near the input until then; the queue survives page reloads); +- While a Task is running the input stays live and the toolbar keeps a single action button: an empty composer shows **Stop**, and typing turns it into **Send**, whose behavior follows the **mid-run send mode** from the toolbar's More-settings popover (a compact extensible settings panel, also available in draft state; the choice is remembered): **Steer** (default) delivers the text **and any attached images** mid-run as a `[user_steering]` user message with the next turn (an image with no caption sends on its own, and the delivered message renders as a steering chip with its images inside it), **Queue** holds the whole message server-side as a follow-up and auto-sends it as an ordinary new message when the run finishes (an "N queued" hint shows near the input until then; the queue survives page reloads). A draft steering cannot carry — selected skills, file attachments, or a staged `/agent` or `/model` pick, with no text and no image — falls back to the queue for that one send, so the button never sits disabled with Stop displaced; +- **Goal mode** (the "+" menu, or `/goal`) turns the draft into an objective the system loops Tasks against. Attached images ride along and are always sent as scratchpad paths (a goal objective is re-injected as text every round — see [Goal mode](/docs/goal-mode)); text is still required, since a picture alone states no objective, and file attachments are refused. Round 1's bubble shows the attachments in full, later rounds collapse them into a chip that expands on click; - `/agent` and `/model` stage their pick instead of acting on it: the chosen Agent or model becomes a chip above the text body and nothing is sent yet, so you keep typing — Enter/Send is what hands the conversation over (a new chat for that Agent) or forks it onto the chosen model, carrying the text along; with an empty composer a default message is filled in, and the chip's × cancels. Both chips are cached with the draft, so they survive a reload or a trip to another conversation together with the text. A model fork additionally waits for the Session to be idle — it continues from the Session's Trace, which a running turn or a compaction is still writing — and a line above the composer says so while it waits; - When human approval is required, tool calls show inline allow/deny buttons in the message stream; the approval mode can be changed mid-Session; - While the engine waits out a reconnect backoff (≥2s), the retry line shows a live countdown to the next attempt with inline **Retry now** (skips the remaining wait) and **Give up** (the ordinary abort) controls; diff --git a/packages/server/src/api/types.ts b/packages/server/src/api/types.ts index dbc7cc1..5e77587 100644 --- a/packages/server/src/api/types.ts +++ b/packages/server/src/api/types.ts @@ -674,8 +674,16 @@ export interface TaskCreateResponse { * task POST). */ export interface SteerRequest { - /** Non-empty message text (trimmed server-side). */ + /** Message text (trimmed server-side); may be empty when `images` carries the message. */ text: string; + /** + * Images sent with the steering message (`data:` or http(s) URLs, same rule as + * `TaskInputPart.image_url`): delivered as user image messages right behind the + * `[user_steering]` text. A model without vision receives them as scratchpad path lines + * instead, exactly as it would a Prompt's images. At least one of `text` / `images` must + * be non-empty. + */ + images?: string[]; } export interface ApprovalDecisionRequest { diff --git a/packages/server/src/http/routes/sessions.ts b/packages/server/src/http/routes/sessions.ts index 5e776fa..44330ac 100644 --- a/packages/server/src/http/routes/sessions.ts +++ b/packages/server/src/http/routes/sessions.ts @@ -79,6 +79,36 @@ const SESSION_CATEGORIES: readonly SessionCategory[] = [ "archived", ]; +/** + * A base64 `data:` URL of an image, in the exact shape core parses it back out of + * (`imagesToScratchpadPaths`): one mime type, the `;base64,` marker, a non-empty base64 body. + * The mime is deliberately unconstrained — core maps the ones it knows to a file extension and + * falls back to `.bin`, and the image tools sniff the magic bytes rather than trusting either. + * + * Checking the body, not just the `data:` prefix, is what keeps the failure here instead of + * three layers down: core turns a data URL it cannot parse into an "[an attached image could + * not be saved and was dropped]" line, which for an HTTP caller means a 202 followed by a + * message quietly missing its picture. The file-attachment field has always validated its own + * payload this way (parseAttachmentPart); this is the same rule for images. + */ +const IMAGE_DATA_URL = /^data:[^;,]+;base64,[A-Za-z0-9+/=\s]+$/; + +/** + * The image-URL rule every image-carrying request field obeys: a `data:` URL the session + * keeps (inline, or written to the scratchpad without vision) or an http(s) URL it references. + * `field` names the offending value in the error, so each caller reads as if it validated + * inline. + */ +function requireImageUrl(url: unknown, field: string): string { + if (typeof url === "string") { + if (url.startsWith("http://") || url.startsWith("https://")) return url; + if (IMAGE_DATA_URL.test(url)) return url; + } + throw badRequest( + `${field} must be an http(s) URL or a base64 data: URL (data:;base64,).`, + ); +} + /** * Resolve a scratchpad file name to an absolute path inside `dir`, or null when it could point * anywhere else (the caller turns that into the same 404 a missing file gets, so a probe learns @@ -133,14 +163,7 @@ function parseTaskInput(body: Record): ParsedTaskInput { return; } if (part.type === "image_url") { - const url = part.imageUrl; - if ( - typeof url !== "string" || - !(url.startsWith("data:") || url.startsWith("http://") || url.startsWith("https://")) - ) { - throw badRequest(`input[${i}].imageUrl only supports data: or http(s) URLs.`); - } - messages.push(imageUrlMessage(url)); + messages.push(imageUrlMessage(requireImageUrl(part.imageUrl, `input[${i}].imageUrl`))); return; } if (part.type === "file") { @@ -157,6 +180,17 @@ function parseTaskInput(body: Record): ParsedTaskInput { return { messages, attachments }; } +/** + * Validate the optional `images` field of a steer request: a list of `data:` / http(s) URLs + * (same rule as a task input's `imageUrl`), absent or empty = a text-only steering message. + */ +function parseSteerImages(body: Record): string[] { + const images = body.images; + if (images === undefined) return []; + if (!Array.isArray(images)) throw badRequest("images must be an array."); + return images.map((url, i) => requireImageUrl(url, `images[${i}]`)); +} + /** * Validate the optional `goal` field of a task request: absent = a regular task (null); * present = goal mode with a token budget (a positive integer, or -1/omitted = unlimited). @@ -491,21 +525,23 @@ export function sessionsRoutes(deps: AppDeps): Hono { // follow-up keeps its level for its auto-start. const thinkingLevel = optionalEnum(body, "thinkingLevel", THINKING_LEVELS); if (goal) { - // Goal mode: the input must be plain non-empty text (its marker-stripped text becomes - // the objective, re-injected every round — images and file attachments have no place in - // the protocol; rejected before any upload is written to disk). + // Goal mode: the input needs non-empty text, since its marker-stripped text becomes the + // objective that every round re-injects and an image on its own doesn't say what the + // goal is. Images can come along — core folds them into `[attached image: ]` lines + // inside the objective (whatever the model's vision) so they survive the rounds. File + // attachments cannot: nothing folds them into the objective, so they are turned away + // here, before any upload is written to disk. const { messages, attachments } = parseTaskInput(body); const text = messages .filter((m) => (m.payload as { type?: string }).type === "text") .map((m) => (m.payload as { text: string }).text) .join("\n") .trim(); - if ( - !text || - attachments.length > 0 || - messages.some((m) => (m.payload as { type?: string }).type !== "text") - ) { - throw badRequest("goal mode requires text-only input (the objective)."); + if (!text) { + throw badRequest("goal mode requires a non-empty text objective."); + } + if (attachments.length > 0) { + throw badRequest("goal mode accepts text and images only (no file attachments)."); } const { sessionId } = await deps.manager.startGoal(row.sessionId, { input: messages, @@ -552,15 +588,27 @@ export function sessionsRoutes(deps: AppDeps): Hono { }); // Mid-run steering: queue a user message for the running Task; core delivers it between - // turns as a standalone `[user_steering]` user message (the model sees it without the loop - // being interrupted). 409 not_running when no Task is in progress — the frontend then falls - // back to a normal task POST. + // turns as a standalone `[user_steering]` user message, with any images following it as + // user image messages (the model sees the whole thing without the loop being interrupted). + // 409 not_running when no Task is in progress — the frontend then falls back to a normal + // task POST. app.post("/:sessionId/steer", async (c) => { const row = resolveSession(c); const body = await readJson(c); const text = typeof body.text === "string" ? body.text.trim() : ""; - if (!text) throw badRequest("text must be a non-empty string."); - deps.manager.steer(row.sessionId, text); + const images = parseSteerImages(body); + // Either half can carry the message on its own: an image with no caption is a complete + // steering message, and so is plain text. + if (!text && images.length === 0) { + throw badRequest("text or images must carry the steering message."); + } + // The wire shape becomes core's: a user text message (omitted when the images are the + // whole message, so the fold's path lines aren't preceded by a blank one) plus one image + // message each — the same input a normal task would carry. + deps.manager.steer(row.sessionId, [ + ...(text ? [userText(text)] : []), + ...images.map((url) => imageUrlMessage(url)), + ]); return c.body(null, 202); }); diff --git a/packages/server/src/runtime/session-manager.ts b/packages/server/src/runtime/session-manager.ts index 19efea6..a1b7a08 100644 --- a/packages/server/src/runtime/session-manager.ts +++ b/packages/server/src/runtime/session-manager.ts @@ -88,8 +88,8 @@ export interface RuntimeSession { compact(opts: { signal: AbortSignal }): AsyncGenerator; /** Whether compaction is possible and why; when not ok, compact() yields no messages (see core ContextEngine.compactability). */ compactability(): CompactAvailability; - /** Queues a mid-run steering message (core `Session.steer`); false when no Task is running. */ - steer(text: string): boolean; + /** Queues a mid-run steering input (core `Session.steer`); false when no Task is running. */ + steer(input: OmniMessage[]): boolean; /** Skips the in-progress reconnect backoff, firing the next retry immediately (core `Session.skipReconnectWait`); false when no wait is in progress. */ skipReconnectWait(): boolean; toolPermission(name: string): "r" | "rw" | undefined; @@ -510,7 +510,7 @@ export class SessionManager { async startGoal( sessionId: string, args: { - /** Round-1 input (text-only, route-validated); its marker-stripped text is the objective. */ + /** Round-1 input (route-validated to carry text; images may ride along); its marker-stripped text is the objective. */ input: OmniMessage[]; budget: number; /** Optional per-goal thinking level: rides every round's Task (route-validated). */ @@ -526,6 +526,10 @@ export class SessionManager { // The objective is the user's own text (leading skill-invocation blocks stripped) — // the same derivation core records in GOAL.yaml; used for the run-state row, the // goal_started event, and as title material. + // `isPlainText` leaves attached images out, so this copy carries no + // `[attached image: ]` lines — which suits its readers, since a status card and a + // generated title read better without absolute scratchpad paths. Core keeps its own + // folded copy (Session.runGoal) for what the rounds actually re-inject. const text = args.input .filter(isPlainText("user")) .map((m) => m.payload.text) @@ -773,15 +777,16 @@ export class SessionManager { } /** - * Mid-run steering: forward the text to the running Session (core delivers it between - * turns as a standalone `[user_steering]` user message — no SSE event of its own; the - * message arrives through the stream the drive loop already publishes). 409 when the - * Session isn't running a Task (idle / compacting / not loaded) or the run finished in the - * race window — the caller falls back to submitting a normal task. + * Mid-run steering: forward the message to the running Session (core delivers it between + * turns as a standalone `[user_steering]` user message followed by its images — no SSE + * event of its own; the messages arrive through the stream the drive loop already + * publishes). 409 when the Session isn't running a Task (idle / compacting / not loaded) + * or the run finished in the race window — the caller falls back to submitting a normal + * task, which carries the same text and images. */ - steer(sessionId: string, text: string): void { + steer(sessionId: string, input: OmniMessage[]): void { const entry = this.entries.get(sessionId); - if (!entry || entry.status !== "running" || !entry.session.steer(text)) { + if (!entry || entry.status !== "running" || !entry.session.steer(input)) { throw new HttpError( 409, "not_running", diff --git a/packages/server/src/services/trace-service.ts b/packages/server/src/services/trace-service.ts index ee49337..97b09de 100644 --- a/packages/server/src/services/trace-service.ts +++ b/packages/server/src/services/trace-service.ts @@ -281,6 +281,8 @@ export class TraceService { // request that resumes after compaction starts yet another Task. let taskIndex = -1; let continuation = false; // The previous round's Request called a tool -> the next request_begin continues the same Task + /** Inside the image run that follows a `[user_steering]` text (see the turn-start rule below). */ + let steeringImages = false; let sawToolCallThisRequest = false; // Compaction interval (compaction_begin..compaction_end): the compaction // request's request_begin/request_end and token_usage all fall inside it (see @@ -360,10 +362,22 @@ export class TraceService { typeof p.text === "string" && parseUserSteeringText(p.text) !== null; if (isSteeringText) continuation = true; + // Images sent with a steering message ride immediately behind its text, exactly as a + // Prompt's images ride behind theirs — and they inherit its exclusion: still the same + // Task, so `steeringImages` keeps the window open across the whole run of them and + // anything else on the main session closes it (an images-only Prompt after a steering + // message is a genuine new turn). A subagent's messages pass through without closing it: + // they belong to another session's stream and say nothing about this one's grouping. + // The Web answers the same "what is one Task" question over the live stream — see + // `openSteering` in web/src/lib/omni/stream-model.ts; the two need to stay in step. + const isImage = !hasOrigin && msg.type === "model_msg" && p.type === "image_url"; + if (!hasOrigin && !isSteeringText && !(isImage && steeringImages)) steeringImages = false; + if (isSteeringText) steeringImages = true; const startsUserTurn = !hasOrigin && !compactionActive && !isSteeringText && + !steeringImages && msg.type === "model_msg" && ((p.type === "text" && p.role === "user") || p.type === "image_url"); const startsCompactionTurn = diff --git a/packages/server/test/goals.test.ts b/packages/server/test/goals.test.ts index 0072a2f..9416ff1 100644 --- a/packages/server/test/goals.test.ts +++ b/packages/server/test/goals.test.ts @@ -10,6 +10,7 @@ import { buildSkillsMessage, emptyTokenCounts, goalFinished, + imageUrlMessage, tokenUsage, userText, } from "@prismshadow/penguin-core"; @@ -154,17 +155,20 @@ describe("SessionManager.startGoal", () => { */ function goalFakeSession( stream: (input: OmniMessage[]) => OmniMessage[], - ): RuntimeSession & { runOpts: RunOpts[] } { + ): RuntimeSession & { runOpts: RunOpts[]; runs: OmniMessage[][] } { const runOpts: RunOpts[] = []; + const runs: OmniMessage[][] = []; return { sessionId: ROW.sessionId, runOpts, + runs, toolPermission: () => "rw", generateTitle: async () => ({ title: null, usage: null }), compactability: () => "ok" as const, steer: () => false, skipReconnectWait: () => false, async *run(input: OmniMessage[], opts) { + runs.push(input); runOpts.push({ ...(opts.thinkingLevel !== undefined ? { thinkingLevel: opts.thinkingLevel } : {}), ...(opts.goal !== undefined ? { goal: opts.goal } : {}), @@ -250,6 +254,31 @@ describe("SessionManager.startGoal", () => { expect((published[0]!.payload as { text: string }).text).toContain("[use_skills]"); }); + it("records the objective without the attached images: the display copy stays path-free", async () => { + // Core folds the attached images into `[attached image: …]` lines inside the objective it + // re-injects each round. The objective recorded here is the one shown to people — status + // card, goal_started, title material — so it keeps the user's words only. + const session = goalFakeSession(() => [goalFinished("complete", 1, 10)]); + const manager = makeManager(session); + const events: ChannelEvent[] = []; + channels.get(ROW.sessionId).subscribe((e) => events.push(e)); + + await manager.startGoal(ROW.sessionId, { + input: [userText("Match this mockup"), imageUrlMessage("data:image/png;base64,aGk=")], + budget: -1, + }); + await waitFor(() => manager.statusOf(ROW.sessionId) === "idle"); + + // The whole input still reaches core (the images included) — only the recorded copy differs. + expect(session.runs[0]).toHaveLength(2); + const started = events + .filter((e) => e.event === "server_event") + .map((e) => JSON.parse(e.data) as { type: string; objective?: string }) + .find((e) => e.type === "goal_started"); + expect(started?.objective).toBe("Match this mockup"); + expect(goals.latestForSession(ROW.sessionId)?.objective).toBe("Match this mockup"); + }); + it("409s while a goal is running (mutual exclusion); runs without a goals repo", async () => { let release: () => void = () => {}; const gate = new Promise((r) => { diff --git a/packages/server/test/session-index.test.ts b/packages/server/test/session-index.test.ts index 7f5d066..339f7ee 100644 --- a/packages/server/test/session-index.test.ts +++ b/packages/server/test/session-index.test.ts @@ -7,7 +7,7 @@ import fs from "node:fs/promises"; import path from "node:path"; import { afterEach, beforeEach, describe, expect, it } from "vitest"; import { sessionMeta, userText } from "@prismshadow/penguin-core"; -import type { SessionMetaPayload } from "@prismshadow/penguin-core"; +import type { OmniMessage, SessionMetaPayload } from "@prismshadow/penguin-core"; import type { ProjectCreateResponse, SessionCreateResponse, @@ -594,14 +594,48 @@ describe("session-index", () => { }); expect(bad.status).toBe(400); } - // Image parts have no place in the re-injected objective. - const image = await api.post(`/api/sessions/${session.sessionId}/tasks`, { + // An image alone states no goal: the objective is re-injected as text every round. + const imageOnly = await api.post(`/api/sessions/${session.sessionId}/tasks`, { + input: [{ type: "image_url", imageUrl: "data:image/png;base64,aGk=" }], + goal: {}, + }); + expect(imageOnly.status).toBe(400); + // With text alongside them the images are fine: they reach the manager, and core folds + // them into the objective as path lines from there. startGoal stands in for the run so + // the assertion is about validation alone — a real goal loop would still be settling + // after the test closed its database. + const started: OmniMessage[][] = []; + t.deps.manager.startGoal = async (sessionId, args) => { + started.push(args.input); + return { sessionId }; + }; + const withText = await api.post(`/api/sessions/${session.sessionId}/tasks`, { input: [ { type: "text", text: "objective" }, { type: "image_url", imageUrl: "data:image/png;base64,aGk=" }, ], goal: {}, }); - expect(image.status).toBe(400); + expect(withText.status).toBe(202); + expect(started[0]?.map((m) => (m.payload as { type: string }).type)).toEqual([ + "text", + "image_url", + ]); + // A file attachment is the one input a goal cannot take, images notwithstanding: nothing + // folds it into the objective every round re-injects, so it is refused before any upload + // is written to disk (startGoal is never reached). + const withFile = await api.post(`/api/sessions/${session.sessionId}/tasks`, { + input: [ + { type: "text", text: "objective" }, + { + type: "file", + fileName: "report.pdf", + dataUrl: `data:application/pdf;base64,${Buffer.from("PDF-BYTES").toString("base64")}`, + }, + ], + goal: {}, + }); + expect(withFile.status).toBe(400); + expect(started).toHaveLength(1); }); }); diff --git a/packages/server/test/session-manager.test.ts b/packages/server/test/session-manager.test.ts index bc69264..ae1783f 100644 --- a/packages/server/test/session-manager.test.ts +++ b/packages/server/test/session-manager.test.ts @@ -23,7 +23,7 @@ import { userText, withOrigin, } from "@prismshadow/penguin-core"; -import type { ApproveFn, OmniMessage } from "@prismshadow/penguin-core"; +import type { ApproveFn, OmniMessage, TextPayload } from "@prismshadow/penguin-core"; import { openDatabase } from "../src/db/database.js"; import { HttpError } from "../src/http/errors.js"; import { SessionsRepo } from "../src/db/repos/sessions.js"; @@ -239,14 +239,14 @@ describe("session-manager", () => { it("steer: forwards to the running session; idle / lost race → 409 not_running", async () => { const steered: string[] = []; const fake = approvalFakeSession("session-1"); - fake.steer = (text: string) => { - steered.push(text); + fake.steer = (input: OmniMessage[]) => { + steered.push((input[0]!.payload as TextPayload).text); return true; }; const manager = makeManager(loaderOf(fake)); const steerErr = (text: string): unknown => { try { - manager.steer("session-1", text); + manager.steer("session-1", [userText(text)]); return null; } catch (e) { return e; diff --git a/packages/server/test/steer.test.ts b/packages/server/test/steer.test.ts index 7336052..12acd7a 100644 --- a/packages/server/test/steer.test.ts +++ b/packages/server/test/steer.test.ts @@ -1,7 +1,7 @@ /** * Integration tests for POST /api/sessions/:id/steer (mid-run steering): - * - 202 while a Task is running, forwarding the trimmed text to the core session; - * - 400 for empty / non-string text; + * - 202 while a Task is running, forwarding the trimmed text and its images to the core session; + * - 400 when neither text nor images carry a message, and for malformed image URLs; * - 409 not_running when the Session is idle (the frontend then falls back to a * normal task POST); * - 404 for foreign/unknown sessions (via the shared resolveSession lookup). @@ -16,15 +16,23 @@ import type { TestApp } from "./helpers.js"; const SID = "session-2026-07-06-10-00-00-ccdd0001"; +/** One recorded steer call: the trimmed text plus the images that rode along with it. */ +/** A recorded steering input, one `text:`/`img:` line per message, in delivered order. */ +const shape = (input: OmniMessage[]): string[] => + input.map((m) => { + const p = m.payload as { type: string; text?: string; image_url?: string }; + return p.type === "image_url" ? `img:${p.image_url}` : `text:${p.text}`; + }); + /** Fake Session that parks on one approval (keeps the Task running) and records steer calls. */ -function steeringFakeSession(sessionId: string, steered: string[]): RuntimeSession { +function steeringFakeSession(sessionId: string, steered: OmniMessage[][]): RuntimeSession { return { sessionId, toolPermission: () => "rw", generateTitle: async () => ({ title: null, usage: null }), compactability: () => "ok" as const, - steer: (text: string) => { - steered.push(text); + steer: (input: OmniMessage[]) => { + steered.push(input); return true; }, skipReconnectWait: () => false, @@ -42,7 +50,7 @@ function steeringFakeSession(sessionId: string, steered: string[]): RuntimeSessi describe("steer route", () => { let t: TestApp; let api: ReturnType; - let steered: string[]; + let steered: OmniMessage[][]; beforeEach(async () => { t = await createTestApp(); @@ -75,17 +83,68 @@ describe("steer route", () => { expect(steered).toEqual([]); }); - it("running → 202, the trimmed text reaches the core session; empty text → 400", async () => { + it("running → 202, the trimmed text reaches the core session; a message with nothing in it → 400", async () => { await t.deps.manager.startTask(SID, [userText("go")]); await waitFor(() => t.deps.manager.pendingApprovalCount(SID) === 1); expect((await api.post(`/api/sessions/${SID}/steer`, { text: " " })).status).toBe(400); expect((await api.post(`/api/sessions/${SID}/steer`, { text: 42 })).status).toBe(400); + expect((await api.post(`/api/sessions/${SID}/steer`, { text: "", images: [] })).status).toBe( + 400, + ); expect(steered).toEqual([]); const ok = await api.post(`/api/sessions/${SID}/steer`, { text: " focus on tests " }); expect(ok.status).toBe(202); - expect(steered).toEqual(["focus on tests"]); + expect(steered.map(shape)).toEqual([["text:focus on tests"]]); + + t.deps.manager.decideApproval(SID, "tc-steer", "allow"); + await waitFor(() => t.deps.manager.statusOf(SID) === "idle"); + }); + + it("images ride along with the steering text — and carry it alone when there is none", async () => { + await t.deps.manager.startTask(SID, [userText("go")]); + await waitFor(() => t.deps.manager.pendingApprovalCount(SID) === 1); + + const png = "data:image/png;base64,AAAA"; + const captioned = await api.post(`/api/sessions/${SID}/steer`, { + text: " look at this ", + images: [png, "https://example.com/shot.png"], + }); + expect(captioned.status).toBe(202); + // An image with no caption is a complete steering message: empty text is accepted here. + const bare = await api.post(`/api/sessions/${SID}/steer`, { text: "", images: [png] }); + expect(bare.status).toBe(202); + // The route hands core the same message list a task input would carry — and drops the + // text message entirely when the images are the whole message, so a fold's path lines + // aren't preceded by an empty line. + expect(steered.map(shape)).toEqual([ + ["text:look at this", `img:${png}`, "img:https://example.com/shot.png"], + [`img:${png}`], + ]); + + // Same URL rule as a task input's imageUrl; a non-array images field is rejected outright. + expect( + (await api.post(`/api/sessions/${SID}/steer`, { text: "x", images: ["/etc/passwd"] })).status, + ).toBe(400); + expect((await api.post(`/api/sessions/${SID}/steer`, { text: "x", images: png })).status).toBe( + 400, + ); + // The data: body is checked here, not left to core: a URL core cannot parse comes back as + // an "could not be saved" line inside the delivered message, which for an HTTP caller is a + // 202 and then a picture quietly missing. These are the shapes that get that far. + for (const bad of [ + "data:image/png", // no ;base64, marker at all + "data:image/png;base64,", // marker, empty body + "data:image/png;base64,not base64!", // body outside the base64 alphabet + "data:,aGk=", // no mime + "data:image/png;charset=utf-8;base64,aGk=", // an extra parameter core's parse rejects + ]) { + expect( + (await api.post(`/api/sessions/${SID}/steer`, { text: "x", images: [bad] })).status, + ).toBe(400); + } + expect(steered).toHaveLength(2); t.deps.manager.decideApproval(SID, "tc-steer", "allow"); await waitFor(() => t.deps.manager.statusOf(SID) === "idle"); diff --git a/packages/server/test/trace-service.test.ts b/packages/server/test/trace-service.test.ts index a596d0d..524af1c 100644 --- a/packages/server/test/trace-service.test.ts +++ b/packages/server/test/trace-service.test.ts @@ -20,6 +20,7 @@ import { toolCall, toolCallOutput, userText, + withOrigin, } from "@prismshadow/penguin-core"; import type { OmniMessage, SessionMetaPayload, TokenCounts } from "@prismshadow/penguin-core"; import { TraceService } from "../src/services/trace-service.js"; @@ -661,6 +662,38 @@ describe("trace-service", () => { expect(a.modelSegments.map((s) => s.taskIndex)).toEqual([0, 0, 0, 1]); }); + // The Web's live-stream twin of this case lives in stream-model.test.ts. + it("Task grouping: images sent with a steering message don't start a Task either (a Prompt's do)", async () => { + const T = (sec: string) => `2026-07-05T10:04:${sec}Z`; + await writeTraceFile(root, P, A, "2026-07-05", S, 13, [ + sessionMeta(metaPayload()), + // Task 0: a normal turn that calls a tool. + at(T("00.000"), userText("q1")), + at(T("01.000"), requestBegin()), + at(T("02.000"), toolCall({ name: "read_file", arguments: "{}", toolCallId: "t1" })), + at(T("02.500"), requestEnd("completed")), + at(T("03.000"), toolCallOutput({ output: "o", toolCallId: "t1" })), + // Steering with two images: core delivers them right behind the text, and the whole + // batch stays inside Task 0 — an image is a turn starter everywhere except here. + at(T("03.200"), userText("[user_steering]\nlike this mock\n[/user_steering]")), + at(T("03.300"), imageUrlMessage("data:image/png;base64,AAAA")), + // A subagent message belongs to another session's stream and says nothing about this + // one's grouping, so it leaves the window open (the Web skips these even earlier). + at(T("03.350"), withOrigin(assistantText("child thinking"), "child-1")), + at(T("03.400"), imageUrlMessage("data:image/png;base64,BBBB")), + at(T("03.500"), requestBegin()), + at(T("04.000"), assistantText("answer 1")), + at(T("04.500"), requestEnd("completed")), + // An images-only Prompt after all that is a genuine new turn. + at(T("20.000"), imageUrlMessage("data:image/png;base64,CCCC")), + at(T("21.000"), requestBegin()), + at(T("22.000"), assistantText("answer 2")), + at(T("22.500"), requestEnd("completed")), + ]); + const a = await service.analyze(P, A, S, 13); + expect(a.requests.map((r) => r.taskIndex)).toEqual([0, 0, 1]); + }); + // Compaction is its own turn: the previous turn called a tool and would // otherwise "continue", but compaction_begin breaks that continuation, so the // compaction request lands on a new taskIndex. A successful compaction splits diff --git a/packages/web/src/features/chat/chat-input.tsx b/packages/web/src/features/chat/chat-input.tsx index d4c44f8..8e15d5e 100644 --- a/packages/web/src/features/chat/chat-input.tsx +++ b/packages/web/src/features/chat/chat-input.tsx @@ -104,6 +104,7 @@ import { skillSlashItems, } from "./skill-use"; import { GOAL_ICON, UNLIMITED_BUDGET, parseBudgetInput } from "./goal-use"; +import { midRunAction } from "./composer-send"; import { PAPERCLIP_ICON } from "./attached-files-banner"; const APPROVAL_MODES: ApprovalMode[] = ["always-ask", "read-only", "allow-all", "deny-all"]; @@ -1178,13 +1179,13 @@ export function ChatInput({ onSend: (input: TaskInputPart[], goal: { budget: number } | null) => Promise; /** * Mid-run steering (session state only): while a Task is running, Enter/send queues the - * trimmed text for the running agent — it is delivered between turns as a standalone - * `[user_steering]` user message. `"queued"` clears the text and shows the queued hint; - * `"not_running"` (409 race with completion) makes the input fall back to its full normal - * send path; `"failed"` keeps the draft. When absent (draft state), the input stays - * send-disabled while running, as before. + * trimmed text **and any attached images** for the running agent — delivered between turns + * as a standalone `[user_steering]` user message followed by its images. `"queued"` clears + * the text and images and shows the queued hint; `"not_running"` (409 race with completion) + * makes the input fall back to its full normal send path; `"failed"` keeps the draft. When + * absent (draft state), the input stays send-disabled while running, as before. */ - onSteer?: (text: string) => Promise<"queued" | "not_running" | "failed">; + onSteer?: (text: string, images: string[]) => Promise<"queued" | "not_running" | "failed">; /** * Count of steering messages already visible in the message stream: the queued hint stays * up until this increases past its value at queue time (i.e. the message was delivered). @@ -1371,9 +1372,18 @@ export function ChatInput({ const running = status === "running"; const compacting = status === "compacting"; + // The draft's "anything sendable at all" rule, shared by canSend / canFollowUp / the + // steer-mode queue fallback and (negated) by the Stop face of the action button. + const draftHasContent = + text.trim().length > 0 || + images.length > 0 || + attachments.length > 0 || + target !== null || + pendingModel !== null || + selectedSkills.length > 0; // Goal mode (engaged via the "+" menu or /goal): the text body becomes the objective. It is - // exclusive with a staged /agent or /model switch (engaging either clears the other) and with - // images (the objective is re-injected every round as plain text); selected skills ride the + // exclusive with a staged /agent or /model switch (engaging either clears the other); attached + // images ride along (core folds them into the objective as path lines) and selected skills ride // round-1 message as a [use_skills] block, exactly like a normal send. const [goalOn, setGoalOn] = useState(false); const [goalBudgetText, setGoalBudgetText] = useState(""); @@ -1407,6 +1417,9 @@ export function ChatInput({ // parseable budget — and an open editor showing an invalid draft disables Send outright: // combined with the editor refusing to close over an invalid draft (below), no click sequence // can fire a goal with a stale committed budget. + // Images may come along with a goal objective (core folds them into `[attached image: …]` + // lines so they survive the rounds), but they don't substitute for the text; file + // attachments cannot — nothing folds those into a re-injected objective. const canSend = !running && !compacting && @@ -1414,16 +1427,10 @@ export function ChatInput({ !modelAuthDead && (goalOn ? text.trim().length > 0 && - images.length === 0 && attachments.length === 0 && goalBudget !== null && !(goalBudgetOpen && goalBudgetDraftInvalid) - : text.trim().length > 0 || - images.length > 0 || - attachments.length > 0 || - target !== null || - pendingModel !== null || - selectedSkills.length > 0); + : draftHasContent); /** * The budget editor is a fixed upward popover. Opening copies the committed value; closing @@ -1482,74 +1489,59 @@ export function ChatInput({ onHandoffTargetChange?.(null); setPendingModel(null); onPendingModelChange?.(null); - // Attachments can't ride a goal (the server rejects non-text goal input): clear any - // already attached, or canSend would stay silently false with the objective looking ready. - setImages([]); + // Images ride a goal (folded into the objective as path lines), file attachments do not + // — the server refuses those, so clear them or canSend would stay silently false with + // the objective looking ready. setAttachments([]); } }, [onHandoffTargetChange, onPendingModelChange], ); - // Mid-run steering: while running, Enter/send queues plain text for the running agent - // (delivered between turns as a [user_steering] user message). Text only — attachments / - // skills / a staged switch stay in the draft for a later normal send (a staged /agent or - // /model chip also blocks steering: the text belongs to the conversation the switch is about - // to open, not to the agent running here). + // Mid-run steering: while running, Enter/send queues the text **and the attached images** + // for the running agent (delivered between turns as a [user_steering] user message followed + // by its images) — so an image with no caption is a complete steering message on its own. + // File attachments and selected skills stay in the draft for a later normal send: a + // [use_skills] block is task-level setup, not something to hand a turn already under way. A + // staged /agent or /model chip also blocks steering: the text belongs to the conversation + // that switch is about to open, not to the agent running here. // `!goalOn`: with the goal chip engaged the text is an OBJECTIVE — steering it into a run // that happens to be active (e.g. a schedule fired) would silently repurpose it. - const canSteer = - running && - !busy && - !goalOn && - !modelAuthDead && - onSteer !== undefined && - target === null && - pendingModel === null && - text.trim().length > 0; - // Mid-run send mode (owner directive): the user chooses between "steer" (delivered - // mid-run as a [user_steering] input) and "follow-up" (held server-side and auto-sent as - // an ordinary next task once this run finishes). Set from the "+" menu's settings row — - // available in draft state and active sessions alike — and **remembered** across - // sessions/reloads (localStorage, see STEER_MODE_KEY); the running-state send simply - // follows the remembered mode. + // + // Mid-run send mode (owner directive): the user chooses between "steer" (delivered mid-run + // as a [user_steering] input) and "follow-up" (held server-side and auto-sent as an ordinary + // next task once this run finishes). Set from the "+" menu's settings row — available in + // draft state and active sessions alike — and **remembered** across sessions/reloads + // (localStorage, see STEER_MODE_KEY). const [steerMode, setSteerModeState] = useState(initialSteerMode); const setSteerMode = (mode: SteerMode): void => { setSteerModeState(mode); localStorage.setItem(STEER_MODE_KEY, mode); }; const followUpMode = steerMode === "followup" && onQueueFollowUp !== undefined; - // A follow-up is a full normal message: the whole draft (text / attachments / skills / a - // staged switch) is eligible, same content rule as canSend. - // `stagedRoute !== "blocked"`: a staged model fork is never eligible mid-run — the follow-up - // path composes the whole draft and then hands it to onSwitchModel rather than to the queue, - // so without this gate Enter would fork off a Trace that is still being written. - const canFollowUp = - running && - !busy && - !goalOn && - !modelAuthDead && - followUpMode && - stagedRoute !== "blocked" && - (text.trim().length > 0 || - images.length > 0 || - attachments.length > 0 || - target !== null || - pendingModel !== null || - selectedSkills.length > 0); - // The single action button's mode: while running, an **empty** composer means Stop - // (abort); as soon as there is something to send it becomes the send button (steer or - // follow-up per the remembered mode). Idle/compacting is always send. - const canMidRunSend = followUpMode ? canFollowUp : canSteer; - const midRunSendLabel = followUpMode ? S.chat.followUpSend : S.chat.steerSend; - const stopAction = - running && - text.trim().length === 0 && - images.length === 0 && - attachments.length === 0 && - target === null && - pendingModel === null && - selectedSkills.length === 0; + // Which of the two channels this draft can use, or Stop when neither will take it — the whole + // decision lives in midRunAction so it can be reasoned about and tested on its own, and so + // that Stop stays the fallthrough rather than a case somebody has to remember to widen. Only + // meaningful while running; idle/compacting is always Send, gated by canSend above. + const midRun = midRunAction({ + sending: busy, + goalOn, + modelAuthDead, + canSteerChannel: onSteer !== undefined, + canQueueChannel: onQueueFollowUp !== undefined, + followUpMode, + stagedRoute, + hasHandoffTarget: target !== null, + hasPendingModel: pendingModel !== null, + hasText: text.trim().length > 0, + hasImages: images.length > 0, + hasContent: draftHasContent, + }); + const steerAction = running && midRun === "steer"; + const queueAction = running && midRun === "queue"; + const canMidRunSend = steerAction || queueAction; + const midRunSendLabel = midRun === "queue" ? S.chat.followUpSend : S.chat.steerSend; + const stopAction = running && midRun === "stop"; // Queued hint: shown after a successful steer until the message shows up in the stream // (steeringDeliveredCount increases past the baseline captured at queue time) or the run // stops being observable (task no longer running). @@ -1902,11 +1894,15 @@ export function ChatInput({ // from being added since), so there is nothing to carry here. setBusy(true); try { - const ok = await onSend([{ type: "text", text: buildSkillsMessage(selectedSkills, t) }], { - budget: goalBudget!, - }); + // Attached images go with the objective (see the goalOn declaration above). + const goalInput: TaskInputPart[] = [ + { type: "text", text: buildSkillsMessage(selectedSkills, t) }, + ]; + for (const url of images) goalInput.push({ type: "image_url", imageUrl: url }); + const ok = await onSend(goalInput, { budget: goalBudget! }); if (ok) { setText(""); + setImages([]); setSelectedSkills([]); toggleGoal(false); } @@ -1970,37 +1966,41 @@ export function ChatInput({ const send = async () => { if (running) { - // Follow-up branch: the whole draft goes out through the normal composition path, - // but posted with queueIfBusy — the server holds it and auto-sends once this run - // finishes (a staged switch still opens its new chat directly: neither the handoff - // target nor the model fork is the session that is running). - if (followUpMode) { - if (!canFollowUp) return; + // Queue branch: the whole draft goes out through the normal composition path, posted + // with queueIfBusy — the server holds it and auto-sends once this run finishes (a staged + // switch still opens its new chat directly: neither the handoff target nor the model fork + // is the session that is running). One branch for both ways of getting here — follow-up + // mode, and steer mode meeting a draft steering cannot carry — since the message sent is + // the same either way. + if (queueAction) { await sendNormal(onQueueFollowUp!); return; } - // Steering branch: queue the trimmed text for the running agent; only the text is - // sent and cleared — attached images / selected skills stay for a normal send (a - // staged switch chip blocks this branch outright, see canSteer). - if (!canSteer) return; + // Steering branch: queue the trimmed text and the attached images for the running agent; + // both are sent and cleared together — file attachments and selected skills stay for a + // normal send (a staged switch chip blocks this branch outright, see midRunAction). + if (!steerAction) return; const steerText = text.trim(); + const steerImages = images; setBusy(true); let res: "queued" | "not_running" | "failed" = "failed"; try { - res = await onSteer!(steerText); + res = await onSteer!(steerText, steerImages); if (res === "queued") { // Show the "queued" hint until the steering message shows up in the stream // (steeringDeliveredCount increases) — see the effect below. steerBaseline.current = steeringDeliveredCount ?? 0; setSteerPending(true); setText(""); + setImages([]); } } finally { setBusy(false); textareaRef.current?.focus(); } - // Completion race (server: no Task running anymore): deliver the whole draft — images, - // skills and all — through the full normal send path instead of a text-only task. + // Completion race (server: no Task running anymore): deliver the whole draft — skills + // and all — through the full normal send path. The draft is untouched in this branch + // (nothing was cleared), so the images go out with it. if (res === "not_running") await sendNormal(); return; } @@ -2051,9 +2051,6 @@ export function ChatInput({ }; const addFiles = (files: Iterable) => { - // Goal mode is text-only (the objective is re-injected each round): drop image attachments - // outright — including pastes — so send never lands in a silently-disabled state. - if (goalOn) return; for (const file of files) { if (!file.type.startsWith("image/")) continue; const reader = new FileReader(); @@ -2607,7 +2604,6 @@ export function ChatInput({ type="file" accept="image/*" multiple - disabled={goalOn} className="hidden" onChange={onPickFiles} /> @@ -2634,11 +2630,11 @@ export function ChatInput({ icon: IMAGE_ICON, label: S.chat.uploadImage, // Without vision the images still send — as scratchpad file paths — so the - // entry stays usable and the hint explains what will happen instead. - desc: vision ? S.chat.uploadImageDesc : S.chat.imagesAsPathHint, + // entry stays usable and the hint says what will happen instead. Goal mode + // sends them that way on any model, since the objective is re-injected as + // text every round. + desc: vision && !goalOn ? S.chat.uploadImageDesc : S.chat.imagesAsPathHint, active: images.length > 0, - // Goal mode is text-only (the objective is re-injected each round). - disabled: goalOn, onSelect: () => imageInputRef.current?.click(), }, { @@ -2650,7 +2646,8 @@ export function ChatInput({ // inlined into the conversation. desc: S.chat.uploadFileDesc, active: attachments.length > 0, - // Same rule as images: goal input is text-only. + // Unlike images, a file cannot ride a goal: nothing folds it into the + // objective that every round re-injects, so the server refuses it. disabled: goalOn, onSelect: () => attachmentInputRef.current?.click(), }, diff --git a/packages/web/src/features/chat/chat-page.tsx b/packages/web/src/features/chat/chat-page.tsx index c913d3a..0395a57 100644 --- a/packages/web/src/features/chat/chat-page.tsx +++ b/packages/web/src/features/chat/chat-page.tsx @@ -707,16 +707,16 @@ export function ChatPage() { [selected, discardSessionDraft, syncHealedSessionId], ); - // Mid-run steering: the text is queued on the server and delivered between turns as a - // standalone `[user_steering]` user message (visible once it arrives over SSE / from the - // Trace). "not_running" (409) means no Task is in progress anymore (race with completion): - // the input area then falls back to its **full** normal send path — images / skills / the - // whole draft included — rather than a text-only task. + // Mid-run steering: the message is queued on the server and delivered between turns as a + // standalone `[user_steering]` user message followed by its images (visible once they + // arrive over SSE / from the Trace). "not_running" (409) means no Task is in progress + // anymore (race with completion): the input area then falls back to its **full** normal + // send path — skills and the whole draft included — rather than a text+images task. const onSteer = useCallback( - async (text: string): Promise<"queued" | "not_running" | "failed"> => { + async (text: string, images: string[] = []): Promise<"queued" | "not_running" | "failed"> => { if (!selected) return "failed"; try { - await api.postSteer(selected.sessionId, { text }); + await api.postSteer(selected.sessionId, { text, ...(images.length > 0 ? { images } : {}) }); return "queued"; } catch (e) { if (e instanceof ApiError && e.status === 409) return "not_running"; diff --git a/packages/web/src/features/chat/composer-send.ts b/packages/web/src/features/chat/composer-send.ts new file mode 100644 index 0000000..135e062 --- /dev/null +++ b/packages/web/src/features/chat/composer-send.ts @@ -0,0 +1,81 @@ +/** + * What the composer's single action button does while a Task is running. + * + * Pulled out of ChatInput for the same reason as `stagedSendRoute`: the decision has more + * inputs than it looks (two send channels, a mode toggle, a staged switch, a dead model key, + * a goal draft) and every one of them can take a channel away. Inline, that made a specific + * mistake easy to repeat — deriving Stop from "is the composer empty" instead of "does this + * send have anywhere to go". Those agree only while every non-empty draft is sendable, and + * several are not: a goal objective is an objective rather than a message for the turn under + * way, a rejected model key refuses everything, and a staged `/model` fork waits for idle. + * Where they disagreed the button showed a permanently disabled Send *in place of* Stop, so + * the run could not be sent to or stopped without emptying the composer first. + * + * Stop is therefore the fallthrough here, not a case: whatever cannot be sent leaves Stop. + */ +export type MidRunAction = + /** Abort the running Task — the button's face whenever this draft has no send channel. */ + | "stop" + /** Deliver text and images to the running agent as a `[user_steering]` message. */ + | "steer" + /** Hold the whole draft server-side and auto-send it as the next ordinary message. */ + | "queue" + /** A send of this draft is already in flight; the button is inert until it settles. */ + | "disabled"; + +export interface MidRunComposerState { + /** A send started from this composer is in flight (ChatInput's `busy`). */ + sending: boolean; + /** The goal chip is engaged: the body is an objective, which no mid-run channel carries. */ + goalOn: boolean; + /** The model API rejected this Session's credentials — nothing can be sent at all. */ + modelAuthDead: boolean; + /** The host wired a steer channel (`onSteer`); the draft page does not. */ + canSteerChannel: boolean; + /** The host wired a follow-up queue (`onQueueFollowUp`); the draft page does not. */ + canQueueChannel: boolean; + /** Mid-run send mode is "follow-up" rather than "steer" (remembered per user). */ + followUpMode: boolean; + /** Where a staged `/agent` / `/model` chip would send this (see stagedSendRoute). */ + stagedRoute: "post" | "handoff" | "model" | "blocked"; + /** An `/agent` handoff target is staged. */ + hasHandoffTarget: boolean; + /** A `/model` fork target is staged. */ + hasPendingModel: boolean; + /** The body has non-whitespace text. */ + hasText: boolean; + /** At least one image is attached. */ + hasImages: boolean; + /** + * The draft carries anything at all — text, images, file attachments, a staged switch chip + * or selected skills. The queue takes a whole message, so this is its content rule. + */ + hasContent: boolean; +} + +/** + * The button's mode while a Task runs. Idle and compacting are not this function's business: + * the button is always Send there, gated by the ordinary `canSend`. + * + * Steering is preferred over the queue when both could carry the draft, since it reaches the + * turn already under way; the queue picks up what steering cannot — a skills-only draft, file + * attachments, a staged chip — so that "unsendable here" never costs the user Stop. + */ +export function midRunAction(s: MidRunComposerState): MidRunAction { + if (s.sending) return "disabled"; + // A goal draft and a dead key close both channels; a blocked `/model` fork closes the queue + // (see stagedSendRoute) and cannot reach steering anyway, since a staged chip rules it out. + const open = !s.goalOn && !s.modelAuthDead; + const canSteer = + open && + s.canSteerChannel && + !s.hasHandoffTarget && + !s.hasPendingModel && + (s.hasText || s.hasImages); + if (canSteer && !s.followUpMode) return "steer"; + const canQueue = open && s.canQueueChannel && s.stagedRoute !== "blocked" && s.hasContent; + if (canQueue) return "queue"; + // Follow-up mode with a draft the queue refuses: steering is not a silent substitute for it, + // because the two put the message in different places. Stop, as with anything unsendable. + return "stop"; +} diff --git a/packages/web/src/features/chat/goal-banner.tsx b/packages/web/src/features/chat/goal-banner.tsx index 966d4f9..e2c9bfa 100644 --- a/packages/web/src/features/chat/goal-banner.tsx +++ b/packages/web/src/features/chat/goal-banner.tsx @@ -10,13 +10,69 @@ * count, token usage against the budget, and the terminal state once the run ends. The * stop control is the regular abort (one signal spans the whole goal loop server-side). */ +import { useState } from "react"; import { S } from "../../lib/strings"; import { humanizeTokens } from "../../lib/format"; import { GlyphIcon } from "../../components/ui/glyph-icon"; +import { ZoomableImage } from "../../components/ui/image-zoom"; import { GOAL_ICON, UNLIMITED_BUDGET } from "./goal-use"; import type { GoalBannerState } from "./goal-use"; -export function GoalRoundBanner({ round, objective }: { round: number; objective?: string }) { +/** Image glyph (24×24 line path) for the collapsed attachment chip on later goal rounds. */ +const ATTACHMENT_ICON = "M3 5h18v14H3zM3 16l5-5 4 4 3-3 6 6M15.5 8.5a1 1 0 1 1-2 0 1 1 0 0 1 2 0"; + +/** The objective's attachments under a round bubble: thumbnails when `showFull`, a chip otherwise. */ +function GoalRoundImages({ + images, + showFull, + onExpand, +}: { + images: string[]; + showFull: boolean; + onExpand: () => void; +}) { + if (images.length === 0) return null; + if (showFull) { + return ( +
+ {images.map((src, i) => ( + + ))} +
+ ); + } + return ( + + ); +} + +export function GoalRoundBanner({ + round, + objective, + images = [], +}: { + round: number; + objective?: string; + /** Images attached to the objective, restored from its `[attached image: …]` path lines. */ + images?: string[]; +}) { + // The path lines ride the re-injected text, so the images really are in every round's input + // — dropping them after round 1 would misreport what was sent. Showing them full-size every + // time would bury a long goal under the same picture, so later rounds collapse to a chip. + const [expanded, setExpanded] = useState(false); + const showFull = images.length > 0 && (round === 1 || expanded); // A regular right-aligned user bubble (same classes as message-item's user_text // rendering), with the round notice under the bubble. if (objective !== undefined && objective !== "") { @@ -27,6 +83,7 @@ export function GoalRoundBanner({ round, objective }: { round: number; objective {objective}

+ setExpanded(true)} />

{S.chat.goalRoundBanner(round)} diff --git a/packages/web/src/features/chat/message-item.tsx b/packages/web/src/features/chat/message-item.tsx index db6f166..6bc179c 100644 --- a/packages/web/src/features/chat/message-item.tsx +++ b/packages/web/src/features/chat/message-item.tsx @@ -203,12 +203,13 @@ export function MessageItem({ item, ctx }: { item: ChatItem; ctx: StreamRenderCo // files come out of the same pass and collapse into one banner naming them. const { text, images, files } = splitAttachments(skills ? skills.rest : afterScheduled); // Every goal round reads like a normal user message: the body in a user bubble with - // the round notice beneath (the system IS re-sending the user's request each round). + // the round notice beneath (the system IS re-sending the user's request each round) — + // the objective's images included, since they ride it as path lines. if (goalRound) { return ( <> {skills && } - + ); } @@ -218,7 +219,7 @@ export function MessageItem({ item, ctx }: { item: ChatItem; ctx: StreamRenderCo {skills && } {/* Files uploaded with this message: named above the bubble, like the other message-level notices — the bytes live in the session scratchpad, the model - opens them by path (goal mode never gets here: it rejects non-text input). */} + opens them by path (goal mode never gets here: it takes text and images only). */} {files.length > 0 && } {text && (

@@ -253,32 +254,53 @@ export function MessageItem({ item, ctx }: { item: ChatItem; ctx: StreamRenderCo ); } - case "user_steering": + case "user_steering": { // Mid-run steering ([user_steering]-wrapped user text delivered between turns): a // compact right-aligned user-styled chip inside the running Task's flow — visually // lighter than a full prompt bubble, since it doesn't start a new Task (the Trace page // still shows the raw marker text as-is). + // Images sent with the message live inside the same chip: `item.images` on a vision + // model (delivered as image messages right behind the text), or restored from the + // [attached image: …] path lines without one — the same two shapes a user_text bubble + // handles, so both render identically here. + const { text: steerText, images: steerImages } = splitAttachments(item.text); + const shown = [...steerImages, ...(item.images ?? [])]; return (
-
- -

- - {S.chat.userSteering} - - {item.text} -

+
+
+ +

+ + {S.chat.userSteering} + + {steerText} +

+
+ {shown.length > 0 && ( +
+ {shown.map((src, i) => ( + + ))} +
+ )}
); + } case "user_image": return (
diff --git a/packages/web/src/lib/omni/stream-model.ts b/packages/web/src/lib/omni/stream-model.ts index f8a8013..54e6aed 100644 --- a/packages/web/src/lib/omni/stream-model.ts +++ b/packages/web/src/lib/omni/stream-model.ts @@ -101,6 +101,14 @@ export interface UserSteeringItem { kind: "user_steering"; id: number; text: string; + /** + * Images sent with this steering message: core delivers them as ordinary user image + * messages right behind the text, and they are folded in here rather than rendered as + * standalone bubbles — they are part of the same message and must not start a Task. + * (Without vision the images arrive as `[attached image: …]` lines inside `text` instead, + * which the chip restores at render time like any user message.) + */ + images?: string[]; /** Message timestamp (milliseconds): shown on footer hover. */ atMs?: number; } @@ -328,6 +336,13 @@ export interface StreamModel { /** A fragment that has stopped and is waiting to be replaced by the complete message. */ pendingText: AssistantTextItem | null; pendingThinking: ThinkingItem | null; + /** + * The steering chip still collecting its images: core delivers a steering message's images + * as user image messages immediately behind its text, so an image arriving while this is + * set belongs to that chip. Any other message closes the window (see pushMessage) — an + * images-only Prompt sent after a steering message is a genuine new Task. + */ + openSteering: UserSteeringItem | null; /** tool_call_id → tool card (shared by both fragment attribution and complete-message replacement). */ toolCards: Map; /** Direct child Session id → nested model. */ @@ -435,6 +450,7 @@ function newModel(nested: boolean, localDecisions: Set): StreamModel { openThinking: null, pendingText: null, pendingThinking: null, + openSteering: null, toolCards: new Map(), subagents: new Map(), localDecisions, @@ -494,6 +510,13 @@ export function pushMessage( routeNested(model, msg, nowMs); return; } + // A steering message's images arrive as user image messages directly behind its text, with + // nothing interleaved (core delivers the batch in one go) — so anything else on this session + // closes the collection window opened by the chip (see openSteering). Subagent messages + // returned above never reach here, so they leave the window alone. + // The server answers the same "what is one Task" question over the Trace — see + // `steeringImages` in server/src/services/trace-service.ts; the two need to stay in step. + if (!isCompleteUserImage(msg)) model.openSteering = null; if (msg.type === "model_msg") { // Internal messages within a compaction range (between begin and end) // (the compaction prompt, summary output): never rendered, never @@ -539,6 +562,11 @@ export function pushMessage( } } +/** Whether the message is a complete user image — the only kind that can join an open steering chip. */ +function isCompleteUserImage(msg: OmniMessage): boolean { + return msg.type === "model_msg" && (msg.payload as { type?: string }).type === "image_url"; +} + /** * Agent id from a session_meta `agent_state` path: the path is * `//agents//agent_state`, so the agent id is the parent directory @@ -931,12 +959,15 @@ function handleComplete( if (steering !== null) { touchTask(model, timestamp); const steerMs = tsOf(timestamp); - model.items.push({ + const item: UserSteeringItem = { kind: "user_steering", id: nextId(model), text: steering, ...(steerMs !== undefined ? { atMs: steerMs } : {}), - }); + }; + model.items.push(item); + // Open the window for the images core delivers right behind this text. + model.openSteering = item; return; } // A complete text message on the main session's user side: starts a new Task. @@ -1000,6 +1031,13 @@ function handleComplete( return; } case "image_url": { + // An image belonging to the steering message just rendered: it joins that chip and + // leaves the running Task alone — unlike a Prompt's image, it starts nothing. + if (model.openSteering) { + touchTask(model, timestamp); + model.openSteering.images = [...(model.openSteering.images ?? []), p.image_url]; + return; + } startTask(model, timestamp, nowMs); const imgMs = tsOf(timestamp); model.items.push({ diff --git a/packages/web/src/lib/strings-en.ts b/packages/web/src/lib/strings-en.ts index ee1f799..a39b326 100644 --- a/packages/web/src/lib/strings-en.ts +++ b/packages/web/src/lib/strings-en.ts @@ -848,6 +848,9 @@ Scenarios: goalBudgetSave: "Save budget", goalRemove: "Exit goal mode", goalRoundBanner: (round: number): string => `Goal · round ${round}`, + /** Later rounds collapse the objective's images into this chip (round 1 shows them in full). */ + goalRoundImages: (count: number): string => + count === 1 ? "1 attached image" : `${count} attached images`, goalProgress: (rounds: number, tokens: string): string => `round ${rounds} · tokens ${tokens}`, goalStatus: { active: "running", diff --git a/packages/web/src/lib/strings.ts b/packages/web/src/lib/strings.ts index 44f18f3..cd8002d 100644 --- a/packages/web/src/lib/strings.ts +++ b/packages/web/src/lib/strings.ts @@ -827,6 +827,8 @@ Benchmark: goalBudgetSave: "保存预算", goalRemove: "退出目标模式", goalRoundBanner: (round: number): string => `目标 · 第 ${round} 轮`, + /** Later rounds collapse the objective's images into this chip (round 1 shows them in full). */ + goalRoundImages: (count: number): string => `${count} 张附图`, goalProgress: (rounds: number, tokens: string): string => `第 ${rounds} 轮 · tokens ${tokens}`, goalStatus: { active: "进行中", diff --git a/packages/web/test/agent-topology.test.ts b/packages/web/test/agent-topology.test.ts index d01c370..cf253a2 100644 --- a/packages/web/test/agent-topology.test.ts +++ b/packages/web/test/agent-topology.test.ts @@ -331,6 +331,12 @@ describe("identity helpers", () => { // Mid-run steering renders as a user_steering item — inside the running Task, not a new one. pushMessage(m, userText("[user_steering]\nnudge\n[/user_steering]")); expect(taskStartCount(m.items)).toBe(1); + // An image sent WITH that steering message follows it directly and joins its chip, so it + // starts nothing either; the assistant reply then closes the chip's collection window. + pushMessage(m, imageUrlMessage("data:image/png;base64,steered")); + expect(taskStartCount(m.items)).toBe(1); + pushMessage(m, assistantText("r2")); + // A standalone Prompt image (no steering message in front of it) still starts a Task. pushMessage(m, imageUrlMessage("data:image/png;base64,xx")); expect(taskStartCount(m.items)).toBe(2); pushMessage(m, userText("t3")); diff --git a/packages/web/test/composer-send.test.ts b/packages/web/test/composer-send.test.ts new file mode 100644 index 0000000..6af4325 --- /dev/null +++ b/packages/web/test/composer-send.test.ts @@ -0,0 +1,107 @@ +/** + * midRunAction: what the composer's single action button does while a Task is running. + * + * The case this file exists for is the last describe below — a draft the mid-run channels + * refuse must still leave Stop. That failed twice in a row when the rule lived inline in + * ChatInput and was written as "Stop when the composer is empty": an unsendable draft is not + * empty, so Stop vanished behind a Send that could never enable, and the only way to abort was + * to delete what you had typed. + */ +import { describe, expect, it } from "vitest"; +import { midRunAction } from "../src/features/chat/composer-send"; +import type { MidRunComposerState } from "../src/features/chat/composer-send"; + +/** An active session with both channels wired, steer mode, and a one-line text draft. */ +const BASE: MidRunComposerState = { + sending: false, + goalOn: false, + modelAuthDead: false, + canSteerChannel: true, + canQueueChannel: true, + followUpMode: false, + stagedRoute: "post", + hasHandoffTarget: false, + hasPendingModel: false, + hasText: true, + hasImages: false, + hasContent: true, +}; + +const act = (over: Partial = {}) => midRunAction({ ...BASE, ...over }); + +/** The draft shapes that carry nothing at all. */ +const EMPTY = { hasText: false, hasImages: false, hasContent: false } as const; + +describe("midRunAction — steering, the preferred channel", () => { + it("takes text, images, or an image with no caption at all", () => { + expect(act()).toBe("steer"); + expect(act({ hasText: false, hasImages: true })).toBe("steer"); + expect(act({ hasText: true, hasImages: true })).toBe("steer"); + }); + + it("is skipped for a draft it cannot carry, which the queue then takes", () => { + // Skills or file attachments only: hasContent without text or images of its own. + expect(act({ ...EMPTY, hasContent: true })).toBe("queue"); + // A staged switch chip: the text belongs to the conversation that switch is about to open. + expect(act({ hasHandoffTarget: true, stagedRoute: "handoff" })).toBe("queue"); + expect(act({ hasPendingModel: true, stagedRoute: "model" })).toBe("queue"); + }); + + it("is skipped in follow-up mode even when it could carry the draft", () => { + // The two channels put the message in different places; the remembered mode decides. + expect(act({ followUpMode: true })).toBe("queue"); + }); +}); + +describe("midRunAction — the queue", () => { + it("takes the whole draft, whatever it is made of", () => { + expect(act({ followUpMode: true, ...EMPTY, hasContent: true })).toBe("queue"); + }); + + it("refuses a staged /model fork while the Session is still writing its Trace", () => { + // stagedRoute "blocked" is the fork waiting for idle (see stagedSendRoute); steering is + // out too, since a staged chip rules it out — so the button is Stop, not a dead Send. + expect(act({ hasPendingModel: true, stagedRoute: "blocked" })).toBe("stop"); + expect(act({ followUpMode: true, hasPendingModel: true, stagedRoute: "blocked" })).toBe("stop"); + }); + + it("is unavailable on a host that wired no queue (the draft page)", () => { + expect(act({ canQueueChannel: false, ...EMPTY, hasContent: true })).toBe("stop"); + // Steering still works there if the draft suits it. + expect(act({ canQueueChannel: false })).toBe("steer"); + }); +}); + +describe("midRunAction — Stop is the fallthrough, not a case", () => { + it("an empty composer is Stop", () => { + expect(act(EMPTY)).toBe("stop"); + }); + + // The regression this file is really for: each of these is a NON-EMPTY draft that no mid-run + // channel will take. Keyed off "is the composer empty" they produced a permanently disabled + // Send standing where Stop should be, so the run could not be stopped from the composer. + it("a goal objective leaves Stop — it is an objective, not a message for this turn", () => { + expect(act({ goalOn: true })).toBe("stop"); + expect(act({ goalOn: true, followUpMode: true })).toBe("stop"); + expect(act({ goalOn: true, hasImages: true })).toBe("stop"); + }); + + it("a rejected model key leaves Stop, for every draft there is", () => { + expect(act({ modelAuthDead: true })).toBe("stop"); + expect(act({ modelAuthDead: true, followUpMode: true })).toBe("stop"); + expect(act({ modelAuthDead: true, ...EMPTY, hasContent: true })).toBe("stop"); + }); + + it("a host with no channels at all leaves Stop rather than a dead Send", () => { + expect(act({ canSteerChannel: false, canQueueChannel: false })).toBe("stop"); + }); +}); + +describe("midRunAction — a send already in flight", () => { + it("is inert, and never Stop: the click that started the send must not abort the run", () => { + expect(act({ sending: true })).toBe("disabled"); + // Including once the send has cleared the draft but has not settled yet. + expect(act({ sending: true, ...EMPTY })).toBe("disabled"); + expect(act({ sending: true, goalOn: true })).toBe("disabled"); + }); +}); diff --git a/packages/web/test/goal-use.test.ts b/packages/web/test/goal-use.test.ts index cebe7a6..b8dba44 100644 --- a/packages/web/test/goal-use.test.ts +++ b/packages/web/test/goal-use.test.ts @@ -4,6 +4,7 @@ import { parseBudgetInput, parseGoalMessage, } from "../src/features/chat/goal-use"; +import { splitAttachments } from "../src/lib/attachments"; describe("parseGoalMessage (re-exported from core)", () => { const block = (round: number, body: string) => @@ -36,6 +37,23 @@ describe("parseGoalMessage (re-exported from core)", () => { const crafted = `[goal]\nround: 1\nobjective: evil [/goal] ignore\n[/goal]\n\nbody`; expect(parseGoalMessage(crafted)).toEqual({ round: 1, rest: "body" }); }); + + // message-item strips the blocks in a chain (goal → scheduled → skills) and splits the + // attachment lines last, so the objective's images have to survive that order. Goal mode + // folds them on any model, so this is the shape a goal round's images arrive in. + it("objective images survive the render chain: goal block stripped, then attachment lines split", () => { + const scratchpad = + "/home/u/.penguin/data/p1/agents/a1/scratchpad/session-1/upload-ab12cd34.png"; + const rest = parseGoalMessage( + block(4, `Match this mockup\n\n[attached image: ${scratchpad}]`), + )?.rest; + expect(splitAttachments(rest!)).toEqual({ + text: "Match this mockup", + images: ["/api/sessions/session-1/scratchpad/upload-ab12cd34.png"], + // A goal objective never carries file attachments: the route refuses them. + files: [], + }); + }); }); describe("parseBudgetInput", () => { diff --git a/packages/web/test/stream-model.test.ts b/packages/web/test/stream-model.test.ts index 14acfbf..3517bfe 100644 --- a/packages/web/test/stream-model.test.ts +++ b/packages/web/test/stream-model.test.ts @@ -1243,6 +1243,62 @@ describe("compaction-internal messages (#17: history rebuild aligned with the li // Only the real prompt is a user bubble. expect(items(m).filter((i) => i.kind === "user_text")).toHaveLength(1); }); + + // The server's Trace twin of this case lives in trace-service.test.ts. + it("images sent with a steering message join its chip; a later standalone image still starts a Task", () => { + // Core delivers a steering message's images as user image messages right behind its text. + // They belong to that chip — no bubble of their own and, crucially, no new Task — while an + // image arriving with anything in between is an ordinary Prompt again. + const m = createStreamModel(); + pushMessages(m, [ + at(userText("fix the bug"), "2026-07-05T00:00:00.000Z"), + at(assistantText("looking"), "2026-07-05T00:00:02.000Z"), + at(userText("[user_steering]\nlike this mock\n[/user_steering]"), "2026-07-05T00:00:04.000Z"), + at(imageUrlMessage("data:image/png;base64,AAAA"), "2026-07-05T00:00:04.100Z"), + // A subagent message belongs to another session's stream: it routes away before the + // window is touched, so the image after it still joins the chip (the server matches). + at(withOrigin(assistantText("child thinking"), "child1"), "2026-07-05T00:00:04.150Z"), + at(imageUrlMessage("data:image/png;base64,BBBB"), "2026-07-05T00:00:04.200Z"), + at(assistantText("matching the mock"), "2026-07-05T00:00:06.000Z"), + at(tokenUsage(counts(200), counts(100)), "2026-07-05T00:00:07.000Z"), + // A new Prompt that is nothing but an image: a Task of its own. + at(imageUrlMessage("data:image/png;base64,CCCC"), "2026-07-05T00:00:20.000Z"), + at(assistantText("on it"), "2026-07-05T00:00:22.000Z"), + ]); + finalizeHistory(m); + // The subagent gets its own card, but no `user_image` bubble appears before the last one: + // both of the steering message's images went into the chip across it. + expect(items(m).map((i) => i.kind)).toEqual([ + "user_text", + "assistant_text", + "user_steering", + "subagent", + "assistant_text", + "task_stats", + "user_image", + "assistant_text", + "task_stats", + ]); + const steering = items(m).find((i) => i.kind === "user_steering") as UserSteeringItem; + expect(steering.text).toBe("like this mock"); + expect(steering.images).toEqual(["data:image/png;base64,AAAA", "data:image/png;base64,BBBB"]); + }); + + it("an images-only steering message keeps an empty chip text (the images are the message)", () => { + const m = createStreamModel(); + pushMessages(m, [ + at(userText("fix the bug"), "2026-07-05T00:00:00.000Z"), + at(assistantText("looking"), "2026-07-05T00:00:02.000Z"), + at(userText("[user_steering]\n\n[/user_steering]"), "2026-07-05T00:00:04.000Z"), + at(imageUrlMessage("data:image/png;base64,AAAA"), "2026-07-05T00:00:04.100Z"), + at(assistantText("got it"), "2026-07-05T00:00:06.000Z"), + ]); + finalizeHistory(m); + const steering = items(m).find((i) => i.kind === "user_steering") as UserSteeringItem; + expect(steering.text).toBe(""); + expect(steering.images).toEqual(["data:image/png;base64,AAAA"]); + expect(items(m).filter((i) => i.kind === "user_image")).toHaveLength(0); + }); }); describe("elapsed comes from Trace timestamps (#5/#20: settled spans, reload-stable live anchor)", () => {