Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
11 KiB
title, description
| title | description |
|---|---|
| The Agent Loop | The context_engine's master flow diagram and a stage-by-stage breakdown — approvals, concurrent tool execution, interrupt carry-over, automatic reconnect and compaction. |
The SDK's single execution entry point is session.run(newMessages, opts?): input is the list of new OmniMessages (the Prompt); the return value is an async generator that streams OmniMessage. One run drives one complete Task, until the model produces a final answer with no tool calls.
This page shows the context_engine's overall flow first, then breaks down each stage; the message-level observable timeline and ordering guarantees are on Message Flow & Ordering. Source: packages/core/src/engine/context-engine.ts.
The loop at a glance
session.run(newMessages, { approve, signal })
│ carry-over from a previous interrupt? → prepend to this run's input
▼
┌── turn loop (≤ max_turns, default 100) ───────────────────────┐
│ │
│ request_begin │
│ LLM.streamGenerate(newMessages) │
│ ├─ streams partial_* fragments + complete msgs │
│ ├─ for each complete tool_call: │
│ │ approve(toolCall) ──deny──► synthetic aborted output │
│ │ │allow (approvals sequential; │
│ │ ▼ decision audited) │
│ │ Environment.executeTool ──► runs concurrently, │
│ │ output streams back │
│ └─ LLMOutcome: │
│ timeout / malformed ──► reconnect within the turn │
│ (≤5, with [turn_retried]; tools not rerun) │
│ token_usage + request_end (at LLM-stream end; not waiting │
│ for tools) │
│ │
│ tool outputs reordered to original call order ──► next turn │
│ no tool_call this turn? ──► Task ends, run returns │
│ compaction trigger (context/turns)? ──► summarize/discard │
│ + Trace rotation │
└───────────────────────────────────────────────────────────────┘
signal fires (any point) ──► emit abort + build carry-over ──► run returns
Every message and event flows to two destinations at once: streamed live to the Human, and written to the Trace.
Inputs and outputs
const agent = await createAgent({ agentId: "default_agent" });
const session = await agent.createSession({ workspaceDir: process.cwd() });
for await (const output of session.run([userText("Clean up the CSV files under data/")], {
approve: async (toolCall) => "allow",
signal: abortController.signal,
})) {
// output: partial_* fragments, complete model_msg, event_msg
}
interface RunOptions {
signal?: AbortSignal; // interrupt (e.g. Ctrl-C)
approve?: ApproveFn; // per-tool approval; denies everything when omitted (conservative default)
thinkingLevel?: ThinkingLevelName; // this run's thinking level (per-turn, carried through reconnect retries; compaction requests keep the default)
}
Lifecycle of a turn
A Task consists of consecutive Requests (turns). Each turn:
- emits
request_begin; - the LLM streams back:
partial_*fragments followed by complete messages; - every complete
tool_calltriggers exactly oneapprovecallback; the decision is recorded as anapproval_decisionevent; - approved calls run concurrently in the Environment (approvals themselves are one at a time); outputs stream out in completion order;
- when the LLM stream ends, its final
token_usageis emitted andrequest_end(status)follows at once — without waiting for tools: still-running tools may emit output afterrequest_end; - once the whole batch is terminal, tool results are reordered to the original call order and become the next turn's input — the next Request never fires before that.
The Task ends when a turn produces no tool_call. A denial produces a synthetic aborted tool output ("Tool call denied by user.") that the model reacts to.
Interruption and carry-over
When signal fires, the engine emits an abort event and returns immediately, while constructing carry-over content for the next run:
- Case A — the model's output had completed (the turn's
tool_calls were committed): finished tool results are re-sent as structuredtool_call_outputs; unfinished calls get an[interrupted: tool aborted by user]placeholder, keepingtool_call/output pairing strictly intact; - Case B — the model's output was incomplete: the whole turn is flattened into one
[turn_aborted]user text carrying whatever partial output existed.
Carry-over enters the model context only — it is never written to the Trace, which records only what actually happened.
Mid-run steering
While a Task is running, the host can queue a user message with session.steer(text) without interrupting the loop: at the next input assembly the engine delivers it as a standalone user text message wrapped in [user_steering]…[/user_steering], sent alongside that turn's tool outputs (or alone as the continuation input when the turn produced no tool calls — the Task keeps going instead of ending). Steering is real user input: written to Trace like any Prompt, yielded to the output stream, and replayed as ordinary turn input on resume; tool outputs are never rewritten. The queue is drained at every input assembly — including right after a mid-run compaction, so steering that arrives during the compaction request is delivered, never swallowed. steer returns false when no Task is running (hosts then submit a normal task); the queue is discarded only when the run exits (abort included).
Automatic reconnect
Only LLM-side timeout (network timeouts, transport disconnects such as a dropped socket, rate limits, 5xx, and transient provider quota/subscription errors like insufficient_user_quota) and malformed (truncated streams, JSON parse failures) trigger an in-run reconnect: the engine re-sends the original input plus a [turn_retried] block carrying the previous partial output, so tools are never re-executed. Default limit is 5 reconnects with exponential backoff under a ceiling (base 250ms, cap 30s: 250ms, 500ms, 1s, 2s, 4s ≈ 7.75s of total patience — one shared schedule growing toward the slower retryable classes, with the first steps as fast as a transport blip needs); beyond that the turn settles as failed. Each failure's request_end announces the planned wait as retry_in_ms (same formula as the sleep), which the Web App renders as a live countdown with "retry now" (skips the remaining wait via Session.skipReconnectWait — the attempt counter is unchanged) and "give up" (the ordinary abort; the engine's abort-during-backoff path ends the turn) controls. Compaction requests use their own tighter cap (3 retries): a failed compaction keeps the original context and retries on the next trigger, so failing fast beats stalling the session. Authentication errors are classified before any retry heuristic and never retry: the request ends with its own terminal status auth (only the model reference is fixed at Session creation — credentials are read from the current Project config when the Session loads), and the Web App disables that Session's composer until the model's credential is updated (which auto-unlocks it) or the notice is dismissed for a retry. Tool errors are never retried — they are fed back to the model as tool_call_output and the model decides what to do next.
Compaction
Compaction settings are filled in from system_config.yaml by the composition layer:
interface CompactionSettings {
maxContextLength: number; // context-token threshold (last token_usage's request.total); <=0 disables
maxSessionTurns: number; // cumulative Session turn threshold (counted across Tasks); <=0 = unlimited
mode: "summarize" | "discard";
prompt: string; // the Prompt used by summarize compaction
}
Three triggers (compaction_begin.reason):
| reason | Condition |
|---|---|
context |
last turn's token_usage.request.total ≥ maxContextLength (default 128000) |
turns |
Session turn count ≥ maxSessionTurns (default -1 = unlimited) |
manual |
the user runs /compact or calls session.compact() |
Two modes: summarize (default) appends the compaction Prompt to the old context, extracts the [summary], wraps it as a [context_summary] user text and continues in a fresh model context; discard simply drops the old context. System markers are written as [tag]…[/tag]; the earlier angle-bracket form (<summary>, <context_summary>, …) is still recognized when reading old Traces and old persisted compaction prompts. Compaction rotates the Trace file (_002, _003, …) — one Trace file always equals one complete model context. compactability() probes feasibility before session.compact() (ok | unsupported | empty | just_compacted).
The compaction request keeps the session's toolset unchanged — the request prefix (tool list included) stays byte-identical to ordinary turns, so the provider's prompt cache remains valid at the moment the context is largest. Compaction still succeeds only with a valid summary: a response that calls a tool or whose extracted summary is empty is rejected — any tool calls are answered with synthesized failed outputs (keeping tool_use/tool_result pairing intact) and the repaired request is resent immediately (a rejection is model behavior, not a transport failure, so no backoff applies), up to 5 rejected attempts; then the compaction ends failed, keeping the original context and Trace file until the next trigger. Transport timeout/malformed attempts follow the compaction-specific reconnect cap and backoff ladder described under "Automatic reconnect" above. The first committed attempt — adopted or rejected — also absorbs whatever turn input was folded into the compaction request (mid-task tool results, or the carry-over a manual /compact folds in) into the old context's history: retries resend only the repairs and the Prompt, and the continuation after an abandoned compaction never resends the absorbed input.
Concurrency model
- Within a turn: approvals are sequential, execution is concurrent, and the next turn's input keeps the original order;
- within a Session: only one Task or one compaction runs at a time (the Server rejects concurrent requests with 409);
- a Subagent is an independent Session with its own Trace and loop; its messages are forwarded to the parent tagged with
origin.
Side channels
- Session titles:
session.generateTitle()is a one-shot out-of-band LLM call (no tools, no system Prompt) that never enters history or Trace; - Usage accounting: each turn's
token_usageevents are persisted row by row by the Server — the raw data behind the cost statistics.