Files
penguin-harness/packages/docs/content/agent-loop.en.md
T

11 KiB

title, description
title description
The Agent Loop The context_engine's master flow diagram and a stage-by-stage breakdown — approvals, concurrent tool execution, interrupt carry-over, automatic reconnect and compaction.

The SDK's single execution entry point is session.run(newMessages, opts?): input is the list of new OmniMessages (the Prompt); the return value is an async generator that streams OmniMessage. One run drives one complete Task, until the model produces a final answer with no tool calls.

This page shows the context_engine's overall flow first, then breaks down each stage; the message-level observable timeline and ordering guarantees are on Message Flow & Ordering. Source: packages/core/src/engine/context-engine.ts.

The loop at a glance

session.run(newMessages, { approve, signal })
  │  carry-over from a previous interrupt? → prepend to this run's input
  ▼
┌── turn loop (≤ max_turns, default 100) ───────────────────────┐
│                                                               │
│  request_begin                                                │
│  LLM.streamGenerate(newMessages)                              │
│    ├─ streams partial_* fragments + complete msgs             │
│    ├─ for each complete tool_call:                            │
│    │     approve(toolCall) ──deny──► synthetic aborted output │
│    │          │allow           (approvals sequential;         │
│    │          ▼                 decision audited)             │
│    │     Environment.executeTool ──► runs concurrently,       │
│    │                                 output streams back      │
│    └─ LLMOutcome:                                             │
│         timeout / malformed ──► reconnect within the turn     │
│                    (≤5, with [turn_retried]; tools not rerun) │
│  token_usage + request_end (at LLM-stream end; not waiting   │
│                              for tools)                       │
│                                                               │
│  tool outputs reordered to original call order ──► next turn  │
│  no tool_call this turn? ──► Task ends, run returns           │
│  compaction trigger (context/turns)? ──► summarize/discard    │
│                                          + Trace rotation     │
└───────────────────────────────────────────────────────────────┘

signal fires (any point) ──► emit abort + build carry-over ──► run returns

Every message and event flows to two destinations at once: streamed live to the Human, and written to the Trace.

Inputs and outputs

const agent = await createAgent({ agentId: "default_agent" });
const session = await agent.createSession({ workspaceDir: process.cwd() });

for await (const output of session.run([userText("Clean up the CSV files under data/")], {
  approve: async (toolCall) => "allow",
  signal: abortController.signal,
})) {
  // output: partial_* fragments, complete model_msg, event_msg
}
interface RunOptions {
  signal?: AbortSignal;    // interrupt (e.g. Ctrl-C)
  approve?: ApproveFn;     // per-tool approval; denies everything when omitted (conservative default)
  thinkingLevel?: ThinkingLevelName;   // this run's thinking level (per-turn, carried through reconnect retries; compaction requests keep the default)
}

Lifecycle of a turn

A Task consists of consecutive Requests (turns). Each turn:

  1. emits request_begin;
  2. the LLM streams back: partial_* fragments followed by complete messages;
  3. every complete tool_call triggers exactly one approve callback; the decision is recorded as an approval_decision event;
  4. approved calls run concurrently in the Environment (approvals themselves are one at a time); outputs stream out in completion order;
  5. when the LLM stream ends, its final token_usage is emitted and request_end(status) follows at once — without waiting for tools: still-running tools may emit output after request_end;
  6. once the whole batch is terminal, tool results are reordered to the original call order and become the next turn's input — the next Request never fires before that.

The Task ends when a turn produces no tool_call. A denial produces a synthetic aborted tool output ("Tool call denied by user.") that the model reacts to.

Interruption and carry-over

When signal fires, the engine emits an abort event and returns immediately, while constructing carry-over content for the next run:

  • Case A — the model's output had completed (the turn's tool_calls were committed): finished tool results are re-sent as structured tool_call_outputs; unfinished calls get an [interrupted: tool aborted by user] placeholder, keeping tool_call/output pairing strictly intact;
  • Case B — the model's output was incomplete: the whole turn is flattened into one [turn_aborted] user text carrying whatever partial output existed.

Carry-over enters the model context only — it is never written to the Trace, which records only what actually happened.

Mid-run steering

While a Task is running, the host can queue a user message with session.steer(text) without interrupting the loop: at the next input assembly the engine delivers it as a standalone user text message wrapped in [user_steering]…[/user_steering], sent alongside that turn's tool outputs (or alone as the continuation input when the turn produced no tool calls — the Task keeps going instead of ending). Steering is real user input: written to Trace like any Prompt, yielded to the output stream, and replayed as ordinary turn input on resume; tool outputs are never rewritten. The queue is drained at every input assembly — including right after a mid-run compaction, so steering that arrives during the compaction request is delivered, never swallowed. steer returns false when no Task is running (hosts then submit a normal task); the queue is discarded only when the run exits (abort included).

Automatic reconnect

Only LLM-side timeout (network timeouts, transport disconnects such as a dropped socket, rate limits, 5xx, and transient provider quota/subscription errors like insufficient_user_quota) and malformed (truncated streams, JSON parse failures) trigger an in-run reconnect: the engine re-sends the original input plus a [turn_retried] block carrying the previous partial output, so tools are never re-executed. Default limit is 5 reconnects with exponential backoff under a ceiling (base 250ms, cap 30s: 250ms, 500ms, 1s, 2s, 4s ≈ 7.75s of total patience — one shared schedule growing toward the slower retryable classes, with the first steps as fast as a transport blip needs); beyond that the turn settles as failed. Each failure's request_end announces the planned wait as retry_in_ms (same formula as the sleep), which the Web App renders as a live countdown with "retry now" (skips the remaining wait via Session.skipReconnectWait — the attempt counter is unchanged) and "give up" (the ordinary abort; the engine's abort-during-backoff path ends the turn) controls. Compaction requests use their own tighter cap (3 retries): a failed compaction keeps the original context and retries on the next trigger, so failing fast beats stalling the session. Authentication errors are classified before any retry heuristic and never retry: the request ends with its own terminal status auth (only the model reference is fixed at Session creation — credentials are read from the current Project config when the Session loads), and the Web App disables that Session's composer until the model's credential is updated (which auto-unlocks it) or the notice is dismissed for a retry. Tool errors are never retried — they are fed back to the model as tool_call_output and the model decides what to do next.

Compaction

Compaction settings are filled in from system_config.yaml by the composition layer:

interface CompactionSettings {
  maxContextLength: number;   // context-token threshold (last token_usage's request.total); <=0 disables
  maxSessionTurns: number;    // cumulative Session turn threshold (counted across Tasks); <=0 = unlimited
  mode: "summarize" | "discard";
  prompt: string;             // the Prompt used by summarize compaction
}

Three triggers (compaction_begin.reason):

reason Condition
context last turn's token_usage.request.total ≥ maxContextLength (default 128000)
turns Session turn count ≥ maxSessionTurns (default -1 = unlimited)
manual the user runs /compact or calls session.compact()

Two modes: summarize (default) appends the compaction Prompt to the old context, extracts the [summary], wraps it as a [context_summary] user text and continues in a fresh model context; discard simply drops the old context. System markers are written as [tag]…[/tag]; the earlier angle-bracket form (<summary>, <context_summary>, …) is still recognized when reading old Traces and old persisted compaction prompts. Compaction rotates the Trace file (_002, _003, …) — one Trace file always equals one complete model context. compactability() probes feasibility before session.compact() (ok | unsupported | empty | just_compacted).

The compaction request keeps the session's toolset unchanged — the request prefix (tool list included) stays byte-identical to ordinary turns, so the provider's prompt cache remains valid at the moment the context is largest. Compaction still succeeds only with a valid summary: a response that calls a tool or whose extracted summary is empty is rejected — any tool calls are answered with synthesized failed outputs (keeping tool_use/tool_result pairing intact) and the repaired request is resent immediately (a rejection is model behavior, not a transport failure, so no backoff applies), up to 5 rejected attempts; then the compaction ends failed, keeping the original context and Trace file until the next trigger. Transport timeout/malformed attempts follow the compaction-specific reconnect cap and backoff ladder described under "Automatic reconnect" above.

Concurrency model

  • Within a turn: approvals are sequential, execution is concurrent, and the next turn's input keeps the original order;
  • within a Session: only one Task or one compaction runs at a time (the Server rejects concurrent requests with 409);
  • a Subagent is an independent Session with its own Trace and loop; its messages are forwarded to the parent tagged with origin.

Side channels

  • Session titles: session.generateTitle() is a one-shot out-of-band LLM call (no tools, no system Prompt) that never enters history or Trace;
  • Usage accounting: each turn's token_usage events are persisted row by row by the Server — the raw data behind the cost statistics.