fix(core,docs): window-derived output caps and compaction threshold, slower retry ladder, retry-all-but-auth (#235)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
@@ -103,7 +103,7 @@ Goal mode is the exception because its objective is re-injected as text every ro
|
||||
|
||||
## Automatic reconnect
|
||||
|
||||
Every LLM-side failure except `auth` triggers an in-run reconnect — `timeout` (network timeouts, transport disconnects, rate limits, 5xx, transient provider quota errors), `malformed` (truncated streams, JSON parse failures), and **`failed` as well**. `failed` retries even though the classifier judged it non-transient: that judgement is an allowlist of known codes, statuses and message vocabulary, so a gateway phrasing a transient fault its own way (`Upstream HTTP/2 stream failed`, say) lands there and used to kill the turn. Retrying a genuinely permanent error costs the ladder and ends the same way; aborting a transient one destroys the turn. Note this changes the *policy*, not the *taxonomy*: a `failed` request is still recorded as `failed` on its `request_end` and in the Cost center, rather than being relabelled a timeout. On a reconnect the engine re-sends the original input plus a `[turn_retried]` block carrying the previous partial output, so tools are never re-executed. Default limit is 5 reconnects with exponential backoff under a ceiling (base 250ms, cap 30s: 250ms, 500ms, 1s, 2s, 4s ≈ 7.75s of total patience — one shared schedule growing toward the slower retryable classes, with the first steps as fast as a transport blip needs); beyond that the turn settles as `failed`. Each failure's `request_end` announces the planned wait as `retry_in_ms` (same formula as the sleep) and stamps `attempt`, the authoritative 1-based ordinal of the request within its retry run (the CLI and Web App display it verbatim); the Web App renders the wait as a live countdown with "retry now" (skips the remaining wait via `Session.skipReconnectWait` — the attempt counter is unchanged) and "give up" (the ordinary abort; the engine's abort-during-backoff path ends the turn) controls; the CLI prints its own `[retry]` line. All three retryable statuses render identically — a retry the user cannot see is a stalled session with no explanation and no way out. A compaction request is an ordinary LLM request and by default retries on the same cap and ladder (an unusable summary draws on the same budget — see "Context compaction"); a compaction that gives up keeps the original context and tries again at the next trigger. Authentication errors are classified before any retry heuristic and never retry: the request ends with its own terminal status `auth` (only the model reference is fixed at Session creation — credentials are read from the current Project config when the Session loads), and the Web App disables that Session's composer until the model's credential is updated (which auto-unlocks it) or the notice is dismissed for a retry. Tool errors are never retried — they are fed back to the model as `tool_call_output` and the model decides what to do next.
|
||||
Every LLM-side failure except `auth` triggers an in-run reconnect — `timeout` (transport-shaped errors: network timeouts, transport disconnects, rate limits, 5xx), `malformed` (truncated streams, JSON parse failures), and **`failed` as well** — every provider rejection that isn't an explicit credential failure, bare 403s and quota/subscription errors included. The statuses are taxonomy, not policy: the classifier only picks the label, and a gateway phrasing a transient fault its own way (`Upstream HTTP/2 stream failed`, say) or a quota that refills mid-ladder retries exactly like a network drop. Retrying a genuinely permanent error costs the ladder and ends the same way; aborting a transient one destroys the turn. Note this changes the *policy*, not the *taxonomy*: a `failed` request is still recorded as `failed` on its `request_end` and in the Cost center, rather than being relabelled a timeout. On a reconnect the engine re-sends the original input plus a `[turn_retried]` block carrying the previous partial output, so tools are never re-executed. Default limit is 5 reconnects with exponential backoff under a ceiling (base 2s, cap 30s: 2s, 4s, 8s, 16s, 30s ≈ 60s of total patience — one shared schedule for every retryable class, sized so transient provider failures such as restarts and rate limits get a real recovery window instead of five retries burning out in about a second, and so every planned wait clears the Web App's 2s countdown floor and stays visible); beyond that the turn settles as `failed`. Each failure's `request_end` announces the planned wait as `retry_in_ms` (same formula as the sleep) and stamps `attempt`, the authoritative 1-based ordinal of the request within its retry run (the CLI and Web App display it verbatim); the Web App renders the wait as a live countdown with "retry now" (skips the remaining wait via `Session.skipReconnectWait` — the attempt counter is unchanged) and "give up" (the ordinary abort; the engine's abort-during-backoff path ends the turn) controls; the CLI prints its own `[retry]` line. All three retryable statuses render identically — a retry the user cannot see is a stalled session with no explanation and no way out. A compaction request is an ordinary LLM request and by default retries on the same cap and ladder (an unusable summary draws on the same budget — see "Context compaction"); a compaction that gives up keeps the original context and tries again at the next trigger. Authentication errors are classified before any retry heuristic and never retry: the request ends with its own terminal status `auth` (only the model reference is fixed at Session creation — credentials are read from the current Project config when the Session loads), and the Web App disables that Session's composer until the model's credential is updated (which auto-unlocks it) or the notice is dismissed for a retry. Tool errors are never retried — they are fed back to the model as `tool_call_output` and the model decides what to do next.
|
||||
|
||||
## Compaction
|
||||
|
||||
@@ -122,7 +122,7 @@ Three triggers (`compaction_begin.reason`):
|
||||
|
||||
| reason | Condition |
|
||||
| --- | --- |
|
||||
| `context` | last turn's `token_usage.request.total` ≥ `maxContextLength` (default 128000) |
|
||||
| `context` | last turn's `token_usage.request.total` ≥ `maxContextLength` (default 128000; the effective threshold is capped at the model's `context_window` − 2048, so a small-window model — a 32k local vLLM, say — compacts at ~30.7k instead of overflowing the window first; an entry without `context_window` derives from the assumed 128000 default) |
|
||||
| `turns` | Session turn count ≥ `maxSessionTurns` (default -1 = unlimited) |
|
||||
| `manual` | the user runs `/compact` or calls `session.compact()` |
|
||||
|
||||
|
||||
Reference in New Issue
Block a user