Files
penguin-harness/changelog/0.2.1/2026-08-04-compaction-retry-observability.md
Yaowei Zheng 88916880eb release: 0.2.1 (#204)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-05 00:30:55 +08:00

1.5 KiB

Compaction extraction tolerance, retry parity, and failure observability

Fixes #170, where deepseek-v4-flash sessions became unusable: the model wrote [summary][/summary] as a title with the body after the closing tag, extraction took the empty pair, and every retry re-showed the model its own committed bad output — which it copied verbatim, forever. extractSummary now applies a tolerance ladder (first non-empty tag pair → whatever remains after stripping empty summary blocks → tagless output verbatim), preserving byte-identical re-extraction of healthy historical Traces.

A committed-but-unusable compaction response now counts as one more failed attempt on the standard compactionMaxReconnects budget and exponential backoff ladder — only auth stops without retrying — with the tool_use pairing repaired and a corrective note prepended before the re-sent prompt (append-only, prompt-cache-safe). The burned cost becomes visible: compaction_end gains attempts, every attempt's token_usage is surfaced between the compaction event pair, and a failed compaction lands as one cost-center error row. Retry state moves to one shared RetryDetail shape (error_message / attempt / retry_in_ms) across request_end and compaction_end, read directly by the CLI and Web retry displays instead of client-side counting. New agents get a shorter compaction prompt that shows the format as a concrete example; existing agents keep their stored prompt and are covered by the extraction rescue and retry guidance.