docs(changelog): 2026-08-06/07 follow-up batch entries (#238)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
@@ -4,12 +4,16 @@ Prices and specs read from each model's provider page on 2026-08-06.
|
||||
|
||||
## New models
|
||||
|
||||
Thinking Machines Lab's Inkling (released 2026-07-14: 1M context, multimodal image + audio input) joins on two gateways: OpenRouter `thinkingmachines/inkling` ($0.95 input / $4.05 output per mtok; the page publishes no cached-input price, so `cache_read` stores the input price per the group's no-discount convention) and Fireworks AI `accounts/fireworks/models/inkling` ($0.17 cached / $1 uncached input / $4.05 output). Fireworks AI also gains `accounts/fireworks/models/deepseek-v4-flash-0731` ($0.028 / $0.14 / $0.28), placed ahead of the undated flash row per the newer-versions-first ordering rule.
|
||||
Thinking Machines Lab's Inkling (released 2026-07-14: 1M context, multimodal image + audio input) joins on two gateways: OpenRouter `thinkingmachines/inkling` ($0.17 cached / $1 uncached input / $4.05 output per mtok, read from the models API — the API is authoritative where a model's web page disagrees) and Fireworks AI `accounts/fireworks/models/inkling` ($0.17 cached / $1 uncached input / $4.05 output). Fireworks AI also gains `accounts/fireworks/models/deepseek-v4-flash-0731` ($0.028 / $0.14 / $0.28), placed ahead of the undated flash row per the newer-versions-first ordering rule.
|
||||
|
||||
## Delisted
|
||||
|
||||
The OpenRouter `z-ai/glm-5.1` and SiliconFlow `Pro/zai-org/GLM-5.1` gateway listings are removed; the Z.AI direct `glm-5.1` stays. Existing Project configs keep working — model entries are copied into `.project_config.toml` at creation and nothing rewrites them on upgrade; delisted presets simply stop being re-added by the models page's "sync presets" action.
|
||||
|
||||
## OpenRouter price refresh (2026-08-07)
|
||||
|
||||
A follow-up one-pass re-read of the whole OpenRouter group from its models API also refreshed rows that had drifted since 2026-08-03: `deepseek/deepseek-v4-flash` ($0.01764 / $0.0882 / $0.1764), `moonshotai/kimi-k2.6` ($0.0992 / $0.589 / $2.48), `qwen/qwen3.6-35b-a3b` (cache-read $0.05, now published), and `z-ai/glm-5.2` ($0.1261 / $0.679 / $2.134). No row carries a cache-write premium; the no-published-cache-price fallback now covers only the `:free` rows. (`inclusionai/ling-3.0-flash:free` no longer appears in the API listing — left in the catalog pending a delisting decision.)
|
||||
|
||||
## Docs and skills
|
||||
|
||||
The agenthub-models skill (v11) adds the Inkling and Fireworks 0731 spellings and stops naming the delisted GLM-5.1 gateway ids; the README (en/zh) supported-models tables gain the Inkling family (OpenRouter, Fireworks AI).
|
||||
|
||||
@@ -1,9 +1,11 @@
|
||||
# Admin "use system HTTP proxy" switch
|
||||
# Admin proxy options: app/agent switches and an explicit proxy address
|
||||
|
||||
The sidebar user menu gains an admin-only, server-global "Use system HTTP proxy" switch (default on, saved immediately, stored in a new `server_settings` table and served by `GET/PUT /api/admin/settings`), implementing the long-specced 出网与系统代理 design.
|
||||
The sidebar user menu gains an admin-only "Proxy options" entry opening a settings dialog — server-global, stored in a new `server_settings` table and served by `GET/PUT /api/admin/settings` — implementing (and then extending) the long-specced 出网与系统代理 design so that using a proxy needs no environment variable at all.
|
||||
|
||||
- On: the server honors `HTTP_PROXY` / `HTTPS_PROXY` / `NO_PROXY` (both spellings). Node's built-in fetch ignores those variables, so the server now routes all of its own outbound traffic — LLM provider requests, the update check, image downloads — through an undici global dispatcher and fetch, installed once at the entry.
|
||||
- Off: the server always connects directly, and the proxy variables (including `ALL_PROXY`; `NO_PROXY` kept) are stripped from agent command subprocesses via a new optional `stripProxyEnv` getter threaded through the SDK's `CreateAgentOptions` → `Environment` (absent = proxy allowed, so standalone SDK/CLI behavior is unchanged).
|
||||
- Either way the effective `NO_PROXY` always includes `localhost,127.0.0.1,::1`, keeping loopback traffic — readiness probes, SSE, workspace previews — off any proxy.
|
||||
- Toggling rebuilds the dispatcher for new connections immediately; no restart. The CLI-hosted server (`penguin web`) inherits the same coverage since it imports the server entry in-process.
|
||||
- Desktop: the shell resolves the OS proxy at launch (Electron `resolveProxy`; PAC `PROXY`/`HTTPS` results, SOCKS deliberately skipped — undici speaks HTTP(S) proxies only) and injects it into the embedded server's environment without overriding explicitly configured values.
|
||||
- Two independent switches share one address:
|
||||
- **Application uses the proxy** (default on) — the server's own outbound traffic (LLM requests, the update check, image fetches). Node's built-in fetch ignores proxy variables, so the server routes all of its own traffic through an undici global dispatcher installed once at the entry. On with an address = that address for both http and https; on without = the `HTTP_PROXY` / `HTTPS_PROXY` environment variables (both spellings); off = always direct.
|
||||
- **Agent environment uses the proxy** (default on) — command-subprocess env policy. On with an address = inject `HTTP_PROXY`/`HTTPS_PROXY` (both spellings) plus the merged `NO_PROXY`, overriding inherited values; on without = pass the host environment through; off = strip the proxy variables (`NO_PROXY` kept). The SDK seam is `proxyEnv?: () => ProxyEnvPolicy | null` (strip / inject / null passthrough; absent = unchanged standalone behavior, subagents inherit).
|
||||
- The proxy address accepts `http://host[:port]`, `https://host[:port]`, or bare `host[:port]` (normalized to `http://…`); anything else is `400 invalid_proxy_url`; empty clears back to "follow the system proxy". The dialog is a form with an explicit Save button (no-op with a toast when nothing changed); validation errors render inline.
|
||||
- In every on-state the effective `NO_PROXY` always includes `localhost,127.0.0.1,::1`, keeping loopback traffic — readiness probes, SSE, workspace previews — off any proxy. Toggling applies to new connections immediately; no restart. The CLI-hosted server (`penguin web`) inherits the same coverage.
|
||||
- Desktop: the shell resolves the OS proxy at launch (Electron `resolveProxy`; PAC `PROXY`/`HTTPS` results, SOCKS deliberately skipped — undici speaks HTTP(S) proxies only) and injects it into the embedded server's environment without overriding explicitly configured values — so on desktop, "follow the system proxy" really means the OS proxy settings.
|
||||
- The interim single `useSystemProxy` switch (never released) is read once as the fallback default for both new switches when their keys are absent.
|
||||
|
||||
@@ -1,4 +1,4 @@
|
||||
# Web App: directory browsing de-raced, Project defaults reach new drafts, update dot explained, focused model catalog, "temporary workspace" unified
|
||||
# Web App: directory browsing de-raced, Project defaults reach new drafts, update dot explained, focused model catalog, protocol hints, "temporary workspace" unified
|
||||
|
||||
## Directory browsing no longer mis-navigates under repeated clicks
|
||||
|
||||
@@ -14,7 +14,11 @@ The sidebar avatar's update indicator now names the available release in a hover
|
||||
|
||||
## Model catalog opens focused on DeepSeek
|
||||
|
||||
The models page now expands only the DeepSeek group by default; other vendor groups and user-defined groups start collapsed. While a search query is active, any group holding matches is force-opened (derived, not persisted) so results can never hide inside a collapsed group; expansion state still resets per visit.
|
||||
The models page now expands only the DeepSeek group by default; other vendor groups and user-defined groups start collapsed. While a search query is active, any group holding matches is force-opened (derived, not persisted) so results can never hide inside a collapsed group. The user's own expand/collapse toggles persist per Project (localStorage, same pattern as the sidebar's group collapse) and are restored on the next visit; with nothing stored, the DeepSeek-only default applies.
|
||||
|
||||
## Model dialog explains protocols instead of warning about them
|
||||
|
||||
The model config dialog's base URL field now carries an in-field grey suffix showing the exact path the client appends to the base URL — `/chat/completions` for OpenAI-compatible entries (gateways, custom groups, DeepSeek/Z.AI/Moonshot/Qwen), `/responses` for OpenAI direct, `/v1/messages` for Anthropic, `/v1beta/models` for Gemini (each verified against the actual SDK request paths) — replacing the wordy "this model speaks the vendor's official protocol" warning. Preset direct-vendor groups' mechanism hints (AgentHub auto-routing, env-var fallback) are unified into one user-facing line: "Only <vendor>'s official API protocol is supported; use a custom model group for OpenAI-compatible endpoints."
|
||||
|
||||
## "Temporary workspace" unified
|
||||
|
||||
|
||||
@@ -0,0 +1,21 @@
|
||||
# Core: unlimited default turn cap, model-window-derived limits, slower and simpler retries, trace integrity
|
||||
|
||||
## Unlimited default turn cap
|
||||
|
||||
A new agent's `system_config.yaml` now defaults to `max_turns: -1` (unlimited) instead of `100`, and the SDK's fallback for an omitted `maxTurns` agrees, so long agent runs are no longer cut off unexpectedly by the per-Task turn cap; a positive integer still caps the Task, and `-1` remains the only accepted non-positive value. Existing agents keep their stored `max_turns` verbatim and adopt the new default via the settings page's "Restore default configuration". Goal mode's 100-round runaway backstop is unchanged, but an explicit `maxRounds: -1` now disables it (internal knob, regression-tested).
|
||||
|
||||
## Limits derived from the model window (vLLM and other small-window endpoints)
|
||||
|
||||
Running against a small-window OpenAI-compatible endpoint (e.g. a local vLLM with `--max-model-len 32768`) used to fail three ways: requests 400'd because the configured `max_tokens` went on the wire regardless of how much of the window the input already occupied, and the compaction threshold (default 128000) was unreachable inside the window. Now (#218):
|
||||
|
||||
- Each request's effective output cap is `min(configured max_tokens, context_window − estimated input − 1024)`, floored at 512 — a no-op for big-window cloud models. Input is estimated from the last request's real `token_usage` plus a character heuristic that errs high; images (including images inside tool outputs) count as a flat allowance rather than raw base64. Entries with no configured `context_window` (or one below 4096) are not clamped; a hard-binding clamp prints a one-line diagnostic.
|
||||
- The effective compaction threshold is `min(configured max_context_length, context_window − 2048)` (headroom for the summary request's own output), replacing the old 75%-of-window rule — a 32k-window model now compacts at ~30.7k instead of never. Derived at use; stored config is never rewritten.
|
||||
- The docs gain a "Local / self-hosted OpenAI-compatible endpoints (e.g. vLLM)" section (`--enable-auto-tool-choice`, `--tool-call-parser`, set the entry's context window to `max_model_len`).
|
||||
|
||||
## Slower, simpler LLM retries
|
||||
|
||||
The retry ladder's base goes 250ms → 2000ms (2s/4s/8s/16s/30s, ≈60s total patience; count and ceiling unchanged), so transient provider failures get a real recovery window and every planned wait clears the web app's countdown display floor. Classification is simplified to "every LLM error retries except auth": explicit auth signals still stop immediately, everything else — bare 403s, 400s, 429s, 5xx, quota/subscription messages, transport errors — rides the ladder and fails only after exhaustion. The quota-detection machinery (`isQuotaExhaustedError` and its message heuristics) is removed. Deliberate tradeoff: genuinely permanent errors now burn the full ladder before surfacing.
|
||||
|
||||
## Trace integrity under concurrent writes
|
||||
|
||||
Trace appends are serialized inside the writer, so concurrent producers (parallel tools, the model stream) can no longer tear a multi-megabyte record — e.g. a base64 image Data URL — into invalid JSONL (#215). Trace reads are best-effort: a malformed middle line in files damaged before the fix is skipped with a truncated stderr diagnostic and every parseable record is kept, so previously corrupted sessions resume and render again (the skip is O(n) even for heavily damaged files); server-side Trace import still validates strictly, and truncated-last-line tolerance is unchanged. The Session-index head reader shares the tolerant path, so a damaged head window no longer blocks reconciliation forever.
|
||||
@@ -1,15 +1,15 @@
|
||||
# Unreleased
|
||||
|
||||
- [2026-08-03] Web App: medium-width chat toolbars keep the running indicator and live statistics separate by collapsing the Agents panel and Workspace actions to accessible icon-only buttons until the large-screen breakpoint. ([details](2026-08-03-chat-toolbar-layout.md))
|
||||
- [2026-08-07] Core: default `max_turns` becomes -1 (unlimited, SDK fallback aligned), request caps and the compaction threshold derive from the model's context window (vLLM-class small windows work; #218), the LLM retry ladder slows to a 2s base with "everything except auth retries" classification, and Trace appends are serialized with best-effort tolerant reads for damaged files (#215). ([details](2026-08-07-core-runtime.md))
|
||||
|
||||
- [2026-08-06] Desktop app: penguin brand icons on every platform, task-completion system notifications (renderer-only, desktop sessions), explicit single-user mode (user/member management rejected with `desktop_single_user` and hidden), and a bundled `penguin` CLI on PATH (automatic on deb; menu-driven install elsewhere, no system Node needed). ([details](2026-08-06-desktop-app.md))
|
||||
|
||||
- [2026-08-06] Admin "use system HTTP proxy" switch: server-wide proxy control (default on, live toggle, loopback exemption), off-state proxy-env stripping for agent subprocesses, OS-proxy injection on desktop. ([details](2026-08-06-system-proxy-switch.md))
|
||||
- [2026-08-06] Admin proxy options: a "Proxy options" dialog (explicit Save) with independent "application" and "agent environment" switches sharing one proxy address (empty = follow the system proxy; no environment variable needed), loopback always exempt, live toggle, OS-proxy injection on desktop. ([details](2026-08-06-system-proxy-switch.md))
|
||||
|
||||
- [2026-08-06] Web App: directory browsing no longer compounds path segments under repeated clicks (listings bound to their own directory, sequenced picker requests, localized `dir_not_found`), saving Project new-chat defaults resets the new-conversation draft (typed text survives, open drafts reseed live), the avatar update dot gains a tooltip and reaches the collapsed rail, the model catalog opens with only DeepSeek expanded (search force-opens matching groups), and the auto-created session workspace is consistently named 临时工作区 / "temporary workspace" (penguin-sdk skill v19). ([details](2026-08-06-web-app.md))
|
||||
- [2026-08-06] Web App: directory browsing no longer compounds path segments under repeated clicks (listings bound to their own directory, sequenced picker requests, localized `dir_not_found`), saving Project new-chat defaults resets the new-conversation draft (typed text survives, open drafts reseed live), the avatar update dot gains a tooltip and reaches the collapsed rail, the model catalog opens with only DeepSeek expanded and remembers the user's toggles per Project, the model dialog shows the protocol path in the base URL field and unified vendor-protocol hints, and the auto-created session workspace is consistently named 临时工作区 / "temporary workspace" (penguin-sdk skill v19). ([details](2026-08-06-web-app.md))
|
||||
|
||||
- [2026-08-06] Landing homepage leads with the desktop app: platform-aware download CTA in the hero, closing CTA repointed, CLI one-liner install moved below the fold to the quick start; `/download` page unchanged. ([details](2026-08-06-landing-desktop-first.md))
|
||||
|
||||
- [2026-08-06] Models: Thinking Machines Lab's Inkling joins on OpenRouter and Fireworks AI, Fireworks AI gains DeepSeek V4 Flash 0731, and the OpenRouter + SiliconFlow GLM-5.1 gateway listings are delisted (Z.AI direct stays; existing Project configs unaffected); agenthub-models skill v11. ([details](2026-08-06-model-catalog-inkling-dsv4-flash-0731.md))
|
||||
- [2026-08-06] Models: Thinking Machines Lab's Inkling joins on OpenRouter and Fireworks AI, Fireworks AI gains DeepSeek V4 Flash 0731, and the OpenRouter + SiliconFlow GLM-5.1 gateway listings are delisted (Z.AI direct stays; existing Project configs unaffected); OpenRouter prices refreshed from the models API on 2026-08-07 (Inkling cached input $0.17, four drifted rows corrected); agenthub-models skill v11. ([details](2026-08-06-model-catalog-inkling-dsv4-flash-0731.md))
|
||||
|
||||
- [2026-08-06] Release tooling: repo versions realigned with the shipped 0.2.1, and the release workflow now refuses a tag push whose version does not match `package.json` (the drift that made every dev build nag about updates); the bump is documented as a release-prep step. ([details](2026-08-06-release-version-guard.md))
|
||||
|
||||
Reference in New Issue
Block a user