Serve "open in new tab" HTML previews from a separate origin (the loopback counterpart, or PENGUIN_PREVIEW_ORIGIN) with a signed, host-bound token, so localStorage/cookies/third-party embeds work while Agent-generated pages stay off the app origin. The app is canonicalized onto localhost and the preview host (127.0.0.1) serves only /preview/* — its /api answers 401 and app routes 302 to the canonical host — so the preview origin can neither set nor honor a session cookie.
Adds three bilingual posts, each grounded in primary sources rather than
secondary summaries, plus a new blog category to hold them.
Simple Harness Is All You Need — opens on the counter-intuitive result in
the Databricks coding-agent benchmark: on their cost-versus-pass-rate Pareto
chart, the highest score on the board belongs to Opus 4.8 on the minimal Pi
harness, ahead of the same model on Claude Code at maximum effort for
roughly half the cost per task. Keeps Databricks' own caution and notes that
Pi at max effort lands well below Claude Code at comparable spend. Maps the
result onto PenguinHarness's measured design: six built-in tools with no
file tools, a 72-line system prompt, a 16,000-character output cap, and
compaction into a fresh context.
The Easiest Way to Build AI Agents in 2026 — argues the cost of building an
agent has moved out of the agent and into the stack around it: LangChain to
build, LangGraph to orchestrate, LangSmith or Langfuse to observe and
evaluate, LangGraph Platform to deploy. Five products, two or three vendors,
and a person who becomes the optimization loop.
AI Infrastructure: Past, Present, and Future — the stack used to build AI
(PyTorch, vLLM, Ollama, LlamaFactory) assumes a human operator who carries
state in their head and treats errors as a starting point. It needs no
reinventing for agents; what was missing is the operating knowledge, which
the ollama, vllm and llamafactory skills encode.
Both benchmark comparisons disclose the results we lose as well as the ones
we win, and the framework post names two cases where you should pick
something else.
Also adds the Perspectives / 观点 category with a teal badge, filter chip and
both dictionaries, keeping Tech practice for the hands-on AMD walkthroughs.
Branch-length batch covering tooling, the model layer, the Web App and the public
surfaces. Highlights:
- Changelog: a per-release `changelog/<version>/` tree, grouped by the surface each
change touches, with a root CHANGELOG.md holding one line per release.
- Dev startup: `scripts/dev-prebuild.mjs` serializes the skills+core prebuild behind a
lock and keeps `pnpm install` current; `pnpm dev` runs server+web together.
- AgentHub 0.3.3 -> 0.4.0: OmniMessage complete payloads carry one opaque `fidelity`
object in place of item-level `signature`/`phase`, threaded verbatim through Trace,
replay and resume; malformed classification adapted to the new error types.
- Model layer: a model is always referenced by an explicit `(provider, model_id)` pair.
The provider is never inferred, guessed or defaulted -- both the catalog inference and
the unique-match config resolution are gone, and CLI, SDK, server routes and
run_subagent all require the complete pair. Catalog gains the Qwen Token Plan, Qwen
Pay-As-You-Go and Fireworks AI gateways, plus an expanded OpenRouter group.
- Web App: catalog preset sync and per-group speed test on the Models page, positional
slash commands, a markdown renderer, skill-library update reminders, and a vertically
centred draft page whose upward menus size themselves to the room available.
- Public surfaces: restructured READMEs, the penguin.ooo landing site and blog, refreshed
benchmark results for both suites, and the demo videos playing on the landing page.
Includes the fixes from a full review of the branch: 23 confirmed findings, among them a
provider-inference bug that could send one vendor's API key to another vendor's endpoint,
and an Escape handler that destroyed the composer's contents unrecoverably.
Verified on the branch head: pnpm test (1127 passing, 7 packages), pnpm typecheck and
pnpm format:check clean, Playwright e2e 14/14.