Adds three bilingual posts, each grounded in primary sources rather than
secondary summaries, plus a new blog category to hold them.
Simple Harness Is All You Need — opens on the counter-intuitive result in
the Databricks coding-agent benchmark: on their cost-versus-pass-rate Pareto
chart, the highest score on the board belongs to Opus 4.8 on the minimal Pi
harness, ahead of the same model on Claude Code at maximum effort for
roughly half the cost per task. Keeps Databricks' own caution and notes that
Pi at max effort lands well below Claude Code at comparable spend. Maps the
result onto PenguinHarness's measured design: six built-in tools with no
file tools, a 72-line system prompt, a 16,000-character output cap, and
compaction into a fresh context.
The Easiest Way to Build AI Agents in 2026 — argues the cost of building an
agent has moved out of the agent and into the stack around it: LangChain to
build, LangGraph to orchestrate, LangSmith or Langfuse to observe and
evaluate, LangGraph Platform to deploy. Five products, two or three vendors,
and a person who becomes the optimization loop.
AI Infrastructure: Past, Present, and Future — the stack used to build AI
(PyTorch, vLLM, Ollama, LlamaFactory) assumes a human operator who carries
state in their head and treats errors as a starting point. It needs no
reinventing for agents; what was missing is the operating knowledge, which
the ollama, vllm and llamafactory skills encode.
Both benchmark comparisons disclose the results we lose as well as the ones
we win, and the framework post names two cases where you should pick
something else.
Also adds the Perspectives / 观点 category with a teal badge, filter chip and
both dictionaries, keeping Tech practice for the hands-on AMD walkthroughs.