Adds three bilingual posts, each grounded in primary sources rather than secondary summaries, plus a new blog category to hold them. Simple Harness Is All You Need — opens on the counter-intuitive result in the Databricks coding-agent benchmark: on their cost-versus-pass-rate Pareto chart, the highest score on the board belongs to Opus 4.8 on the minimal Pi harness, ahead of the same model on Claude Code at maximum effort for roughly half the cost per task. Keeps Databricks' own caution and notes that Pi at max effort lands well below Claude Code at comparable spend. Maps the result onto PenguinHarness's measured design: six built-in tools with no file tools, a 72-line system prompt, a 16,000-character output cap, and compaction into a fresh context. The Easiest Way to Build AI Agents in 2026 — argues the cost of building an agent has moved out of the agent and into the stack around it: LangChain to build, LangGraph to orchestrate, LangSmith or Langfuse to observe and evaluate, LangGraph Platform to deploy. Five products, two or three vendors, and a person who becomes the optimization loop. AI Infrastructure: Past, Present, and Future — the stack used to build AI (PyTorch, vLLM, Ollama, LlamaFactory) assumes a human operator who carries state in their head and treats errors as a starting point. It needs no reinventing for agents; what was missing is the operating knowledge, which the ollama, vllm and llamafactory skills encode. Both benchmark comparisons disclose the results we lose as well as the ones we win, and the framework post names two cases where you should pick something else. Also adds the Perspectives / 观点 category with a teal badge, filter chip and both dictionaries, keeping Tech practice for the hands-on AMD walkthroughs.
2.8 KiB
Version 0.2.0
Unreleased.
-
[2026-07-22] Models and core: empty tool lists are omitted from LLM requests (fixing 400s from strict OpenAI-compatible servers), the default system prompt gains service-protection and API-key retry guardrails with the default port as a core SDK constant, a per-model max output tokens cap lands on the Models page, the thinking level moves to a conversation-time picker (low and above) that writes through to Agent settings, subagents inherit the parent session's model and thinking level, sessions record their origin in
session_meta.sourceas the single source of truth, the SDK moves to AgentHub 0.4.1 and its supported-model registry drives a catalog refresh across every provider group (with the READMEs trimmed to the newest generation per vendor), and session-title generation folds into core's internal module. (details) -
[2026-07-22] Web App: the chat sidebar groups conversations by Workspace (with an Agent-mode toggle, group pinning, subagent/scheduled folders and paged loading), the collapsed sidebar becomes an eight-entry navigation rail with bilingual tooltips, the model picker lists key-configured models first, chat renders links in a new tab with clean CJK/URL wrapping and the subagent expansion below the tool's own output, mobile dropdowns stay inside the viewport, the Cost center's daily-token tooltip follows the pointer and shows the cache hit rate, the model and Agent settings forms were tightened, the copied task-stats line is localized, and custom model groups and Agents get initial-letter avatars. (details)
-
[2026-07-22] Skills: vLLM and Ollama deployment plus LlamaFactory fine-tuning join the AI App Development group with a guided serving workflow that follows the user's engine preference, and
agenthub-modelstracks the AgentHub 0.4.1 API. (details) -
[2026-07-22] Sites: the docs and landing navbars are now identical, the blog gains Tech-practice and Perspectives categories, pinned posts, author/date/copy-link metadata, a second AMD practice post and three bilingual Perspectives posts (harness minimalism against the Databricks benchmark, a sourced comparison of five agent frameworks, and the AI development stack — vLLM, Ollama, LlamaFactory — as infrastructure now driven by agents rather than people), and the built-in Skills are listed in the READMEs and on the landing page. (details)
-
[2026-07-22] Docs and examples: two new README roadmap items, and the self-improvement example reworked to genuinely evolve itself. (details)
-
[2026-07-22] Tooling: server query-parameter validation hardening, and unit tests for two previously uncovered core modules. (details)