feat(blog): a Perspectives category and three posts on harness design and agent infrastructure (#45)

Adds three bilingual posts, each grounded in primary sources rather than
secondary summaries, plus a new blog category to hold them.

Simple Harness Is All You Need — opens on the counter-intuitive result in
the Databricks coding-agent benchmark: on their cost-versus-pass-rate Pareto
chart, the highest score on the board belongs to Opus 4.8 on the minimal Pi
harness, ahead of the same model on Claude Code at maximum effort for
roughly half the cost per task. Keeps Databricks' own caution and notes that
Pi at max effort lands well below Claude Code at comparable spend. Maps the
result onto PenguinHarness's measured design: six built-in tools with no
file tools, a 72-line system prompt, a 16,000-character output cap, and
compaction into a fresh context.

The Easiest Way to Build AI Agents in 2026 — argues the cost of building an
agent has moved out of the agent and into the stack around it: LangChain to
build, LangGraph to orchestrate, LangSmith or Langfuse to observe and
evaluate, LangGraph Platform to deploy. Five products, two or three vendors,
and a person who becomes the optimization loop.

AI Infrastructure: Past, Present, and Future — the stack used to build AI
(PyTorch, vLLM, Ollama, LlamaFactory) assumes a human operator who carries
state in their head and treats errors as a starting point. It needs no
reinventing for agents; what was missing is the operating knowledge, which
the ollama, vllm and llamafactory skills encode.

Both benchmark comparisons disclose the results we lose as well as the ones
we win, and the framework post names two cases where you should pick
something else.

Also adds the Perspectives / 观点 category with a teal badge, filter chip and
both dictionaries, keeping Tech practice for the hands-on AMD walkthroughs.
This commit is contained in:
Yaowei Zheng
2026-07-23 05:04:54 +08:00
committed by GitHub
parent 9767fe60e9
commit 7617374d58
15 changed files with 723 additions and 11 deletions
+11 -1
View File
@@ -6,7 +6,17 @@ The two sites' navbars differed in container width (6xl vs 7xl), the docs-only b
## Blog categories, pinned posts, and page metadata
The blog list stays a single flat list with category badges and filter chips, now across three categories — Product news, Release notes, and the new Tech practice, which the AMD local-agents post moved into. Posts can be pinned to the top via `pinned: true` frontmatter; the launch post introducing PenguinHarness is pinned. A second practice post joined the blog: implementing agent self-improvement with PenguinHarness on an AMD GPU (en + zh), adopted into the same category and author conventions. The detail page moves its metadata below the title: a locale-formatted date ("July 20, 2026" / "2026年7月20日"), the author line (frontmatter `author`, defaulting to Yaowei Zheng (PrismShadow AI)), and a copy-page-link button with a safe clipboard fallback and a transient "Copied" state.
The blog list stays a single flat list with category badges and filter chips, now across four categories — Product news, Release notes, Tech practice (which the AMD local-agents post moved into), and Perspectives. Perspectives holds analysis and opinion rather than hands-on walkthroughs, keeping the practice category for posts you follow along with; it is labelled "Perspectives" / "观点" and carries a teal badge, set apart from the brand-blue product and practice badges and the neutral release-notes one. Posts can be pinned to the top via `pinned: true` frontmatter; the launch post introducing PenguinHarness is pinned. A second practice post joined the blog: implementing agent self-improvement with PenguinHarness on an AMD GPU (en + zh), adopted into the same category and author conventions. The detail page moves its metadata below the title: a locale-formatted date ("July 20, 2026" / "2026年7月20日"), the author line (frontmatter `author`, defaulting to Yaowei Zheng (PrismShadow AI)), and a copy-page-link button with a safe clipboard fallback and a transient "Copied" state.
## Three technical posts on harness design and agent infrastructure
The blog gains three bilingual posts in the new Perspectives category, each sourced against primary material rather than summary:
- **Simple Harness Is All You Need** — opens on the counter-intuitive result in the Databricks coding-agent benchmark: on their cost-versus-pass-rate Pareto chart (reproduced with credit as `blog-assets/databricks-pareto.png`), the highest score on the board belongs to Opus 4.8 running on the minimal Pi harness, ahead of the same model on Claude Code at maximum effort for roughly half the cost per task, with most of the frontier held by Pi. It keeps Databricks' own caution — Pi at `max` effort lands well below Claude Code at comparable spend — and their explanation that Pi sent about a third of the context per turn. The post then maps this onto PenguinHarness's measured design: six built-in tools with no file tools at all, a 72-line system prompt, a 16,000-character output cap, and compaction into a fresh context — before arguing where minimalism must stop, since Pi ships no permission system while per-call approval and Trace auditing are load-bearing here.
- **The Easiest Way to Build AI Agents in 2026** — argues the cost of building an agent has moved out of the agent itself and into the stack assembled around it: on the most popular option that means LangChain to build, LangGraph to orchestrate, LangSmith or Langfuse to observe and evaluate, and LangGraph Platform to deploy — five products across two or three vendors, each with its own concepts and documentation, plus a person who becomes the optimization loop. Compares five representative options against PenguinHarness on facts checked 2026-07-22, notes the field's convergence on thin (AutoGen in maintenance mode, LangChain's legacy surface moved to `langchain-classic`, "harness" adopted as vendor vocabulary by AWS, Microsoft and Anthropic within two months), and sets against it a single install that ships chat, skills, models, usage, Trace and the evaluation center together, then hands the tuning loop to the agent. Includes a section on when not to use PenguinHarness.
- **AI Infrastructure: Past, Present, and Future** — argues that the AI development stack (PyTorch, vLLM, Ollama, LlamaFactory) was built for a human operator who carries state in their head, treats errors as a starting point for investigation, and reads documentation once. It needs no reinventing for agents, since it is already commands and config files; what was missing is the operating knowledge around it, which the shipped `ollama`, `vllm` and `llamafactory` skills encode — check before you change, preflight the constraint that actually binds, verify with an observation, and register the served model so the job is finished rather than started. Closes on what remains unsolved: ML-stack errors still written for humans, GPUs as a shared resource with no reservation protocol, and reproducibility.
The landing blog list test moves to the new post count and the practice-category ordering that follows from it.
## The built-in Skills, listed where people look
+1 -1
View File
@@ -8,7 +8,7 @@ Unreleased.
- [2026-07-22] Skills: vLLM and Ollama deployment plus LlamaFactory fine-tuning join the AI App Development group with a guided serving workflow that follows the user's engine preference, and `agenthub-models` tracks the AgentHub 0.4.1 API. ([details](2026-07-22-skills.md))
- [2026-07-22] Sites: the docs and landing navbars are now identical, the blog gains a Tech-practice category, pinned posts, author/date/copy-link metadata and a second AMD practice post, and the built-in Skills are listed in the READMEs and on the landing page. ([details](2026-07-22-sites-and-blog.md))
- [2026-07-22] Sites: the docs and landing navbars are now identical, the blog gains Tech-practice and Perspectives categories, pinned posts, author/date/copy-link metadata, a second AMD practice post and three bilingual Perspectives posts (harness minimalism against the Databricks benchmark, a sourced comparison of five agent frameworks, and the AI development stack — vLLM, Ollama, LlamaFactory — as infrastructure now driven by agents rather than people), and the built-in Skills are listed in the READMEs and on the landing page. ([details](2026-07-22-sites-and-blog.md))
- [2026-07-22] Docs and examples: two new README roadmap items, and the self-improvement example reworked to genuinely evolve itself. ([details](2026-07-22-docs-and-examples.md))