更新数据分析 Skill,优化多 Run 评估流程 (#194)
This commit is contained in:
@@ -24,11 +24,11 @@ Before the first dispatch of every new or changed Case, the Builder checks that
|
||||
|
||||
Before each calibration dispatch, the Builder predicts the result produced by the observed Trace strategy, the different result produced by the desired behavior, and the score range affected. Adding another public rule, exception, source, or check that the model can directly execute does not automatically increase difficulty. If both strategies still reach the same scored result, the Builder chooses another refinement.
|
||||
|
||||
The Pilot score is a desired target: meeting it permits an early Freeze; otherwise the Builder completes the configured number of valid Pilot iterations and freezes the lowest-scoring valid revision. The Builder temporarily retains only the current lowest valid revision, then removes that copy and other calibration scaffolding after recording the Formal Baseline. Freeze is followed by a fresh complete Formal matrix. Every valid Formal Baseline is recorded even when its score misses the desired target.
|
||||
Every Pilot iteration runs each Case exactly once. The Pilot score is a desired target: meeting it permits an early Freeze; otherwise the Builder completes the configured number of valid Pilot iterations and freezes the lowest-scoring valid revision. The Builder temporarily retains only the current lowest valid revision and its complete result. After the final consistency review, it records that selected Pilot's one-Run result directly as the Formal Baseline without rerunning or backfilling more Runs, then removes the temporary copy and other calibration scaffolding. A Formal score that misses the desired target does not invalidate the Benchmark.
|
||||
|
||||
After the user confirms that step is complete, they start the second top-level Session in a new conversation. The Optimizer checks the Benchmark and its first complete Formal Baseline before following `agent-optimization`:
|
||||
After the user confirms that step is complete, they start the second top-level Session in a new conversation and specify `runs` per Case for every Candidate. The Optimizer checks the Benchmark and its first complete Formal Baseline before following `agent-optimization`:
|
||||
|
||||
1. orchestrate Evaluators in parallel through `run_subagent`, covering the Case × runs matrix;
|
||||
1. orchestrate Evaluators in parallel through `run_subagent`, covering the Case × user-specified `runs` matrix;
|
||||
2. use scores and linked Traces to propose one bounded Candidate;
|
||||
3. edit the Target Agent's editable state — `AGENTS.md`, Skills, config — to produce version N+1;
|
||||
4. keep the Candidate only when its Evaluation score strictly improves; otherwise roll it back;
|
||||
@@ -36,7 +36,7 @@ After the user confirms that step is complete, they start the second top-level S
|
||||
|
||||
Invalid evaluations and correction reruns do not count toward the round limit. On an execution failure, the Optimizer keeps the same Candidate and repairs only the missing cell; it keeps trying while each attempt follows a new diagnosis and applies a distinct safe repair. Both Builder and Optimizer validate that the complete Evaluator response is plain protocol YAML before reading status or score; if formatting is invalid, that same Evaluator resends from its existing result without rerunning the Target Agent.
|
||||
|
||||
Every accepted Candidate is appended to and verified in the Scoreboard immediately. A strictly higher Evaluation score decides acceptance; whether the predicted Case behavior changed is reported separately so unrelated single-run variation is not presented as causal evidence. Agent optimization requires a complete Formal Baseline in the Scoreboard — without one there is no improvement to compare against.
|
||||
Every accepted Candidate is appended to and verified in the Scoreboard immediately. A strictly higher Evaluation score decides acceptance; the first comparison directly compares the Candidate's multi-Run average with the Formal Baseline's one-Run score without backfilling the Baseline. Whether the predicted Case behavior changed is reported separately so unrelated single-run variation is not presented as causal evidence. Agent optimization requires a complete Formal Baseline in the Scoreboard — without one there is no improvement to compare against.
|
||||
|
||||
## Benchmark storage
|
||||
|
||||
@@ -44,7 +44,7 @@ Benchmarks are stored per Agent under `benchmarks/<id>/`:
|
||||
|
||||
```text
|
||||
benchmarks/<id>/
|
||||
├── benchmark_config.toml # Benchmark configuration (e.g. runs per Case)
|
||||
├── benchmark_config.toml # Benchmark configuration (Builder runs is fixed at 1)
|
||||
├── <case-id>/
|
||||
│ ├── statement/ # the task given to the Target Agent
|
||||
│ └── rubric/ # private scoring rubric, isolated from the Target Agent
|
||||
|
||||
@@ -24,11 +24,11 @@ Builder 和 Optimizer 在各自的顶层 Session 中直接遵循对应 Skill。E
|
||||
|
||||
每轮校准都要在派发前预测:当前 Trace 中的策略会产生什么结果、期望行为会产生什么不同结果,以及会影响多少分。增加一条模型可以直接执行的公开规则、例外、来源或检查项并不会自动增加难度;如果两种策略仍会得到相同的计分结果,就应选择其他改法。
|
||||
|
||||
Pilot 分数是期望目标:达到后可以提前 Freeze;未达到时完成设定数量的有效 Pilot iteration,并选择其中分数最低的有效版本 Freeze。Builder 在临时目录只保留当前最低有效版本,Formal Baseline 记录后清理该副本和校准脚手架。Freeze 后必须运行全新完整的 Formal matrix;只要 Formal 有效就记录 Baseline,分数没有达到期望也不会使 Benchmark 作废。
|
||||
每个 Pilot iteration 对每个 Case 固定只运行一次。Pilot 分数是期望目标:达到后可以提前 Freeze;未达到时完成设定数量的有效 Pilot iteration,并选择其中分数最低的有效版本 Freeze。Builder 在临时目录只保留当前最低有效版本及其完整结果。最终一致性检查通过后,直接把被选中 Pilot 的单次运行结果记录为 Formal Baseline,不再重新运行或补齐更多 Runs;记录后清理临时副本和校准脚手架。Formal 分数没有达到期望也不会使 Benchmark 作废。
|
||||
|
||||
用户确认第一步完成后,在新对话中启动第二个顶层 Session。Optimizer 先检查 Benchmark 和第一条完整 Formal Baseline,再使用 `agent-optimization`:
|
||||
用户确认第一步完成后,在新对话中启动第二个顶层 Session,并指定每个 Candidate、每个 Case 的 `runs`。Optimizer 先检查 Benchmark 和第一条完整 Formal Baseline,再使用 `agent-optimization`:
|
||||
|
||||
1. 通过 `run_subagent` 并行编排 Evaluator,覆盖 Case × 运行次数矩阵;
|
||||
1. 通过 `run_subagent` 并行编排 Evaluator,覆盖 Case × 用户指定 `runs` 的矩阵;
|
||||
2. 根据得分和关联 Trace 提出一个有界 Candidate;
|
||||
3. 编辑 Target Agent 的可编辑状态——`AGENTS.md`、Skills、配置——产出版本 N+1;
|
||||
4. Evaluation 分数严格提升才保留 Candidate,否则回滚;
|
||||
@@ -36,7 +36,7 @@ Pilot 分数是期望目标:达到后可以提前 Freeze;未达到时完成
|
||||
|
||||
无效评测和修复重跑不计入轮数。出现执行失败时,Optimizer 保持同一个 Candidate,只补齐失败单元;只要还能根据新诊断提出不同的安全修复,就继续尝试。Builder 和 Optimizer 都先验证 Evaluator 的完整响应是否为纯协议 YAML,再读取状态或分数;格式不合规时,由同一个 Evaluator 基于已有结果重发,不重新运行 Target Agent。
|
||||
|
||||
每个 Accepted Candidate 立即写入并校验 Scoreboard。Evaluation 分数严格提高决定是否接受;假设是否在预期 Case 上得到支持单独报告,避免把单次运行中的无关波动解释为改动因果。Agent 优化要求 Scoreboard 中已有完整 Formal Baseline——没有基线,就没有可比较的提升。
|
||||
每个 Accepted Candidate 立即写入并校验 Scoreboard。Evaluation 分数严格提高决定是否接受;第一次比较允许直接用 Candidate 的多 Run 平均分比较 Formal Baseline 的单 Run 分数,不为 Baseline 补跑。假设是否在预期 Case 上得到支持单独报告,避免把单次运行中的无关波动解释为改动因果。Agent 优化要求 Scoreboard 中已有完整 Formal Baseline——没有基线,就没有可比较的提升。
|
||||
|
||||
## Benchmark 存储
|
||||
|
||||
@@ -44,7 +44,7 @@ Benchmark 按 Agent 存放在 `benchmarks/<id>/` 下:
|
||||
|
||||
```text
|
||||
benchmarks/<id>/
|
||||
├── benchmark_config.toml # Benchmark 配置(如每个 Case 的运行次数 runs)
|
||||
├── benchmark_config.toml # Benchmark 配置(Builder 的 runs 固定为 1)
|
||||
├── <case-id>/
|
||||
│ ├── statement/ # 交给 Target Agent 的任务描述
|
||||
│ └── rubric/ # 私有评分标准,对 Target Agent 隔离
|
||||
|
||||
@@ -51,7 +51,7 @@ BENCHMARK_DIR = <test_agent_dir>/benchmarks/<benchmark_id>
|
||||
|
||||
Reject path traversal, symlink escape, or any resolved path outside the requested Test Agent. Inspect only the requested Agent State, Benchmark config and Case, isolated Test Workspace, and Traces needed to verify this execution. Do not inspect another Agent, Project secrets, hidden configuration, or unrelated Workspaces or Traces.
|
||||
|
||||
Require `agent_state/system_config.yaml`, `benchmark_config.toml`, `<case_id>/statement/README.md`, and `<case_id>/rubric/README.md`. Check that `runs` is positive and `run` is within `1..runs`. The top-level Agent State `version`, defaulting to 1, must equal `expected_version`; otherwise return `version_changed`. Read and snapshot `model.thinking_level` from this Target Agent config, using the normal Agent-config default `medium` only when the field is absent. This configured value is the evaluation `thinking_level`; do not require or read thinking metadata from a Trace.
|
||||
Require `agent_state/system_config.yaml`, `benchmark_config.toml`, `<case_id>/statement/README.md`, and `<case_id>/rubric/README.md`. Treat `run` only as the caller-owned label for this evaluation and return it unchanged; do not read or validate the total Run count. The top-level Agent State `version`, defaulting to 1, must equal `expected_version`; otherwise return `version_changed`. Read and snapshot `model.thinking_level` from this Target Agent config, using the normal Agent-config default `medium` only when the field is absent. This configured value is the evaluation `thinking_level`; do not require or read thinking metadata from a Trace.
|
||||
|
||||
Before launch, snapshot every file under the Case's `statement/` and `rubric/` directories. Require a usable Rubric whose scoring items total exactly 100 points. Create a unique Workspace under `<test_agent_dir>/workspaces/`, resolve it to an absolute canonical path, and verify that the resolved path remains under that directory. Copy only `statement/` into it. The Test Agent may see the Statement and its own State, but never the Rubric, Gold answers, scoring rules, or Evaluator reasoning.
|
||||
|
||||
|
||||
@@ -3,8 +3,8 @@ name: agent-optimization
|
||||
description: Improve an Agent State through versioned scores and score-linked Traces from a frozen Benchmark.
|
||||
short_description: Improve an Agent from measured Benchmark results.
|
||||
short_description_zh: 根据 Benchmark 结果改进 Agent。
|
||||
version: 8
|
||||
updated: 2026-07-31T04:10:53Z
|
||||
version: 9
|
||||
updated: 2026-08-04T00:00:00Z
|
||||
---
|
||||
|
||||
# Agent Optimization
|
||||
@@ -13,15 +13,15 @@ Improve one Test Agent through an evidence → hypothesis → Candidate → eval
|
||||
|
||||
## Before you start
|
||||
|
||||
If the request does not identify the Test Agent, frozen Benchmark, desired target score, and round limit, ask for the missing inputs. When they are already supplied, proceed without asking the user to restate them.
|
||||
If the request does not identify the Test Agent, frozen Benchmark, desired target score, positive Run count, and round limit, ask for the missing inputs. When they are already supplied, proceed without asking the user to restate them.
|
||||
|
||||
## Goal and contract
|
||||
|
||||
Require an explicit Test Agent, a frozen Benchmark with a complete valid Formal Baseline, a desired target score, and a positive round limit. Read the evaluation `(provider, model_id, thinking_level)` from the complete Evaluation that matches the current Agent State; do not require the user to repeat it. An Evaluation without any part of this runtime is incomplete and cannot be used as a Reference. The top-level Session must provide `run_subagent`, and the current Agent must have the `agent-evaluation` Skill. If a prerequisite is missing, stop and explain what is needed. Do not create the missing Agent, Benchmark, or Baseline, and do not evaluate the Test Agent directly.
|
||||
Require an explicit Test Agent, a frozen Benchmark with a complete valid Formal Baseline, a desired target score, a positive `runs` value, and a positive round limit. `runs` is the number of Runs per Case for every Candidate in this optimization Session. Freeze it for the Session; do not infer it from `benchmark_config.toml` or the Formal Baseline. Read the evaluation `(provider, model_id, thinking_level)` from the complete Evaluation that matches the current Agent State; do not require the user to repeat it. An Evaluation without any part of this runtime is incomplete and cannot be used as a Reference. The top-level Session must provide `run_subagent`, and the current Agent must have the `agent-evaluation` Skill. If a prerequisite is missing, stop and explain what is needed. Do not create the missing Agent, Benchmark, or Baseline, and do not evaluate the Test Agent directly.
|
||||
|
||||
A **Reference** is the Agent State currently kept as best, together with its complete Evaluation on the frozen Benchmark.
|
||||
|
||||
Each round starts from the Reference and tests a bounded, general **Candidate**. Evaluate every Candidate on the same frozen Case × Run matrix and evaluation runtime. Accept it only when the change is admissible, the matrix is complete and valid, and its Evaluation's top-level `score` is strictly higher than the Reference Evaluation's `score`. An accepted Candidate and its Evaluation become the next Reference; otherwise restore the previous Reference. Stop early when the Reference reaches the desired target; otherwise run no more than the requested number of complete valid Candidate rounds.
|
||||
Each round starts from the Reference and tests a bounded, general **Candidate**. Evaluate every Candidate on the frozen Case set with the requested `runs` count and the Reference evaluation runtime. The initial Formal Baseline has one Run per Case; do not rerun or backfill it to the requested count. Compare each Candidate's stored top-level average directly with the current Reference score even when their Run counts differ. Accept the Candidate only when the change is admissible, its Evaluation is complete and valid, and its top-level `score` is strictly higher than the Reference Evaluation's `score`. An accepted Candidate and its Evaluation become the next Reference; otherwise restore the previous Reference. Stop early when the Reference reaches the desired target; otherwise run no more than the requested number of complete valid Candidate rounds.
|
||||
|
||||
## Access and changes
|
||||
|
||||
@@ -49,12 +49,12 @@ Modify only the Test Agent State and the versioned snapshot required to protect
|
||||
|
||||
For each round:
|
||||
|
||||
1. **Establish the Reference.** Confirm that its complete Evaluation uses the frozen Case × Run matrix and evaluation runtime and that its version matches the current Agent State.
|
||||
1. **Establish the Reference.** Confirm that its complete Evaluation covers the frozen Case set, uses the frozen evaluation runtime, and matches the current Agent State version. Do not require its Run count to equal the requested Candidate `runs` count.
|
||||
2. **Diagnose capability gaps.** Compare each Case's `runs[].score` on the fixed `0..100` scale; use the Evaluation's top-level average `score` only for whole-version comparison. Use public Statements, score-linked Test Traces, and prior accepted or rejected attempts to identify observable behaviors that general Agent State changes could improve. Use repeated Runs to distinguish stable behavior from variation.
|
||||
3. **State a falsifiable hypothesis.** Choose the related gaps to address, connect them to a bounded Candidate, and state which observable decisions or artifacts should change and why. A change that only adds analysis steps without predicting a behavioral change is not a useful hypothesis. If the current diagnosis is exhausted, use the remaining public evidence and prior attempts to construct a different admissible Candidate.
|
||||
4. **Create one Candidate from the Reference.** Apply the change and its Candidate version under the construction and rollback rules below. Do not carry rejected Candidate files into the next attempt.
|
||||
5. **Check admissibility.** Confirm that the change is general, uses no private evaluation information, and modifies only permitted Test Agent State.
|
||||
6. **Evaluate the Candidate.** Delegate the complete frozen Case × Run matrix in parallel under the evaluation rules below and assemble all returned cells. Do not modify the Candidate while any cell is in flight.
|
||||
6. **Evaluate the Candidate.** Delegate the complete frozen Case set × requested `runs` matrix in parallel under the evaluation rules below and assemble all returned cells. Do not modify the Candidate while any cell is in flight.
|
||||
7. **Decide.** Accept the Candidate only when every cell is valid and its Evaluation's top-level average `score` is strictly higher than the Reference Evaluation's `score`. Otherwise restore the Reference. Record separately whether the predicted Case behavior changed; a higher Evaluation score accepts the Candidate even when the stated hypothesis was not supported.
|
||||
8. **Persist and continue.** Immediately append and verify every accepted Candidate Evaluation before starting another round. An accepted Candidate becomes the next Reference. Use valid results from rejected Candidates only as evidence for a later hypothesis. Stop when the Reference reaches the desired target. Otherwise complete the requested number of valid Candidate rounds unless infrastructure, contamination, concurrent State changes, or the inability to construct any admissible Candidate creates a concrete blocker. At the round limit, retain the highest-scoring accepted Reference.
|
||||
|
||||
@@ -74,7 +74,7 @@ If the Candidate is rejected or cannot be evaluated, restore the Reference files
|
||||
|
||||
## Delegate evaluation
|
||||
|
||||
For each Case × Run cell, call `run_subagent` with:
|
||||
For each frozen Case, dispatch exactly the requested number of Run cells, using one-based Run indices `1..runs`. Call `run_subagent` for each cell with:
|
||||
|
||||
```text
|
||||
Use the `agent-evaluation` Skill. Run the specified Test Agent on the specified Case exactly once, then score that single execution.
|
||||
@@ -129,4 +129,4 @@ After writing, parse the complete `scoreboard.yaml` and verify the appended Eval
|
||||
|
||||
Every Run and Case score is on the fixed `0..100` scale. Do not write `max_score`. Calculate and write every Case and Evaluation average directly in the Scoreboard: ignore `null` values when averaging cost and write `null` only when all contributing costs are unknown; round `score` averages to two decimal places, `cost` averages to six decimal places, and `duration_ms` averages to the nearest integer. These stored values are authoritative—do not add a server, frontend, script, or consistency check that recomputes or validates them. Do not add an `aggregate` object or use `case_id`, `mean_score`, `mean_cost`, or `mean_duration_ms`. Do not record rejected Candidates in the Scoreboard.
|
||||
|
||||
Report the Baseline and every fully evaluated Candidate with its score, version, change, decision, and Test Session ids. For each Candidate, distinguish the acceptance decision from whether its stated hypothesis was supported by the predicted Case behavior. Include the final retained version, stop reason, and known limitations. Never report a score for an Agent State that was not evaluated.
|
||||
Report the Baseline and every fully evaluated Candidate with its score, Run count, version, change, decision, and Test Session ids. Make the one-Run Formal Baseline and requested Candidate `runs` count explicit. For each Candidate, distinguish the acceptance decision from whether its stated hypothesis was supported by the predicted Case behavior. Include the final retained version, stop reason, and known limitations. Never report a score for an Agent State that was not evaluated.
|
||||
|
||||
@@ -3,13 +3,13 @@ name: benchmark-design
|
||||
description: Design and calibrate a multi-Case capability Benchmark and establish a traceable Formal Baseline.
|
||||
short_description: Design and calibrate an Agent capability Benchmark.
|
||||
short_description_zh: 设计并校准 Agent 能力评测 Benchmark。
|
||||
version: 6
|
||||
updated: 2026-07-31T04:10:53Z
|
||||
version: 7
|
||||
updated: 2026-08-04T00:00:00Z
|
||||
---
|
||||
|
||||
# Benchmark Design
|
||||
|
||||
Build a multi-Case Benchmark for one Test Agent, calibrate its difficulty, and record a complete Formal Baseline.
|
||||
Build a multi-Case Benchmark for one Test Agent, calibrate its difficulty with one Run per Case, and record the selected frozen Pilot as the Formal Baseline.
|
||||
|
||||
This Skill changes the Benchmark, never the Test Agent. It does not run or score the Test Agent. Delegate every evaluation with `run_subagent`, and tell each worker to use `agent-evaluation`. Stop after the Baseline; do not begin optimization.
|
||||
|
||||
@@ -19,9 +19,9 @@ If the request does not identify a Test Agent, target capability, desired baseli
|
||||
|
||||
## Workflow
|
||||
|
||||
- A **Pilot** is a provisional evaluation used to improve the Benchmark. Its results never enter the Scoreboard.
|
||||
- A **Pilot** is a one-Run-per-Case evaluation used to improve the Benchmark. Unselected Pilot results never enter the Scoreboard; the selected result becomes the Formal Baseline after Freeze.
|
||||
- **Freeze** means the Benchmark revision and evaluation settings stop changing.
|
||||
- A **Formal Baseline** is the accepted result of a fresh, complete Case × Run evaluation of the frozen Benchmark on one unchanged Agent State version.
|
||||
- A **Formal Baseline** is the accepted result of the selected complete valid Pilot revision, recorded after that exact revision is frozen on one unchanged Agent State version.
|
||||
|
||||
Follow this order:
|
||||
|
||||
@@ -31,7 +31,7 @@ Follow this order:
|
||||
4. Complete one valid evaluation for every planned Case. Together these results form Pilot iteration 1; finish this complete set before refining any Case.
|
||||
5. For later Pilot iterations, use scores and Traces to reconstruct how the Test Agent solved each Case. A single iteration may refine multiple Cases or difficulty dimensions; rerun every affected Case.
|
||||
6. Freeze the first valid Pilot revision that meets the desired baseline score. If none does within the requested valid-iteration limit, restore and freeze the lowest-scoring valid Pilot revision.
|
||||
7. Run a fresh, complete Case × Run matrix and save it as the Formal Baseline when every cell is valid, the Agent State version remains unchanged, and no known design defect remains. The Formal score does not determine validity.
|
||||
7. Freeze the selected revision and record its complete one-Run-per-Case Pilot result as the Formal Baseline when every cell is valid, the Agent State version remains unchanged, and no known design defect remains. Do not rerun or backfill it. The Formal score does not determine validity.
|
||||
|
||||
## Setup and access
|
||||
|
||||
@@ -72,9 +72,9 @@ Each Case contains:
|
||||
|
||||
Both directories require a `README.md` and may contain supporting files. Do not put Gold answers for evaluated instances, hidden mappings, or private scoring conditions in `statement/`.
|
||||
|
||||
Create `benchmark_config.toml` with `title`, `description`, and `runs = 3`. Use another positive Run count only when requested. Initialize `scoreboard.yaml` with `evaluations: []`.
|
||||
Create `benchmark_config.toml` with `title`, `description`, and `runs = 1`. Benchmark design always uses one Run per Case; do not ask for or accept another Run count. Initialize `scoreboard.yaml` with `evaluations: []`.
|
||||
|
||||
Pass the resolved `(provider, model_id)` explicitly in every Pilot and Formal Evaluator request, starting with the first cell. Freeze that pair and the Test Agent's configured `thinking_level` for the complete Benchmark workflow. Every scored Evaluator result must report the requested pair and the same configured thinking level. A mismatch invalidates the matrix.
|
||||
Pass the resolved `(provider, model_id)` explicitly in every Pilot Evaluator request, starting with the first cell. Freeze that pair and the Test Agent's configured `thinking_level` for the complete Benchmark workflow. Every scored Evaluator result must report the requested pair and the same configured thinking level. A mismatch invalidates the matrix.
|
||||
|
||||
Before planning Cases, state the Capability Contract:
|
||||
|
||||
@@ -121,9 +121,9 @@ model_id: <model_id>
|
||||
|
||||
Inspect the complete streamed and final worker response. Before reading `status`, `score`, or any other protocol field, verify that the worker-authored text is exactly one plain protocol YAML document. Narration, headings, code fences, summaries, or scoring details are not valid protocol. Ask the same Evaluator to resend only the clean YAML from its existing result; do not rerun the Test Agent for a formatting repair and do not extract YAML from the invalid response yourself. Transport metadata added by `run_subagent` is not worker-authored text. A wrong or missing Test Agent artifact is a valid scored result and must not be retried.
|
||||
|
||||
For every scored result, require non-empty `provider`, `model_id`, and `thinking_level`. Require the model pair to equal the explicitly resolved pair and the thinking level to equal the Test Agent configuration read before dispatch. Reject a Pilot or Formal matrix whose cells report mixed or mismatched runtimes. The Evaluator verifies provider/model from the root Trace and reports thinking from the unchanged Target Agent configuration; it does not require Trace metadata for thinking.
|
||||
For every scored result, require non-empty `provider`, `model_id`, and `thinking_level`. Require the model pair to equal the explicitly resolved pair and the thinking level to equal the Test Agent configuration read before dispatch. Reject a Pilot result whose cells report mixed or mismatched runtimes. The Evaluator verifies provider/model from the root Trace and reports thinking from the unchanged Target Agent configuration; it does not require Trace metadata for thinking.
|
||||
|
||||
Correct and resend an `invalid_request`. For `benchmark_invalid`, repair and rerun the affected Case during Pilot; during Formal, abandon the matrix and return to Pilot. For `version_changed`, discard the matrix and restart after the Agent version is stable.
|
||||
Correct and resend an `invalid_request`. For `benchmark_invalid`, repair and rerun the affected Case during Pilot. For `version_changed`, discard the current Pilot result and restart after the Agent version is stable.
|
||||
|
||||
For `evaluation_failed`, keep the same Benchmark revision and cell. Diagnose the failure and retry only when evidence proves the Test Agent did not start and the retry applies a new, specific repair. Do not set a numeric retry limit or repeat an unchanged launch. Stop when no new safe repair remains, external configuration is required, or it is unclear whether the Test Agent started. Never treat an evaluation failure as score zero.
|
||||
|
||||
@@ -131,7 +131,7 @@ For `evaluation_failed`, keep the same Benchmark revision and cell. Diagnose the
|
||||
|
||||
Treat the first draft as a hypothesis. The first valid result from every planned Case together forms Pilot iteration 1. A later iteration starts after a difficulty refinement and completes when every affected Case has a valid new result. Request corrections, validity repairs, and evaluation reruns stay in the current iteration and do not consume the requested iteration budget. Use the recorded Agent State version and fixed evaluation runtime.
|
||||
|
||||
Keep Pilot results out of the Scoreboard. During calibration, retain only one temporary restorable copy: the lowest-scoring complete valid revision seen so far. Store it outside `benchmarks/`, replace it only when a lower valid revision completes, and never retain invalid revisions.
|
||||
Keep unselected Pilot results out of the Scoreboard. During calibration, retain only one temporary restorable copy: the lowest-scoring complete valid revision seen so far, including its one-Run-per-Case result. Store it outside `benchmarks/`, replace it only when a lower valid revision completes, and never retain invalid revisions.
|
||||
|
||||
Use the Pilot to find the current Test Agent's capability boundary.
|
||||
|
||||
@@ -156,13 +156,13 @@ More rows, fields, distractors, files, near-duplicate examples, or explicit rule
|
||||
|
||||
Freeze immediately when a complete valid Pilot iteration meets the desired baseline score and no known design defect remains. Do not run another difficulty refinement merely to create more score margin. Otherwise continue through the requested valid-iteration limit. If the desired score is still unmet, restore the temporary lowest-scoring valid revision and proceed to Freeze. Report `calibration_failed` only when no valid Pilot revision can be produced or evaluation failures prevent a valid selection; missing the desired score alone is not a failure.
|
||||
|
||||
## Freeze and run the Formal Baseline
|
||||
## Freeze and record the Formal Baseline
|
||||
|
||||
After selecting the Pilot revision, restore that exact revision if needed. Run a complete consistency review and final leak check across every Case, fix any defect, then freeze the Benchmark and record the current Agent State version. Run a fresh, complete Case × Run matrix and never reuse a Pilot result. Once the first Formal cell is dispatched, do not change the Benchmark.
|
||||
After selecting the Pilot revision, restore that exact revision and its complete result if needed. Run a complete consistency review and final leak check across every Case. If the review finds a defect, repair it and produce a complete valid one-Run-per-Case Pilot result for the repaired revision before selecting and freezing it. Freeze the Benchmark and record the current Agent State version. Do not launch a fresh Formal matrix, rerun the selected Pilot, or backfill it to another Run count.
|
||||
|
||||
Accept the matrix when every cell is valid, every cell reports the frozen evaluation runtime, the Agent State version remains unchanged, the private scoring standard remained fixed, and every score loss reflects the Capability Contract. Record the Formal Baseline even when its score does not meet the desired baseline score.
|
||||
Accept the selected Pilot result as the Formal Baseline when every Case has exactly one valid Run, every cell reports the frozen evaluation runtime, the Agent State version remains unchanged, the private scoring standard remained fixed, and every score loss reflects the Capability Contract. Record the Formal Baseline even when its score does not meet the desired baseline score.
|
||||
|
||||
If Formal reveals a design defect, abandon the matrix and repair the frozen candidate revision before freezing and rerunning the complete matrix. Report `calibration_failed` only when no valid revision remains or evaluation failures prevent a complete Formal matrix. Never record a partial, abandoned, or invalid Formal matrix.
|
||||
Report `calibration_failed` only when no valid revision remains or evaluation failures prevent a complete selected Pilot result. Never record a partial, abandoned, invalid, or non-selected Pilot result as the Formal Baseline.
|
||||
|
||||
## Record and finish
|
||||
|
||||
@@ -198,7 +198,7 @@ After writing, parse the complete `scoreboard.yaml` and verify the appended Eval
|
||||
|
||||
Every Run and Case score is on the fixed `0..100` scale. Do not write `max_score`. Calculate and write every Case and Evaluation average directly in the Scoreboard: ignore `null` values when averaging cost and write `null` only when all contributing costs are unknown; round `score` averages to two decimal places, `cost` averages to six decimal places, and `duration_ms` averages to the nearest integer. These stored values are authoritative—do not add a server, frontend, script, or consistency check that recomputes or validates them. Do not add an `aggregate` object or use `case_id`, `mean_score`, `mean_cost`, or `mean_duration_ms`.
|
||||
|
||||
Report the Benchmark path, configuration, Agent State version, Evaluation average and Case Run scores, Test Session ids, and known limitations. Include one compact row per Pilot iteration with its score, diagnosed capability gap, difficulty adjustment, and freeze or stop decision.
|
||||
Report the Benchmark path, configuration, Agent State version, Evaluation average and Case Run scores, Test Session ids, and known limitations. Include one compact row per Pilot iteration with its score, diagnosed capability gap, difficulty adjustment, and freeze or stop decision. Identify which one-Run-per-Case Pilot result was recorded as the Formal Baseline.
|
||||
|
||||
After the accepted Formal Baseline is recorded, delete the temporary lowest-revision copy and other Builder calibration scaffolding. Keep the frozen Benchmark, Scoreboard, evaluation Workspaces, and score-linked Traces.
|
||||
|
||||
|
||||
@@ -1,39 +1,86 @@
|
||||
---
|
||||
name: data-analysis
|
||||
description: Complete data-analysis tasks with bounded evidence inspection, explicit answer-changing decisions, native artifact handling, and final output verification.
|
||||
short_description: Complete data analysis with a bounded, verified method.
|
||||
short_description_zh: 以有界、可验证的方法完成数据分析任务。
|
||||
version: 1
|
||||
updated: 2026-07-18T00:00:00Z
|
||||
description: Complete data-analysis tasks with bounded inspection, correct data semantics, native artifact handling, complete delivery, and risk-based verification.
|
||||
short_description: Deliver correct data-analysis artifacts with bounded verification.
|
||||
short_description_zh: 以正确数据语义完成产物并进行有界核查。
|
||||
version: 2
|
||||
updated: 2026-08-03T00:00:00Z
|
||||
---
|
||||
|
||||
# Data Analysis
|
||||
|
||||
Deliver the requested result and artifacts. Do not turn the task into a proof
|
||||
exercise or add evidence, reports, explanations, or intermediate files that
|
||||
were not requested.
|
||||
|
||||
## Before you start
|
||||
|
||||
Require a concrete data-analysis task, its available inputs, and the requested deliverable, location, and format. If the task or required materials are missing or ambiguous in an answer-changing way, ask before calculating or creating outputs. Otherwise proceed without forcing a fixed analysis template.
|
||||
Require a concrete data-analysis task, its available inputs, and the requested
|
||||
deliverable, location, and format. Ask only when missing information prevents a
|
||||
defensible result and would materially change the deliverable; otherwise proceed.
|
||||
|
||||
## Success criteria
|
||||
## Contract
|
||||
|
||||
- Derive the final result from one committed method: the selected evidence, answer-changing decisions, transformations, calculations, and judgments.
|
||||
- Produce every output the task asks for, in the requested location and format, and ensure it reflects the committed method.
|
||||
- Ground each answer-changing choice in the task materials rather than in a merely plausible nearby match.
|
||||
- Prefer a complete, simple, defensible deliverable over an elaborate or exhaustive analysis that risks not being delivered.
|
||||
Read the task, supplied inputs, and relevant data documentation. Identify every
|
||||
required output path and format, plus only the definitions that can change the
|
||||
result: scope, observation grain, keys, units, operators, ordering, coverage,
|
||||
and explicit formatting rules. Treat examples as illustrative unless the task
|
||||
makes them normative.
|
||||
|
||||
## Constraints
|
||||
If information is incomplete or ambiguous, first resolve it from the supplied
|
||||
materials. Ask only when the missing choice prevents a defensible result and
|
||||
would materially change the deliverable. Otherwise choose the best-supported
|
||||
interpretation and proceed.
|
||||
|
||||
- Use bounded probes first for large, unfamiliar, or expensive-to-read inputs. Narrow the scope, cap output, or use a timeout before expanding.
|
||||
- Before calculating or producing final outputs, identify the few decisions that can change the answer, such as inclusion or exclusion, matching, boundary choices, transformations, formulas, and ranking criteria. Adapt this check to the task; do not force a fixed analysis template.
|
||||
- Compare materially plausible alternatives only when they would change the final output. Use the smallest comparison needed to resolve the choice from the task materials, then commit to a method.
|
||||
- Once the evidence supports a defensible answer, create the requested output files promptly. Do not continue open-ended thinking, searching, or polishing over alternatives that would not change the delivered answer.
|
||||
- Use intermediate files only when they help compute or verify the result. Preserve the requested output format and structure unless the task asks for a change.
|
||||
- Satisfy the requested deliverable with the simplest sufficient artifact and implementation. Avoid optional structure, styling, helper code, or reimplementation that is not required by the task.
|
||||
- Keep final delivery steps short and robust. Avoid putting a long report, large dataset, or large script into one fragile streamed command when the deliverable can be created by shorter steps. Create the required artifact first, then refine only if the required outputs already exist.
|
||||
- For spreadsheet or Office deliverables, prefer the tool path that preserves the native artifact contract. If a task depends on Excel formula recalculation, data tables, workbook formatting, or Office export, first check whether native Excel automation such as `xlwings` is available before falling back to LibreOffice or a hand-rolled model. Record formulas or inputs that will be temporarily overwritten, restore them before finalizing, and verify that the requested workbook, deck, document, or PDF can be opened and contains the expected tables, sheets, slides, or sections. Do not dump or reimplement an entire workbook when bounded input changes and output reads can answer the task.
|
||||
- When creating spreadsheet tables, keep each table as one contiguous rectangle: title or caption, row-axis labels, column-axis labels, and data cells should be adjacent and inspectable together. Do not place row labels or key headers outside the visible table bounds or to the left of the title anchor. If you add decorative axis labels, keep the actual machine-readable row and column values inside the same table rectangle.
|
||||
## Bounded inspection
|
||||
|
||||
## Stop rules
|
||||
For large or unfamiliar inputs, begin with a bounded inventory, schema check,
|
||||
targeted sample, or narrow query. Expand inspection only when it can change a
|
||||
selection, transformation, calculation, or output. Do not exhaustively read or
|
||||
render data merely to increase confidence.
|
||||
|
||||
- Before finalizing, inspect the requested outputs and check their location, format, shape, and content against the committed method and the task.
|
||||
- If a later finding changes the method, assumptions, selections, transformations, calculations, decisions, or coverage, regenerate the affected outputs before finalizing.
|
||||
- Once the requested outputs exist and the relevant checks pass, stop instead of continuing open-ended exploration.
|
||||
## Data semantics
|
||||
|
||||
Compute at the correct row or entity grain. Evaluate conjunctive conditions on
|
||||
the same record or entity; do not replace row-level matching with unions of
|
||||
separate field values. Preserve nulls, exclusions, and explicit prohibitions.
|
||||
Enumerated outputs must cover the complete requested universe.
|
||||
|
||||
Ground answer-changing choices in the task and supplied data. Preserve
|
||||
documented source semantics, units, mappings, and native workflow behavior when
|
||||
they define the requested result. Do not reproduce an apparent source or tool
|
||||
defect merely for consistency. When plausible methods disagree, compare only
|
||||
the smallest answer-changing difference, choose the best-supported method, and
|
||||
use it consistently.
|
||||
|
||||
## Native artifacts
|
||||
|
||||
Preserve the requested artifact type and structure. When correctness depends on
|
||||
spreadsheet formulas, recalculation, formatting, database semantics, document
|
||||
layout, or export behavior, prefer a tool path that preserves and can verify
|
||||
those native properties. Restore temporarily changed inputs or formulas before
|
||||
finalizing. Use intermediate files only when they help produce or verify the
|
||||
requested deliverable.
|
||||
|
||||
## Delivery
|
||||
|
||||
As soon as a complete best-supported result exists, write every requested
|
||||
artifact at its exact path. For a multi-artifact task, establish a valid version
|
||||
of every artifact before refining any one of them. Do not leave a required
|
||||
artifact missing while pursuing additional certainty, polish, or diagnostics.
|
||||
If later evidence changes the result, update the artifact.
|
||||
|
||||
## Verification
|
||||
|
||||
Choose checks in proportion to answer-changing risk. Use the smallest
|
||||
independent check that can falsify each load-bearing assumption or computation.
|
||||
If a check disagrees, isolate and resolve the concrete difference. Do not repeat
|
||||
equivalent searches, calculations, renders, or inspections once remaining
|
||||
uncertainty cannot change the deliverable.
|
||||
|
||||
## Final check
|
||||
|
||||
Reopen the actual deliverables and verify their path, format, schema or
|
||||
structure, values, coverage, and openability as applicable. Confirm that every
|
||||
requested artifact exists and reflects the chosen method. Report the output
|
||||
paths concisely and stop.
|
||||
|
||||
@@ -182,17 +182,17 @@ describe("librarySkill", () => {
|
||||
expect(librarySkill("no-such-skill")).toBeUndefined();
|
||||
});
|
||||
|
||||
it("benchmark-design selects a valid Pilot revision before the Formal Baseline", () => {
|
||||
it("benchmark-design records the selected one-Run Pilot as the Formal Baseline", () => {
|
||||
const content = librarySkill("benchmark-design")?.content;
|
||||
expect(content).toBeDefined();
|
||||
|
||||
const pilotIndex = content!.indexOf("## Refine the Benchmark");
|
||||
const baselineIndex = content!.indexOf("## Freeze and run the Formal Baseline");
|
||||
const baselineIndex = content!.indexOf("## Freeze and record the Formal Baseline");
|
||||
const normalizedContent = content!.replace(/\s+/g, " ");
|
||||
|
||||
expect(pilotIndex).toBeGreaterThan(-1);
|
||||
expect(baselineIndex).toBeGreaterThan(pilotIndex);
|
||||
expect(normalizedContent).toContain("Keep Pilot results out of the Scoreboard");
|
||||
expect(normalizedContent).toContain("Keep unselected Pilot results out of the Scoreboard");
|
||||
expect(normalizedContent).toContain("requested valid-iteration limit");
|
||||
expect(normalizedContent).toContain("lowest-scoring valid Pilot revision");
|
||||
expect(normalizedContent).toContain("retain only one temporary restorable copy");
|
||||
@@ -245,7 +245,11 @@ describe("librarySkill", () => {
|
||||
"Do not run another difficulty refinement merely to create more score margin",
|
||||
);
|
||||
expect(normalizedContent).toContain("missing the desired score alone is not a failure");
|
||||
expect(normalizedContent).toContain("never reuse a Pilot result");
|
||||
expect(normalizedContent).toContain("`runs = 1`");
|
||||
expect(normalizedContent).toContain(
|
||||
"Do not launch a fresh Formal matrix, rerun the selected Pilot, or backfill it",
|
||||
);
|
||||
expect(normalizedContent).toContain("Accept the selected Pilot result as the Formal Baseline");
|
||||
expect(normalizedContent).toContain(
|
||||
"Record the Formal Baseline even when its score does not meet the desired baseline score",
|
||||
);
|
||||
@@ -293,7 +297,7 @@ describe("librarySkill", () => {
|
||||
"Read `thinking_level` from the Test Agent's `model.thinking_level`",
|
||||
);
|
||||
expect(benchmarkDesign).toContain(
|
||||
"Pass the resolved `(provider, model_id)` explicitly in every Pilot and Formal Evaluator request",
|
||||
"Pass the resolved `(provider, model_id)` explicitly in every Pilot Evaluator request",
|
||||
);
|
||||
expect(benchmarkDesign).toContain("Every Case Rubric has a fixed maximum of 100 points");
|
||||
expect(benchmarkDesign).toContain("average of the Case scores");
|
||||
@@ -311,6 +315,17 @@ describe("librarySkill", () => {
|
||||
);
|
||||
expect(benchmarkDesign).toContain("Ask the same Evaluator to resend only the clean YAML");
|
||||
expect(optimization).toContain("Delegate every evaluation to an `agent-evaluation` subagent");
|
||||
expect(optimization).toContain("a positive `runs` value");
|
||||
expect(optimization).toContain(
|
||||
"do not infer it from `benchmark_config.toml` or the Formal Baseline",
|
||||
);
|
||||
expect(optimization).toContain(
|
||||
"The initial Formal Baseline has one Run per Case; do not rerun or backfill it",
|
||||
);
|
||||
expect(optimization).toContain(
|
||||
"Compare each Candidate's stored top-level average directly with the current Reference score even when their Run counts differ",
|
||||
);
|
||||
expect(optimization).toContain("frozen Case set × requested `runs` matrix");
|
||||
expect(optimization).toContain("Before reading `status`, `score`, or any other protocol field");
|
||||
expect(optimization).toContain("Ask the same Evaluator to resend only the clean YAML");
|
||||
expect(optimizationRaw).toMatch(/summary_title:\s*>-\n\s+<public title>/);
|
||||
|
||||
@@ -657,7 +657,6 @@ Agent:
|
||||
Benchmark:
|
||||
- id: \`contextual-choice-adaptation\`
|
||||
- capability: form and transfer a stable finite-choice decision process from public rules, historical examples, and current facts
|
||||
- runs: \`1\`
|
||||
- desired_baseline_score: \`<75\`
|
||||
- pilot_iteration_limit: \`5\`
|
||||
|
||||
@@ -674,6 +673,7 @@ Scenarios:
|
||||
- test_agent_id: \`finite_choice_agent\`
|
||||
- benchmark_id: \`contextual-choice-adaptation\`
|
||||
- capability_direction: improve stability under incomplete information, conflicting rules, and finite choices
|
||||
- runs: \`3\`
|
||||
- desired_score: \`>=95\`
|
||||
- candidate_round_limit: \`5\``,
|
||||
},
|
||||
|
||||
@@ -640,7 +640,6 @@ Agent:
|
||||
Benchmark:
|
||||
- id:\`contextual-choice-adaptation\`
|
||||
- capability:从公开规则、历史案例和当前事实中形成并迁移稳定的有限选择决策过程
|
||||
- runs:\`1\`
|
||||
- desired_baseline_score:\`<75\`
|
||||
- pilot_iteration_limit:\`5\`
|
||||
|
||||
@@ -657,6 +656,7 @@ Benchmark:
|
||||
- test_agent_id:\`finite_choice_agent\`
|
||||
- benchmark_id:\`contextual-choice-adaptation\`
|
||||
- capability_direction:提高信息不完整、规则冲突和有限选项决策中的稳定性
|
||||
- runs:\`3\`
|
||||
- desired_score:\`>=95\`
|
||||
- candidate_round_limit:\`5\``,
|
||||
},
|
||||
|
||||
@@ -33,12 +33,13 @@ describe("draft example tasks", () => {
|
||||
"售后政策与工单事实",
|
||||
"投资策略、历史市场与当前指标",
|
||||
],
|
||||
buildForbiddenMarkers: ["thinking_level", "provider", "model_id"],
|
||||
buildForbiddenMarkers: ["thinking_level", "provider", "model_id", "runs:"],
|
||||
optimizationMarkers: [
|
||||
"使用 `agent-optimization`",
|
||||
"test_agent_id:`finite_choice_agent`",
|
||||
"benchmark_id:`contextual-choice-adaptation`",
|
||||
"提高信息不完整、规则冲突和有限选项决策中的稳定性",
|
||||
"runs:`3`",
|
||||
"desired_score:`>=95`",
|
||||
"candidate_round_limit:`5`",
|
||||
],
|
||||
@@ -58,12 +59,13 @@ describe("draft example tasks", () => {
|
||||
"policy and ticket facts",
|
||||
"strategy, historical markets, and current indicators",
|
||||
],
|
||||
buildForbiddenMarkers: ["thinking_level", "provider", "model_id"],
|
||||
buildForbiddenMarkers: ["thinking_level", "provider", "model_id", "runs:"],
|
||||
optimizationMarkers: [
|
||||
"Use `agent-optimization`",
|
||||
"test_agent_id: `finite_choice_agent`",
|
||||
"benchmark_id: `contextual-choice-adaptation`",
|
||||
"improve stability under incomplete information, conflicting rules, and finite choices",
|
||||
"runs: `3`",
|
||||
"desired_score: `>=95`",
|
||||
"candidate_round_limit: `5`",
|
||||
],
|
||||
|
||||
Reference in New Issue
Block a user