> ## Documentation Index
> Fetch the complete documentation index at: https://opencompass-docs-preview-pr-335-0.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# 实现概览

实现 Benchmark 前，先确定由谁执行 agent、在哪里评测，以及如何聚合结果。

Benchmark 没有一套适用于所有场景的固定模板。大多数集成把执行交给 Harness；只有上游交互循环本身就是评测定义的一部分时，Benchmark 才接管执行。评分既可以在 AgentCompass 进程中完成，也可以在任务 Environment 或独立的评测 Environment 中完成。

## 第一步：确定由谁执行

这一步决定由 Harness 还是 Benchmark 执行 agent 循环：

| 场景 | <span style={{ display: "inline-block", minWidth: "10.5rem" }}>实现类型</span> | 谁实现执行阶段的 `run_task()` | 命令中的 Harness |
| - | - | - | - |
| 已有 Harness 可以执行准备后的任务 | `BaseBenchmark` | Harness | 已注册的 Harness ID |
| 上游 Benchmark 定义了不可复用的专属交互循环 | `HarnessFreeBenchmark` | Benchmark | `none` |

`HarnessFreeBenchmark` 是 `BaseBenchmark` 的子类，两者共享任务加载、准备、产物收集、评测和聚合契约，只是在执行阶段由不同组件负责。优先使用 [Harness 驱动](/zh/developer_guide/extensions/benchmark/code_implementation/harness_driven)；只有上游交互协议本身属于 Benchmark 定义时，才使用 [Benchmark 驱动](/zh/developer_guide/extensions/benchmark/code_implementation/benchmark_driven)。

## 第二步：选择评测位置

这一步决定 `evaluate()` 在哪里运行，不受第一步选择的影响：

| 评测器需要什么 | `evaluation_environment_mode` | `evaluate()` 收到的 `env` |
| - | - | - |
| 只根据答案或已收集的数据评分 | `none` | `None` |
| 检查任务执行后的同一工作区或进程状态 | `reuse` | 仍然存活的任务 Environment |
| 在隔离的 Environment 中运行验证器 | `fresh` | 新建的评测 Environment |

`BaseBenchmark` 和 `HarnessFreeBenchmark` 都可以使用这三种模式。通过 `TaskSpec.artifacts` 使用公共采集与 fresh 恢复：普通任务没有隐式路径，Harbor 适配层在未声明或空列表时仍补入 `/logs/artifacts/`。执行参数可以覆盖这些路径和命令列表，通过 `execution.save_artifacts` 控制本地保存。需要生成或整理文件时声明 `TaskSpec.artifact_collect` 命令；否则直接收集已有文件。具体实现见[评测模式与产物](/zh/developer_guide/extensions/benchmark/code_implementation/evaluation_modes)。

## 第三步：选择聚合方式

这一步决定如何把各次任务尝试的判定汇总为请求级指标，但不会改变由谁执行或在哪里评测：

| 结果形态 | `aggregate_metrics()` |
| - | - |
| 官方结果是 Contract 中任务级 reduction 的均值 | 继承 `BaseBenchmark` 的默认聚合 |
| 官方结果使用比值之和、总和、排名、奖牌或其他语料级公式 | 覆盖并返回 `MetricReport` |

具体结果字段和聚合实现见[结果与聚合](/zh/developer_guide/extensions/benchmark/code_implementation/results_and_aggregation)。无论最终组合是什么，都先阅读[共同契约](/zh/developer_guide/extensions/benchmark/code_implementation/shared_contracts)，确定任务字段、可见性边界和逐任务计划。

## 共享生命周期

每次请求先加载并选择任务，之后再为每个任务和每次尝试执行以下步骤：

```text theme={"system"}
load_tasks
  → select_tasks
  → build_plan
  → 打开任务 Environment
  → prepare_task
  → Harness.run_task 或 Benchmark.run_task
  → artifact_collect 命令（已声明且启用时）
  → download_artifacts (runtime)
  → evaluate
  → 持久化尝试结果
  → aggregate_metrics
```

产物准备和收集期间，任务 Environment 始终仍在运行。`evaluate()` 的执行位置由评测模式决定；所有任务完成后，`aggregate_metrics()` 读取持久化结果并完成汇总。

## 方法分工

| <span style={{ display: "inline-block", minWidth: "10.5rem" }}>方法</span> | 实现要求或默认行为 | 责任 |
| - | - | - |
| `load_tasks()` | 必须实现 | 将固定版本的数据集转换为 `TaskSpec` |
| `select_tasks()` | 默认实现 | 对已加载任务应用通用选择逻辑；只有特殊选择语义才覆盖 |
| `build_plan()` | 默认实现 | 为一次任务尝试解析类型化的 Benchmark 状态 |
| `prepare_task()` | 必须实现 | 在任务 Environment 中准备 Harness 或 Benchmark 执行输入 |
| `evaluate()` | 必须实现 | 保留执行状态，并把已声明观测写入 `RunResult.metrics` |
| `aggregate_metrics()` | 默认 Contract 驱动聚合 | 将任务级 reduction 转换为 `MetricReport` |
| `run_task()` | 仅 `HarnessFreeBenchmark` 必须实现 | 执行由 Benchmark 负责的推理或交互循环 |

实现完成后，按[文档更新](/zh/developer_guide/extensions/benchmark/documentation_update)补充用户入口，再按[验证与对齐](/zh/developer_guide/extensions/benchmark/validation_and_alignment)验证真实数据、Environment 和官方结果。

在 Benchmark 类上声明 `collect_artifacts = False`，表示不支持公共产物收集。runtime 会在创建环境前拒绝与此冲突的任务/Recipe 声明及用户路径/命令覆盖。默认值为 `True`；没有默认产物路径本身不代表禁用该能力。


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.