Step 1: Choose the Execution Owner
This choice determines whether the Harness or the Benchmark runs the agent loop:HarnessFreeBenchmark is a subclass of BaseBenchmark. They share the task-loading, preparation, artifact-collection, evaluation, and aggregation contracts; only the owner of the execution stage differs. Prefer a Harness-driven Benchmark. Use a Benchmark-driven implementation only when the upstream interaction protocol itself belongs to the Benchmark definition.
Step 2: Choose the Evaluation Location
This choice determines whereevaluate() runs and is independent of the execution owner:
Both
BaseBenchmark and HarnessFreeBenchmark can use all three modes. Use TaskSpec.artifacts for shared collection and fresh restore: generic tasks have no implicit paths, while the Harbor adapter adds /logs/artifacts/ even for empty declarations. Execution parameters may replace those path and command lists; use execution.save_artifacts to control local storage. Declare TaskSpec.artifact_collect commands when files need preparation; otherwise runtime collects existing files directly. See Evaluation Modes and Artifacts.
Step 3: Choose the Aggregation Strategy
This choice determines how attempt-level verdicts become request-level metrics. It does not change the execution owner or evaluation location:
See Results and Aggregation for the result fields and aggregation patterns. Regardless of the final combination, read Shared Contracts first to define task fields, visibility boundaries, and per-task plans.
Shared Lifecycle
Each request loads and selects tasks first, then performs these steps for every task and attempt:evaluate() runs. After all tasks finish, aggregate_metrics() reads persisted results and produces the summary.
Method Responsibilities
After the implementation works, add the user-facing entry point through Documentation Update, then follow Validation and Alignment with real data, Environments, and official results.
Declare
collect_artifacts = False on the Benchmark class to mark shared artifact collection as unsupported. Runtime rejects conflicting task/Recipe declarations and user path/command overrides before environment allocation. The default is True; having no default artifact paths alone does not disable the capability.