Skip to main content
Choose none, reuse, or fresh from the state the verifier must access, and extract any artifacts needed later before the task Environment closes.

Three Modes

In all three modes, runtime executes declared collection commands while the environment is alive; without commands it collects existing files directly. Collection follows the resolved path and command lists, using Benchmark defaults unless execution parameters replace them. execution.save_artifacts controls local storage; fresh requires it to be enabled. Artifact preparation and collection do not decide the score.

Set Defaults and Per-Task Overrides

Set a class-level default when every sample uses the same mode:
Set the mode on individual TaskSpec values when samples require different verifier paths:
The runtime Planner initially resolves the mode in this order:
This is the selection order while building the initial ExecutionPlan. Matching Recipes can then adjust the plan according to their own contracts. If a Recipe changes the evaluation Environment, its implementation and tests must state which fields it may override and which it must preserve. When debugging the effective mode, inspect resolved_execution_plans in run_info.json; it records the plan after Recipe adjustments.

none: Evaluate in the AgentCompass Process

Use none when evaluation depends only on TaskSpec, PreparedTask, and RunResult. Do not read the task workspace in this mode:
If a controller-side evaluator calls another Model, configure that judge Model explicitly and record its version and inference parameters. Do not silently let the Model under test score its own result.

reuse: Inspect the Task Environment

reuse evaluates before the task Environment closes, so the verifier can see the workspace left by the agent:
The verifier path and its dependencies must exist in the task Environment. Evaluation uses the resolved evaluation_network_policy; do not assume the execution-stage network permissions remain active. For a complete production reuse implementation, inspect terminalbench2.py.

fresh: Prepare, Collect, Then Verify in Isolation

fresh does not copy the entire workspace. Declare the files or directories needed by the verifier; shared runtime code collects them before closing the agent environment and restores them at the same absolute paths in the fresh environment:
If the agent already writes these files, no preparation commands are needed. Otherwise declare TaskSpec.artifact_collect commands; Harbor’s verifier.collect maps to the same contract. For example:
The evaluator receives an environment with the declared files already restored. Do not download or upload them again:
Preparation and collection have independent execution budgets. Failed preparation or transport retains diagnostics but does not proceed to grading or publish a resumable checkpoint; a normally missing submission remains a verifier scoring decision. See Shared Contracts for size bounds, integrity checks, and preparation-hook responsibilities. DeepSWE is a production example of preparation plus shared transfer.

Check Before Choosing a Mode

  • Choose none when scoring depends only on an answer or in-memory objects; do not create an extra sandbox for simple scoring.
  • Choose reuse when the verifier must see the original filesystem modified by the agent, and ensure the verifier does not corrupt results that must be retained.
  • Choose fresh when the verifier must not trust dependencies or processes left by the agent, and transfer only the minimum required submission.
  • Set explicit timeouts for commands in prepare_task(), artifact_collect, and the verifier, and make every step safe to repeat during retries.
  • Use different statuses for evaluation failure and a valid zero score; see Results and Aggregation for the mapping.