none, reuse, or fresh from the state the verifier must access, and extract any artifacts needed later before the task Environment closes.
Three Modes
In all three modes, runtime executes declared collection commands while the environment is alive; without commands it collects existing files directly. Collection follows the resolved path and command lists, using Benchmark defaults unless execution parameters replace them.
execution.save_artifacts controls local storage; fresh requires it to be enabled. Artifact preparation and collection do not decide the score.
Set Defaults and Per-Task Overrides
Set a class-level default when every sample uses the same mode:TaskSpec values when samples require different verifier paths:
ExecutionPlan. Matching Recipes can then adjust the plan according to their own contracts. If a Recipe changes the evaluation Environment, its implementation and tests must state which fields it may override and which it must preserve. When debugging the effective mode, inspect resolved_execution_plans in run_info.json; it records the plan after Recipe adjustments.
none: Evaluate in the AgentCompass Process
Use none when evaluation depends only on TaskSpec, PreparedTask, and RunResult. Do not read the task workspace in this mode:
reuse: Inspect the Task Environment
reuse evaluates before the task Environment closes, so the verifier can see the workspace left by the agent:
evaluation_network_policy; do not assume the execution-stage network permissions remain active.
For a complete production reuse implementation, inspect terminalbench2.py.
fresh: Prepare, Collect, Then Verify in Isolation
fresh does not copy the entire workspace. Declare the files or directories needed by the verifier; shared runtime code collects them before closing the agent environment and restores them at the same absolute paths in the fresh environment:
TaskSpec.artifact_collect commands; Harbor’s verifier.collect maps to the same contract. For example:
Check Before Choosing a Mode
- Choose
nonewhen scoring depends only on an answer or in-memory objects; do not create an extra sandbox for simple scoring. - Choose
reusewhen the verifier must see the original filesystem modified by the agent, and ensure the verifier does not corrupt results that must be retained. - Choose
freshwhen the verifier must not trust dependencies or processes left by the agent, and transfer only the minimum required submission. - Set explicit timeouts for commands in
prepare_task(),artifact_collect, and the verifier, and make every step safe to repeat during retries. - Use different statuses for evaluation failure and a valid zero score; see Results and Aggregation for the mapping.
