The last two fields belong under
execution. The examples below show how to pass them through each interface.
Set Execution and Evaluation Budgets
This example gives the agent 3,600 seconds and scoring 600 seconds, with a 7,200-second limit for the complete evaluation. It uses the same task, Model environment variables, and Docker setup as the Quick Start.- CLI
- YAML
- Python SDK
launch, set the overall limit with --timeout-seconds. Put shared task budgets under defaults.execution, or a request’s overrides under requests[].execution.
The overall limit covers task loading, preparation, execution, scoring, analysis, and summarization after component preflight. Its default is 360000 seconds (100 hours); explicitly set 0 to disable it. On expiry, unfinished work is cancelled and resource cleanup begins. Cleanup can continue beyond the deadline.
Task budgets do not add to the overall limit. Allow enough total time for scheduling, all tasks and retries, result collection, and evaluation. The overall limit can expire during any of these phases.
Understand the Defaults
All time values are in seconds. Execution settings accept finite positive values;0 does not disable a task timeout.
Omitting a field or setting it to
null inherits the Benchmark plan’s budget, then the task’s budget. Execution can also fall back to the Harness default. If none of these sources supplies a budget, that phase has no time limit. Setting null does not clear an inherited budget.
Each attempt and retry starts a new phase budget. Execution and scoring use separate clocks; unused execution time does not carry over to scoring.
Use a multiplier when tasks already have suitable relative budgets and you want to extend them together. For example, this doubles agent execution time while leaving scoring unchanged:
The resolved budget is
base seconds × effective multiplier. Phase-specific multipliers replace the common multiplier; they do not multiply together. They also apply to explicit budgets. If there is no base budget, a multiplier cannot create one. Collection budgets and execution compensation are not scaled.
See Which Phase a Timeout Covers
The diagram shows one attempt with a Harness and a separate evaluation Environment. Read it from left to right; column widths do not represent duration. Some Benchmarks skip artifact transfer or use the execution Environment for scoring.run_timeout_seconds.
Request and Command Limits
run_timeout_seconds and evaluation_timeout_seconds bound an entire phase. Individual model requests, tool commands, or judge requests within that phase may have their own limits. Increasing the phase budget does not increase those limits. The component determines whether an operation timeout triggers a retry, allows execution to continue, or ends execution.
These examples link to the relevant component’s parameter documentation. Fields and defaults depend on the selected component; not every Harness or Benchmark accepts these fields.
For example, this
mini_swe_agent command gives the entire agent execution 3,600 seconds, each model request 120 seconds, and each tool command 300 seconds:
model_params, harness_params, and benchmark_params. With --config, put Harness and Benchmark fields directly under their component ID in harnesses or benchmarks. Supply Model parameters through the run command, orchestration request, or SDK. See configuration file structure for the complete format.
Query the selected Harness or Benchmark’s complete field list and defaults with:
Advanced Settings
Keep the defaults unless a failure points to one of these operations. Set these fields underexecution, through CLI --execution-params, YAML, or SDK execution_params.
Compensation, result collection, individual preparation commands, and transfer budgets require finite positive values. See Save and Prepare Artifacts for examples of preparation and transfer limits.
Allow a Harness to Finish Handling Its Timeout
Some Harnesses start their internal execution timer after preparing the agent. The runtime allows an additional 120 seconds by default around Harness execution so that preparation and stopping do not immediately cut off the internal timeout handling. For a resolved execution budget ofT seconds, the Harness receives T seconds and the runtime’s outer execution budget is T + 120 seconds.
Increase compensation if agent preparation is unusually slow or logs show the runtime timing out before the Harness finishes stopping. For example:
Allow More Time to Collect the Result
If execution ends but collecting the answer, logs, or trajectory exceeds 60 seconds, increase the result collection budget:Troubleshoot a Timeout
Inspect the task’s issues andrun.log to identify the failing operation before changing a setting. See Diagnose a Failed Run for the files to check.
