Skip to main content
Set a budget for the complete evaluation, each agent execution, or each scoring phase. Start with these three controls; the other settings on this page are for specific collection or timeout problems. The last two fields belong under execution. The examples below show how to pass them through each interface.

Set Execution and Evaluation Budgets

This example gives the agent 3,600 seconds and scoring 600 seconds, with a 7,200-second limit for the complete evaluation. It uses the same task, Model environment variables, and Docker setup as the Quick Start.
For launch, set the overall limit with --timeout-seconds. Put shared task budgets under defaults.execution, or a request’s overrides under requests[].execution. The overall limit covers task loading, preparation, execution, scoring, analysis, and summarization after component preflight. Its default is 360000 seconds (100 hours); explicitly set 0 to disable it. On expiry, unfinished work is cancelled and resource cleanup begins. Cleanup can continue beyond the deadline.
Task budgets do not add to the overall limit. Allow enough total time for scheduling, all tasks and retries, result collection, and evaluation. The overall limit can expire during any of these phases.

Understand the Defaults

All time values are in seconds. Execution settings accept finite positive values; 0 does not disable a task timeout. Omitting a field or setting it to null inherits the Benchmark plan’s budget, then the task’s budget. Execution can also fall back to the Harness default. If none of these sources supplies a budget, that phase has no time limit. Setting null does not clear an inherited budget. Each attempt and retry starts a new phase budget. Execution and scoring use separate clocks; unused execution time does not carry over to scoring. Use a multiplier when tasks already have suitable relative budgets and you want to extend them together. For example, this doubles agent execution time while leaving scoring unchanged:
The resolved budget is base seconds × effective multiplier. Phase-specific multipliers replace the common multiplier; they do not multiply together. They also apply to explicit budgets. If there is no base budget, a multiplier cannot create one. Collection budgets and execution compensation are not scaled.

See Which Phase a Timeout Covers

The diagram shows one attempt with a Harness and a separate evaluation Environment. Read it from left to right; column widths do not represent duration. Some Benchmarks skip artifact transfer or use the execution Environment for scoring. Timeout settings for agent execution, Harness result collection, artifact preparation, separate batch download and restore, and scoring within one task attempt. The execution phase includes the Harness’s agent preparation, model requests, and tool loop. Result collection starts afterward with its own budget. The evaluation budget starts after its Environment is ready and artifacts have been restored. Gray columns are outside the six phase settings shown. Environment creation has a separate startup limit; Harness closing and resource cleanup are not covered by run_timeout_seconds.

Request and Command Limits

run_timeout_seconds and evaluation_timeout_seconds bound an entire phase. Individual model requests, tool commands, or judge requests within that phase may have their own limits. Increasing the phase budget does not increase those limits. The component determines whether an operation timeout triggers a retry, allows execution to continue, or ends execution. These examples link to the relevant component’s parameter documentation. Fields and defaults depend on the selected component; not every Harness or Benchmark accepts these fields. For example, this mini_swe_agent command gives the entire agent execution 3,600 seconds, each model request 120 seconds, and each tool command 300 seconds:
The SDK uses the corresponding model_params, harness_params, and benchmark_params. With --config, put Harness and Benchmark fields directly under their component ID in harnesses or benchmarks. Supply Model parameters through the run command, orchestration request, or SDK. See configuration file structure for the complete format. Query the selected Harness or Benchmark’s complete field list and defaults with:

Advanced Settings

Keep the defaults unless a failure points to one of these operations. Set these fields under execution, through CLI --execution-params, YAML, or SDK execution_params. Compensation, result collection, individual preparation commands, and transfer budgets require finite positive values. See Save and Prepare Artifacts for examples of preparation and transfer limits.

Allow a Harness to Finish Handling Its Timeout

Some Harnesses start their internal execution timer after preparing the agent. The runtime allows an additional 120 seconds by default around Harness execution so that preparation and stopping do not immediately cut off the internal timeout handling. For a resolved execution budget of T seconds, the Harness receives T seconds and the runtime’s outer execution budget is T + 120 seconds. Increase compensation if agent preparation is unusually slow or logs show the runtime timing out before the Harness finishes stopping. For example:
With this configuration, the agent’s budget remains at 5,400 seconds and the runtime’s outer execution budget is 5,700 seconds. Preparation consumes part of that extra time. Result collection still has its own 60-second budget afterward; it is not included in compensation. Without a Harness, execution uses the resolved budget directly. Compensation does not create a deadline when execution is unlimited.

Allow More Time to Collect the Result

If execution ends but collecting the answer, logs, or trajectory exceeds 60 seconds, increase the result collection budget:
Collection is attempted after execution succeeds, fails, or times out. This budget does not extend agent execution, artifact transfers, or the overall evaluation limit. A partial result preserves available evidence; it does not remove the timeout failure or guarantee a score.

Troubleshoot a Timeout

Inspect the task’s issues and run.log to identify the failing operation before changing a setting. See Diagnose a Failed Run for the files to check.