Skip to main content
agentcompass run creates and executes one evaluation request with the selected benchmark, harness, model, and environment. BENCHMARK HARNESS MODEL are positional arguments in a fixed order; select the environment with --env.

Run a Minimal Evaluation

The example uses sample_ids in --benchmark-params to select one benchmark task by its stable task ID so you can quickly verify the component and endpoint configuration:
One run command corresponds to one evaluation request. To coordinate multiple explicit requests, use agentcompass launch.

Parameter Reference

This section explains how to use agentcompass run parameters and lists their complete signatures, defaults, and owners. See Run Controls for guidance on concurrency, timeouts, retries, output, and debugging. “Built-in default” means the value used before user-level, project-level, or explicit configuration files override it. A conditional parameter is required only when the selected component or endpoint needs it.

Component Selection and Parameters

Configuration and Recipes

Execution Controls

Configure Repeated Attempts

For example, the following options plan three independent attempts per task and aggregate them with avg:
When these options are omitted, AgentCompass uses k=1 and avg: each task runs once and its metrics use native@1. With k>1, attempts are independent execution units and share the --task-concurrency limit; the two strategies differ as follows: Keep these boundaries in mind:
  • A scalar primary with pass raises an error before tasks start.
  • k counts logical attempts and excludes retries created within an attempt by --max-retries.
  • See Metrics and Aggregation for exact avg@k and pass@k definitions and missing-attempt behavior.
With launch, each request can set its own k and strategy; see Configure k and Strategy per Request.

Output and Reuse

Process Settings

Analysis

Component-Specific JSON Parameters

The four JSON parameter flags do not share one schema. Their available fields and defaults depend on the selected component: sample_ids belongs in --benchmark-params; repeated attempts use the execution controls above. Provider CPU, memory, image, and network settings belong in --env-params. See Configure an Evaluation for the ownership map.