run_info.json records the request configuration, metric-artifact provenance, and final state. params.json keeps the compact parameter set needed to save and re-aggregate results. progress.json provides the latest snapshot, progress.jsonl preserves the full event sequence, and the log records readable execution messages and errors.
When Files Are Created
Not every invocation leaves these files behind. The CLI and SDK check the request before creating its run directory; a failure at that point creates no run directory.
agentcompass launch --dry-run also creates no output.
A preparation error after the directory exists will usually leave the log and run_info.json. If error handling completes normally, AgentCompass also writes the final state and a run_finished event. If the process is forcibly terminated, the final state, the last progress events, or params.json may not have been written.
run_info.json
run_info.json records which request configuration the evaluation used, which plan produced the current metric artifacts, and how the request ended. It is created before tasks are loaded and updated throughout the run.
Top-Level Fields
Current run-info v3 uses structured issues. Previous v2 results are adapted only at read boundaries; classification-unknown failures cannot be reused for current scoring. Historical directories are not modified.
request Structure
request is divided into model, Benchmark, Harness, Environment, execution, runtime, output, and metadata sections. Each component’s params is an open object whose fields depend on the selected component.
Each network-policy object contains
network_mode, which selects the network-access mode, and allowed_hosts, which lists permitted hosts. When AgentCompass writes JSON, it removes null values and empty objects or lists. An empty allowed_hosts may therefore be absent from the file.
request is not a copy of the original command line. It also excludes process-level settings such as results_dir, the whole-request timeout, log levels, and Environment provider concurrency limits. To verify these settings, consult the invocation, configuration, and log together. See the Model, Benchmark, Harness, and Environment documentation for component-specific fields.
When present, reused_from has these fields:
Reuse validation
Reuse requires a supported run schema, the same Benchmark ID, and the sameexecution.attempts plan. Tasks and checkpoints are matched by task ID and attempt number. Category labels, complete task inputs, Model parameters, Harness settings, Environment settings, and Recipe configuration do not need to match the source run.
New records do not write execution_fingerprints or task_fingerprints. These fields in older records are ignored; their absence does not prevent reuse or summary generation. The sanitized request and resolved plans remain available for provenance.
Selecting reuse means using the saved agent outputs and prepared task context. Pending fresh evaluations use the current request and a newly resolved execution plan, including evaluation timeouts, resources, and environment variables. Complete results still take precedence and are not evaluated again when parameters change.
Checkpoint recovery validates task/attempt identity, a completed agent run, fresh evaluation mode, artifact coverage, and file integrity. Saved artifacts may include more declarations than the current verifier needs; declaration order and collection settings do not need to match. A newly required artifact declaration that was never captured cannot be restored. Recorded missing or excluded outputs retain their original meaning.
metric_artifacts Structure
Whenever the Benchmark metric files are written, AgentCompass replaces this provenance record:
source is evaluation for normal run finalization and summary after a non-dry-run agentcompass summary. A summary regeneration also records a redacted benchmark_params_override object, including an empty object when no override was passed. report binds the files to the exact attempt plan and run-level aggregation that produced them; it does not replace the original request.
resolved_execution_plans Structure
resolved_execution_plans records the Environment, network policies, and Recipes resolved for each task attempt. Its structure is:
The plan summary is written after resolution but before the Environment is opened. It tells you what the attempt planned to use; it does not prove that the Environment was created successfully. It also excludes the complete Recipe-resolved image, snapshot, working directory, resources, and Environment provider parameters.
A task or attempt reused without execution receives no new resolved-plan entry. The task detail keeps the metric attempt plan in
attempt_plan; Environment and Recipe plans remain in the source run’s run_info.json.
Evaluation Checkpoints
task.json stores shared task metadata, the attempt plan, and logical attempt numbers mapped to directory names such as "1": "attempt-1". Each attempt’s result.json stores its own result and retry count. Readers derive task summaries and total retries; there is no task-level result.json. Evaluation checkpoints record agent completion and references to saved artifacts.
checkpoint.json uses schema agentcompass.attempt_checkpoint.v1. Its evaluation section stores task/attempt identity, agent completion, source run and evaluation mode, and the artifact manifest and references the same artifacts/ tree used by the result. The runtime no longer writes an additional ZIP. Its scheduler section records terminal attempt state until the task result is durable; an in-progress download can also record artifact_transfer diagnostics. Metadata updates are atomic.
Recovery validates task/attempt identity, declarations, file sizes and checksums before uploading artifacts into the fresh verifier. Cross-run reuse preserves the logical attempt number, copies validated artifacts into the target’s attempt-<n> directory and updates references before publishing results. When a task moves between running/ and its state directory, artifact references are updated with it. The target run can then be used independently of its source.
Default artifact limits remain 16 GiB, 100,000 entries and 600 seconds per operation. Files retain their original bytes. When an attempt reruns, previous local outputs are moved beneath that attempt’s retries/ directory, so stale files cannot mix with the new submission. No per-attempt logs directory is created automatically.
Evaluation checkpoint v4 saves a classified RunResult, PreparedTask, network policy and artifact manifest. Each none/fresh scoring retry receives an isolated copy. Legacy v3 records remain readable; missing classified snapshots prevent cross-run materialization and cause normal execution with an explicit reason.
The Benchmark rebuilds evaluation context from the current task and plan through prepare_evaluation(). SWE-bench evaluators load their patch from the saved artifact. The runtime does not require or validate a trajectory to resume evaluation. Benchmarks that need additional in-memory run state must implement artifact-based recovery before they can resume this way; otherwise the agent reruns. Missing required inputs also cause a rerun. Successful complete results remain reusable without reevaluation.
params.json
params.json stores only the parameters needed to write task details and regenerate summaries. AgentCompass rewrites it when saving task details or generating the final summary. The file may not exist if the request fails before either operation. Running agentcompass summary separately leaves an existing params.json unchanged, while updating the metric files and their provenance in run_info.json.
Unset fields directly under
model, benchmark, execution, and output are omitted; values such as empty strings can still remain inside nested params objects. params.json does not contain the Harness, Environment, reuse settings, metadata, or complete Recipe-resolved configuration. You therefore cannot use it to reconstruct the complete evaluation configuration.
When regenerating a summary, AgentCompass reads run_info.json.request first and uses params.json to fill in missing values. The two files serve these purposes:
progress.json
progress.json stores the latest run state and task counts. AgentCompass replaces the snapshot with the latest state whenever a progress event occurs, so a status page or script can poll it.
Each
active_tasks.<task-id> object contains category, phase, attempt, and updated_at. After a task starts but before it enters a specific phase, phase is running. A missing category or attempt number is stored as null.
A task’s highest issue severity is resolved as fatal > error > warning, and every finished task counts toward exactly one severity. fatal maps to failed_tasks; error and warning map to error_tasks and warning_tasks while also being included in completed_tasks.
completed_tasks means that execution ended normally; it does not mean that the Benchmark marked the answer correct. Use task details and canonical metrics.json for Benchmark observations and aggregate values.progress.jsonl
progress.jsonl stores the complete progress event stream. Each line is one JSON object, appended in emission order. To reconstruct a task’s phases, attempts, and retries, read this file instead of relying only on the latest snapshot.
The CLI’s --progress auto|plain|none and the SDK’s progress="auto"|"plain"|"none" control only the live terminal display. They do not disable progress.json or progress.jsonl. When an SDK caller provides a custom progress reporter, file creation depends on that reporter’s output configuration.
In the fields below, an orchestration means one launch invocation that schedules multiple evaluation requests. For a standalone request, the orchestration-related fields are null.
Fields on Every Event
All fields above are always serialized. Missing values are written as
null, and payload is always an object.
Events and Event-Specific Fields
Matching
task_started and task_finished events use the same payload.index and payload.total. These values are the sequence number and task count used during scheduling; they are not task identifiers. Always use task_id to identify a task. A multi-evaluation orchestration normally preserves positions in the original selected list, so reuse can leave gaps. A single evaluation request may instead renumber the remaining tasks.
The fields in attempt_retry.payload mean:
Current
phase_changed.phase values are:
Events from concurrent tasks can interleave. Filter by
task_id and attempt to follow one task. Do not assume that every task enters the same phases, and do not infer dependencies from adjacent events belonging to different tasks.
When agentcompass analysis runs, AgentCompass removes both existing progress files from the target result directory before recording the new analysis events. Without --override, the target is a newly created result copy, so the source directory is not modified.
Re-analysis keeps the original request’s run_id, but it does not rebuild run_info.json, params.json, or the run-directory log. Those files continue to describe the original evaluation request.
run.log
Each run or launch request writes framework logs to run.log. Reopening the same run directory appends to that file without overwriting earlier messages.
Logging begins after the run directory is created, before run_info.json and later runtime checks. Earlier output from the CLI or SDK is not copied into this file.
Each line uses this structure:
--file-log-levelcontrols the minimum level in the run-directory log and defaults toDEBUG;--log-levelcontrols console output independently.- Third-party loggers are filtered to
WARNINGand above by default, even when the file level isDEBUG. - The file contains messages recorded by AgentCompass and integrated components. It does not guarantee every shell command, provider response, or internal third-party event.
- Logs are not structured results and are not inputs to summary generation, reuse, or analysis regeneration.
run_info.json and params.json redact recognized credential fields and remove underscore-prefixed runtime fields from parameter objects. This is not a general sensitive-data scanner, and it does not apply to logs.
Custom fields, free-form text, progress events, and logs may still contain paths, URLs, task data, provider information, or tracebacks. Review and remove sensitive content before sharing a run directory.
Diagnose a Failed Run
Check the files in this order to narrow down a failure:- Inspect
progress.jsonfor the request state and task counts. While the request is active, it also shows current phases. - Filter
progress.jsonlbytask_idto reconstruct the failing task’s last phase, attempts, and retry path. A terminal snapshot clears active tasks, so use the event stream for the last phase after the request ends. - Inspect
run_info.jsonfor the merged request, reuse source, and that attempt’s recipe and network-policy summary. - If the failure involves result persistence or summary regeneration, inspect
params.json. - Search
run.logby task ID, phase, or exception type for detailed messages and tracebacks.
run_info.json and the progress files in different states. Use persisted task details and summaries to determine the final evaluation results.
See Task Results for task-level fields, Summary and Analysis Results for aggregate metrics, and Run Controls for log-level and progress-display options.
