> ## Documentation Index
> Fetch the complete documentation index at: https://opencompass-docs-preview-pr-335-0.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# Results and Reuse

AgentCompass checkpoints completed attempts, writes each completed task detail atomically, and then aggregates those details into request-level metrics. Reuse always materializes compatible data into a newly reserved run directory; it never modifies the source run.

The main implementation is `RunStore` in `src/agentcompass/runtime/results/store.py`. Detail shaping lives in `detail.py`, summary construction in `summary.py`, rendering in `render.py`, attempt checkpoints under `runtime/attempts/`, and validated metric types under `runtime/metrics/`.

## Result Directory

`RunStore` reserves a directory according to the entry point. A single `run` uses:

```text theme={"system"}
<results_dir>/[<output.run_name>/]<model>_<benchmark>_<harness>/<run_id>/
```

A named request in `launch` uses:

```text theme={"system"}
<results_dir>/[<output.run_name>/]<requests[].name>/<run_id>/
```

Component IDs and request names are normalized to safe directory components. Launch validates that request output namespaces are distinct after normalization, even if their explicit run IDs differ. Distinct namespaces can use the same run ID.

An explicit run ID must not already exist. Without one, the store generates a timestamp-like ID and advances it when necessary to reserve a unique directory.

| Artifact | Purpose |
| - | - |
| `run_info.json` | schema marker, sanitized request, terminal state, per-attempt resolved plans, reuse source, and metric-artifact provenance |
| `params.json` | compact sanitized Benchmark, Model, execution-attempt, and output identity used by result tooling |
| `details/<state>/<task-id>--<sha256>/task.json` | Shared task metadata, attempt plan, and attempt-to-directory mapping such as `"1": "attempt-1"` |
| `details/<state>/<task-id>--<sha256>/attempt-<n>/result.json` | One attempt result and its retry count; task summaries are derived on read |
| `details/<state>/<task-id>--<sha256>/attempt-<n>/checkpoint.json` | Transient scheduler state and the post-agent evaluation recovery snapshot |
| `details/<state>/<task-id>--<sha256>/attempt-<n>/retries/` | Retry diagnostics and previous execution outputs; excluded from metrics |
| `metrics.json` | authoritative validated `MetricReport` |
| `summary.md` | concise human-readable presentation of the same report |
| `analysis_summary.json` and `analysis_summary.md` | optional Analyzer aggregation |

Temporary files are created next to their target and atomically replaced; no permanent staging directory is required. Sensitive configuration is recursively redacted. Metric observations, ground truth, and final answers are evaluation facts and remain unchanged by configuration-key redaction.

Task directories use the readable task ID plus its full SHA-256 suffix. `<state>` is `running` until the task result is written, then `fatal`, `error` or `normal` according to the most severe final issue of any attempt; warning-only results are `normal`. Moving a task rewrites the run-relative artifact directories stored in its attempt records. Reuse writes to a new run directory and places tasks whose attempts will be retried in `running/`. New output has no task-level `result.json` or error filename prefix. Shared fields live in `task.json`; each attempt owns its result. The attempt mapping must use `attempt-<n>`. See [Legacy](/en/user_guide/other_features/results/overview#legacy) for the previous layout and reuse support.

## Result Layers

Each result layer has one producer and one responsibility:

| Layer | Producer | Contract |
| - | - | - |
| Raw execution result | `BaseHarness.run_task` or `HarnessFreeBenchmark.run_task` | `RunResult` records status, answer, trajectory, artifacts, Harness telemetry, and execution errors; it has no authoritative Benchmark verdict yet |
| Evaluated attempt | `Benchmark.evaluate` | preserves the execution result, writes Contract-declared observations to `RunResult.metrics`, and records evaluation errors or Benchmark evidence |
| Attempt checkpoint | attempt scheduler | stores one terminal attempt independently, allowing the other attempts of the same task to resume without rerunning it |
| Task detail | `build_detail_record` | stores task identity, category, ground truth, `attempt_plan`, retry counts, and the attempt map in the strict persisted shape |
| Request metric report | reducers and `Benchmark.aggregate_metrics` | reduces attempts per task, applies the official cross-task formula, and returns a validated `MetricReport` |
| Request presentation | `RunStore.save_results` | writes `metrics.json` and `summary.md`; `report.html` is retired |
| Orchestration result | `Orchestrator` | records each named request's terminal status, error, and available paths |

Every persisted attempt always contains `status`, `metrics`, `final_answer`, `trajectory`, `error`, `artifacts`, `analysis_result`, and `meta.benchmark`/`meta.harness`. Optional namespace contents may be empty, but the fields remain present. Each attempt result stores its own `retry_count`; the logical task total and `retry_counts` map are derived on read.

`MetricReport` records the resolved attempt plan and metric series. Counts satisfy `evaluated + unavailable + invalidated = total`; `error` is an overlapping diagnostic. ERROR may retain valid observations. Final FATAL invalidates the whole task across all series, sets `evaluation_failed=true`, and suppresses every official `value`. `reference_value` reports only valid tasks, with explicit coverage; missing observations are never filled with zero.

## Resume and Reuse

Resume and reuse share the same persisted records but select their source differently:

* Resume reads details and attempt checkpoints already present in the current run directory.
* Reuse resolves another run through `reuse_run_id` or the newest run within the same output namespace, validates compatibility, then copies usable details and checkpoints into a new run directory.

Existing result directories are not moved or renamed. Automatic reuse searches only the current output hierarchy; it does not fall back to the former Benchmark/Model hierarchy.

Before reuse, the store requires the same Benchmark ID and `execution.attempts` plan. Tasks are matched by task ID and checkpoints by attempt number. Request parameters and task-input fingerprints are not compared, so pending fresh evaluations can use updated evaluation settings. Complete results retain ordinary reuse behavior.

Current run-info v3 uses structured issues. Previous v2 results are adapted only at read boundaries; classification-unknown failures cannot be reused for current scoring. Historical directories are not modified.

A complete compatible detail without errors is copied as a unit. A detail containing execution or evaluation errors is not reused as a whole, and its `k`-attempt plan is scheduled again. A fresh-mode attempt with a clean post-agent completion record can resume evaluation from its artifacts; otherwise the agent runs again.

If an interrupted source task has no complete detail, compatible attempt checkpoints are materialized independently. For example, if attempts 1 and 2 of `k=3` were completed before interruption, resume or reuse schedules only attempt 3, then builds the final detail from all three attempts. Runtime retries within an attempt also leave its completed sibling attempts intact.

Resolved execution plans for materialized attempts are copied into the target `run_info.json`, keeping the new run self-contained. Incompatible Benchmark IDs, attempt plans, or task/attempt identities, and malformed details or checkpoints that must be read, fail explicitly rather than being translated.

## Change Checklist

* Keep every selected `task_id` stable, non-empty, and unique.
* Keep attempt observations in `metrics`, Benchmark evidence in `meta.benchmark`, and Harness diagnostics in `meta.harness.telemetry`.
* Preserve atomic writes, task/attempt identity checks, artifact integrity checks, retry isolation, and per-attempt checkpoints.
* Validate changes against both a new run and an interrupted `k>1` run.
* Keep every `MetricSeries.counts` aligned with the series' actual denominator.

Continue with [Runtime Contracts and Planning](/en/developer_guide/architecture/contracts) for in-memory types, or [Execution, Scheduling, and Cleanup](/en/developer_guide/architecture/execution_lifecycle) for attempt and retry behavior.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.