> ## Documentation Index
> Fetch the complete documentation index at: https://opencompass-docs-preview-pr-335-0.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# Summary and Analysis Results

Run-level outputs are split by purpose. Benchmark metric files describe measured performance; analysis files describe patterns found in trajectories, errors, and diagnostics without changing Benchmark observations.

## Files at a Glance

<table style={{width:'100%', tableLayout:'fixed'}}>
  <thead>
    <tr><th style={{width:'36%'}}>File</th><th style={{width:'34%'}}>When it is generated</th><th style={{width:'30%'}}>What to read it for</th></tr>
  </thead>

  <tbody>
    <tr><td><code>summary.md</code></td><td>Benchmark metric aggregation succeeds.</td><td>Readable metrics adapted to single or repeated attempts.</td></tr>
    <tr><td><code>metrics.json</code></td><td>From the same metric report as <code>summary.md</code>.</td><td>Canonical values, per-series counts, categories, and hierarchy.</td></tr>
    <tr><td><code>analysis\_summary.json</code></td><td>At least one saved attempt contains aggregatable analyzer output.</td><td>Analyzer statistics, analyzer error counts, bad-case file indexes, and distributions.</td></tr>
    <tr><td><code>analysis\_summary.md</code></td><td>From the same analyzer aggregation as the JSON file.</td><td>Readable overall, category, and distribution analysis.</td></tr>
  </tbody>
</table>

A run that stops before aggregation can lack some or all of these files. The Markdown and JSON files are written separately, so an interrupted write can also leave an incomplete set; rerun the corresponding `summary` or `analysis` command.

## Benchmark Metric Outputs

Both Benchmark metric files are projections of one strictly validated metric report. They do not run independent aggregators.

### `summary.md`

Start here for a readable result. For every `k`, the Markdown starts with the Model, `Total`, `Evaluated`, `Error`, and `Metrics`. The top-level counts come from the primary series selected by the current strategy; other series can use different valid samples, so use the counts on their own rows when reading them.

For `k=1`, `Metrics` keeps the traditional two-column name/value table, followed by available category or hierarchy details. For `k>1`, only the metric area expands: the summary adds a one-line attempt plan, and the `Metrics` table distinguishes headline and auxiliary series with `Role` while listing each reducer, actual run-level formula, value, and independent coverage counts. Complete structured breakdowns remain in `metrics.json`.

For `k>1`, the summary does not include a `first` result. Under the `avg` strategy, a binary primary can show both `avg@k` and `pass@k`, while a scalar primary itself produces only `avg@k`. If that scalar-primary contract also declares binary auxiliary metrics, each binary auxiliary can still produce its own `avg@k` and `pass@k` series. The Benchmark still cannot use the `pass` execution strategy because that strategy is controlled by the scalar primary. See [Metrics and Aggregation](/en/user_guide/other_features/results/metrics_aggregation#know-which-series-are-produced).

### `metrics.json`

`metrics.json` is the machine-readable source of truth:

```json theme={"system"}
{
  "k": 3,
  "strategy": "avg",
  "aggregation": "micro_weighted",
  "series": [
    {
      "series_id": "correct.avg@3",
      "metric_id": "correct",
      "kind": "binary_success",
      "reducer": "avg",
      "role": "headline",
      "aggregation": "micro_weighted",
      "k": 3,
      "value": 0.61,
      "counts": {
        "total": 100,
        "evaluated": 99,
        "error": 3,
        "unavailable": 1,
        "invalidated": 0
      },
      "categories": {
        "example": {
          "value": 0.61,
          "reference_value": null,
          "counts": {
            "total": 100,
            "evaluated": 99,
            "error": 3,
            "unavailable": 1,
            "invalidated": 0
          }
        }
      },
      "hierarchy": {},
      "extra": {},
      "reference_value": null
    },
    {
      "series_id": "correct.pass@3",
      "metric_id": "correct",
      "kind": "binary_success",
      "reducer": "pass",
      "role": "headline",
      "aggregation": "micro_weighted",
      "k": 3,
      "value": 0.8,
      "counts": {
        "total": 100,
        "evaluated": 99,
        "error": 3,
        "unavailable": 1,
        "invalidated": 0
      },
      "categories": {
        "example": {
          "value": 0.8,
          "reference_value": null,
          "counts": {
            "total": 100,
            "evaluated": 99,
            "error": 3,
            "unavailable": 1,
            "invalidated": 0
          }
        }
      },
      "hierarchy": {},
      "extra": {},
      "reference_value": null
    }
  ],
  "extra": {},
  "evaluation_failed": false
}
```

<table style={{width:'100%', tableLayout:'fixed'}}>
  <thead>
    <tr><th style={{width:'38%'}}>Field</th><th style={{width:'62%'}}>Meaning</th></tr>
  </thead>

  <tbody>
    <tr><td><code>k</code>, <code>strategy</code></td><td>Resolved repeated-attempt plan.</td></tr>
    <tr><td><code>aggregation</code></td><td>Actual run-level policy: <code>micro\_weighted</code>, <code>category\_mean</code>, or <code>category\_hierarchy</code>.</td></tr>
    <tr><td><code>series\[].series\_id</code></td><td>Stable <code>\<metric>.\<reducer>@\<k></code> identity.</td></tr>
    <tr><td><code>series\[].kind</code></td><td><code>binary\_success</code> or <code>scalar</code>.</td></tr>
    <tr><td><code>series\[].role</code></td><td><code>headline</code> for the Benchmark primary metric, otherwise <code>auxiliary</code>.</td></tr>
    <tr><td><code>series\[].aggregation</code></td><td>The formula that produced this series, such as <code>micro\_weighted</code>, <code>ratio\_of\_sums</code>, <code>sum</code>, or <code>benchmark</code>.</td></tr>
    <tr><td><code>series\[].value</code></td><td>Aggregated value, or <code>null</code> when it cannot be computed exactly.</td></tr>
    <tr><td><code>series\[].counts</code></td><td>Independent <code>total</code>, <code>evaluated</code>, <code>error</code>, <code>unavailable</code>, and <code>invalidated</code> task counts for this series.</td></tr>
    <tr><td><code>series\[].categories</code></td><td>Category keys mapped to their own <code>value</code>, <code>counts</code>, and optional formula <code>aggregation\_weight</code>.</td></tr>
    <tr><td><code>series\[].hierarchy</code></td><td>Hierarchy paths mapped to the same breakdown fields; populated only for explicit hierarchy aggregation.</td></tr>
    <tr><td><code>series\[].extra</code></td><td>Auditable inputs for this formula, such as numerator and denominator metric IDs and totals.</td></tr>
    <tr><td><code>extra</code></td><td>Benchmark-level structured diagnostics, such as rank or medal comparison details.</td></tr>
  </tbody>
</table>

Counts are not global run aliases. Always use the counts beside the series you are reading; missing observations can make denominators differ across metrics or reducers. `evaluated + unavailable + invalidated` equals `total`, while `error` is an independent diagnostic and may overlap either count.

## Analysis Summaries

Analyzer output first appears under `attempts.<N>.analysis_result.<analyzer-family>`. Analysis aggregation is independent of the Benchmark Metric Contract.

### Combine Multiple Attempts

AgentCompass does not choose a generic “best attempt.” For each task and analyzer family it combines saved attempts as follows:

1. `is_badcase` is true when any attempt reports true.
2. Numeric analyzer scores are averaged across attempts that provide one.
3. The latest non-empty analyzer payload supplies diagnostic fields used by distributions.
4. The task is counted at most once for that analyzer.

When combining analyzer families, a task is a bad case if any family marks it. The combined `avg_score` uses the maximum available family score for each task before averaging tasks.

### `analysis_summary.json`

<table style={{width:'100%', tableLayout:'fixed'}}>
  <thead>
    <tr><th style={{width:'38%'}}>Field</th><th style={{width:'62%'}}>Contents</th></tr>
  </thead>

  <tbody>
    <tr><td><code>per\_category\_per\_analyzer</code></td><td>One statistics row for each retained category and analyzer combination.</td></tr>
    <tr><td><code>per\_category\_overall</code></td><td>One row per category, combining analyzer families.</td></tr>
    <tr><td><code>overall\_per\_analyzer</code></td><td>One row per analyzer across categories, with an <code>items</code> list of matching bad-case detail files.</td></tr>
    <tr><td><code>overall</code></td><td>Statistics combined across categories and analyzers.</td></tr>
    <tr><td><code>distributions</code></td><td>Analyzer-declared value counts or numeric distributions.</td></tr>
  </tbody>
</table>

Statistics rows use these fields:

<table style={{width:'100%', tableLayout:'fixed'}}>
  <thead>
    <tr><th style={{width:'38%'}}>Field</th><th style={{width:'62%'}}>Meaning</th></tr>
  </thead>

  <tbody>
    <tr><td><code>category</code></td><td>Task category; overall rows use <code>**overall**</code>, and uncategorized tasks use <code>(no category)</code>.</td></tr>
    <tr><td><code>analyzer</code></td><td>Analyzer family ID; combined rows use <code>**overall**</code>.</td></tr>
    <tr><td><code>total</code></td><td>Tasks in this scope that contain the analysis result.</td></tr>
    <tr><td><code>badcase\_count</code>, <code>badcase\_ratio</code></td><td>Number and fraction of tasks marked as bad cases.</td></tr>
    <tr><td><code>error\_count</code></td><td>Tasks whose selected analysis result contains a non-empty <code>error</code>.</td></tr>
    <tr><td><code>avg\_score</code></td><td>Average available analyzer score, or <code>null</code>.</td></tr>
    <tr><td><code>items</code></td><td>Only in <code>overall\_per\_analyzer</code>; filenames marked by that analyzer.</td></tr>
  </tbody>
</table>

Analyzers can declare distribution fields using `value_counts` or `numeric_stats`. Value counts keep the 50 most frequent values. Numeric output contains `count`, `min`, `mean`, `p50`, `p90`, `p95`, and `max` when data exists.

In rows that combine all analyzers, `badcase_count` is the number of tasks marked by at least one analyzer and `error_count` is the number with at least one analyzer error; neither is the sum of the per-analyzer rows. Bad-case analyzers with neither bad cases nor errors can be omitted, while statistics-only analyzers and analyzers with errors remain. The Markdown version presents `Total`, `Badcase`, `Error`, `Badcase Ratio`, and `Avg Score` tables plus distributions, but omits the full `items` indexes.

## Generate or Regenerate Outputs

[`agentcompass run`](/en/user_guide/using_agentcompass/cli/run) and [`agentcompass launch`](/en/user_guide/using_agentcompass/cli/launch) generate Benchmark outputs after aggregation. [`agentcompass summary`](/en/user_guide/using_agentcompass/cli/summary) strictly reloads task details and the persisted attempt plan, then replaces `summary.md` and `metrics.json` without rerunning attempts or analyzers. Both paths update `run_info.json.metric_artifacts` with the files' source and report plan.

[`agentcompass analysis`](/en/user_guide/using_agentcompass/cli/analysis#re-run-on-existing-results) can update per-attempt `analysis_result` and regenerate the two analysis summary files without rerunning the agent or Benchmark verifier. It does not update Benchmark metric outputs.

## Share Results Safely

Generated files can contain task IDs, categories, analyzer values, or other integration data. They do not receive a complete content-safety or Markdown sanitization pass. Inspect them before sharing.

## Related Pages

* [Results Overview](/en/user_guide/other_features/results/overview)
* [Task Results](/en/user_guide/other_features/results/task_results)
* [Metrics and Aggregation](/en/user_guide/other_features/results/metrics_aggregation)
* [`agentcompass summary`](/en/user_guide/using_agentcompass/cli/summary)
* [`agentcompass analysis`](/en/user_guide/using_agentcompass/cli/analysis)

A final FATAL sets `evaluation_failed=true`. All official `value` fields become null; `reference_value` uses only non-invalidated tasks and is null when none can be scored. The CLI exits nonzero and the failed run retains its result paths.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.