> ## Documentation Index
> Fetch the complete documentation index at: https://opencompass-docs-preview-pr-335-0.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# DeepSearchQA

DeepSearchQA ([arxiv](https://arxiv.org/abs/2601.20975)) evaluates a deep-research agent's ability to retrieve and answer across multiple knowledge domains: given a question that requires web search and multi-step evidence gathering, the agent produces a final answer, which an **LLM judge \*\* then grades as correct or not against the official rubric. The dataset contains \*\* 900 tasks** spanning 17 categories, with questions split by answer form into Single Answer and Set Answer.

Unlike pairwise-judged benchmarks such as GDPval, DeepSearchQA uses single-sided judging. The judge only compares the agent-under-test's answer against the ground truth, checking item by item whether it is hit, without comparing to any baseline. Both inference and judging run in the local process (`host_process`) — the harness first drives the model under test through the search loop to produce a final answer, then the judge model grades it.

## How it works

A DeepSearchQA run has two stages — inference and judging — where the judging stage applies different criteria based on the task's answer form.

### Inference and judging

* **Inference.** The model under test acts as a search agent and, driven by the harness (default [`naive_search_agent`](/en/user_guide/modules/harnesses/naive_search_agent)), completes multi-turn tool loops such as search / visit per task, producing a natural-language answer.
* **Judging.** The judge model (`judge_model`) receives "question + ground truth + answer form + answer under test" and grades it with the official rubric template. The judge and the model under test are two separate endpoints; `judge_model` must be specified explicitly.

### How the two answer forms are judged

The judge applies different criteria based on each task's `answer_type`:

* **Single Answer (316 tasks):** the answer under test is judged correct if it semantically hits the ground truth; verbatim matching is not required.
* **Set Answer (584 tasks):** the ground truth is a set of items, and the answer under test must \*\* hit every item **; the judge also checks whether the answer includes \*\* excessive answers** beyond the ground truth.

The judge outputs three parts: `Correctness Details` (a per-item boolean dictionary of hits), `Excessive Answers` (a list of extra answers), and `Explanation` (the grading rationale). A task is judged **correct \*\* if and only if \*\* all expected items are hit \*\* and \*\* no excessive answers exist**; any missing item or any excessive answer counts as incorrect.

## Parameters

Pass a JSON object via `--benchmark-params '{...}'`, or a `benchmark.params` block in the YAML given to `--config`; the CLI wins on shared keys. See the [Benchmark overview](/en/user_guide/modules/benchmarks/overview) for merge precedence.

### Parameter reference

<div style={{overflowX:'auto'}}>
  <table style={{minWidth:'1040px', width:'100%'}}>
    <colgroup>
      <col width="18%" />

      <col width="16%" />

      <col width="15%" />

      <col width="20%" />

      <col width="31%" />
    </colgroup>

    <thead>
      <tr><th style={{whiteSpace:'nowrap'}}>Parameter</th><th style={{whiteSpace:'nowrap'}}>Type</th><th style={{whiteSpace:'nowrap'}}>Default</th><th>Choices / values</th><th>Description</th></tr>
    </thead>

    <tbody>
      <tr><td style={{whiteSpace:'nowrap'}}><code>judge\_model</code></td><td style={{whiteSpace:'nowrap'}}>dict</td><td style={{whiteSpace:'nowrap'}}><code>null</code></td><td><code>id</code>, <code>base\_url</code>, <code>api\_key</code>, <code>api\_protocol</code>, <code>params</code></td><td>Judge model spec, <strong>required</strong> (see <a href="#judge-model-spec">Judge model spec</a>). It decides grading, and is not the CLI <code>--model-\*</code>.</td></tr>
      <tr><td style={{whiteSpace:'nowrap'}}><code>category</code></td><td style={{whiteSpace:'nowrap'}}>string / list</td><td style={{whiteSpace:'nowrap'}}><code>"all"</code></td><td><code>"all"</code>, a single category name, or a list of category names (17 listed below)</td><td>Filter tasks by category; <code>"all"</code> = no filter. A list takes the union.</td></tr>
      <tr><td style={{whiteSpace:'nowrap'}}><code>answer\_type</code></td><td style={{whiteSpace:'nowrap'}}>string</td><td style={{whiteSpace:'nowrap'}}><code>"all"</code></td><td><code>all</code> / <code>Single Answer</code> / <code>Set Answer</code></td><td>Filter tasks by answer form; <code>all</code> = no filter. Case and full name must match exactly.</td></tr>
    </tbody>
  </table>
</div>

Shared Benchmark fields such as `sample_ids` follow [Benchmark Parameters](/en/user_guide/modules/benchmarks/overview). Configure repeated attempts with `--k` and `--attempt-strategy`; see [Metrics and Aggregation](/en/user_guide/other_features/results/metrics_aggregation).

<Accordion title="All 17 category values (click to expand)">
  `Politics & Government` (148), `Finance & Economics` (132), `Geography` (95), `Education` (94), `Health` (92), `Science` (90), `Other` (65), `History` (44), `Travel` (36), `Media & Entertainment` (29), `Arts` (26), `Technology` (22), `Sports` (20), `Current Events` (3), `Biology` (2), `Linguistics` (1), `Arts & Entertainment` (1). Numbers in parentheses are the task count per category (900 total).
</Accordion>

<a id="judge-model-spec" />

### Judge model spec

`judge_model` is passed as a dict with the fields `id`, `base_url`, `api_key`, `api_protocol`, and `params`, pointing to the judge model's own endpoint, with inference parameters under `params`.

We recommend **fixing a single judge** across all models under test. Grading directly decides the scores, so switching judges makes scores no longer comparable across models; likewise, the model under test should not serve as its own judge, as that is neither fair nor comparable. The judge need not be especially strong — DeepSearchQA's criteria (semantic hit + excess check) are relatively objective, so a mid-sized model suffices. AgentCompass recommends `Qwen3.6-35B-A3B`.

## Run examples

The three positional arguments to `agentcompass run` are Benchmark, Harness, and Model. These examples use `deepsearchqa`, [`naive_search_agent`](/en/user_guide/modules/harnesses/naive_search_agent), and `$MODEL_NAME`, with [`host_process`](/en/user_guide/modules/environments/providers/host_process) as the Environment.

Set the following environment variables in your terminal before running the examples:

* Model under test: `MODEL_NAME`, `MODEL_BASE_URL`, and `MODEL_API_KEY`; see [Model connection details](/en/user_guide/modules/models/overview#configure-connection-details).
* Judge Model: `JUDGE_MODEL_NAME`, `JUDGE_MODEL_BASE_URL`, and `JUDGE_MODEL_API_KEY`. Use a fixed, independent judge configuration.
* Search tools: `SERPER_API_KEY` for `search` and `JINA_API_KEY` for `visit`.

See the [run command](/en/user_guide/using_agentcompass/cli/run) for configuration ownership and CLI overrides.

<Tabs>
  <Tab title="Smoke test (single task end-to-end)">
    Use `sample_ids` to evaluate a single task, verifying that the end-to-end inference and judging flow works; defaults for the rest.

    ```bash wrap theme={"system"}
    agentcompass run \
      deepsearchqa \
      naive_search_agent \
      "$MODEL_NAME" \
      --env host_process \
      --benchmark-params '{
        "judge_model": {
          "id": "'"$JUDGE_MODEL_NAME"'",
          "base_url": "'"$JUDGE_MODEL_BASE_URL"'",
          "api_key": "'"$JUDGE_MODEL_API_KEY"'"
        },
        "sample_ids": ["1"]
      }' \
      --harness-params '{
        "serper_api_key": "${SERPER_API_KEY}",
        "jina_api_key": "${JINA_API_KEY}"
      }' \
      --model-base-url "$MODEL_BASE_URL" \
      --model-api-key "$MODEL_API_KEY" \
      --model-api-protocol openai-chat
    ```
  </Tab>

  <Tab title="Custom parameters">
    Evaluate only a subset of categories and answer forms to focus analysis on a specific domain; also demonstrates lowering the iteration limit in `--harness-params`.

    ```bash wrap theme={"system"}
    agentcompass run \
      deepsearchqa \
      naive_search_agent \
      "$MODEL_NAME" \
      --env host_process \
      --benchmark-params '{
        "judge_model": {
          "id": "'"$JUDGE_MODEL_NAME"'",
          "base_url": "'"$JUDGE_MODEL_BASE_URL"'",
          "api_key": "'"$JUDGE_MODEL_API_KEY"'"
        },
        "category": ["Science", "Geography"],
        "answer_type": "Set Answer"
      }' \
      --harness-params '{
        "max_iterations": 40,
        "serper_api_key": "${SERPER_API_KEY}",
        "jina_api_key": "${JINA_API_KEY}"
      }' \
      --model-base-url "$MODEL_BASE_URL" \
      --model-api-key "$MODEL_API_KEY" \
      --model-api-protocol openai-chat \
      --task-concurrency 16
    ```
  </Tab>

  <Tab title="AgentCompass recommended config">
    Evaluate all 900 tasks. `--benchmark-params` only needs the judge model `judge_model`; use `--task-concurrency` to raise cross-task concurrency.

    ```bash wrap theme={"system"}
    agentcompass run \
      deepsearchqa \
      naive_search_agent \
      "$MODEL_NAME" \
      --env host_process \
      --benchmark-params '{
        "judge_model": {
          "id": "'"$JUDGE_MODEL_NAME"'",
          "base_url": "'"$JUDGE_MODEL_BASE_URL"'",
          "api_key": "'"$JUDGE_MODEL_API_KEY"'"
        }
      }' \
      --harness-params '{
        "serper_api_key": "${SERPER_API_KEY}",
        "jina_api_key": "${JINA_API_KEY}"
      }' \
      --model-base-url "$MODEL_BASE_URL" \
      --model-api-key "$MODEL_API_KEY" \
      --model-api-protocol openai-chat \
      --task-concurrency 16
    ```
  </Tab>
</Tabs>

<a id="outputs" />

## Evaluation Results

For shared result conventions, see [Run Directory](/en/user_guide/other_features/results/overview#directory-layout), [Aggregate Scores](/en/user_guide/other_features/results/summary_analysis), and [Task Files and Shared Fields](/en/user_guide/other_features/results/task_results).

<a id="metric-contract-and-aggregate-series" />

### Scoring Metrics

DeepSearchQA's primary metric is binary `correct`: under the [answer-form rules](#how-the-two-answer-forms-are-judged), it is `true` only when all expected items are matched with no excessive answers. Item-level judgments explain the result but do not earn partial credit. An empty answer is graded incorrect.

With the default configuration, each task has one attempt and the overall score is accuracy over tasks with valid scores, ranging from 0 to 1; higher is better. Single-answer and set-answer tasks each count as one task, regardless of the number of items in the answer set. Scores cover only the tasks selected by `category`, `answer_type`, or `sample_ids`.

See [Metrics and Aggregation](/en/user_guide/other_features/results/metrics_aggregation) for repeated attempts, category aggregation, and scoring failure rules.

<a id="per-task-details-details" />

### Task Results and Scoring Evidence

Each attempt's scoring record is stored in `scoring` under `meta.benchmark`. A successful judge response produces:

| Field | Contents |
| - | - |
| `evaluation_type` | Fixed as `deepsearchqa_judge`. |
| `correct` | The task-level boolean verdict, matching `metrics.correct`. |
| `all_expected_correct` | Whether every expected answer item was matched. |
| `has_excessive_answers` | Whether answers beyond the reference set are present. |
| `correctness_details` | The match result for each expected item. |
| `excessive_answers` | The list of answer items judged excessive. |
| `explanation` | The judge's grading rationale. |

An empty answer skips the judge and records `correct=false` with `reason=empty_model_response`, without item-level judgments. For judge-call or response-parsing failures, inspect `error` in the same record; parsing failures may also preserve a truncated `raw_response`.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.