Skip to main content
Implement a Harness as the owner of one agent loop: validate compatibility, start its runtime, execute one PreparedTask, normalize a public RunResult, and release everything it created. The tutorial adapter below returns a configured answer instead of calling a model. That makes registry and lifecycle smoke tests deterministic. Replace only its execution body when integrating a real SDK or CLI; keep the same public contracts.

Record the Upstream Contract

Record the official framework or CLI version, supported model protocols, configuration format, prompt flow, tool and workspace behavior, installation method, timeouts, termination rules, trajectory format, and credential handling. Pin the version when it affects commands, prompts, parsing, or reproducibility, and prefer the public SDK or CLI over private functions.

Create the Minimal File

The smallest complete integration needs one implementation file and one package export:
Create src/agentcompass/harnesses/example_answer.py:
This implements supports(), start_session(), and the execution hook execute_task(). The inherited collect_result() returns the supplied result unchanged when no additional recovery is needed. It also shows close_session() explicitly even though the base class provides a no-op. BaseHarness.build_plan() copies matching config fields into plan_class, including the inherited inject_network_restriction_notice field.

Export and Inspect the Registration

Add the import to src/agentcompass/harnesses/__init__.py:
Then inspect discovery and the generated config schema:
The first command should contain example_answer and its description. The second should show answer with default Paris and the inherited network-notice field. An absent ID indicates an import or registration failure; a present ID does not prove that a real upstream runtime can install or launch.

Run One Task

Use example_exact_match from the Harness-driven Benchmark tutorial:
The command needs no endpoint because this tutorial Harness does not call req.model. The terminal result should complete with one selected task and report paths.run_info; its parent directory is the run directory. That directory should contain run_info.json, params.json, progress.json, progress.jsonl, run.log, task and attempt result.json files, metrics.json, and summary.md. After Benchmark evaluation, the detail attempt should contain status: "completed", final_answer: "Paris", metrics.correct: true, and meta.harness.telemetry.answer_characters: 5. For a real Harness, repeat the smoke with one bounded upstream task and its actual credentials. A registry check alone never exercises installation, launch, parsing, cleanup, or the model endpoint.

Replace the Tutorial Execution Body

Keep public parameters in RuntimeHarnessConfig so the CLI, Python SDK, config files, and generated documentation resolve the same fields. Put normalized runtime choices such as version, launch mode, install strategy, step limit, command timeout, and cost behavior in a typed HarnessPlan. Do not mutate RunRequest, inspect private Benchmark fields, or persist secrets in a plan. Implement supports(environment, model) around capabilities rather than Benchmark IDs. Validate model protocol, shell and filesystem needs, endpoint forwarding, browser or GUI needs, workspace assumptions, credential location, and whether installation can work in the selected Environment. Reject unsupported combinations before Environment startup and never silently switch protocols, providers, install modes, or models. The runtime calls the lifecycle in this order:
Use start_session() for trusted installation, config generation, uploads, clients, or background services. execute_task() executes exactly one PreparedTask and consumes only public fields such as prepared.input.prompt, messages, files, media, tools, and workspace. close_session() releases Harness-owned clients, processes, servers, temporary config, and background tasks after success, timeout, cancellation, or error; the runtime, not the Harness, closes the Environment.

Recover Results After Execution

The runtime calls collect_result(session, prepared, req, plan, result) after normal completion, a phase timeout, or an execution error, before closing the Harness session. Override this hook to load saved output, trajectory, or candidate files. Initialize recovery state in the session before awaiting the agent. The hook must only recover existing outputs; it must not resume the agent, issue model requests, or generate another candidate. Return a RunResult; the runtime preserves the original execution error even when recovery succeeds. execution.harness_result_timeout_seconds bounds recovery separately (default 60 seconds, finite and positive). Run/evaluation multipliers do not affect it. Recovery failure is recorded without replacing the original run error. Explicit caller cancellation propagates to cleanup instead of starting recovery. Existing extensions that implement run_task() remain supported through the default execute_task() bridge, but must implement collect_result() to recover outputs independently of their execution coroutine. For new extensions, implement execute_task(); direct BaseHarness.run_task() calls compose execution and collection without the runtime’s phase watchdogs. Declare a Harness fallback with default_run_timeout_seconds on the class. A legacy non-null HarnessPlan.timeout default without an explicit class declaration now fails planning with migration guidance. Use an explicit None for no Harness default. The Planner alone maps the final shared run budget to native plan.timeout; do not expose another user-facing timeout field.

Normalize Results Without Scoring

A Harness reports execution, not Benchmark correctness. Return the best available final_answer, requested files, ordered trajectory, token usage, timing, artifacts, and an accurate TaskStatus. Preserve timeout, refusal, invalid-output, termination, installation, launch, parsing, and model API errors. Do not write Benchmark observations to RunResult.metrics in the Harness, and do not turn evaluator failure into Harness failure. Put Harness diagnostics in RunResult.telemetry. A process exit code of zero is not automatically a correct result. Ensure the public answer is assigned to RunResult.final_answer; a Benchmark must not need to recover it from a Harness-private artifact. Support only reproducible installation strategies: a pinned preinstalled image, controlled installation before restricted execution, or an isolated driver-side optional extra. Do not assume every image has a package manager or compiler, and do not broaden the run network policy to make installation convenient. Inject model, judge, search, and provider credentials through supported Environment or configuration mechanisms. Recursively redact them from commands, files, logs, trajectories, URLs, exceptions, dataclass representations, and persisted metadata. Give each limit one owner: the shared execution deadline bounds the agent phase, command timeout bounds one tool command, step limits bound the agent loop, model-client retries handle request transport, and runtime retries repeat a failed task attempt. Unknown model pricing must follow the Harness cost contract and should not terminate a run when the user explicitly selected an ignore-errors or disabled-cost mode. Use the scoped Environment for agent variables: run_env_variables is active during start_session() and execute_task(), and is removed before collect_result() and close_session(). Nested native tool configurations should receive env.get_task_env_variables() (or merge their defaults via env._merge_exec_env()). Do not add a Harness env config/plan field. Pass native protocol requirements as env.exec(..., required_env={...}); the shared merge injects missing values and rejects conflicting bindings before starting the command. Keep overridable defaults in env. Prefer runner arguments or a Python bootstrap for internal import paths so the user’s PYTHONPATH remains available. Custom Environment sessions must implement and forward the required_env keyword.

Diagnose Failures by Stage

For a compact real lifecycle and result adapter, read qwen3vl_gui.py. For an agent that runs directly in the AgentCompass process with host_process, collects the final answer, and converts a trajectory, read naive_search_agent/harness.py. Harness config dataclasses define the accepted parameter keys. Unknown keys fail during config construction; Recipe hints must be declared fields or handled through the typed plan. Use command_with_path_updates from agentcompass.utils.command to prepend or append internal POSIX search paths. It expands inherited values inside the target environment after the common/phase merge, quotes paths and arguments, and uses exec to preserve provider process control. Do not read host PATH to construct a sandbox command or start a login shell that resets the resolved environment. Use the shared launchers in agentcompass.utils.command for shipped modules, scripts, or Python console scripts; they add source paths to that interpreter’s sys.path while retaining the user’s PYTHONPATH. Script launchers preserve the script’s import directory and remove the implicit working-directory entry introduced by python -c; an explicit working-directory entry in PYTHONPATH remains effective. These helpers are internal and introduce no CLI fields. Never mutate host os.environ for a task; configure HTTP clients directly and pass subprocess settings through Environment.