Choose a Policy for Each Phase
Network policy is resolved independently for every task. A Benchmark loader can declare the policy fields on eachTaskSpec, so two samples in the same run can use different policies. Values passed through --env-params have
run-wide scope and explicitly override the corresponding phase for every selected sample:
These are not four successive phases. Shared evaluation has one environment baseline plus run/evaluation overrides. Fresh evaluation has two environments, each with its own baseline, and still only two execution-phase overrides. Harbor maps
[environment].network_mode to the agent baseline and [verifier.environment].network_mode to the separate verifier baseline; [agent].network_mode and [verifier].network_mode remain the phase overrides. Request-level evaluation_baseline_network_policy requires fresh mode, and overrides a common request baseline before falling back to the task’s verifier baseline or the inherited agent baseline. An explicit empty Harbor verifier environment uses its own default public baseline.
The existing phase fields are resolved with this precedence:
Omitting a field means “continue to the next source”; if no source supplies that phase, it becomes
public. Explicitly
passing public and leaving a phase unset therefore produce the same effective policy when no other source supplies
one. Use --env-params when intentionally forcing one policy across a run. To preserve official per-sample
behavior, omit those overrides and let the Benchmark load its task policy. A compatible Recipe may validate
execution-required endpoints, but it does not change the resolved policy; the effective policy and applied_recipes
record the resulting decision.
Harness setup happens before
run_network_policy is applied. This lets a trusted harness install its runtime under
the baseline policy and then execute the untrusted agent under a stricter policy. Network phases are function-level
trust boundaries, not a separate policy for every progress stage.Understand Function-Level Boundaries
AgentCompass applies the same evaluation rule to both evaluation Environment modes: only the completebenchmark.evaluate() call uses evaluation_network_policy. The modes differ because fresh creates another
Environment, while reuse evaluates in the agent Environment. The evaluate_environment progress phase in a fresh
run is not a separate Benchmark hook; its setup boundary is the evaluation provider’s open() call.
The tables below show the active policy for Environment operations. Provider control-plane requests made by the
AgentCompass host remain outside sandbox network enforcement. A function can contain several internal operations; the
runtime does not split those operations into finer policy scopes.
Shared Environment
Withreuse, the same Environment remains active while its lifecycle functions move through the baseline, run, and
evaluation policies:
Immediately before step 7, the runtime switches directly from
run_network_policy to evaluation_network_policy,
calls the entire benchmark.evaluate() function, and restores baseline_network_policy in finally. This keeps the
evaluation policy scoped to the evaluation function without an intermediate baseline transition.
Separate Environment
fresh completes the agent Environment lifecycle first, then creates a separate evaluation Environment:
Step 8 is the conditional evaluation Environment setup introduced by selecting
fresh; it is not an optional
Benchmark hook. The runtime switches to the evaluation policy only after evaluation_provider.open() completes,
restores the evaluation Environment baseline immediately after benchmark.evaluate(), and then releases or retains
that Environment.
Runtime executes declared collection commands directly, then collects existing files when storage is enabled. Without commands, it skips preparation. Commands and downloads remain inside the run trust boundary; they do not create another network phase.
Network Modes
Each phase accepts one of four modes:
Use a string for
public or no-network:
api.*.example.com. Both allowlists and denylists accept a
bare host such as www.example.com, a host:port shorthand such as www.example.com:443, or a target object such as
{"host": "www.example.com", "port": 443}. Bare hosts retain all-port semantics; port values must be integers from 1
through 65535. Use [2001:db8::1]:443 when adding a port to an IPv6 address. Providers that cannot enforce a
port-specific target reject it.
Tell the Agent About Rollout Restrictions
By default, each Harness appends an English network restriction statement to its final user instruction when the effectiverun_network_policy is no-network, allowlist, or denylist. This prevents an agent from treating an
intentional restriction as a transient network failure and repeatedly retrying blocked operations. A public policy
does not add a statement.
The statement is based on the provider-resolved policy, not only the requested value. Recipe-added targets therefore
appear in an allowlist statement, and allowlist or denylist targets are expanded from the effective policy. A provider
that cannot enforce the requested policy rejects the run before rollout instead of injecting a misleading statement.
no-network uses the same wrapper without a target list:
denylist, the effective denied targets replace the list:
benchmark.evaluate() continue to receive the original input. Standard Harness paths support plain prompts,
structured user messages, multimodal user messages, and the JSON-encoded messages accepted by openai_chat. The
harness-free TauBench inference path does not currently consume this notice.
Disable injection through the selected Harness configuration:
inject_network_restriction_notice: false under the selected harnesses.<id> entry in a
YAML or JSON configuration file. Python callers pass the same field in harness_params. This setting changes only the
Harness input notice; it does not change or disable network-policy enforcement.
Pass Phase Policies from the CLI
Pass the three runtime policies as JSON fields in--env-params:
The same values can be supplied as
environment_params through the Python API, or by a
Benchmark loader on TaskSpec for sample-level policy.
Select the Narrowest Practical Policy
Use this decision sequence:- Check the benchmark page for an official or recommended policy.
- Identify where the harness is installed and where it calls the model API.
- Keep the baseline
publicif the sandbox must install a package or executable; otherwise prefer an allowlist or a prebuilt image. - Set the run phase to
no-networkwhen the task should use only local evidence. - Include any endpoints needed by Harness close or artifact collection in the run policy; do not broaden it between rollout and verification.
- Add only the exact model, search, judge, package, or artifact hosts required by a network-dependent phase.
- Run one task and inspect the resolved execution plan before scaling.
harness.start_session run inside the environment and therefore use
the baseline policy. If the baseline must also be no-network, put those dependencies in the task image or snapshot first.
Whether a model endpoint needs to be allowlisted depends on where the harness makes its request:
- A local harness process calls the model from the AgentCompass host, outside the task environment policy.
- A harness running inside the sandbox needs the model endpoint in the run-phase allowlist.
- Some benchmark recipes, including DeepSWE recipes, infer the resolved model endpoint. Do not assume every custom benchmark or external recipe does so; inspect the resolved plan.
- The meaning of
run_network_policyis independent of where it came from. DeepSWE validates that its required model endpoints are permitted and fails during planning if the effective task or CLI policy blocks them; it never broadens that policy. A remote Harness therefore needs an explicit CLI allowlist containing the model host; a local Harness can keepno-network.
Provider Support
Daytona accepts at most 20 domain entries or 10 IPv4 network entries. Docker cannot use dynamic phase policies with
network set to none, host, or container:<id>. Its default egress proxy image is
python:3.12-alpine; make sure the Docker daemon can pull it or pre-pull it on an offline host.
Outbound policy accepts only the shared fields above. Provider-native block-all and allowlist parameters are rejected; adapters generate those API options from the resolved shared policy.
Docker is currently the only documented public provider that enforces
denylist. Other public providers reject that
policy instead of silently broadening it to public.Verify the Effective Policy
Run one known task with persistent debug logs:baseline_network_mode, run_network_mode, and evaluation_network_mode when each task execution
plan is built. Per-task details retain the resolved policies and applied_recipes. Verify those effective values rather
than relying only on the original command, because a Benchmark Recipe may add an inferred endpoint or provider
adaptation.
For an adversarial isolation test, ask the agent to access a known external URL and confirm both outcomes:
- the request fails during the restricted run phase; and
- the same environment can still perform the trusted setup work allowed by its baseline policy.
NetworkOperationAnalyzer can summarize commands such
as curl, wget, package installation, or git clone. It observes agent behavior but does not enforce the policy and
cannot replace provider transition logs.
