> ## Documentation Index
> Fetch the complete documentation index at: https://docs.boxd.sh/llms.txt
> Use this file to discover all available pages before exploring further.

# Harbor

> The agent evaluation framework. Every benchmark trial runs in its own isolated KVM microVM with a full Docker daemon.

<img src="https://mintcdn.com/azin/ZeNTyQegk-ndMEAW/images/boxd-x-harbor.png?fit=max&auto=format&n=ZeNTyQegk-ndMEAW&q=85&s=8535f5634dc35b8ad94c22d1c750ec44" alt="boxd.sh and Harbor, the agent evaluation framework" width="2400" height="1640" data-path="images/boxd-x-harbor.png" />

[Harbor](https://harborframework.com) is a framework for evaluating and optimizing AI agents and models: it runs agents like Claude Code, Codex, or OpenHands against benchmarks like Terminal-Bench and SWE-Bench across many parallel sandboxes. The boxd environment provider gives every trial its own isolated KVM microVM with a full Docker daemon, so Dockerfile and multi-container `docker-compose.yaml` tasks run unmodified:

```bash theme={"theme":"github-dark"}
export BOXD_API_KEY=<your-key>
harbor run --dataset terminal-bench@2.0 --agent claude-code --model anthropic/claude-opus-4-1 -e boxd
```

For benchmarks we recommend starting from snapshots: prepare a machine to your liking, snapshot it, and start every sandbox from there with `--ek from_snapshot=<name>`. This shortens the startup sequence drastically, because each trial restores a machine whose Docker cache already holds the task image instead of building it from scratch.

The provider is [awaiting merge upstream](https://github.com/harbor-framework/harbor/pull/2778). You don't have to wait for it: install Harbor from the pull request's branch and everything above works today.

```bash theme={"theme":"github-dark"}
uv tool install "harbor[boxd] @ git+https://github.com/MichielMAnalytics/harbor.git@add-boxd-environment"
```
