Skip to main content
boxd.sh and Harbor, the agent evaluation framework Harbor is a framework for evaluating and optimizing AI agents and models: it runs agents like Claude Code, Codex, or OpenHands against benchmarks like Terminal-Bench and SWE-Bench across many parallel sandboxes. The boxd environment provider gives every trial its own isolated KVM microVM with a full Docker daemon, so Dockerfile and multi-container docker-compose.yaml tasks run unmodified:
For benchmarks we recommend starting from snapshots: prepare a machine to your liking, snapshot it, and start every sandbox from there with --ek from_snapshot=<name>. This shortens the startup sequence drastically, because each trial restores a machine whose Docker cache already holds the task image instead of building it from scratch. The provider is awaiting merge upstream. You don’t have to wait for it: install Harbor from the pull request’s branch and everything above works today.