reset / step (and optionally evaluate), so any environment speaking the
protocol can serve any trainer.
Miles integrates OpenEnv as an
agent-function integration: a Miles-side agent
function drives the agentic loop — reset(task_id), repeated steps, then
scoring the episode with the task’s own tests — against an unmodified OpenEnv
server, and the score becomes the sample’s reward through a custom reward
hook.
Try it
The maintained end-to-end recipe is Terminal-Bench-2 GRPO inexamples/experimental/openenv.
It gives every episode its own cloud sandbox, built from that task’s official
image so no resident infrastructure is left behind, on a choice of providers —
AgentENV (self-hosted, E2B-compatible),
Daytona, E2B, or
Modal. One shared Docker env server is supported as well,
for running without any sandbox platform.
Follow the
recipe README
for prompt-data preparation, environment options, launcher flags, and
operational notes.
