Harnesses
A harness is the agent program that acts as the robot's policy; the model is what it calls. Any program that can run shell commands can drive the robot, because the only interface is the robo command. Pick one with --harness, and a model with --model (for molmoact2 the policy server decides).
--harness | Runs | Default model | Needs |
|---|---|---|---|
oracle | the task's oracle/solve.sh | none | nothing |
noop | robo done at once (negative control) | none | nothing |
claude-code | Claude Code, claude -p | claude-sonnet-5 | claude on PATH; ANTHROPIC_API_KEY or CLAUDE_CODE_OAUTH_TOKEN |
codex | Codex, codex exec --json | gpt-6-astra | codex on PATH; OPENAI_API_KEY or a codex login |
claude-code-glm | Claude Code through a local proxy to Baseten | zai-org/GLM-5.3 | claude on PATH, BASETEN_API_KEY |
codex-glm | Codex with Baseten as the provider | zai-org/GLM-5.3 | codex on PATH, BASETEN_API_KEY |
mini-swe-agent-glm | mini-swe-agent, mini | zai-org/GLM-5.3 | mini on PATH, BASETEN_API_KEY |
molmoact2 | MolmoAct2 client, see below | allenai/MolmoAct2-LIBERO-LeRobot | ROBOUSE_MOLMOACT2_URL |
Credentials come from your environment and are passed to the harness only:
claude-code:ANTHROPIC_API_KEYif set, elseCLAUDE_CODE_OAUTH_TOKEN(a Claude subscription token fromclaude setup-token).codex:OPENAI_API_KEYif set (the runner logs in with it inside Robo Use's own Codex folder), else a copy of your~/.codex/auth.jsonfromcodex login. Your own file is only read, never written.- the GLM harnesses:
BASETEN_API_KEY.
Without one, the trial stops before it starts, with an exception in result.json such as no Claude credentials: set ANTHROPIC_API_KEY, or CLAUDE_CODE_OAUTH_TOKEN (create one with claude setup-token). An API error during the run is not an exception: the harness exits and the episode scores 0.
What the harness gets#
- A fresh temporary workspace as its working directory, with
instruction.mdandobservations/. robofirst on itsPATH, andROBOUSE_SOCKETset.- The prompt: one short paragraph ("You are controlling a simulated robot... call
robo doneonce") followed by the task's instruction. - A wall-clock limit: the task's
agent.timeout_sec(900 s if unset), or--timeout. On timeout the process group is killed and the episode ends asagent_timeout.
Claude Code runs with its own configuration folder, ~/.cache/robouse/claude-home, instead of yours (no user settings, CLAUDE.md, MCP servers or plugins), without permission prompts, with at most 400 turns, and with WebFetch, WebSearch and Task disabled. Codex runs with CODEX_HOME=~/.cache/robouse/codex-home, which holds its login and a config.toml that sets medium reasoning effort. Both folders are shared by all trials.
Isolation#
On macOS, model harnesses run under sandbox-exec with a profile that denies reading and writing the paths in ROBOUSE_SANDBOX_DENY. Nothing is protected unless you list it, so list your virtual environment (the simulator code and the bundled reference solutions are in its site-packages), your --out folder (the episode record is written there) and any credential files:
export ROBOUSE_SANDBOX_DENY="$PWD/.venv:$PWD/runs:$HOME/.aws"Paths are separated by colons, ~ is expanded and symlinks are resolved. ROBOUSE_SANDBOX=0 turns the sandbox off; config.json records whether it was on. It is a file fence, not a full security boundary, and there is none on Linux. For container isolation, export the tasks to BenchFlow (Run a suite).
MolmoAct2#
molmoact2 runs a vision-language-action model instead of an LLM agent. Its client drives robo like any harness: it sends the camera images, the robot state and the task's language instruction to a MolmoAct2 policy server and applies the action chunk it gets back. The server, robouse/molmoact2/server.py, runs on an NVIDIA GPU inside a LeRobot install and takes the same arguments as lerobot-eval. Point ROBOUSE_MOLMOACT2_URL at it. It targets the LIBERO suite; in the v1 runs it solved 20 of 20 LIBERO tasks.