Robo Use

Harnesses

A harness is the agent program that acts as the robot's policy; the model is what it calls. Any program that can run shell commands can drive the robot, because the only interface is the robo command. Pick one with --harness, and a model with --model (for molmoact2 the policy server decides).

--harnessRunsDefault modelNeeds
oraclethe task's oracle/solve.shnonenothing
nooprobo done at once (negative control)nonenothing
claude-codeClaude Code, claude -pclaude-sonnet-5claude on PATH; ANTHROPIC_API_KEY or CLAUDE_CODE_OAUTH_TOKEN
codexCodex, codex exec --jsongpt-6-astracodex on PATH; OPENAI_API_KEY or a codex login
claude-code-glmClaude Code through a local proxy to Basetenzai-org/GLM-5.3claude on PATH, BASETEN_API_KEY
codex-glmCodex with Baseten as the providerzai-org/GLM-5.3codex on PATH, BASETEN_API_KEY
mini-swe-agent-glmmini-swe-agent, minizai-org/GLM-5.3mini on PATH, BASETEN_API_KEY
molmoact2MolmoAct2 client, see belowallenai/MolmoAct2-LIBERO-LeRobotROBOUSE_MOLMOACT2_URL

Credentials come from your environment and are passed to the harness only:

  • claude-code: ANTHROPIC_API_KEY if set, else CLAUDE_CODE_OAUTH_TOKEN (a Claude subscription token from claude setup-token).
  • codex: OPENAI_API_KEY if set (the runner logs in with it inside Robo Use's own Codex folder), else a copy of your ~/.codex/auth.json from codex login. Your own file is only read, never written.
  • the GLM harnesses: BASETEN_API_KEY.

Without one, the trial stops before it starts, with an exception in result.json such as no Claude credentials: set ANTHROPIC_API_KEY, or CLAUDE_CODE_OAUTH_TOKEN (create one with claude setup-token). An API error during the run is not an exception: the harness exits and the episode scores 0.

What the harness gets#

  • A fresh temporary workspace as its working directory, with instruction.md and observations/.
  • robo first on its PATH, and ROBOUSE_SOCKET set.
  • The prompt: one short paragraph ("You are controlling a simulated robot... call robo done once") followed by the task's instruction.
  • A wall-clock limit: the task's agent.timeout_sec (900 s if unset), or --timeout. On timeout the process group is killed and the episode ends as agent_timeout.

Claude Code runs with its own configuration folder, ~/.cache/robouse/claude-home, instead of yours (no user settings, CLAUDE.md, MCP servers or plugins), without permission prompts, with at most 400 turns, and with WebFetch, WebSearch and Task disabled. Codex runs with CODEX_HOME=~/.cache/robouse/codex-home, which holds its login and a config.toml that sets medium reasoning effort. Both folders are shared by all trials.

Isolation#

On macOS, model harnesses run under sandbox-exec with a profile that denies reading and writing the paths in ROBOUSE_SANDBOX_DENY. Nothing is protected unless you list it, so list your virtual environment (the simulator code and the bundled reference solutions are in its site-packages), your --out folder (the episode record is written there) and any credential files:

export ROBOUSE_SANDBOX_DENY="$PWD/.venv:$PWD/runs:$HOME/.aws"

Paths are separated by colons, ~ is expanded and symlinks are resolved. ROBOUSE_SANDBOX=0 turns the sandbox off; config.json records whether it was on. It is a file fence, not a full security boundary, and there is none on Linux. For container isolation, export the tasks to BenchFlow (Run a suite).

MolmoAct2#

molmoact2 runs a vision-language-action model instead of an LLM agent. Its client drives robo like any harness: it sends the camera images, the robot state and the task's language instruction to a MolmoAct2 policy server and applies the action chunk it gets back. The server, robouse/molmoact2/server.py, runs on an NVIDIA GPU inside a LeRobot install and takes the same arguments as lerobot-eval. Point ROBOUSE_MOLMOACT2_URL at it. It targets the LIBERO suite; in the v1 runs it solved 20 of 20 LIBERO tasks.