Robo Use

Run a suite

robouse run-many runs one harness on a set of tasks, several at a time. It is Robo Use's own local runner, the same one as robouse run.

A suite#

A suite is a subfolder of the bundled tasks. This runs the first 4 Meta-World reference solutions, 4 at a time:

$ robouse run-many \
    --tasks "$(robouse tasks --path)/metaworld" \
    --harness oracle \
    --out runs/mw-oracle \
    --concurrency 4 \
    --limit 4
metaworld-bin-picking                    reward=1 outcome=success_reached steps=131 9.0s
metaworld-assembly                       reward=1 outcome=success_reached steps=93 9.1s
metaworld-basketball                     reward=1 outcome=success_reached steps=91 9.1s
metaworld-box-close                      reward=1 outcome=success_reached steps=105 9.2s
oracle/-: 4/4 solved

The suite folders are the subfolders of robouse tasks --path; Suites says what each needs. Without --tasks, run-many runs all bundled tasks.

Pick tasks#

# ids containing "drawer"
robouse run-many \
  --filter drawer \
  --harness oracle \
  --out runs/drawer

# one task id per line
robouse run-many \
  --task-list picks.txt \
  --harness noop \
  --out runs/picks

# any folder of task folders
robouse run-many \
  --tasks mytasks \
  --harness oracle \
  --out runs/mine

Stop and resume#

Each finished trial is appended to <out>/run_log.jsonl. After an interruption, run the same command with --resume: tasks that already have a finished trial in --out are skipped, and the rest run.

With an agent#

robouse run-many \
  --tasks "$(robouse tasks --path)/metaworld" \
  --harness codex \
  --model gpt-6-astra \
  --out runs/mw-codex \
  --concurrency 4

Each trial is one agent session, so cost and rate limits grow with --concurrency. A trial that failed for an infrastructure reason has an EXC note on its line and exception_info in its result.json; rerunning with --resume retries it.

Run under BenchFlow#

robouse run and run-many do not use BenchFlow. They write BenchFlow's trial layout and ATIF trajectories, and the tasks are BenchFlow-native task.md folders, but the runner, the episode server and the agent all run on your machine.

To run tasks under BenchFlow's bench eval run instead, export them with benchflow/export.py. Each exported task runs in two Docker containers: the agent with only the robo client, and the simulator with no network. They share only the episode socket, and the verifier runs in the simulator container. The exporter is in the Robo Use repository, not in the PyPI package, so this needs a checkout:

$ python benchflow/export.py \
    --tasks tasks/metaworld \
    --filter reach,push \
    --out dist/robouse
...
7 task(s) -> /path/to/robouse/dist/robouse
$ cd dist/robouse && bench tasks check tasks/metaworld-reach
✓ metaworld-reach — valid (structural)
$ bench eval run --config robouse-oracle.yaml

bench eval run needs Docker with Compose 2.24.4 or newer.

To run published sets of exported tasks by name, without a checkout, see Datasets.