Robo Use

Datasets

A dataset is a named, versioned set of Robo Use tasks that BenchFlow's bench runs by name, for example benchflow/robouse-core@0.1. The tasks are exported as native BenchFlow task packages into a hub repository, and a registry file pins each version to one commit of that repository together with a content digest of every task. A published version is not changed: a new export gets a new version.

The hub is robouse.ai/hub. It lists every dataset with its task count, robots and simulators, and the command to run it.

Names#

A dataset name is <org>/<name>, where the org is whoever designed the benchmark's tasks. Benchmarks adapted from other groups sit under that group's GitHub organisation; these are BenchFlow's adaptations, not published or reviewed by the original authors. Tasks BenchFlow designed stay under benchflow/, even when they run on someone else's simulator or robot models.

DatasetTasksOriginal benchmark
benchflow/robouse-core@0.1100a sample of all Robo Use suites
farama-foundation/metaworld@0.150Meta-World MT50 (Farama Foundation, MIT)
farama-foundation/gymnasium-robotics@0.112Gymnasium-Robotics (Farama Foundation, MIT)
benchflow/robouse-tabletop@0.1119Robo Use's own tabletop scenes
lifelong-robot-learning/libero@0.240LIBERO: Spatial, Object, Goal and Long (Lifelong Robot Learning, MIT); 0.1 had 20 of them
lifelong-robot-learning/libero-90@0.190LIBERO-90 (Lifelong Robot Learning, MIT)
arise-initiative/robosuite@0.121robosuite (ARISE Initiative, MIT)
benchflow/robouse-menagerie@0.115Robo Use tasks on MuJoCo Menagerie robots
brave-eai/dexjoco@0.114DexJoCo (brave-eai, MIT)
benchflow/robouse-drone@0.110Robo Use drone tasks
benchflow/robouse-harm@0.180Robo Use harm scenarios (first published as benchflow/robouse-roboharm@0.1)
benchflow/quadruped@0.111Robo Use tasks for Unitree Go2 and Go1, Spot, ANYmal C
benchflow/humanoid@0.19Robo Use tasks for Unitree G1 and H1, Apptronik Apollo, Booster T1
benchflow/mobile-manip@0.110Robo Use tasks for Stretch 3, TIAGo, Google Robot
benchflow/dexhand@0.17Robo Use tasks for the Shadow Hand and LEAP Hand
benchflow/crazyflie@0.17Robo Use tasks for Crazyflie 2 quadrotors
benchflow/driving@0.17Robo Use traffic tasks on MetaDrive
farama-foundation/franka-kitchen@0.110Franka Kitchen in Gymnasium-Robotics (Farama Foundation, MIT)
farama-foundation/adroit-hand@0.112Adroit hand in Gymnasium-Robotics (Farama Foundation, MIT)
google-deepmind/dm-control@0.117dm_control Control Suite and manipulation (Google DeepMind, Apache-2.0)
myohub/myosuite@0.115MyoSuite (MyoHub, Apache-2.0)
carlosferrazza/humanoid-bench@0.17HumanoidBench (Carmelo Sferrazza, MIT)
mani-skill/maniskill@0.112ManiSkill3 on its CPU backend (ManiSkill, Apache-2.0)
robocasa/robocasa@0.123RoboCasa (RoboCasa, MIT)
benchflow/robouse-noop-control@0.1100negative control for benchflow/robouse-core

The earlier names (robouse-core@0.1, robouse-metaworld@0.1, benchflow/robouse-roboharm@0.1 and so on) still resolve: each is an alias in the registry with exactly the same pinned tasks and digests. lifelong-robot-learning/libero@0.1 (20 tasks) stays published next to 0.2.

Run a dataset#

$ gh api repos/benchflow-ai/robohub/contents/registry.json -H 'Accept: application/vnd.github.raw' > registry.json
$ bench eval run \
    -d robouse-core@0.1 \
    --registry registry.json \
    --agent oracle \
    --concurrency 4
✓ robouse-core@0.1: 100 tasks, digests verified (9e672aa1e7f0)
...
[PASS] arc-gravity (reward=1.00, tools=0)
[PASS] metaworld-reach (reward=1.00, tools=0)
...
Job complete: 100/100 (100.0%), mean_reward=1.00, errors=0, idle_timeouts=0, time=97.8min
✓ Score: 100/100 (100.0%), mean reward 1.00, errors=0

bench reads the registry, clones the hub repository at the pinned commit, recomputes every task's digest and stops if one does not match. Then each task runs in two Docker containers: the agent with only the robo client, and the simulator with no network, which also runs the verifier. The reward is 1 if the episode server judged the task solved, else 0.

  • --include TASK (repeatable) runs a subset:
$ bench eval run \
    -d farama-foundation/metaworld@0.1 \
    --registry https://robouse.ai/hub/registry.json \
    --agent oracle \
    --include metaworld-reach \
    --include metaworld-drawer-open
✓ farama-foundation/metaworld@0.1: 50 tasks, digests verified (9e672aa1e7f0)
...
[PASS] metaworld-reach (reward=1.00, tools=0)
[PASS] metaworld-drawer-open (reward=1.00, tools=0)
Job complete: 2/2 (100.0%), mean_reward=1.00, errors=0, idle_timeouts=0, time=1.0min
  • --agent oracle runs each task's reference solution. Other BenchFlow agents take --agent NAME --model ID; no model-driven run on the hub has been checked yet.
  • --registry takes a URL or a file. The registry is in the hub repository and mirrored at https://robouse.ai/hub/registry.json, which bench can read without a token (the no-op run below used it). Running a dataset needs read access to github.com/benchflow-ai/robohub.
  • Run bench from a folder under your home folder: on a Mac with Colima, Docker can only mount paths there.

benchflow/robouse-noop-control@0.1 is a negative control: the same tasks as benchflow/robouse-core, but each reference solution only looks at the robot and exits. Run with --agent oracle, every task must score 0.

The simulator images are built from the hub repository the first time a dataset needs them (the largest, for LIBERO, is about 4 GB), so the first run takes longer. Scenes render on the CPU inside the container, so some tasks are much slower than a local robouse run: a DexJoCo reference solution took about 8 minutes.