Robo Use

Fit the square nut over the square peg

robosuite-nut-square-pandarobosuiteFranka PandaManipulationhard

Reference solution, 322 of 500 steps.

Instruction

A Franka Emika Panda (7-DoF) with its parallel-jaw gripper works at a table; its base is at the -x end, facing +x. The scene is robosuite's NutAssemblySquare environment (MuJoCo).

Goal: Pick up the square nut by its handle and fit it over the square peg so that it slides down around the peg to the table. robosuite's NutAssembly check: the nut's centre is within 3 cm of the peg's axis in x and y, below table_height + 0.05, and the gripper is at least about 4 cm away from it. The square nut's hole only fits the square peg when the nut is turned to match it (the peg's faces are aligned with the x and y axes).

Call robo done when finished; the check is made after the robot then holds still for 10 steps, so the result must last.

This robot

  • The action has 7 numbers, not 4: robo act DX DY DZ DROLL DPITCH DYAW GRIP, each in [-1, 1]: robosuite's OSC_POSE end-effector command plus the gripper, 20 steps per second, in the world frame. DX/DY/DZ = 1.0 asks for a 5 cm move (held at 1.0 the hand travels about 1.1 cm per step); DROLL/DPITCH/DYAW rotate the hand about the world x/y/z axes (1.0 asks for 0.5 rad; leave them at 0 unless you need to turn the hand, e.g. DYAW to line the fingers up with an object); GRIP +1 closes, -1 opens, 0 keeps the fingers as they are (they move over about 5 steps).
  • robo move-to X Y Z [--grip G] and robo grip G are available; they move only the position and keep the hand's orientation (it can drift slowly under load; correct it with DROLL/DPITCH/DYAW).
  • World frame, metres: +x points away from the robot's base, +y to the robot's left, +z up.
  • robo observe fields: hand_pos (the gripper's grasp point between the fingertips, metres), hand_quat (hand orientation quaternion x, y, z, w; about [1, 0, 0, 0] when the gripper points straight down), gripper_open (distance between the two finger pads in metres: largest when open, smallest when closed on nothing, in between when holding something), gripper_yaw_deg (direction of the line through the two fingertips in the table plane, degrees from +x, in [-90, 90); the fingers close along this line); squarenut_pos, squarenut_quat, squarenut_yaw_deg (the nut's centre, i.e. its hole; the nut's local x axis points from the hole toward the handle), squarenut_handle_pos (the flat handle tab, about 5.4 cm from the centre); square_peg_pos, round_peg_pos (the two pegs; use the square one); table_height. Every <object>_quat (x, y, z, w) comes with <object>_yaw_deg, the object's heading about the vertical axis (its local x axis, degrees from +x).
  • robo observe --image saves the agentview camera; add --camera robot0_eye_in_hand for the wrist camera. Images are 320x320.
  • The step budget is 500 steps.
  • Holding the nut by its handle, the hole hangs about 5 cm from the hand, so align the nut's centre (not the hand) over the peg. The pegs are near the far edge of the robot's reach: reaching past them is hard, so turn the nut so its handle points toward the robot or sideways. Lower it slowly; if it jams on the peg top, lift and re-centre.
How the robot is controlled and scored

You are controlling a simulated robot. Read the task below, then solve it by running the robo command in your shell (start with robo info and robo observe). Keep going until the task is done, then call robo done once. Do not stop to ask questions; there is no human to answer.

How to control the robot

You are the robot's policy. You act only through the robo command in your shell. There is no other way to move the robot, and you cannot read or change the simulator, the scoring, or other files to succeed; the episode server judges the final physical state itself.

robo info                         # action space, available skills, step budget
robo observe                      # robot and object state as numbers
robo observe --image              # also saves a camera image and prints its path (open it to look)
robo act DX DY DZ GRIP [--repeat N]   # low-level action, applied N times (N <= 50)
robo move-to X Y Z [--grip G]     # skill: move the gripper toward a point (if enabled for this task)
robo grip G [--steps N]           # skill: hold position and set the gripper (+1 close, -1 open)
robo done "short summary"         # end the episode and ask for scoring
robo give-up "reason"             # end the episode without claiming success
  • Positions are in metres in the world frame (x, y on the table plane, z up).
  • The episode has a fixed step budget (see robo info); every simulated step counts, including skills.
  • Unless the task says otherwise, success is judged about 10 steps after you call robo done, with the robot holding still, so the goal must still be true when the robot stops.
  • Work in small steps and re-observe after each motion. Call robo done exactly once when finished.

Run this task

ROBOUSE_ORACLE_TOKEN=$(openssl rand -hex 16) \
  bench eval run \
    -d arise-initiative/robosuite@0.2 \
    --registry https://robouse.ai/hub/registry.json \
    --agent oracle \
    --include robosuite-nut-square-panda

Pinned to robohub commit e472b1a1e041. The verifier and the reference solution are not published.

Trial
GPT-6 Astra · CodexSolved217 / 50081 sTrial Compare
Kimi K3 · Claude CodeSolved359 / 50012 minTrial Compare
GLM-5.3 · mini-swe-agentNot solved498 / 50015 minTrial Compare
GLM-5.3 · Claude CodeNot solved459 / 50015 minTrial Compare
GLM-5.3 · CodexNot solved464 / 5003 minTrial Compare
Kimi K3 · mini-swe-agentInfra failure147 / 5002.2 hTrial Compare

These trials ran on version 0.1, now deprecated (superseded by 0.2).