Robo Use

Rule inference from pictures: completing a pattern

hard-arc-img-quadrantsHardFloating gripperRule inferenceVisionhard

Reference solution, 781 of 1,450 steps.

Instruction

A two-finger parallel gripper hangs over a table (a floating gripper, no arm; it cannot rotate). x points to the right as seen from the front of the table, y away from the front (toward the back of the table), z up; the table top is at z = 0 and spans x from -0.35 to 0.35 and y from -0.30 to 0.30. The gripper's tool point can reach x in [-0.34, 0.34], y in [-0.27, 0.28], z in [0.012, 0.35]. The fingers close along the y axis; the fully open gap is 10 cm.

Goal: Coloured objects stand on a 4 x 4 grid of square cells marked on the table (cells 7 cm apart, centre to centre; row 0 is the back row, col 0 the left column). The cell positions are NOT given as data: find the grid in the camera image. A hidden rule transforms one arrangement into another. The rule is shown ONLY as pictures: at the start of the episode the files observations/example_1.png, observations/example_2.png are placed in your working directory. Each picture shows one example, BEFORE on the left and AFTER on the right, seen from straight above (the back of the table is at the top of the picture) in this same scene with the gripper hidden. Nothing about the examples is given as data. Work out the rule and apply it to the objects in front of you by moving them. Objects of the same colour and shape are interchangeable. Only the final arrangement on the grid is judged: every object in the grid area must rest on the table within 2.5 cm of a cell centre, and objects off the grid are ignored. Objects must be released at the end.

Spare blocks stand in a row in front of the grid; spare blocks that are not needed must stay off the grid.

What you can observe

robo observe does not give object or goal positions. It returns only hand_pos, gripper_open, constraint_violations and saves one image per camera (top; 320 x 320 pixels), printing the paths. Open the images to see the scene. hand_pos is the gripper's tool point (between the fingertips) in metres, gripper_open is 0 (closed) to 1 (10 cm gap) (there is no contact sensor that says what the fingers hold: a finger opening above zero after closing means something is between them, and the images show what), and constraint_violations lists broken rules (any entry means the task has failed).

There is only one camera, and robo info does NOT give its projection matrix (only its image size and field of view). You know where the gripper is from hand_pos, so you can relate image positions to world positions by moving the gripper and looking where it appears.

Gripper

The grip value is a finger position target, not a hold command: +1 = fully closed, -1 = fully open (10 cm gap), and values in between give a partial opening (finger gap = 10 cm * (1 - G) / 2, so 0 = half open, a 5 cm gap). There is no separate "hold" value: to keep holding an object, keep sending +1 (the fingers then squeeze it). robo move-to X Y Z --grip G applies G on every step of the move, starting with the first, so --grip 1 closes the fingers at the start of the move (they take about 15 steps to close fully); without --grip, the last grip value is kept. robo grip G --steps N holds the hand still for N steps while applying G. The same values apply to the GRIP argument of robo act.

Task rules

  • Skills are disabled on this task: robo move-to and robo grip return an error. Use robo act DX DY DZ GRIP only (DX/DY/DZ move the tool point by about 1 cm per unit per step; GRIP +1 closes, -1 opens, values in between give a partial opening: finger gap = 10 cm * (1 - GRIP) / 2).
  • robo act --repeat N is limited to N <= 10 on this task (larger values are cut to 10).
  • Step budget: 1450 simulated steps (one step = 20 ms). Wall-clock limit: 30 minutes.
  • Objects must be released (not touching the fingers) when the episode is judged, about 10 steps after robo done.
How the robot is controlled and scored

You are controlling a simulated robot. Read the task below, then solve it by running the robo command in your shell (start with robo info and robo observe). Keep going until the task is done, then call robo done once. Do not stop to ask questions; there is no human to answer.

How to control the robot

You are the robot's policy. You act only through the robo command in your shell. There is no other way to move the robot, and you cannot read or change the simulator, the scoring, or other files to succeed; the episode server judges the final physical state itself.

robo info                         # action space, available skills, step budget
robo observe                      # robot and object state as numbers
robo observe --image              # also saves a camera image and prints its path (open it to look)
robo act DX DY DZ GRIP [--repeat N]   # low-level action, applied N times (N <= 50)
robo move-to X Y Z [--grip G]     # skill: move the gripper toward a point (if enabled for this task)
robo grip G [--steps N]           # skill: hold position and set the gripper (+1 close, -1 open)
robo done "short summary"         # end the episode and ask for scoring
robo give-up "reason"             # end the episode without claiming success
  • Positions are in metres in the world frame (x, y on the table plane, z up).
  • The episode has a fixed step budget (see robo info); every simulated step counts, including skills.
  • Unless the task says otherwise, success is judged about 10 steps after you call robo done, with the robot holding still, so the goal must still be true when the robot stops.
  • Work in small steps and re-observe after each motion. Call robo done exactly once when finished.

Run this task

bench eval run \
  -d benchflow/robouse-tabletop@0.1 \
  --registry https://robouse.ai/hub/registry.json \
  --agent oracle \
  --include hard-arc-img-quadrants

Pinned to robohub commit 9e672aa1e7f0. The verifier and the reference solution are not published.

Trial
GPT-6 Astra · CodexSolved760 / 1,4503 minTrial Compare
Kimi K3 · Claude CodeSolved1,064 / 1,4508 minTrial Compare
Kimi K3 · mini-swe-agentNot solved1,450 / 1,45025 minTrial Compare
GLM-5.3 · mini-swe-agentNot solved535 / 1,45010 minTrial Compare
GLM-5.3 · Claude CodeNot solved485 / 1,45030 minTrial Compare