Lift the box onto the shelf with both arms
h1-box-to-shelfhumanoidUnitree G1Manipulationmedium
Instruction
You control a Unitree H1 humanoid (1.8 m; MuJoCo Menagerie unitree_h1, BSD-3-Clause). It has no hands: each 4-joint arm ends in a round fist (6.6 cm ball). The robot is mounted on a support stand (its pelvis is welded to the stand): it does not balance or walk, its legs hold a standing pose and only the arms move. A counter stands in front of it. World frame in metres: +x points forward from the robot, +y to the robot's left, +z up; the floor is at z = 0 and the origin is on the floor below the robot's pelvis. A cardboard box (16 x 22 x 16 cm, 0.5 kg) stands on the left of the counter; a blue shelf block, 10 cm high, stands on the right. The fists cannot grasp: lift the box by squeezing it between both fists or forearms. All positions and sizes are listed under scene and objects in robo observe. Knocking an object onto the floor is a safety event and fails the task. Each hand's control point is the centre of its fist. The arms have four joints, so the forearm's angle follows from where the fist is; the forearms also touch things. Objects are pushed, pressed or squeezed between the two fists or forearms.
Task
Move the box onto the top of the blue shelf block.
Success: the box sits flat (tilt at most 8 degrees) with its centre over the shelf top, resting on it, and the robot does not touch it; and no safety event. Judged by the episode server from the simulated state after you call robo done and the robot has held still for 10 steps.
Controls. robo act LDX LDY LDZ RDX RDY RDZ, each in [-1, 1]. LDX LDY LDZ and RDX RDY RDZ move the goal point of the left and right hand by up to 2 cm per step along world x, y, z; an arm controller (inverse kinematics and joint servos) moves the hand after its goal, stops when the hand is blocked and keeps the goal within reach. Zero keeps the goal where it is, so an all-zero action holds the robot still. One step is 40 ms; --repeat N applies an action N times. Skills (robo info lists them with their arguments): reach SIDE X Y Z [TOL] (move one hand's control point in a straight line to a point; reports whether it arrived, was blocked or hit the edge of its reach, and what the hand touches), reach_both XL YL ZL XR YR ZR [TOL] (both hands at once, for bimanual holds), home (both hands back to their start points), wait [STEPS]. Skills do not plan around obstacles.
Observation. robo observe reports robot.hands.left / robot.hands.right (pos: the control point, goal: its current goal, touching: what the hand touches); objects (each free object's centre pos, tilt_deg from upright, touched_by_robot, and its size); scene (the table and every fixture and target zone, with positions); progress where the task tracks something; and safety_events. Cameras: front (default), top, side and head (fixed at the robot's head, looking at the table).
The step budget is 700 steps (28 s of robot time).
How the robot is controlled and scored
You are controlling a simulated robot. Read the task below, then solve it by running the robo command in your shell (start with robo info and robo observe). Keep going until the task is done, then call robo done once. Do not stop to ask questions; there is no human to answer.
How to control the robot
You are the robot's policy. You act only through the robo command in your shell. There is no other way to move the robot, and you cannot read or change the simulator, the scoring, or other files to succeed; the episode server judges the final physical state itself.
robo info # the robot, its sensors, action groups, skills and step budget
robo observe # robot and scene state as numbers
robo observe --image [--camera C] # also saves a camera image and prints its path (open it to look)
robo act V1 V2 ... [--repeat N] # one low-level action (the action groups under Controls), applied N times (N <= 50)
robo skill NAME ARG ... # run a skill listed by `robo info`; it runs until it finishes and reports the result
robo done "short summary" # end the episode and ask for scoring
robo give-up "reason" # end the episode without claiming success- Positions are in metres in the world frame (+z up); angles are in degrees unless a field says otherwise.
- The episode has a fixed step budget (see
robo info); every simulated control step counts, including the steps a skill runs. - Skills are ordinary controllers: they can fail, stop early or be blocked by the scene. Read what they report and re-observe.
- Success is judged about 10 steps after you call
robo done, with the robot holding still (each action group's hold value: zero for velocity and delta commands, full brake for a car), so the goal must still be true when the robot stops. - Call
robo doneexactly once when finished.
Run this task
ROBOUSE_ORACLE_TOKEN=$(openssl rand -hex 16) \
bench eval run \
-d benchflow/humanoid@0.2 \
--registry https://robouse.ai/hub/registry.json \
--agent oracle \
--include h1-box-to-shelfPinned to robohub commit e472b1a1e041. The verifier and the reference solution are not published.