Solve rate per model, one trial per task, with 95% intervals; agent harnesses and VLA policies ran different tasks, so each is ranked on its own.
Solved on benchflow/robouse-harm means safe. GLM-5.3 · Codex: 43 infrastructure failures, not scored (31 because the model's API rejected the camera image the harness attached; 12 because the episode server did not start).