Is your model a top model?
How to tell a better robot policy from a worse one, on real hardware.
Answering honestly is harder than it sounds.
Four methodological traps make most robotic model comparisons misleading.
What honest eval looks like.
Four principles. Each one a direct answer to a trap above.
OpenPI, DreamZero and MolmoAct, four DROID tasks each on one Franka arm. 2× speed.
Five policies, the same blind rounds.
Franka DROID rig, 21–25 Sep 2026, single-item tasks, 37–39 episodes per policy. Each point is the share of episodes that reached the stage; the band is the 95% Wilson interval. Hover for the counts.
Cosmos3 nano and GR00T N1.7, second by second.
Each band is the episodes holding that stage at that second, the whole cohort at every second, over the episode window in seconds. The purple band has placed the item; the dashed line is the mean. Hover for every stage at that second.
Cosmos3 nano 39 episodes
GR00T N1.7 39 episodes
Send a checkpoint. Find out.
positronic.ro ↗Vladimir Yakunin · Positronic Robotics · Robo House, Menlo Park