The new standard of Physical AI eval

We measure whether your model is getting better: your policy on our robots. Video and a score for every run, back the same day.

Design partner
PhAILOur physical AI leaderboard
Founding partner of PhAIL
Problem

Every training run ends with the same question

Is this checkpoint better than the last one? On real robots that question has no cheap answer.

Someone has to put the world back after every try. Operators on shift, robots to keep alive. It still buys tens of rollouts a day, not thousands.

Binary success rate throws away most of what the rollout showed. The difference you care about ends up smaller than the noise.

Nothing stays still. Lighting, placement, wear, the operator. Two checkpoints run a week apart were never compared under the same conditions.

There is no single "better". Change the robot, the simulator or the metric and the winner changes. One rig and one number cannot settle it.

Method

How teams solve it with us

Checkpoint in, rollouts out. The lab, the operators and the resets are ours. An infrastructure problem becomes an API call.

Many embodiments and simulators, same tasks, scoring fixed before anyone sees a result. Both checkpoints meet the same conditions, and nobody picks the metric that wins.

More signal from each rollout. Time-to-milestone scoring, so a call that needed hundreds of runs takes tens: the ~30x trial reduction in the PhAIL paper.

Protocol

What you send, what comes back

You send

A served endpoint, or the weights. Keep them on your own servers if you prefer. Plenty of teams do.

We run

Your policy, blind, against a maintained baseline or your own previous checkpoint. Same tasks, same scenes, scoring fixed up front.

You get

Video for every run, the time-to-milestone scores, and the comparison itself, back the same day.

Positronic lets us evaluate checkpoints continuously as we train, so we can course-correct our research quickly. Day-to-day evaluation doesn’t compete with our robot fleet or operators for time, which keeps them focused on data collection and real-world deployment.
Andy Chen, Head of Robotics, Runway
Talk to us

Bring us the checkpoint you cannot score

Leave an address and a line about what you are training, or take half an hour now. Either way we work out what is worth measuring on your setup, and what a first round would tell you.