The new standard of Physical AI eval
We measure whether your model is getting better: your policy on our robots. Video and a score for every run, back the same day.
Every training run ends with the same question
Is this checkpoint better than the last one? On real robots that question has no cheap answer.
Someone has to put the world back after every try. Operators on shift, robots to keep alive. It still buys tens of rollouts a day, not thousands.
Binary success rate throws away most of what the rollout showed. The difference you care about ends up smaller than the noise.
Nothing stays still. Lighting, placement, wear, the operator. Two checkpoints run a week apart were never compared under the same conditions.
There is no single "better". Change the robot, the simulator or the metric and the winner changes. One rig and one number cannot settle it.
How teams solve it with us
Checkpoint in, rollouts out. The lab, the operators and the resets are ours. An infrastructure problem becomes an API call.
Many embodiments and simulators, same tasks, scoring fixed before anyone sees a result. Both checkpoints meet the same conditions, and nobody picks the metric that wins.
More signal from each rollout. Time-to-milestone scoring, so a call that needed hundreds of runs takes tens: the ~30x trial reduction in the PhAIL paper.
What you send, what comes back
A served endpoint, or the weights. Keep them on your own servers if you prefer. Plenty of teams do.
Your policy, blind, against a maintained baseline or your own previous checkpoint. Same tasks, same scenes, scoring fixed up front.
Video for every run, the time-to-milestone scores, and the comparison itself, back the same day.
Positronic lets us evaluate checkpoints continuously as we train, so we can course-correct our research quickly. Day-to-day evaluation doesn’t compete with our robot fleet or operators for time, which keeps them focused on data collection and real-world deployment.
Bring us the checkpoint you cannot score
Leave an address and a line about what you are training, or take half an hour now. Either way we work out what is worth measuring on your setup, and what a first round would tell you.