Open evaluations
Benchmarks, methods, and results anyone can test.
TESTABLE BY ANYONE
Infrastructure, evaluations, and research for
Physical AI that is capable, safe, and open.
01 / Our approach
Benchmarks, methods, and results anyone can test.
TESTABLE BY ANYONEMore capable systems. Stronger human control.
HUMAN CONTROLOpen code and research to build on, together.
BUILT IN THE OPEN02 / Open tools
Open infrastructure for Physical AI.
From world models to reproducible evaluation.
03 / Research
A VLA answers 'what should I do next?'. A world-action model answers 'what happens if I do this?' — and only the second question supports search, recovery, and calibrated refusal. We set out the latent dynamics objective, the rollout-horizon limit that governs it, and how a WAM composes with a VLA rather than replacing it.
A VLA is a vision-language model with the action space bolted onto its output head. We walk through the three design decisions that actually determine whether one works on a real arm — action tokenization, chunking, and the latency budget — and where each of them breaks.
04 / Community
Explore the code. Share your ideas.
Help shape what comes next.