World-action models: planning in a learned latent
A VLA answers 'what should I do next?'. A world-action model answers 'what happens if I do this?' — and only the second question supports search, recovery, and calibrated refusal. We set out the latent dynamics objective, the rollout-horizon limit that governs it, and how a WAM composes with a VLA rather than replacing it.
No. 028 minWAM, World Models, PlanningVision-language-action models, from the inside out
A VLA is a vision-language model with the action space bolted onto its output head. We walk through the three design decisions that actually determine whether one works on a real arm — action tokenization, chunking, and the latency budget — and where each of them breaks.
No. 018 minVLA, Manipulation, Foundation Models