Two arXiv papers propose different fixes for the same problem: a robot’s planned movement is not always the movement its body achieves.
Humanoid robots must predict what happens next, then turn that forecast into stable movement. The hard part is that a high-level command is usually a motion reference—not direct motor torque—and a low-level controller may alter it for balance, contact, or body dynamics.
Being-M0.7 uses prediction mainly as transferable context. Its authors describe three stages: pretraining on more than 10,000 hours of human video and motion data, adapting that model to robot data, and training an action expert to produce whole-body commands. The expert combines predicted future visual information with current camera images and body-state measurements.
That approach aims to reduce dependence on scarce robot demonstrations. On the authors’ tests, Being-M0.7 recorded 128 successes in 180 simulated trials and 13 in 15 real-world Unitree G1 trials. Those were the paper’s comparisons, not a common leaderboard shared with HWAM.
HWAM targets a different failure point: the gap between a command and the motion actually realized. The LimX Dynamics paper makes the robot’s post-execution body state an explicit prediction target alongside the reference action. Its policy path generates both; its forward-dynamics path predicts future images from actions and realized states; its inverse-dynamics path works backward from visual changes to body motion.
That design is grounded in measured error. The authors report nonzero command-to-state differences on both the LimX OLI humanoid and ALOHA platform, including 4.420 degrees of direct error on OLI.
Both systems remain research demonstrations, evaluated on specific robots and tasks. The practical next test is matched, longer-horizon evaluation: same hardware, same tasks, and enough physical variation to show whether prediction improves execution beyond controlled trials.
