Simulation-trained navigation policy guides a Unitree G1 through narrow passages, overhead obstacles and floor clutter—but a separate tracker still turns those predictions into motion.
Researchers from the University of California, Berkeley, Peking University, Tsinghua University, the University of Hong Kong and Princeton University report TANGO as a research demonstration, not a deployed product. The system combines a natural-language instruction, front- and downward-facing RGB images, and proprioceptive data—the robot’s measured body state—to predict 29-degree-of-freedom joint-space actions.
That is the key engineering choice. Rather than outputting only a flat waypoint, TANGO predicts whole-body motion references that can account for the robot’s articulated geometry. Its simulation-only Plan–Edit–Track pipeline first plans a route, edits the motion around obstacles, then checks the result with the SONIC motion tracker. The process generated 64,633 trajectories and required 211 RTX PRO 6000 GPU-hours for motion generation and rendering.
From prediction to physical movement
On a Unitree G1, the authors report zero-shot transfer to real indoor scenes without real-world navigation training. Qualitative demonstrations included side-stepping through a narrow passage, bending below overhead obstacles and stepping over floor obstacles. Separately, the quantitative comparison covered three settings—short-horizon navigation, roughly 30-metre long-horizon navigation and cluttered 3D navigation—with three scenes and five trials per method in each setting.
TANGO does not replace the walking controller. The VLA model and action expert run on an RTX PRO 6000 server, while SONIC runs on an onboard Jetson Orin NX. The robot sends RGB observations and proprioception with approximately 20 milliseconds of latency; inference runs every 0.5 seconds, producing 15 actions at 30 Hz. Those chunks are resampled to 50 Hz, and SONIC closes the low-level loop at approximately 200 Hz.
The practical next test is broader physical validation: the authors identify RGB-only perception and tracker capability as constraints in visually ambiguous, low-light or more complex environments.