An arXiv paper reports simulation and real-robot tests, including 95.86% success in physical simulation.
A research team from Nanjing University proposes BeyondRetarget, a system that converts ordinary single-camera color video directly into motion for humanoid robots. Instead of first rebuilding a human’s movements and then translating them, the method learns a robot-oriented motion pattern from the video.
The paper says robot-specific decoders then turn that shared pattern into motion for different machines. The team trained the main model on the Unitree G1 and adapted it to seven other humanoids, including Unitree R1, Fourier robots, Unitree H1, Atlas, and Tienkung. The paper reports that adapting another robot took about one hour on one Nvidia RTX 4090.
BeyondRetarget also adds contact-aware refinement. It predicts when each foot touches the ground, corrects the robot’s height, and uses inverse kinematics—a calculation of joint positions—to keep supporting feet from sliding. This targets a common failure in video-driven robots: a pose can look right while being physically unstable.
In physical simulation, the authors report 347 successful executions out of 362, or 95.86%, with an executed motion error of 35.04 millimeters. Their real-time comparison reports 192.8 milliseconds of median latency, 50.02 frames per second at maximum throughput, and 7.03 gigabytes of peak graphics memory on the same graphics processor.
The paper also reports tests on real humanoid robots, but it does not establish the full safety procedures or operating conditions. Its authors note remaining problems with large movements, moving cameras, complex terrain, and contact-heavy tasks.
This is a research result, not a product people can buy or a production deployment. The next practical test is whether independent teams can reproduce the results across varied environments and longer-running tasks.
