The research demo combines audio, word timing, and robot-specific motion before running it through a whole-body tracker.

ECHO-G is a research framework that generates full-body motion for a humanoid from speech audio and a transcript marked with word timings. The authors describe a physical demonstration on a Unitree G1; this is not a consumer robot people can buy.

Its core model, the Speech-Grounded Diffusion Transformer, handles the two inputs at different time scales. It matches frame-by-frame audio features to motion while using transcript words to guide gestures toward the spoken content. A generative model can produce different movements for the same sentence.

ECHO-G generates motion directly for the robot instead of creating human movement first and converting it afterward. A fixed whole-body motion tracker turns the generated joint-position references into actions.

For training and evaluation, the authors retargeted motion to the 29-degree-of-freedom Unitree G1. After filtering, their dataset contained 14,987 training clips and 3,242 validation clips, according to the paper published on arXiv.org.

On the paper’s comparison table, ECHO-G recorded a Fréchet Gesture Distance of 2.278 and end-to-end processing time of 5.96 milliseconds per output frame. Fréchet Gesture Distance is a score comparing generated motion patterns with reference motion; lower scores indicate closer matches in this study.

The team also reports a physical G1 trial that played speech while the robot executed generated references through the tracker. In a separate video-rating study, 45 participants rated the joint audio-and-text version highest overall: 3.49 out of five, versus 3.05 for pooled single-input versions.

The current workflow needs complete audio and timed transcripts, so it cannot yet generate gestures causally as speech arrives. Its next practical hurdle is streaming operation, along with more reliable gestures for explicit instructions.