The research demo generates emotion-aware motion in short windows, then reshapes and stabilizes it for a Unitree G1 humanoid.
SocialHumanoid is a research system, not a product people can buy. The team from Shanghai Jiao Tong University and ZTE Corporation reports that it can turn response speech and a chosen emotion into continuous full-body movement on a physical robot.
The process starts in human motion space. A speech signal, emotion setting, and a short history of earlier movement enter the generator. It produces each 128-frame motion window in one forward pass, then reuses motion history so the next window connects smoothly instead of starting from scratch.
That output cannot run directly on a robot. The authors’ online General Motion Retargeting step converts human-body motion into positions the robot’s different joints and proportions can support. Think of it as adapting a person-sized animation to a machine with a different skeleton and limited joint ranges.
A whole-body controller then treats the retargeted motion as a reference, not as raw motor commands. In the reported setup, the controller runs at 50 hertz while a lower-level thread sends joint commands at 500 hertz. The system was demonstrated on a 29-degree-of-freedom Unitree G1.
The authors also report AffectMoCap, a four-hour dataset with synchronized speech, full-body and hand motion, and eight emotion labels from two professional actors. They say physical experiments showed continuous affect-conditioned behavior and stable execution during a speech sequence longer than 10 minutes.
That is a useful systems demonstration: generation, robot adaptation, and balance control are separate layers. The next practical test is whether the same pipeline remains reliable during interruptions, unusual robot body shapes, disturbances, or infeasible motions.
