The ShanghaiTech framework links memory, visual cues, speech, facial motion, and gestures in a physical robot demo.
ARISE is a research framework that helps Sophia decide not only what to say, but when to engage and how to move while speaking. A paper from ShanghaiTech University describes the system running on the Sophia humanoid robot developed by Hanson Robotics.
The system takes conversational audio and real-time camera images. Its multimodal agent—software that combines several types of input—can use long-term memory, online search, and proactive interaction planning to create a spoken response and a symbolic gesture plan.
That plan does not directly control every motor. Instead, ARISE proposes several candidate gesture sequences from a library of validated robot postures. A review stage checks whether each choice fits the meaning, social setting, timing, and physical limits before execution, according to the paper published on arXiv.
A streaming coordinator then synchronizes the robot’s speech with facial animation. It uses that audio-and-face timing as a clock for scheduling upper-body gestures, sending the resulting commands to Sophia’s actuators. The reported setup used an RGB camera, a microphone, and a host computer with an NVIDIA RTX 3090 graphics processor.
The paper reports two user studies with 20 participants each, plus a latency test using 40 short dialogue segments. However, the supplied results do not include the numerical table values, so the size of any reported improvement cannot be judged here.
This remains a research demonstration, not a consumer product or announced commercial deployment. The authors say the cascaded system still cannot reliably deliver sub-second conversation, and Sophia’s mechanical face limits subtle expressions. The next practical test is whether ARISE remains useful during longer, unscripted interactions rather than prepared study scenarios.
