A UAE University research system uses a prebuilt gesture catalog and state-aware replanning on a physical G1 robot.

Researchers at United Arab Emirates University propose TRABot, a research framework evaluated on a physical G1 humanoid robot. It coordinates streaming speech, a digital face, and body gestures in one real-time interaction loop.

The system first builds a catalog of “motion atoms”—short, approved gesture units with defined meanings. The offline process gives each atom a communicative label, such as emphasis or greeting, and checks whether the G1 can execute its trajectory within tested movement and contact limits, according to the researchers’ paper on arXiv.

During a conversation, the dialogue agent produces the spoken reply, an ordered list of communicative functions, and an estimated reply duration. The motion planner then selects the longest matching sequence that fits the available time, including transitions and a return to a neutral pose.

Facial animation runs from the same streaming audio used for speech. Meanwhile, the body planner assembles gestures from the prebuilt catalog, so the robot’s movement is coordinated with the reply but not generated freely from scratch.

The key engineering feature is state-aware execution. TRABot commits only to the next motion atom, observes the robot’s resulting state, and replans the remaining gestures. That design is meant to respond to differences between planned and reached poses; the reported evaluation does not establish performance during unexpected contact or interruptions.

In a 12-turn evaluation, the authors report 84.72% average speech-body overlap and completion without failure on 100% of turns. A 20-person user study ranked the complete system highest among the compared conditions.

This remains author-reported research, not a product available for purchase or a demonstrated public deployment. Its practical range depends on the existing gesture catalog, making broader interaction and safety testing the next important steps.