A Berkeley research paper shows a Unitree G1 sorting, catching, pushing, and climbing—but its success rates come from simulation.
The arXiv paper from University of California, Berkeley proposes Workhorse, a research system that learns whole-body behavior from human demonstrations recorded without a robot. It has not been presented as a commercial product.
Workhorse separates the job into two parts. A visual planner watches the robot’s camera and predicts the poses of five body parts: the torso, both wrists, and both feet. A reinforcement-learning tracker then turns those target poses into joint movements while trying to keep the robot balanced.
The key design choice is that the human data stays in this five-part pose format. The system does not first convert the demonstration into robot-specific joint angles, a process called retargeting. That shared interface is meant to make the behavior easier to transfer between humanoid bodies.
On a real Unitree G1, the Berkeley team demonstrated three task-specific systems. The robot sorted boxes with its hands and a kick, caught a thrown box after a flight of about 0.56 seconds, and pushed, toppled, and climbed onto a 14-kilogram suitcase. In box sorting, it recovered after people pushed the robot or removed a box.
The numbers require a sharper distinction. In a simulated copy of the demonstration room, the full system completed box sorting in 77% of episodes without pushes and 64% with pushes. The paper reports no real-robot success rate.
The second humanoid, Unitree’s H2, also appears only in simulation. Retrained from the same demonstrations, it completed 83% of 100 simulated box-sorting episodes without pushes, versus 77% for the G1. The next practical test is running that transfer on a real H2 and beyond the single room and demonstrator used here.
