Two research methods target different control problems: composing whole-body skills and adapting human-object contact to a robot’s body.
A humanoid needs more than a list of joint angles. It must coordinate its body, keep balance, and stay connected to floors, tables, and objects while moving through a task.
Two research papers posted on arXiv describe complementary ways to handle those problems. Georgia Tech’s spectral skills method gives a controller a compact command interface for whole-body movement. OTRetarget, from researchers including Inria and Stanford, transfers human demonstrations while allowing the robot and manipulated objects to adjust together.
Both are research demonstrations, not consumer products or ready-made workplace robots. Their experiments include simulation and Unitree G1 hardware, but they do not show broad deployment.
Turning motion into reusable commands
A robot controller turns commands into joint actions. For a humanoid, those commands may involve 29 degrees of freedom, meaning 29 independently controlled movement axes.
The Georgia Tech work focuses on the layer between a high-level planner and a low-level controller. Instead of asking the planner to predict every joint position at every moment, the system gives it a shorter motion code called a spectral skill.
The code represents a short movement segment. Its unusual feature is how researchers train it. The system learns by predicting what motion comes next, rather than reconstructing only the segment it just received.
That encourages the code to capture a movement’s future effects. A skill can therefore describe more than a pose: it can help communicate how a step, turn, or arm movement should continue.
This matters because long sequences of joint targets are difficult for a planner to predict. A compact skill acts more like a reusable instruction. One code can represent a base movement, while other directions modify it.
The paper calls this process steering. Researchers add directions in a latent space—a hidden numerical space used by the model—to change part of a base movement. In the reported demonstrations, the robot could raise an arm or turn while continuing another motion.
The same frozen controller could also chain independently encoded skills without a separate transition policy. In plain terms, it could switch from one learned movement to another instead of requiring one long motion sequence for every combination.
In the reported evaluation, a 29-degree-of-freedom humanoid reduced global tracking error by 62 percent compared with the paper’s state-of-the-art reference. A language-conditioned planner’s success rose from 77.1 percent to 91.1 percent when it predicted skills instead of explicit trajectories.
Those figures come from the authors’ evaluations across simulation and demonstrations on Unitree G1 hardware. They do not show that a humanoid can reliably invent any movement combination in a home, warehouse, or factory. The hardware work used selected research tasks and a deployment version of the tracking controller trained with hardware-oriented adjustments.
Preserving the relationships that make tasks work
OTRetarget addresses a different failure. A robot can copy a person’s posture while missing the physical relationship that makes the task succeed.
Consider lifting a box from the floor and placing it on a table. Matching the person’s hand and body positions may place the box too high, too low, or beyond a shorter robot’s reach. Shrinking the entire demonstration can also break the box’s relationship with the floor or table.
OTRetarget treats contact as geometry, not just skeletal motion. It samples points from the human, robot, objects, and surrounding surfaces. For each point, it records a signed distance, which indicates its position relative to a surface; the closest point on that surface; and the direction between them.
The method then uses entropic optimal transport, a mathematical technique for matching points between different shapes. Those matches transfer contact relationships from the human body to the robot’s body.
The important engineering choice is that object poses can change during retargeting. A constrained inverse-kinematics solver—software that finds joint positions satisfying movement and physical requirements—jointly adjusts the robot and objects at each frame.
The solver balances several goals: preserve contact, retain the demonstrated motion style, and obey limits such as joint ranges, velocity limits, and collision constraints. This allows the robot to move a box to a reachable height instead of blindly following the human’s exact box trajectory.
The approach also handles object-to-object relationships, such as placing a box on a table. That is useful because many real tasks depend on several contacts at once: hands with a box, a box with a table, and feet with the ground.
On the OMOMO evaluation reported by the OTRetarget authors, the method achieved an 87 percent robot-object interaction Jaccard score and an 8.7-millimeter depth error. A Jaccard score measures how much the robot’s contact set overlaps the demonstrated contact set.
For comparison, the OTRetarget authors report 28 percent and 29.3 millimeters for OmniRetarget. Those are figures from OTRetarget’s comparison, not results reported by OmniRetarget’s authors.
The researchers also trained a whole-body policy from retargeted references and ran it on a physical Unitree G1. It completed six of eight two-handed box pick-and-place trials without retraining or object perception.
The paper describes that result as a single-robot, single-clip demonstration rather than a hardware study. Its evaluation was limited to one robot, flat ground, and one two-object capture. The method is also kinematic and frame-by-frame, so it does not explicitly model forces or optimize over a full time horizon.
Where the methods might meet
The two methods operate at different stages. OTRetarget creates robot-and-object reference motions that preserve important contacts. Spectral skills give a controller a compact way to execute, chain, or modify whole-body movements.
That suggests a possible pipeline: first create a contact-aware reference, then encode parts of it as skills for a controller. But the papers do not report such an integrated system, so the combination remains a research direction.
For operators, the practical question is whether future humanoids can handle both sides of physical work: planning a useful sequence and maintaining the contacts that make the sequence succeed.
The next meaningful test is broader hardware evaluation with different robots, uneven surfaces, moving loads, forces, and multiple objects. Until then, these methods are promising control components—not robots people can buy or deploy at scale.
