Skip to content
SUNDAY, AUGUST 2, 2026
HumanoidsLegacy Report1 recorded source

One GPT model tracks full humanoid motion across bodies

Visual status: no verified article image is available. The reporting remains text-first.

A single GPT-based tracker follows full humanoid motion across multiple bodies in simulation.

The VENOM paper, VENOM: Versatile Embodied Network for Omni-bodied Motion tracking, introduces a cross-embodiment full-body motion tracking model for humanoids that operates entirely in simulated environments. The model is trained on a multi-humanoid dataset called the VENOM dataset, which catalogs states, actions, and rewards, and is designed to avoid the common practice of splitting control into separate upper and lower body streams. In other words, VENOM aims to track the entire body with one unified policy, rather than stitching together two halves of a control problem.

Testing shows that VENOM achieves stable motion tracking across different humanoids and does so with capabilities that surpass a multi-humanoid MLP baseline trained with supervised learning alone. Remarkably, VENOM also closely matches the tracking performance of experts trained using asymmetric-actor critic reinforcement learning, even in the absence of explicit reward feedback during execution. The work argues that the model’s cross-embodiment generalization comes from learning directly from a diverse, multi-humanoid data distribution and from keeping the tracking problem cohesive rather than partitioned.

In practical terms, the study demonstrates a higher bar for what a single model can accomplish in sim-to-sim environments: a unified tracker that stays stable when the morphologies change, rather than requiring tailor-made controllers for each form. The VENOM dataset is central to this, providing the states, actions, and reward signals used to train and evaluate the system. The authors emphasize that this is a simulation-focused milestone, with the demonstrated performance grounded in how the model handles varied embodiment rather than real-world sensor noise or hardware idiosyncrasies.

From a practitioner standpoint, the result is provocative for robotics engineering and systems design. First, data diversity across humanoid morphologies matters: the cross-embodiment success hinges on exposing the model to a range of forms during training, so a single tracker can generalize to unseen bodies. Second, removing the split between upper and lower body control reduces software complexity and integration friction, but it places heavier demand on the sequence model’s capacity to reason over long horizons and varied kinematics. Third, translating this from simulated success to real robots will require attention to sim-to-real gaps, including sensor realism and hardware limits; domain randomization and robust system identification are likely to be the next essential steps. Fourth, the approach raises questions about data efficiency and compute: a GPT-based tracker trained on multiple morphologies may demand substantial training data and compute, so practitioners will watch for improvements in data efficiency and real-time performance as the method moves toward practical use.

The VENOM result is notable not as a finished product for factory floors, but as a concrete engineering demonstration of cross-embodiment learning for full-body tracking in simulation. It signals a shift in how engineers think about building motion trackers that work across forms rather than for a single chassis, while underscoring the gap that remains before such systems can routinely operate on real humanoid hardware with all the frictions that entails.

Sources & methodology
  1. VENOM: Versatile Embodied Network for Omni-bodied Motion tracking
    arXiv Humanoid/Bipedal Query / Primary source / Published JUN 15, 2026 / Accessed JUN 16, 2026

Newsletter

The Robotics Briefing

New signups are closed while external email delivery is being verified. No email address is collected here.

Follow the live RSS feeds