The research system avoids copying human joints, but still needs robot-specific calibration and demonstrations.

IronMind is a research system that learns humanoid manipulation from first-person human video without directly mapping human body movements onto robot joints. In an arXiv paper, its authors report closed-loop tests on the IRON-R01 humanoid robot.

The central idea is a camera-space action representation. Instead of describing a person’s movement in a human torso frame, IronMind describes what the hand does relative to the camera. Robot trajectories are converted into that same camera-based frame using camera calibration, then the system aligns actions by meaning, such as wrist or finger motion.

This sidesteps a difficult engineering problem. Human hands and robot end-effectors have different shapes, joints, and limits, so directly copying human poses can produce noisy or physically unsuitable commands.

The authors report training on more than 10,000 hours of egocentric human video and heterogeneous robot data. They filtered tracking failures, divided recordings into single-task clips, refined captions, and gave clearer frames more weight. The model combines vision-and-language processing with a continuous action generator.

IronMind does not transfer directly to every humanoid. The paper says IRON-R01 demonstrations entered only during post-training, when the system was adapted with teleoperated robot data and extrinsic camera calibration.

After that adaptation, the authors report a 55.0% average success rate for the 10,000-hour model across six unfamiliar real-robot tasks. A torso-frame version using the same training budget reached 26.7%. Each task used 10 trials, and the evaluation covered one robot and six tasks.

That is a useful research result, not evidence of household reliability. Longer tasks, broader contact-heavy work, and transfer to another humanoid remain the next practical tests.