Korea University’s method removes the training aid before testing, then compares simulation results with limited physical-robot trials.
Korea University researchers propose a reinforcement-learning curriculum that temporarily changes how a humanoid experiences gravity. Early in training, a command-conditioned tilted-gravity field gives the robot slope-like help for forward motion, while added rewards discourage unnecessary positive motor work.
The virtual slope is not part of the finished controller. The tilt fades, and training continues on flat ground with normal gravity before evaluation. In engineering terms, the team changes the conditions used to search for a gait, not the robot’s hardware or runtime controller.
In controlled simulation, the full curriculum lowered mechanical cost of transport—the actuator work needed to move a given distance—by 6.8% to 15.2% at tested speeds from 0.5 to 2.0 meters per second, according to the paper on arXiv. Command tracking did not measurably worsen.
The savings mainly came from positive actuator work, which supplies energy to the motion. With reward terms held constant, the paper reports that tilted gravity did not make stepping appear earlier. Instead, it affected the gait the policy retained after training.
The separate physical-robot evaluation used a 29-degree-of-freedom Unitree G1. The strongest result required a walking-specific motion prior, or reference-based training signal: reported mechanical cost fell 16.3% for forward commands and 17.0% for lateral commands, while backward walking was essentially unchanged.
That hardware test used one deployed policy per condition and only two to six steady-state traversals per direction. Its distance estimate came from onboard joint and inertial data, and the metric measured actuator mechanical work—not total electrical power.
The next practical test is whether independent trials on other humanoids preserve the gain during longer operation and real disturbances.
