An IIT Mandi paper proposes a two-level controller that links route planning with gait limits, tested only in PyBullet simulation.
A biped cannot always turn or step sideways simply because a navigation planner requests it. Its balance, joint limits, and walking pattern may make that command unsafe.
A paper by researchers at the Indian Institute of Technology Mandi proposes a hierarchical reinforcement-learning controller for this problem. The high-level policy chooses a body-velocity command—forward speed, sideways speed, and turning rate—using the robot’s pose, a local goal, proximity readings, and the positions and speeds of nearby moving obstacles.
A low-level policy then turns that command into targets for the robot’s eight actuated leg joints. It runs at 50 hertz, while the navigation policy updates at 5 hertz and holds each command for 10 control steps. Both policies are trained together with Soft Actor-Critic, a reinforcement-learning method.
The controller also includes a time-to-collision signal. In plain terms, it can penalize a predicted crash before an obstacle becomes immediately close, giving the walking robot time to change course.
In 100 trials per method and condition, the authors report 98% goal success in static simulated environments and 88% with moving obstacles. Planner combinations using the same learned gait reached no more than 78% and 68%, respectively. Path lengths were within 4% of the A* reference, although averages covered successful episodes only.
Those results are simulation results, not a physical humanoid demonstration. The PyBullet system used simulator raycasts and obstacle states standing in for LiDAR and object tracking. The next practical test is whether noisy onboard sensing, timing delays, and real leg mechanics preserve the controller’s advantage.
