The simulation benchmark checks climbing, bulb handling, and safety—but climbing was not demonstrated.
A humanoid replacing a light bulb from a ladder must solve several problems at once. It must position the ladder, climb while balancing, handle a fragile bulb, install the replacement, and dispose of the old one without dropping or crushing anything.
Fiatlux turns that chore into a structured simulation benchmark for a Unitree G1 humanoid. The paper was developed by researchers affiliated with the University of Hawaiʻi at Mānoa and collaborators.
The benchmark is not a physical-robot demonstration. It tests whether current control methods can combine whole-body movement, ladder contact, manipulation, and fragile-object safety inside NVIDIA Isaac Lab, a physics simulation platform.
What the benchmark asks
In the complete simulated episode, the G1 positions a step ladder under a ceiling or wall fixture. It then climbs, removes the spent bulb, handles a fresh bulb, seats it in the socket, and places the old bulb in a disposal crate.
That sequence matters because each action changes the conditions for the next one. A robot may walk well on the floor but lose balance when its feet move onto ladder rungs. It may climb successfully but fail when one hand must hold a bulb. It may insert the new bulb but crush it by applying too much force.
Fiatlux breaks the workflow into 12 atomic subtasks. They cover flat-floor navigation, manipulation, and climbing. Some combine those modes, such as climbing while carrying a bulb or working from the ladder.
The subtasks are separate simulation environments, with starting states designed to resemble the previous stage’s endpoint. They are not an automatically chained run in which one stage’s actual mistakes become the next stage’s starting condition. That separation helps researchers isolate failures, but it is less like a complete maintenance job.
How Fiatlux defines safe success
Fiatlux does not count bulb insertion alone as success. A clean full-credit episode must seat the fresh bulb, place the spent bulb in the crate, drop neither bulb, and remain below the hand-force limit for the fragile payload.
This distinction matters. A robot that forces a bulb into place might appear successful if the only question is whether the bulb reached the socket. Fiatlux separately checks contact force, so an apparent success can also be marked as a broken or unsafe run.
The benchmark gives partial credit for progress toward a subtask. A failed attempt may still receive credit for conditions it achieved, such as moving the ladder or holding a stable posture. But the main safety measure is clean success, not partial progress.
Fiatlux also separates two kinds of simulated observations. Standard observations use signals a physical robot could sense or estimate, including joint movement, body orientation, contact forces, camera features, and distance readings from LiDAR. LiDAR measures distance with light.
Privileged observations contain exact simulator information, such as the precise positions of the robot, ladder, fixture, bulbs, and disposal crate. The benchmark uses those signals for critic training and privileged baselines rather than treating them as ordinary robot senses.
What the validation actually tested
The researchers used teleoperation to check whether each subtask could be completed in simulation. An operator controlled the simulated robot through a virtual-reality headset, while the benchmark used the same scene, physics, success gates, and scoring rules applied to policy evaluations.
There was one important difference in how those demonstrations ran: failure terminations and the episode timeout were cleared. That allowed the operator to continue after a mistake. The success predicates still ran, and the recorded takes were scored offline using the benchmark’s normal rules.
Teleoperation produced clean, full-credit takes for eight of the 12 subtasks. These were the non-climbing navigation and manipulation tasks, including moving the ladder, handling bulbs, and placing objects in their required locations.
The four climbing subtasks had no teleoperated take that satisfied its full success gate. For those tests, the robot was initialized on ladder steps while holding the ladder. The setup therefore tested stability and task conditions from an established ladder position, not the full process of approaching the ladder and climbing onto it from the floor.
That qualification is central. Fiatlux established that eight subtasks could be completed under its simulated setup and scoring rules. It did not establish that the G1 can independently approach and climb the ladder, climb while carrying a bulb, descend safely, or complete the entire errand.
The paper also reports results from three released autonomous baselines: a zero-action policy, a random policy, and zero-shot NVIDIA GR00T N1.7. None completed any subtask. GR00T’s weighted score was not separable from the zero-action policy within the reported variation, meaning the measured difference was smaller than either reported standard deviation.
Those results do not show that humanoid maintenance is impossible. They show that the released policies did not solve this benchmark under the tested conditions.
Why the result matters
Fiatlux exposes the transitions that simple robot demonstrations can hide. Walking, climbing, grasping, force control, and safe release may each look manageable in isolation. Combining them while maintaining balance and protecting a fragile object is a harder systems problem.
For operators and investors, the benchmark is therefore a measurement tool, not a deployment claim. The paper says that an adapter connecting its joint-position targets to Unitree software commands remains future work, and the reported validation does not establish physical-robot execution.
The next useful milestone is clear: demonstrate the climbing subtasks from realistic floor-level starts, connect the stages into a continuous handoff, and repeat the clean-safety test on hardware. Until then, Fiatlux identifies the difficult engineering boundary in humanoid maintenance rather than showing a robot ready to change bulbs.
