An NVIDIA-reported joint evaluation with Nscale found 49.2% more aggregate throughput, while P99 time to first token rose 17%.
NVIDIA’s DSX MaxLPS software shifts unused power between managed GPU resources. It keeps the site within an operator-approved power budget; it does not increase the facility’s power supply.
The control system uses telemetry—power measurements from GPUs, nodes, racks, and resource groups—to find available headroom. Operator policies set limits, reserves, priorities, and responses to emergencies. When one resource draws less power, the software can raise another resource’s limit while checking the shared budget.
Nscale deployed the software at its data center in Keflavík, Iceland, while NVIDIA ran the workloads and collected telemetry. The evaluation used Kimi K2.5 inference workloads on NVIDIA Blackwell Ultra GPUs, with 8K input and 1K output sequences. It compared 140 GPUs under static provisioning with 192 GPUs using DSX MaxLPS.
NVIDIA reports that both configurations used the same 264.4-kilowatt provisioned power budget. Aggregate throughput rose from 1.085 million to 1.618 million tokens per second, a 49.2% increase. Throughput per provisioned watt increased from 4.10 to 6.12 tokens per second per watt.
The trade-off appeared in slow requests. Median and P75 latency stayed within 5% of baseline, but P99 time to first token increased 17%, from 15.7 seconds. P99 describes the slowest 1% of requests, which can matter for interactive services.
Independent replication is not established. Operators should stage their own tests with representative workloads, checking telemetry accuracy, power compliance, per-instance performance, P99 latency, and behavior during faults or reduced power availability.
