The multi-GPU solver is a developer tool for supply-chain and energy models, with gains that depend on size and sparsity.
NVIDIA has introduced cuOpt mPDLP, a GPU-based solver for very large linear-programming problems. These models help planners choose actions under limits, such as assigning products to factories or expanding an energy system.
The release matters to AI and machine-learning infrastructure teams because it targets a common computing bottleneck: moving data between processors. PDLP repeatedly performs sparse matrix-vector multiplication, which calculates only the relationships present in a model.
cuOpt mPDLP uses min-cut partitioning to place closely connected calculations on the same GPU. That leaves fewer relationships crossing between GPUs over NVLink, NVIDIA’s high-speed GPU connection. Less cross-device traffic can reduce waiting during each solver iteration.
NVIDIA compared mPDLP with single-GPU cuOpt PDLP and the multi-GPU D-PDLP implementation across more than 100 linear-programming instances. The company reports that mPDLP’s speedups become noticeable above 10 million nonzero matrix entries. On most larger instances, it was 1.2 to 2.5 times faster than D-PDLP.
The headline result needs context. On one benchmark, the PDLP steps were up to 11.4 times faster, but the full end-to-end run was 4.2 times faster after including setup, partitioning, and final processing. NVIDIA also says three ultra-large benchmark cases ran slower than D-PDLP.
The practical test is workload-specific. Smaller models can lose time to synchronization and data transfers, while heavily connected models can create too much cross-GPU traffic. NVIDIA offers a tutorial for running a reader’s own model, making that comparison—using full runtime, not just solver steps—the next useful step.
