A solver that splits the work
NVIDIA released a new solver called mPDLP in its cuOpt decision-optimization stack on October 7, 2026, described in a company technical blog post by Bulle Mostovoi, Preethi Maulik, Burcin Bozkaya and Rajath Narasimha. The change is aimed at a specific bottleneck: linear programming problems so large they take hours to solve or will not fit in a single GPU's memory at all. The new solver spreads one problem across several GPUs connected by NVIDIA's NVLink interconnect.
For the people who build these models, that distinction decides what kind of planning is possible. Supply chain teams and electricity grid planners run enormous what-if calculations. When a single run takes overnight, planners cut the number of scenarios they examine. When a model will not fit on one machine, they simplify it further. NVIDIA's post states the two headline outcomes plainly: reducing solve time for large problems that are impractical to solve within the time window available for a planner, and up to 6x lower peak memory usage per GPU compared with the single-GPU PDLP solver, with LP problems capped at 2.1 billion nonzeros.
Every speedup in this story is a vendor-reported benchmark, not an independent measurement. That includes the partner figures from Kinaxis and PSR discussed below, which this article attributes solely to NVIDIA's post because the partner blog posts themselves could not be retrieved for verification at publication time.
What a linear program actually is
A linear program, or LP, is a mathematical way of finding the best decision under constraints. Imagine a distribution company deciding how many units to ship along each route. Every route has a cost, every warehouse has a capacity, every customer has a demand. An LP encodes all of that as one equation to minimize (the total cost) plus a list of rules the answer must obey. The solver's job is to find the cheapest answer that still satisfies every rule.
The size of these problems is measured in nonzeros, the individual nonzero entries in the big table of constraints. A realistic supply chain or grid model can easily hold tens or hundreds of millions of them. More nonzeros means more realistic detail: more products, more routes, more hours of the day, more possible generation sites on the grid.
Why does speed matter so much? Because planning is a conversation. A planner does not want one answer; she wants to compare twenty answers: what if this factory closes, what if fuel prices double, what if a heat wave shifts demand. A model that answers overnight supports one question per day. A model that answers in minutes supports an actual discussion. NVIDIA's framing is exactly this: teams need to evaluate larger models and more uncertainty within a practical planning window.
The bottleneck, and the two taxes
cuOpt's existing single-GPU PDLP solver, already shipping, reportedly delivers more than 10x speedups over CPU solvers on large LP problems, according to the same post. The limitation is physics, not cleverness: the core computation of this algorithm family, a repeated sparse matrix-vector multiplication, is limited by memory bandwidth, and one GPU only has so much.
The obvious fix is to use more GPUs, which multiplies available bandwidth. The catch is that distributed algorithms pay two taxes. The first is communication overhead: time spent waiting on information from other GPUs. The second is load imbalance: uneven distribution of work across GPUs, causing some GPUs to wait for others. Both taxes can erase the gains from adding hardware, which is why multi-GPU solvers are hard to build well.
Prior work, the academic method D-PDLP, partitioned the constraint matrix into a grid of sub-blocks and computed each multiplication piece independently, then summed the pieces. It demonstrated that multi-GPU acceleration was possible. The new mPDLP does something more refined.
Min-cut partitioning, without the mathematics
The key insight is that consecutive sparse matrix-vector multiplications in the PDLP algorithm share dependencies through the two solution vectors being iterated. Instead of treating each multiplication independently, mPDLP builds a bipartite dependency graph from the sparsity pattern of the constraint matrix, showing exactly which variables feed which constraints in both directions of the loop.
It then partitions that graph so that tightly connected dependencies are kept on the same GPU whenever possible. Each edge that stays inside a partition can be computed locally; each edge that is cut between partitions requires cross-GPU communication. Min-cut partitioning, then, is essentially graph bookkeeping: it counts the cuts and rearranges the assignment to minimize them, so GPUs spend less time waiting on each other.
The result depends on the structure of each problem. NVIDIA notes the method's performance varies with the sparsity pattern of the matrix and the number of edge cuts, a qualification worth carrying through any reading of the numbers that follow. NVLink and NVSwitch provide the fast hardware links between GPUs, and the NCCL library handles the communication, so that multiple distinct GPUs become one fast shared machine, in the post's description.
The benchmarks, with their fine print
NVIDIA benchmarked mPDLP on more than 100 LP instances on DGX B200 systems, comparing against both the single-GPU cuOpt PDLP and the multi-GPU D-PDLP. Vendor-reported results: speedups correlate strongly with problem size, becoming noticeable when the nonzero count exceeds 10 million and reaching up to 11.4x on PDLP steps alone for the largest problems. On most large-scale instances above that threshold, mPDLP achieved 1.2x to 2.5x speedup over the prior D-PDLP approach.
Two cautions apply. First, the 11.4x figure is the largest case, on the internal PDLP step count rather than full end-to-end solve time, and results vary with sparsity structure, as NVIDIA itself states. Most instances will see the smaller 1.2x to 2.5x range. Second, these are the vendor's own benchmarks on the vendor's own hardware. They are best read as a claim about the shape of the improvement, not a guarantee for any particular model.
What partners reported
Two NVIDIA partners reported results, both relayed through NVIDIA's post rather than verified independently here. Kinaxis, a supply chain planning company, reported a 3.3x speedup on a consumer-goods supply chain model with over 135 million variables, running on eight NVLink-connected H100 GPUs. PSR, an energy systems company, reported more than 5x speedup on a stochastic energy capacity-expansion model with 185 million variables on eight B200 GPUs.
The PSR case deserves a plain-language note: a stochastic expansion model is one that plans the grid's future buildout under many possible futures at once, rather than a single forecast. That is precisely the kind of model that gets trimmed when compute runs short. If a planner can keep all 185 million variables and finish inside a working session, the scenarios that used to be cut for feasibility can come back into the study.
Neither partner blog was retrievable through this newsroom's verification pipeline at publication time, so those figures rest on NVIDIA's account of what its partners reported. The numbers are plausible within the ranges of NVIDIA's own benchmarks, but plausibility is not verification.
The 2008 benchmark that keeps getting faster
One benchmark problem carries symbolic weight in this field. zib03, introduced by researcher Thorsten Koch in 2008, contains over 104 million nonzeros and has served for years as a measuring stick for new solvers. NVIDIA's post reports that cuOpt with mPDLP now solves it nearly 10x faster than a year prior, shown in a chart of progression of solve times by linear programming solvers on the problem.
A roughly tenfold improvement in one year is a statement about pace, not a solution to every planning problem. But it illustrates the compounding effect of algorithmic work like min-cut partitioning landing on top of faster hardware. A problem that represented the frontier of the impractical in 2008 is now, on this vendor's evidence, an overnight job compressed toward minutes.
For a planner, the practical translation is blunt: larger models and more scenarios inside the same planning window, which in practice means cheaper logistics studies and more realistic grid buildout analyses, if the vendor's results hold up in third-party use.
What is verified, and what is not
The claims in this article rest on one primary source: NVIDIA's own technical blog. The math behind the method is published and the cuOpt source code is open on GitHub, with a tutorial for reproducing the multi-GPU comparisons, so independent confirmation is possible in principle. Until independent groups publish their own numbers, the honest summary is this: NVIDIA says its new distributed solver makes 100-million-variable decision problems tractable on multi-GPU machines, and the benchmark history it cites is consistent with that claim. Whether those speedups transfer to a specific company's messier, real-world model is the question only their own runs can answer.
