What happened: more GPUs inside the same power limit

On Sept 27, 2026, NVIDIA published a joint NVIDIA and Nscale evaluation describing a software approach called DSX MaxLPS that increased the number of GPUs running inside an unchanged approved power budget. In the evaluation, a fleet of 140 GPUs grew to 192 GPUs under the same 264.4 kW provisioned power limit, and measured aggregate inference throughput rose about 49 percent. The findings come from a vendor-run test and have not been independently audited, but the numbers, the stated trade-offs and the underlying idea are specific enough to matter for anyone following AI infrastructure costs and energy constraints.

The article begins with a short, plain statement of the problem NVIDIA says it is solving:

Every unused watt is capacity left on the table.

Sarah McKenney and Harry Petty, NVIDIA Technical Blog, Sept 27, 2026, "How NVIDIA DSX MaxLPS Maximizes AI Factory Throughput and Efficiency," opening line.

This report explains what dynamic power sharing is, what the evaluation measured, where the trade-offs showed up, and how much of this should be treated as a vendor claim rather than established fact.

Why reserved power sits idle

Data centers are wired with layers of electrical limits: the utility connection feeding the building, the substations and power-distribution gear, the racks, the individual nodes, and finally the GPUs themselves. Operators have to stay under each of these approved limits, and the safest way to do that with static planning is to size each layer for the rare worst case, when every GPU draws its maximum rated power at the same time.

That worst-case sizing is why the headroom exists, but it also means the headroom usually sits unused. In reality, AI workloads fluctuate. Training cycles through compute-heavy phases, communication waits, synchronization and checkpointing. Inference alternates between processing incoming prompts, generating output tokens, memory-bound work and idle moments. Even two copies of the same model can draw different amounts of power depending on the requests they receive.

NVIDIA says AI factories are typically provisioned for "the unlikely moment when every GPU reaches peak power" (Sarah McKenney and Harry Petty, NVIDIA Technical Blog, Sept 27, 2026). The consequence is a persistent gap between the power a facility reserved and the power it actually consumes. Under static per-node reservations, one node cannot lend its unused allowance to a neighbor, so the facility can be well under its total limit while additional GPUs sit switched off or never get installed at all.

The evaluation numbers illustrate the size of that gap. In the static configuration, mean GPU power was 97.0 kW and total measured power was 166.2 kW against a 264.4 kW provisioned budget, meaning the site used about 62.9 percent of what had been approved. More than a third of the budget was reserved but idle.

How the control loop reallocates power

NVIDIA's answer is a control layer it calls DSX MaxLPS, built on its Dynamic Power Software. Rather than increasing the site's power supply, the software monitors how much power each participating resource is actually drawing and continuously adjusts GPU power limits so that unused allowance flows to resources that can use it.

The blog describes the mechanism in one careful sentence:

This is coordinated allocation, not an increase in the site's power supply.

Sarah McKenney and Harry Petty, NVIDIA Technical Blog, Sept 27, 2026, section "Inside the DSX MaxLPS control loop." The emphasized distinction is the vendors' own: no new electrical capacity is created, only more productive use of the existing limit.

According to the blog and NVIDIA's Dynamic Power Software documentation (version 0.8, currently offered as a Developer Preview), the control loop has five elements. First, operators map their power topology and define a managed resource group with an aggregate budget. Second, telemetry is collected at GPU, node, rack and group level to detect available headroom and emerging power events. Third, operator-defined policies set node and group limits, allocation priorities, reserve requirements and responses to maintenance or emergencies. Fourth, when some resources draw less than their allocation, the software raises the power limits of participating GPUs so others can use the slack. Fifth, validation compares measured power against the approved budget and adjusts as consumption approaches the limit.

The critical safety property is that the approved aggregate budget is preserved at all times. The software reallocates within it; it does not expand it. NVIDIA's own framing stresses that this requires reliable telemetry, because delayed or mis-mapped measurements could undermine fleet-level decisions.

What the numbers measured

The evaluation was run at Nscale's data center at the Verne campus in Keflavík, Iceland, which the blog describes as powered entirely by renewable energy. Nscale deployed the MaxLPS software and collected telemetry while NVIDIA ran the workloads. The test used NVIDIA GB300 NVL72 systems with Blackwell Ultra GPUs, running Kimi K2.5 in FP4 precision through NVIDIA Dynamo and TensorRT LLM, with 8,000-token inputs and 1,000-token outputs.

The setup deliberately mixed two service profiles: high-throughput inference instances and a low-latency instance, because real AI factories serve heterogeneous workloads with different power profiles. Distributed jobs were confined to single racks to control for cross-rack performance differences.

The static baseline used 35 four-GPU nodes, or 140 GPUs, split into two high-throughput instances of 52 GPUs each and one low-latency instance of 36 GPUs. The DSX MaxLPS configuration used 48 four-GPU nodes, or 192 GPUs, adding a third 52-GPU high-throughput instance while keeping the same low-latency instance.

The measured results, as reported in the blog:

Managed GPUs: 140 to 192, up 37.1 percent.

Aggregate throughput: 1,084,503 tokens per second to 1,618,443 tokens per second, up 49.2 percent.

High-throughput output per instance: 59,153 to 59,220 tokens per second, up 0.1 percent.

Low-latency output per instance: 2,265 to 2,265 tokens per second, unchanged.

Mean GPU power: 97.0 kW to 131.8 kW, up 35.9 percent.

Total measured power: 166.2 kW to 198.9 kW, up 19.7 percent.

Power-budget utilization: 62.9 percent to 75.2 percent, up 12.3 percentage points.

Throughput per provisioned watt: 4.10 to 6.12 tokens per second per watt, up 49.2 percent.

The headline working figure of 37 percent more GPUs matches the 37.1 percent fleet increase. The per-instance figures are the important nuance: the existing high-throughput and low-latency services did not visibly slow down, so the new throughput came from the added 52-GPU instance rather than from degrading existing ones. The 49.2 percent throughput-per-watt figure simply mirrors the throughput increase, because the 264.4 kW denominator did not change.

What the trade-offs mean

The evaluation is candid about a cost. While median and 75th-percentile latency stayed within 5 percent of baseline, the 99th-percentile time to first token rose 17 percent from a 15.7-second baseline. P99 captures the slowest 1 percent of requests, so the worst-tail experience of the low-latency service measurably degraded even though typical requests were barely affected.

This is the concrete trade-off readers should remember: dynamic power sharing converts unused electrical headroom into extra capacity, but it does so by letting the whole fleet run hotter on average, and tail latency is where the squeeze shows up first. The blog itself recommends that production acceptance criteria include tail-latency checks, noting that "Stable median latency can mask changes in tail latency" (Sarah McKenney and Harry Petty, NVIDIA Technical Blog, Sept 27, 2026).

The blog also points out that the size of the opportunity depends on workload mix. Workloads with complementary power profiles leave more headroom to harvest than workloads that peak at the same time, which is why the team chose a mixed workload rather than a homogeneous one. This means results from this specific Kimi K2.5 test cannot be assumed to transfer to every deployment; operators are advised to run representative production workloads through a five-stage validation process before adding capacity.

Vendor-run evidence and how to read it

Every performance number in this article comes from the joint NVIDIA and Nscale vendor-run evaluation described in NVIDIA's own technical blog. NVIDIA supplies both the software being evaluated and one of the parties running the test, so this is self-reported evidence, not an independent audit. The documentation confirms the software is still in Developer Preview status, meaning it may change significantly and is not recommended for general production use except within authorized collaborations.

That said, the internal consistency of the reported numbers is notable. The 37 percent GPU increase, the 49.2 percent throughput increase, the unchanged per-instance outputs and the power-budget utilization figures all fit together arithmetically, and the authors explicitly publish the unflattering P99 latency regression alongside the gains. Publishing a trade-off against your own product is a reasonable, though not conclusive, sign of good faith.

The distinction that matters for assessment: the measurements are observations from one controlled deployment of a specific workload on specific hardware at one site. The claims that this generalizes to arbitrary AI factories, or delivers up to 40 percent more GPUs in other power environments, are NVIDIA's predictions, not demonstrated results. Whether renewable-powered sites in Iceland, with their particular cooling and grid conditions, generalize to other geographies is an open question no source in this article answers.

The reader takeaway is a measured one. Data centers around the world hold significant reserved-but-unused electrical headroom, and software that can safely reallocate it would raise AI capacity from infrastructure and grid connections that already exist, without new construction or new power procurement. The cost is a quantified degradation in worst-case response time, and the honest bottom line from the vendors' own test is that this is an engineering trade-off to be validated per deployment, not a free lunch.