Summary
NVIDIA says its DSX platform can increase AI token throughput within fixed power budgets and help data centres respond automatically to electricity-grid signals. The company cites a Lambda validation on HGX B200 servers and a commercial-scale Emerald AI demand-response deployment.
NVIDIA has detailed DSX, a suite of software, simulation tools and reference designs intended to increase useful AI work from a fixed electricity supply. The company says early deployments have increased token throughput within the same power budget and allowed an AI factory to reduce demand automatically when the local grid was constrained.
The announcement frames the central measure for large AI facilities as work per megawatt rather than simply the size of the power connection. In this context, tokens are the small units used to represent model input or output, while token throughput measures how quickly an AI system processes them.
More tokens from the same power budget
NVIDIA DSX MaxLPS monitors power use at GPU and rack level and reallocates available power headroom between computing nodes. The system is designed to account for the different power patterns of training and inference workloads, rather than provisioning every node statically at its maximum.
NVIDIA reports that cloud provider Lambda validated MaxLPS on a five-rack, 19-node cluster using NVIDIA HGX B200 GPU servers. Lambda ran 19 nodes at 85% power within the same facility budget previously used by 16 nodes at full power. Cluster token throughput increased from roughly 4 million tokens per second to 5 million, which NVIDIA reports as a 24% gain. Performance per watt improved by 23%.
The result is specific to that deployment and workload configuration. NVIDIA also projects that MaxLPS, combined with data-centre power planning, could provide up to 40% more GPU capacity in next-generation Vera Rubin NVL72 AI factories operating within the same megawatt budget in suitable environments.
AI factories can respond to grid signals
DSX Flex is intended to connect an AI facility’s workload management to external grid events, including demand-response requests, load-shedding signals and pricing events. Instead of shutting down the facility, the system ranks workloads so that lower-priority jobs can be slowed or rescheduled while higher-priority services continue operating.
NVIDIA describes a deployment involving Emerald AI’s Conductor platform and Silicon Valley Power. When the utility sent a signal to the AI factory, Conductor automatically reduced power from 4 megawatts to 3 megawatts in under a minute. High-priority inference continued running and no operator had to intervene. According to NVIDIA, Silicon Valley Power has sent more than 200 demand signals to the facility and each response worked as intended.
The Santa Clara deployment is not described as a completed DSX Flex installation. NVIDIA presents it as an earlier commercial-scale demonstration of the flexibility that DSX Flex is designed to support, with Emerald AI Conductor intended to integrate into the platform as it develops. The company says the first dedicated DSX Flex commercial deployment is planned for a 96-megawatt Vera Rubin AI factory at NVIDIA’s AI Factory Research Center in Manassas, Virginia.
A whole-factory approach to efficiency
DSX extends beyond workload scheduling. DSX OS is described as open-source, modular software for AI-factory lifecycle management, runtime consistency, health automation and resiliency. DSX Sim is intended to let operators model factory designs before construction, while DSX Reference Designs combine compute, networking, storage and facility systems into validated architectures for particular generations of hardware.
The platform also incorporates NVIDIA’s 800V DC power architecture. NVIDIA says the higher-voltage design is intended to reduce conversion complexity, improve power delivery and support denser accelerated-computing racks. Its current projection is a 3% to 5% end-to-end efficiency gain compared with 54V distribution, with availability alongside Vera Rubin NVL72 in 2027.
This system-level focus reflects the constraints of large AI facilities. A faster GPU can remain underused if networking, cooling or power distribution becomes the bottleneck. NVIDIA cites direct-liquid-cooled GB200 NVL72 racks handling about 120 kilowatts of heat, illustrating why facility engineering is part of the computing problem rather than a separate concern.
The practical significance of DSX is therefore twofold: operators may be able to produce more AI output from an existing power connection, while utilities may gain a way to treat some data-centre workloads as flexible demand. The reported gains are tied to particular hardware and deployment conditions, and the larger Vera Rubin and 800V figures remain company projections or planned capabilities.