AWS and NVIDIA announced on 26 August 2026 that they plan to deploy 2 million additional NVIDIA GPUs across AWS’s global infrastructure during 2027–2028. The companies describe the plan as an expansion of their 16-year relationship across not only GPU instances, but also CPUs, networking, memory, managed AI models, data processing and robotics.

The new GPU commitment is additional to AWS’s previously announced plan to add more than 1 million NVIDIA GPUs starting in 2026. It remains a future deployment plan: the announcement does not establish how many of the GPUs have been installed, when customers will receive access, or where the capacity will be located.

Contents

What changed

The planned expansion covers several NVIDIA GPU families, including Blackwell Ultra, Rubin and Rubin Ultra, as well as additional Blackwell capacity. AWS also plans to add NVIDIA RTX PRO 4500 Blackwell Server Edition GPUs to Amazon EC2 G7 instances.

Amazon EC2 is AWS cloud infrastructure through which customers can use computing instances and accelerators. GPUs are particularly useful for AI because they can perform many similar mathematical operations in parallel. That makes them important for both training, where model parameters are adjusted using data, and inference, where a trained model produces outputs.

AWS says G7 instances provide 4.6 times the AI inference performance and 2.1 times the graphics performance of previous-generation G6 instances. Those are company-reported comparisons; the supplied announcement does not provide independent testing, detailed benchmark methodology or workload configurations.

The wider collaboration includes:

  • NVIDIA Spectrum networking for large-scale AI training across GPU clusters;
  • planned NVIDIA Vera CPU-based infrastructure on AWS for agentic AI workloads;
  • support for NVIDIA’s NVLink Fusion with NVIDIA’s custom high-bandwidth memory technology, called NVHBM;
  • integration of that memory and interconnect work with Amazon’s Annapurna Labs Trainium chips;
  • 100,000 GPUs on secure AWS infrastructure for U.S. government workloads;
  • broader use of NVIDIA Nemotron open models through Amazon Bedrock and Amazon SageMaker; and
  • collaboration between Amazon Robotics and NVIDIA on simulation, synthetic data, robot training, route optimisation, functional safety and real-to-sim validation.

The government infrastructure is intended to support workloads classified at Impact Level 6 and above. The announcement does not provide independent details about particular authorisations, deployment locations or security certifications.

How the expanded stack fits together

The announcement is significant less because of a single new GPU than because it describes a more integrated computing stack.

In a large AI system, accelerator count is only one constraint. GPUs must receive data, exchange intermediate results and access memory quickly enough to remain useful. A high-bandwidth interconnect such as NVLink moves data between processors and memory faster than ordinary host connections, which is important when a workload is divided across many accelerators.

High-bandwidth memory, meanwhile, uses very wide memory interfaces placed close to a compute package. This can increase the rate at which large tensors—arrays of numbers used by AI models—move between memory and processors. AWS and NVIDIA’s proposed NVHBM and NVLink Fusion work is intended to connect NVIDIA technologies with AWS’s Trainium accelerators in a common rack-scale architecture.

The companies are also combining different processor types. GPUs handle highly parallel workloads, while CPUs are general-purpose processors that can manage broader sequential tasks and system orchestration. The planned Vera CPU infrastructure would add another NVIDIA-based option alongside GPU and Trainium instances.

AWS and NVIDIA identify the AWS Nitro System and Elastic Fabric Adapter as part of the infrastructure foundation for NVIDIA GPU-based and Trainium-based EC2 instances. The announcement presents these technologies as part of the foundation for large-scale AI systems.

The collaboration also targets bottlenecks outside model computation. AWS and NVIDIA report up to 3.7 times faster data processing and 30% better price performance for certain Amazon EMR workloads when using EC2 G7 instances and NVIDIA cuDF, compared with CPU-based configurations. For Amazon OpenSearch workloads, they report up to nine times faster vector indexing at one-quarter of the cost with GPU acceleration.

Vector indexing organises numerical representations of data so that systems can efficiently search for similar items. It is used in applications such as semantic search and retrieval-augmented generation, where a model retrieves relevant information before producing an answer.

What the plan could mean for customers and governments

If the planned capacity is delivered, AWS customers could eventually have access to substantially more NVIDIA-accelerated computing for model training, inference, scientific workloads and enterprise automation. The intended users include frontier AI laboratories, enterprises, startups and government organisations.

For developers using NVIDIA’s Nemotron open models, the announcement describes two AWS routes. Amazon Bedrock provides fully managed, serverless access to the models, while Amazon SageMaker supports deployment and fine-tuning on customer infrastructure. These are the two deployment paths identified in the announcement.

The government programme extends the plan into secure public-sector and national-security computing. AWS and NVIDIA say they intend to build AI factories with 100,000 GPUs on secure AWS infrastructure for U.S. federal and national-security workloads at Impact Level 6 and above.

The robotics work similarly broadens the scope beyond cloud-based software. Physical AI refers to systems that perceive, model or act in the physical world, including robots and autonomous machines. The announcement identifies simulation, synthetic-data generation, robot training and real-to-sim validation as parts of the collaboration, but it does not provide evidence showing how the proposed work performs on deployed robots outside simulation.

Limits and what to watch

The announcement establishes what AWS and NVIDIA plan to do.

Several practical details remain unspecified:

  • the allocation among Blackwell Ultra, Rubin and Rubin Ultra GPUs;
  • the AWS regions and data-centre locations involved;
  • the instance types and access mechanisms through which customers will receive the new capacity;
  • prices, power requirements, system counts and cooling requirements;
  • the production timetable for Vera-based infrastructure and the NVHBM integration with Trainium; and
  • the security authorisations and locations for the proposed government infrastructure.

NVIDIA identifies projections about deployment, performance, availability, demand, benefits and technology development as forward-looking statements subject to risks and uncertainties. These include manufacturing and supply constraints, competition, defects, demand changes and regulatory issues. The supplied material contains no detailed assessment of how power, cooling, networking or manufacturing capacity could affect delivery.

More compute capacity may make AI workloads easier to run, but the announcement does not establish that additional GPUs will by themselves produce more capable AI systems or specific advances in agentic or physical AI. The most meaningful evidence will come from actual deployment: customer availability during 2027–2028, independently measured performance, pricing, the production status of the Trainium integration and results from physical-AI systems operating outside controlled simulations.

Sources