AWS & NVIDIA Expand Partnership Across AI Infrastructure
AWS and NVIDIA are deepening integration across the AI stack, bringing Vera CPUs, advanced networking, Nemotron open models, and physical AI technologies to AWS as customer demand accelerates.
You're reading Entrepreneur India, an international franchise of Entrepreneur Media.
Amazon Web Services (AWS), and NVIDIA announced a major expansion of their strategic collaboration to meet surging global demand for AI infrastructure as demand continues to accelerate. Building on already-rapid customer adoption of NVIDIA-accelerated compute on AWS, the companies plan to deploy 2 million additional NVIDIA GPUs across AWS’s global infrastructure and deepen their work together across AI factories, CPUs, networking, open models, data processing and robotics.
This will deliver co-engineered AI solutions that enable customers to accelerate AI development and deployment at unprecedented scale. Customers are moving from pilot to production and scaling workloads across agentic AI, scientific discovery, enterprise automation and robotics.
AI workloads are scaling at a swift pace, from how models are trained and run, to how data is processed, indexed and used to power intelligent applications. They need broader model choice, faster data pipelines and new capabilities for emerging use cases like physical AI. They also need confidence that the underlying infrastructure can keep pace with their own ability to innovate while maintaining the highest level of security and reliability for mission-critical workloads.
“Customers want the freedom to choose the best tools for their AI workloads, and they want confidence that everything works seamlessly together,” said Matt Garman, CEO of AWS. “That’s why we’ve invested deeply with NVIDIA to make AWS the best place to run NVIDIA AI technologies, optimizing performance across our infrastructure from networking and security to deployment. This expanded collaboration gives frontier labs, enterprises and governments even more ways to build and deploy AI on AWS.”
“NVIDIA and AWS have built one of the great growth engines of the AI era, and demand is running ahead of every forecast,” said Jensen Huang, founder and CEO of NVIDIA. “For 16 years, we have scaled NVIDIA computing in the cloud together. Now, we are expanding our partnership across the full stack, GPUs, CPUs, networking, open models and software, to make agentic and physical AI real at an unprecedented pace and scale that only AWS and NVIDIA can deliver. This expansion reflects customers’ demand for NVIDIA’s platform on AWS.”
Additionally, NVIDIA introduced NVIDIA NVHBM, a next-generation high-bandwidth memory technology developed in collaboration with leading memory partners that brings higher memory performance and efficiency to XPUs. NVHBM will deliver up to 30 per cent greater memory bandwidth and 15 per cent lower HBM power consumption, and frees up to 25 per cent more area on XPU compute die compared to standard HBM4E. Amazon’s Annapurna Labs will be the first to work on NVHBM as part of its broader collaboration with NVIDIA around NVLink Fusion
AI factories must evolve to handle ever‑larger models and increasingly complex reasoning workloads. To meet the insatiable compute demands, hyperscalers and AI‑native firms are building custom accelerators, or XPUs. Scaling these at data‑center level requires high‑bandwidth memory to keep compute fed, ample silicon and packaging area, efficient power delivery, and a resilient supply chain, all underpinned by rack‑scale architectures that seamlessly integrate XPUs into modern infrastructure.
AWS–NVIDIA alliance underscores how AI infrastructure has become the defining growth engine of this era. AWS and NVIDIA are not just keeping pace with surging demand, they are delivering choice, speed, and reliability for mission‑critical workloads. This collaboration signals a new phase where AI factories and intelligent applications can scale seamlessly, backed by infrastructure designed to match the speed of innovation itself.
Amazon Web Services (AWS), and NVIDIA announced a major expansion of their strategic collaboration to meet surging global demand for AI infrastructure as demand continues to accelerate. Building on already-rapid customer adoption of NVIDIA-accelerated compute on AWS, the companies plan to deploy 2 million additional NVIDIA GPUs across AWS’s global infrastructure and deepen their work together across AI factories, CPUs, networking, open models, data processing and robotics.
This will deliver co-engineered AI solutions that enable customers to accelerate AI development and deployment at unprecedented scale. Customers are moving from pilot to production and scaling workloads across agentic AI, scientific discovery, enterprise automation and robotics.
AI workloads are scaling at a swift pace, from how models are trained and run, to how data is processed, indexed and used to power intelligent applications. They need broader model choice, faster data pipelines and new capabilities for emerging use cases like physical AI. They also need confidence that the underlying infrastructure can keep pace with their own ability to innovate while maintaining the highest level of security and reliability for mission-critical workloads.