AMD, Cisco and HUMAIN said on August 31 that an AI infrastructure deployment using AMD Instinct MI355X GPUs is live in Saudi Arabia and serving HUMAIN customers in the Kingdom and beyond. The production system combines AMD accelerators, AMD EPYC processors and Cisco networking infrastructure.

The announcement marks a move from a planned collaboration to what the companies describe as an operating, customer-serving deployment. It does not disclose the number of GPUs, the system’s total power capacity, its location, customer identities, workloads or independent performance results.

The companies also say their joint venture remains on track to expand to up to 250 megawatts (MW) of AI infrastructure beginning in 2027, with capacity expected to start coming online in the second half of that year. Its longer-term target is up to 1 gigawatt (GW) by 2030.

Contents

What changed

The live deployment is designed to support HUMAIN’s GPU-as-a-service offering. In this model, customers access computing capacity remotely instead of buying and operating their own GPU servers.

The stated use cases include AI model training and inference. Training uses accelerator hardware to adjust a model’s parameters using large datasets. Inference is the later execution of a trained model to produce predictions or responses. These workloads can have different requirements for memory, compute capacity, latency and utilisation.

HUMAIN is described as a Public Investment Fund, or PIF, company providing full-stack AI capabilities. Those capabilities include data centres, infrastructure and cloud platforms, AI models and AI solutions. In the arrangement described by the companies, HUMAIN operates the service, AMD supplies the compute platforms and software, and Cisco supplies networking and related infrastructure.

The deployment is therefore more than an announcement of individual accelerator hardware. It is an integrated infrastructure and service offering intended to provide managed access to AI computing capacity.

How the infrastructure fits together

The production system uses AMD Instinct MI355X GPUs alongside AMD EPYC CPUs. GPUs are used as AI accelerators because they can perform many mathematical operations in parallel, while CPUs handle general-purpose processing and coordinate other parts of a server and data-centre system.

The GPUs are connected through an AI-optimised network fabric built on Cisco’s N9000 Series platform. AMD says the platform is based on Cisco Silicon One and Cisco 800G optics. In a distributed AI system, networking connects accelerator nodes so that data and intermediate computations can move between GPUs. Bandwidth and latency can affect how effectively the overall system uses its processors.

For the planned expansion, the companies say they will use AMD Instinct MI400 Series GPUs, AMD EPYC CPUs and AMD ROCm open software, together with Cisco networking and related infrastructure. The announcement does not identify the specific MI400 products or provide their final system configurations.

The companies describe the proposed platform as combining open models, open software and locally operated infrastructure. They say this approach is intended to give governments, enterprises, research institutions and developers greater control over data location, model customisation, deployment and governance.

Why local GPU capacity matters

AI organisations often need large amounts of accelerator capacity for model development, training and deployment. A locally operated service could give Saudi Arabian and regional organisations access to that capacity without relying entirely on data centres located elsewhere.

This is also a data-governance issue. Data sovereignty generally refers to keeping data and related processing under the physical, legal or organisational control required by a jurisdiction or customer. The companies present local operation as a way to provide greater control over where data is stored and processed, how models are adapted and how deployments are governed.

The project also illustrates the infrastructure stack required for large-scale AI services. Accelerators alone are not sufficient: a production platform also requires servers, high-speed networking, software, cooling, power systems, operations and customer-facing services.

However, the announcement does not establish how broadly the platform is available, which customers can use it or what governance conditions will apply. It provides no pricing, service-level commitments, onboarding details, supported-model list or public access terms.

What to watch

The immediate milestone will be whether the planned expansion begins in 2027 and whether capacity starts coming online during the second half of that year as stated.

Useful measures of progress will include the number of deployed GPUs, usable compute capacity, data-centre locations, customer workloads, utilisation, reliability and energy efficiency. Independent performance data will be particularly important for assessing how the MI355X-based deployment performs in real workloads and how future MI400-based systems compare.

The longer-term test is whether the joint venture reaches its stated target of up to 1 GW by 2030. The availability and terms of HUMAIN’s GPU-as-a-service offering will also determine whether the project becomes broadly useful to governments, businesses, research institutions and developers rather than remaining limited to a small group of customers.

Sources