NVIDIA said on September 10, 2026, that d-Matrix’s next-generation Raptor AI-inference XPUs will be integrated with NVIDIA NVLink Fusion, a platform intended to connect third-party processors to NVIDIA’s rack-scale AI infrastructure. The planned configuration includes NVLink scale-up networking, Spectrum-X Ethernet networking, the NVIDIA MGX rack architecture and related compute, cooling, power and software components.
NVIDIA describes this as a planned integration, and the announcement provides no evidence of a named or completed customer deployment. It does not disclose Raptor’s architecture, performance, power consumption, price or release date.
Contents
What is changing
d-Matrix is an AI-inference chipmaker. Its Raptor products are described as XPUs, a general term for processors or accelerators designed for particular workloads. In this case, the processors are specialised for inference: running trained AI models to produce outputs, rather than training those models by adjusting their parameters.
NVIDIA says Raptor XPUs will connect to its AI infrastructure platform through NVLink Fusion. The platform is presented as a way to integrate custom XPUs and CPUs with:
- NVLink scale-up networking;
- the NVIDIA MGX rack architecture;
- Spectrum-X scale-out networking;
- NVIDIA CPUs, GPUs, SuperNICs and DPUs; and
- associated storage, security, power, cooling and software components.
The stated change is therefore broader than adding an interconnect to a processor. NVIDIA is positioning NVLink Fusion as a rack-level integration path through which a company such as d-Matrix can contribute custom silicon while using an existing infrastructure design.
The planned d-Matrix configuration is also intended to operate alongside NVIDIA GPU systems, including the Vera Rubin NVL72 platform, for what NVIDIA calls disaggregated inference. In such a setup, different parts of an inference workload can be distributed across different processor types or systems instead of requiring every stage to run on one accelerator architecture.
How the planned system fits together
The distinction between scale-up and scale-out networking is important.
Scale-up networking connects processors within a tightly coupled system or rack. It is intended to provide the bandwidth and latency characteristics needed when processors share data within one compute domain. Scale-out networking connects separate systems or racks, allowing a larger installation to operate across multiple units.
According to NVIDIA, d-Matrix plans to connect its XPUs into a single high-bandwidth, low-latency scale-up domain using NVLink Fusion. The broader design would then use Spectrum-X Ethernet for scale-out communication between systems or racks.
NVIDIA’s description also names Vera CPUs, ConnectX-9 SuperNICs and BlueField-4 DPUs as planned components of the integration.
NVIDIA describes NVLink Fusion as supporting Arm, x86 and RISC-V CPU architectures alongside NVIDIA GPUs. It also presents the platform as covering more than data transfer, including the rack’s compute, networking, storage, security, power, cooling and software layers.
NVIDIA says sixth-generation NVLink can provide 3 TB/s of all-to-all bandwidth per XPU, three-times lower XPU-to-XPU latency than off-the-shelf Ethernet and ten-times higher packet rates. Bandwidth measures how much data can be transferred over time, while latency measures communication delay; a higher bandwidth figure does not by itself establish lower latency or better inference performance.
The announcement does not provide the test conditions, workload, system configuration or independent measurements behind those figures.
Why the integration matters
A specialised accelerator is only one part of a data-centre system. Operators also need a way to connect processors, move data between racks, manage infrastructure workloads, supply power, remove heat and run compatible software. A custom-chip company that has to develop every one of those layers itself faces a larger integration task than a company supplying the processor alone.
If the planned integration works as described, d-Matrix’s Raptor silicon could be deployed within a rack architecture that also supports NVIDIA CPUs, GPUs and networking hardware. For data-centre operators, the stated goal is a common infrastructure approach rather than a separate rack design for every processor type.
NVIDIA characterises this model as a semi-custom AI factory: NVIDIA supplies the surrounding infrastructure platform while a partner contributes differentiated XPU silicon. The arrangement could allow d-Matrix to retain a custom processor design while using NVIDIA’s existing rack, networking, software and supply-chain components.
The source presents reduced integration time, cost and deployment risk as intended benefits. Those benefits remain claims about the planned platform, however; the announcement supplies no independent schedule, cost estimate or deployment evidence.
The design is also positioned as complementary to NVIDIA’s own GPU systems, not as a demonstrated replacement for them. The potential value would depend on how effectively Raptor XPUs handle particular inference workloads, how software distributes work across the different processors and whether the complete system offers advantages in throughput, latency, energy use or cost.