d-Matrix adopts NVIDIA NVLink Fusion for rack-scale XPU deployment
d-Matrix, a maker of inference-focused AI chips, will connect its next-generation Raptor XPUs to NVIDIA's AI infrastructure platform using NVLink Fusion, NVIDIA said in a blog post. The move gives d-Matrix's chips a direct path into NVIDIA's rack-scale ecosystem rather than requiring the company to build its own proprietary interconnect and rack infrastructure from scratch.
What's new
NVIDIA confirmed the integration directly: "d-Matrix will connect its next-generation Raptor XPUs to NVIDIA's AI infrastructure platform with NVLink Fusion." NVLink Fusion is NVIDIA's program for letting third-party silicon plug into its rack-scale systems, and the d-Matrix integration runs through two specific pieces of that stack:
- MGX rack architecture — NVIDIA's modular rack design, which standardizes power, cooling, and mechanical specs so partner chips can be deployed in existing rack designs rather than bespoke ones.
- Spectrum-X networking — NVIDIA's Ethernet-based AI networking platform, paired here with sixth-generation NVLink for the chip-to-chip links inside the rack.
On performance, the companies cite three figures for what sixth-gen NVLink delivers over commodity alternatives: 3x lower XPU-to-XPU latency than off-the-shelf Ethernet, 10x higher packet rates, and 3 TB/s of per-XPU all-to-all bandwidth.
d-Matrix CEO Sid Sheth framed the benefit as speed-to-market: the arrangement gives customers "a faster, lower-risk path to deploy and scale ultralow-latency inference" by letting d-Matrix skip building its own rack-scale infrastructure and instead build on NVIDIA's existing, proven platform.
Context
d-Matrix has positioned itself as a specialist in low-latency inference hardware, competing in a crowded field of inference-focused chip startups (alongside companies like Groq, Cerebras, and SambaNova) that argue purpose-built silicon can beat general-purpose GPUs on latency and cost-per-token for serving models at scale. Until now, going after that market meant either integrating with a hyperscaler's existing infrastructure piecemeal or building proprietary rack-scale networking in-house — a costly, slow path for a chip startup.
NVLink Fusion itself is NVIDIA's answer to a specific competitive pressure: as more chip startups and hyperscalers design custom silicon (XPUs, TPUs, custom ASICs) to reduce dependence on NVIDIA GPUs, NVIDIA has opened parts of its interconnect and rack stack to those same competitors' chips. It's a hedge — even if a customer chooses non-NVIDIA compute, NVIDIA can still capture the networking and rack-infrastructure layer around it.
Why it matters
For d-Matrix, this is a distribution and credibility shortcut. Rack-scale AI infrastructure is expensive and slow to build from zero, and enterprise buyers are wary of unproven interconnects at scale. By adopting NVIDIA's MGX and Spectrum-X stack, d-Matrix can offer customers a rack-scale deployment path that looks and operates like the infrastructure they already trust, while still differentiating on its own chip design.
For NVIDIA, each NVLink Fusion partner reinforces the strategy of owning the plumbing regardless of whose chips end up in the rack. If d-Matrix's Raptor XPUs gain traction with inference customers, NVIDIA still captures value through the interconnect and rack architecture — insulating its infrastructure business from any erosion in raw GPU share. It's a small but telling data point in a broader trend: as the inference chip market fragments, the fight is increasingly over who controls the rack, not just who makes the fastest individual chip.
Corroborating sources
- Blogs.nvidia
https://blogs.nvidia.com/blog/d-matrix-nvlink-fusion/
“d-Matrix will connect its next-generation Raptor XPUs to NVIDIA’s AI infrastructure platform with NVLink Fusion, joining a growing ecosystem of partners using NVIDIA’s rack-scale architecture to accelerate time to market.”