Skip to content
AI & Hardware

d-Matrix Is Building an AI Chip That Plugs Straight Into NVIDIA’s AI Racks

d-Matrix is bringing NVIDIA NVLink Fusion into its next-generation Raptor inference accelerator, giving the chip a direct path into NVIDIA-connected rack-scale AI systems. Raptor combines d-Matrix's 3D in-memory compute architecture with RISC-V orchestration, while NVLink Fusion extends that design into a broader scale-up infrastructure.

Official d-Matrix accelerator hardware image from the company's media kit.

01Raptor is being designed for NVIDIA-connected AI racks

d-Matrix is taking its next-generation Raptor inference accelerator beyond the accelerator card itself.

On September 10, the company said Raptor will adopt NVIDIA NVLink Fusion, giving the custom AI chip a direct path into NVIDIA-connected data-center systems.

The move is important because NVLink Fusion is not just a cable or a standalone interconnect.

NVIDIA built it as a platform for custom CPUs and XPUs that need to participate inside rack-scale AI infrastructure. By adopting that platform, d-Matrix is designing Raptor to become one component inside a larger heterogeneous system rather than an isolated accelerator.

That turns the story from a new chip announcement into a rack-level architecture story.

03Raptor starts with a very different memory architecture

The core of Raptor is d-Matrix’s 3DIMC architecture.

3DIMC stands for 3D In-Memory Compute. Instead of treating memory and compute as distant parts of the accelerator, d-Matrix is stacking DRAM and compute much more tightly together.

The company first validated the approach through its Pavehawk test silicon and plans to bring the technology into a commercial accelerator with Raptor.

Raptor is also chiplet-based, which lets d-Matrix combine multiple specialized pieces inside one accelerator design.

The result is a processor built around fast movement between memory and inference compute, with the memory system treated as a first-class architectural component rather than a supporting block.

04RISC-V becomes the orchestration layer inside Raptor

Raptor also includes a separate control and orchestration layer.

d-Matrix selected Andes’ AX46MPV, a 64-bit Linux-capable RISC-V processor, for the architecture.

According to the companies, the RISC-V CPU will manage workload distribution, memory coordination and runtime control across Raptor’s compute fabric. It can also handle supporting vector operations around the inference pipeline.

That gives Raptor multiple levels of specialization.

The 3DIMC layer focuses on the memory-compute path. The RISC-V cores coordinate the system. The chiplet structure connects those pieces into the accelerator.

NVLink Fusion then extends the design one level higher, connecting the accelerator into rack-scale infrastructure.

05The architecture now stretches from silicon to the rack

Put those layers together and the shape of the system becomes clearer.

At the lowest level, Raptor combines compute and stacked DRAM.

Above that, RISC-V processors coordinate work across the accelerator.

At the device level, chiplets form the larger Raptor package.

Then NVLink Fusion provides a path from the custom accelerator into NVIDIA-connected scale-up infrastructure.

That is a much broader design space than building a standalone PCIe card.

Each layer has a different job, but the full system is designed around one continuous inference path from local memory access to rack-scale communication.

06d-Matrix has already been moving toward heterogeneous inference

The NVLink Fusion announcement also fits d-Matrix’s wider infrastructure direction.

Earlier this year, Parasail announced that it was deploying d-Matrix Corsair accelerators alongside NVIDIA Hopper and Blackwell systems.

In that deployment, different processors can be assigned different parts of an inference workflow while operating inside the same broader service.

Raptor pushes that idea further into the hardware architecture itself.

Instead of only placing separate accelerator types next to each other in a data center, the next generation can connect more directly through a common rack-scale fabric.

That makes heterogeneous inference less about separate islands of compute and more about coordinated infrastructure.

07The next milestone is a 2027 rack-scale deployment path

According to Reuters, d-Matrix expects Raptor to complete its final design stage by the end of 2026.

NVIDIA-compatible rack systems using the new integration are expected in 2027.

d-Matrix is also working with Astera Labs on connectivity around the broader system, adding another infrastructure layer around high-speed data movement.

The timing matters because Raptor is still being built as a next-generation platform.

Today’s announcement is therefore about the architecture that d-Matrix wants ready before those systems arrive: specialized inference silicon, tightly integrated memory, local orchestration and a direct scale-up path into NVIDIA-connected racks.

08The Upgrade Feeling

Raptor shows how the definition of an AI chip is expanding.

The accelerator itself still matters, but the surrounding system now matters just as much.

d-Matrix is combining its own 3D memory-compute architecture with RISC-V control, chiplets and NVIDIA’s rack-scale connectivity layer.

That creates a stack where each company contributes a different piece without forcing every part of the system into one processor architecture.

The upgrade is the connection between those layers.

Raptor is not only being designed to run inference. It is being designed to arrive already connected to the kind of rack-scale infrastructure where modern AI services are increasingly built.

Get smarter updates

Straight to your inbox. No spam. Only useful tech.

Discover more from That Upgrade Feeling

Subscribe now to keep reading and get access to the full archive.

Continue reading

Discover more from That Upgrade Feeling

Subscribe now to keep reading and get access to the full archive.

Continue reading