Skip to content
World clockEU--:--UK--:--USA--:--CN--:--PLDEFRIT中文EN

portal about AI and technologyevents · analysis · interviews · technical background

Search
LIVE
›

The memory wall in AI accelerators: d-Matrix, Microsoft and the 3D DRAM push

AI accelerators are hitting a memory problem that designers increasingly solve by stacking chips vertically. At Hot Chips 2026, d-Matrix argued that its Raptor design, which puts a TSMC N4 logic die on top of a 3D DRAM die, can beat the bandwidth ceiling it says HBM packages reach at roughly 20 TB/s.

TechnologyExplainerRachel NwosuPublished: 27 September 20265 min readSources 2
The memory wall in AI accelerators: d-Matrix, Microsoft and the 3D DRAM push

Model weights keep growing, and the KV cache scales with context length multiplied by batch size. ServeTheHome, reporting from d-Matrix's Hot Chips 2026 presentation on 14 September, lays out the arithmetic: 64 users at 1M context can mean roughly 935 GB of KV cache. Weights and cache together create a problem of both capacity and bandwidth, and both keep growing.

SRAM hits the bandwidth target, but only on a tiny scale. A pair of Corsair SRAM accelerator cards reaches roughly 300 TB/s at about 1 ns latency, yet holds only about 4 GB. A 6T SRAM cell is around 10 times larger than a DRAM cell, and leakage runs to tens of watts at GB scale. That makes SRAM a good fit for a draft model in speculative decoding, not for holding frontier model weights.

HBM solves capacity, not bandwidth

HBM solves the capacity half of the problem but struggles with bandwidth. Pin speed and I/O width per base die improve slowly, and the number of stacks is limited by available package beachfront, roughly 8 to 16 stacks per package. d-Matrix cites a practical bandwidth ceiling around 20 TB/s for HBM4 packages such as the NVIDIA Vera Rubin and AMD Instinct MI455.

Bandwidth that high carries a power price. At 2.4 pJ/bit, pushing 100 TB/s through HBM eats about 1.92 kW before any fabric traffic is counted. According to ServeTheHome's write-up, packages today lack both the beachfront and the power budget to reach SRAM-class bandwidth with HBM.

d-Matrix's answer is to stack compute directly on top of DRAM dies. Stacking creates a thermal challenge, because hundreds of watts must escape through TSVs in a temperature-sensitive DRAM stack. It also creates a power-delivery challenge from IR drop. d-Matrix says a 1-Hi logic-on-top stack at no more than 0.5 W/mm2 can be liquid cooled and keep DRAM under 100 C.

The company places 3D DRAM between the two extremes on an energy ladder. On-die SRAM costs roughly 50 fJ, while 2.5D HBM4 systems run in the 2.5 to 5 pJ range when chip-level energy is included. Vertical 3D IO comes in at around 0.3 to 0.4 pJ, about 10 times lower than HBM, because it is a PHY-less millimeter-scale path rather than a centimeter-scale interposer route. Fewer stacked layers than HBM also means a larger die and better yield.

Decode is where the time goes

Prefill processes many prompt tokens in parallel and is bound by compute throughput. Decode produces one token at a time and is typically bound by memory bandwidth. Attention can flip to compute-bound with high GQA and speculative decoding, and MoE stays memory-bound even at modest batch sizes. Decode is the phase that wants huge bandwidth.

Because decode dominates wall-clock runtime, the memory-bound portion matters most. d-Matrix notes that most inference time is spent in the decode phase, so improving decode bandwidth improves overall inference performance. At 32GB per card, with 4-bit weights and an 8-bit KV cache, d-Matrix sizes the system to fit in one rack. A 72-card scale-up can host a frontier model such as Kimi K3 at 1M context, and disaggregation and multi-rack setups extend beyond a single Raptor rack.

The implementation itself is called Raptor. A TSMC N4 logic die sits on top of a 3D DRAM die using 36 um face-to-face stacking, a process d-Matrix describes as proven, low-cost, high-volume and high-yield. Each tensor engine needs a 128B flit per access, and with 32B delivered per column access from 32B banks, that works out to needing 4 banks per channel. The die has 840 banks, 768 after 72 spares, spread across 256 channels for just 3 banks per channel. With 3 banks per channel, a single access returns 96B, so delivering a 128B flit takes two accesses and fetches 192B, wasting about 33 percent of bandwidth near 33 TB/s.

Stream blocking reclaims that waste. d-Matrix shares one partial 32B access across three flits, so 4 accesses at 96B feed 3 flits at 128B, with 384B in matching 384B out. Overfetch drops to zero and no shifting network is required. Moving 100 TB/s at 0.37 pJ/bit works out to 296 W just for I/O, and conventional DBI could save 20 percent. Stream flipping delivers that 20 percent without the pin, comparing each flit to the previous one and inverting when needed, with a single metadata bit per flit carried alongside ECC.

Microsoft's Maia 200 takes the other route

Microsoft used Hot Chips 2026 to detail its second-generation Maia 200 accelerator, which ServeTheHome covered on 26 August. The 3nm chip has 140 billion transistors, six stacks of HBM3e, 7TB/s of HBM bandwidth and a 750 Watt TDP, with 10,000 TFLOPS of FP4 performance on an 820mm2 SoC die. It uses a Software Defined Local Access dataflow architecture, where the dataflow is set at compile time and data access stays within compute elements. There is no scale-out networking: Microsoft's slide shows 128 racks and 6,000 chips in a scale-up domain over unified Ethernet.

Two different bets, one shared constraint. Whether the fix is stacking DRAM under logic or pairing HBM with a software-defined dataflow, both designs are shaped by the cost of moving bits to and from memory.

Comments 0

Sources

2
  1. 01d-Matrix Raptor 3D-DRAM Accelerator for Generative Inference at Hot Chips 2026EN
  2. 02Microsoft's Maia 200 AI Accelerator at Hot Chips 2026EN

All figures and quotations in this text come from the sources listed below.

Content prepared by the editorial team with AI assistance.

Rachel Nwosu

Rachel Nwosu

AI, models and technology

Rachel Nwosu covers AI, models and technology for FLASH24, working from public model documentation, benchmark releases and repository histories rather than press summaries, and she skips announcements that arrive without reproducible numbers. She checks training-data claims against dataset cards and reruns reported metrics where code is available. She spends much of her week interviewing researchers and engineers, tracking model launch calendars, and comparing vendor benchmarks with independent evaluations. Outside the desk she runs 3D printers, restores old computers, and tests how models learn from internet junk. She does not publish benchmark figures she cannot trace to a source.

Newsroom →

Comments

0
  1. No comments yet — be the first.

Write a comment

Comments are public. We do not publish abuse, spam or advertising.