Skip to content
World clockEU--:--UK--:--USA--:--CN--:--PLDEFRIT中文EN

portal about AI and technologyevents · analysis · interviews · technical background

Search
LIVE
›

How AI accelerators are built: the packaging, memory and design bottlenecks

The AI accelerator race is no longer just about who can print the fastest logic die. Three supply chain stories from 2026 show the constraints sit in the material, the memory stack and the engineering workflow around the chip.

TechnologyExplainerRachel NwosuPublished: 27 September 20267 min readSources 4
How AI accelerators are built: the packaging, memory and design bottlenecks

AI accelerators are usually described by their compute. Nvidia's GPUs, Microsoft's Maia, d-Matrix's Raptor: the numbers that get quoted are teraflops, transistor counts and memory bandwidth. The harder story is what surrounds that silicon. In 2026 that story got crowded.

Three separate reports, from Tom's Hardware, ServeTheHome and The Next Platform, describe three different chokepoints. A specialist glass cloth sits inside every advanced package. A memory architecture trades capacity for bandwidth. An engineering pipeline is itself being rebuilt around AI agents.

A 90% share in a material most people have never heard of

The first bottleneck is a material called T-glass. According to Tom's Hardware, a Japanese company called Nittobo controls roughly 90% of the global supply of it. T-glass is a low-CTE glass-fiber cloth used in the organic core of IC substrates, the interconnect layer between a chip and its printed circuit board. It keeps large, high-heat packages dimensionally stable as they get denser.

That matters more as AI processors grow. Tom's Hardware cites Nvidia data showing interposer sizes rising from 814mm² for Hopper to 1,700mm² for Blackwell, an increase of 109%, with the forthcoming Rubin and Feynman generations scaling further still. Each generation consumes more T-glass per unit. And a new T-glass line is not something anyone can spin up quickly: it needs specialised electric melting furnaces running between 1,600 and 1,700°C, plus years of investment and expertise to melt silica-rich glass, spin it into yarn and weave it into ultrathin cloth.

The numbers already show the squeeze. Tom's Hardware reports prices have risen between 20 and 30%, while lead times for downstream materials such as copper-clad laminates have stretched from a normal 8 to 10 weeks to beyond 20. "With T-glass supply even more constrained now, suppliers are no longer providing lead times," said Bill Ho, an analyst at Yuanta, in the article.

Nittobo is tripling capacity at its Fukushima plant, but the new supply will not reach the market until mid-2027. It is also doubling raw yarn capacity at its Taiwan plant and importing yarn back to Japan for cloth manufacturing. The company has struck a collaboration deal with Nanya Plastics to outsource some weaving: by 2027, roughly 20% of Nittobo's glass cloth is expected to be woven by Nanya, according to the company's disclosure. Partnering with a competitor is itself a signal of how tight the market is.

Bilal Hachemi, an analyst at Yole Group who tracks the IC substrate supply chain, told Tom's Hardware Premium that replacing T-glass is not easy because it "has specific dielectric and CTE values that work better for the AI chips, especially for the organic core." He also pointed to thin industry margins: "Any increase in demand for build-up materials, ABF material, or T-glass can cause potential shortage, because it's against the basics of this industry," he said.

Bank of America estimates Nittobo's electronic materials segment will nearly double sales from ¥40.9 billion ($266 million) in 2025 to ¥87.7 billion by March 2028, with operating margins approaching 48%. "Demand for T-glass cloth seems likely to grow more than originally expected," said Takashi Enomoto, a research analyst at Bank of America, in the same piece.

The most striking detail may be who is now knocking on the door. Hachemi said Nvidia reaching out directly to an upstream material supplier like Nittobo is unprecedented. "For the first time, we are seeing Nvidia, the end customer, reaching out to the upstream material suppliers to secure the capacity and make sure they will get it," he said. The worry follows directly: once the largest buyer locks down its share, rival chip buyers compete for what is left.

Memory is the other wall

The second bottleneck is inside the accelerator package. At Hot Chips 2026, d-Matrix presented its Raptor 3D-DRAM accelerator for generative inference, and ServeTheHome's write-up lays out the trade-offs cleanly. Model weights keep growing, and the KV cache scales with context length multiplied by batch size. ServeTheHome gives the example of 64 users at 1M context, which can mean roughly 935 GB of KV cache. That is both a capacity problem and a bandwidth problem.

SRAM hits the bandwidth target but holds almost nothing: a Corsair SRAM accelerator card pair reaches roughly 300 TB/s at about 1 ns latency, yet holds only about 4 GB. HBM solves capacity but struggles on bandwidth, with d-Matrix citing a practical ceiling around 20 TB/s for HBM4 packages such as Nvidia's Vera Rubin and AMD's Instinct MI455. Pushing 100 TB/s through HBM at 2.4 pJ/bit would consume about 1.92 kW before any fabric traffic, according to ServeTheHome.

d-Matrix's answer is to stack compute directly on top of DRAM dies. Vertical 3D IO comes in at around 0.3 to 0.4 pJ, roughly 10 times lower than HBM, because it is a millimetre-scale path rather than a centimetre-scale interposer route. The company uses a TSMC N4 logic die stacked on a 3D DRAM die with 36 um face-to-face stacking. At 32GB per card, with 4-bit weights and an 8-bit KV cache, a 72-card scale-up can host a frontier model such as Kimi K3 at 1M context, per ServeTheHome.

None of this is free. d-Matrix says a 1-Hi logic-on-top stack at no more than 0.5 W/mm² can be liquid cooled and keep DRAM under 100 C. It also had to solve an awkward arithmetic problem: each tensor engine needs a 128B flit per access, but its die has 840 banks, 768 after 72 spares, spread across 256 channels, which works out to 3 banks per channel. A single access returns 96B, so a 128B flit takes two accesses and fetches 192B, wasting about 33% of bandwidth near 33 TB/s. The fix, stream blocking, shares one partial 32B access across three flits so 4 accesses at 96B feed 3 flits at 128B.

The design loop itself is being automated

The third bottleneck is not a material or a memory type but the engineering process. The Next Platform reported on 27 July that Nvidia is pushing AI agents into chip engineering, quoting Tim Costa, vice president and general manager of computational engineering at Nvidia, who told journalists that by 2030 the industry is expected to produce 2 trillion chips and process about 41 million wafers a month, with individual packages approaching a trillion transistors.

"The key point is not any one number; it's the interaction of scale, architecture, packaging, and system complexity," Costa said. "The traditional design process just can't keep pace with that scale of challenge."

Nvidia is deploying its Arm-based Vera CV100 CPU across the EDA workflows used to create its future CPUs and GPUs, including simulation, formal verification and physical implementation. According to Costa, early engineering testing shows Vera running on Synopsys' VCS and Cadence's Jasper platforms delivering 1.5 times the performance of AMD's Epyc Torrent systems. Vera holds 88 custom Olympus CPU cores and a 1.2 TB/sec LPDDR5X memory subsystem. Nvidia's next-generation CPU, Rosa, is due in 2028 as part of the Feynman datacenter platform.

The tool vendors are making their own claims. Cadence said its AI Super Agents can implement hundreds of simulations in less than a day, work that currently takes five weeks, providing 40-times faster Register-Transfer Level validation cycles. Synopsys demonstrated a Fully Autonomous Design Verification Workflow at the Design Automation Conference and said it can deliver validated RTL 50 times faster than other platforms.

Treat those multiples with care: they come from the vendors selling the tools, and The Next Platform notes the broader industry challenge of orchestration, integration and guardrails, with AI-generated results still needing verification before sign-off.

Microsoft's Maia 200, presented at Hot Chips 2026 and covered by ServeTheHome, shows what the resulting designs look like. It is a 3nm chip with 140 billion transistors, 6 stacks of HBM3e, 7TB/second of HBM bandwidth, a 750 Watt TDP, 10,000 TFLOPS of FP4 performance and an 820mm² SoC die. Microsoft uses a fully connected quad topology with no scale-out networking, just scale-up: its slide shows 128 racks with 6,000 chips. The company also co-designed the software stack, with the Microsoft Collective Communication Library at the core, because the hardware's software-defined dataflow requires kernels tuned to it.

Put the three stories together and the picture is consistent. The accelerator is no longer a chip. It is a package that depends on a near-monopoly material supplier, a memory architecture that has to be co-designed with the workload, and a design pipeline that is being rebuilt with AI to keep up with its own complexity. Each layer has its own lead times, and the slowest one sets the pace.

Comments 0

Sources

4
  1. 01Shortages of crucial chip packaging material threatens AI accelerator supply chainsEN
  2. 02d-Matrix Raptor 3D-DRAM Accelerator for Generative Inference at Hot Chips 2026EN
  3. 03Microsoft's Maia 200 AI Accelerator at Hot Chips 2026EN
  4. 04Nvidia Accelerates Chip Engineering With AI AgentsEN

All figures and quotations in this text come from the sources listed below.

Content prepared by the editorial team with AI assistance.

Rachel Nwosu

Rachel Nwosu

AI, models and technology

Rachel Nwosu covers AI, models and technology for FLASH24, working from public model documentation, benchmark releases and repository histories rather than press summaries, and she skips announcements that arrive without reproducible numbers. She checks training-data claims against dataset cards and reruns reported metrics where code is available. She spends much of her week interviewing researchers and engineers, tracking model launch calendars, and comparing vendor benchmarks with independent evaluations. Outside the desk she runs 3D printers, restores old computers, and tests how models learn from internet junk. She does not publish benchmark figures she cannot trace to a source.

Newsroom →

Comments

0
  1. No comments yet — be the first.

Write a comment

Comments are public. We do not publish abuse, spam or advertising.