Skip to content
World clockEU--:--UK--:--USA--:--CN--:--PLDEFRIT中文EN

portal about AI and technologyevents · analysis · interviews · technical background

Search
LIVE
›

AI accelerator supply chain: T-glass, 3D DRAM and the packaging bottleneck

AI accelerators are running into a supply chain most data centre operators have never heard of: T-glass, a specialist glass-fibre cloth made almost entirely by one Japanese company, Nittobo, which controls roughly 90% of global supply according to Tom's Hardware.

TechnologyExplainerRachel NwosuPublished: 28 September 20266 min readSources 4
AI accelerator supply chain: T-glass, 3D DRAM and the packaging bottleneck

Tom's Hardware reported on 9 March that demand for T-glass is so far ahead of supply that prices have risen between 20 and 30%. Lead times for downstream materials such as copper-clad laminates have stretched from the usual 8 to 10 weeks to beyond 20. Nittobo is tripling capacity at its Fukushima plant, but the extra material will not reach the market until mid-2027.

What is T-glass? It is not a chip. It is a low-CTE glass cloth used in the organic core of IC substrates, the interconnect layer between a processor and its printed circuit board. It keeps large, hot packages dimensionally stable, and as AI processors get bigger and run hotter, that job gets harder.

E-glass does a similar job more cheaply, but it is mainly used in lower-end chips such as microcontrollers and older mobile processors. For 2.5D and 3D packaging, and therefore for advanced AI accelerators, T-glass is the preferred option. According to Tom's Hardware, Nvidia's own data shows interposer sizes growing from 814mm² on Hopper to 1,700mm² on Blackwell, a 109% increase, with the Rubin and Feynman generations scaling further.

Why substitution is hard

Bilal Hachemi, an analyst at Yole Group who tracks the IC substrate supply chain, told Tom's Hardware Premium that replacing T-glass is not straightforward. It "has specific dielectric and CTE values that work better for the AI chips, especially for the organic core," he said. Hachemi also noted that the IC substrate industry has historically run on thin margins, meaning even modest demand surges can trigger shortages. "Any increase in demand for build-up materials, ABF material, or T-glass can cause potential shortage, because it's against the basics of this industry," he said.

Bill Ho, an analyst at Yuanta, put the state of the market bluntly: "With T-glass supply even more constrained now, suppliers are no longer providing lead times."

The scramble for allocation has already changed behaviour. Hachemi said Nvidia reaching out directly to an upstream material supplier is unprecedented: "For the first time, we are seeing Nvidia, the end customer, reaching out to the upstream material suppliers to secure the capacity and make sure they will get it." The concern for rival chip buyers is what happens once Nvidia has locked down its share.

Nittobo is not standing still. Beyond the Fukushima expansion, it is doubling raw yarn capacity at its Taiwan plant and importing yarn back to Japan for weaving. It has also struck a collaboration deal with Nanya Plastics to outsource some weaving, partnering with one of its biggest competitors to ease the bottleneck. By 2027, roughly 20% of Nittobo's glass cloth is expected to be woven by Nanya, according to Tom's Hardware.

Bank of America estimates Nittobo's electronic materials segment will nearly double sales from ¥40.9 billion ($266 million) in 2025 to ¥87.7 billion by March 2028, with operating margins approaching 48%. Takashi Enomoto, a research analyst at the bank, said demand for T-glass cloth "seems likely to grow more than originally expected", adding that attention had been on thick T-glass for GPU and CPU packages but that ultra-thin T-glass demand is now likely to rise as leading-edge devices shift away from E-glass.

The memory side of the same problem

Packaging material is only one constraint. At Hot Chips 2026, d-Matrix presented Raptor, an accelerator that stacks compute directly on top of DRAM dies, according to ServeTheHome's write-up published on 14 September. The argument starts with arithmetic: model weights keep growing, and the KV cache scales with context length multiplied by batch size, so 64 users at 1M context can mean roughly 935 GB of KV cache.

SRAM meets the bandwidth target but only at tiny scale. ServeTheHome reports that a pair of Corsair SRAM accelerator cards reaches roughly 300 TB/s at about 1 ns latency but holds only about 4 GB, because a 6T SRAM cell is around 10 times larger than a DRAM cell and leakage runs to tens of watts at GB scale. HBM solves capacity but struggles on bandwidth, with d-Matrix citing a practical ceiling around 20 TB/s for HBM4 packages such as Nvidia's Vera Rubin and AMD's Instinct MI455.

Bandwidth at that level carries a power cost. At 2.4 pJ/bit, pushing 100 TB/s through HBM consumes about 1.92 kW before any fabric traffic. d-Matrix's alternative puts a TSMC N4 logic die on top of a 3D DRAM die using 36 µm face-to-face stacking. Vertical 3D IO comes in at around 0.3 to 0.4 pJ, roughly 10 times lower than HBM.

Thermals are the catch. Hundreds of watts must escape through TSVs in a temperature-sensitive DRAM stack. d-Matrix says a 1-Hi logic-on-top stack at no more than 0.5 W/mm² can be liquid cooled and keep DRAM under 100 C. At 32GB per card, with 4-bit weights and an 8-bit KV cache, a 72-card scale-up can host a frontier model such as Kimi K3 at 1M context, according to the company's Hot Chips presentation.

Microsoft used the same conference to detail Maia 200, its second-generation accelerator for Azure. ServeTheHome reported on 26 August that the chip is built on TSMC 3nm with 140 billion transistors, six stacks of HBM3e, 7TB/second of HBM bandwidth, a 750 Watt TDP and 10,000 TFLOPS of FP4 performance. The SoC die is 820mm². Microsoft's networking is scale-up only, with 128 racks and 6,000 chips in the slide shown, using tier-0 switches inside racks and tier-1 switches between them.

China's parallel track

SemiEngineering's weekly review, published on 25 September, lists several moves that point in the same direction: more domestic supply and more packaging capacity. Alibaba unveiled the Zhenwu V900 AI accelerator with 216GB of memory and 1,200GB/s inter-chip bandwidth, claiming 3× the performance of its M890 predecessor, with mass production planned for Q1 2027. CXMT said its fifth-gen G5 DRAM entered mass production alongside a new LPDDR5X line, claiming an 11.95nm active-area half-pitch and at least 50% more dies per wafer than the previous generation, though it did not disclose actual manufacturing yields.

The same review notes that Samsung reportedly plans to at least double HBM4-family output in 2027, according to Seoul Economic Daily, and that GUC said its HBM4E PHY and controller are design-ready on TSMC N2P and have already been adopted in customer AI ASICs. Imec researchers modelled a 3D DRAM-on-GPU architecture with vertically oriented DRAM dies and cooling cavities, cutting peak temperature from 121.7°C to 103.4°C against an HBM-on-GPU baseline.

None of this fixes the T-glass timeline. New melting furnaces operating between 1,600 and 1,700°C, spinning and weaving lines take years to bring up, and Tom's Hardware's reporting puts the first meaningful new volume in mid-2027. Until then, the constraint sits upstream of the fabs, in a material most data centre operators will never see.

Comments 0

Sources

4
  1. 01Shortages of crucial chip packaging material threatens AI accelerator supply chainsEN
  2. 02d-Matrix Raptor 3D-DRAM Accelerator for Generative Inference at Hot Chips 2026EN
  3. 03Microsoft's Maia 200 AI Accelerator at Hot Chips 2026EN
  4. 04Chip Industry Week In ReviewEN

All figures and quotations in this text come from the sources listed below.

Content prepared by the editorial team with AI assistance.

Rachel Nwosu

Rachel Nwosu

AI, models and technology

Rachel Nwosu covers AI, models and technology for FLASH24, working from public model documentation, benchmark releases and repository histories rather than press summaries, and she skips announcements that arrive without reproducible numbers. She checks training-data claims against dataset cards and reruns reported metrics where code is available. She spends much of her week interviewing researchers and engineers, tracking model launch calendars, and comparing vendor benchmarks with independent evaluations. Outside the desk she runs 3D printers, restores old computers, and tests how models learn from internet junk. She does not publish benchmark figures she cannot trace to a source.

Newsroom →

Comments

0
  1. No comments yet — be the first.

Write a comment

Comments are public. We do not publish abuse, spam or advertising.