Skip to content
World clockEU--:--UK--:--USA--:--CN--:--PLDEFRIT中文EN

portal about AI and technologyevents · analysis · interviews · technical background

Search
LIVE
›

Why AI Accelerators Are Getting Harder to Build: T-Glass, HBM and 3D Stacking

A Japanese company called Nittobo controls roughly 90% of the global supply of specialist glass-fiber cloth known as T-glass, and according to Tom's Hardware the squeeze is now hitting AI accelerator supply chains, with prices up 20 to 30% and lead times stretching past 20 weeks.

TechnologyExplainerGrace OkonkwoPublished: 28 September 20266 min readSources 5
Why AI Accelerators Are Getting Harder to Build: T-Glass, HBM and 3D Stacking

Most coverage of AI hardware starts and ends with Nvidia and TSMC. Tom's Hardware reported on 9 March that the bottleneck has moved one layer down, into a material most people have never heard of. T-glass is a low-CTE glass cloth used in the organic core of IC substrates, the interconnect layer between a chip and its printed circuit board.

What T-glass actually does

T-glass keeps large, high-heat chip packages dimensionally stable. AI processors are growing bigger and running hotter, and keeping them flat is now a manufacturing problem rather than a theoretical one. Another glass fiber, E-glass, is cheaper, but it mainly goes into lower-end chips like microcontrollers and older mobile processors. For massive 2.5D and 3D packaging, T-glass is the preferred option.

Nittobo dominates that market. Neither it nor any other company can spin up a new production line overnight. The material requires specialised electric melting furnaces running at 1,600 to 1,700°C. The process involves melting silica-rich glass, spinning it into yarn and weaving it into an ultrathin cloth. Scaling that takes years of investment and expertise.

Nittobo is tripling capacity at its Fukushima plant in Japan, but the new supply will not reach the market until mid-2027. In the meantime, prices have risen between 20 and 30%, while lead times for downstream materials like copper-clad laminates have stretched from a normal 8 to 10 weeks to beyond 20. Bill Ho, an analyst at Yuanta, told Tom's Hardware: "With T-glass supply even more constrained now, suppliers are no longer providing lead times."

Bilal Hachemi, an analyst at Yole Group who tracks the IC substrate supply chain, said T-glass "has specific dielectric and CTE values that work better for the AI chips, especially for the organic core." He also pointed out that the IC substrate industry has historically operated on thin margins, which means even modest demand surges can trigger shortages. "Any increase in demand for build-up materials, ABF material, or T-glass can cause potential shortage, because it's against the basics of this industry," he said.

What makes this crunch acute is who is driving it. Hyperscalers are ordering ever-larger chip packages, and each generation uses a bigger interposer and a more complex substrate. According to data from Nvidia cited by Tom's Hardware, interposer sizes have grown from 814mm² for the Hopper architecture to 1,700mm² for Blackwell, a 109% increase. The forthcoming Rubin and Feynman generations scale further.

Bigger packages, more material

Bank of America estimates that Nittobo's electronic materials segment will see sales nearly double from ¥40.9 billion ($266 million) in 2025 to ¥87.7 billion by March 2028, with operating margins approaching 48%. Takashi Enomoto, a research analyst at Bank of America, said demand for ultra-thin T-glass is now rising too, on a shift away from E-glass in leading-edge devices.

The scramble for allocation has begun. Hachemi described Nvidia reaching out directly to an upstream material supplier as unprecedented: "For the first time, we are seeing Nvidia, the end customer, reaching out to the upstream material suppliers to secure the capacity and make sure they will get it." The worry is that once Nvidia locks down its share, rival chip buyers will be left fighting over the remainder.

Nittobo is not sitting still. Beyond the Fukushima expansion, it is doubling raw yarn capacity at its Taiwan plant and importing yarn back to Japan for cloth manufacturing. It has also struck a collaboration deal with Nanya Plastics to outsource some weaving. That is a sign of how tight the market is: Nittobo is partnering with one of its biggest competitors to ease the bottleneck. By 2027, roughly 20% of Nittobo's glass cloth is expected to be woven by Nanya.

The memory wall, and the stacking answer

Packaging material is one constraint. Memory bandwidth is the other, and it is the subject of a separate argument running through this year's chip conferences. ServeTheHome's write-up of d-Matrix's Hot Chips 2026 presentation lays out the problem: model weights keep growing, and the KV cache scales with context length multiplied by batch size. Sixty-four users at 1M context can mean roughly 935 GB of KV cache.

SRAM meets the bandwidth target but only at tiny scale. A Corsair SRAM accelerator card pair reaches roughly 300 TB/s at about 1 ns latency, yet holds only about 4 GB. HBM solves the capacity half but struggles on bandwidth, with d-Matrix citing a practical ceiling around 20 TB/s for HBM4 packages. Bandwidth that high carries a power price: at 2.4 pJ/bit, pushing 100 TB/s through HBM eats about 1.92 kW before any fabric traffic is counted.

d-Matrix's answer is to stack compute directly on top of DRAM dies. A TSMC N4 logic die sits on a 3D DRAM die using 36 um face-to-face stacking, a process the company describes as proven, low-cost, high-volume and high-yield. The thermal challenge is real. d-Matrix says a 1-Hi logic-on-top stack at no more than 0.5 W/mm² can be liquid cooled and keep DRAM under 100 C. Vertical 3D IO comes in at around 0.3 to 0.4 pJ, about 10 times lower than HBM.

There are engineering trade-offs. Each tensor engine needs a 128B flit per access, and with 3 banks per channel a single access returns 96B, wasting about 33% of bandwidth. d-Matrix's fix, called stream blocking, shares one partial 32B access across three flits so overfetch drops to zero. At 32GB per card, a 72-card scale-up can host a frontier model such as Kimi K3 at 1M context.

Microsoft and Rebellions show the same trend

Microsoft's Maia 200, presented at Hot Chips 2026 and covered live by ServeTheHome, is a 3nm chip with 140 billion transistors, 6 HBM3e stacks, 7TB/second of HBM bandwidth, 750 Watts TDP and 10,000 TFLOPS of FP4 performance. The SoC die is 820mm². Microsoft is deploying it inside Azure data centers with a scale-up-only networking design: 128 racks and 6,000 chips, connected by tier-0 and tier-1 switches over unified Ethernet. Attention benchmarks show an effective peak of 1.65 PFLOPS, and collective benchmarking reaches almost 1.3TB/second in BF16 AllReduce.

At Hot Chips 2025, Rebellions showed a different approach to the same packaging problem. The REBEL-Quad uses four HBM3E sites for 144GB of memory and UCIe as its chiplet interconnect, built on Samsung SF4X and CoWoS-S with four compute ASICs and four integrated silicon capacitors per package. ServeTheHome noted it is a dual PCIe Gen5 x16 card, and that the company ran a live Llama 3.3 70B demo at 35.5 msec average per output token on a development board.

The academic side is moving too. Stanford, Carnegie Mellon, Penn, MIT and SkyWater Technology reported the first monolithic 3D chip built in a U.S. foundry, presented at the 71st IEEE International Electron Devices Meeting in December 2025. Early hardware tests show the prototype outperforms comparable 2D chips by roughly a factor of four, and simulations of taller versions point to up to a twelve-fold improvement on AI workloads including those derived from Meta's LLaMA. Subhasish Mitra, the Stanford professor who led the work, said breakthroughs like this are how the industry gets to the 1,000-fold hardware performance improvements future AI systems will demand.

None of this arrives on the timescale the current shortage demands. The T-glass constraint is a years-long problem, and the stacking and packaging research that might relieve pressure later is still largely in presentation slides and development boards.

Comments 0

Sources

5
  1. 01Shortages of crucial chip packaging material threatens AI accelerator supply chainsEN
  2. 02d-Matrix Raptor 3D-DRAM Accelerator for Generative Inference at Hot Chips 2026EN
  3. 03Rebellions REBEL-Quad UCIe and 144GB HBM3E Accelerator at Hot Chips 2025EN
  4. 04Scientists and U.S. foundry achieve 3D chip breakthrough to accelerate AIEN
  5. 05Microsoft's Maia 200 AI Accelerator at Hot Chips 2026EN

All figures and quotations in this text come from the sources listed below.

Content prepared by the editorial team with AI assistance.

Grace Okonkwo

Grace Okonkwo

AI, models and technology

Grace Okonkwo covers AI, models and technology for FLASH24, working from primary sources such as model cards, API documentation and benchmark papers rather than vendor summaries. She checks training data provenance, evaluation conditions and reported scores against the underlying datasets before any figure reaches print. She interviews researchers and engineers directly, tracks release calendars from major labs, and compares successive model versions on the same tests. Her own self-hosting, home-network and documentation-reading habits feed straight into that desk, since she tests tools on her own hardware first. She does not publish benchmark claims without a reproducible method.

Newsroom →

Comments

0
  1. No comments yet — be the first.

Write a comment

Comments are public. We do not publish abuse, spam or advertising.