What Is Really Inside an AI Accelerator: T-Glass, HBM and the Decode Bottleneck
One Japanese company controls roughly 90% of the global supply of a glass-fiber cloth that sits inside every advanced AI chip package. The shortage is now shaping accelerator roadmaps as much as GPU architecture does.

Nittobo controls about 90% of the world market for T-glass, a low-CTE glass cloth used in the organic core of IC substrates, according to Tom's Hardware. The company is tripling capacity at its Fukushima plant, but the new supply will not reach the market until mid-2027.
That gap is already visible in prices and schedules. T-glass prices have risen between 20% and 30%, and lead times for downstream materials such as copper-clad laminates have stretched from a normal 8 to 10 weeks to beyond 20. "With T-glass supply even more constrained now, suppliers are no longer providing lead times," Bill Ho, an analyst at Yuanta, told Tom's Hardware. The squeeze shows up first in quotes that sales teams cannot hold for more than a few days, then in the schedules that fabless designers build their launch windows around. Nobody in the chain will say publicly how long the tightness lasts, because nobody wants to be the one who called it wrong.
Why the substrate matters more than the transistor count
An AI accelerator is not one chip. It is a package: logic dies, memory stacks, an interposer and a substrate that connects all of it to a printed circuit board. The substrate's organic core has to stay flat while the package heats and cools, because a warped package breaks the microscopic connections between dies. T-glass is what keeps it dimensionally stable, and its dielectric and thermal-expansion values are tuned for exactly that job.
Bilal Hachemi, an analyst at Yole Group who tracks the IC substrate supply chain, told Tom's Hardware Premium that replacing it is not simple. T-glass "has specific dielectric and CTE values that work better for the AI chips, especially for the organic core," he said. He also pointed to an industry structure that turns small demand shocks into shortages. The IC substrate business has historically run on thin margins, so "any increase in demand for build-up materials, ABF material, or T-glass can cause potential shortage, because it's against the basics of this industry." That structure predates the AI boom. It was built for a market where substrate demand grew slowly and predictably, and it has not been rebuilt for one where a single product cycle can double orders in a year.
The demand side is not subtle. Each generation of AI accelerator uses a bigger interposer and a more complex substrate, which means more material per unit. According to data from Nvidia cited by Tom's Hardware, interposer sizes have grown from 814mm² for Hopper to 1,700mm² for Blackwell, a 109% increase, with the forthcoming Rubin and Feynman generations scaling further.
Bank of America estimates that Nittobo's electronic materials segment will see sales nearly double from ¥40.9 billion ($266 million) in 2025 to ¥87.7 billion by March 2028, with operating margins approaching 48%. "Demand for T-glass cloth seems likely to grow more than originally expected," said Takashi Enomoto, a research analyst at Bank of America. He added that the focus had been on thick T-glass for GPU and CPU packages, but that ultra-thin T-glass demand is likely to rise as leading-edge devices shift away from E-glass.
Nittobo is not waiting. Beyond Fukushima, it is doubling raw-yarn capacity at its Taiwan plant and shipping yarn back to Japan for weaving, and it has a collaboration deal with Nanya Plastics to outsource some weaving. Tom's Hardware notes the symbolism: Nittobo is partnering with one of its biggest competitors. By 2027, roughly 20% of Nittobo's glass cloth is expected to be woven by Nanya.
Memory is the other half of the bottleneck
If the substrate decides whether a package can be built, memory decides how fast it can answer. d-Matrix used its Hot Chips 2026 presentation on the Raptor accelerator, covered by ServeTheHome, to lay out the arithmetic. Model weights keep growing, and the KV cache scales with context length multiplied by batch size, so 64 users at 1M context can mean roughly 935 GB of KV cache.
The split between compute and memory shows up in the two phases of inference. Prefill processes many prompt tokens in parallel and is compute-throughput-bound. Decode produces one token at a time and is typically memory-bandwidth-bound. Since decode dominates wall-clock runtime, bandwidth is where the wins are. SRAM meets the bandwidth target but only at tiny scale: a Corsair SRAM accelerator card pair reaches roughly 300 TB/s at about 1 ns latency and holds only about 4 GB. A 6T SRAM cell is around 10 times larger than a DRAM cell.
HBM solves capacity but not bandwidth. d-Matrix cites a practical ceiling around 20 TB/s for HBM4 packages such as the Nvidia Vera Rubin and AMD Instinct MI455, limited by pin speed, I/O width and package beachfront of roughly 8 to 16 stacks. Bandwidth also carries a power price: at 2.4 pJ/bit, pushing 100 TB/s through HBM costs about 1.92 kW before any fabric traffic is counted.
"For the first time, we are seeing Nvidia, the end customer, reaching out to the upstream material suppliers to secure the capacity and make sure they will get it." Bilal Hachemi, Yole Group, to Tom's Hardware
d-Matrix's answer is to stack compute directly on DRAM dies. A TSMC N4 logic die sits on a 3D DRAM die using 36 um face-to-face stacking. Vertical 3D IO comes in at around 0.3 to 0.4 pJ, roughly 10 times lower than HBM, because it is a millimeter-scale path rather than a centimeter-scale interposer route. The company says a 1-Hi logic-on-top stack at no more than 0.5 W/mm² can be liquid cooled and keep DRAM under 100 C.
Integration is where the elegance runs out. Each tensor engine needs a 128B flit per access, and with 32B delivered per column access from 32B banks, that works out to 4 banks per channel. The die has 840 banks, 768 after 72 spares, spread across 256 channels, or just 3 banks per channel. Stream blocking fixes the mismatch by sharing one partial 32B access across three flits, so 4 accesses at 96B feed 3 flits at 128B. At 32GB per card with 4-bit weights and an 8-bit KV cache, d-Matrix sizes a frontier model such as Kimi K3 at 1M context into a single 72-card rack.
The design loop is being automated, unevenly
Packaging and memory are physical constraints. Design time is an economic one. Nvidia's Tim Costa, vice president and general manager of computational engineering, told journalists that by 2030 the industry is expected to produce 2 trillion chips and process about 41 million wafers a month, with individual packages approaching a trillion transistors. "The traditional design process just can't keep pace with that scale of challenge," he said, according to The Next Platform.
Nvidia is deploying its Arm-based Vera CV100 CPU across its own EDA workflows, including simulation, formal verification and physical implementation. Costa said early testing shows Vera running Synopsys VCS and Cadence Jasper at 1.5 times the performance of AMD Epyc Torrent systems. Cadence claims its AI Super Agents can run hundreds of simulations simultaneously and deliver 40-times faster register-transfer-level validation cycles. Synopsys says its autonomous verification workflow can deliver validated RTL 50 times faster. Those are vendor numbers, and SemiEngineering notes that the shift toward specialised agents is creating new problems around orchestration, integration and guardrails, with AI-generated results still requiring verification before sign-off.
One data point from Cadence's ChipStack AI Super Agent, reported by SemiEngineering, suggests the gains are not only about speed. The company reports about 24% lower area and 18% lower power versus pure foundation-model code generation.
None of this removes the upstream constraint. A larger interposer needs more T-glass. A bigger model needs more memory bandwidth. Nvidia going directly to Nittobo, which Hachemi described as unprecedented in that part of the supply chain, is a signal about where the real scarcity sits. The worry he names is straightforward: once Nvidia locks down its share, rival chip buyers are left fighting over whatever remains. New fabs take years, and so does a new glass-cloth furnace running between 1,600 and 1,700°C.
Sources
4- 01Shortages of crucial chip packaging material threatens AI accelerator supply chainsEN
- 02d-Matrix Raptor 3D-DRAM Accelerator for Generative Inference at Hot Chips 2026EN
- 03Nvidia Accelerates Chip Engineering With AI AgentsEN
- 04Chip Industry Week In ReviewEN
All figures and quotations in this text come from the sources listed below.
Content prepared by the editorial team with AI assistance.
Comments
0- No comments yet — be the first.