AI accelerators: the packaging bottleneck nobody prices in, and the startups betting against HBM
A Japanese company most datacenter operators have never heard of controls roughly 90% of the global supply of a glass cloth that sits inside every advanced AI chip package, and its next production line will not arrive until mid-2027.

Nittobo's T-glass is a low-CTE glass cloth used in the organic core of IC substrates, the interconnect layer between a chip and its printed circuit board. According to Tom's Hardware, the company holds roughly 90% of the global supply. Demand has outrun it badly enough that downstream lead times for copper-clad laminates have stretched from a normal 8 to 10 weeks to beyond 20. Prices are up 20 to 30%.
That number matters more than it looks. Every generation of AI accelerator packs a bigger interposer and a more complex substrate, so each chip consumes more of the material. Nvidia's own data, cited by Tom's Hardware, puts interposer sizes at 814mm² for Hopper and 1,700mm² for Blackwell, a 109% increase, with Rubin and Feynman scaling further. The supply side cannot answer quickly. T-glass needs electric melting furnaces running between 1,600 and 1,700°C, and Nittobo is tripling capacity at its Fukushima plant. New output will not reach the market until mid-2027.
Analysts quoted in that report describe a market where substitution is hard. "It is not easy to replace the T-glass," Bilal Hachemi of Yole Group told Tom's Hardware, pointing to specific dielectric and CTE values that suit AI chips, especially the organic core. Yuanta analyst Bill Ho said suppliers have stopped quoting lead times at all. Bank of America's Takashi Enomoto expects Nittobo's electronic materials sales to nearly double from ¥40.9 billion ($266 million) in 2025 to ¥87.7 billion by March 2028, with operating margins approaching 48%.
One detail stands out. Hachemi said Nvidia going directly to an upstream material supplier was unprecedented in that part of the chain. If the largest buyer locks capacity first, everyone else negotiates over the remainder.
Memory bandwidth is the other wall
While packaging tightens, accelerator designers are arguing about memory. At Hot Chips 2026, d-Matrix presented Raptor, an accelerator that stacks a TSMC N4 logic die on top of a 3D DRAM die using 36 micron face-to-face stacking, according to ServeTheHome. The pitch is that neither SRAM nor HBM alone fits inference. SRAM delivers bandwidth but holds little capacity. HBM delivers capacity but runs into pin speed, stack count and package beachfront, and d-Matrix cites a practical ceiling around 20 TB/s for HBM4 packages.
Bandwidth has a power bill attached. ServeTheHome's write-up puts HBM at 2.4 pJ/bit, meaning 100 TB/s costs about 1.92 kW before fabric traffic. Vertical 3D IO comes in around 0.3 to 0.4 pJ, roughly ten times lower, because it is a millimeter-scale path rather than a centimeter-scale interposer route.
d-Matrix's numbers are specific and unglamorous. Each tensor engine needs a 128B flit per access, but 32B banks across 256 channels give 3 banks per channel after spares, returning 96B per access. Two accesses fetch 192B to deliver 128B, wasting about 33% of bandwidth near 33 TB/s. The company's fix, stream blocking, shares one partial access across three flits so 384B in matches 384B out, with overfetch dropping to zero. At 32GB per card with 4-bit weights and an 8-bit KV cache, d-Matrix says a 72-card scale-up can host a frontier model at 1M context.
Chiplets and the RISC-V licensing shift
Packaging constraints and memory trade-offs are pushing vendors toward different integration strategies. Rebellions showed its REBEL-Quad at Hot Chips 2025: four compute ASICs, four HBM3E sites for 144GB, and UCIe-A as the die-to-die interconnect, built on Samsung SF4X and CoWoS-S, per ServeTheHome. The card runs dual PCIe Gen5 x16. The company demoed Llama 3.3 70B on a development board at an average 35.5 msec per output token.
ServeTheHome noted the PCIe choice sits awkwardly against Nvidia's move toward Gen6, but treated the live demo as the more significant fact: a working UCIe-based chip in public, which is rare enough to be worth pointing out. On the compute side, SiFive has moved from licensing RISC-V CPU cores into other companies' AI silicon to selling its own accelerator design. The Register reported on 19 September 2024 that SiFive's Intelligence XM clusters pair four Intelligence X CPU cores with an in-house matrix math engine, with up to 1TB/sec of memory bandwidth per cluster and up to 16 TOPS of INT8 or 8 teraFLOPS of BF16 per gigahertz.
John Ronco, SiFive's SVP and GM for the UK, told The Register that some customers still want their own hardware, but others wanted a one-stop shop. He expected most chips built on the design to use four to eight clusters, and was hesitant about how far it scales, though the company's slide deck suggests 512 clusters is possible. For context, Nvidia's top Blackwell parts are rated at 2.5 petaFLOPS of BF16.
SiFive already supplies RISC-V designs used in Google's tensor processing units and in Tenstorrent's Blackhole. CEO Patrick Little claimed in a canned statement that the company supplies five of the "Magnificent 7" firms, a claim The Register flagged as probably not all AI-related.
Sanctions keep reshaping the map
The supply chain is also being redrawn by export controls. The Register reported on 28 October 2024 that TSMC had allegedly cut off Chinese designer Sophgo over suspicions it was trying to supply components to Huawei. Reuters, citing two people familiar with the matter, identified the customer that ordered a chip resembling Huawei's Ascend 910B. Sophgo denied any direct or indirect business relationship with Huawei and said it had submitted a detailed investigation report to TSMC. Bitmain, associated with Sophgo, called the allegations false and baseless.
The Register noted Bitmain's record is not clean. In 2021, prosecutors raided its Taiwan offices and accused two affiliates of trying to illegally recruit Taiwanese engineers. TSMC stopped doing business with Huawei in 2020 under US rules.
Huawei's Ascend 910B is rated at 320 teraFLOPS of FP16 or 640 teraFLOPS of Int8, roughly on par with Nvidia's four-year-old A100, according to that report. Slower than the current frontier, but available, which is the part that matters when export caps keep tightening.
Sources
5- 01Shortages of crucial chip packaging material threatens AI accelerator supply chainsEN
- 02d-Matrix Raptor 3D-DRAM Accelerator for Generative Inference at Hot Chips 2026EN
- 03Rebellions REBEL-Quad UCIe and 144GB HBM3E Accelerator at Hot Chips 2025EN
- 04SiFive expands from RISC-V cores for AI chips to designing its own full-fat acceleratorEN
- 05TSMC reportedly cuts off RISC-V chip designer linked to Huawei acceleratorsEN
All figures and quotations in this text come from the sources listed below.
Content prepared by the editorial team with AI assistance.
Comments
0- No comments yet — be the first.