Skip to content
World clockEU--:--UK--:--USA--:--CN--:--PLDEFRIT中文EN

portal about AI and technologyevents · analysis · interviews · technical background

Search
LIVE
›

AI accelerators: the datacenter fight moves to the packaging supply chain

AI accelerator vendors are redesigning how compute talks to memory, but the deeper bottleneck sits upstream: a Japanese glass-fibre cloth maker controls roughly 90% of a material that every advanced package needs, and new capacity will not arrive until mid-2027.

TechnologyAnalysisRachel NwosuPublished: 27 September 20267 min readSources 5
AI accelerators: the datacenter fight moves to the packaging supply chain

Every AI accelerator argument in 2026 eventually runs into the same wall: memory bandwidth. Compute is expensive but available. Moving weights and KV cache in and out of memory is where the power budget, the package size and the shipping date get decided. That is why SiFive, d-Matrix and Rebellions all presented designs this cycle that attack the memory interface rather than the multiply-accumulate array.

The Register reported on 19 September 2024 that SiFive, long a supplier of RISC-V CPU cores into other companies' AI silicon, moved to licensing its own accelerator. The Intelligence XM cluster pairs four Intelligence X RISC-V cores with an in-house matrix math engine. Each cluster supports up to 1TB/sec of memory bandwidth and is rated at up to 16 TOPS of INT8 or 8 teraFLOPS of BF16 per gigahertz. SiFive's John Ronco told the publication that most chips built on the design will use four to eight clusters. That works out to between 4 and 8TB/sec of peak memory bandwidth and up to 32 to 64 teraFLOPS of BF16 at 1GHz.

Those numbers are modest next to an Nvidia H100.

The interesting part is who SiFive is selling to. Ronco said some customers still want to build their own hardware, but others wanted a one-stop shop. SiFive is not targeting Google or Tenstorrent, which already craft their own accelerators. It is targeting organisations that want an off-the-shelf block to customise and send to a fab.

The ceiling is unclear. Ronco was hesitant about how far the design scales, though the product slide deck suggests 512 XM clusters is within possibility. At a 1GHz clock, that would be roughly four petaFLOPS of BF16 matrix compute, against 2.5 petaFLOPS for Nvidia's top-specced Blackwell GPUs. SiFive also said it will offer an open source reference implementation of its SiFive Kernel Library.

Stacking DRAM on the logic die

ServeTheHome's coverage of Hot Chips 2026 describes d-Matrix taking a different route: putting logic directly on top of DRAM. The company's Raptor design stacks a TSMC N4 logic die on a 3D DRAM die using 36 micron face-to-face stacking. d-Matrix argues SRAM hits the bandwidth target, roughly 300 TB/s at about 1 ns latency for a Corsair SRAM accelerator card pair, but holds only around 4GB, and a 6T SRAM cell is about ten times larger than a DRAM cell. HBM solves capacity, but d-Matrix cites a practical bandwidth ceiling around 20 TB/s for HBM4 packages such as Nvidia Vera Rubin and AMD Instinct MI455.

Power is the reason. At 2.4 pJ/bit, pushing 100 TB/s through HBM consumes about 1.92 kW before any fabric traffic is counted. Vertical 3D IO, by contrast, comes in at roughly 0.3 to 0.4 pJ, about ten times lower than HBM, because it is a PHY-less millimetre-scale path rather than a centimetre-scale interposer route. The trade-off is thermal: d-Matrix says a 1-Hi logic-on-top stack at no more than 0.5 W/mm2 can be liquid cooled and keep DRAM under 100C.

The engineering detail is where the claims get testable. Each tensor engine needs a 128B flit per access, and with 32B delivered per column access from 32B banks that means four banks per channel. d-Matrix's die has 840 banks, 768 after 72 spares, across 256 channels: three banks per channel. A single access returns 96B, so delivering a 128B flit takes two accesses and fetches 192B, wasting about 33% of bandwidth near 33 TB/s. d-Matrix's fix, stream blocking, shares one partial 32B access across three flits, so four accesses at 96B feed three flits at 128B. Overfetch drops to zero. A second trick, stream flipping, inverts each flit against the previous one to cut toggles, recovering about 20% of I/O power without the sideband pin that conventional data bus inversion needs.

At 32GB per card, with 4-bit weights and an 8-bit KV cache, d-Matrix sizes a 72-card scale-up to host a frontier model such as Kimi K3 at 1M context. ServeTheHome notes the company presented the work at Hot Chips 2026.

UCIe reaches a running board

Rebellions showed something rarer than a slide: silicon running. ServeTheHome reported on 25 August 2025 that the Rebellions REBEL-Quad was demonstrated at Hot Chips 2025 on a development board running Llama 3.3 70B at 35.5 msec average per output token. The package is built on Samsung SF4X and CoWoS-S, with four compute ASICs, four HBM3E sites for 144GB of memory and four integrated silicon capacitors. It uses UCIe-A as its chiplet interconnect and a dual PCIe Gen5 x16 interface card.

ServeTheHome flagged the PCIe choice as a possible miss, given that Nvidia's GB300 ushers in PCIe Gen6. That is a fair criticism of a part aimed at scale-out inference. But the demo matters more than the spec sheet. UCIe has been discussed for years, and vendors have been slow to market that they use it because the interconnect sits inside the package. Rebellions integrating that many pieces of silicon in one large package and showing it working is a data point the chiplet ecosystem needed.

Not every accelerator needs a datacenter rack. Draw Things published results on 11 November 2025 for Metal FlashAttention v2.5 with Neural Accelerators on Apple's M5, claiming up to 4.6x improvement over M4 in its own image and video generation app. The company reports 3.6x to 5.5x raw performance gains over M4 iPad and 3.3x to 4.6x end-to-end, with an M5 iPad performing in the range of an M2 Max and often around 80% slower than an M3 Ultra (60c). It says a 16GiB M5 iPad can generate five-second 480p video with Wan 2.2 A14B models. The release is a preview: Draw Things says BF16 support was initially off due to unsolved bugs, later restored in a 17 November build, and that odd attention sequence lengths and large head dimensions still cause performance cliffs.

The material nobody outside packaging has heard of

All of this design work runs into a supply chain problem that has nothing to do with architecture. Tom's Hardware reported on 9 March 2026 that Nittobo controls roughly 90% of the global supply of T-glass, a low-CTE glass-fibre cloth used in the organic core of IC substrates, the interconnect layer between a chip and its printed circuit board. T-glass keeps large, hot packages dimensionally stable. E-glass is cheaper but used mainly in lower-end chips such as microcontrollers and older mobile processors; for massive 2.5D and 3D packaging, T-glass is preferred.

Nittobo is tripling capacity at its Fukushima plant, but the new supply will not reach the market until mid-2027. Prices have risen between 20% and 30%, and lead times for downstream materials such as copper-clad laminates have stretched from a normal 8 to 10 weeks to beyond 20. Yuanta analyst Bill Ho told Tom's Hardware that suppliers are no longer providing lead times. Bilal Hachemi of Yole Group said T-glass is not easy to replace because it has specific dielectric and CTE values that work better for AI chips, especially the organic core. He added that the IC substrate industry's historically thin margins mean even modest demand surges trigger shortages.

The demand driver is package size. According to data from Nvidia cited by Tom's Hardware, interposer sizes have grown from 814mm2 for Hopper to 1,700mm2 for Blackwell, a 109% increase, with Rubin and Feynman scaling further. Bank of America estimates Nittobo's electronic materials segment will nearly double sales from 40.9 billion yen in 2025 to 87.7 billion yen by March 2028, with operating margins approaching 48%. Nittobo is doubling raw yarn capacity at its Taiwan plant and has struck a collaboration deal with Nanya Plastics to outsource some weaving. By 2027, roughly 20% of Nittobo's glass cloth is expected to be woven by Nanya.

Nittobo is partnering with one of its biggest competitors to ease the bottleneck. That is the clearest signal yet of where the constraint sits. Accelerator architects can keep cutting picojoules per bit, but if the cloth inside the substrate is allocated years ahead, the shipping date is decided somewhere else entirely.

Comments 0

Sources

5
  1. 01SiFive shifts from RISC-V cores for AI chips to designing its own acceleratorEN
  2. 02d-Matrix Raptor 3D-DRAM Accelerator for Generative Inference at Hot Chips 2026EN
  3. 03Rebellions REBEL-Quad UCIe and 144GB HBM3E Accelerator at Hot Chips 2025EN
  4. 04Metal FlashAttention v2.5 with Neural Accelerators on the Apple M5 ChipEN
  5. 05Shortages of crucial chip packaging material threatens AI accelerator supply chainsEN

All figures and quotations in this text come from the sources listed below.

Content prepared by the editorial team with AI assistance.

Rachel Nwosu

Rachel Nwosu

AI, models and technology

Rachel Nwosu covers AI, models and technology for FLASH24, working from public model documentation, benchmark releases and repository histories rather than press summaries, and she skips announcements that arrive without reproducible numbers. She checks training-data claims against dataset cards and reruns reported metrics where code is available. She spends much of her week interviewing researchers and engineers, tracking model launch calendars, and comparing vendor benchmarks with independent evaluations. Outside the desk she runs 3D printers, restores old computers, and tests how models learn from internet junk. She does not publish benchmark figures she cannot trace to a source.

Newsroom →

Comments

0
  1. No comments yet — be the first.

Write a comment

Comments are public. We do not publish abuse, spam or advertising.