Skip to content
World clockEU--:--UK--:--USA--:--CN--:--PLDEFRIT中文EN

portal about AI and technologyevents · analysis · interviews · technical background

Search
LIVE
›

Inside AI datacenter accelerators: RISC-V cores, T-glass shortages and 3D DRAM

AI accelerator designs are shifting on three fronts at once: who licenses the building blocks, who can get the packaging materials, and how memory is stacked under the compute. The Register reported on 19 September 2024 that SiFive is now offering its own accelerator blueprints, not just RISC-V CPU cores.

TechnologyExplainerGrace OkonkwoPublished: 27 September 20266 min readSources 3
Inside AI datacenter accelerators: RISC-V cores, T-glass shortages and 3D DRAM

SiFive built its name on RISC-V CPU cores that other companies bolt onto their AI chips. This week the company said it will license a complete machine-learning accelerator of its own. The Intelligence XM series is a cluster of four SiFive Intelligence X RISC-V CPU cores wired to an in-house matrix math engine, according to The Register.

John Ronco, SiFive's SVP and GM for the UK, told The Register that some customers still want to build their own hardware. Others, he said, "wanted more of a one-stop shop from SiFive." The company's earlier work put its CPU cores next to third-party matrix engines, as in Google's tensor processing units. The XM clusters flip that. SiFive now supplies the whole accelerator design, and customers customize it and send it to a fab.

What one cluster actually delivers

Each base XM cluster supports up to 1TB/sec of memory bandwidth through a coherent hub interface. SiFive expects it to deliver up to 16 TOPS of INT8 or 8 teraFLOPS of BF16 performance per gigahertz.

Per-gigahertz numbers can mislead, because this is not a finished chip. Final performance depends on how many clusters a customer places, how they are wired, what else sits on the die, and the power and cooling budget. Ronco expects most designs to use between four and eight clusters. At a 1GHz clock, that works out to between 4 and 8TB/sec of peak memory bandwidth and up to 32 to 64 teraFLOPS of BF16. An Nvidia H100 churns out nearly a petaFLOPS of dense BF16, so the comparison flatters Nvidia. But FLOPS are not everything for bandwidth-bound work like inference. Ronco said the XM clusters probably will not be used much for AI training.

Scale is the open question. Ronco was hesitant to say how far the design can stretch, and The Register noted that process technology and die area will set limits. The company's product slide deck suggests 512 XM clusters is possible. At 1GHz with no thermal or power problems, that would come to roughly four petaFLOPS of BF16 matrix compute. Nvidia's top-specced Blackwell GPUs are listed at 2.5 petaFLOPS of BF16.

SiFive also said it will publish an open source reference implementation of its SiFive Kernel Library to lower adoption barriers for RISC-V. Its CEO, Patrick Little, claimed in a canned statement that the company supplies RISC-V-based chip designs to five of the "Magnificent 7" companies, a group that includes Microsoft, Apple, Nvidia, Alphabet, Amazon, Meta and Tesla. The Register noted it suspects not all of that silicon involves AI.

The packaging bottleneck nobody watches

Designs are one thing. Materials are another. A Japanese company called Nittobo controls roughly 90% of the global supply of specialist glass-fiber cloth known as T-glass, which sits inside every advanced AI chip package, according to Tom's Hardware. T-glass is a low-CTE glass cloth used in the organic core of IC substrates, the interconnect layer between a chip and its printed circuit board. It keeps large, hot chip packages dimensionally stable as they get denser.

Demand has outrun supply. Nittobo is tripling capacity at its Fukushima plant, but the new output will not reach the market until mid-2027. Prices have already risen 20% to 30%, and lead times for downstream materials such as copper-clad laminates have stretched from a normal 8 to 10 weeks to beyond 20. "With T-glass supply even more constrained now, suppliers are no longer providing lead times," Bill Ho, an analyst at Yuanta, told Tom's Hardware.

Replacing T-glass is not simple. Bilal Hachemi, an analyst at Yole Group who tracks the IC substrate supply chain, said the material "has specific dielectric and CTE values that work better for the AI chips, especially for the organic core." The substrate industry runs on thin margins, so even modest demand surges can trigger shortages. "Any increase in demand for build-up materials, ABF material, or T-glass can cause potential shortage, because it's against the basics of this industry," Hachemi said.

Hyperscalers are making it worse by ordering ever-larger packages. According to Nvidia data cited by Tom's Hardware, interposer sizes grew from 814mm2 for Hopper to 1,700mm2 for Blackwell, a 109% increase, with the Rubin and Feynman generations scaling further. Bank of America estimates Nittobo's electronic materials segment will nearly double sales from 40.9 billion yen, about $266 million, in 2025 to 87.7 billion yen by March 2028, with operating margins approaching 48%.

Nittobo is doubling raw yarn capacity at its Taiwan plant, importing yarn back to Japan for cloth manufacturing, and outsourcing some weaving to Nanya Plastics, one of its biggest competitors. By 2027, roughly 20% of Nittobo's glass cloth is expected to be woven by Nanya, the company disclosed. Hachemi said Nvidia reaching out directly to an upstream material supplier like Nittobo is unprecedented, and warned that once Nvidia locks down its share, rival chip buyers will fight over what is left.

3D DRAM and the memory wall

Memory is the third front. At Hot Chips 2026, d-Matrix presented its Raptor accelerator, which stacks compute directly on top of DRAM dies, according to ServeTheHome. Model weights keep growing, and the KV cache scales with context length multiplied by batch size. ServeTheHome's example: 64 users at 1M context can mean roughly 935GB of KV cache. That is both a capacity and a bandwidth problem.

SRAM hits the bandwidth target but only at tiny scale. A Corsair SRAM accelerator card pair reaches roughly 300TB/s at about 1 nanosecond of latency, yet holds only about 4GB. HBM solves capacity but struggles on bandwidth, with a practical ceiling around 20TB/s for HBM4 packages such as Nvidia's Vera Rubin and AMD's Instinct MI455, according to d-Matrix. Stacked 3D DRAM lands between the two: vertical 3D IO costs around 0.3 to 0.4 pJ, about 10 times lower than HBM, because it is a PHY-less millimeter-scale path rather than a centimeter-scale interposer route.

Raptor puts a TSMC N4 logic die on top of a 3D DRAM die using 36 micron face-to-face stacking, a process d-Matrix describes as proven, low-cost, high-volume and high-yield. Each card holds 32GB, and d-Matrix says a 72-card scale-up can host a frontier model such as Kimi K3 at 1M context with 4-bit weights and an 8-bit KV cache.

The engineering is not finished. d-Matrix says a 1-Hi logic-on-top stack at no more than 0.5 W/mm2 can be liquid cooled and keep DRAM under 100C. Its die has 840 banks, 768 after 72 spares, across 256 channels, which gives 3 banks per channel. A 128B flit does not divide evenly across that, so a single access returns 96B and two accesses fetch 192B, wasting about 33% of bandwidth near 33TB/s. Its stream blocking scheme shares one partial 32B access across three flits, cutting overfetch to zero.

Then there is I/O power. Moving 100TB/s at 0.37 pJ/bit works out to 296W just for I/O, and conventional DBI could save 20%. HBM gets away with DBI because its multi-cycle bursts let the PHY see the full burst. d-Matrix's single-cycle 256-bit 3D-DRAM link has no burst and no sideband pin to signal the inversion choice, so it uses stream flipping instead, comparing each flit to the previous one and inverting when needed.

None of this arrives on its own. RISC-V accelerator blueprints, a near-monopoly on one packaging cloth, and stacked DRAM all have to line up before the next generation of datacenter silicon ships.

Comments 0

Sources

3
  1. 01SiFive shifts from RISC-V cores for AI chips to designing its own acceleratorEN
  2. 02Shortages of crucial chip packaging material threatens AI accelerator supply chainsEN
  3. 03D-Matrix Raptor 3D-DRAM Accelerator for Generative Inference at Hot Chips 2026EN

All figures and quotations in this text come from the sources listed below.

Content prepared by the editorial team with AI assistance.

Grace Okonkwo

Grace Okonkwo

AI, models and technology

Grace Okonkwo covers AI, models and technology for FLASH24, working from primary sources such as model cards, API documentation and benchmark papers rather than vendor summaries. She checks training data provenance, evaluation conditions and reported scores against the underlying datasets before any figure reaches print. She interviews researchers and engineers directly, tracks release calendars from major labs, and compares successive model versions on the same tests. Her own self-hosting, home-network and documentation-reading habits feed straight into that desk, since she tests tools on her own hardware first. She does not publish benchmark claims without a reproducible method.

Newsroom →

Comments

0
  1. No comments yet — be the first.

Write a comment

Comments are public. We do not publish abuse, spam or advertising.