How AI Accelerators Actually Work: Cores, Cluster Building Blocks and Memory
SiFive is now licensing a drop-in AI accelerator cluster built around four RISC-V CPU cores and its own matrix math engine, The Register reported on 19 September 2024.

SiFive spent years designing RISC-V CPU cores for other companies' AI chips. Now it sells a complete accelerator design of its own. Each Intelligence XM cluster pairs four of its X-series RISC-V CPU cores with an in-house matrix math engine, according to The Register.
That is a shift. Google's tensor processing units already use SiFive's X280 cores to keep their matrix multiplication units fed. John Ronco of SiFive UK told The Register that SiFive designs also underpin CPU cores in Tenstorrent's Blackhole accelerator. The difference now is that SiFive brings its own matrix engine instead of attaching cores to a third party's.
What the numbers say, and what they do not
Each cluster supports up to 1TB/sec of memory bandwidth through a coherent hub interface. SiFive expects it to deliver up to 16 TOPS of INT8 or 8 teraFLOPS of BF16 per gigahertz. That per-gigahertz framing matters, because a cluster is not a chip. Performance depends on how many clusters a customer places on the die, how they are wired, what else sits alongside them, and how fast the final part is clocked.
Ronco expects most designs to use four to eight clusters. At 1GHz, that would mean 4 to 8TB/sec of peak memory bandwidth and 32 to 64 teraFLOPS of BF16. SiFive's product slides suggest 512 clusters is possible. At 1GHz that would be roughly four petaFLOPS of BF16, above the 2.5 petaFLOPS Nvidia quotes for its top Blackwell GPUs. An Nvidia H100, by comparison, delivers nearly a petaFLOPS of dense BF16.
The catch is physics. Holding 1GHz across 512 clusters without hitting thermal or power limits is the open question, and Ronco was hesitant to say how far the design scales. He also expects the clusters will not be used widely for AI training. SiFive says it will publish an open source reference implementation of its SiFive Kernel Library.
Memory bandwidth is the recurring bottleneck
Accelerators are built this way because inference, especially the decode phase, is memory-bound. d-Matrix, presenting at Hot Chips 2026, put the problem bluntly: 64 users at 1M context can mean roughly 935GB of KV cache. Its Raptor design stacks a TSMC N4 logic die on 3D DRAM using 36 micrometre face-to-face stacking at 32GB per card, with a 72-card rack sized for frontier models.
Rebellions showed a different approach at Hot Chips 2025. Its REBEL-Quad packs four compute ASICs and four HBM3E sites for 144GB of memory on a Samsung SF4X and CoWoS-S package, using UCIe-A as the chiplet interconnect. ServeTheHome saw it running Llama 3.3 70B on a development board at 35.5 milliseconds average per output token.
Suppliers, packaging and the upstream squeeze
None of this ships without packaging materials. Tom's Hardware reported in March 2026 that Nittobo controls roughly 90% of the global supply of T-glass, the low-CTE glass cloth used in IC substrate cores. Prices have risen 20 to 30%, and lead times for copper-clad laminates stretched beyond 20 weeks from a normal 8 to 10. Nittobo is tripling capacity at its Fukushima plant, but new supply will not reach the market until mid-2027. Bank of America estimates the electronic materials segment will nearly double sales from 40.9 billion yen in 2025 to 87.7 billion yen by March 2028.
The policy layer is just as tangled. TSMC reportedly cut off Chinese designer Sophgo in October 2024 after a customer ordered a chip resembling Huawei's Ascend 910B. Sophgo denied any Huawei relationship and said it submitted an investigation report to TSMC.
Sources
5- 01SiFive shifts from RISC-V cores for AI chips to designing its own acceleratorEN
- 02Shortages of crucial chip packaging material threatens AI accelerator supplyEN
- 03d-Matrix Raptor 3D-DRAM Accelerator for Generative Inference at Hot Chips 2026EN
- 04Rebellions REBEL-Quad UCIe and 144GB HBM3E Accelerator at Hot Chips 2025EN
- 05TSMC reportedly cuts off RISC-V chip designer linked to Huawei acceleratorsEN
All figures and quotations in this text come from the sources listed below.
Content prepared by the editorial team with AI assistance.
Comments
0- No comments yet — be the first.