Skip to content
World clockEU--:--UK--:--USA--:--CN--:--PLDEFRIT中文EN

portal about AI and technologyevents · analysis · interviews · technical background

Search
LIVE
›

Inference costs split two ways: open engines tune themselves, chipmakers cut hardware

An open source inference engine launched on GitHub on 30 September claims open models run up to 2x faster than llama.cpp by compiling and tuning its kernels on the user's own device, while MaxLinear says its new Puma 9 DOCSIS chip can cut customer premises equipment costs by 30% to 50%.

AI & modelsAnalysisRachel NwosuPublished: 1 October 20265 min readSources 11
Inference costs split two ways: open engines tune themselves, chipmakers cut hardware

Magnitude, a Y Combinator S25 company, published its repository on 30 September under an Apache 2.0 licence. The desktop app tunes kernels on Apple Silicon, NVIDIA, AMD or plain CPUs before a model runs. MaxLinear used a 29 September briefing to claim its Puma 9 silicon cuts modem and gateway costs by 30% to 50% against its predecessor, the Puma 8 introduced in 2023.

Magnitude's numbers are specific: the README claims 92% faster decode on Metal and 19% on CUDA against llama.cpp, 27% less memory per agent that is freed when agents stop, and prefix caches shared across concurrent sessions. It connects to Pi, OpenCode, Hermes, OpenClaw, Codex, Claude Code, Oh My Pi and Cline, with anything else handled through an OpenAI-compatible API. The company says prompts, files and models stay on the machine and no internet is required once a model is downloaded. That is the pitch. Whether it holds on your hardware is another matter: the benchmark is the vendor's own, and the repository does not publish a third-party replication.

MaxLinear's claim rests on a different kind of tuning. The Puma 9 supports DOCSIS 4.0, DOCSIS 3.1+ and DOCSIS 3.1, adds Wi-Fi 8 support and moves to an ARM architecture for the first time in the Puma line, which MaxLinear inherited through its 2020 acquisition of Intel's Home Gateway Platform unit.

An industry source familiar with the silicon told Light Reading the ARM switch helps cut power draw and some licensing costs. MaxLinear says it will keep supporting x86 products and that the RDK software platform makes the underlying CPU architecture largely transparent to customers. In the face of DDR4 shortages, the Puma 9 also supports DDR5. Puneet Sethi, SVP and GM of MaxLinear's network infrastructure and carrier business unit, said the company thinks the DDR5 ecosystem, pricing and supply will become more relaxed than the stress seen with DDR4. Jeff Heynen of Dell'Oro Group told Light Reading most vendors will move to DDR5 for more advanced Wi-Fi 8 units because of the extra onboard memory apps and diagnostics may demand.

MaxLinear says it has multiple commitments for the Puma 9, names Askey, Fritz! GmbH and Gemtek as CPE partners, and has support from Australian operator NBN Co. Modems and gateways based on it are expected in 2027.

The cost pressure is not confined to modems. In healthcare, Endpoints News asked on 1 October whether AI is raising costs rather than cutting them, framing the question against a wave of funding: Parakeet Health raised $10M to automate patient outreach and Devoted Health raised $1.2B, both reported on 1 October. CleanTechnica reported on 29 September that Transport & Environment analysis puts EV running costs at less than half of diesel per kilometre, with diesel drivers paying €32 more per tank than at the start of the year, around €16 of it extra refinery margin. The same outlet reported on 28 September on data centre operators turning to methane-fired portable generators, and on the older diesel exhaust fight that followed the World Health Organization's 2012 classification of diesel exhaust as a carcinogen.

Hardware makers are also trying to strip cost out of batteries. Electrek reported on 29 September that Ultium Cells, GM's joint venture with LG Energy Solution, will mass-produce prismatic LMR cells at Spring Hill, Tennessee, claiming 33% higher energy density than LFP at comparable cost. GM vice president Kurt Kelty said adding LMR positions the company to leapfrog today's more affordable chemistries. Upgrades start later this year and are due for completion by 2028, adding 500 jobs for a workforce of 1,700; the first vehicle with LMR batteries is expected in 2028.

Cheaper running, more expensive building

The pattern in the dossier is that inference and operations get cheaper while the assets underneath stay costly. Xpeng consolidated four product lines into two to focus R&D and cut costs, 36Kr reported on Monday via CnEVPost. The numbers behind that decision: deliveries fell 10.49% to 243,111 vehicles in the first eight months of 2026, R&D expenses rose 32.1% year-on-year to 2.91 billion yuan in the second quarter, and the quarterly net loss widened to 1.34 billion yuan from 480 million yuan a year earlier. Telecoms operators are taking the same route. Light Reading reported on 30 September that StarHub and M1 are in merger talks in Singapore, where Singtel reportedly held a 43% mobile share in June, M1 22%, StarHub 21% and Simba 14%. In New Zealand, 2degrees and OneNZ have proposed combining their radio access networks into a jointly owned entity, with the Commerce Commission expected to decide next year. The China Telecom-China Unicom 5G arrangement, the largest such partnership, covers 1.5 million basestations and claimed $56 billion in capex savings.

Research is chipping at the software side too. A paper submitted to arXiv on 29 September, Demographic Pluralism, reports reducing Jensen-Shannon distance by 8.4% to 26.4% over Modular Pluralism across four backbones on GlobalOpinionQA and VITAL without opinion-distribution training data or task-specific fine-tuning. A second paper accepted at NeurIPS 2026, FluxLite, reports improvements over standard Feynman-Kac sequential Monte Carlo baselines of up to two orders of magnitude in terminal KL on a finite-state CTMC benchmark, and row-correlation MSE reductions on 2D Ising sampling by 5-7x in geometric mean and up to 55x at peak.

Microsoft Research argued on 23 September that running physical AI inference only on onboard GPUs limits robot performance, battery life and scale, and that offloading to edge or cloud GPUs improved task success rates and operating time across mobile manipulation workloads. It also introduced Kubernetes-based tooling to containerise and orchestrate robotics workloads across robots, edge and cloud.

One caveat runs through all of it. The performance figures from Magnitude, Ultium Cells and MaxLinear are the vendors' own, and the efficiency gains in the two arXiv papers are measured on academic benchmarks rather than production traffic.

Comments 0

Sources

11
  1. 01Launch HN: Magnitude (YC S25) - Self-optimizing inference engine for agentsEN
  2. 02MaxLinear claims new 'Puma 9' DOCSIS chip is a big cost-cutterEN
  3. 03Is AI jacking up healthcare costs?EN
  4. 04Driving An EV Now Costs Half As Much As Diesel, New Analysis ShowsEN
  5. 05The Human Cost Of Data Center PollutionEN
  6. 06GM's new EV battery tech will cut costs without sacrificing rangeEN
  7. 07Xpeng reportedly consolidates product lines to cut R&D costsEN
  8. 08Singapore, New Zealand operators seek partnerships to cut costsEN
  9. 09Demographic Pluralism: Inference-Time Modeling of Pluralistic Human Preference DistributionsEN
  10. 10FluxLite: Inference-Time Proposal Control for Discrete Diffusion ModelsEN
  11. 11Offloaded inference for real-world physical AI roboticsEN

All figures and quotations in this text come from the sources listed below.

Content prepared by the editorial team with AI assistance.

Rachel Nwosu

Rachel Nwosu

AI, models and technology

Rachel Nwosu covers AI, models and technology for FLASH24, working from public model documentation, benchmark releases and repository histories rather than press summaries, and she skips announcements that arrive without reproducible numbers. She checks training-data claims against dataset cards and reruns reported metrics where code is available. She spends much of her week interviewing researchers and engineers, tracking model launch calendars, and comparing vendor benchmarks with independent evaluations. Outside the desk she runs 3D printers, restores old computers, and tests how models learn from internet junk. She does not publish benchmark figures she cannot trace to a source.

Newsroom →

Comments

0
  1. No comments yet — be the first.

Write a comment

Comments are public. We do not publish abuse, spam or advertising.