Inference Price Signals Point Both Ways as Hardware Bet Splits from Cloud
Q/C Technologies said on 1 October it is working with Sandia National Laboratories' Center for Integrated Nanotechnologies on optical processing for AI inference, the same day MaxLinear's Puma 9 pitch and a $32-per-tank diesel gap kept the cost argument running.

Two announcements landed within hours of each other on 1 October. Neither settles the question of what inference costs. Q/C Technologies said it will collaborate with the Center for Integrated Nanotechnologies at Sandia National Laboratories on nanophotonic components for a planned optical processing unit aimed at AI inference, according to HPCwire.
Separately, MaxLinear's new Puma 9 DOCSIS chip, announced 29 September, is billed as cutting customer premises equipment costs by 30% to 50%. Q/C's arrangement with CINT is research, not a product. The company says the project will identify which optical operations and device technologies matter for a scalable OPU architecture, and will address nonlinear operations, memory, precision, optical loss and integration with electronics. Sandia, which operates CINT, has roughly 17,000 employees and an annual budget of about $5 billion, HPCwire reports. Nobody has published a benchmark for an optical inference part at scale.
Cheaper silicon, cheaper cable boxes
MaxLinear's numbers are more concrete, though still vendor-supplied. The Puma 9 supports DOCSIS 3.1, D3.1+ and D4.0. The company claims the D3.1+ configuration can deliver up to 12 Gbit/s downstream. MaxLinear expects modems and gateways based on the chip to appear in 2027. It says it has secured "multiple commitments." The chip moves from Intel x86 to ARM, which an industry source told Light Reading will reduce power draw and some licensing costs.
"As manufacturers are incentivized to move their capacity to DDR5, we think the DDR5 ecosystem and pricing and supply will become more relaxed than the stress that we see with DDR4," said Puneet Sethi, SVP and GM of MaxLinear's network infrastructure and carrier business unit.
The memory point matters beyond cable. Dell'Oro Group's Jeff Heynen told Light Reading that most vendors will move to DDR5 for advanced Wi-Fi 8 units because of the onboard memory those devices need. That is a supply chain signal, not an AI inference one. But it is the same pressure: capacity shifting toward newer memory, with older parts getting tighter.
On the cloud side, the picture is messier. Recent headlines in the dossier, which we are not treating as citable facts, point to Nebius acquiring Inferize and raising prices, AWS lifting GPU capacity prices, and OpenAI cutting model prices. Those pull in opposite directions. The dossier does not resolve them.
Hardware bets versus fuel bills
Magnitude, a Y Combinator S25 company, published an open source inference engine on GitHub on 30 September that it says tunes kernels on the user's own device. The README claims up to 2x faster performance than llama.cpp, with 92% faster decode on Metal and 19% on CUDA, and 27% less memory per agent. It works on Apple Silicon, NVIDIA, AMD or CPU only, and connects to Pi, OpenCode, Hermes, Codex, Claude Code, Cline and others. The claims are self-reported and the repository carries an Apache 2.0 licence.
That is a local-hardware argument. The cloud argument is different, and it is being made with price tags. Transport & Environment, in analysis published by CleanTechnica on 29 September, says EV drivers in Europe pay less than half what diesel drivers pay per kilometre, based on average consumption of 6.9 L/100 km for diesel and 20.2 kWh/100 km for battery electric, with electricity at €0.343/kWh. Diesel drivers are paying €32 more per tank than at the start of the year, of which about €16 is extra refinery margin, per T&E. The EU-27 weighted average pump price comes from the EC Weekly Oil Bulletin dated 21 September 2026, while the refinery margin figure, from ECB estimates, is a week older.
The comparison is not like-for-like with GPU pricing, and T&E says its charging cost estimate is likely an upper bound. But it is the same argument in a different sector: electrify the thing that burns fuel, and the per-unit cost falls.
What the other numbers say
In telecoms, operators are chasing cost cuts through sharing rather than new silicon. StarHub and M1 are in merger talks in Singapore, where they already share 5G spectrum and radio access network through a joint company called Antina. Singtel reportedly held a 43% share of the mobile market in June, with M1 at 22%, StarHub at 21% and Simba at 14%, per Light Reading. In New Zealand, 2degrees and OneNZ have proposed combining their radio networks into a jointly owned entity, with a Commerce Commission decision expected next year.
The largest example cited is the China Telecom and China Unicom 5G arrangement, which Light Reading says covers 1.5 million shared basestations and claimed $56 billion in capex savings. If that number holds, network sharing has delivered more measurable cost reduction than any inference chip announcement in this dossier.
Elsewhere, the cost pressure is showing up in R&D budgets. Xpeng consolidated its product lines from four to two, local outlet 36Kr reported on 28 September, citing industry sources. The company delivered 243,111 vehicles in the first eight months of 2026, down 10.49% year on year, while second-quarter R&D expenses rose 32.1% to 2.91 billion yuan. Its second-quarter net loss was 1.34 billion yuan ($199 million), wider than 480 million yuan a year earlier but narrower than 1.78 billion yuan in the first quarter. Vehicle margin was 12.1%, flat quarter on quarter and below 14.3% a year earlier.
GM is making a different calculation. Ultium Cells, its joint venture with LG Energy Solution, said on Tuesday it will produce prismatic lithium manganese rich cells at Spring Hill, Tennessee, which the company says will be the first plant in the world to mass-produce them. GM claims 33% higher energy density than LFP at comparable cost, with over 400 miles of range planned for electric trucks and full-size SUVs, and a first vehicle launch expected in 2028. Facility upgrades start later this year and are due for completion by 2028, adding 500 jobs to bring the workforce to 1,700.
Microsoft Research made a related argument on 23 September, saying that running physical AI inference exclusively on onboard GPUs can limit robot performance and battery life, and that offloading to edge or cloud GPUs improved task success rates and enabled larger models across representative mobile manipulation workloads.
None of this produces a single inference price. The dossier offers vendor claims, one research collaboration, a grid of delivery and margin numbers, and a diesel comparison that is explicitly an upper bound. What it does show is that the cost argument is being fought on at least four fronts: local kernels, optical research, cloud pricing and network sharing. The newest of them, Q/C's Sandia collaboration, is also the furthest from a benchmark.
Sources
10- 01Q/C Technologies Collaborates with Sandia on Optical Computing for AI InferenceEN
- 02MaxLinear claims new 'Puma 9' DOCSIS chip is a big cost-cutterEN
- 03Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agentsEN
- 04Driving An EV Now Costs Half As Much As Diesel, New Analysis ShowsEN
- 05Singapore, New Zealand operators seek partnerships to cut costsEN
- 06Xpeng reportedly consolidates product lines to cut R&D costsEN
- 07GM's new EV battery tech will cut costs without sacrificing performance or rangeEN
- 08Offloaded inference for real-world physical AI roboticsEN
- 09Is AI jacking up healthcare costs?EN
- 10The Human Cost Of Data Center PollutionEN
All figures and quotations in this text come from the sources listed below.
Content prepared by the editorial team with AI assistance.
Comments
0- No comments yet — be the first.