Skip to content
World clockEU--:--UK--:--USA--:--CN--:--PLDEFRIT中文EN

portal about AI and technologyevents · analysis · interviews · technical background

Search
LIVE
›

Magnitude launches self-tuning inference engine claiming up to 2x faster than llama.cpp

Magnitude, a YC S25 startup, launched an open source inference engine on 30 September that it says compiles and tunes its kernels on the user's own hardware, running open models up to 2x faster than llama.cpp.

AI & modelsAnalysisGrace OkonkwoPublished: 1 October 20264 min readSources 5
Magnitude launches self-tuning inference engine claiming up to 2x faster than llama.cpp

The launch landed on Hacker News on 30 September and by the next morning had drawn 169 points and 85 comments. Magnitude's pitch is narrow and specific: instead of shipping kernels precompiled for broad hardware classes, the engine compiles and tunes them on the device before a model runs. The company claims 92% faster decode on Apple Silicon's Metal API and 19% on CUDA, according to its GitHub repository.

That framing matters for anyone paying inference bills, because the speed claim is a cost claim. Fewer seconds per token on your own hardware is money you are not sending to a hosted API.

What the repository actually says

Magnitude describes itself as Apache 2.0 licensed, with prompts, files and models staying on the machine and no internet needed once a model is downloaded. It ships as a desktop app that bundles the magnitude CLI, so there is no separate install step. Supported hardware is deliberately broad: any Apple Silicon, NVIDIA or AMD GPU, or a CPU alone. The repo says there is no fixed minimum, with smaller machines running smaller models.

The project lists hand-optimized kernels for popular open-weight families and claims 27% less memory per agent, freed when agents stop. It also says concurrent sessions share prefix caches to prevent slowdown. One click connects Pi, OpenCode, Hermes, OpenClaw, Codex, Claude Code, Oh My Pi and Cline, with anything else reachable through an OpenAI-compatible API.

The memory claim is the quieter one. Agent frameworks are memory-hungry because each session holds its own context; if Magnitude genuinely frees that memory when a session ends, it changes how many agents you can run on one box.

The billing problem it does not solve

Running locally removes token costs but not the cost of the machine, the electricity or the engineer tuning it. Independent cost tooling published on 23 September by developer Flavio Copes tries to make that comparison concrete. His inference cost calculator compares 27 models across OpenAI, Anthropic, Google, xAI, Mistral, OpenRouter and Workers AI, turning daily active users, calls per user and token counts into a monthly bill.

Copes is explicit that the numbers are estimates. The calculator, he writes, ignores batch pricing, enterprise discounts, image or tool surcharges and any caching layer you built yourself. It also flags that one provider publishes no cached-input discount, so the cache rate is ignored for it. Everything runs in the browser.

The dominant cost of an AI app is the LLM calls, and almost nobody can estimate it.

That gap between what a token costs and what a feature costs is where most planning arguments happen.

Context, memory and the rest of the stack

Two other projects in the same Hacker News cycle attack adjacent parts of the bill. FastRecall, posted on 28 September, positions itself as an OpenRouter for memory, letting context move between models and providers. It quotes 100 contexts with 500 messages each at a $7.00 monthly minimum before model-provider charges, and says it avoids LLM-based memory compaction because those techniques add latency, cost and retrieval errors.

Caspian, also posted on 28 September, goes after the build side. The desktop app turns a typed idea into a plan and then a live project, and it claims to match each change to the right model while showing cost, tokens and time before anything runs. Its free tier publishes up to five projects with a badge.

All three point the same direction: the interesting work has moved from raw model quality to the plumbing around it, where the meter actually spins.

Where the claims stop

None of these figures come with independent verification. The 2x headline and the 92% and 19% splits are Magnitude's own benchmarks against llama.cpp, published by the company. The $7.00 FastRecall figure is a platform estimate excluding provider charges, and the $0.00 shown for 500 recalls is an estimate, not a quote. Copes says outright that his output is a planning estimate, not a vendor quote.

Treat all of it as vendor arithmetic until someone reproduces it. That is not a knock on the projects; it is the normal state of inference pricing, where published rates, cache discounts and hardware assumptions change faster than anyone's spreadsheet.

Comments 0

Sources

5
  1. 01Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agentsEN
  2. 02Inference Cost CalculatorEN
  3. 03Show HN: HN.watch – Videos of all Hacker News postsEN
  4. 04Show HN: FastRecall, OpenRouter for MemoryEN
  5. 05Show HN: Caspian – Desktop app to build and launch small vibe-coded projectsEN

All figures and quotations in this text come from the sources listed below.

Content prepared by the editorial team with AI assistance.

Grace Okonkwo

Grace Okonkwo

AI, models and technology

Grace Okonkwo covers AI, models and technology for FLASH24, working from primary sources such as model cards, API documentation and benchmark papers rather than vendor summaries. She checks training data provenance, evaluation conditions and reported scores against the underlying datasets before any figure reaches print. She interviews researchers and engineers directly, tracks release calendars from major labs, and compares successive model versions on the same tests. Her own self-hosting, home-network and documentation-reading habits feed straight into that desk, since she tests tools on her own hardware first. She does not publish benchmark claims without a reproducible method.

Newsroom →

Comments

0
  1. No comments yet — be the first.

Write a comment

Comments are public. We do not publish abuse, spam or advertising.