Skip to content
World clockEU--:--UK--:--USA--:--CN--:--PLDEFRIT中文EN

portal about AI and technologyevents · analysis · interviews · technical background

Search
LIVE
›

Self-hosted inference, not API pricing, is where the cost gap is widest

A September 2026 comparison of seven self-hosted inference orchestrators found the cheapest option offers no clustering at all, while the tools that span multiple GPUs carry caveats their vendors' READMEs do not hide.

AI & modelsNewsRachel NwosuPublished: 27 September 20264 min readSources 1
Self-hosted inference, not API pricing, is where the cost gap is widest

Nexlab published the survey on 20 September. It lines up LocalAI, exo, GPUStack, Xinference, Ollama, vLLM and CoderAI on the things a buyer actually cares about: multi-machine support, cache-aware routing, cloud burst and training. Star counts come from the GitHub API on the same date. Feature cells come from the projects' own README and docs. Where the author could not confirm something, the cell says so rather than guessing.

The ranking is not what a procurement deck would predict. Ollama, with 181,000 stars, is the right answer for a laptop or one desktop and nothing more. One machine, one model at a time, and no cluster story beyond Open WebUI round-robining several Ollama URLs. Anyone who needs more than that is told to stop reading at that point. Mozilla's llamafile sits in the same column: a single executable for Linux, macOS, Windows and the BSDs, single machine by design.

The projects that do span hardware are the expensive ones to operate. vLLM does tensor and pipeline parallelism over Ray and prefix caching per instance, but manages no models, users or placement. llama.cpp, at 129,000 stars, reaches across boxes with an rpc-server on each machine and a --rpc flag on the host. Neither is an orchestrator in the sense most buyers mean. They are what LocalAI, GPUStack, Xinference and CoderAI run underneath.

The macOS exception and the Linux gap

exo gets the strongest single endorsement in the piece. Zero-config discovery, ring, pipeline and tensor partitioning proportional to each device's memory, MLX underneath, and RDMA over Thunderbolt 5 on recent Macs. The vendor's own number is 3.2x scaling on four devices for tensor parallelism. But as of September 2026 exo is CPU-only on Linux, with NVIDIA and AMD support listed as under development. It serves language models, with image generation behind a feature flag. The survey's verdict is blunt: if your hardware is Apple Silicon, nothing else comes close; if it is not, exo is not for you yet.

LocalAI is the broadest of the general-purpose options. Text, image, video, audio, embeddings and rerank, each backend a gRPC service in its own OCI image, no GPU required, Helm charts, and since June 2026 a real distributed mode. A --p2p flag generates a shared token, instances discover each other over libp2p and EdgeVPN, requests federate to the least loaded node, and a NATS-based v3 router is aware of VRAM and prefix caches. Backend images are cosign-signed. What it does not do: rent a GPU, split a non-LLM request over machines, or train. Raw LLM throughput trails a dedicated engine by some tens of percent, the author writes, because the generality costs.

GPUStack and Xinference are grouped as supervisor-plus-workers designs with web consoles, both backed by companies, and both what the author sees deployed as clusters in Asia. GPUStack, 5.7k stars, is the more operational: users and roles, API keys with metering, Prometheus and Grafana, automatic recovery of failed models, Ray worker logs in the UI, and support for nine accelerator vendors including Ascend, Hygon and MThreads. Xinference, at 9.6k stars, adds shared KV cache across replicas when it runs vLLM on the back end.

It sits in front of a hundred providers and your own endpoints with keys, budgets and spend tracking, and never runs a model.

That line is about LiteLLM, which the survey says is often mistaken for a self-hosting solution. It is not one. It routes between endpoints and to clouds, and runs no weights itself.

The two entries with the strongest production framing are not primarily about cost. NVIDIA Dynamo, 8.1k stars, and llm-d, 4.6k, do disaggregated prefill and decode with KV-aware routing, and both require Kubernetes on NVIDIA hardware. SkyPilot and dstack, 10.6k and 2.3k stars, schedule your jobs across clusters and can burst to cloud, which is a different problem from serving a model efficiently. CoderAI, new and unstarred, is the author's own project, disclosed up front, and claims per-model budgeted cloud burst plus distributed LoRA, with cosign-signed images on Linux CUDA and Vulkan.

Two projects are named only to be excluded. Petals was last committed to in 2024. Hugging Face TGI was archived in March 2026. The survey also notes that vLLM's production-stack is the only Kubernetes option in the table marked as production-ready among the engines themselves, and that signed images are rare: LocalAI and CoderAI use cosign, and most of the field does not.

For teams weighing where inference money goes, the survey offers no price table. It offers something less comfortable: a map of which orchestrators can actually use a second GPU, which cannot, and which ones will tell you so in their own documentation.

Comments 0

Sources

1
  1. 01Self-hosted inference orchestrators compared: LocalAI, exo, GPUStack, vLLMEN

All figures and quotations in this text come from the sources listed below.

Content prepared by the editorial team with AI assistance.

Rachel Nwosu

Rachel Nwosu

AI, models and technology

Rachel Nwosu covers AI, models and technology for FLASH24, working from public model documentation, benchmark releases and repository histories rather than press summaries, and she skips announcements that arrive without reproducible numbers. She checks training-data claims against dataset cards and reruns reported metrics where code is available. She spends much of her week interviewing researchers and engineers, tracking model launch calendars, and comparing vendor benchmarks with independent evaluations. Outside the desk she runs 3D printers, restores old computers, and tests how models learn from internet junk. She does not publish benchmark figures she cannot trace to a source.

Newsroom →

Comments

0
  1. No comments yet — be the first.

Write a comment

Comments are public. We do not publish abuse, spam or advertising.