Skip to content
World clockEU--:--UK--:--USA--:--CN--:--PLDEFRIT中文EN

portal about AI and technologyevents · analysis · interviews · technical background

Search
LIVE
›

Volunteers Train a 145M-Parameter Language Model With No Servers and No Funding

A volunteer project called Coop pushed its second-stage model live on 30 September, a 145M-parameter language model pretrained entirely on donated consumer hardware and free cloud tiers.

AI & modelsNewsGrace OkonkwoPublished: 1 October 20263 min readSources 5
Volunteers Train a 145M-Parameter Language Model With No Servers and No Funding

The repository lives at commonsense-ai/coop on GitHub. Its README says the whole training loop runs without a server, a daemon or any funding. Workers download the current checkpoint from a Hugging Face model repo, run local AdamW steps on a personal data shard, and submit a pseudo-gradient as a pull request against a public Hugging Face dataset repo. A GitHub Actions cron job runs every five minutes. It aggregates whatever arrived, takes one Nesterov outer step and uploads a new checkpoint.

Stage 2 is a decoder-only transformer with 12 layers, 14 heads, d=896 and a 1024-token context, trained on FineWeb-Edu.

Stage 1 is a 15M-parameter model trained for six days on TinyStories. It finished past its Chinchilla-optimal budget, with validation loss falling from 9.01 to 2.8, according to the project's own README.

What the volunteers actually proved

The README is unusually specific about failure modes, and that is where the engineering detail sits. One submission raced an aggregation tick and was accepted one step later at reduced staleness weight. Repeat rounds from one user merged into a single vote, which the project describes as a guard against token farming. A coop stop command flushed a half-finished round instead of discarding it. The inbox drains to zero every tick. Several volunteers on Apple Silicon and plain CPUs have trained the same outer step, and the results were averaged into one update.

Contributors do not need permission. Any free Hugging Face account with a write token can open a submission, and the maintainer grants no special access.

Weights and optimizer state live only on Hugging Face in safetensors format. Git holds code, config and a contributor ledger on a dedicated branch.

The project is also careful about what its numbers mean. Each tick evaluates the new checkpoint on a fixed held-out slice with a fixed seed, so a change between steps reflects the model rather than the sampling. The leaderboard fits a slope over validation loss against tokens, not outer steps, because a step is however much work happened to show up that tick. The README states that going down means the slope clears two standard errors. Anything less is reported as such.

Where it sits against the rest of the field

Small-model work is crowded. TechCrunch reported on 30 September that Google released Gemini 4 Argon to a select group of cyber partners through its Fairwind Program, a frontier model rather than a small one. CNBC reported the same day that Argon is already used internally to optimise memory at Google data centers. The contrast is the point. Coop's 145M parameters are roughly three orders of magnitude below a frontier release, and it runs on hardware volunteers already own.

The project is not the only effort in that direction, though the dossier only supports the Coop claims directly.

What it does show is a working distributed training mechanism with no central infrastructure budget. The README says the loop is production-proven rather than designed, and lists the specific edge cases it survived.

Coop ships a CLI installable through uv. A coop run latest command downloads the current checkpoint and generates text from a prompt.

The project recommends running the finished stage-1 TinyStories model instead, calling it far more coherent than a run still in progress.

Comments 0

Sources

5
  1. 01Coop: A small language model pretrained by volunteersEN
  2. 02Google releases Gemini 4 Argon, called its most powerful model yetEN
  3. 03Google rolls out Gemini 4 Argon, its most advanced AI modelEN
  4. 04OpenAI and Synopsys team up to build an AI model that designs chips like a seasoned engineerEN
  5. 05OpenAI announces 'dots' agent after scrapping launch of new AI model over safety concernsEN

All figures and quotations in this text come from the sources listed below.

Content prepared by the editorial team with AI assistance.

Grace Okonkwo

Grace Okonkwo

AI, models and technology

Grace Okonkwo covers AI, models and technology for FLASH24, working from primary sources such as model cards, API documentation and benchmark papers rather than vendor summaries. She checks training data provenance, evaluation conditions and reported scores against the underlying datasets before any figure reaches print. She interviews researchers and engineers directly, tracks release calendars from major labs, and compares successive model versions on the same tests. Her own self-hosting, home-network and documentation-reading habits feed straight into that desk, since she tests tools on her own hardware first. She does not publish benchmark claims without a reproducible method.

Newsroom →

Comments

0
  1. No comments yet — be the first.

Write a comment

Comments are public. We do not publish abuse, spam or advertising.