Small models, big claims: volunteers train Coop, Google ships a restricted Gemini 4
A 145-million-parameter model pretrained by volunteers on donated hardware went live on 30 September, the same week Google restricted its most advanced model to vetted cyber partners and Amazon open-sourced a 2B decision model.

Coop, a project published on GitHub on 30 September, says it is pretraining a roughly 145-million-parameter language model on FineWeb-Edu. No server. No funding. No daemon.
Training runs on donated consumer hardware plus the free tiers of Hugging Face and GitHub Actions, according to the project's README. The mechanism is the interesting part: volunteers download a checkpoint, run a handful of local AdamW steps on a personal data shard, and submit a pseudo-gradient as a pull request against a public Hugging Face dataset repo. A GitHub Actions cron job then reads the checkpoint and the open pull requests, drops stale submissions, clips and gates the rest, aggregates them with a trimmed mean or geometric median, takes one Nesterov step, uploads a new checkpoint and closes the processed PRs. Coop says the loop is production-proven, not designed.
Multiple volunteers on Apple Silicon and plain CPU machines trained the same outer step and had it averaged into one update. A submission that raced a tick was accepted a step later at reduced staleness weight. Repeat rounds from one user merged into a single vote. A "coop stop" flushed a half-finished round rather than discarding it.
The project is careful about evaluation. Each tick scores the new checkpoint on a fixed held-out slice with a fixed seed, so a change between steps reflects the model rather than eval sampling, and appends one point to a ledger file. A single validation loss says nothing, the README notes: outer steps move it up as often as down, so the leaderboard fits a slope against tokens rather than outer steps.
Frontier labs move the other way
The contrast with the commercial labs is stark. Alphabet launched Gemini 4 Argon on 30 September, calling it its most advanced model, and CNBC reported that it is being rolled out only to select cyber partners first. The Guardian reported on 1 October that Google is withholding the model from the public, citing misuse risk, and quoting Koray Kavukcuoglu, Google's chief AI architect, writing that "safely releasing frontier capabilities at this level requires a phased approach". Tulsee Doshi, Google's Gemini model product lead, told CNBC that starting the rollout this way "gives us more confidence, but also enables us to put a model that is trained and strong in cyber defense in the hands of defenders as soon as possible". Google says Argon ties with OpenAI's GPT-6 Astra and Grok 4.7 on cybersecurity benchmarks and leads on the Vals Index. It also says the model is already used internally to optimize data center memory, freeing hundreds of terabytes.
Small, open and local is the other pole. TechCrunch reported on 1 October that AWS released Strands Decider 2B, an open source decision model built on the torso of Qwen3.5-2B that returns calibrated choices instead of text. Amazon distinguished engineer Marc Brooker told TechCrunch the model came out of customer conversations and offers "a workflow step that can be structured in a way that is more reliable, thanks to the confidence scores, thanks to the closed domain of answers, [and is] lower latency, potentially lower cost".
"I get that people think it's a gold rush, but they might be underestimating the difficulty of making the models actually smart," TypeSafe CEO Diogo Almeida told TechCrunch.
TypeSafe's Jev is the reference point. InfoQ reported on 1 October that Jev returns typed probabilistic decisions, prices input at $0.042 per million tokens, makes output free, carries a 32,000-token context window and quotes 70ms to 500ms end-to-end latency. Vercel said Jev reached nearly 13% of paid teams within 24 hours, twice the share of the GPT-5.6 family. An analysis of 12,759 launch tweets by OpenChamber put user-reported speedups at a median of 7x against the 193.6x headline, and cost savings at a median of 30x.
Security is the awkward part
Cheap, local models also mean cheap, local extraction. OpenAI said in a blog post on 30 September that it disrupted an "adversarial distillation" campaign that began on 1 July, spiked to 16,000 requests from more than 4,000 users on 24 and 25 July, and was fully shut down by 28 July. CNBC reported that OpenAI linked a core cluster to people associated with China's Moonshot AI, the Kimi developer, and that the operators did not breach encryption, databases or stored conversations.
The Register noted on 30 September that OpenAI itself trained on vast amounts of web content, and that Anthropic's Claude Opus 5.5 carries a distillation defense called "preserved thinking". The Decoder reported on 1 October that the same trick kept working for weeks on Microsoft Azure, citing researcher Joachim Schaeffer, who wrote on X: "We stole reasoning. Again."
Meanwhile the hardware keeps shrinking. Tom's Hardware reported on 30 September that a developer using the handle "stmonty" trained a roughly 12.5-million-parameter world model on a single RTX 3080 Ti to play Pokémon Red, learning button functions by predicting compressed 192-number summaries of the next screen across 42,382 grayscale frames. And Ideogram said on 1 October that Ideogram 4.5 offers four quality tiers from 0.8 to 22 cents per image at native 2K resolution, with an open-weight release coming soon.
Sources
11- 01Coop: A small language model pretrained by volunteersEN
- 02Google rolls out Gemini 4 Argon, its most advanced AI modelEN
- 03Google rolls out new Gemini AI model but restricts access over safety concernsEN
- 04Amazon releases its own Jev clone as decision models flood the webEN
- 05TypeSafe AI Releases Jev: A Decision-Only Model That Returns Typed Probabilities Instead of TextEN
- 06OpenAI says actors linked to China-based Moonshot AI spearheaded a campaign to extract its models' hidden reasoningEN
- 07AI race heats up as OpenAI flags alleged model-copying campaignEN
- 08Irony alert: OpenAI whines that Chinese model stole its special IP that it stole from everybody elseEN
- 09OpenAI says it stopped a campaign to steal its models' reasoning, but the trick still worked on AzureEN
- 10Developer trains a small AI on a single RTX 3080 Ti gaming GPU to 'play' Pokémon RedEN
- 11Ideogram says its new model can edit part of an image without messing up the restEN
All figures and quotations in this text come from the sources listed below.
Content prepared by the editorial team with AI assistance.
Comments
0- No comments yet — be the first.