Skip to content
World clockEU--:--UK--:--USA--:--CN--:--PLDEFRIT中文EN

portal about AI and technologyevents · analysis · interviews · technical background

Search
LIVE
›

Small models move to the edge as Google locks down its biggest one

Google began rolling out Gemini 4 Argon on Wednesday, but only to vetted cybersecurity partners, while the same week produced an open 145M-parameter model trained by volunteers and a 2B decision model from AWS. The frontier and the edge are pulling apart.

AI & modelsAnalysisGrace OkonkwoPublished: 2 October 20267 min readSources 15
Small models move to the edge as Google locks down its biggest one

Google unveiled Gemini 4 Argon on Wednesday, calling it its most advanced model yet and restricting it to a vetted group of cybersecurity partners. CNBC and The Guardian both reported that restraint, and it is now the organising fact of the week in AI. The largest labs are slowing down their biggest models.

Smaller ones keep shipping. The Argon rollout is phased by design. Koray Kavukcuoglu, Google's chief AI architect, wrote in a blog post that "safely releasing frontier capabilities at this level requires a phased approach", according to The Guardian. Google says it is giving the US government early access and will gather tester feedback before a wider release. The company also claims Argon sets a record in real-world software engineering, ties for first in cybersecurity and leads the Vals Index, where CNBC reports it is ahead of OpenAI's GPT-6 Astra and recent Anthropic models.

The small-model counter-current

Against that, three releases in the same window went the other way: open, small, and available now. On 30 September a GitHub project called Coop published a ~145M-parameter model pretraining from scratch on FineWeb-Edu. According to the project's README, the whole training loop runs on donated consumer hardware plus the free tiers of Hugging Face and GitHub Actions. No servers. No funding. No daemon.

"Stage 1 (15M on TinyStories) completed past its Chinchilla-optimal budget, proof that the whole mechanism works."

The mechanics matter more than the parameter count. Volunteers download the current checkpoint, run local AdamW steps on a personal data shard, and submit a pseudo-gradient as a pull request against a public Hugging Face dataset repo. A stateless GitHub Actions cron job, scheduled every five minutes but firing anywhere from minutes to hours apart, reads the checkpoint and open inbox PRs, drops over-stale submissions, aggregates the rest, takes one Nesterov outer step and closes the processed PRs. Weights and optimizer state live only on Hugging Face; Git holds code, config and the contributor ledger.

The project claims the loop is production-proven rather than designed: multiple volunteers on different machines, Apple Silicon and plain CPU, have trained the same outer step and been averaged into one update. A submission that raced a tick was accepted a step later at reduced staleness weight. Repeat rounds from one user merged into a single vote. The README is explicit that a single validation loss says nothing, and that the leaderboard fits a slope against tokens rather than outer steps.

Decision models, and a crowded field

AWS released Strands Decider 2B, an open source decision model inspired by TypeSafe's Jev, TechCrunch reported on 1 October. It sorts between pre-decided options and returns a confidence measure, built on the torso of Qwen3.5-2B rather than generating text. Marc Brooker, an Amazon distinguished engineer, told TechCrunch he built it after seeing Jev and that it briefly reached the top spot on the Jevbench ranking for models of its size. "What originally piqued my interest in this class of models was that they make a perfect decider for a workflow step," he said.

TypeSafe's own position is less triumphal. CEO Diogo Almeida told TechCrunch: "I get that people think it's a gold rush, but they might be underestimating the difficulty of making the models actually smart." TypeSafe released Jev itself, described by InfoQ as returning typed probabilities instead of text. A separate arXiv paper from 1 October, "More Choices, Fewer Decisions", examines ordinal-scale bias in JEV-like direct-decision models. The category is still being stress-tested rather than settled.

Elsewhere in the same 24 hours: Ideogram said its Ideogram 4.5 image editor touches only the region a user selects, with four quality tiers from 0.8 to 22 cents per image at native 2K, and an open-weight release promised, per The Decoder. BMW confirmed at its Capital Markets Day that an entry-level Neue Klasse EV arrives in 2028 with a focus on Europe, Electrek reported, with BMWBlog-sourced reports pointing to a five-door hatchback codenamed NBO.

What the security news says about capability

The most consequential small-model story of the week may not be a release at all. OpenAI said it disrupted an adversarial distillation campaign that ran through July, with spikes of 16,000 requests from more than 4,000 users on 24 and 25 July and related activity across 15,000 accounts, fully shut down by 28 July. CNBC, Tom's Hardware and The Register all carried the company's account, which links a core cluster of the activity to people associated with Moonshot AI. Moonshot did not immediately respond to requests for comment, per CNBC and The Register.

The technical detail is the interesting part. OpenAI says its encryption was not broken and no database was compromised. Instead, operators moved encrypted reasoning between conversations and asked a model in a second session to decrypt and write it out, a technique researchers Joachim Schaeffer and colleagues had already documented. OpenAI credits them by name. The Decoder reported that the researchers published an update the same day arguing that securing your own API does not secure the cloud providers that also sell access to the same models, and that the trick kept working for weeks on Microsoft Azure.

Smaller models are not a footnote to that story, they are the enabling layer. A cheaper sibling model from the same family is what turns an encrypted reasoning packet into readable text. That is the same economics AWS is selling to customers: a workflow step that is lower latency and potentially lower cost, as Brooker put it, with confidence scores and a closed domain of answers.

Policy, price and the gap between markets

Distributed deployment is spreading beyond phones. ILSR's quarterly Big Impact of Small Solar report, covered by pv magazine on 1 October, put US small-scale solar additions at almost 1.6 GW in Q2 2026, alongside over 2.5 GWh of behind-the-meter storage. MIT Technology Review reported on 1 October that startups including PopWheels, Copper, Every Electric and David Energy are deploying small batteries in e-bike swap cabinets, induction stoves and plug-in units, avoiding grid upgrades and permitting. David Hammer of PopWheels said at a New York Climate Week event on 24 September that about 50 cabinets and 2,500 batteries are in circulation in New York City; Sam Calisch of Copper said shipping every US stove with a battery would add up to tens of gigawatts.

The same split runs through cars. Tesla upgraded every Model 3 it sells in China to a 16-inch screen and enabled up to 2.2 kW of AC power export across its Model 3 and Model Y lineup there, announced on 1 October, CnEVPost reported. The US Model 3 keeps a 15.4-inch display and is excluded from the V2L feature. Tesla's US Model 3 starts at $36,990, Electrek noted. GM, meanwhile, sold 670,974 vehicles in the quarter, down 5.5%, with the Equinox EV down 92.4% to 1,905 cars after the $7,500 US purchase credit ended, The Next Web reported. Toyota's sales rose 0.6% to 633,223 with electrified models up 28.5%.

None of this is a clean story about small models winning. Google's Argon is being held back for safety reasons, and Anthropic keeps Claude Mythos Preview restricted to trusted organisations, The Guardian noted. Florida attorney general James Uthmeier has asked a judge to bar OpenAI from developing new models without third-party approval, Tom's Hardware reported on 30 September, in a suit that also covers the July Hugging Face hack. The same outlets that cover volunteer-trained 145M models also cover distillation campaigns at 16,000 requests a day.

What the week does show is that capability is diffusing faster than control. A 2B decider runs locally. A 145M model trains on donated laptops. The encrypted reasoning that leaks between sessions is worth stealing precisely because a cheap model can read it. The frontier labs are answering with phased releases and vetted partner lists. Whether that holds depends less on benchmark tables than on what a 2B model on a desk can do next month.

Comments 0

Sources

15
  1. 01Coop: A small language model pretrained by volunteersEN
  2. 02Google rolls out Gemini 4 Argon, its most advanced AI modelEN
  3. 03Google rolls out new Gemini AI model but restricts access over safety concernsEN
  4. 04Amazon releases its own Jev clone as decision models flood the webEN
  5. 05TypeSafe AI Releases Jev: A Decision-Only Model That Returns Typed Probabilities Instead of TextEN
  6. 06More Choices, Fewer Decisions: Ordinal-Scale Bias in JEV-like Direct-Decision ModelsEN
  7. 07Ideogram says its new model can edit part of an image without messing up the restEN
  8. 08BMW is launching a smaller, more affordable EVEN
  9. 09OpenAI says it stopped a campaign to steal its models' reasoning, but the trick still worked on AzureEN
  10. 10OpenAI says actors linked to China-based Moonshot AI spearheaded a campaign to extract its models' hidden reasoningEN
  11. 11Irony alert: OpenAI whines that Chinese model stole its special IP that it stole from everybody elseEN
  12. 12AI race heats up as OpenAI flags alleged model-copying campaignEN
  13. 13U.S. small-scale solar adds 1.6 GW in Q2EN
  14. 14How smaller, distributed batteries could help the gridEN
  15. 15Tesla gives Model 3 minor upgrades in China, keeps prices unchangedEN

All figures and quotations in this text come from the sources listed below.

Content prepared by the editorial team with AI assistance.

Grace Okonkwo

Grace Okonkwo

AI, models and technology

Grace Okonkwo covers AI, models and technology for FLASH24, working from primary sources such as model cards, API documentation and benchmark papers rather than vendor summaries. She checks training data provenance, evaluation conditions and reported scores against the underlying datasets before any figure reaches print. She interviews researchers and engineers directly, tracks release calendars from major labs, and compares successive model versions on the same tests. Her own self-hosting, home-network and documentation-reading habits feed straight into that desk, since she tests tools on her own hardware first. She does not publish benchmark claims without a reproducible method.

Newsroom →

Comments

0
  1. No comments yet — be the first.

Write a comment

Comments are public. We do not publish abuse, spam or advertising.