Open weights week: Cloudflare Clef, Amazon, Google and DeepSeek move on the open model market
Cloudflare released two open-weight decision models on Thursday, claiming they beat TypeSafe's Jev on accuracy while admitting they cost close to six times more per million tokens. The launch caps a week in which Amazon, Google and Huawei all moved on open or low-cost model tooling.

Cloudflare announced the Clef family of decision models on Thursday, 1 October, hosted on its Workers AI platform and published as open weight on Hugging Face under an Apache 2.0 licence. The Register reported the same day that Clef and Clef-flash answer three types of bounded questions: yes/no, multiple choice and rankings.
That puts Cloudflare directly against TypeSafe's Jev, the decision model released two weeks earlier that turned into a technical phenomenon. According to The Register, Clef has an LLM backbone: specially post-trained, frozen versions of Qwen3.8-27B for Clef and Qwen3.5-9B for Clef-flash, each running a prefill-only pass and then scoring choices in parallel.
Cloudflare self-reported its benchmark scores against the Jev Decision Index, and The Register notes they have not yet been reproduced on the official leaderboard. Cloudflare also claims Clef beat Jev in three of four areas on TypeSafe's own benchmarks, losing only on agent trace observability.
The price is the awkward part
Clef costs $0.24 per million tokens, nearly six times Jev's $0.042 per million, The Register reported. Cloudflare's pitch is that it handles images and video as well as text, and that it supports a 64k context window. Jev can also handle up to 64k tokens across a request, though its state plus longest individual question is limited to 32k.
Cloudflare's own benchmark post describes a threat-intelligence test in which Clef fetched, rendered and classified a website in 2.2 seconds, against 4.7 seconds for the company's fastest general LLM, gpt-oss-120b, in the same workflow. The company says Clef classified the domain with a 95% chance it was a fashion site, 85% ecommerce and under 1% phishing. It also says it is open-sourcing the weights so customers can run them locally.
Two other decision-model launches landed in the same 48-hour window. Vercel's AI Gateway lists Laya, a "System One evaluation model" from Convai Innovations, with a free hosted API, an 8K context and a promotional pricing end date of 31 October 2026. TechCrunch reported that Amazon released its own Jev clone as decision models flooded the web.
Gemini 4 Argon, with the door half shut
Google's week cut the other way. The company rolled out Gemini 4 Argon, which TechCrunch called its most powerful model yet, on 30 September. The Guardian reported on 1 October that Google restricted access to the model over safety concerns, and The Decoder's benchmark write-up said Argon closes the gap with OpenAI and Anthropic but does not take a clear lead.
Anthropic, meanwhile, is arguing about training data rather than weights. The Guardian reported on 2 October that Anthropic is pushing for an opt-out model for Australian content, as ABC warned of "cannibalisation" of news.
The underlying tension in the open-weight debate is that published weights and published evaluation are different things. The UK AI Security Institute's Inspect framework, an open-source evaluation stack built with Meridian Labs, ships over 200 pre-built evaluations and supports more than 20 model providers, according to its documentation. It exists precisely because vendors still grade their own homework.
Independent work keeps arriving on the same schedule. A paper submitted to arXiv on 29 September by Loay Mualem and six co-authors introduces GoldiMask, a fine-tuning method for discrete diffusion language models that selects context tokens by approximately maximising a submodular objective and weights the remaining targets. The authors report the highest average accuracy in most evaluated settings across three backbones and three training datasets, with gains on reasoning and code generation.
Context as a file, and a 15M-parameter volunteer run
A second paper, submitted the same day by Rulin Shao and twelve co-authors, describes Context Language Models: models that treat their context as a file they can update themselves. The authors report 11.4% higher accuracy with 21.5% fewer FLOPs on BrowseComp-Plus, 5% higher scores with 59% fewer FLOPs on a 12-hour EdgeBench task, and a 47.6% improvement on Qwen3.5-9B when trained with their online reinforcement learning method. The code is on GitHub under CC BY-NC 4.0.
At the small end, a volunteer project called Coop is pretraining a roughly 145M-parameter model on FineWeb-Edu using donated consumer hardware, with pseudo-gradients submitted as Hugging Face pull requests and aggregated by a stateless GitHub Actions cron job. Stage 1, a 15M-parameter run on TinyStories, finished past its Chinchilla-optimal budget.
Hardware politics is moving too. Tom's Hardware reported on 1 October that DeepSeek and Huawei released open-source Ascend AI programming tools, including compute and communication libraries and Ascend support for TileLang, aimed at reducing reliance on Nvidia's ecosystem. Nvidia's answer to agent risk, an open Agent Safety Platform that can quarantine agents in milliseconds, was also announced on 1 October.
Not every claim this week survives contact with a benchmark. The Register's 30 September column noted that OpenAI, which trained on large volumes of web data during years of copyright fights, now calls distillation of its own reasoning a national security risk. OpenAI says a campaign that began on 1 July peaked at 16,000 requests from more than 4,000 users on 24 and 25 July and was fully disrupted by 28 July, with a core cluster linked to people associated with Moonshot AI. CNBC reported the company says its encryption, databases and stored user conversations were not breached.
The one number that argues against the open-weight cheerleading is $0.24 per million tokens. Open weights do not mean cheap inference, and Cloudflare's own comparison with Jev makes that plain.
Sources
17- 01Cloudflare tries to outplay Jev with open-weight Clef modelsEN
- 02Introducing Clef: our open-source decision models, and new RL fine-tuning platformEN
- 03Free hosted API for Laya, the open-weight decision modelEN
- 04Amazon releases its own Jev clone as decision models flood the webEN
- 05Google releases Gemini 4 Argon, called its most powerful model yetEN
- 06Google rolls out new Gemini AI model but restricts access over safety concernsEN
- 07Google Gemini 4 Argon closes the gap with OpenAI and Anthropic but doesn't take a clear leadEN
- 08Anthropic pushes for opt-out model for Australian content as ABC warns of 'cannibalisation' of newsEN
- 09Inspect: An open-source framework for large language model evaluationsEN
- 10Fine-Tuning Diffusion Language Models with Context Selection and Target WeightingEN
- 11Context Language ModelsEN
- 12Context Language Models (CLMs) official repositoryEN
- 13Coop: A small language model pretrained by volunteersEN
- 14DeepSeek and Huawei release open-source Ascend AI programming tools to reduce reliance on Nvidia ecosystemEN
- 15Nvidia launches Open Agent Safety Platform to physically restrain rogue AI agentsEN
- 16Irony alert: OpenAI whines that Chinese model stole its special IP that it stole from everybody elseEN
- 17AI race heats up as OpenAI flags alleged model-copying campaignEN
All figures and quotations in this text come from the sources listed below.
Content prepared by the editorial team with AI assistance.
Comments
0- No comments yet — be the first.