Open weights, open questions: Cloudflare's Clef takes on Jev as Amazon and OpenAI pile in
Cloudflare released two open-weight decision models, Clef and Clef-flash, on Thursday 1 October, claiming the pair beat Typesafe AI's Jev on accuracy while staying downloadable under an Apache 2.0 licence.

Cloudflare announced the two models on its blog. The Register reported the launch on 1 October, noting Clef costs $0.24 per million tokens, nearly six times Jev's $0.042 per million. Clef and Clef-flash are hosted on Workers AI and also published on Hugging Face under Apache 2.0.
The company says Clef is built on post-trained, frozen versions of Qwen3.8-27B and Qwen3.5-9B, running a prefill-only pass and then scoring choices in parallel. According to Cloudflare, Clef topped its own run against the Jev Decision Index, and beat Jev in three of four of Typesafe's own benchmarks, losing only on agent trace observability. The Register adds a caveat worth keeping: Cloudflare self-reported those scores, and they have not yet been reproduced on the official Decision Index. Clef also handles images and video, and supports a 64k context window. Jev handles up to 64k tokens per request, though its state plus longest individual question is capped at 32k, according to The Register.
The decision-model pile-up
Cloudflare is not alone. Amazon released its own Jev clone the same day, TechCrunch reported on 1 October, and Vercel's AI Gateway lists Convai Innovations' Laya as a free hosted endpoint as of 1 October, with promotional pricing that ends on 31 October 2026.
Sebastian Raschka's Substack post on 29 September walked through the history of text classification to explain why Jev became a phenomenon in technical communities, arguing its selling point is generality rather than raw classification quality. Typesafe has kept Jev's underlying architecture secret, while Cloudflare names its Qwen backbones outright. The economics are the subplot. Open-weight models are getting cheaper to run and easier to download, but Cloudflare's own price comparison shows open weights do not automatically mean cheap inference. Clef is open weight and open source, and still costs more per token than the closed model it is trying to beat.
Open weights on the research side
Academic work published this week leans the same way. A paper submitted to arXiv on 29 September introduces GoldiMask, a fine-tuning method for discrete diffusion language models that selects which tokens to reveal as context and weights the remaining prediction targets. The authors report the highest average accuracy in most evaluated settings across three backbones and three training datasets, with gains on reasoning and code generation, and fewer decoding iterations on GSM8K and MATH-500 under confidence-threshold parallel decoding.
A second paper, submitted the same day by authors from the University of Washington, Meta Superintelligence Labs, MIT and Trillium Labs, proposes Context Language Models, which treat the context as a file the model updates itself. The team reports 11.4% higher accuracy with 21.5% fewer FLOPs on BrowseComp-Plus, 5% higher scores with 59% fewer FLOPs on a 12-hour EdgeBench task, and an online reinforcement learning method that improves Qwen3.5-9B on BrowseComp-Plus by 47.6% while using 12% fewer FLOPs. Code is on GitHub under CC BY-NC 4.0.
Elsewhere, FLOORA, a family of small domain-specific models for architectural layout generation, reports that its 0.6B model beats much larger frontier models, with VLM judge win rates up to 92.0% on out-of-distribution real-world buildings and 96.0% on synthetic ones, and human evaluators picking it as best in 89.3% of cases. Dyad, another arXiv paper, adds an environment-conditioned action encoder to a frozen LLM and reports a 3.80% average absolute gain on ALFWorld with a 9B model.
Small models, small budgets
Two community projects published this week test how far open weights can stretch without corporate money. Coop, on GitHub, is pretraining a roughly 145M-parameter model from scratch on FineWeb-Edu using donated consumer hardware plus free tiers of Hugging Face and GitHub Actions, aggregating pseudo-gradients through pull requests with a DiLoCo-style outer step. The maintainers say stage 1, a 15M model on TinyStories, completed past its Chinchilla-optimal budget.
No server, no funding, no daemon: the whole training loop runs on donated consumer hardware plus the free tiers of Hugging Face and GitHub Actions.
PSSA, meanwhile, is a 2.7-billion-parameter model written in Rust with no transformer and no ML framework underneath, using a recurrent state-space layer plus an episodic memory bank. The project's README claims it learns faster than a transformer at matched parameters and generates text about twelve times quicker on the same CPU, and says gradients are checkable against a scalar reference path to around 3e-8. Treat those numbers as the authors' claims: the repository is the primary source and there is no independent reproduction in the dossier.
Inspect, an evaluation framework from the UK AI Security Institute and Meridian Labs, sits underneath several of these efforts. Its documentation lists more than 200 pre-built evaluations, support for over 20 model providers, and sandboxing via Docker, Kubernetes, Modal, Proxmox and Vagrant, plus local inference with HuggingFace, vLLM and SGLang.
The safety counter-current
Open weights are being released into a week of enforcement news. California's attorney general issued an investigative subpoena to OpenAI on Thursday, The Guardian reported, as part of a broader inquiry into cybersecurity incidents involving its models, including the July Hugging Face hack. Florida attorney general James Uthmeier filed a motion for a temporary injunction against five OpenAI entities and Sam Altman, Tom's Hardware reported on 30 September, asking a court to stop model development without independent third-party approval, among other requests.
Nvidia launched an Open Agent Safety Platform on 1 October that can quarantine agents in milliseconds, according to Tom's Hardware, and the FTC is running an industry-wide investigation into OpenAI, Anthropic and other labs, CNBC reported on 30 September. OpenAI also parted ways with three employees over mishandled sensitive information, first reported by the Wall Street Journal and covered by TechCrunch on 1 October. None of this is aimed at open weights specifically, but it shapes the market they are entering. The same week Clef shipped under Apache 2.0, the two most scrutinised closed labs were answering subpoenas and motions.
Sources
18- 01Clef: Open-source decision models, and new RL fine-tuning platformEN
- 02Cloudflare tries to outplay Jev with open-weight Clef modelsEN
- 03Amazon releases its own Jev clone as decision models flood the webEN
- 04Free hosted API for Laya, the open-weight decision modelEN
- 05Language models for text classification: From bag-of-words to JevEN
- 06Fine-Tuning Diffusion Language Models with Context Selection and Target WeightingEN
- 07Context Language ModelsEN
- 08Context Language Models (CLMs) official repositoryEN
- 09FLOORA: A Human-Aligned Domain-Specific Language Model for Architectural DesignEN
- 10Dyad: Extending Large Language Models with Native Typed Decision-MakingEN
- 11Coop: A small language model pretrained by volunteersEN
- 12PSSA: A non-transformer language model written from scratch in RustEN
- 13Inspect: An open-source framework for large language model evaluationsEN
- 14California issues investigative subpoena to OpenAI over rogue agents' hackingEN
- 15Florida attorney general asks judge to bar OpenAI from developing new AI models without third-party approvalEN
- 16Nvidia launches Open Agent Safety Platform to physically restrain rogue AI agentsEN
- 17FTC is investigating OpenAI, Anthropic and other AI companies over product risksEN
- 18OpenAI cuts ties with 3 safety researchers, WSJ reportsEN
All figures and quotations in this text come from the sources listed below.
Content prepared by the editorial team with AI assistance.
Comments
0- No comments yet — be the first.