Skip to content
World clockEU--:--UK--:--USA--:--CN--:--PLDEFRIT中文EN

portal about AI and technologyevents · analysis · interviews · technical background

Search
LIVE
›

Open weights dominate AI token processing as Google restricts free access

Open-weight models processed the majority of AI tokens for the first time on 28 September, according to industry data, a shift occurring as Google restricts free access to its proprietary Gemini models starting 9 October.

AI & modelsExplainerGrace OkonkwoPublished: 4 October 20267 min readSources 7
Open weights dominate AI token processing as Google restricts free access

The balance of power in large language models is shifting toward open architectures.

On 28 September, open-weight models processed the majority of AI tokens for the first time, marking a significant milestone in the industry's transition away from closed, proprietary systems. This development coincides with a wave of new releases and strategic moves by major players who are either embracing open standards or tightening the screws on their paid services.

Google is leading the charge in restricting free access.

According to an updated support document cited by 9to5Google on 3 October, the company will limit which Gemini models free users and AI Plus subscribers can access starting 9 October. Users without a subscription will be restricted to the (3.5) Flash-Lite model, losing access to the more powerful (3.6) Flash and (3.1) Pro models. AI Plus subscribers, who pay $4.99 per month, will also lose access to the Pro tier, retaining only Flash-Lite and Flash. The company justifies these changes by citing compute-based usage limits introduced in May. The move mirrors OpenAI's freemium model, where higher-tier models are reserved for paying customers. The Decoder reported on 4 October that this setup effectively locks out budget users from the most capable models, pushing them toward the $19.99 AI Pro tier or the $99.99 AI Ultra plan.

This restriction on proprietary models has accelerated the adoption of open-weight alternatives.

Aleph Alpha, a German AI company, released Kolibri, a 78-billion parameter model, under the Apache 2.0 license. According to Dealroom on 3 October, the full weights were placed on Hugging Face, allowing researchers and developers to download, fine-tune, and deploy the model locally. This is part of a broader trend where companies are releasing their most capable models as open weights to build trust and foster an ecosystem around their technology. Startup Fortune confirmed the release on the same day, noting that the Apache 2.0 license permits commercial use without restriction.

Google has also joined the open-weight movement, albeit with smaller models. On 6 October, the company released EmbeddingGemma 2 under the Apache 2.0 license.

While this model is designed for embedding tasks rather than general conversation, its release signals a strategic commitment to contributing to the open-source ecosystem. This approach allows Google to maintain a foothold in the open-weight space without directly competing with its proprietary Gemini models for the most demanding general-purpose tasks.

Meanwhile, Mistral, a French AI company, unveiled its 1.05-trillion parameter 'Le Chonk' Mixture-of-Experts (MoE) model in public preview on 5 October. The model is open-weight, and its massive parameter count, combined with the MoE architecture, aims to deliver high performance while managing inference costs. This release is significant because it demonstrates that open-weight models can scale to the same levels of complexity as proprietary systems. The model's availability in public preview allows developers to test its capabilities and provide feedback before a full release.

These developments are not just about model releases; they are about the infrastructure and tools that support them.

On 4 October, Jerry Gamblin published a detailed analysis of using local models for security tasks. He tested Cloudflare's Clef and Clef Flash models, which are open-weight and fine-tuned from Qwen, on a MacBook. Gamblin used these models to analyze 27,489 CVEs (Common Vulnerabilities and Exposures) published in the last 60 days. The experiment demonstrated that open-weight models can perform complex classification tasks locally, without sending data to the cloud. This is a critical use case for security-sensitive industries where data privacy is paramount. The fact that these models run locally on consumer hardware demonstrates the efficiency gains in modern model architecture.

The rise of open weights is also driving innovation in tooling and infrastructure.

Spinifex, an open-source project built by Mulga, aims to recreate the AWS service surface on bare-metal and edge devices. According to its GitHub repository, Spinifex allows users to run existing AWS CLI, SDK, and Terraform workflows against EC2, EBS, S3, VPC, IAM, EKS, and RDS on their own servers. This project is vital for the open-weight ecosystem because it provides a way to deploy AI models in air-gapped environments, such as government facilities or industrial sites, where cloud access is restricted. By making it possible to run cloud-native software without a hyperscaler, Spinifex lowers the barrier to entry for organizations that want to use open-weight models but cannot rely on public cloud services.

Another tool that supports the open-weight ecosystem is Docent, an open-source private AI assistant that runs in the terminal. According to its GitHub repository, Docent allows users to chat with their PDFs and Office files, search the web, and connect MCP servers, all on their own machine. The project is built in Rust and uses a local agent for GPU inference. This type of tool is essential for developers who want to use open-weight models for personal or professional tasks without sending their data to third-party servers. The ability to run a full-featured AI assistant locally, with privacy as a core feature, is a strong selling point for open-weight models.

The trend toward open weights is also influencing the way models are evaluated and compared.

On 4 October, a paper titled 'No Model Required: Text Entropy Rate Filtering Mitigates Iterative Fine-Tuning Collapse' was published on arXiv. The paper proposes a new approach to mitigating model collapse, a phenomenon where output diversity narrows during iterative fine-tuning. The method uses a non-parametric entropy rate estimator computed entirely from raw text, without requiring access to model log-probabilities. This is significant because it provides a way to evaluate and improve open-weight models without needing access to the model's internal parameters. The paper's findings suggest that information-theoretic approaches to collapse mitigation are efficient and can help maintain diversity in multi-agent systems.

However, the shift to open weights is not without its challenges.

On 3 October, Google suspended product vulnerability submissions to its Open Source Software Vulnerability Reward Program (OSS VRP) due to an influx of invalid AI-driven reports. According to Tom's Hardware, the suspension went into effect on 1 October and is expected to last until the first quarter of 2027. The company cited a flood of 'AI slop' submissions, which are reports generated by AI models that are often inaccurate or nonsensical. This incident highlights a potential downside of the widespread use of AI tools: the increased volume of low-quality data and reports. It also highlights the need for strong filtering and verification mechanisms in open-source projects that rely on community contributions.

Despite these challenges, the momentum behind open-weight models is strong.

On 28 September, the fact that open-weight models processed the majority of AI tokens for the first time indicates a significant shift in market share. This trend is likely to continue as more companies release their models under permissive licenses and as the infrastructure to support local deployment improves. The ability to run powerful models locally, with privacy and control, is a compelling proposition for many users and organizations. The coming months will be vital in determining whether open weights can maintain their momentum or if proprietary models will reassert their dominance.

The environment for large language models is changing rapidly. The restriction of free access to proprietary models, combined with the release of powerful open-weight models, is creating a new dynamic in the AI industry. Organizations and individuals are increasingly turning to open weights for their privacy, control, and cost-effectiveness. The tools and infrastructure to support this transition are maturing, making it easier than ever to deploy and use open-weight models. The direction of AI may well be open, and the developments of the last week suggest that this direction is closer than we thought.

Comments 0

Sources

7
  1. 01Gemini app limiting what models free and AI Plus users can accessEN
  2. 02Google's new Gemini tiers cut free users to its weakest modelEN
  3. 03Local models on 27,489 CVEs: 4.2% omit security impactEN
  4. 04Spinifex: The open, AWS-compatible cloud you run yourselfEN
  5. 05Docent: An Open Source Private AI assistant in your terminalEN
  6. 06No Model Required: Text Entropy Rate Filtering Mitigates Iterative Fine-TuningEN
  7. 07Google freezes open-source bug bounty program amid flood of invalid AI slopEN

All figures and quotations in this text come from the sources listed below.

Content prepared by the editorial team with AI assistance.

Grace Okonkwo

Grace Okonkwo

AI, models and technology

Grace Okonkwo covers AI, models and technology for FLASH24, working from primary sources such as model cards, API documentation and benchmark papers rather than vendor summaries. She checks training data provenance, evaluation conditions and reported scores against the underlying datasets before any figure reaches print. She interviews researchers and engineers directly, tracks release calendars from major labs, and compares successive model versions on the same tests. Her own self-hosting, home-network and documentation-reading habits feed straight into that desk, since she tests tools on her own hardware first. She does not publish benchmark claims without a reproducible method.

Newsroom →

Comments

0
  1. No comments yet — be the first.

Write a comment

Comments are public. We do not publish abuse, spam or advertising.