Skip to content
World clockEU--:--UK--:--USA--:--CN--:--PLDEFRIT中文EN

portal about AI and technologyevents · analysis · interviews · technical background

Search
LIVE
›

Reflection AI launches Beam, a 501B open-weight model targeting Chinese rivals

Reflection AI officially unveiled Beam, its first frontier open-weight model, on Monday, October 5. The startup claims the 501-billion-parameter system matches leading Chinese open models on advanced reasoning benchmarks while using 3-4x less inference compute.

AI & modelsExplainerGrace OkonkwoPublished: 5 October 20265 min readSources 5
Reflection AI launches Beam, a 501B open-weight model targeting Chinese rivals

Brooklyn-based startup Reflection AI released Beam on Monday, October 5.

According to TechCrunch, Reflection confirmed earlier reporting from Axios that a launch was imminent. The company detailed the model's architecture in a blog post, describing Beam as a text-only mixture-of-experts (MoE) model. It features 501 billion total parameters, with only 23 billion active during inference. The model was pretrained on 23.8 trillion tokens and supports a context window of one million tokens. For context, Z.ai's GLM-5.2 has roughly 744 billion total parameters with 40 billion active, making Beam significantly smaller in its active footprint.

Reflection asserts that Beam scores on par with Z.ai's GLM-5.2 on advanced reasoning benchmarks.

The startup labels Beam a "workhorse model" designed for enterprises, the public sector, and developers who require high-volume processing without the premium costs associated with closed-source APIs.

The technical foundation of Beam relies heavily on high-compute reinforcement learning (RL). Reflection's blog post, published on its own website, notes that the RL run generated over 100 million rollouts. This training phase took place on 10,500 NVIDIA GB300 GPUs over four weeks. The company emphasizes that this investment in RL infrastructure is what allows Beam to achieve competitive performance with larger models like Qwen 3.8-Max on coding and agentic tasks, despite its lower parameter count.

Beam is currently undergoing final red-teaming and evaluations.

The release occurs against a backdrop of narrowing gaps between US and Chinese AI capabilities. Bloomberg Intelligence reported on Sunday, October 4, that the US lead in AI over China has narrowed, with DeepSeek gains closing the performance gap to just 3 percent. This context makes Reflection's claim of parity with Z.ai's GLM-5.2 particularly significant. If accurate, Beam would challenge the narrative that only the largest Chinese models can deliver frontier-level reasoning at competitive prices.

Reflection positions Beam against both closed labs like Anthropic and OpenAI, and open models from Western players like Mistral and Meta. Its most direct US rival appears to be Inkling, the open model from Mira Murati's Thinking Machines Lab, which was released in July. According to Reflection's internal benchmarks, Beam outscores Inkling on four coding tests where both report results. However, a key distinction remains: Inkling is a multimodal model, whereas Beam is text-only. This limitation may restrict Beam's utility in applications requiring image or video processing, a feature that is becoming standard in frontier deployments.

Benchmark comparisons in the AI sector are often contested.

The Veris AIsolutions leaderboard, released on October 5, provides a contrasting view of performance in specific verticals. In their VAmoS Pro Bench for voice agents, Grok Voice topped the leaderboard with a 44.7% task completion rate, while other models trailed behind. While this benchmark focuses on voice interactions rather than general reasoning or coding, it illustrates the fragmentation of performance metrics. A model that excels in text-based reasoning may not necessarily lead in agentic or voice-driven workflows, complicating direct comparisons between releases like Beam and existing market leaders.

GitHub's new ReviewBench, also published on October 5, offers another perspective on code review capabilities. The benchmark analyzed over 100 million real pull requests to model the distribution of code review workloads. It includes 219 public pull requests across 19 languages. GitHub stated that their offline evaluation of Copilot code review became more effective at anticipating production experiment directions with this new tool. While Reflection claims Beam is strong in coding, the lack of independent verification on standardized, publicly available benchmarks like ReviewBench leaves the full picture of Beam's coding prowess unclear.

Reflection was founded in 2024 by two former Google DeepMind researchers.

Per PitchBook data cited by TechCrunch, the startup has raised roughly $4.7 billion from backers including Nvidia, Sequoia Capital, and Lightspeed Venture Partners. Its last round valued the company at a $25 billion pre-money valuation. This substantial capital injection is critical for securing the compute resources necessary to train frontier models. The industry consensus is that compute access is a key ingredient for luring customers away from closed models and cheaper open-weight alternatives from Chinese labs.

This summer, Reflection signed deals worth more than $7 billion with SpaceX and Nebius. These agreements secure access to Nvidia's GB300 chips through 2029. The reliance on Nvidia hardware is a double-edged sword; it ensures top-tier performance but ties the company's roadmap closely to its largest investor. Nvidia CEO Jensen Huang has long championed the "AI factory" idea, a vision that would benefit Nvidia as its GPUs power these systems. Reflection aims to sell "AI factories" to enterprises and sovereign nations, allowing institutions to build customized, local AI systems by training Reflection's models on proprietary data.

Axios reported that hedge funds and trading firms are among those eager to build such systems.

The "AI factory" model offers a distinct value proposition compared to standard API services. By enabling local training on private data, Reflection addresses data sovereignty concerns that are paramount for financial institutions and government agencies. This strategic pivot suggests that Reflection's long-term goal is not just to release competitive models, but to establish an infrastructure layer for private AI deployment.

The timing of the Beam release also coincides with other significant developments in the AI ecosystem. On October 5, Liquid AI introduced d1, a decision model that supports both text and images. Liquid AI claims d1 matches or beats GPT-6.1 Sol on four of six real applications while costing 19x to 200x less. This aggressive pricing strategy by smaller players highlights the pressure on frontier models to justify their compute costs. Similarly, GitHub released ReviewBench on the same day, signaling a push for more rigorous and standardized evaluation methods in the industry.

Reflection's entry into the market adds to the complexity of the current AI landscape. The company must prove that its efficiency claims translate into real-world savings for enterprises. If Beam delivers on its promise of frontier performance at a fraction of the cost, it could disrupt the pricing models of both closed and open competitors. However, without independent verification, the industry remains skeptical. The coming weeks will be critical as the weights are released and third-party benchmarks are conducted. For now, Reflection has set a high bar for the Western open-weight sector, challenging Chinese leaders and setting the stage for a new phase in the global AI race.

Comments 0

Sources

5
  1. 01Reflection debuts Beam, an open-weight AI model to rival Chinese models at lower compute costEN
  2. 02Beam: Reflection's 501B open-weight modelEN
  3. 03Grok Voice tops new benchmarkEN
  4. 04ReviewBench: An open benchmark for AI code reviewEN
  5. 05d1: The most capable decision model, now with visionEN

All figures and quotations in this text come from the sources listed below.

Content prepared by the editorial team with AI assistance.

Grace Okonkwo

Grace Okonkwo

AI, models and technology

Grace Okonkwo covers AI, models and technology for FLASH24, working from primary sources such as model cards, API documentation and benchmark papers rather than vendor summaries. She checks training data provenance, evaluation conditions and reported scores against the underlying datasets before any figure reaches print. She interviews researchers and engineers directly, tracks release calendars from major labs, and compares successive model versions on the same tests. Her own self-hosting, home-network and documentation-reading habits feed straight into that desk, since she tests tools on her own hardware first. She does not publish benchmark claims without a reproducible method.

Newsroom →

Comments

0
  1. No comments yet — be the first.

Write a comment

Comments are public. We do not publish abuse, spam or advertising.