Skip to content
World clockEU--:--UK--:--USA--:--CN--:--PLDEFRIT中文EN

portal about AI and technologyevents · analysis · interviews · technical background

Search
LIVE
›

Hallucination benchmarks sharpen as Gemini 4 Argon ships to a vetted few

Google began rolling out Gemini 4 Argon, its most powerful model to date, to a vetted group of cybersecurity experts on Wednesday, 30 September, holding the model back from the public over misuse concerns.

AI & modelsAnalysisGrace OkonkwoPublished: 2 October 20268 min readSources 13
Hallucination benchmarks sharpen as Gemini 4 Argon ships to a vetted few

Google began rolling out Gemini 4 Argon, its most powerful model to date, to a vetted group of cybersecurity experts on Wednesday, 30 September, holding the model back from the public over misuse concerns. The company said in a blog post that it was voluntarily giving the US government early access and would gather tester feedback before wider availability, according to The Guardian.

That cautious posture sits awkwardly next to the claims Google is making for the model. In its launch blog, cited by TechCrunch, the company said Argon scored significantly higher than OpenAI's GPT-6 Astra and Anthropic's Fable and Opus across a variety of AI benchmarks, citing the benchmarking startup Vals. It also said Argon can autonomously find, validate and patch critical software vulnerabilities. Those are capability claims, not reliability claims. The two are not the same thing.

This is where the hallucination question bites. A model that scores well on reasoning or coding benchmarks can still produce confident wrong answers, and the dossier's newest hallucination-adjacent material points in both directions. A controlled audit of gradient-conflict metrics in unified multimodal models, submitted to arXiv on 29 September, is the more interesting of the recent papers. It tests whether a widely used proxy metric actually predicts the thing researchers assume it predicts. It finds it does not.

The authors built a testbed called GRIDUMM that mirrors key structural ingredients of unified multimodal model training while making the ground-truth trade-off between understanding and generation exactly computable. Across 63 configurations and 372 measured checkpoints, no directional conflict metric reached an absolute Spearman correlation of 0.3 with a confidence interval excluding zero. A dose-response intervention that monotonically suppresses conflict left the trade-off flat, separating correlation from causation. Functional interference measures outperformed directional conflict metrics, and training loss tracked the trade-off strongly. The authors do not claim conflict is useless. They claim its validity as a diagnostic target has to be established, not assumed.

That is a narrow result about a narrow class of models. It is also the kind of finding that ought to make anyone reading a leaderboard pause.

What the newer benchmarks actually measure

Two other arXiv papers posted in the same window take a similarly skeptical line. ChartDensity-Bench, submitted on 30 September, evaluates multimodal large language models on reconstructing numerical data from scientific charts under controlled visual density, varying the number of charts shown at once across k equal to 1, 3, 6 and 9. Experiments on five recent MLLMs show numerical reconstruction degrades as visual density increases, and the magnitude of that degradation varies substantially across models. The paper's chart-level paired comparisons found the same source chart can incur higher reconstruction error when embedded in a denser visual context. That is a hallucination-adjacent failure mode with a clean measurement. The model is not refusing. It is producing numbers.

A third paper, Reverse Scoring for Diffusion Language Model Agents, submitted on 29 September, traces a specific failure in diffusion-based language model agents. They fall into retry loops, re-issuing an action long after it has failed. The authors attribute this to masked decoding committing the positions it is most confident about while deferring the uncertain ones, so at a failure state the context already offers a confident fill for the deferred decision, namely the failed action itself. They model the distortion as a task-blind corruption and propose a training-free remedy called Reflect Reverse, which they evaluate on four multi-turn embodied benchmarks.

There is a pattern across these three papers. Each one isolates a specific way a model can be confidently wrong, measures it under controlled conditions, and finds that the popular proxy metric does not track the outcome. None of them offers a general hallucination score. All of them suggest that general hallucination scores are doing less work than their publishers imply.

For a sense of how much a benchmark can move the conversation, look at the decision-model space. InfoQ reported that TypeSafe AI released Jev, which does not generate text at all but returns typed, probabilistic decisions with a confidence value, so calling code can act above a threshold and escalate below it. Input costs $0.042 per million tokens, output is free, the context window is 32,000 tokens, and TypeSafe quotes end-to-end latency of 70ms to 500ms. Vercel added Jev to AI Gateway on day two and said it reached nearly 13% of paid teams within 24 hours, twice the share of the GPT-5.6 family.

Amazon followed. TechCrunch reported on 1 October that AWS released Strands Decider 2B, an open source decision model inspired by Jev, built on the torso of Qwen3.5-2B. It sorts between pre-decided options and returns a confidence measure. Amazon distinguished engineer Marc Brooker told TechCrunch the model briefly reached the top spot on the Jevbench ranking for models of its size, and that the appeal is a workflow step that is more reliable thanks to confidence scores and a closed domain of answers, with lower latency and potentially lower cost.

TypeSafe founder Diogo Almeida was blunt about the crowding. "I get that people think it's a gold rush, but they might be underestimating the difficulty of making the models actually smart," he told TechCrunch. The current batch, he said, "seems more like" a wave of imitators. Whether or not that holds, the direction of travel is clear: some builders are responding to the reliability problem by narrowing the output space rather than by promising a better general model.

The safety argument arrives before the evidence

Google's own framing for restricting Argon is that frontier capability at this level needs a phased approach. Koray Kavukcuoglu, Google's chief AI architect, wrote that in the blog post announcing the model. Google said Argon was designed to refuse requests that could help carry out cyber-attacks or develop chemical, biological or nuclear weapons, and that it monitors the model's reasoning to stop it straying beyond what users intended, a risk researchers call misalignment.

The Guardian reported that the cautious rollout mirrors Anthropic, which has kept Claude Mythos Preview restricted to a small number of trusted organizations. Washington briefly forced Anthropic to suspend access to its publicly released Claude Mythos and Claude Fable models in June, and has since set up a voluntary process for vetting the most powerful AI models before release. The announcement came a day after Donald Trump hosted top tech executives, including Sundar Pichai and Anthropic's Dario Amodei, at the White House, where they signed a voluntary accord pledging to police the risks of their own AI systems.

Google said early testers used Argon to uncover a flaw in software used by hospitals around the world that exposed sensitive personal information, something other advanced models had missed. That is one anecdote, from the vendor, about a model not in general circulation. Treat it as a capability claim from an interested party.

The provenance fight is running in parallel. CNBC reported that OpenAI said it identified and disrupted a coordinated campaign to extract protected reasoning from its models, linking a core cluster of the activity to Chinese startup Moonshot AI. The activity began in early July and surged to 16,000 requests from more than 4,000 users over two days, with related activity across more than 15,000 users. OpenAI said operators did not breach its encryption, databases or stored user conversations, and that it had fully disrupted the campaign by July 28.

The Register noted the awkwardness of the complaint, given OpenAI's own history of training on web data, and reported that The Register asked OpenAI which of its models were targeted in July and did not hear back. It also reported that Michael Kratsios, Trump's Assistant for Science and Technology, accused Moonshot AI in late July of creating its Kimi K3 model by distilling Anthropic's Fable. Anthropic's Claude Opus 5.5, released a week before that report, ships with a defense against distillation called preserved thinking, introduced with Fable 5.1.

Moonshot did not immediately respond to CNBC's requests for comment, and The Register said it also received no immediate response. So the central accusation remains one-sided in the public record.

The other thread worth watching is where the training data comes from and on what terms. The Guardian reported on 1 October that Anthropic urged the Albanese government to consider conditional approval for training on Australian copyrighted works under an opt-out model, after conceding it would not secure a blanket copyright exemption. The ABC and SBS pushed back in submissions to a federal parliamentary inquiry, with the ABC warning of a cannibalisation of the Australian news industry and both broadcasters suggesting AI companies be included in the news bargaining incentive.

Anthropic's submission claimed the technology could transform the Australian economy while also warning of serious implications for national security, critical infrastructure resilience and public safety. The joint select committee on AI holds hearings in Sydney next week, with Anthropic and OpenAI executives due to attend.

None of this settles whether any given model is reliable. It does suggest the industry is now shipping reliability claims faster than it is shipping ways to check them, and that the checks which do exist tend to be narrow, controlled and unflattering to the proxy metrics everyone quotes.

Comments 0

Sources

13
  1. 01Google rolls out new Gemini AI model but restricts access over safety concernsEN
  2. 02Google releases Gemini 4 Argon, called its most powerful model yetEN
  3. 03Does Gradient Conflict Predict the Understanding-Generation Trade-off? A Controlled Audit of Conflict-Metric Validity in Unified Multimodal ModelsEN
  4. 04ChartDensity-Bench: Benchmarking MLLMs for Numerical Data Reconstruction under Visual DensityEN
  5. 05Does This Action Still Explain the Task? Reverse Scoring for Diffusion Language Model AgentsEN
  6. 06TypeSafe AI Releases Jev: A Decision-Only Model That Returns Typed Probabilities Instead of TextEN
  7. 07Amazon releases its own Jev clone as decision models flood the webEN
  8. 08AI race heats up as OpenAI flags alleged model-copying campaignEN
  9. 09Irony alert: OpenAI whines that Chinese model stole its special IP that it stole from everybody elseEN
  10. 10Anthropic pushes for opt-out model for Australian content as ABC warns of 'cannibalisation' of newsEN
  11. 11Microsoft AI releases new transcription and text-to-speech models for voice agentsEN
  12. 12Ideogram says its new model can edit part of an image without messing up the restEN
  13. 13ThreatsDay: AI-Powered Zero-Day Chain, 543K Live Secrets, Model Inspection RCE and 13 More StoriesEN

All figures and quotations in this text come from the sources listed below.

Content prepared by the editorial team with AI assistance.

Grace Okonkwo

Grace Okonkwo

AI, models and technology

Grace Okonkwo covers AI, models and technology for FLASH24, working from primary sources such as model cards, API documentation and benchmark papers rather than vendor summaries. She checks training data provenance, evaluation conditions and reported scores against the underlying datasets before any figure reaches print. She interviews researchers and engineers directly, tracks release calendars from major labs, and compares successive model versions on the same tests. Her own self-hosting, home-network and documentation-reading habits feed straight into that desk, since she tests tools on her own hardware first. She does not publish benchmark claims without a reproducible method.

Newsroom →

Comments

0
  1. No comments yet — be the first.

Write a comment

Comments are public. We do not publish abuse, spam or advertising.