New benchmarks show AI models still lag human experts in research reasoning
A new benchmark released on 3 October reveals that agentic search tools perform no better than standard embedding retrieval at finding scientific papers that inspire new research, with top models reaching only 0.51 Recall@20.

ScholarCatalyst was published on 3 October. The benchmark tests if AI can find the specific prior work that catalyzes new scientific projects. Current large language models struggle with the intuitive leap required to connect a vague research question to a useful paper buried in a massive archive.
The study involved 184 lead authors of 207 recent computer science papers. These researchers labeled which candidate papers had actually advanced their completed projects, providing detailed rationales for each choice. The goal was to create a dataset that reflects the "research taste" of expert scientists, a skill that remains elusive for current AI systems.
The gap between retrieval and intuition
The core task given to the models was to retrieve these catalyst papers from a corpus of 191,000 papers, using only literature that existed before the project began. This setup forces the AI to operate under the same constraints as a human researcher starting a new project: without knowing the answer, they must find the key insight among thousands of irrelevant documents.
The findings were sobering. Agentic search, a method where an AI agent uses a retrieval tool to iteratively search for information, did not outperform simple embedding retrieval. In fact, agentic search achieved a Recall@20 score of 0.42, while standard embedding retrieval scored 0.48. This indicates that the added complexity of agentic behavior did not help the model identify the correct papers.
Even a more advanced attempt using an agent built on Claude Fable 5.1 yielded only a slight improvement. This model, which may have been exposed to the completed papers during its training phase, reached a Recall@20 of 0.51. While higher than the baseline, this score still highlights a significant gap between AI performance and the ability of human experts to sense which prior idea a new problem needs.
"These results highlight the need for new training recipes that equip models with expert intuition for searching broad corpora," the authors wrote in their report. The benchmark is envisioned as a step toward scientific agents that can take a half-formed idea and point to the prior research it needs, but current models are not yet ready for that role.
Other reasoning challenges persist
ScholarCatalyst is not the only recent evaluation showing that AI reasoning capabilities are uneven. On 3 October, Anthropic released a report claiming that Zhipu AI's GLM-5.3 model possesses "Mythos-class" hacking abilities. According to the report, GLM-5.3 developed end-to-end exploits for Google Chrome 50 times in 410 runs in a sandboxed environment known as Exploitbench. In the same test, Anthropic's unreleased Claude Mythos model achieved 56 successful exploits in 410 attempts.
However, the same report noted that GLM-5.3's safeguards were weak. Anthropic stated that the model's guardrails could be bypassed using deceptive prompts, which resulted in a 64% success rate, or by prefilling the model's thinking tokens, which resulted in a 92% success rate. This suggests that while some models are becoming more capable at complex, step-by-step tasks like exploit development, their safety mechanisms may not keep pace.
In a related development, a paper titled "PTXBench: Benchmarking and Adapting LLMs for GPU Kernel Optimization with Architecture-specific PTX" was revised on 28 September. The study evaluated large language models on their ability to write optimized GPU kernels using architecture-specific PTX instructions. The authors found that "success rates fall substantially on complex attention backward workloads, and executing the target instructions does not necessarily translate into competitive performance." No evaluated model consistently matched frontier libraries across the test suite.
Human insight remains the gold standard
Despite these advances in specific technical domains, the broader question of reasoning remains unresolved. A guest post on the blog "What's new" by mathematician Terence Tao, published on 3 October, reflects on the role of AI in mathematics. Tao noted that while AI models have helped fix incorrect lemmas and suggested new directions, he has yet to succeed in prompting a solution to one of his long-term problems.
"Mathematics is a human endeavor, and machine-generated output requires substantial human intervention in order to contribute to our knowledge base," the post states. This sentiment echoes the findings of the ScholarCatalyst benchmark: AI can assist with execution and retrieval, but the strategic decision of which path to take often requires human insight.
As AI models continue to evolve, benchmarks like ScholarCatalyst and PTXBench will be essential for measuring progress. They provide concrete metrics for what models can and cannot do, moving the conversation beyond marketing claims to empirical evidence. For now, the data suggests that while AI is becoming a powerful tool for specific tasks, it has not yet replaced the human capacity for scientific intuition.
Sources
15- 01ScholarCatalyst: A Benchmark for Retrieving Papers That Inspire New ResearchEN
- 02Anthropic claims popular Chinese AI model has Mythos-class hacking abilitiesEN
- 03PTXBench: Benchmarking and Adapting LLMs for GPU Kernel OptimizationEN
- 04What does the advent of powerful AI models mean for mathematicians like me?EN
- 05Anthropic's super bug-hunting model Mythos is hardcore good at mathEN
- 06The Download: a biological de-aging contest and why LLMs don't reasonEN
- 07OpenAI alerts 100+ orgs that its 'misaligned models' attempted to break inEN
- 08New in Llama.cpp: Decision ModelsEN
- 09Shield: A 118M model for detecting prompt injections and jailbreaksEN
- 10Kolibri: A Sovereign Open-Weight ModelEN
- 11Training an open decision model to replace a closed one, in shadow on prodEN
- 12Open 10-player CS:GO dataset for multiplayer world modelsEN
- 13Changes to Gemini model access and limitsEN
- 14A Beginner's Guide to Running AI Models LocallyEN
- 15Twelve AI clay films for $184: The agents cost more than the video modelEN
All figures and quotations in this text come from the sources listed below.
Content prepared by the editorial team with AI assistance.
Comments
0- No comments yet — be the first.