UC Riverside study maps internal features driving AI confidence and accuracy
Researchers at UC Riverside have identified distinct internal features within large language models that govern confidence and correctness separately, offering a path to reduce hallucinations without retraining entire systems.

On 23 September, a team led by UC Riverside computer scientists published a study revealing that a model's self-reported confidence does not reliably predict whether its answer is accurate. The finding challenges a foundational assumption in the field: that high confidence correlates with high correctness. Instead, the researchers found these traits can arise from entirely different internal mechanisms.
"The main assumption in the field is that when the model is confident, it is likely to be correct, and when the model is unsure, it is more likely to be incorrect," said Het Patel, a doctoral student and lead author, as reported by UC Riverside. "But we often see the counterexamples that are well documented."
The study examined two open-weight models: Meta's Llama-3.1-8B and Google's Gemma-2-9B. Using multiple-choice questions, the team categorized responses into four groups based on accuracy and confidence levels. They then employed sparse autoencoders to analyze the models' internal activity, identifying features that activated during confident, uncertain, correct, or incorrect states.
The mechanism of the mismatch
The analysis uncovered three types of features: those associated primarily with uncertainty, those tied to incorrect answers, and "confounded" features linked to both. When the researchers suppressed features associated purely with uncertainty, model accuracy dropped sharply. This suggests that uncertainty-related features actually play a role in generating good answers, contrary to the idea that they are merely noise or hedging.
Conversely, suppressing features associated solely with incorrect answers had a different effect. This separation allows for targeted interventions. Patel described the process as "turning knobs" inside the model to observe behavioral changes. The goal is to adjust these features so models become more confident when they are right and more cautious when they are likely to be wrong, all without the costly process of retraining the entire network.
This work fits into a broader trend of researchers trying to make AI systems more trustworthy. As large language models are increasingly used to inform decisions and complete tasks, developers need better ways to determine when an answer can be trusted. The UCR-led research points toward a method where AI systems can be adjusted to align their internal confidence signals with actual performance.
Context from recent benchmarks and models
The challenge of balancing speed and accuracy is also visible in recent model releases. Reflection AI debuted Beam, a 501-billion-parameter open-weight model, on 5 October. The company claims Beam matches the performance of leading Chinese open models on reasoning benchmarks while using 3 to 4 times less inference compute, according to TechCrunch. While Beam is not a "decision model" in the same sense as the UCR study, it highlights the industry's push for efficiency in reasoning tasks where reliability is paramount.
Similarly, GitHub released ReviewBench on 5 October, an open benchmark for AI code review. The benchmark was designed to measure the quality of agentic code review, a task where hallucinated or incorrect code suggestions can have real-world consequences. ReviewBench analyzes 219 public pull requests across 19 languages, aiming to provide a rigorous evaluation methodology that tracks whether changes in review systems lead to meaningful production gains.
These developments highlight the dual pressure on AI developers: to improve raw capability while simultaneously engineering out the uncertainty that leads to hallucinations. The UCR study provides a technical roadmap for the latter, suggesting that the internal features responsible for doubt and accuracy are separable and modifiable. This could allow future models to explicitly flag their own doubts, creating a layer of transparency that is currently missing from many commercial systems.
Sources
15- 01Study: Path discovered to make AI models red-flag their doubtful answersEN
- 02Reflection debuts Beam, an open-weight AI model to rival Chinese models at lower compute costEN
- 03Beam: Reflection's 501B open-weight modelEN
- 04ReviewBench: An open benchmark for AI code reviewEN
- 05Jeff-Code: a 0.8B model makes Qwen 3.8-27B coding 47% faster, same pass rateEN
- 06d1: The most capable decision model, now with visionEN
- 07Interfaze-1-lite: the first open-weight model for deterministic taskEN
- 08Grok Voice tops new benchmarkEN
- 09NASA-IBM Lunar Foundation Model and fine-tuning codeEN
- 10Precogly (The open-source alternative to commercial threat modeling platforms)EN
- 11Sales of sub-€25,000 electric car models set to rise sevenfoldEN
- 12Could a Large Language Model Be Conscious? – David J. Chalmers (2024)EN
- 13System One models like Jev can train their own replacementsEN
- 14Study of 2.5M children finds no link between MMR vaccine and autismEN
- 15Benchmark in MillisecondsEN
All figures and quotations in this text come from the sources listed below.
Content prepared by the editorial team with AI assistance.
Comments
0- No comments yet — be the first.