Skip to content
World clockEU--:--UK--:--USA--:--CN--:--PLDEFRIT中文EN

portal about AI and technologyevents · analysis · interviews · technical background

Search
LIVE
›

Brittle decisions: how easily a correct model choice can be overturned

Four papers from September expose weak spots in language models. A short, natural-looking piece of context can flip a correct decision, and the gap between what a model encodes and what it actually uses can be very wide.

ScienceExplainerSofia MarchettiPublished: 25 September 20267 min readSources 4
Brittle decisions: how easily a correct model choice can be overturned

Decision models map text onto a probability distribution over a finite set of choices, and they increasingly drive real actions: routing tickets, picking tools, triggering operations. The paper "JevOut" asks what happens to such a decision when a short piece of text is added to the input, one that fits the context naturally and leaves the correct answer untouched.

Context versus answer

The authors kept the experiment simple. For every initially correct task they fixed one wrong target option. Then they optimised additions to the context so that the additions stayed fluent while pushing the choice toward that option. The source, the question, the answer set and the correct answer all stayed the same. Of 508 initially correct decisions by the Jev model, they redirected 312, or 61.4 percent. In 229 cases the model gave the fixed wrong option a probability of at least 0.7. Tests on three further decision systems produced redirection rates from 64.9 to 73.2 percent. The authors draw a blunt conclusion: if these models turn language directly into decisions, their probabilities should not be treated as a reliable decision interface.

Does a justification do anything at all

A second project looks at something even more basic. A model rejects one candidate and cites a fact that was not in the profile, for instance a missing director or date of death. That sentence can be checked without any judge: insert into the profile a sentence from the corpus that states the fact, then ask again. In the largest of three runs, with six open models on the 2WikiMultihopQA set, supplying the cited fact shifted the choice more strongly than a length-matched irrelevant addition. The odds ratio was 3.57, in a range from 1.54 to 8.26, with a Holm-corrected level of 0.0210. The result held after removing any single model.

The most interesting effect, though, is not about content. The same irrelevant addition shifted the choice more strongly when it sat next to the cited rival than when it sat next to the third option the model never mentioned, with a Holm-corrected level of 0.0008. Part of the work, in other words, is done by where the sentence sits, not by what it says. The author also reports that validation of the measurement rules turned up eight defects. The most serious was a choice-parsing rule that returned the option just rejected in 17.1 percent of resolvable answers; it would have inflated the number of surviving contrasts from four to six.

Linearity and repetition

Two further papers concern properties of the models themselves. Paweł Tichonow's team shows that transformers, though built from strongly nonlinear components, display a fundamental linearity. Sum the inputs from different text streams linearly, and the next-token distribution becomes a superposition of the distributions of the individual streams. The authors call this the linearity of superposition hypothesis and argue it is a property of the architecture rather than an effect of training. The phenomenon weakens as pretraining proceeds, and light fine-tuning can restore linearity. That leads to a practical use: controlled decoding can untangle the superposition and generate two coherent continuations from a single forward pass.

The last paper deals with text degeneration, the tendency of a model to repeat itself. The author does not settle the dispute between a data explanation and a network explanation. He runs a different kind of measurement instead: the structure of fixed points of the maximum-probability map in a short window, across 96 random two-token starts and seventeen ready-made models. The class of behaviour was stable in seventeen out of seventeen cases, but with a fixed corpus and scale it is not settled: some families form funnels, others do not. Eight of the seventeen models and seven families out of five corpora show a funnel. The conclusion is methodological: calling repetitiveness a trait of a particular model is premature.

Comments 0

Sources

4
  1. 01JevOut: Natural Context Can Flip Decision ModelsEN
  2. 02Does a model's stated reason for rejecting a candidate do any work?EN
  3. 03Superposition Linearity HypothesisEN
  4. 04What a Cross-Model Fixed-Point Census Can and Cannot Arbitrate About RepetitionEN

All figures and quotations in this text come from the sources listed below.

Content prepared by the editorial team with AI assistance.

Sofia Marchetti

Sofia Marchetti

Science and health

Sofia Marchetti covers science and health for FLASH24, working from primary literature, preprints, and agency data rather than press releases. She checks sample sizes, confidence intervals, and whether a study's numbers match its abstract before filing. She interviews researchers and clinicians directly, tracks conference calendars for embargoed results, and compares new findings with earlier trials on the same question. Outside the newsroom she works on materials physics and stargazes through a home telescope, which keeps her close to how measurement error actually behaves. She does not publish a health claim without a named source and the underlying data.

Newsroom →

Comments

0
  1. No comments yet — be the first.

Write a comment

Comments are public. We do not publish abuse, spam or advertising.