Skip to content
World clockEU--:--UK--:--USA--:--CN--:--PLDEFRIT中文EN

portal about AI and technologyevents · analysis · interviews · technical background

Search
LIVE
›

Amazon Ships Strands Decider 2B as Decision Models Crowd the Benchmark Field

Amazon Web Services released Strands Decider 2B on 1 October, an open-source decision model that sorts between pre-decided options and returns a confidence score, the same week OpenAI announced a similar offering and a month after TypeSafe's Jev started the category.

AI & modelsAnalysisGrace OkonkwoPublished: 1 October 20267 min readSources 11
Amazon Ships Strands Decider 2B as Decision Models Crowd the Benchmark Field

Amazon Web Services released Strands Decider 2B on 1 October, a small open-source model that does not write prose. It picks between options a workflow has already defined and reports how confident it is in the pick. TechCrunch reported the release the same day, noting the model is small enough to run locally and is free to download. It is the latest entrant in a category that barely existed a month ago. TypeSafe AI, a San Francisco lab founded by former OpenAI researcher Diogo Almeida, shipped Jev and called it the first of its System One Models. The pitch is narrow: no chat, no text generation, just typed probabilistic answers software can act on.

What a decision model actually returns

According to InfoQ, a caller sends Jev a state, either a string or structured data, plus a set of typed questions. The model evaluates them in a single parallel pass and returns Choice, Score and Noul answers with a probability distribution and a confidence value. Calling code acts above a threshold and escalates below it.

The commercial details are public. Input costs $0.042 per million tokens, output is free, the context window is 32,000 tokens, and TypeSafe quotes end to end latency of 70ms to 500ms. Training uses a method it calls Reinforcement Learning for Calibrated Decisions. Those numbers come from InfoQ's 1 October write-up.

Amazon's version is built on the torso of Qen3.5-2B. Marc Brooker, an Amazon distinguished engineer, built it after seeing Jev and trying his own take. The homebrew project briefly reached the top spot on the Jevbench ranking for models of its size, TechCrunch reported, and Amazon engineers cleaned it up for release under Strands Labs.

"What originally piqued my interest in this class of models was that they make a perfect decider for a workflow step, 'what is the next thing for me to do here, based on where I am?'" Brooker told TechCrunch.

He described the appeal as reliability from confidence scores and a closed domain of answers, plus lower latency and potentially lower cost. That is the whole sales argument for the category, and it is the same one TypeSafe makes.

Adoption numbers, and their limits

Jev's early traction is the most concrete evidence that the category has buyers. Vercel added Jev to its AI Gateway on day two and said it reached nearly 13% of paid teams within 24 hours, twice the share of the GPT-5.6 family, per InfoQ. Netlify followed, LangChain shipped a TypeSafeClassifier integration with model routing and an AutoMode middleware that screens tool calls before they run, and five independent Elixir clients appeared within days.

The performance claims deserve scepticism. Vercel engineer Pranit Sharma found a safety classifier ran five to 18 times faster than the LLM it replaced. Bryo AI CTO Nikhil Mudholkar rated Gemini slightly more accurate on email classification but 10 to 20 times more expensive, and valued Jev as the only one handing back a real probability.

An analysis of 12,759 launch tweets by OpenChamber put user reported speedups at a median of 7x against the 193.6x headline. Cost savings came in at a median of 30x, latency at a median of 76ms with an upper quartile of 270ms. Those are self-reported numbers from social posts, aggregated, and they are far below the marketing figure.

Armin Ronacher, CTO of Earendil, told TechCrunch the design "delegates the hallucination problem a little bit to the user", who has to decide whether a 50% probability is worth acting on. He pointed to model routing as another good fit. On Hacker News, one developer noted that Jev cannot emit an invalid type but can still emit a completely wrong valid value.

A crowded field with unclear value

Dozens of similar models have appeared since TypeSafe debuted the idea. That shows interest. It also raises the question of how much any single one is worth, especially when the underlying method is public and the build cost is low.

Brooker told TechCrunch he does not necessarily expect the frontier labs to dominate the space, since with smaller markets the cost to build something interesting runs into the hundreds or thousands of dollars. TypeSafe executives say they are keeping their heads down and improving future models.

"I get that people think it's a gold rush, but they might be underestimating the difficulty of making the models actually smart," CEO and founder Diogo Almeida told TechCrunch, saying that for now, he did not see real competition for his company emerging yet.

Brooker framed the engineering trade-off as a balance between accuracy and calibration on decision tasks, without degrading the model's language understanding and general knowledge. Push one, and the other moves.

The benchmark problem underneath

Decision models land in a measurement environment that is already contested. A paper posted to arXiv on 29 September, More Choices, Fewer Decisions, examines ordinal-scale bias in JEV-like direct-decision models, which is a direct challenge to how these systems are scored. Another arXiv paper submitted the same day, ArgGYM, argues that success under fixed problem specifications and stable evaluation criteria may not transfer to reasoning under incomplete and revisable information, and introduces a procedural benchmark of 1,440 verified instances across fifteen curriculum configurations.

Clinical work points the same way. A paper accepted to Findings EMNLP 2026, Sense and Sensitivity, compared LLMs to practicing physicians across more than 6,000 clinical scenarios and 7,000 physician annotations. It found LLMs more likely than physicians to recommend unnecessary care at baseline, with the tendency increasing under perturbed inputs, and greater sensitivity to gender and tone changes. Benchmark scores did not predict that.

Hardware is another axis. AgBench, also posted on 29 September, evaluated local, hybrid and cloud agent execution across 162.07 million data points and found that no single architecture performs best across task success, goodput, cloud cost and data exposure. Local-only execution eliminates cloud API costs and sensitive-information exposure but generally has lower task success and longer completion times.

That finding matters for Amazon's local-first pitch. Strands Decider 2B can run on your own machine. Whether it should, for a given workflow, depends on the numbers AgBench says no one is publishing in a single comparable table.

Meanwhile the same week produced a very different kind of release. Google launched Gemini 4 Argon on 30 September, calling it its most powerful model yet, but restricted it to a select group of cyber partners through its Fairwind Program. The Guardian reported on 1 October that Koray Kavukcuoglu, Google's chief AI architect, wrote that safely releasing frontier capabilities at this level requires a phased approach, and that Google is voluntarily giving the US government early access.

The contrast is the story of the month. At one end, a 2B model anyone can download and run on a laptop. At the other, a frontier system held back from the public over misuse risk, following OpenAI's decision to scrap GPT-6.1 Astra after it failed alignment tests, as CNBC and WIRED reported on 28 and 29 September.

Decision models are not competing on that axis. They are competing on whether a workflow step can be made cheaper, faster and more predictable than calling a general model. Jevstiller, a GitHub project published on 30 September, takes that logic to its conclusion: it sits in front of repeated Jev classification calls, learns from Jev's own probability distributions, and answers most requests locally once the local model agrees with Jev within a set budget. On Banking77, replayed against Jev's recorded answers, the local model answered 71.9% of held-out messages at 99.50% agreement with Jev, taking over most traffic from about 5,000 messages on.

The category's economics are still being tested. Vercel's 13% of paid teams in 24 hours is real adoption, but it is adoption of a free-tier gateway integration, not proof of durable spend. OpenChamber's median 7x speedup is well short of the headline. And the benchmark papers arriving this week suggest the field has not settled on how to tell a good decider from a fast one.

For now, the models are cheap, open and multiplying. The measurement is not keeping up.

Comments 0

Sources

11
  1. 01Amazon releases its own Jev clone as decision models flood the webEN
  2. 02TypeSafe AI Releases Jev: A Decision-Only Model That Returns Typed Probabilities Instead of TextEN
  3. 03More Choices, Fewer Decisions: Ordinal-Scale Bias in JEV-like Direct-Decision ModelsEN
  4. 04ArgGYM: A Procedural, Engine-Verified Benchmark for Structured Defeasible ReasoningEN
  5. 05Sense and Sensitivity: Benchmarking LLM Clinical Triage Recommendations with Physician ExpertsEN
  6. 06AgBench: Agentic AI Benchmarks for Personal AI DevicesEN
  7. 07Jevstiller: Distill a repeated Jev classification task into a local modelEN
  8. 08Google rolls out new Gemini AI model but restricts access over safety concernsEN
  9. 09Google releases Gemini 4 Argon, called its most powerful model yetEN
  10. 10OpenAI abandons plan to release upcoming model as safety concerns escalateEN
  11. 11OpenAI Delays Release of Latest Model Over Safety ConcernsEN

All figures and quotations in this text come from the sources listed below.

Content prepared by the editorial team with AI assistance.

Grace Okonkwo

Grace Okonkwo

AI, models and technology

Grace Okonkwo covers AI, models and technology for FLASH24, working from primary sources such as model cards, API documentation and benchmark papers rather than vendor summaries. She checks training data provenance, evaluation conditions and reported scores against the underlying datasets before any figure reaches print. She interviews researchers and engineers directly, tracks release calendars from major labs, and compares successive model versions on the same tests. Her own self-hosting, home-network and documentation-reading habits feed straight into that desk, since she tests tools on her own hardware first. She does not publish benchmark claims without a reproducible method.

Newsroom →

Comments

0
  1. No comments yet — be the first.

Write a comment

Comments are public. We do not publish abuse, spam or advertising.