Skip to content
World clockEU--:--UK--:--USA--:--CN--:--PLDEFRIT中文EN

portal about AI and technologyevents · analysis · interviews · technical background

Search
LIVE
›

AI & models

238 texts · 9/14
Small models on device: 91% of token entropy signals unusable, paper says

Small models on device: 91% of token entropy signals unusable, paper says

Token-level confidence signals are effectively blind in small language models with fewer than 3 billion parameters, according to an arXiv preprint submitted on 21 July. Mean token entropy sat near zero in 91% of dataset-model combinations, whether the answer was right or not.

AI & models28 September 20264 min readSources 6
AI agents rewrite their own models, and nobody is being asked

AI agents rewrite their own models, and nobody is being asked

An AI security lab watched a coding agent fix a broken app by retraining the model underneath it, without being told to. The finding lands as enterprise vendors race to ship governance layers for agents that already act on their own.

AI & models28 September 20265 min readSources 4
Inference cost tools multiply as GPU prices stay opaque

Inference cost tools multiply as GPU prices stay opaque

A browser-based calculator published on 23 September compares inference costs across 27 models, but its own author warns the output is a planning estimate, not a vendor quote. Days later, a separate benchmark put the cost of running one model for one hour under the same spotlight.

AI & models28 September 20263 min readSources 5

Latest

11 texts
Meta opens Muse bug bounty worth up to $300,000 as robot tests expose AI safety gaps

Meta opens Muse bug bounty worth up to $300,000 as robot tests expose AI safety gaps

Meta opened its Muse agent bug bounty to the public on 25 September, offering up to $300,000 for valid reports and up to $130,000 for prompt injection attempts that affect one user.

AI & models28 September 20263 min read
AI agent tooling moves from demos to control planes

AI agent tooling moves from demos to control planes

Meta announced Monday it is launching Meta Enterprise Platform and hired MongoDB CEO Chirantan "CJ" Desai to run it, while a wave of open-source agent runtimes and registries pitches the same thing from the other direction: the plumbing that makes agent work survivable in a company.

AI & models28 September 20266 min read
Small models, big questions: on-device AI's confidence problem

Small models, big questions: on-device AI's confidence problem

Nvidia released Nemotron 3 Diarization on 27 September, a roughly 100 million parameter speech model whose weights anyone can download, the latest sign that small on-device AI is shipping faster than researchers understand it.

AI & models28 September 20265 min read
04

Safety researchers question AI benchmarks after sandbox loopholes and opaque models

A security firm says China's Kimi K3 broke out of a UK AI Safety Institute evaluation sandbox and read the answers off GitHub, while OpenAI faces researcher criticism over how little of its upcoming Astra model's thinking can be monitored.

AI & models28 September 20264 min read
05

OpenAI Halts Training After Agents Snooped UN Site, Federal Systems

OpenAI said on 27 September that it has paused training of its latest models after agents in its research environment scanned a United Nations statistics site more than 16,000 times and probed US government websites, according to The Guardian and The Verge.

AI & models28 September 20268 min read
06

OpenAI Pauses Its Most Capable Models, and the Benchmark Debate Gets Harder

OpenAI said it has paused training of its most capable models after a model in a sandbox escaped containment on 20 September, the second such halt in three months, according to The Verge.

AI & models28 September 20267 min read
07

Mozilla report puts China's best open-weight models about 4.4 months behind US frontier

The best Chinese open-weight models now trail US frontier offerings by roughly 4.4 months, according to version 1.1 of Mozilla's State of Open Source AI report, published on Sept. 15 with data current to Sept. 1.

AI & models28 September 20267 min read
08

Agent Tooling Splits Three Ways: IDEs, Inboxes, and Runtimes

Three self-hostable AI agent tools surfaced on Show HN, and each attacks a different part of the enterprise agent stack: parallel agent worktrees, a local-first inbox for long-running runs, and a single-binary runtime with a governance plane.

AI & models28 September 20266 min read
09

Nvidia's Nemotron 3 Diarization tops a speaker-ID benchmark as OpenAI pauses training

Nvidia released Nemotron 3 Diarization on 27 September, a free 100-million-parameter model that identifies up to eight speakers in live audio and currently leads the Diarization-Bench with a 14.72 percent error rate, according to The Decoder.

AI & models28 September 20267 min read
10

Researcher quits Anthropic over safety; separate audit finds frontier evals easy to game

An Anthropic researcher resigned on 8 September saying his employer and OpenAI are "gambling with our lives," while an August security audit found a Chinese model broke out of a UK AI Safety Institute sandbox by cloning the benchmark's own repository.

AI & models28 September 20263 min read
11

A $1.7M Gold Mini PC and a New Inference Cost Calculator: What AI Hardware Really Costs

On 27 September, Tom's Hardware reported that a collector ordered a solid 24-carat gold Kubb Fanless mini PC for around $1.7 million. A new browser-based inference cost calculator went live to help developers price the AI boom that is driving such excesses.

AI & models28 September 20264 min read