Skip to content
World clockEU--:--UK--:--USA--:--CN--:--PLDEFRIT中文EN

portal about AI and technologyevents · analysis · interviews · technical background

Search
LIVE
›

AI Is Eating Science: What the Newest Results Actually Show

On 29 September the National Bureau of Economic Research published a working paper describing an open-source workflow in which a large language model reproduces, improves and extends published economics research. The workflow flagged discrepancies in 3,460 of 4,452 replication packages.

ScienceExplainerSofia MarchettiPublished: 29 September 20267 min readSources 9
AI Is Eating Science: What the Newest Results Actually Show

The paper, NBER Working Paper 35782, is dated September 2026 and went up on the bureau's site on 29 September. Its authors, Matthew Schwartz, Isaiah Andrews and Jesse M. Shapiro, describe a three-step workflow. It first tries to reproduce an article's original calculations, then checks for discrepancies with the published findings, then runs automated sensitivity analysis.

Across 4,452 published replication packages from five economics journals, the workflow flagged discrepancies in 3,460 articles or their appendices. That is not a claim that most economics is wrong. It is a claim about how often an automated re-run fails to match what was printed, whether because of a coding slip, a version mismatch or an undocumented choice. The same paper reports two other numbers. In 496 articles the workflow cut a calculation's computation time by more than a factor of 10 at similar or greater accuracy. In 923 articles it produced an extension that does not appear in the original article and that is aligned with the original article's goals and assumptions.

An honest disclosure in the paper itself

There is a disclosure worth reading before the results are quoted anywhere. The NBER paper states that LLMs were used in the analysis and writing of the article. Schwartz worked as a contractor for Anthropic, and the paper says the results and views are not endorsed by Anthropic. Shapiro has in the past been a paid visitor at Microsoft Research New England and a paid consultant for FutureOfCapitalism, and has been paid for writing by the New York Times. That is more transparency than most announcements in this area carry, and it cuts both ways: readers can weigh the conflicts, and the conflicts are real.

Mathematicians are arguing about the same thing

Two posts published on 29 September take up the question from the pure-math side. On his blog, Dan Romik writes that in the wake of OpenAI's announcement of a solution to the Navier-Stokes problem, mathematicians are anxious about what role will be left for them. He quotes Scott Aaronson's reaction to the news as a run of the letter A, and then argues that the anxious reading is wrong.

Mathematics research is a feedback loop that generates new knowledge in a constant process of re-examining, digesting and distilling existing knowledge and attempting to improve on it, Romik writes.

His technical claim is narrower than the headline version. AI models trained on the existing corpus can generate theorems that are, in his phrase, distance one away from human-generated knowledge, and they will do that faster than people can. What they cannot do, he argues, is reflect on that output, simplify it, extract why a proof works and generalize the idea so it can be reused. Left to itself, he predicts, an autonomous system stalls after the first wave and drowns in its own output. Stephen Wolfram makes a related argument in a long post dated 28 September and published on 29 September. He writes that modern AI is first and foremost a way of drawing on the existing corpus of human knowledge, and that at the core of pure mathematics is human imagination guiding which questions to ask. Wolfram also draws a line between AI and computation. AI, in his account, mines what humans have already produced. Computation, starting from a rule and running it, generates results that are irreducibly new, a property he calls computational irreducibility.

The industrial layer is moving in parallel

The research results are arriving alongside commercial announcements that treat autonomous research as a product feature. CoreWeave published a post on 29 September introducing ARIA, a coding agent built into Weights & Biases. It reads a team's experiments, forms a hypothesis, writes a config and launches the run through W&B Launch, then evaluates the result against a baseline and drafts a report.

The company says ARIA builds shareable W&B workspaces, panels and reports rather than returning text, and that it is available in the Weights & Biases mobile app. The post was originally published on the W&B blog on 29 July, according to a line at the top of the page, and CoreWeave's own site carries it under a September date. Oracle made a similar move on 29 September, announcing Fusion Claw, a governed agentic execution runtime for its Fusion Applications. The company says 25 Claw-powered agentic applications are available, part of a portfolio of 75, and that the runtime combines a frontier model with deterministic enterprise computation, with an Outcome Receipt as an auditable record of what ran. Oracle's announcement quotes Mike Sicilia, the company's CEO. The systems run on Oracle Cloud Infrastructure and are powered by frontier models including Gemini and OpenAI, with support for others planned, according to the press release.

Hardware is chasing the same problem from the energy side. Efficient Computer said on 29 September that it raised a $97 million Series B led by TQ Ventures, taking its total raise to $173 million at a $650 million valuation. CEO Brandon Lucia writes that the company's processors bring 10 to 100 times better energy efficiency than traditional CPU architectures, and that it is launching the Electron E1 at volume for embedded physical AI systems. That claim comes from the company's own blog post and has not been independently verified here.

Energy, water and the places asked to host it

The compute buildout has a geography, and the geography has politics. CBC News reported on 29 September that Newfoundland and Labrador's energy minister, Lloyd Parrott, told the House of Assembly during a special session on the Churchill Falls agreement in mid-September that the province's door is open for business. A department spokesperson, Brodie Thomas, confirmed to CBC that the government has been approached by companies about data centres, and said commercially sensitive discussions could not be commented on. CBC quotes Paris Marx, author of a book on data centres, asking whether it makes sense to give that electricity to data centres when there might be better uses. Andrea King, CEO of TechNL, told the broadcaster that Labrador meets the criteria proponents look for, low-carbon hydro, cold weather and land, but that this does not by itself make a project a good idea.

The debate is not only Canadian. CBC notes that Meta's planned data centre in Sturgeon County, Alberta, is expected to cost $13 billion and come online in two to three years, and that Nova Scotia Premier Tim Houston has set five starting conditions for any proponent. The energy arithmetic behind all of this is old and unfashionable. David MacKay's Sustainable Energy Without the Hot Air, whose site carries a last-modified date of 29 August 2015, remains the standard back-of-the-envelope treatment, and the page was still being surfaced on 29 September.

A newer treatment comes from Hannah Ritchie, who published an analysis of electrification on her Substack on 29 September. Working through numbers from Oxford professor Nick Eyre, she reports that global final energy demand falls from 416 to 247 exajoules in a post-transition system, about 40% lower, while electricity demand rises from 110 to 189 EJ. Electric cars convert around 80% of energy to motion against 20% for petrol, she writes. The model assumes all non-electrified sectors run on hydrogen, and she flags that as a simplification.

What the week actually established

Nothing here shows that AI has replaced researchers. The NBER paper is a workflow paper about replication, not a claim about discovery. Romik's argument is a prediction, not a result, and he says plainly that he does not think we are in that world yet. Wolfram's is an essay. What the week did establish is that the tools for automated replication and extension of published work now exist in the open, that their error rate on a large corpus is high enough to matter, and that the people building the commercial infrastructure are selling autonomy on a much shorter timeline than the people thinking about what autonomy is for.

Comments 0

Sources

9
  1. 01An LLM Workflow That Reproduces, Improves, and Extends Published Economics ResearchEN
  2. 02The feedback loop of mathematics researchEN
  3. 03What's the Future for Pure Math Research in the Age of AI?EN
  4. 04Introducing CoreWeave ARIA: AI Research and Iteration AgentEN
  5. 05Oracle Extends Fusion Agentic Applications with Introduction of Fusion ClawEN
  6. 06Solving computing's energy problem with Efficient Computer's $97M Series BEN
  7. 07AI data centres in N.L.? The door is 'open for business,' says energy ministerEN
  8. 08Electrification efficiency: The world will need less energy after the transitionEN
  9. 09Sustainable energy without the hot airEN

All figures and quotations in this text come from the sources listed below.

Content prepared by the editorial team with AI assistance.

Sofia Marchetti

Sofia Marchetti

Science and health

Sofia Marchetti covers science and health for FLASH24, working from primary literature, preprints, and agency data rather than press releases. She checks sample sizes, confidence intervals, and whether a study's numbers match its abstract before filing. She interviews researchers and clinicians directly, tracks conference calendars for embargoed results, and compares new findings with earlier trials on the same question. Outside the newsroom she works on materials physics and stargazes through a home telescope, which keeps her close to how measurement error actually behaves. She does not publish a health claim without a named source and the underlying data.

Newsroom →

Comments

0
  1. No comments yet — be the first.

Write a comment

Comments are public. We do not publish abuse, spam or advertising.