Medical AI regulation: what the research actually shows about diagnosis tools
Medical AI models used for diagnosis can be tricked into revealing whether a specific patient's data trained them, German researchers reported in Nature on 24 June 2026. A separate Microsoft system claimed 85.5% diagnostic accuracy against doctors. Regulators are still catching up.

Two research results, published roughly a year apart, sit awkwardly next to each other. One says medical AI can beat doctors at hard diagnoses. The other says the same class of models leaks information about the patients it learned from. Both feed the same unresolved question: what should regulation of medical AI actually require?
The first result came from Microsoft. According to GeekWire, the company announced the MAI Diagnostic Orchestrator, or MAI-DxO, on 30 June 2025. Microsoft tested it against 21 experienced physicians from the US and UK on complex cases drawn from the New England Journal of Medicine. The tool produced a correct diagnosis in 85.5% of cases. The doctors hit 20%.
What the Microsoft benchmark measured, and what it did not
The test set was 304 recent NEJM case studies, turned into something Microsoft called the Sequential Diagnosis Benchmark. WIRED reported that a language model broke each case into the step-by-step process a doctor would follow. MAI-DxO then queried several frontier models, including OpenAI's GPT, Google's Gemini, Anthropic's Claude, Meta's Llama and xAI's Grok, in a setup meant to resemble several experts debating. MAI-DxO paired with OpenAI's o3 performed best, GeekWire said.
WIRED put the accuracy figure at 80%, not 85.5%, and said the system cut costs by 20% by ordering cheaper tests. The two numbers come from the same project but different write-ups. GeekWire cites 85.5% and 20% for doctors, WIRED cites 80% and 20%. Anyone citing this benchmark should check which figure they are using.
"We're taking a big step towards medical superintelligence," Mustafa Suleyman, CEO of Microsoft AI, said in a LinkedIn post, according to GeekWire.
Outside researchers were more measured. MIT scientist David Sontag told WIRED the paper was strong on methodology, but warned that the doctors in the study were told not to use any additional tools, which does not reflect real practice. Eric Topol of the Scripps Research Institute called it an impressive report on complex cases. Both said a clinical trial comparing the system with real doctors treating real patients would be the next validation step. Microsoft has not decided whether to commercialise the tool, WIRED reported, and clinical deployment would require safety testing and regulatory approval.
The privacy finding that complicates everything
On 24 June 2026, The Register reported on a Nature paper from German researchers. The team tested membership inference attacks, or MIAs, against seven medical AI datasets containing images, ECG records and general electronic health records. These attacks query a model to work out whether a particular data point was in its training set.
The finding: for individual patients, the attacks succeeded with what the researchers called near-perfect success. The paper argues that standard evaluation protocols miss this because they measure attack success in aggregate across records. Underrepresented groups were easier to identify. Race, insurance status, sex, imaging protocol and certain disease statuses all act as outliers that make individuals stand out.
Moritz Knolle, who leads the AI in Healthcare and Medicine chair at the Technical University of Munich and is lead author on the paper, told The Register that privacy risks grow as a training cohort becomes more specific. He gave the example of membership in a dataset revealing a dormant genetic condition such as Huntington's disease, depression, or attendance at a specialised treatment clinic.
There is a practical limit. To run the attack, an attacker needs at least partial data belonging to the person they want to identify. Knolle told The Register that the paper shows partial access is enough, contrary to earlier assumptions. He suggested a scenario where an attacker gains unauthorised access to a GP's database after a routine blood test.
"I hope that the medical AI community will start to take privacy risks seriously and that risk mitigation techniques are used in situations where they are necessary," Knolle said.
The researchers recommend differential privacy frameworks, which aim to give mathematical guarantees that training data stays anonymous, and changes to privacy audit standards so they assess individual-level risk rather than aggregate risk. They also suggest compiling training data so underrepresented groups are better represented. Knolle noted that many successful attacks represent a small or negligible privacy violation, for instance where models are trained on large general populations containing both healthy and diseased people.
The evidence that models fail early, where it matters
A third study cuts against the enthusiasm. In April 2026, The Register reported on research published in JAMA Network Open and led by Harvard medical student Arya Rao. The team tested 21 off-the-shelf AI models on 29 standardised clinical vignettes. When given a full portfolio of medical information and asked for a final diagnosis, leading models were correct 91% of the time. Early differential diagnosis, the stage where clinicians rule conditions in and out while uncertainty is highest, failed in more than 8 out of 10 cases.
Rao told The Register that every model tested failed on the vast majority of cases, and that early differential diagnosis is where the systems are weakest. Co-author Dr Marc Succi, a radiologist at Massachusetts General Hospital, said the results suggest off-the-shelf LLMs should not be trusted for patient-facing diagnostic reasoning without structured comprehensive human review. He added that the models can project confidence without showing reliable reasoning, which can inflame anxiety in patients.
Rao also offered a caveat: failure under the paper's strict definition did not always mean the model was wrong. Measured as raw accuracy by proportion of correct answers per case, scores ranged from 63% to 78%, meaning models were often partially correct. The team still argues the stricter metric matters, because LLMs are being marketed as frontline tools that narrow down diagnoses before a human takes over.
Where regulation stands
None of these three papers is a regulatory document, and the dossier behind this article does not contain the text of any AI-in-medicine rule. What it does show is the shape of the problem regulators face. Diagnostic accuracy on curated cases can look strong while early-stage reasoning fails, and privacy audits can pass while individual patients remain identifiable.
The Microsoft work points one way: prove the system in a clinical trial, then seek approval. The Nature paper points another: change how privacy audits are written, and adopt techniques such as differential privacy before deployment. The JAMA study points at labelling and marketing, arguing that describing LLMs as diagnostic agents creates false confidence exactly where they are least reliable.
For anyone reading about a medical AI product, three questions follow from the evidence above. Was it tested on real patients in a clinical setting, or on published case studies? Does its privacy evaluation look at individual patients or only at aggregate risk? And does it claim to diagnose, or to assist a clinician who remains responsible? The answers are usually in the paper, not the press release.
Sources
4- 01Medical diagnosis AIs can be tricked into telling whose data trained themEN
- 02Microsoft AI tool outperforms doctors in diagnosing complex medical casesEN
- 03Microsoft Says Its New AI System Diagnosed Patients 4 Times More Accurately Than Human DoctorsEN
- 04LLMs fail in 8 out of 10 early differential diagnosis casesEN
All figures and quotations in this text come from the sources listed below.
Content prepared by the editorial team with AI assistance.
Comments
0- No comments yet — be the first.